Unmanned vehicle confrontation control method and device based on reinforcement learning

By building a self-driving vehicle simulation system and multi-neural network model, designing reward functions and introducing regular networks, optimizing the unmanned vehicle confrontation control strategy, solving the problems of high training cost, low robustness and weak confrontation in complex environments, and realizing intelligent control of unmanned vehicles in complex scenarios.

CN120297355APending Publication Date: 2025-07-11SUN YAT SEN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510282809.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing unmanned vehicle confrontation control methods have high training costs, low robustness and weak confrontation in complex and dynamic environments, making it difficult to cope with the complexity of interactions of multiple unmanned vehicles.

Method used

The unmanned vehicle confrontation control method based on reinforcement learning is designed by building an unmanned vehicle simulation system, introducing agents and multi-neural network models, designing reward functions, and iterative training is carried out to optimize the strategy and introduce regular networks for constraints, improving the robustness and adaptability of the strategy.

Benefits of technology

It realizes that unmanned vehicles quickly learn effective strategies in complex scenarios, improves confrontation ability and adaptability, and can maximize the use of environmental status information for intelligent control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297355A_ABST
    Figure CN120297355A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned vehicle confrontation control method and device based on reinforcement learning, and the method comprises the steps: constructing an unmanned vehicle simulation system, and creating a corresponding agent for each unmanned vehicle in the unmanned vehicle simulation system; based on initialization of a plurality of neural networks, constructing a first unmanned vehicle confrontation control strategy model of each agent; constructing a reward function corresponding to the first unmanned vehicle confrontation control strategy model on the basis of reward design of a plurality of unmanned vehicle training tasks corresponding to each agent; creating an unmanned vehicle game based on the unmanned vehicle simulation system, and collecting agent experience data of the unmanned vehicle game; and performing iterative training on the first unmanned vehicle confrontation control strategy model based on the agent experience data to obtain a second unmanned vehicle confrontation control strategy model, and determining the second unmanned vehicle confrontation control strategy model as a corresponding unmanned vehicle confrontation decision model. According to the invention, efficient autonomous behaviors of the unmanned vehicle in a complex environment can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous vehicles, and in particular, to an autonomous vehicle adversarial control method and device based on reinforcement learning. Background Art

[0002] Autonomous vehicle confrontation usually occurs in extreme driving simulation environments. Multiple autonomous vehicles achieve their respective goals through mutual competition, attack, and avoidance strategies, demonstrating the potential of autonomous vehicle technology under extreme conditions. The key to autonomous vehicle confrontation lies in how to make intelligent decisions in a dynamic, complex, and information-incomplete environment. Autonomous vehicles must analyze the changes in the surrounding environment, the behaviors of other vehicles, and potential threats in real time and adopt appropriate confrontation strategies, which requires highly intelligent algorithms.

[0003] Traditional autonomous vehicle adversarial control methods usually adopt fixed rules or path search algorithms, such as the A* algorithm and the Dijkstra algorithm. These methods are suitable for known or relatively stable environments. Due to the lack of adaptability, they are difficult to cope with the constantly changing attack strategies of opponent vehicles, resulting in poor performance of autonomous vehicles in complex scenarios of multi-vehicle adversarial attacks.

[0004] Reinforcement learning methods represented by the Deep Q-Network (DQN) and the Proximal Policy Optimization (PPO) algorithm have shown remarkable effects in discrete and continuous space environments, but there are some bottlenecks in autonomous vehicle control tasks: 1) Low sample efficiency: The control tasks of autonomous vehicles usually require a large amount of actual or simulated data for training. Traditional reinforcement learning converges slowly in complex environments, resulting in high training costs; 2) Low robustness: The autonomous vehicle control environment is full of uncertainties. Traditional reinforcement learning depends on fixed states and rewards and is difficult to quickly adapt and update strategies in a constantly changing environment; 3) Weak adversarial ability: In adversarial scenarios, current autonomous vehicles need to predict the behaviors of other autonomous vehicles. However, the strategies generated by traditional reinforcement learning are relatively fixed, difficult to handle the complexity of multi-autonomous vehicle interactions, and prone to failure in adversarial scenarios. Summary of the Invention

[0005] In order to solve the above technical problems existing in the prior art in autonomous vehicle control tasks, the present invention provides an autonomous vehicle adversarial control method and device based on reinforcement learning.

[0006] In a first aspect, an embodiment of the present invention provides an autonomous vehicle adversarial control method based on reinforcement learning, including:

[0007] Construct an unmanned vehicle simulation system based on the unmanned vehicle motion control parameters and the unmanned vehicle confrontation environment variables, and create a corresponding agent for each unmanned vehicle in the unmanned vehicle simulation system. Among them, the unmanned vehicle motion control parameters include the unmanned vehicle steering control parameters and the unmanned vehicle speed control parameters, and the unmanned vehicle confrontation environment variables include the ground friction coefficient and obstacles;

[0008] Based on the initialization of a number of neural networks, construct the first unmanned vehicle confrontation control strategy model for each of the agents, where the number of neural networks includes a policy network, a regularization network, an advantage network, and a Q-value network;

[0009] Based on the reward design for a number of unmanned vehicle training tasks corresponding to each agent, construct a reward function corresponding to the first unmanned vehicle confrontation control strategy model, where the number of unmanned vehicle training tasks includes an unmanned vehicle navigation training task and an unmanned vehicle confrontation training task;

[0010] Create an unmanned vehicle game confrontation based on the unmanned vehicle simulation system, and collect the agent experience data of the unmanned vehicle game confrontation. Among them, the agent experience data includes the agent's current moment observation data, the agent's executed action, the agent's corrected reward value, the agent's next moment observation data, and the game end judgment condition. The agent's corrected reward value is obtained by calculating the reward function corresponding to the agent;

[0011] Iteratively train the first unmanned vehicle confrontation control strategy model based on the agent experience data to obtain a second unmanned vehicle confrontation control strategy model, and determine the second unmanned vehicle confrontation control strategy model as the unmanned vehicle confrontation decision model of the unmanned vehicle entity system corresponding to the unmanned vehicle simulation system. Among them, the iterative training includes updating the regularization network at a fixed interval according to the iteration times based on the policy network.

[0012] Preferably, the constructing the first unmanned vehicle confrontation control strategy model for each of the agents based on the initialization of a number of neural networks includes:

[0013] Determine the input dimension and output dimension of each neural network. Among them, the input dimension of each neural network is D-dimensional, the output dimensions of the policy network, the regularization network, and the advantage network are all K-dimensional, and the output dimension of the Q-value network is one-dimensional. Both D and K are integers greater than or equal to 2;

[0014] Initialize the policy network so that the policy network is configured to generate the probability distribution of each agent selecting each executed action in the corresponding unmanned vehicle current state;

[0015] Initialize the regularization network based on the policy network, so that the regularization network is configured to perform regularization constraints on the policy network, the advantage network, and the Q-value network respectively;

[0016] Initialize the advantage network, so that the advantage network is configured to evaluate the advantage degree of each agent in selecting each execution action in the current state of the corresponding autonomous vehicle;

[0017] Initialize the Q-value network, so that the Q-value network is configured to evaluate the Q-value of each execution action of each agent in the current state of the corresponding autonomous vehicle;

[0018] Integrate the policy network, the regularization network, the advantage network, and the Q-value network to obtain the first autonomous vehicle confrontation control policy model for each agent.

[0019] Preferably, the reward function corresponding to the first autonomous vehicle confrontation control policy model is constructed based on the reward design for each agent corresponding to several autonomous vehicle training tasks, including:

[0020] Based on the reward design for each agent corresponding to the autonomous vehicle navigation training task, obtain the first reward function corresponding to the first autonomous vehicle confrontation control policy model;

[0021] Based on the reward design for each agent corresponding to the autonomous vehicle confrontation training task, obtain the second reward function corresponding to the first autonomous vehicle confrontation control policy model;

[0022] Based on the combination of the first reward function and the second reward function, construct the reward function corresponding to the first autonomous vehicle confrontation control policy model.

[0023] Preferably, the first reward function is represented by the following formula:

[0024]

[0025] where, represents the first reward function at time t, r collision represents the autonomous vehicle collision penalty, r distance represents the autonomous vehicle distance reward, r velocitysmooth represents the autonomous vehicle speed smoothing penalty;

[0026] The autonomous vehicle collision penalty is calculated by the following formula:

[0027]

[0028] The autonomous vehicle distance reward is calculated by the following formula:

[0029]

[0030] Among them, α d represents the weight coefficient of the distance reward of the unmanned vehicle, representing the distance between the unmanned vehicle and the target unmanned vehicle;

[0031] The following formula is used to calculate the speed smoothing penalty of the unmanned vehicle:

[0032] r velocitysmooth =-α v ∥v t -v t-1 ∥

[0033] Among them, α v represents the coefficient of the speed smoothing penalty of the unmanned vehicle, v t represents the speed of the unmanned vehicle at time t, v t-1 represents the speed of the unmanned vehicle at time t-1.

[0034] Preferably, the following formula is used to characterize the second reward function:

[0035]

[0036] Among them, represents the second reward function at time t, r aim represents the attack orientation reward of the unmanned vehicle, r hit represents the attack hit reward of the unmanned vehicle, r miss represents the attack miss penalty of the unmanned vehicle, r risk represents the risk reward of the unmanned vehicle, and the risk reward of the unmanned vehicle is configured to quantify the risk degree of the unmanned vehicle when facing the target unmanned vehicle.

[0037] Preferably, the risk reward of the unmanned vehicle is obtained by prediction through supervised learning based on the state of the unmanned vehicle and the state of the target unmanned vehicle.

[0038] Preferably, the second unmanned vehicle confrontation control strategy model is obtained by iteratively training the first unmanned vehicle confrontation control strategy model based on the intelligent agent experience data, including:

[0039] Randomly extracting the intelligent agent experience data to obtain a training set;

[0040] Based on the training set, the advantage network in the first unmanned vehicle confrontation control strategy model is iteratively trained by using a first update strategy to obtain the trained advantage network, wherein the first update strategy includes updating the advantage network based on the mean square error loss function;

[0041] Iteratively train the Q-value network in the first unmanned vehicle adversarial control policy model using a second update strategy based on the training set to obtain the trained Q-value network, where the second update strategy includes updating the Q-value network based on the corrected reward value of the agent;

[0042] Iteratively train the policy network in the first unmanned vehicle adversarial control policy model using a third update strategy based on the training set to obtain the trained policy network, where the third update strategy includes updating the policy network based on the difference between the advantage network and the Q-value network;

[0043] Based on the training set, train the regularization network in the first unmanned vehicle adversarial control policy model at fixed intervals according to the iteration times based on the policy network to obtain the trained regularization network;

[0044] Integrate the trained advantage network, the trained Q-value network, the trained policy network, and the trained regularization network to obtain the second unmanned vehicle adversarial control policy model for each agent.

[0045] Preferably, the first update strategy is represented by the following formula:

[0046]

[0047] where A target represents the target value of the advantage network, l represents the hyperparameter of the clipping operator, A represents the advantage network, I represents the agent's observation data, a represents the agent's executed action, θ t represents the parameters of the advantage network at time step t, Q represents the Q-value network, ω t-1 represents the parameters of the Q-value network at time step t-1, G t represents the corrected reward value of the agent, τ represents the hyperparameter of the influence degree of the regularization network, π t (a|I) represents the probability of the policy network outputting the agent's executed action at time step t, μ(a|I) represents the probability of the regularization network outputting the agent's executed action, α actor represents the hyperparameter of the advantage network, represents the gradient at time step t, MSE loss represents the mean square error loss function;

[0048] The second update strategy is represented by the following formula:

[0049]

[0050] where α critic represents the learning rate parameter of the Q-value network;

[0051] The third update strategy is characterized by the following formula:

[0052] π t+1 (a|I) ∝ exp(A(I, a; θ t ) - Q(I; ω t ))

[0053] where μ represents the regular network, π t represents the policy network corresponding to time step t, iter represents the number of iterations, and N reg represents the preset hyperparameters of the regular network.

[0054] In a second aspect, an embodiment of the present invention provides an unmanned vehicle adversarial control device based on reinforcement learning, including:

[0055] A simulation system construction module, configured to construct an unmanned vehicle simulation system based on unmanned vehicle motion control parameters and unmanned vehicle adversarial environment variables, and create corresponding agents for each unmanned vehicle in the unmanned vehicle simulation system, where the unmanned vehicle motion control parameters include unmanned vehicle steering control parameters and unmanned vehicle speed control parameters, and the unmanned vehicle adversarial environment variables include ground friction coefficient and obstacles;

[0056] A first model construction module, configured to construct a first unmanned vehicle adversarial control policy model for each of the agents based on the initialization of a number of neural networks, where the number of neural networks includes a policy network, a regular network, an advantage network, and a Q-value network;

[0057] A reward function construction module, configured to construct a reward function corresponding to the first unmanned vehicle adversarial control policy model based on the reward design for a number of unmanned vehicle training tasks corresponding to each of the agents, where the number of unmanned vehicle training tasks includes an unmanned vehicle navigation training task and an unmanned vehicle adversarial training task;

[0058] An experience data acquisition module, configured to create an unmanned vehicle game session based on the unmanned vehicle simulation system and collect the agent experience data of the unmanned vehicle game session, where the agent experience data includes agent current moment observation data, agent executed actions, agent corrected reward values, agent next moment observation data, and game end judgment conditions, and the agent corrected reward value is obtained by calculating the reward function corresponding to the agent;

[0059] The adversarial decision-making model determination module is used to iteratively train the first unmanned vehicle adversarial control strategy model based on the agent experience data to obtain a second unmanned vehicle adversarial control strategy model, and determine the second unmanned vehicle adversarial control strategy model as the unmanned vehicle adversarial decision-making model of the unmanned vehicle entity system corresponding to the unmanned vehicle simulation system. Among them, the iterative training includes updating the regularization network at fixed intervals according to the number of iterations based on the policy network.

[0060] Compared with the prior art, the beneficial effects of an unmanned vehicle adversarial control method and device based on reinforcement learning according to an embodiment of the present invention are as follows: By designing a reward function applicable to the unmanned vehicle navigation and adversarial scenarios through a reinforcement learning framework, and introducing a regularization network independent of the environment to constrain and optimize the policy, it is avoided that the policy falls into a local optimum, enabling the unmanned vehicle to quickly learn effective policies in different scenarios, being able to maximize the use of state information in the environment, and realizing intelligent control of the unmanned vehicle in complex scenarios. Description of the Drawings

[0061] Figure 1 is a schematic flowchart of an unmanned vehicle adversarial control method based on reinforcement learning according to an embodiment of the present invention;

[0062] Figure 2 is a schematic flowchart of constructing the first unmanned vehicle adversarial control strategy model according to an embodiment of the present invention;

[0063] Figure 3 is a schematic flowchart of constructing the reward function according to an embodiment of the present invention;

[0064] Figure 4 is a schematic flowchart of obtaining the second unmanned vehicle adversarial control strategy model according to an embodiment of the present invention

[0065] Figure 5 is a schematic structural diagram of an unmanned vehicle adversarial control device based on reinforcement learning according to an embodiment of the present invention. Detailed Embodiments

[0066] The following further describes in detail the specific embodiments of the present invention with reference to the drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0067] In the description of the present invention, it should be understood that the terms "first" and "second" etc. used in the present invention are used to distinguish different objects, rather than to describe a specific order.

[0068] In the description of the present invention, it should be noted that unless otherwise defined, all the technical and scientific terms used in the present invention have the same meanings as those commonly understood by those skilled in the art. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0069] As Figure 1 shown, an embodiment of the present invention provides an adversarial control method for an autonomous vehicle based on reinforcement learning, including the steps of:

[0070] S1. Construct an autonomous vehicle simulation system based on the motion control parameters of the autonomous vehicle and the adversarial environment variables of the autonomous vehicle, and create corresponding agents for each autonomous vehicle in the autonomous vehicle simulation system;

[0071] To construct the autonomous vehicle simulation system, the control parameters of the actual autonomous vehicle need to be considered. For example, the steering control system determines the angle and speed at which the vehicle can turn, and the speed control system determines the acceleration and deceleration performance of the vehicle. That is to say, the motion control parameters of the autonomous vehicle include the steering control parameters and the speed control parameters of the autonomous vehicle. At the same time, the application environment variables are also crucial. For example, the friction coefficient in the environment affects the braking distance and driving stability of the vehicle, and the presence of obstacles determines that the vehicle needs to perform obstacle avoidance operations, that is, the adversarial environment variables of the autonomous vehicle include the ground friction coefficient and obstacles. By incorporating these actual parameters and variables into the simulation system, the simulation can be made closer to the real scenario.

[0072] In the autonomous vehicle simulation system, the agent is a key element, which is responsible for guiding all the actions of the autonomous vehicle. Creating corresponding agents for each autonomous vehicle means that each agent will be specifically responsible for controlling the behavior decision of an autonomous vehicle.

[0073] During the driving process, the autonomous vehicle will obtain a large amount of information through various sensors. The on-vehicle image information allows the vehicle to observe the surrounding environment and identify other vehicles and obstacles, etc.; the on-vehicle radar information can accurately measure the distance between the vehicle and the surrounding objects and provide more accurate spatial position data; the current state of the vehicle includes its own operating parameters such as vehicle speed, steering angle, and acceleration. These information together constitute the state of the autonomous vehicle and are the basis for the agent to make decisions.

[0074] Since the forms and dimensions of these state information vary, in order to facilitate the processing and understanding by the agent, it is necessary to represent them as fixed D-dimensional vectors through various embedding techniques. At the same time, in order to obtain the specific content of the state information when needed, after representing the state information as a fixed-dimensional vector, the vector can still be sliced through slicing techniques to restore the corresponding specific information. Meanwhile, in the unmanned vehicle simulation system, the policy stipulates the actions that the agent can take in each state, such as the steering, acceleration / deceleration, and shooting actions of the unmanned vehicle, and limits the number of behaviors to K.

[0075] S2. Based on the initialization of several neural networks, construct the first unmanned vehicle adversarial control policy model for each agent;

[0076] Specifically, as Figure 2 shown, step S2 includes:

[0077] S201. Determine the input dimension and output dimension of each neural network;

[0078] The several neural networks include a policy network, a regularization network, an advantage network, and a Q-value network. Among them, the input dimension of each neural network is D-dimensional, the output dimensions of the policy network, the regularization network, and the advantage network are all K-dimensional, and the output dimension of the Q-value network is one-dimensional. Both D and K are integers greater than or equal to 2.

[0079] S202. Initialize the policy network so that the policy network is configured to generate the probability distribution of each agent selecting each execution action in the current state of the corresponding unmanned vehicle;

[0080] The policy network is responsible for generating the probability distribution of the agent selecting various actions in a given state. It guides the agent to make optimal decisions by learning and optimizing the policy, aiming to maximize the expected long-term reward.

[0081] S203. Initialize the regularization network based on the policy network so that the regularization network is configured to perform regularization constraints on the policy network, the advantage network, and the Q-value network respectively;

[0082] The regularization network is used to provide a regularization signal. By performing regularization constraints on the policy network, the advantage network, and the Q-value network respectively, it balances the stability and flexibility of the first unmanned vehicle adversarial control policy model. It helps reduce the noise in policy updates during the training process, especially in the face of incomplete information or adversarial environments, and prompts the agent to learn more robust policies.

[0083] S204. Initialize the advantage network so that the advantage network is configured to evaluate the advantage degree of each agent selecting each execution action in the current state of the corresponding unmanned vehicle;

[0084] The Advantage Network is used to calculate the advantage function and evaluate the superiority of a certain action in a given state. It helps the agent identify which actions are more valuable compared to other possible actions, thereby optimizing the output of the Policy Network.

[0085] S205. Initialize the Q-value Network so that the Q-value Network is configured to evaluate the Q-value of each action executed by each agent in the current state of the corresponding driverless vehicle.

[0086] The Q-value Network is used to estimate the value of the state-action pair, that is, the expected return obtained by taking a certain action in a given state, namely the Q-value. It is continuously updated through training to provide an evaluation of the long-term returns of different action choices.

[0087] S206. Integrate the Policy Network, the Regularization Network, the Advantage Network, and the Q-value Network to obtain the first driverless vehicle adversarial control policy model for each agent.

[0088] It can be understood that the first driverless vehicle adversarial control policy model is designed based on the reinforcement learning framework. In particular, by introducing a regularization network independent of the environment, it can guide the policy to effectively complete convergence.

[0089] The Experience Replay Buffer is a commonly used technique in reinforcement learning for storing data on the interaction between the agent and the environment. In the driverless vehicle simulation system, at each time step, the agent selects an action based on the current state. After executing the action, a new state and the corresponding reward are obtained. This information (including the current state, action, reward, next state, etc.) is stored in the Experience Replay Buffer. Initialize the Experience Replay Buffer to prepare for subsequent data storage.

[0090] Set the training times calculator iter = 0. The training times counter is used to record the training times of the neural network. In reinforcement learning, it is usually necessary to perform multiple iterative trainings to optimize the network parameters so that the agent can learn a better policy. The training times counter can help control the training progress and termination conditions. For example, when the training times reach the preset maximum value, stop the training.

[0091] Set the data reuse times k = 0. The data reuse times specify the number of times the data in the Experience Replay Buffer can be reused. During each training, data can be randomly sampled from the Experience Replay Buffer for training. By reusing these data multiple times, the data utilization rate can be improved, the dependence on environmental interaction can be reduced, and at the same time, it is also helpful to improve the training stability.

[0092] The experience replay pool and the training data are both stored in containers. Storing data in containers allows for convenient random extraction of one or more pieces of data for training, which can break the correlation between data and improve the efficiency and stability of training. Additionally, when the experience replay pool is saturated, data is updated in a queue form, i.e., the earliest inserted data is discarded and the latest inserted data is saved. This update mechanism ensures that the experience replay pool always stores the latest interaction data, enabling the agent to learn and make decisions based on the latest environmental information.

[0093] S3. Based on the reward design for each agent corresponding to several unmanned vehicle training tasks, construct a reward function for the corresponding first unmanned vehicle adversarial control strategy model.

[0094] Existing unmanned vehicle adversarial training methods usually have difficulty effectively dealing with complex and dynamic adversarial environments because they often adopt a training method with a single task or a single time granularity and cannot fully handle multi-task collaboration and high-complexity adversarial scenarios. To solve this problem, the present invention proposes a multi-stage strategy reward design. Several unmanned vehicle training tasks include unmanned vehicle navigation training tasks and unmanned vehicle adversarial training tasks. By starting from simple navigation tasks, gradually transitioning to more complex adversarial tasks, and then combining the two, the agent can optimize its behavior in the enemy-opponent adversarial environment while gradually increasing the task complexity. This design not only improves the learning efficiency of the tasks but also avoids the agent falling into a local optimal solution when facing complex situations, thereby enhancing the adversarial ability and adaptability of the unmanned vehicle.

[0095] Specifically, as Figure 3 shown, step S3 includes:

[0096] S301. Based on the reward design for each agent corresponding to the unmanned vehicle navigation training task, obtain the first reward function for the corresponding first unmanned vehicle adversarial control strategy model.

[0097] In the navigation task, the design of the reward function takes into account multiple factors to enhance the stability of the control strategy and the navigation effect.

[0098] Specifically, the first reward function is represented by the following formula:

[0099]

[0100] where represents the first reward function at time t, r collision represents the unmanned vehicle collision penalty, r distance represents the unmanned vehicle distance reward, r velocitysmooth represents the unmanned vehicle speed smoothness penalty.

[0101] If a collision occurs in the autonomous vehicle at a certain moment, increase the collision penalty for the autonomous vehicle. Further, the following formula is used to calculate the collision penalty for the autonomous vehicle:

[0102]

[0103] Further, the following formula is used to calculate the distance reward for the autonomous vehicle:

[0104]

[0105] where α d represents the weight coefficient of the distance reward for the autonomous vehicle, represents the distance between the autonomous vehicle and the target autonomous vehicle.

[0106] By controlling the speed difference between two moments, the skidding phenomenon caused by drastic speed changes is reduced. Further, the following formula is used to calculate the speed smoothing penalty for the autonomous vehicle:

[0107] r velocitysmooth = -α v ∥v t -v t-1 ∥

[0108] where α v represents the coefficient of the speed smoothing penalty for the autonomous vehicle, v t represents the speed of the autonomous vehicle at time t, v t-1 represents the speed of the autonomous vehicle at time t-1.

[0109] S302. Based on the reward design for the confrontation training task of the autonomous vehicle corresponding to each agent, obtain the second reward function corresponding to the first autonomous vehicle confrontation control strategy model;

[0110] In the confrontation task, the design of the reward function takes into account multiple factors to improve the stability and confrontation effect of the control strategy.

[0111] Specifically, the following formula is used to represent the second reward function:

[0112]

[0113] where, represents the second reward function at time t, r aim represents the attack orientation reward of the autonomous vehicle, r hit represents the attack hit reward of the autonomous vehicle, r miss represents the attack miss penalty of the autonomous vehicle, r risk represents the risk reward of the autonomous vehicle, and the risk reward of the autonomous vehicle is configured to quantify the risk degree of the autonomous vehicle when facing the target autonomous vehicle.

[0114] For the reward of the attack direction of the driverless vehicle, if the driverless vehicle is facing the target driverless vehicle and there is no obstacle on the line connecting the two vehicles, a reward r is given aim 。

[0115] For the risk reward of the driverless vehicle, by effectively analyzing the state s of the driverless vehicle in the driverless vehicle simulation system t and the state es of the target driverless vehicle t , calculate the risk reward r risk =R risk (s t ,es t ). The R risk function is used to analyze the risk situation of the friendly driverless vehicle facing the target driverless vehicle. If it is judged according to the states of the two vehicles that the friendly driverless vehicle faces a high risk of being attacked, collision risk, etc., the R risk function will output a lower value, indicating a greater risk; on the contrary, if the risk is lower, a higher value will be output. The risk reward of the driverless vehicle is obtained by prediction through supervised learning based on the state of the driverless vehicle and the state of the target driverless vehicle, where the prediction label is obtained from the actual data of the two vehicles recorded during the simulation environment training process. In the simulation environment, various actual operation data of the friendly driverless vehicle and the target driverless vehicle will be recorded in real time, as well as the actual situations that occur under these data, such as whether a collision occurs, whether it is attacked, etc. After these actual situations are processed and quantified, they become the label data in supervised learning.

[0116] S303. Based on the combination of the first reward function and the second reward function, construct a reward function corresponding to the first driverless vehicle confrontation control strategy model.

[0117] Specifically, the following formula is used to represent the reward function of the first driverless vehicle confrontation control strategy model:

[0118]

[0119] Among them, R t represents the reward function at time t. At time t, the reward calculation of reinforcement learning is mainly based on the specific state information collected from the driverless vehicle simulation system. Since the reward is the feedback of the driverless vehicle simulation system on the actions made by the agent based on the current information, the calculation of the reward can be regarded as being carried out under complete information, that is, the agent can obtain all relevant state information including both friendly and target driverless vehicles and the environment. These state information include the state information of the friendly vehicle, the vehicle body collision information, the vehicle body speed, the distance between the vehicle body and the target, and the state information of the target vehicle. This design ensures that the calculation of the reward fully considers all aspects of the current environment and task, helps the agent make more optimized decisions, and the reward only plays a role during the decision training of the driverless vehicle and does not affect the decision-making process of the final control decision.

[0120] S4. Create a game session for the driverless vehicle based on the driverless vehicle simulation system, and collect the agent experience data of the driverless vehicle game session;

[0121] Empty the experience replay pool, create a game session for the driverless vehicle based on the driverless vehicle simulation system, and store the agent experience data collected in the current simulation system game session through creating a temporary variable pool. It should be noted that emptying the experience replay pool before starting a new game session is to ensure that the new round of data collection will not be interfered by the data of the previous game session, and to guarantee the independence and pertinence of the subsequent training data. This can ensure that the agent learns based on brand-new interaction information and avoid learning biases caused by the influence of old data.

[0122] Specifically, the agent experience data includes the agent's current moment observation data, the agent's executed actions, the agent's corrected reward value, the agent's next moment observation data, and the game end judgment condition. Further, the agent's current moment observation data includes all relevant state information of both driverless vehicles and the environment. These state information include the state information of the vehicle itself, vehicle body collision information, vehicle body speed, the distance between the vehicle body and the target, and the state information of the target vehicle. The game end judgment condition is whether the target driverless vehicle is damaged, that is, when the target driverless vehicle is damaged, it is determined that the current game session ends.

[0123] The agent's corrected reward value is obtained by calculating the corresponding reward function of the agent. Specifically, it is judged whether the final point of the current game session is reached according to the feedback of the adversarial environment. If the final point is not reached, the agent is controlled to continue making decisions. If the final point is reached, the revenue value of the end state is obtained, and the Monte Carlo method is used to correct the rewards fed back to the agent in the entire adversarial round to obtain the agent's corrected reward value, and the agent experience data in the temporary variable pool is stored in the experience replay pool.

[0124] Further, the following formula is used to calculate the agent's corrected reward value:

[0125] G t =R t +γR t+1 +γ 2 R t+2 +...+γ T-t R t

[0126] where G t represents the agent's corrected reward value at time step t, γ represents the discount factor, which is used to measure the importance of future rewards, and T represents the final time step in the current adversarial round.

[0127] S5. Iteratively train the first unmanned vehicle adversarial control strategy model based on the agent experience data to obtain a second unmanned vehicle adversarial control strategy model, and determine the second unmanned vehicle adversarial control strategy model as the unmanned vehicle adversarial decision-making model corresponding to the unmanned vehicle entity system of the unmanned vehicle simulation system.

[0128] Traditional reinforcement learning does not consider the impact of incomplete information that may occur in the unmanned vehicle adversarial environment, resulting in poor training effects and unstable convergence. The present invention proposes a new reinforcement learning training strategy, where the iterative training includes updating the regular network at fixed intervals based on the policy network according to the number of iterations.

[0129] Specifically, as Figure 4 shown, step S5 includes:

[0130] S501. Randomly extract agent experience data to obtain a training set;

[0131] Randomly extract agent experience data from the experience replay pool to construct a training set. The experience replay pool is a container for storing agent experience data. Randomly extracting data can break the correlation between data and make the training more stable.

[0132] S502. Iteratively train the advantage network in the first unmanned vehicle adversarial control strategy model using a first update strategy to obtain a trained advantage network;

[0133] The first update strategy includes updating the advantage network based on the mean squared error loss function. Specifically, the first update strategy is represented by the following formula:

[0134]

[0135] where, A target represents the advantage network target value, l represents the hyperparameter of the clipping operator, used to control the error range of the Q-value network, A represents the advantage network, I represents the agent observation data, a represents the agent execution action, θ t represents the advantage network parameter at time step t, Q represents the Q-value network, ω t-1 represents the Q-value network parameter at time step t - 1, G t represents the agent corrected reward value, τ represents the hyperparameter of the influence degree of the regular network, π t (a|I) represents the probability of the policy network outputting the agent execution action corresponding to time step t, μ(a|I) represents the probability of the regular network outputting the agent execution action, α actor represents the advantage network hyperparameter, represents the gradient at time step t, MSE lossDenotes the mean squared error loss function, which is used to minimize the difference between the predicted value of the advantage network and the target value of the advantage network.

[0136] S503. Iteratively train the Q-value network in the first unmanned vehicle adversarial control policy model using the second update strategy based on the training set to obtain the trained Q-value network.

[0137] The second update strategy includes updating the Q-value network based on the agent-corrected reward value. Specifically, the second update strategy is represented by the following formula:

[0138]

[0139] where α critic denotes the learning rate parameter of the Q-value network, and MSE loss denotes the mean squared error loss function, which is used to minimize the difference between the predicted value of the Q-value network and the agent-corrected reward.

[0140] S504. Iteratively train the policy network in the first unmanned vehicle adversarial control policy model using the third update strategy based on the training set to obtain the trained policy network.

[0141] The third update strategy includes updating the policy network based on the difference between the advantage network and the Q-value network. Specifically, the third update strategy is represented by the following formula:

[0142] π t+1 (a|I) ∝ exp(A(I,a;θ t ) - Q(I;ω t ))

[0143] where ∝ means proportional, and exp() represents the exponential function. The third update strategy is a probability distribution calculated based on A(I,a;θ t ) - Q(I;ω t ), which is used to guide the agent to make the best action selection in each state.

[0144] S505. Based on the training set, train the regularization network in the first unmanned vehicle adversarial control policy model at fixed intervals according to the number of iterations of the policy network to obtain the trained regularization network.

[0145] The adaptive regularization coefficient is used to control the strength of regularization. In the reinforcement learning of the unmanned vehicle adversarial environment, by adjusting this coefficient, the fitting degree of the first unmanned vehicle adversarial control policy model to the current data and the adaptability to new data can be balanced.

[0146] Specifically, the adaptive regularization coefficient is updated using the following formula:

[0147] τ = τ * n decayrate**(iter / step decayrate )

[0148] where τ represents the adaptive regularization coefficient, n decayrate and step decayrate represent preset descent parameters for controlling the change of the adaptive regularization coefficient, and ** represents the exponentiation symbol.

[0149] Furthermore, the regularization network is updated using the following formula:

[0150] μ ← π t , if iter mod N reg = 0

[0151] where μ represents the regularization network, π t represents the policy network corresponding to time step t, iter represents the number of iterations, and N reg represents the preset hyperparameter of the regularization network. It can be seen from the update condition that the regularization network is updated only when the number of iterations satisfies a specific modulo condition. This setting is to update the regularization network periodically rather than at each iteration. Because overly frequent updates may cause the model to be unstable, while periodic updates can maintain a certain stability while ensuring that the model can adapt to data changes.

[0152] S506. Integrate the trained advantage network, Q-value network, policy network, and regularization network to obtain the second unmanned vehicle confrontation control strategy model for each agent.

[0153] After subsequently determining the second unmanned vehicle confrontation control strategy model as the unmanned vehicle confrontation decision model for the corresponding unmanned vehicle entity system of the unmanned vehicle simulation system, the efficient autonomous behavior of the unmanned vehicle in a complex environment can be achieved.

[0154] In the embodiment of the present invention, a method for controlling the confrontation of unmanned vehicles based on reinforcement learning designs a reward function applicable to the navigation and confrontation scenarios of unmanned vehicles through a reinforcement learning framework, and introduces a regularization network independent of the environment to constrain and optimize the policy, avoiding the policy from falling into local optima, enabling the unmanned vehicle to quickly learn effective policies in different scenarios, being able to maximize the use of state information in the environment, and realizing intelligent control of the unmanned vehicle in complex scenarios.

[0155] Based on the above method for controlling the confrontation of unmanned vehicles based on reinforcement learning, as Figure 5 shown, the embodiment of the present invention provides a device for controlling the confrontation of unmanned vehicles based on reinforcement learning, including:

[0156] The simulation system construction module 1 is used to construct an unmanned vehicle simulation system based on the unmanned vehicle motion control parameters and the unmanned vehicle confrontation environment variables, and create corresponding agents for each unmanned vehicle in the unmanned vehicle simulation system. Among them, the unmanned vehicle motion control parameters include the unmanned vehicle steering control parameters and the unmanned vehicle speed control parameters, and the unmanned vehicle confrontation environment variables include the ground friction coefficient and obstacles.

[0157] The first model construction module 2 is used to construct the first unmanned vehicle confrontation control strategy model for each agent based on the initialization of a number of neural networks. Among them, the number of neural networks includes a policy network, a regularization network, an advantage network, and a Q-value network.

[0158] The reward function construction module 3 is used to construct a reward function corresponding to the first unmanned vehicle confrontation control strategy model based on the reward design for a number of unmanned vehicle training tasks corresponding to each agent. Among them, the number of unmanned vehicle training tasks includes an unmanned vehicle navigation training task and an unmanned vehicle confrontation training task.

[0159] The experience data acquisition module 4 is used to create an unmanned vehicle game session based on the unmanned vehicle simulation system and collect the agent experience data of the unmanned vehicle game session. Among them, the agent experience data includes the agent's current moment observation data, the agent's executed action, the agent's corrected reward value, the agent's next moment observation data, and the game end judgment condition. The agent's corrected reward value is obtained by calculating the reward function corresponding to the agent.

[0160] The confrontation decision model determination module 5 is used to iteratively train the first unmanned vehicle confrontation control strategy model based on the agent experience data to obtain a second unmanned vehicle confrontation control strategy model, and determine the second unmanned vehicle confrontation control strategy model as the unmanned vehicle confrontation decision model of the unmanned vehicle entity system corresponding to the unmanned vehicle simulation system. Among them, the iterative training includes updating the regularization network at a fixed interval according to the number of iterations based on the policy network.

[0161] It should be noted that each module in the above-mentioned unmanned vehicle confrontation control device based on reinforcement learning can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules. For the specific limitations of an unmanned vehicle confrontation control device based on reinforcement learning, refer to the limitations of an unmanned vehicle confrontation control method in the above text. The two have the same functions and effects, and will not be elaborated here.

[0162] In summary, in the embodiment of the present invention, a method and device for adversarial control of an autonomous vehicle based on reinforcement learning design a reward function applicable to the navigation and adversarial scenarios of the autonomous vehicle through a reinforcement learning framework, and introduce a regularization network independent of the environment to constrain and optimize the policy, avoiding the policy from falling into local optimality, enabling the autonomous vehicle to quickly learn effective policies in different scenarios, being able to maximize the use of state information in the environment, and realizing intelligent control of the autonomous vehicle in complex scenarios.

[0163] Each embodiment in this specification is described in a progressive manner. For parts that are the same or similar in each embodiment, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the partial description of the method embodiment for the relevant parts. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0164] The above is only the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art in the technical field, without departing from the technical principle of the present invention, several improvements and replacements can still be made, and these improvements and replacements should also be regarded as the protection scope of the present invention.

Claims

1. An adversarial control method for autonomous vehicles based on reinforcement learning, characterized in that Comprising: Construct a self-driving vehicle simulation system based on the self-driving vehicle motion control parameters and the self-driving vehicle confrontation environment variables, and create a corresponding intelligent agent for each self-driving vehicle in the self-driving vehicle simulation system. Among them, the self-driving vehicle motion control parameters include the self-driving vehicle steering control parameters and the self-driving vehicle speed control parameters, and the self-driving vehicle confrontation environment variables include the ground friction coefficient and obstacles; Based on the initialization of several neural networks, construct the first self-driving vehicle confrontation control strategy model for each intelligent agent. Among them, the several neural networks include a policy network, a regularization network, an advantage network, and a Q-value network; Based on the reward design for several self-driving vehicle training tasks corresponding to each intelligent agent, construct a reward function corresponding to the first self-driving vehicle confrontation control strategy model. Among them, the several self-driving vehicle training tasks include self-driving vehicle navigation training tasks and self-driving vehicle confrontation training tasks; Create a self-driving vehicle game confrontation based on the self-driving vehicle simulation system, and collect the intelligent agent experience data of the self-driving vehicle game confrontation. Among them, the intelligent agent experience data includes the intelligent agent's current moment observation data, the intelligent agent's executed actions, the intelligent agent's corrected reward value, the intelligent agent's next moment observation data, and the game end judgment condition. The intelligent agent's corrected reward value is obtained by calculating the reward function corresponding to the intelligent agent; Iteratively train the first self-driving vehicle confrontation control strategy model based on the intelligent agent experience data to obtain a second self-driving vehicle confrontation control strategy model, and determine the second self-driving vehicle confrontation control strategy model as the self-driving vehicle confrontation decision-making model of the self-driving vehicle entity system corresponding to the self-driving vehicle simulation system. Among them, the iterative training includes updating the regularization network at a fixed interval according to the iteration times based on the policy network.

2. The unmanned vehicle confrontation control method according to claim 1, wherein, The constructing the first self-driving vehicle confrontation control strategy model for each intelligent agent based on the initialization of several neural networks includes: Determine the input dimension and output dimension of each neural network. Among them, the input dimension of each neural network is D-dimensional, the output dimensions of the policy network, the regularization network, and the advantage network are all K-dimensional, and the output dimension of the Q-value network is one-dimensional. Both D and K are integers greater than or equal to 2; Initialize the policy network so that the policy network is configured to generate the probability distribution of each intelligent agent selecting each executed action in the corresponding self-driving vehicle current state; Based on the policy network, initialize the regularization network so that the regularization network is configured to perform regularization constraints on the policy network, the advantage network, and the Q-value network respectively; Initialize the advantage network so that the advantage network is configured to evaluate the advantage degree of each intelligent agent selecting each executed action in the corresponding self-driving vehicle current state; Initialize the Q-value network so that the Q-value network is configured to evaluate the Q-value of each intelligent agent for each executed action in the corresponding self-driving vehicle current state; Integrate the policy network, the regularization network, the advantage network, and the Q-value network to obtain the first self-driving vehicle confrontation control strategy model for each intelligent agent.

3. The unmanned vehicle confrontation control method according to claim 1, characterized in that Based on the reward design for each of the agents corresponding to several unmanned vehicle training tasks, a reward function corresponding to the first unmanned vehicle adversarial control strategy model is constructed, including: Based on the reward design for each of the agents corresponding to the unmanned vehicle navigation training task, a first reward function corresponding to the first unmanned vehicle adversarial control strategy model is obtained; Based on the reward design for each of the agents corresponding to the unmanned vehicle adversarial training task, a second reward function corresponding to the first unmanned vehicle adversarial control strategy model is obtained; Based on the combination of the first reward function and the second reward function, a reward function corresponding to the first unmanned vehicle adversarial control strategy model is constructed.

4. The method for unmanned vehicle confrontation control according to claim 3, wherein The first reward function is represented by the following formula: Among them, represents the first reward function at time t, r collision represents the unmanned vehicle collision penalty, r distance represents the unmanned vehicle distance reward, r velocitysmooth represents the unmanned vehicle speed smoothing penalty; The following formula is used to calculate the unmanned vehicle collision penalty: The following formula is used to calculate the unmanned vehicle distance reward: Among them, α d represents the weight coefficient of the distance reward of the driverless vehicle, represents the distance between the driverless vehicle and the target driverless vehicle; The following formula is used to calculate the unmanned vehicle speed smoothing penalty: r velocitysmooth = -α v ∥v t -v t-1 ∥ Among them, α v represents the coefficient of the smoothness penalty of the speed of the driverless vehicle, v t represents the speed of the driverless vehicle at time t, v t-1 represents the speed of the driverless vehicle at time t - 1.

5. The method for controlling the confrontation of driverless vehicles according to claim 3, characterized in that, The second reward function is represented by the following formula: Among them, represents the second reward function at time t, r aim represents the unmanned vehicle attack orientation reward, r hit represents the unmanned vehicle attack hit reward, r miss represents the unmanned vehicle attack miss penalty, r risk represents the unmanned vehicle risk reward, and the unmanned vehicle risk reward is configured to quantify the risk level of the unmanned vehicle when facing the target unmanned vehicle.

6. The method for controlling the confrontation of unmanned vehicles according to claim 5, characterized in that, The unmanned vehicle risk reward is obtained by prediction through supervised learning based on the unmanned vehicle state and the target unmanned vehicle state.

7. The method for controlling the confrontation of driverless vehicles according to claim 1, characterized in that, The second unmanned vehicle adversarial control strategy model is obtained by iteratively training the first unmanned vehicle adversarial control strategy model based on the agent experience data, including: Randomly extract the agent experience data to obtain a training set; Based on the training set, the advantage network in the first unmanned vehicle adversarial control strategy model is iteratively trained using a first update strategy to obtain the trained advantage network, where the first update strategy includes updating the advantage network based on the mean squared error loss function; Based on the training set, the Q-value network in the first unmanned vehicle adversarial control strategy model is iteratively trained using a second update strategy to obtain the trained Q-value network, where the second update strategy includes updating the Q-value network based on the agent corrected reward value; Based on the training set, the policy network in the first unmanned vehicle adversarial control strategy model is iteratively trained using a third update strategy to obtain the trained policy network, where the third update strategy includes updating the policy network based on the difference between the advantage network and the Q-value network; Based on the training set, the regularization network in the first unmanned vehicle adversarial control strategy model is trained according to the policy network at fixed intervals of the iteration times to obtain the trained regularization network; Integrate the trained advantage network, the trained Q-value network, the trained policy network, and the trained regularization network to obtain the second unmanned vehicle adversarial control strategy model for each of the agents.

8. The method for controlling the confrontation of driverless vehicles according to claim 7, characterized in that, The first update strategy is represented by the following formula: A target ← min{l, max{0, A(I, a; θ t-1 ) - Q(I; ω t-1 )}} + G t -τ log(π t (a|I) / μ(a|I)) Among them, A target represents the advantage network target value, l represents the hyperparameter of the clipping operator, A represents the advantage network, I represents the agent's observation data, a represents the agent's executed action, θ t represents the advantage network parameters at time step t, Q represents the Q-value network, ω t-1 represents the Q-value network parameters at time step t - 1, G t represents the agent's corrected reward value, τ represents the hyperparameter of the influence degree of the regularization network, π t (a|I) represents the probability of the policy network outputting the agent's executed action corresponding to time step t, μ(a|I) represents the probability of the regularization network outputting the agent's executed action, α actor represents the hyperparameter of the advantage network, represents the gradient at time step t, MSE loss represents the mean squared error loss function; The second update strategy is represented by the following formula: where α critic represents the learning rate parameter of the Q-value network; The third update strategy is represented by the following formula: π t+1 (a|I) ∝ exp(A(I, a; θ t ) - Q(I; ω t )) Where, ∝ represents proportional, and exp() represents the exponential function.

9. The unmanned vehicle confrontation control method according to claim 7, characterized in that The regularization network is updated by the following formula: μ ← π t , if iter mod N reg = 0 Among them, μ represents the regular network, and π t represents the policy network corresponding to the time step t, iter represents the number of iterations, and N reg represents the preset hyperparameter of the regular network.

10. An unmanned vehicle countermeasure control device based on reinforcement learning, characterized in that, Including: A simulation system construction module is used to construct an unmanned vehicle simulation system based on unmanned vehicle motion control parameters and unmanned vehicle confrontation environment variables, and create corresponding agents for each unmanned vehicle in the unmanned vehicle simulation system. Among them, the unmanned vehicle motion control parameters include unmanned vehicle steering control parameters and unmanned vehicle speed control parameters, and the unmanned vehicle confrontation environment variables include ground friction coefficient and obstacles; A first model construction module is used to construct a first unmanned vehicle confrontation control strategy model for each of the agents based on the initialization of a number of neural networks. Among them, the number of neural networks includes a policy network, a regularization network, an advantage network, and a Q-value network; A reward function construction module is used to construct a reward function corresponding to the first unmanned vehicle confrontation control strategy model based on the reward design for a number of unmanned vehicle training tasks corresponding to each of the agents. Among them, the number of unmanned vehicle training tasks includes unmanned vehicle navigation training tasks and unmanned vehicle confrontation training tasks; An experience data acquisition module is used to create an unmanned vehicle game confrontation based on the unmanned vehicle simulation system and collect the agent experience data of the unmanned vehicle game confrontation. Among them, the agent experience data includes the agent's current moment observation data, the agent's executed actions, the agent's corrected reward value, the agent's next moment observation data, and the game end judgment condition. The agent's corrected reward value is obtained by calculating the reward function corresponding to the agent; An adversarial decision model determination module is used to iteratively train the first unmanned vehicle confrontation control strategy model based on the agent experience data to obtain a second unmanned vehicle confrontation control strategy model, and determine the second unmanned vehicle confrontation control strategy model as the unmanned vehicle confrontation decision model of the unmanned vehicle entity system corresponding to the unmanned vehicle simulation system. Among them, the iterative training includes updating the regularization network at fixed intervals according to the number of iterations based on the policy network.

Citation Information

Cited By

  • Humanoid robot gait control method and system based on imitation and inverse reinforcement learning

    CN120524975A

  • Gradient-free reinforcement learning method based on acceleration and deceleration strategy

    CN121480599A

  • A gradient-free reinforcement learning method based on acceleration / deceleration strategies

    CN121480599B