An Unmanned Vehicle Confrontation and Obstacle Avoidance Method Based on Progressive Deep Reinforcement Learning

Through progressive deep reinforcement learning and progressive self-game SAC algorithm, the problems of environmental complexity and sparse rewards in autonomous decision-making of unmanned vehicles are solved, efficient learning and generalization capabilities are achieved, and robust and generalized autonomous decision-making strategies are obtained.

CN116243727BActive Publication Date: 2025-06-20XIAMEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310260597.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-17
Publication Date
2025-06-20
Estimated Expiration
2043-03-17

AI Technical Summary

Technical Problem

When the existing technology applies reinforcement learning algorithms to make autonomous decisions with unmanned vehicles, it faces the problems of complex environmental models, high state space dimensions and sparse reward information, which leads to low training efficiency and difficulty in obtaining the optimal strategy.

Method used

The method of progressive deep reinforcement learning is adopted, and the progressive self-game SAC algorithm is designed, and the critic neural network and the executor neural network are used, combining the entropy increase mechanism and the difficulty of training courses, and automatic entropy and course learning mechanisms are designed to gradually improve the autonomous decision-making ability of unmanned vehicles.

Benefits of technology

Improve learning efficiency and enhance generalization capabilities, so that unmanned vehicles can generalize from static tasks to random dynamic tasks, and obtain a final strategy that is robust and generalized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116243727B_ABST
    Figure CN116243727B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning, comprising the following steps: S1, according to the kinematic model of the unmanned vehicle, solving and modeling through the Runge-Kutta method; S2, simulating the autonomous decision-making process of multiple unmanned vehicle systems by computer; S3, designing and optimizing the form, size and number of the critic neural network and the actor neural network of the progressive self-play SAC algorithm, constructing the objective loss function of the actor neural network and the policy entropy coefficient α for the real motion situation of the unmanned vehicle, and designing an automatic entropy in combination with the entropy increase mechanism and the training course difficulty to obtain the progressive self-play SAC algorithm; S4, using the progressive self-play SAC algorithm to self-regulate the training course difficulty and execute the self-play process to complete a learning course; S5, repeating step S4 to obtain the trained actor neural network for generating real-time decisions for unmanned vehicle confrontation and obstacle avoidance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep reinforcement learning, and specifically refers to an unmanned vehicle confrontation and obstacle avoidance method for progressive deep reinforcement learning. Background Art

[0002] With the rapid development of sensor technology, computer technology, and communication technology, the performance of both military and civilian unmanned vehicles has been significantly improved. Autonomous decision-making is one of the core research contents in the current research of unmanned vehicle systems, and it has very important value for expanding the application scenarios and functions of unmanned vehicles. In the military field, unmanned vehicles can complete more difficult and complex tasks than manned vehicles, so they have become weapons and equipment that countries are competing to develop. Unmanned vehicles have far more advantages than manned vehicles in terms of product types, application fields, and the ability to perform tasks. However, most current unmanned vehicles still cannot perform tasks without the operation and decision-making of remote control personnel. This working mode makes the application of unmanned vehicles still highly dependent on wireless communication technology and the decision-making ability of remote control personnel, and it is easily restricted by communication conditions and the decision-making ability of remote operators, making it difficult to adapt to highly dynamic application scenarios, especially the complex and changeable battlefield situations in the military field. In many research works on unmanned vehicle autonomous decision-making systems, autonomous decision-making schemes usually adopt technologies such as optimization principles and artificial intelligence to automatically generate autonomous decision-making instructions in various application scenarios. At the theoretical level, the theoretical methods to solve the autonomous decision-making problem can be roughly divided into three categories, namely: game theory, optimization theory, and artificial intelligence methods. Among them, the method based on game theory mainly reflects the situation in the confrontation process by establishing a mathematical model, and forms the optimal decision through differential game and influence diagram algorithms. When facing highly dynamic battlefield situations, due to the complexity of the model, it is often difficult to solve the optimal decision in real time, so there are still great difficulties in practical applications. Methods such as genetic algorithms, Bayesian inference, and statistical theory based on optimization theory transform the problem into an optimization problem for mathematical solution to obtain the autonomous optimal strategy. However, when facing large-scale problems, there is also a problem of poor real-time decision-making solution, and it is also difficult to guarantee the optimality of the solution when facing a large number of non-convex optimization problems. In addition, the above methods are mostly used for offline tactical optimization research. Artificial intelligence-based methods include expert systems, neural networks, and reinforcement learning methods. The core of the expert system method is to describe the decision-making behavior as a rule base according to expert experience, and then form control instructions through rule reasoning according to specific situations. The establishment of the rule base is relatively complex, and as a fixed strategy, it is also easy to be cracked. The neural network method regards the autonomous decision-making behavior as a "black box" and forms an adversarial strategy through learning a large number of effective adversarial sample data. However, in practical applications, it is difficult to obtain effective learning samples, and the performance of autonomous decision-making is limited by the performance of sample data, making it difficult to achieve further optimization. Compared with the above methods, the method based on reinforcement learning neither requires experts to provide a rule base nor depends on an environmental model. Instead, on the basis of the optimization principle, through the interaction between the agent and the environment, using the state information and reward signals feedback by the environment, it continuously optimizes the strategy through online or offline learning algorithms and finally obtains the optimal strategy.In addition, the decision-making behavior of reinforcement learning is generally expressed by neural networks. On the premise of sufficient training, it not only has strong non-linear expression ability, but also has good generalization.

[0003] Performance can make the final autonomous decision-making scheme have both optimality in performance and good robustness in environmental adaptability. Therefore, at present, the reinforcement learning method has become an effective solution to the problem of autonomous decision-making of unmanned vehicles.

[0004] In the field of research on autonomous decision-making of unmanned vehicles, the problems of autonomous maneuvering and adversarial decision-making of unmanned aerial vehicles have received wide attention. At present, in this research direction, most of the reinforcement learning-based solutions are mainly based on the DQN algorithm. By decomposing the decision-making behavior of the unmanned aerial vehicle into a series of discrete actions, the complexity of solving and optimizing the autonomous decision-making problem is reduced. However, this simplification results in a large difference from the real situation, and the adversarial performance is difficult to guarantee. If we want to be as close as possible to the real situation, the design problem often needs to face continuous and high-dimensional state and action spaces, which easily causes the dimensionality disaster and sparse reward problems in the reinforcement learning process, and the learning efficiency is extremely low. Although the DDPG algorithm can be used for the policy optimization problem of continuous state and action spaces, there are many hyperparameters in the algorithm design, and the training is prone to falling into local optima. As a more advanced reinforcement learning algorithm, the original SAC algorithm has fewer hyperparameters, but it is also difficult to solve the problems of sparse rewards, complex and changeable environments, and autonomous adversarial decision-making with multiple tasks at the same time. To sum up, when the intelligent decision-making technology is currently used for the autonomous decision-making of unmanned vehicles, there are still problems such as complex environmental models, high-dimensional state spaces, and sparse reward information, which make the training efficiency of traditional reinforcement learning algorithms generally low and it is difficult to obtain the optimal strategy.

[0005] The purpose of the research of the present invention is to design an adversarial and obstacle avoidance method for unmanned vehicles based on progressive deep reinforcement learning to solve the problems existing in the above-mentioned prior art. Summary of the Invention

[0006] Aiming at the problems existing in the above-mentioned prior art, the present invention provides an adversarial and obstacle avoidance method for unmanned vehicles based on progressive deep reinforcement learning, which can effectively solve at least one of the problems existing in the above-mentioned prior art.

[0007] The technical solution of the present invention is as follows:

[0008] An adversarial and obstacle avoidance method for unmanned vehicles based on progressive deep reinforcement learning, comprising the following steps:

[0009] S1. Based on the kinematic bicycle model, solve and model through the Runge-Kutta method, construct it into a standard gym environment class in the Python environment, and mathematically represent and describe in computer language the state data of the unmanned vehicle's own state and environmental observations as necessary elements according to the actual situation;

[0010] S2. Through computer simulation of the autonomous decision-making process of multiple unmanned vehicle systems, generate simulation data of the movement process and decision-making behavior of the unmanned vehicle;

[0011] S3. Design and optimize the form, size, and quantity of the critic neural network and the actor neural network of the progressive self-play SAC algorithm, construct the actor neural network and the policy entropy coefficient for the real movement situation of the unmanned vehicle, design the automatic entropy in combination with the entropy increase mechanism and the training course difficulty, and design a course learning mechanism that increases with the training process and the strategy type and intensity of the adversarial opponent for the complexity of the unmanned vehicle decision-making scenario to obtain the progressive self-play SAC algorithm;

[0012] S4. Use the progressive self-play SAC algorithm to self-regulate the training course difficulty and execute the self-play process, generate multiple decision-making data of the unmanned vehicle and put them into the experience replay pool to update the data of different course learning, average-sample the latest data in the experience replay pool and update the parameters of the critic neural network and the actor neural network to complete one learning course;

[0013] S5. Repeat step S4 to enable the critic neural network and the actor neural network to complete several learning courses, and obtain the trained actor neural network for generating real-time decisions for the unmanned vehicle to confront and avoid obstacles.

[0014] Furthermore, in S1, the state data of the own state and environmental observations includes one or more of the position coordinates, real-time speed, yaw angle of the unmanned vehicle, the distance to the obstacle, and the position coordinates, real-time speed, yaw angle of the adversarial opponent.

[0015] Furthermore, the results obtained by designing and optimizing the form, size, and quantity of the critic neural network and the actor neural network of the progressive self-play SAC algorithm are:

[0016] The critic neural network adopts a fully connected neural network structure with two hidden layers, and the number of neurons in each layer is 256. The number of critic neural networks is two or more, and each critic neural network corresponds to a target critic neural network with low-frequency updates.

[0017] Furthermore, the target loss function for constructing the actor neural network for the real movement situation of the unmanned vehicle includes:

[0018] Provide exploration ability based on the policy entropy mechanism and design a loss function that balances exploration ability and policy optimization;

[0019] The designed loss function Satisfies formula (1):

[0020] Formula (1);

[0021] Wherein, is the mathematical expectation, is the current state data of the unmanned vehicle, is the experience replay pool, is the action output by the executor neural network. Through the powerful expressive ability of the neural network, is modeled as an executor policy network that generates a mapping from state to specific action, is the parameter of the executor neural network, represents that when the given state is present, the executor policy outputs a certain action with a probability of, is the policy entropy coefficient, initialized to 1, represents the critic's long-term discounted return evaluation of the current state-action value of the unmanned vehicle.

[0022] Furthermore, a target loss function for constructing the policy entropy coefficient for the actual motion situation of the unmanned vehicle includes:

[0023] Define as a monotonically increasing function of the difficulty coefficient k. As the training course progresses, the policy entropy coefficient is updated according to the different course difficulties. The designed loss function of the policy entropy coefficient Satisfies formula (2):

[0024] Formula (2);

[0025] Wherein, is the target information entropy of the policy. In the early stage of training, as the course difficulty increases and the sparse reward problem intensifies, the policy entropy coefficient is increased to enhance the exploration ability. In the later stage of training, as self-play optimization progresses and the policy converges stably, the policy entropy coefficient is decreased to obtain a stable and reliable policy network.

[0026] Furthermore, according to the kinematic bicycle model, it is solved and modeled by the Runge-Kutta method, and constructed into a standard gym environment class in the Python environment, including:

[0027] The kinematic bicycle model is based on the kinematics of the unmanned vehicle system. The differential equation of the kinematic bicycle model is solved precisely by the fourth-order Runge-Kutta algorithm to obtain the state observation of the unmanned vehicle at the next moment, and the kinematic bicycle model is encapsulated as a standard class function and unified under the gym framework.

[0028] Further, in S4, the process of self-regulating the training course difficulty and performing the self-play process by using the progressive self-play SAC algorithm includes:

[0029] In the later stage of training, an iterative adversarial training strategy is adopted. When the strategy is optimized to a certain stage through the progressive self-play SAC framework, the opponent continuously uses the optimized strategy in the previous round to optimize the strategy in the next round, finds the defects in the previous round of strategy, and makes targeted checks and supplements to the strategy network in an autonomous manner to perform deeper-level strategy optimization.

[0030] Further, after S5, execute:

[0031] S6, evaluate the overall advantage of the executor neural network that has completed training through formula (3), and define the overall advantage evaluation function of the th step of the unmanned vehicle as:

[0032] Formula (3);

[0033] where represents the performance index of tracking and confrontation, represents the obstacle avoidance performance index, and are proportionality coefficients, measuring the proportion of different indicators in the overall advantage indicator, represents the comprehensive advantage performance index, and adaptively adjusts the proportionality coefficients in the overall advantage evaluation function at different stages of the course with different difficulties.

[0034] Further, in S4, the parameters of the critic neural network and the executor neural network are updated by using the gradient descent algorithm based on historical momentum gradient.

[0035] Further, in S5, generating real-time decisions for the unmanned vehicle to confront and avoid obstacles includes: one or more of static obstacle avoidance, trajectory tracking, comprehensive confrontation, generalization to dynamic obstacle avoidance, and random trajectory tracking.

[0036] Therefore, the present invention provides the following effects and / or advantages:

[0037] This application adopts progressive curriculum learning, avoids sparse appreciation, improves learning efficiency, enhances generalization ability, and can generalize from static tasks to random dynamic tasks. Specifically, as the training process progresses, from a scenario without obstacles and a single and fixed opponent strategy to a scenario with a small number of fixed obstacles, and then to a scenario with a large number of randomly moving obstacles and a diverse and random opponent strategy, through multi-curriculum learning in complex scenarios, the unmanned vehicle gradually masters complex comprehensive autonomous decision-making capabilities. Through progressive learning, it can ultimately make decisions in various types of tasks, and the obtained strategy has a certain degree of robustness and generalization ability.

[0038] The present invention utilizes the progressive self-play SAC algorithm, which does not require an environmental model. By continuously interacting with the environment in a designed training curriculum with an increasing difficulty gradient, a large amount of new training data is generated. Therefore, it has the advantage of being model-free and purely data-driven.

[0039] The progressive self-play SAC algorithm provided by the present invention, through progressive curriculum learning, avoids sparse appreciation, improves learning efficiency, enhances generalization ability, and can generalize from static tasks to random dynamic tasks. Specifically, as the training process progresses, from a scenario without obstacles and a single and fixed opponent strategy to a scenario with a small number of fixed obstacles, and then to a scenario with a large number of randomly moving obstacles and a diverse and random opponent strategy, through multi-curriculum stage learning in complex scenarios, the unmanned vehicle gradually masters complex comprehensive autonomous decision-making capabilities. Through progressive learning, it can ultimately make decisions in various types of tasks, and the obtained strategy has a certain degree of robustness and generalization ability.

[0040] The present invention trains the strategy by adopting an iterative adversarial method in the later stage of training to ensure the in-depth optimization of the strategy. It can find the defects in the previous round of the strategy and conduct targeted checking and filling of the strategy network in an autonomous manner to perform deeper-level strategy optimization.

[0041] The present invention comprehensively considers multi-task scenarios such as tracking, confrontation, and obstacle avoidance, adaptively allocates task weights, and can make decisions for more complex tasks closer to the real world.

[0042] It should be understood that the above summary and the following detailed description of the present invention are exemplary and explanatory, and are intended to provide further explanation of the present invention as claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic flow chart provided for one embodiment of the present invention.

[0044] Figure 2 It is a vehicle three-degree-of-freedom particle model diagram established for one embodiment of the present invention.

[0045] Figure 3It is a diagram of a vehicle's one-on-one confrontation and obstacle avoidance model.

[0046] Figure 4 It is a schematic diagram of the autonomous decision-making framework of the progressive self-play SAC algorithm provided by one embodiment of the present invention.

[0047] Figure 5 It is a schematic diagram of the autonomous decision-making learning framework of a multi-course stage unmanned vehicle provided by one embodiment of the present invention.

[0048] Figure 6 It is a schematic diagram of the comparison of the learning efficiency and performance between the progressive self-play SAC algorithm and the original SAC algorithm.

[0049] Figure 7 It is a schematic diagram of decision-making strategy evaluation, trajectory tracking, obstacle avoidance and advantage evaluation indicators.

[0050] Figure 8 It is a diagram of the final strategy performance of the present invention, comparing the decision-making tracking situation with and without obstacles. Detailed implementation method

[0051] For the convenience of those skilled in the art to understand, the embodiments will now be further described in detail with reference to the accompanying drawings: It should be understood that the steps mentioned in this embodiment, unless specifically stating their order, can be adjusted according to actual needs in their front and back order, and can even be executed simultaneously or partially simultaneously.

[0052] Refer to Figure 1 , a method for the confrontation and obstacle avoidance of an unmanned vehicle based on progressive deep reinforcement learning, comprising the following steps:

[0053] S1, according to the kinematic model of the unmanned vehicle, solve and model through the Runge-Kutta method, construct a standard gym environment class in the Python environment, and mathematically express and describe in computer language the state data of the unmanned vehicle itself and the environmental observation as necessary elements according to the actual situation;

[0054] In this embodiment, the unmanned vehicle can be an unmanned aerial vehicle, an unmanned vehicle, etc. In this embodiment, if it is an unmanned vehicle, the kinematic model of the unmanned vehicle adopts a kinematic bicycle model, and the kinematic bicycle model is a prior art.

[0055] In this embodiment, reinforcement learning is a model-independent intelligent decision-making technology that can be driven only by pure data. Among them, the original SAC algorithm is one of the most advanced algorithms in current reinforcement learning. Due to the mechanism of policy entropy, the original SAC algorithm has strong policy space exploration ability, can handle continuous action spaces, and the final policy has high robustness and can make real-time decisions based on state feedback information. The original SAC algorithm is a reinforcement learning algorithm based on the Actor-Critic framework proposed by the Pieter Abbeel and Sergey Levine teams, and the reference is "Soft Actor-Critic Algorithms and Applications" (Haarnoja et al., 29, Jan, 2019).

[0056] The optimal policy of the original SAC algorithm is defined as follows:

[0057] ;

[0058] where, is the information entropy of the policy and is the temperature coefficient, which is used to adjust the weights of the information entropy and during the process of optimizing the policy.

[0059] In the original SAC algorithm, the optimization objective of the value network is to minimize the following Bellman residual metric:

[0060] ;

[0061] where the state value function involved is defined as follows:

[0062] ;

[0063] The above objective function is optimized using the stochastic gradient algorithm.

[0064] The objective function of the policy network in the original SAC algorithm is defined as follows:

[0065] ;

[0066] where, are the parameters of the policy network. In order to optimize the network parameters using the backpropagation algorithm, it is necessary to reparameterize the action . For this, it can be sampled from some fixed distribution functions, such as spherical Gaussian distribution, etc. The above objective function is optimized using the stochastic gradient algorithm.

[0067] Furthermore, the temperature coefficient It also needs to be updated automatically, and its objective function is defined as follows:

[0068] ;

[0069] It can be seen from the core formula of the original SAC algorithm above that while maximizing the cumulative reward, the original SAC algorithm also maximizes the policy information entropy. It is precisely by maximizing the policy information entropy that the exploration ability of the algorithm is guaranteed, so it is not easy to fall into local optima. At the same time, the temperature coefficient is automatically adjusted during the training process At the early stage of training is relatively large to ensure that the agent has good exploration ability, and it gradually decreases after learning a certain policy in the later stage to maintain the stability of training.

[0070] S2, through computer simulation of the autonomous decision-making process of multiple unmanned vehicle systems, generate simulation data of the movement process and decision-making behavior of the unmanned vehicle;

[0071] S3, design and optimize the forms, sizes, and numbers of the critic neural network and the actor neural network of the progressive self-play SAC algorithm, construct the actor neural network and the target loss function of the policy entropy coefficient for the real movement situation of the unmanned vehicle, and combine the entropy increase mechanism and the training course difficulty to design an automatic entropy, and design a course learning mechanism that increases with the training process and the types and intensities of the strategies of the adversarial opponents for the complexity of the unmanned vehicle decision-making scenario, to obtain the progressive self-play SAC algorithm;

[0072] S4, use the progressive self-play SAC algorithm to self-regulate the training course difficulty and execute the self-play process, generate multiple decision-making data of the unmanned vehicle and put them into the experience replay pool for updating different course learning data, average-sample the latest data in the experience replay pool and update the parameters of the critic neural network and the actor neural network to complete one learning course;

[0073] S5, repeat step S4 to make the critic neural network and the actor neural network complete several learning courses, and obtain the trained actor neural network for generating real-time decisions for the unmanned vehicle to confront and avoid obstacles.

[0074] Furthermore, in S1, the state data of the self-state and environmental observation include one or more of the position coordinates, real-time speed, yaw angle of the unmanned vehicle, the distance to the obstacle, and the position coordinates, real-time speed, yaw angle of the adversarial opponent.

[0075] Furthermore, the results obtained by designing and optimizing the forms, sizes, and numbers of the critic neural network and the actor neural network of the progressive self-play SAC algorithm are:

[0076] The critic neural network adopts a fully connected neural network structure with two hidden layers, and the number of neurons in each layer is 256. The number of critic neural networks is two or more, and each critic neural network corresponds to a target critic neural network with low-frequency updates.

[0077] Furthermore, the target loss function for constructing the actor neural network according to the actual motion of the unmanned vehicle includes:

[0078] Based on the policy entropy mechanism to provide exploration ability, design a loss function that balances exploration ability and policy optimization;

[0079] The designed loss function satisfies formula (1):

[0080] Formula (1);

[0081] where, is the mathematical expectation, is the current state data of the unmanned vehicle, is the experience replay pool, is the action output by the actor neural network. Through the powerful expression ability of the neural network, is modeled as an actor policy network that generates a state-to-specific action mapping, is the parameter of the actor neural network, represents that when the given state is given, the actor policy outputs a certain action with a probability of, is the policy entropy coefficient, initialized to 1, represents the critic's long-term discounted return evaluation of the current state-action value of the unmanned vehicle.

[0082] Furthermore, the target loss function for constructing the policy entropy coefficient according to the actual motion of the unmanned vehicle includes:

[0083] Define as a monotonically increasing function of the difficulty coefficient k. As the training course progresses, the policy entropy coefficient is updated according to the different course difficulties. The designed loss function of the policy entropy coefficient satisfies formula 2:

[0084] Formula (2);

[0085] where, is the target information entropy of the policy. In the early stage of training, as the course difficulty increases and the sparse reward problem intensifies, the policy entropy coefficient will increase To increase the exploration ability, during the later stage of training, as the self-play optimization progresses and the strategy converges stably, the strategy entropy coefficient is decreased to obtain a stable and reliable policy network.

[0086] Furthermore, according to the kinematic bicycle model, solving and modeling through the Runge-Kutta method, the construction of the standard gym environment class in the Python environment includes:

[0087] The kinematic bicycle model is based on the kinematics of the unmanned vehicle system. The differential equation of the kinematic bicycle model is solved precisely through the fourth-order Runge-Kutta algorithm to obtain the state observation of the unmanned vehicle at the next moment, and the kinematic bicycle model is encapsulated as a standard class function, unified under the gym framework.

[0088] Furthermore, in S4, the process of using the progressive self-play SAC algorithm to self-regulate the training course difficulty and perform the self-play process includes:

[0089] In the later stage of training, an iterative adversarial training strategy is adopted. When the strategy is optimized to a certain stage through the progressive self-play SAC framework, the opponent continuously uses the optimized strategy of the previous round to optimize the strategy of the next round, finds the defects in the previous round of strategy, and makes targeted checks and supplements to the policy network in an autonomous manner to perform deeper-level strategy optimization.

[0090] Furthermore, after S5, execute:

[0091] S6, evaluate the overall advantage of the executor neural network that has completed training through formula 3, and define the overall advantage evaluation function for the nth step of the unmanned vehicle as:

[0092] Formula (3);

[0093] where represents the performance index of tracking and confrontation, represents the obstacle avoidance performance index, and are proportionality coefficients, measuring the proportion of different indicators in the overall advantage indicator, represents the comprehensive advantage performance index, and adaptively adjusts the proportionality coefficient in the overall advantage evaluation function at different stages of the difficulty course.

[0094] Furthermore, in S4, the parameters of the critic neural network and the executor neural network are updated using the gradient descent algorithm based on historical momentum gradient.

[0095] Further, in S5, generating real-time decisions for the confrontation and obstacle avoidance of the unmanned vehicle includes one or more of the following: static obstacle avoidance, trajectory tracking, comprehensive confrontation, generalization to dynamic obstacle avoidance, and random trajectory tracking.

[0096] The following is the specific description of S1 - S5 of this application.

[0097] First, for the convenience of subsequent description, discussion, and verification, as Figure 2 shown, the present invention directly considers the autonomous confrontation and autonomous obstacle avoidance of an unmanned vehicle on a two-dimensional plane as the actual application scenario. However, from the proposed design scheme, the proposed design method can be fully extended to more complex application scenarios, such as the automatic following driving of unmanned vehicles and the air combat confrontation of unmanned aerial vehicles, after appropriate expansion. The present invention considers the application scenario as the autonomous decision-making of a system with two small vehicles for short-range tracking confrontation and obstacle avoidance in an obstacle-containing two-dimensional environment. This scenario is usually studied as a simplified scenario for the autonomous decision-making confrontation problem of unmanned fighter jets in three-dimensional space, as Figure 3 shown. Assume that in this scenario the vehicle represents the tracking vehicle, the vehicle is the vehicle to be tracked. What needs to be designed is the autonomous motion and obstacle avoidance strategy of the vehicle. The design goal is to enable the vehicle to avoid dangerous collisions with obstacles during the movement process and, relative to the vehicle, to maintain the best tracking posture as much as possible.

[0098] In S1, in order to establish a kinematic bicycle model, first establish a two-dimensional coordinate system including axis and axis. In this coordinate system, the motion of the vehicle can be described by the following differential equations:

[0099] ;

[0100] ;

[0101] Formula (4);

[0102] ;

[0103] ;

[0104] Among them, is the position vector of the small vehicle in the two-dimensional coordinate, represents the magnitude of the small vehicle speed, , respectively represent the speed in axis and The component on the axis represents the azimuth angle of the vehicle, that is, the angle between the vehicle body direction and the axis represents the distance between the rear of the vehicle and the steering center represents the distance between the front of the vehicle and the steering center represents the angle between the velocity direction of the steering center and the vehicle body. Assume that the control steering angle of the front wheels of the vehicle relative to the vehicle body direction is , and the rear wheels are the controlled driving wheels, and the magnitude of the driving force is described by the acceleration . In the above model is the operating variable for controlling the movement of the vehicle

[0105] Then, the present invention establishes a model of the autonomous decision-making process for tracking, confrontation and obstacle avoidance of two vehicles in a two-dimensional environment with obstacles. This scenario is usually studied as a simplified scenario of the autonomous decision-making confrontation problem of unmanned combat aircraft in three-dimensional space. Assume that in this scenario the vehicle represents the tracking vehicle the vehicle is the vehicle to be tracked, and what needs to be designed is the autonomous movement and obstacle avoidance strategy of the vehicle. The design goal is to make the vehicle avoid dangerous collisions with obstacles during the movement process, and relative to the vehicle can also maintain the best tracking posture as much as possible

[0106] As Figure 3 shows the position schematic diagram of the vehicle , the vehicle and the obstacle at any moment. Among them represents the spatial position of the vehicle represents the spatial position of the vehicle represents the velocity vector of the vehicle represents the velocity vector of the vehicle represents the position of the nearest obstacle in the velocity direction of the vehicle . represents the distance between the vehicle and the nearest obstacle in the positive velocity direction (varies with the velocity direction of the vehicle and is infinite when there is no obstacle in the moving direction). The gray fan-shaped area behind the vehicle body represents the favorable confrontation (attack) area for the vehicle . This area changes with the spatial position and velocity direction of the vehicle is The best confrontation (attack) position of the vehicle is generally the center point of the fan-shaped confrontation area.

[0107] A large amount of autonomous decision-making data generated by the above simulation process mainly consists of three parts: state vectors, reward signals, and decision-making actions.

[0108] In the present invention, based on the characteristics of the original SAC algorithm having policy entropy, combined with the curriculum design of different training stages, the policy entropy is adaptively adjusted to improve the learning and exploration efficiency.

[0109] According to the basic components of reinforcement learning, at each moment , a set of state quantities is defined as the state information that the intelligent agent the vehicle can observe, and is also used to calculate the advantage evaluation function value to evaluate the current situation. For the confrontation system composed of two vehicles, the following environmental state information is defined:

[0110] Formula (5);

[0111] And the action vector is defined:

[0112] Formula (6);

[0113] Among them, represents the acceleration of the vehicle, corresponding to the throttle control in a real car, represents the steering angle of the vehicle, corresponding to the steering wheel control in a real car.

[0114] And based on the advantage evaluation function in the obstacle avoidance and tracking confrontation process, the single-step reward value is calculated as the real-time reward for the current maneuver decision. During the movement of the vehicle, when the distance between the two vehicles exceeds a certain range, it will fall into the reward sparse area, generating a large number of invalid samples, resulting in the inability to complete subsequent learning. Therefore, a large penalty should be given to guide the vehicle and the distance between the vehicles to be kept within a reasonable range. For this, a penalty function is defined:

[0115] Formula (7);

[0116] Among them, is an adjustable coefficient, and . , respectively represent the minimum coordinate boundary values that the vehicle can move along the axis and the axis during the vehicle movement, , respectively represent the maximum coordinate boundary values that can be reached along the axis and the axis during the vehicle's movement. The specific numerical values can be set according to the limitations of the actual application scenario.

[0117] According to the comprehensive advantage evaluation function and the penalty function , the reward value of the i-th step of the agent is defined as:

[0118] Formula (8);

[0119] where is the weighted adjustment coefficient of the penalty term.

[0120] Next, according to S3, construct the progressive self-play SAC algorithm framework as shown in Figure 4 . The type, size, and number of the critic neural networks of the progressive self-play SAC algorithm are as follows: The critic neural network is designed to adopt a fully connected neural network structure with two hidden layers, and the number of neurons in each layer is 256. The number of critic neural networks is two or more, and each critic neural network corresponds to a target critic neural network with low-frequency updates.

[0121] Furthermore, the target loss function of the actor neural network of the progressive self-play SAC algorithm includes:

[0122] Based on the excellent exploration ability provided by the policy entropy mechanism, design a loss function that balances exploration ability and policy optimization:

[0123] Formula (1);

[0124] where is the mathematical expectation, is the current state data of the unmanned vehicle, is the experience replay pool, is the action output by the actor neural network. Through the powerful expression ability of the neural network, is modeled as an actor policy network that generates a state-to-specific action mapping, is the parameter of the actor neural network, represents that when the given state is , the probability that the actor policy outputs a certain action is , the policy entropy coefficient, which is initialized to 1, represents the long-term discounted return evaluation of the critic for the current state-action value of the unmanned vehicle.

[0125] As the training process progresses, the policy entropy coefficient It will also be updated according to the different course difficulties. Therefore, it is necessary to design the loss function of the policy entropy coefficient as follows:

[0126] Formula (2);

[0127] Among them, is the target information entropy of the policy. In the early stage of training, as the course difficulty increases and the sparse reward problem intensifies, it is necessary to increase the policy entropy coefficient to enhance the exploration ability. In the later stage, with the progress of self-play optimization and the stable convergence of the policy, it is necessary to decrease the policy entropy coefficient to obtain a stable and reliable policy network.

[0128] Design a training course with increasing difficulty. If the original SAC algorithm is directly used to solve the problems of autonomous confrontation and obstacle avoidance of the vehicle, the following two phenomena will occur in the initial stage of agent training:

[0129] When the vehicle has not yet learned the basic tracking strategy, it is easy to have:

[0130] Formula (9);

[0131] It can be seen from formula (10) that , resulting in the problem of sparse rewards.

[0132] When the vehicle has not yet learned the basic obstacle avoidance strategy and is simultaneously trained for high-difficulty obstacle avoidance and confrontation, at this time the vehicle is extremely likely to collide with obstacles, so there is:

[0133] Formula (10);

[0134] It can be seen from formula (10) that at this time , a large amount of extremely negative reward information will be generated and the training will be interrupted. Moreover, it is difficult to make the final policy robust and generalizable for the learning of a fixed course.

[0135] To address the above problems and mimicking the human learning method from simple to complex, the present invention designs a progressive self-play SAC algorithm, which adopts a learning course with increasing difficulty. The increasing difficulty specifically refers to the increase in scene complexity, such as the increase in the number, size, and randomness of obstacles, and the increase in the randomness, complexity, and pertinence of the opponent's strategy. For specific reference, see Figure 5 for the learning difficulty design.

[0136] Furthermore, considering multi-task scenarios such as tracking, confrontation, and obstacle avoidance comprehensively and adaptively allocating task weights, more complex task decisions closer to the real world can be made. Define the single step (the The overall advantage metric function for step

[0137] Formula (3);

[0138] where represents the performance metric for tracking and confrontation, represents the obstacle avoidance performance metric, and are proportionality coefficients that measure the proportion of different metrics in the overall advantage metric, represents the comprehensive advantage performance metric, which adaptively adjusts the proportionality coefficients in the metric function at different stages of courses with different difficulties. At the same time, this set of metrics can also be used to evaluate and verify strategies.

[0139] Furthermore, repeat step S4 until several learning courses are completed. When the confrontation win rates of both sides in 3 consecutive training courses are close to 50%, it can be proved that the strategy has converged, and the final executor neural network is obtained for real-time decision-making.

[0140] The progressive self-play SAC algorithm adopted in this application can repeatedly perform quantitative evaluation of the existing strategy during training to give cumulative rewards, use Bellman temporal difference learning to update the value function, which is used to guide the evaluation of the policy network and indicate the direction of policy improvement, thereby promoting the optimization and update of the policy network. Then, through self-play, the policy network of the previous round is used as the opponent for iterative self-play learning, and the policy gradually approaches the theoretical optimal policy during the continuous iteration process.

[0141] In this embodiment, the size of the experience replay pool is pre-filled to be approximately 1 million, which can store 1 million state-action-reward-state sequence data for learning.

[0142] In this embodiment, imitating the learning method of humans from simple to complex, the present invention designs a progressive self-play SAC algorithm. This algorithm adopts learning courses with increasing difficulty and divides the training process of the maneuver decision-making progressive self-play SAC algorithm proposed in the previous section into the following 4 courses:

[0143] (1) Fixed tracking learning course: The vehicle adopts a fixed strategy , such as simple linear motion and curvilinear motion, aiming to enable the vehicle to learn basic tracking strategies and be familiar with the confrontation situation.

[0144] (2) Random tracking learning course: The vehicle adopts a random maneuver strategy , aiming to enable the vehicle to learn basic chasing and confrontation maneuver strategies.

[0145] (3) Obstacle Avoidance and Tracking Learning Course: Continuously add obstacles to the environment. The number and size of the obstacles gradually increase with the number of training rounds, and the positions where the obstacles appear also randomly change with the number of training rounds. The purpose of the training is to enable the vehicle to gradually acquire the performance of autonomous obstacle avoidance.

[0146] (4) Iterative Adversarial Learning Course: After the vehicle obtains good adversarial performance each time, transfer the adversarial strategy of the vehicle to the vehicle, and then continue to improve the adversarial performance of the vehicle through the confrontation between the vehicle and the vehicle. The training purpose of this course is to continuously optimize the maneuvering strategy of the vehicle through continuous strategy iteration until satisfactory autonomous adversarial performance is obtained. By means of the above training courses from easy to difficult, not only can the efficiency of reinforcement learning training be greatly improved, but also different training objectives can be set in different learning courses, so that the small vehicle has the ability to generalize from static simple tasks to dynamic complex tasks, ensuring good performance of learning efficiency and the final strategy.

[0147]

[0148] Figure 5 The training scheme block diagram of the progressive self-play SAC algorithm is given. Specifically, first conduct Course 1. In this course, let the vehicle adopt a fixed maneuvering strategy, such as simple uniform linear, accelerating linear, and curvilinear motions, or set a small number of obstacles with fixed sizes. Then, based on the training scheme of the progressive self-play SAC algorithm, conduct simple tracking training on the vehicle. The goal is to let the vehicle be familiar with the battlefield situation and learn basic tracking and obstacle avoidance strategies. Next, enter Course 2. In this stage, gradually increase the number of obstacles while keeping the size unchanged. At the same time, the vehicle adopts a random maneuvering motion strategy, increase the setting range of the initial positions of the two vehicles, and conduct autonomous tracking and obstacle avoidance training on the vehicle based on the existing strategy. The goal is to further improve the autonomous adversarial and obstacle avoidance performance of the vehicle. After obtaining satisfactory performance, enter Course 3. At this stage, the vehicle already has certain autonomous adversarial performance. The training goal is to further improve the autonomous obstacle avoidance performance of the vehicle. For this purpose, randomly and dynamically generate obstacles of the same size during training to increase the training difficulty until satisfactory obstacle avoidance performance is obtained. Finally, enter Course 4. In this stage, use the The learned policy network of the vehicle is used as the maneuvering policy of the vehicle, enabling the vehicle to also have certain tracking, confrontation, and obstacle avoidance capabilities, and then conduct confrontation training between the two vehicles. In subsequent courses, the size, position, and quantity of the obstacles change randomly and dynamically. At the same time, use the learned policy network of the vehicle as the policy selection of the vehicle, and then implement confrontation and obstacle avoidance training. After obtaining satisfactory performance, transfer the learned policy of the vehicle to the vehicle, and then conduct the training and policy optimization of the vehicle. Repeatedly iterate the above process to continuously optimize the policy network of the vehicle.

[0149] Furthermore, comprehensively considering multi-task scenarios such as tracking, confrontation, and obstacle avoidance, adaptively allocate task weights to make more complex task decisions closer to the real world. Define the overall advantage metric function for a single step (the th step) of the unmanned vehicle as:

[0150] ;

[0151] where represents the performance metric for tracking and confrontation, represents the obstacle avoidance performance metric, and are proportionality coefficients that measure the proportion of different metrics in the overall advantage metric, represents the comprehensive advantage performance metric, and the proportionality coefficients in the metric function will be adaptively adjusted at different stages of courses with different difficulties. At the same time, this set of metrics can also be used for policy evaluation and verification.

[0152] Experimental data

[0153] Refer to Figure 6 , in order to actually test the learning efficiency of the proposed progressive self-play SAC algorithm in this paper, the original SAC training scheme was also adopted in the simulation. Figure 6The curves of the cumulative rewards of the non - progressive training scheme (the original SAC algorithm) (the lower curve) and the curriculum progressive training scheme (the progressive self - play SAC algorithm) (the upper curve) with respect to the number of training episodes are given. It can be seen from the figure that due to the relatively simple initial training curriculum, the learning efficiency of the training scheme with the progressive curriculum is significantly higher than that of the non - progressive training scheme in the first 20,000 episodes. After 20,000 episodes, as the difficulty of the curriculum increases, the rising rate of the upper curve decreases, but the overall return is always higher than that of the lower curve, and the lower curve even stops rising. This indicates that the proposed progressive self - play SAC algorithm can not only achieve faster training efficiency but also ensure a better decision - making scheme.

[0154] Reference Figure 7 , 7(a) shows the A trajectory of vehicle B and vehicle A with a random policy, 7(b) shows the B distance change between vehicle A and vehicle B , 7(c) shows the A speed change between vehicle B and vehicle Figure 7 , and 7(d) shows the Figure 7 change of the dominant angle between vehicle Figure 7 and vehicle . For the obstacle - confrontation process shown in (a), Figure 7 (b) and (c) respectively give the distance - change curve and the speed - change curve between vehicle and vehicle

[0155] Reference Figure 8 , Figure 8 shows that after learning through a progressive curriculum with increasing difficulty, when two vehicles start from the same initial state and the same environmental state each time, and vehicle adopts the same random policy, when there are no obstacles and when there are large cylindrical obstacles in the scenario, the confrontation results obtained by vehicle

[0156] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0157] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0158] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.

[0159] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present invention.

[0160] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms should not be understood as necessarily referring to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

Claims

1. A method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning, characterized in that: Including the following steps: S1. According to the kinematic bicycle model, solve and model through the Runge-Kutta method, construct a standard gym environment class in the Python environment, and mathematically represent and describe in computer language the state data of the unmanned vehicle's own state and environmental observations as necessary elements; S2. Through computer simulation of the autonomous decision-making process of multiple unmanned vehicle systems, generate simulation data of the movement process and decision-making behavior of the unmanned vehicle; S3. Design and optimize the form, size, and number of the critic neural network and the actor neural network of the progressive self-play SAC algorithm, construct the target loss function of the actor neural network and the policy entropy coefficient α for the actual movement of the unmanned vehicle, and combine the entropy increase mechanism and the training course difficulty to design the automatic entropy, and design a course learning mechanism that increases with the training process and the strategy type and intensity of the adversarial opponent for the complexity of the unmanned vehicle decision-making scenario to obtain the progressive self-play SAC algorithm; The construction of the target loss function of the actor neural network for the actual movement of the unmanned vehicle includes: Based on the policy entropy mechanism to provide exploration ability, design a loss function that balances exploration ability and policy optimization; The designed loss function J π (φ) satisfies formula (1): where E is the mathematical expectation, s t is the current state data of the unmanned vehicle, D is the experience replay pool, a t is the action output by the executor neural network. Through the powerful expressive ability of the neural network, π φ is modeled as an executor policy network that generates a mapping from states to specific actions. φ is the parameter of the executor neural network, π φ (a t |s t ) represents the probability that when the given state is s t , the executor policy outputs a certain action a t . α is the policy entropy coefficient, initialized to 1. Q(s t , a t ) represents the critic's long-term discounted return evaluation of the current state-action value of the unmanned vehicle; S4. Use the progressive self-play SAC algorithm to self-regulate the training course difficulty and execute the self-play process, generate multiple decision-making data of the unmanned vehicle and put them into the experience replay pool to update different course learning data, average sample the latest data in the experience replay pool and update the parameters of the critic neural network and the actor neural network to complete one learning course; S5. Repeat step S4 to make the critic neural network and the actor neural network complete several learning courses, and obtain the trained actor neural network for generating real-time decisions for the unmanned vehicle to confront and avoid obstacles.

2. The method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning according to claim 1, characterized in that: In S1, the state data of the own state and environmental observations includes one or more of the position coordinates, real-time speed, yaw angle of the unmanned vehicle, and the distance to the obstacle, as well as the position coordinates, real-time speed, and yaw angle of the adversarial opponent.

3. The method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning according to claim 1, characterized in that: The results obtained by designing and optimizing the form, size, and number of the critic neural network and the actor neural network of the progressive self-play SAC algorithm are: The critic neural network adopts a fully connected neural network structure with two hidden layers, and the number of neurons in each layer is 256. The number of critic neural networks is two or more, and each critic neural network corresponds to a target critic neural network with low-frequency updates.

4. The method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning according to claim 1, characterized in that: The construction of the target loss function of the policy entropy coefficient α for the actual movement of the unmanned vehicle includes: Define α as a monotonically increasing function of the difficulty coefficient k. As the training course progresses, the policy entropy coefficient α is updated according to different course difficulties. The designed loss function J(α) of the policy entropy coefficient satisfies formula (2); Among them, is the target information entropy of the policy. In the early stage of training, as the course difficulty increases and the sparse reward problem intensifies, the policy entropy coefficient α will be increased to enhance the exploration ability. In the later stage of training, as the self-play optimization progresses and the policy converges stably, the policy entropy coefficient α will be decreased to obtain a stable and reliable policy network.

5. The method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning according to claim 1, characterized in that: The construction of the standard gym environment class in the Python environment by solving and modeling through the Runge-Kutta method according to the kinematic bicycle model includes: The kinematic bicycle model is based on the kinematics of the unmanned vehicle system. The differential equation of the kinematic bicycle model is solved precisely by the fourth-order Runge-Kutta algorithm to obtain the state observation of the unmanned vehicle at the next moment, and the kinematic bicycle model is encapsulated as a standard class function and unified under the gym framework.

6. The method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning according to claim 1, characterized in that: In S4, the use of the progressive self-play SAC algorithm to self-regulate the difficulty of the training course and perform the self-play process includes: In the later stage of training, an iterative adversarial training strategy is adopted. When the strategy is optimized to a certain stage through the progressive self-play SAC framework, the opponent continuously uses the optimized strategy in the previous round to optimize the strategy in the next round, finds the defects in the previous round of strategy, and makes targeted checks and supplements to the strategy network in an autonomous manner to perform deeper-level strategy optimization.

7. The method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning according to claim 1, characterized in that: After S5, execute: S6. Evaluate the overall advantage of the executor neural network that has completed training through formula (3). Define the overall advantage evaluation function for the i-th step of the unmanned vehicle as: Among them represents the performance indicators of tracking and confrontation represents the obstacle avoidance performance indicator. k1 and k1 are proportionality coefficients, which measure the proportion of different indicators in the overall dominant indicator represents the comprehensive dominant performance indicator, and adaptively adjusts the proportionality coefficient in the overall dominant evaluation function at different stages of courses with different difficulties 8. A method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning according to claim 1, characterized in that: In S4, the parameters of the critic neural network and the executor neural network are updated using the gradient descent algorithm based on historical momentum gradient.

9. A method for unmanned vehicle confrontation and obstacle avoidance based on progressive deep reinforcement learning according to claim 1, characterized in that: In S5, generating real-time decisions for the unmanned vehicle to confront and avoid obstacles includes one or more of the following: static obstacle avoidance, trajectory tracking, comprehensive confrontation, generalization to dynamic obstacle avoidance, and random trajectory tracking.

Citation Information

Patent Citations

  • Air combat maneuvering method based on parallel self-gaming

    CN113095481A