Reinforcement Learning-Based Pursuit Strategy Training Method, Device, Medium and Product

By introducing the homogenization deep deterministic strategy gradient algorithm (MDPG) of integrated value network structure into the reinforcement learning algorithm, the problem of slow learning and poor results in complex pursuit and fugitive game scenarios is solved, and efficient strategy training and improved pursuit performance are achieved.

CN118095340BActive Publication Date: 2025-06-10NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410244720.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-05
Publication Date
2025-06-10
Estimated Expiration
2044-03-05

AI Technical Summary

Technical Problem

Classic reinforcement learning algorithms have problems of slow learning and poor results when facing complex pursuit and escape game scenarios.

Method used

The homogenized deep deterministic strategy gradient algorithm (MDPG) is adopted based on the integrated value network structure. This algorithm introduces an integrated value network structure based on the traditional deep deterministic strategy gradient algorithm (DDPG), and achieves more efficient strategy training through multi-step returns and integrated target value functions.

Benefits of technology

It realizes efficient and autonomous training of pursuit strategies, improves the pursuit performance and success rate of agents, and overcomes the problems of low learning efficiency and poor decision-making performance of traditional algorithms in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118095340B_ABST
    Figure CN118095340B_ABST
Patent Text Reader

Abstract

The present invention discloses a pursuit strategy training method, device, medium and product based on reinforcement learning, which relates to the technical fields of reinforcement learning and pursuit-evasion game control. The method involves a game scenario among an interceptor, a pursuer and a target. The interceptor uses a proportional guidance strategy to intercept the pursuer, while the pursuer uses an averaged deep deterministic policy gradient algorithm based on an integrated value network structure to pursue the target. The MDPG algorithm introduces an integrated value network structure, where each value network corresponds to a target value function and is independently trained using different sample probability distributions. The target uses an escape strategy to avoid being pursued by the pursuer. By different training samples, the distances between agents and the change amount of the heading angle of the pursuer in each pursuit-evasion game scenario are calculated to obtain the return value of the pursuer in each scenario. The MDPG algorithm provided by the present invention can achieve efficient autonomous training of the pursuit strategy, improving the pursuit performance and success rate of the agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of reinforcement learning and pursuit-evasion game control, and particularly to a training method, device, medium and product for a pursuit strategy based on reinforcement learning. Background Art

[0002] As an important topic in the field of control, the pursuit-evasion game is widely applied in various fields such as military and industrial processes. The three-body confrontation is a classic pursuit-evasion game scenario, which includes three parties: a pursuer, an interceptor, and a target. The pursuer needs to approach and capture the target as much as possible while avoiding the interceptor. The interceptor is responsible for intercepting the pursuer, and the target moves away from the pursuer according to its own escape strategy. With the continuous development of reinforcement learning technology, using reinforcement learning algorithms to solve control problems has demonstrated its advantages of model-free dependence, fast response, and good performance. The pursuit-evasion game has become a classic test scenario for reinforcement learning algorithms, and research programs for realizing autonomous pursuit training of agents based on reinforcement learning algorithms and improving the pursuit performance of agents have received extensive attention. However, classical reinforcement learning algorithms still have defects such as slow learning and poor effects when facing complex problems. Summary of the Invention

[0003] The purpose of the present invention is to provide a training method, device, medium and product for a pursuit strategy based on reinforcement learning, which can realize efficient autonomous training of the pursuit strategy and improve the pursuit performance and success rate of the agent.

[0004] To achieve the above purpose, the present invention provides the following solutions:

[0005] In a first aspect, the present invention provides a training method for a pursuit strategy based on reinforcement learning, including:

[0006] Obtaining the simulation environment related parameters of each agent in the pursuit strategy; the agents include a pursuer, an interceptor, and a target; the simulation environment related parameters include the initial coordinates, speed, maximum range, maximum heading angle change amount, and collision judgment distance of the agent.

[0007] Setting the interceptor to intercept the pursuer by using a proportional guidance strategy.

[0008] Setting the pursuer to pursue the target by using an MDPG strategy; the MDPG strategy is an averaged deep deterministic policy gradient algorithm based on an integrated value network structure; the averaged deep deterministic policy gradient algorithm based on the integrated value network structure is an algorithm obtained by introducing an integrated value network structure on the basis of the traditional deep deterministic policy gradient algorithm; the integrated value network structure includes multiple value networks, and each value network corresponds to a target value function; each value network uses a different sample probability distribution and independently extracts training samples for training.

[0009] Set the target to adopt an escape strategy to avoid being pursued by the pursuer.

[0010] Establish a two-dimensional particle model according to the simulation environment related parameters of each agent and the corresponding strategies of each agent.

[0011] Randomly generate multiple training samples; the initial coordinates of each agent in each training sample are different.

[0012] Based on each training sample, calculate the distance between each agent and the change in the heading angle of the pursuer in each pursuit-evasion game scenario, and obtain the return value of the pursuer in each pursuit-evasion game scenario.

[0013] Optionally, the averaged deep deterministic policy gradient algorithm implements temporal difference update based on the averaged integrated objective value function.

[0014] Optionally, for the averaged deep deterministic policy gradient algorithm, based on network random initialization, determine the difference of each value network, and use multi-step return to calculate the objective value function.

[0015] Optionally, using multi-step return to calculate the objective value function specifically includes:

[0016] According to the formula Calculate the objective function.

[0017] where μ′(s|ω μ′ ) is the target + action network, ω μ′ are the parameters of the target action network; N is the number of value networks; s is the state; are the parameters of network Q i ; (s t , a t , r t , s t+1 ) is a given experience sequence.

[0018] Optionally, each value network uses a different sample probability distribution and independently extracts experience samples for training. The loss function for optimizing the value network is:

[0019]

[0020] where is the m-step TD target of the integrated network.

[0021] Optionally, the formula for obtaining the return value of the pursuer in the pursuit-evasion game scenario is as follows:

[0022]

[0023] Among them, A represents the pursuer, D represents the interceptor, and T represents the target; d(k 1 ,k 2 ,t) represents the geometric distance of the agent k 1 ,k 2 ∈{A,D,T} at time t; d norm is a scaling constant used to scale the distance; θ A (t) is the change in the heading angle of the pursuer at time t; W i ,i∈{1,2,3} represents the weights of each item.

[0024] Optionally, the escape strategy includes a fixed-position escape strategy and an escape strategy away from the pursuer.

[0025] In a second aspect, the present invention provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the method for training a pursuit strategy based on reinforcement learning described in the first aspect.

[0026] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for training a pursuit strategy based on reinforcement learning described in the first aspect are implemented.

[0027] In a fourth aspect, the present invention provides a computer program product, including a computer program, and when the computer program / instruction is executed by a processor, the steps of the method for training a pursuit strategy based on reinforcement learning described in the first aspect are implemented.

[0028] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0029] The present invention discloses a pursuit strategy training method, device, medium and product based on reinforcement learning. The method includes: obtaining the simulation environment parameters of the pursuer, interceptor and target in the pursuit strategy, such as the initial coordinates, speed, maximum range, maximum heading angle change amount and collision judgment distance. The interceptor uses a proportional guidance strategy to intercept the pursuer, while the pursuer uses the averaged deep deterministic policy gradient algorithm (MDPG) based on the integrated value network structure to pursue the target. The MDPG algorithm introduces an integrated value network structure, where each value network corresponds to a target value function and is independently trained using different sample probability distributions. The target uses an escape strategy to avoid the pursuer. Through the simulation environment parameters and corresponding strategies of each agent, a two-dimensional particle model is established, and multiple training samples are randomly generated, where the initial coordinates of the agents in each sample are different. Based on these training samples, the distances between the agents and the heading angle change amount of the pursuer in each pursuit-evasion game scenario are calculated to obtain the reward value of the pursuer in each scenario. Based on the traditional DDPG algorithm, the present invention introduces an integrated network structure and proposes the averaged deep deterministic policy gradient algorithm MDPG based on the integrated value network structure. The MDPG algorithm can achieve efficient autonomous training of the pursuit strategy, improving the pursuit performance and success rate of the agent. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0031] Figure 1 It is a flowchart of a pursuit strategy training method based on reinforcement learning provided in Embodiment 1 of the present invention;

[0032] Figure 2 It is an interaction and training flowchart of the MDPG algorithm in the pursuit-evasion scenario provided in Embodiment 1 of the present invention;

[0033] Figure 3 It is the pseudo-code of the MDPG algorithm provided in Embodiment 1 of the present invention;

[0034] Figure 4 It is an example of the fixed-mode pursuit simulation trajectory provided in Embodiment 1 of the present invention;

[0035] Figure 5 It is an example of the straight-line-mode pursuit simulation trajectory provided in Embodiment 1 of the present invention;

[0036] Figure 6 It is an example of the escape-mode pursuit simulation trajectory provided in Embodiment 1 of the present invention;

[0037] Figure 7 The change curve of the average cumulative return of the model in the fixed mode with the number of evaluations in the test set scenario provided by Embodiment 1 of the present invention;

[0038] Figure 8 The change curve of the average cumulative return of the model in the linear mode with the number of evaluations in the test set scenario provided by Embodiment 1 of the present invention;

[0039] Figure 9 The change curve of the average cumulative return of the model in the escape mode with the number of evaluations in the test set scenario provided by Embodiment 1 of the present invention;

[0040] Figure 10 The internal structure diagram of the computer device provided by the embodiment of the present invention. Detailed implementation manners

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0042] The purpose of the present invention is to provide a pursuit strategy training method, device, medium and product based on reinforcement learning, which can realize efficient autonomous training of the pursuit strategy and improve the pursuit performance and success rate of the intelligent agent.

[0043] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0044] Embodiment 1

[0045] As Figure 1 shown, this embodiment provides a pursuit strategy training method based on reinforcement learning, including:

[0046] Step 101: Obtain the simulation environment related parameters of each intelligent agent in the pursuit strategy; the intelligent agents include pursuers, interceptors and targets; the simulation environment related parameters include the initial coordinates, speeds, maximum voyage distances, maximum heading angle change amounts and collision judgment distances of the intelligent agents.

[0047] Step 102: Set the interceptor to intercept the pursuer using the proportional guidance strategy.

[0048] Step 103: Set the pursuer to adopt the MDPG strategy to pursue the target; the MDPG strategy is the averaged deep deterministic policy gradient algorithm based on the integrated value network structure; the averaged deep deterministic policy gradient algorithm based on the integrated value network structure is an algorithm obtained by introducing the integrated value network structure on the basis of the traditional deep deterministic policy gradient algorithm; the integrated value network structure includes multiple value networks, and each value network corresponds to a target value function; each value network uses different sample probability distributions and independently extracts training samples for training.

[0049] Step 104: Set the target to adopt an escape strategy to avoid the pursuit of the pursuer.

[0050] Step 105: Establish a two-dimensional particle model according to the simulation environment related parameters of each intelligent agent and the strategies corresponding to each intelligent agent.

[0051] Step 106: Randomly generate multiple training samples; the initial coordinates of each intelligent agent in each training sample are different.

[0052] Step 107: Based on each training sample, calculate the distances between the intelligent agents and the heading angle change of the pursuer in each pursuit-evasion game scenario, and obtain the return value of the pursuer in each pursuit-evasion game scenario.

[0053] Among them, the pursuit strategy training method based on reinforcement learning solves the Markov decision process in the form of (S, A, R, P, γ), where S is the state space, which is a set of certain characteristic expressions s t of the environmental state information at any time; A is the action space, which is a set of actionable actions a t at any time; r t = R(s t , a t ) is the reward function; P is the state transition probability, which determines the probability Pr(s t | s t ) of transferring to the state s t+1 when the action a t+1 is executed in the state s t ; γ ∈ (0, 1] is the discount factor, which is used to balance the immediate reward and the future reward. Each intelligent agent makes a decision based on the current state s t , executes the action a t , obtains the next state s t and the reward r t+1 after interacting with the environment, and then conducts the interaction at the next moment, iterating in this way until the termination condition is met. t

[0054] In this embodiment, the goal of reinforcement learning is to maximize the discounted cumulative return under the current policy u. The calculation formula of the discounted cumulative return is shown in Equation (1):

[0055]

[0056] In the formula, r t is the immediate return obtained by taking action a t in state s t ; γ ∈ (0, 1] is the discount factor, which is used to balance the immediate return and the future return; G t is the discounted cumulative return.

[0057] In this embodiment, the Mean Deep Deterministic Policy Gradient (MDPG) algorithm based on the integrated network is used to realize the autonomous training of the above pursuit strategy.

[0058] Specifically, the Mean Deep Deterministic Policy Gradient algorithm based on the integrated value network structure is an algorithm obtained by introducing the integrated value network structure on the basis of the traditional Deep Deterministic Policy Gradient algorithm. The Deep Deterministic Policy Gradient (DDPG) algorithm is a reinforcement learning algorithm based on the Actor-Critic framework proposed by Timothy in 2015. Among them, Actor, that is, the action network μ(s|ω μ ), is a neural network with s as the input and ω μ as the parameter, which is used to output actions; Critic, that is, the value network Q(s, a|ω Q ), is a neural network with (s, a) as the input and ω Q as the parameter, which is used to estimate the state value function under the current policy. The state value function is shown in Equation (2):

[0059]

[0060] Among them, represents the expectation of a random variable under the policy μ, and Q μ (s t , a t ) is the state value function.

[0061] DDPG introduces target networks μ′(s|ω μ′ ) and Q′(s, a|ω Q′ ) with the same structure as the action network and the value network but different parameters, and uses the Time Difference (TD) method to realize the bootstrap update of the value network. The Time Difference method is shown in Equation (3):

[0062] Q μ (s t ,a t ) ← R t +γQ′ μ (s t+1 ,μ′(s t+1 )) (3).

[0063] DDPG uses an experience replay buffer to store the experience sequence (s t ,a t ,r t ,s t+1 ) generated during the interaction. The capacity of the buffer is denoted as When updating the network parameters using gradient descent, a batch of experience sequences is randomly sampled from the experience replay buffer as training samples to calculate the loss function, where is the experience batch size. The loss function for optimizing the value network is shown in Equation (4):

[0064]

[0065] where, q i is the TD target, and its calculation formula is shown in Equation (5):

[0066] q i = r i +γQ′(s i+1 ,μ′(s i+1 |ω μ′ )|ω Q′ ) (5).

[0067] In the formula, (s i ,a i ,r i ,s i+1 ) is an experience sequence, representing the current state, current action, immediate reward, and successor state respectively. Q′ and μ′ represent the target value network and target action network respectively. DDPG introduces target networks μ′(s|ω μ′ ) and Q′(s,a|ω Q′ ) with the same structure as the action network and value network but different parameters

[0068] The goal of the action network is to maximize the return of the policy. Therefore, the loss function for optimizing the action network is shown in Equation (6):

[0069]

[0070] In the formula, is the loss function of the action network, Is a training sample drawn from the experience pool.

[0071] Among them, the parameter optimization of the value network and the action network can be achieved by performing gradient backpropagation on the losses shown in formulas (4) and (6).

[0072] The target network implements delayed update based on the soft update method, where the update method is shown in Equation (7):

[0073]

[0074] In the formula, τ∈(0,1) is the soft update rate.

[0075] According to DDPG in the above, DDPG provides a technical framework for solving continuous-action Markov processes. However, in complex environments, there are defects such as low learning efficiency and poor decision-making performance. The core problem lies in the low sample efficiency and the inability to accurately estimate the value function.

[0076] Therefore, when improving the DDPG algorithm in this embodiment, the DDPG algorithm is improved from the following aspects.

[0077] (1) Use multi-step returns

[0078] MDPG uses multi-step TD to update the value function. When calculating the TD target, the state after m steps is used as the successor state to calculate the target value, and the immediate returns of the remaining m - 1 steps are used to calculate the corresponding discounted cumulative return. For the experience sequence (s t , a t , r t , s t+1 ), the formula for calculating the m-step TD target is shown in Equation (8):

[0079]

[0080] When m = 1, Equation (8) is equivalent to the single-step TD adopted by DDPG. A too large value of m will increase the variance of the estimation error and cause unstable network updates. Appropriately selecting the value of m can enable the network to approximate the true value quickly and stably.

[0081] (2) Introduce an integrated network structure

[0082] MDPG integrates N parallel value networks And equips each value network with a corresponding target network When updating the value function, MDPG uses the mean of the N target networks to calculate the target value function of the successor state. For the experience sequence (s t , a t , r t , st+1 ) The formula for the m-step TD target of the MDPG calculation integration network is shown in Equation (9):

[0083]

[0084] Calculating the target value function based on the integration network can reduce the variance of the estimation bias, improve the network estimation ability, and thus improve the performance of the reinforcement learning decision-making.

[0085] (3) Separately extract network training samples

[0086] The estimation ability of the integration network is positively correlated with the estimation ability of each network and negatively correlated with the correlation between networks. To improve the diversity of the integration network, MDPG randomly initializes the N value networks, so that the networks have different parameters at the beginning of training. During the training process, MDPG performs a sampling update for each value network, so the empirical samples used for updating each network are different, enhancing the network diversity.

[0087] For the i-th value network Extract a batch of empirical sequences from the empirical cache pool As training samples, update the network Q i The loss function of is shown in Equation (10):

[0088]

[0089] In the formula, is the m-step TD target of the integration network.

[0090] Then, based on the samples Update the loss function of the action network, and the loss function is shown in Equation (11):

[0091]

[0092] Perform gradient backpropagation on the loss function, and the gradient descent update operations of the integration network parameters μ and are shown in Equation (12):

[0093]

[0094] Among them, α u and α Q are the learning rates of the action network and the value network, respectively.

[0095] Integrate the target network parameters μ′ and Implement delayed update based on the soft update method, and the update method is shown in Equation (13):

[0096]

[0097] Updating each value network requires recalculating the integrated TD target, so the time complexity of the MDPG algorithm is O(N 2 ). When N = 1, the network structure of MDPG is equivalent to that of DDPG. A too large value of N will increase the computing power cost. Selecting an appropriate number of networks can improve the network estimation ability while ensuring the training speed.

[0098] (4) Use prioritized experience replay.

[0099] DDPG uses a uniform distribution for sample extraction. MDPG introduces prioritized experience replay, and uses the absolute TD error of the integrated network to characterize the weights of the sequences in the experience pool. Sequences with higher weights, that is, sequences with larger estimation biases, are more likely to be extracted as training samples.

[0100] To ensure the diversity of the integrated network, MDPG equips each value network with an independent sampling probability distribution p i (·|D). The probability that the sequence τ j =(s j ,a j ,r j ,s j+1 ) is extracted when updating the network Qi is shown in Equation (14):

[0101]

[0102] In the formula, is the absolute TD error of the sequence τ j under the network Q i , and the calculation formula of is shown in Equation (15):

[0103]

[0104] The extracted sequence is used as a sample, and the network loss function is calculated through Equations (10) and (11), and then gradient backpropagation is performed to update the network parameters.

[0105] Among them, the interaction and training process of the MDPG algorithm in the pursuit-evasion scenario is as Figure 2 shown, and the algorithm pseudocode is as Figure 3 shown.

[0106] This example demonstrates the interaction and training process of the MDPG algorithm in the pursuit-evasion scenario. Taking the three-body confrontation scenario as an example, it details how to use the MDPG algorithm to train the pursuit strategy. Specifically as follows:

[0107] This example applies the MDPG algorithm to the training of pursuit strategies in a three-body confrontation scenario. The classic three-body confrontation scenario includes a pursuer A, an interceptor D, and a target T. The interceptor adopts a proportional guidance strategy to intercept the pursuer. The target selects a fixed position or an escape strategy away from the pursuer according to its own maneuvering mode. The pursuer is controlled by a deterministic strategy u. The pursuer A, the interceptor D, and the target T are modeled as two-dimensional particle models. x k (t) and y k (t) represent the horizontal and vertical coordinates of the agent k ∈ {A, D, T} at time t, respectively; θ k (t) ∈ [-π, π] represents the heading angle of the agent at time t (the angle between the velocity direction and the horizontal coordinate axis); v k is a constant value, representing the speed of the agent; Δθ k (t) is used as a control variable, representing the change in the heading angle of the agent at time t. To sum up, the discrete motion model of the agent is shown in Equation (16):

[0108]

[0109] Among them, the voyage of the agent k ∈ {A, D, T} at time t is shown in Equation (17):

[0110]

[0111] Set the maximum voyage of the agent to which measures the farthest distance that the agent can move before running out of energy. When the current voyage of the agent exceeds , the agent stops moving, which represents the maximum voyage of the agent. Set The maximum voyage of the pursuer is greater than that of the interceptor. The pursuer can first take evasive actions and then capture the target after the interceptor stops moving.

[0112] The reward function R determines the optimization direction of the decision. In this embodiment, according to the characteristics of the pursuit problem, a reward function is designed as shown in Equation (18):

[0113]

[0114] In the formula, represents the geometric distance of the agent k 1 , k 2 ∈ {A, D, T} at time t; d norm is a scaling constant; W i, where \(i\in\{1,2,3\}\) represents the weights of each item. During the pursuit process, the reward value is calculated based on the distances of each agent and the change in the heading angle of the pursuer. This value is positively correlated with the distance between the pursuer and the interceptor, and negatively correlated with the distance between the pursuer and the target and the change in the pursuer's heading angle; when the pursuit ends, a default fixed reward is given according to the game result - if the pursuer successfully captures the target, the default reward is \(R \) t = 300; if the pursuer is intercepted or fails to capture the target before the voyage is exhausted, the default reward is \(R \) t = -100. When the geometric distance between the two agents is less than the collision judgment distance \(d \) collide , it is determined that the two agents are destroyed.

[0115] The pursuer is controlled by the MDPG policy network \(\mu(s \) t ), which takes the state \(s \) t at time \(t\) as the input and outputs the action \(a \) t of the pursuer at time \(t\), where \(a\in[-1,1]\). During the training process, Gaussian exploration noise is added to the action \(a \) t , that is and the exploration noise is removed during testing. The action \(a_t\) is linearly mapped to the change in the heading angle \(\Delta\theta \) A (t) of the pursuer. The mapping formula is shown in formula (19):

[0116]

[0117] The state \(s \) t contains the motion information of all agents at this moment, which is expressed as shown in formula (20):

[0118] s t = \{x A (t), y A (t), \(\theta A (t), x D (t), y D (t), \(\theta D (t), x T (t), y T (t), \(\theta T (t)\} (20).

[0119] In the formula, each value is normalized. After \(\Delta\theta A (t) acts on the environmental model shown in formula (16), the state \(s t+1 and the reward \(r t at the next moment are obtained. The generated sequence \((s t , a t , r t , s t+1 ) is saved to the experience replay buffer and sample participation in network update based on the prioritized experience replay mechanism during the training process.

[0120] In this embodiment, the pursuit strategy is trained based on a randomly initialized simulation environment, and three maneuver modes (i.e., escape strategies) are designed for the target, including: fixed mode, straight-line mode, and escape mode. In each initial scenario, the initial coordinates of the pursuer are (15, 15).

[0121] In the fixed mode, the initial coordinates of the target are (75, 75), and the target remains stationary during the capture process.

[0122] In the straight-line mode, the initial coordinates of the target are (60, 60), and the target moves along the extension line direction from the initial coordinates of the pursuer to its own coordinates.

[0123] In the escape mode, the initial coordinates of the target are (60, 60), and the target moves along the extension line direction from the current coordinates of the pursuer to its own coordinates.

[0124] The horizontal and vertical coordinates of the interceptor are randomly selected within (0, 100).

[0125] The pursuit and escape simulation trajectories of the agent in the fixed mode, straight-line mode, and escape mode are respectively as Figure 4 , Figure 5 and Figure 6 shown.

[0126] In this embodiment, the relevant parameters of the simulation environment are set as shown in Table 1, and the hyperparameters of the training algorithm are set as shown in Table 2.

[0127] Table 1 Setting of relevant parameters of the simulation environment

[0128]

[0129] Table 2 Setting of hyperparameters of the training algorithm

[0130]

[0131] For each target maneuver mode, 50 scenarios are randomly generated as the test set. After the model is trained every 1000 steps, it is evaluated based on the scenarios in the test set.

[0132] In this embodiment, the classical DDPG algorithm and its improved variant TD3 are used as the comparison models of MDPG. All three algorithms use a fully connected network with three layers and 256 nodes as the basic structure of the action network and the value network. The average value of the data obtained from the last 50 evaluations is used as the final performance of the model. The average cumulative rewards (ACR) and the success rates (SR) of the DDPG, TD3, and MDPG output models in the test set scenarios are shown in Table 3.

[0133] Table 3 Final Performance of the Model

[0134]

[0135]

[0136] The change curve of the average cumulative rewards in the test set scenarios with the number of evaluations is as Figure 6 , Figure 7 and Figure 8 shown (the curve is the result after mean filtering with a window size of 20).

[0137] Comprehensively Figure 7 , Figure 8 , Figure 9 and the results shown in Table 3, it can be seen that the learning efficiency of the MDPG algorithm and the performance of the final output model are both better than those of DDPG and TD3, indicating that an integrated network of a certain scale can effectively improve the network estimation ability and thus improve the model decision-making level.

[0138] To further verify the influence of the integrated network scale on the model performance, in this embodiment, the performance of the MDPG algorithm output model and the average training time (ATT) are tested when the number N of integrated networks is 5, 10, 15, and 20 respectively. The simulation results are shown in Table 4.

[0139] As can be seen from Table 4, increasing the integrated network scale does not necessarily bring a significant improvement in model performance while increasing the computing power cost. When solving practical problems, it is necessary to reasonably select the integration parameters in combination with the task background and hardware conditions. In this embodiment, considering the computing power cost and model performance comprehensively, the optimal parameter for the number of integrated networks of MDPG is N = 5.

[0140] Table 4 Final Performance and Training Duration of the Model under Different Integrated Network Scales

[0141]

[0142]

[0143] In one embodiment, a computer device is provided. The computer device may be a database, and its internal structure diagram may be as shown in Figure 10 . The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store transactions to be processed. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a data processing method.

[0144] In one embodiment, a computer device is further provided, including a memory and a processor to store a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the above method embodiments.

[0145] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0146] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps in the above method embodiments.

[0147] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the object or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0148] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0149] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0150] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A hunting strategy training method based on reinforcement learning, characterized in that: include: Acquire simulation environment related parameters of each intelligent agent in the pursuit strategy; the intelligent agents include pursuers, interceptors and targets; the simulation environment related parameters include initial coordinates, speed, maximum range, maximum heading angle change and collision judgment distance of the intelligent agent; Setting the interceptor to adopt a proportional guidance strategy to intercept the pursuer; The pursuer is set to use the MDPG strategy to pursue the target; the MDPG strategy is an averaged deep deterministic policy gradient algorithm based on an integrated value network structure; the averaged deep deterministic policy gradient algorithm based on an integrated value network structure is an algorithm obtained by introducing an integrated value network structure on the basis of a traditional deep deterministic policy gradient algorithm; the integrated value network structure includes a plurality of value networks, each of which corresponds to a target value function; each of the value networks uses a different sample probability distribution and independently extracts training samples for training; Setting the target to adopt an escape strategy to avoid being pursued by the pursuer; Establishing a two-dimensional particle model according to simulation environment related parameters of each intelligent agent and the corresponding strategy of each intelligent agent; Randomly generate multiple training samples; the initial coordinates of each of the intelligent agents in each of the training samples are different; Based on the training samples, the distance between the agents and the change in the course angle of the pursuer in each pursuit-escape game scenario are calculated to obtain the reward value of the pursuer in each pursuit-escape game scenario; The formula for calculating the reward value of the pursuer in the pursuit-escape game scenario is as follows: Where A refers to the pursuer, D refers to the interceptor, and T refers to the target; d(k1,k2,t) represents the geometric distance between agents k1,k2∈{A,D,T} at time t; d norm is the scaling constant used to scale the distance; θ A (t) is the change in the pursuer’s heading angle at time t; W i ,i∈{1,2,3} represents the weight of each item; The averaged deep deterministic policy gradient algorithm implements time difference update based on the averaged integrated objective value function; The averaged deep deterministic policy gradient algorithm determines the difference of each value network based on random network initialization and uses multi-step returns to calculate the target value function; Among them, the target value function is calculated using multi-step returns, including: According to the formula Calculate the objective function; Among them, μ′(s|ω μ′ ) is the target + action network, ω μ′ is the parameter of the target action network; N is the number of value networks; s is the state; It's Network Q i Parameters; (s t ,a t ,r t ,s t+1 ) is a given experience sequence, γ∈(0,1) is a discount factor, Q′ i Indicates the target network corresponding to the value network; Each value network uses a different sample probability distribution and independently extracts experience samples for training. The loss function of the value network optimized during training is: in, is the m-step TD target of the integrated network, |B| is the experience batch capacity pool and B={(s j , a j , r j ,s j+1 ), j∈(1, 2, ...|B|)}; The escape strategies include a fixed position escape strategy and a strategy of escaping from the pursuer.

Citation Information

Patent Citations

  • Power distribution network reactive power optimization method based on depth deterministic strategy gradient algorithm

    CN116207750A

  • Multi-agent collaborative pursuit confrontation method based on P3C-MADDPG algorithm

    CN117131770A