MADDPG multi-agent motion control method mixed with three types of experience
By introducing inferior and excellent experience into the MADDPG algorithm, three types of experience hybrid mechanisms are generated, and the adaptability and stability of multi-agent motion planning in a dynamic environment is solved, and efficient motion control effect is achieved.
Patent Information
- Application Number
- CN202510429690.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Traditional multiagent motion planning methods lack adaptability and flexibility in dynamic and uncertain environments, MADDPG algorithm training efficiency is low and the experience pool quality is poor, making it difficult to achieve stable and efficient motion control in complex environments.
Introduce inferior experience and excellent experience, generate inferior experience through counterexample modules, use HER modules to generate excellent experience, build a three-type experience hybrid mechanism, improve the MADDPG algorithm, enhance training efficiency and stability, and improve sample diversity and quality.
Significantly improves the adaptability and robustness of multi-agent systems in complex dynamic environments, ensuring efficient and reliable motion control in uncertain environments.
Smart Images

Figure CN120354876A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-agent deep reinforcement learning, and particularly to a MADDPG multi-agent motion control method. Background Art
[0002] With the continuous progress of technology, the cooperative motion control technology of multi-agent systems has been widely applied in many fields. Especially in tasks such as emergency rescue and military operations, agents and intelligent vehicles have successfully executed complex tasks through efficient path planning and cooperative strategies. In the complex and changeable battlefield environment, the autonomous navigation technology of multi-agents has become a research hotspot. The core goal of the research is to plan a path for each agent that can avoid collisions and be completed within the task time constraint. However, traditional multi-agent motion planning methods mostly rely on the condition that the environment is completely known and relatively fixed, and use search algorithms such as the A* algorithm, Artificial Potential Field (APF), and Vector Field Histogram Plus (VFH+) to calculate paths. Although they perform excellently in static and stable environments, when the environment changes dynamically, their adaptability and flexibility are insufficient, and it is difficult to cope with the rapid changes and uncertainties in the battlefield environment.
[0003] Multi-agent reinforcement learning involves multiple agents interacting in the same environment to coordinate their behaviors to learn optimal strategies. For example, the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm uses a framework with an Actor-Critic structure to solve the motion planning problem in multi-agent systems. The core idea of the MADDPG algorithm is centralized training and decentralized execution. Each agent independently generates actions based on its own state information and uses the joint state information of all agents to estimate the value of the actions. This method can theoretically achieve efficient decision-making and good coordination. However, the MADDPG algorithm also faces some challenges. First, in the initial stage of training, agents rely on random exploration to learn the environment, which is not only inefficient but also leads to a low training speed of the neural network and poor quality of the experience pool due to the need to avoid collisions simultaneously. Second, most samples are concentrated in specific states, lacking high-quality positive feedback and extremely poor negative experiences, which will affect the learning effect of the strategy. Third, the Actor-Critic network is updated after each iteration, which leads to frequent changes in the policy direction, making it difficult for agents to stably learn and adapt to the environment.
[0004] To address these issues, some researchers have proposed methods to optimize the MADDPG algorithm. For example, the published patent CN113341958A mentions a method to optimize MADDPG through a mixed experience pool, namely the ME-MADDPG algorithm. This method improves the efficiency and stability of training by introducing high-quality experiences. In this patent, in addition to the exploration experiences of the agents themselves, expert experiences generated by the artificial potential field method and improved experiences after post-processing are added to increase sample diversity and improve learning efficiency. However, although the method proposed in CN113341958A improves the learning efficiency to a certain extent, it still has some limitations and problems. First, although the method in the patent increases high-quality experiences, it does not handle the low-quality experiences in the experience pool well, which may cause the agents to be less sensitive to some important negative situations during training. Second, when generating expert experiences, the method may rely too much on the artificial potential field method, which may not be flexible or adaptable enough in some highly dynamic and unstructured environments. In addition, the method may be too fixed when dealing with the mixing ratio of exploration experiences and high-quality experiences, lacking the ability to dynamically adjust to environmental changes, which may limit the performance of the algorithm in a changing environment.
[0005] The present invention proposes a multi-agent motion control method of MADDPG (Three Experience Mixed Multi-Agent Deep Deterministic Policy Gradient, 3E-MADDPG) that mixes three types of experiences. By improving the MADDPG algorithm, it solves the multi-agent motion control problem and enhances the stability and decision-making ability of the multi-agent system in complex dynamic environments. Summary of the Invention
[0006] To solve the above problems, the present invention proposes a multi-agent motion control method of MADDPG that mixes three types of experiences, belonging to the field of multi-agent deep reinforcement learning. The MADDPG algorithm introduces low-quality experiences and excellent experiences on the basis of expert experiences; the low-quality experiences are generated by the counterexample module, and the excellent experiences are generated by the Hindsight Experience Replay (HER) module; the present invention improves the MADDPG algorithm by integrating three types of experiences, namely expert experiences, low-quality experiences, and excellent experiences, and constructs a mixing mechanism for the three types of experiences, thereby improving the training efficiency and stability of the MADDPG algorithm; enhancing the diversity and quality of training samples, and enhancing the stability and decision-making ability of the multi-agent system in complex dynamic environments; by integrating the counterexample module to generate low-quality experiences, it can perform efficient and reliable motion control in uncertain environments, significantly improving the adaptability and robustness of the multi-agent system.
[0007] Such asFigure 1 As shown in:
[0008] The present invention provides a MADDPG multi-agent motion control method that combines three types of experiences, including the following steps:
[0009] Step 1: Set the environmental model and the agent motion model of the agents;
[0010] Step 2: Construct a multi-agent joint state model;
[0011] Step 3: Construct the Actor-Critic network of the MADDPG algorithm, and initialize the parameters of the Actor-Critic network and the experience pool;
[0012] Step 4: Determine the way the agents select actions;
[0013] Step 5: Update the agent states;
[0014] Step 6: Use the HER algorithm to generate excellent experiences, and use the counterexample module to generate poor experiences;
[0015] Step 7: Determine whether the multi-agent enters the training stage;
[0016] Step 8: Set the sampling strategy and experience replay.
[0017] As Figure 2 shown in:
[0018] Furthermore, in Step 1, the environmental model is set as follows:
[0019] Each agent in the multi-agent system is a circular body, and the radius of the agent is r i ; A series of dynamic obstacles are included in the external environment of the multi-agent system, and the agents need to avoid circular obstacles with a radius of r obs during flight to safely reach the target position; The collision distance D io between the agent and the obstacle is = r i + r obs ; When the Euclidean distance d between the agent and the obstacle satisfies d ≤ D io , it means that the agent has collided;
[0020] When the obstacle is in a moving state, the obstacle motion model is the same as the agent motion model.
[0021] The agent motion model is set as follows:
[0022] The agent motion model simulates the motion of an agent in a two-dimensional plane, flying at a constant altitude and turning with inertial coordination; the agent obtains its own data through an inertial measurement device, and the motion state of the agent at each time step t is updated by the acceleration velocity position heading angle and angular velocity determine:
[0023] The angular velocity is calculated as follows:
[0024]
[0025] where, is the rate of change of the heading angle at time t of the agent, that is, the angular velocity; is the velocity in the x direction; is the velocity in the y direction; the heading angle ranges from [0, 2π];
[0026] The agent obtains the acceleration and velocity
[0027] through the inertial measurement device, and obtains the current position
[0028] according to the navigation data. The acceleration is determined by the resultant force acting on the agent and the mass m u of the agent. The acceleration is expressed as:
[0029]
[0030] and respectively represent the accelerations in the x and y directions; and respectively represent the resultant forces in the x and y directions;
[0031] The velocity of the agent at time t is determined by the velocity at the previous time step t - 1 and the acceleration ; similarly, the position information is also determined by the position information at the previous time step and the current velocity. The velocity and information The formula is as follows:
[0032]
[0033] and respectively represent the positions in the x-direction and y-direction; δt represents the time interval from the previous moment t - 1 to the current moment t;
[0034] To ensure the physicality of the agent system, the maximum acceleration a max and the maximum speed v max of the agent are set to avoid the motion parameters of the agent exceeding the actual constraints;
[0035] Furthermore, in step 2, the process of constructing the multi-agent joint state model is as follows:
[0036] Step 2.1, using the Markov decision model for multi-agent path planning:
[0037] The Markov decision model is described by the five-tuple <N, S, A, C, R>, where N = {1, 2,..., n}, representing a set of n agents; n represents the number of agents; S = s1 × s2 × s i ×...× s n , S represents the joint state space, s i is the state of agent i, and S is composed of the state combinations of all agents, because the current state of an agent is not only related to its previous state but also closely related to the states of other agents; A = A1 × A2 × A i ×...× A n , A represents the joint action space, A i is the action of agent i; each agent selects an action based on its current state and the cooperation or competition information from other agents; C is the state transition model, and the state transition model is: S × A × S → [0, 1], indicating the probability of the multi-agent transferring to the next joint state after taking a joint action in the current joint state; R = r1 × r2 × r i ×...× r n , R is the joint reward function, r i represents the reward value obtained by agent i under a specific state-action pair;
[0038] In the Markov decision model, the optimal policy J(θ i ) of the agent is to maximize the joint cumulative return of the multi-agent system, and J(θ i ) is achieved through the expected return:
[0039]
[0040] Among them, θ i is the policy parameter of agent i; G i is the expected cumulative return of agent i under the current policy; γ t is the discount factor at time t, and the discount factor represents the importance that the agent attaches to future rewards; is the immediate reward of agent i at time t;
[0041] Step 2.2, multi-agent joint state space construction:
[0042] The sensors of each agent can detect detection information in 7 different directions to obtain the detected distance value; when detecting other agents or obstacles, the sensor returns the straight-line distance between the agent and the detected object; if no object is detected, the sensor returns the default value of the maximum detection distance; through the sensor, the agent can obtain multi-dimensional information about itself, the environment, and the target point, so as to make decisions in complex task scenarios;
[0043] Agent state information S i includes its own state information s u , sensor detection information s r and environmental information s e ; that is, S i =(s u , s r , s e );
[0044] s u =[v i , ψ i Equation (6)
[0045]
[0046] Among them, the own state information s u includes the current speed v i and the heading angle ψ i of the agent; the sensor detection information s r is composed of the detection information returned by the sensor, and respectively represent the distance detection values in 7 directions detected by the sensor; the target position P ig =(P igx , P igy ), P igx and P igy respectively represent the positions of P ig in the x direction and the y direction; the relative position between agent i and the target The relative distance between agent i and other agents Represents the heading angle difference between agent i and the target; Represents the heading angle difference between agent i and other agents, s e The range of is [0, π];
[0047] Normalize the state information S of the agent i and map it to the range of [0, 1]; S i The normalization process is as follows:
[0048] First, the normalization of the self-state information s u is:
[0049]
[0050] where, v i represents the speed of agent i; ψ i represents the heading angle of agent i; π represents the set of strategies selected by all agents during the entire movement process;
[0051] Second, the normalization of the sensor detection information s r is:
[0052]
[0053] where, d R represents the maximum detection distance of the sensor of agent i;
[0054] Finally, the normalization of the environmental information s e is:
[0055]
[0056] where, L max is the boundary length of the environment;
[0057] The final agent state space S i is a 15-dimensional vector that contains all the information related to the multi-agent cooperation and competition tasks, ensuring that the agent can effectively learn and adapt to the changes in different task scenarios;
[0058] Step 2.3, construction of the multi-agent joint action space:
[0059] By controlling the acceleration and angular velocity of the agent, the motion posture and flight speed of the agent can be controlled; the agent action space A i is:
[0060] A i = [a i , ω i Equation (12)
[0061] Among them, a i is the acceleration of agent i; ω i is the angular velocity of agent i, and a i ∈[-a max , a max , ω i ∈-ω max , ω max , a max represents the maximum acceleration of agent i, and ω max represents the maximum angular velocity of agent i;
[0062] To ensure that the agent moves under physical constraints, the velocity v i and the heading angle ψ i of agent i both need to satisfy v i ∈[0, v max , ψ i ∈[0, ψ max constraints, v max represents the maximum velocity of the agent, and ψ max represents the maximum heading angle of the agent;
[0063] The action space A i describes the motion adjustment ability of the agent at any time. Due to the complex cooperation and competition relationships among agents, the joint action space A = A1×A2×A i ×…×A n is crucial for the overall motion control of the agent system;
[0064] Step 2.4, multi-agent objective function construction:
[0065] The goal of path planning is to ensure that the agent reaches the target position in the shortest time and path without collision, so as to achieve the optimal motion strategy π * :
[0066]
[0067] P i = P ig Equation (14)
[0068]
[0069] Among them, P i represents the position of agent i; P ig represents the target position; T i represents the time for agent i to reach the target position; L i represents the total path length for agent i to reach the target position; the optimal motion strategy π* , that is, the strategy with the minimum total time for all agents to reach the target and the total path length; D io is the collision distance; It means that at any moment, agent i cannot collide with obstacles or other agents; j represents obstacles or other agents, that is, adjacent objects; represents the position of agent i at time t; represents the position of j at time t;
[0070] Step 2.5, construction of multi-agent reward function:
[0071] Based on the framework of deep reinforcement learning, the reward function also sets a sparse reward mechanism to guide agents to avoid obstacles and other agents; the reward function is divided into sparse rewards and dense rewards;
[0072] When an agent achieves the target without collision, that is, ρ ig <d goal , a positive reward r arrive is given:
[0073] r arrive = C arrive + σ t Equation (16)
[0074] where C arrive is the reward constant for the agent to successfully reach the target; σ t is the time-related dense reward used to encourage the agent to reach the target quickly, σ t = -W(ΔT - t min ), W is the penalty parameter, ΔT is the consumption time, t min = [||P0, P ig || - d goal / v max , t min is the shortest time consumed by the agent at the maximum speed v max , P0 is the initial position of the agent, P ig is the target position, d goal is the distance threshold for the agent to reach the target;
[0075] When an agent collides, a negative reward r collision is given:
[0076] r collision = c collision When min(d1, d2,..., d7) < d collision Equation (17)
[0077] C collision = -C arriveEquation (18)
[0078] where C collision represents the penalty for the agent's collision; d collision is the distance threshold for collision;
[0079] When the agent neither reaches the target nor collides, four non-sparse reward functions are set:
[0080]
[0081] r4 = C t Equation (22)
[0082] r1 represents the distance difference of agent i relative to the target from time t - 1 to t. If r1 > 0, it means the agent is approaching the target and a positive reward is given; otherwise, a penalty is imposed. r2 represents the heading angle difference of agent i relative to the target from time t - 1 to t. If r2 > 0, it means the agent is adjusting the angle towards the target and a positive reward is given; otherwise, a penalty is imposed;
[0083] where represents the relative position of agent i and the target at time t; represents the heading angle difference between agent i and the target at time t;
[0084] r3 represents the agent's collision warning. When an obstacle is detected within the sensor detection range, r3 = 0. When an obstacle is detected, a penalty is given to warn the agent to stay away from the obstacle;
[0085] C t represents the time penalty per moment, hoping that the agent can reach the target point as soon as possible;
[0086] The reward function r when the agent is moving normally normal :
[0087]
[0088] where represents the normalization of relevant values; μ1 + μ2 + μ3 + μ4 = μ = 1, and μ represents the contribution rate to the reward function.
[0089] Such as Figure 3 shown:
[0090] Furthermore, in step 3, the construction process of the Actor - Critic network is as follows:
[0091] In the MADDPG algorithm, an independent Actor policy network and a Critic evaluation network are used for each agent; the Actor network generates the optimal policy for the agent, and the Critic network evaluates the value function of each agent's action; the Actor network optimizes the policy through the policy gradient method based on the feedback provided by the Critic network to maximize the long-term cumulative reward; when there are n agents in the agent system, there are n pairs of independent Actor-Critic networks;
[0092] Step 3.1, construct the Actor policy network:
[0093] The Actor network takes the state s of the agent i and outputs the action a i ; the update of the Actor network is carried out by the policy gradient method, aiming to continuously optimize the policy parameter θ i so that the action a generated by the agent in the state s i can obtain the maximized long-term reward; the gradient of the cumulative reward expectation of agent i i is
[0094]
[0095] where is the Q value evaluated by the Critic network, representing the expected long-term reward of agent i in the joint state s and all agents' actions A1,…,A n ; represents the gradient of the Q value with respect to the action of agent i; μ i (A i ∣s i ) represents the policy function μ i , which takes the input state s i and outputs the action A i ; represents the gradient of the policy function μ i with respect to θ i ; θ i is the policy parameter of agent i, used to generate the policy function μ i ;
[0096] In the joint state, the multi-agent system learns from the experiences in the experience pool and continuously optimizes its policy using the Actor network to improve the coordination ability and task completion efficiency of the entire system;
[0097] Step 3.2, construct the Critic evaluation network:
[0098] In the MADDPG algorithm, the Critic network Q i (s,A|φi ) Evaluate the action A of each agent i and output the long-term return estimate of the agent in state s i ; The Critic network combines the joint state S and joint action A of all agents and calculates the Q-value based on this, that is, evaluates the long-term return of the current action in the global state; The update of the Critic network is based on the Temporal Difference error. By minimizing this error, the Critic network can gradually learn the long-term return of the agent performing different actions in each state;
[0099] The loss function of the Critic network represents the mean square error between the current Q-value prediction and the target Q-value, and the formula is as follows:
[0100]
[0101] where y i represents the target Q-value, which combines the immediate reward and the future discounted reward; γ represents the discount factor; is the immediate reward obtained by agent i at time t; represents the Q-value estimate of the target network for the next state s' and the next actions A'1, A' i …, A' n ; A' i = μ' i (s i ) represents the action of agent i at the next time, which is determined by its policy function μ' i (s i );
[0102] Furthermore, the parameters and the experience pool initialization method are as follows:
[0103] Step 3-1, set the number of training steps and the iteration period:
[0104] Set the total number of training steps M = 15000. During the training process, the simulation step size is set to Δt = 0.2s, and the maximum movement time of each agent within a single training episode is t max = 300×Δt = 60s to ensure that the agent fully explores the environment within a reasonable time range;
[0105] Step 3-2, set the sampling probability and related hyperparameters:
[0106] Set the number of samples η = 64, and the update frequency ξ of the Actor-Critic network = 100; To ensure that the agent learns from different types of experiences, it is necessary to set the sampling probability p0 = 0.1, and the sampling probability p0 controls the proportion of sampling from the excellent experience pool D' and the poor experience pool D''; Generate a random number p, p ∈ [0, 1]; If p ≤ p0, select the next action to execute according to the artificial potential field method; If p > p0, select the action according to the policy network;
[0107] Both the Actor network and the Critic network are fully connected neural networks, both including an input layer, two hidden layers and an output layer; The activation function of the hidden layer of the Actor network uses ReLU; The first hidden layer of the Actor network contains 64 nodes, and the second hidden layer contains 32 nodes; The output layer of the Actor network is 1 node, representing the action selection without manual intervention, and the activation function used is the tanh function. The learning rate of the Actor network is 0.001;
[0108] The number of nodes in both hidden layers of the Critic network is 64, and the output layer is 1 node, outputting the Q value, using the linear activation function y = x + b, where b is the bias parameter. The learning rate of the Critic network is 0.0001;
[0109] Both the Actor network and the Critic network use the Adam optimizer, and set the discount factor γ = 0.97 to balance the influence of the current reward and the future reward;
[0110] Step 3-3, initialize the noise model:
[0111] To increase the exploration ability of the agent, exploration noise needs to be added to each agent; Use the Ornstein-Uhlenbeck (OU) noise model to generate random perturbations when generating actions for each agent, thereby enhancing the exploration ability; The OU noise dx t The expression is as follows:
[0112] dx t = θ(μ - x t )dt + ρdW Equation (26)
[0113] where, x t represents the current value of the OU noise; θ represents the noise mean regression rate; μ represents the long-term mean of the noise; ρ represents the noise fluctuation intensity; dW represents the differential of the noise Brownian motion;
[0114] Step 3-4, initialize the experience pool:
[0115] In the 3E-MADDPG algorithm, the learning of the agent depends on three experience pools to store different types of experiences. The regular experience pool D stores the regular experiences collected by the agent when exploring the environment; the excellent experience pool stores the high-quality experiences generated by the HER algorithm and the artificial potential field method; the poor experience pool stores the failed experiences during the training process, and the poor experiences are generated by the counterexample model.
[0116] For the regular experience pool D: It stores the regular experiences collected by the agent when exploring the environment; each sample is a five-tuple (s t , A t , r t , s t+1 , g), where s t , A t , r t , s t+1 and g represent the state of the agent at the current moment, the action generated by the agent according to the policy network in the current state, the immediate reward obtained by the agent from the environment after executing the action A t , the next state of the agent after executing the action A t , and the task goal respectively.
[0117] For the excellent experience pool D': It stores the excellent experiences generated by the HER algorithm and the artificial potential field method; the excellent experiences are used to handle the sparse reward problem and the goal achievement problem in complex environments; when initializing the policy parameter θ i of the Actor network, set the excellent experience pool D' generated by HER.
[0118] For the poor experience pool D": It stores the failed experiences encountered by the agent during the training process. The failed experiences include the agent colliding and deviating from the path; the counterexample model helps the agent recognize and avoid repeating the same mistakes through the poor experiences; when initializing the policy parameter θ i of the Actor network, set the excellent experience pool D" generated by HER.
[0119] The capacity of the regular experience replay buffer is D = 1000000. Since D' and D" are auxiliary experience pools, their capacities are both set to 50000, and training starts after the experience pool capacity reaches 5000.
[0120] Furthermore, in step 4, the determination process of the action selection method of the agent is as follows:
[0121] Determine the action selection method of the agent according to the result of step 3;
[0122] If p ≤ p0, then go to step 4.1 and select the next action to execute according to the artificial potential field method;
[0123] If p > p0, go to step 4.2 and select an action according to the policy network;
[0124] Step 4.1, the agent generates a decision action according to the artificial potential field method: It is necessary to obtain the position Pi of the agent i at time t i The gravitational force received and the repulsive force Let Pi i = [Pi ix , Pi iy represent the position of the agent i, and Pi ig = [Pi igx , Pi igy be the target position. Then the normalized gravitational force component received by the agent is expressed as:
[0125]
[0126] and respectively represent the gravitational force in the x - direction and the y - direction;
[0127] The expression of the normalized repulsive force component received by the agent i from the adjacent object j is as follows:
[0128]
[0129] The position of the adjacent object j is Pj j = [Pj jx , Pj jy ; Pj jx and Pj jy respectively represent the positions of Pj j in the x - direction and the y - direction; and respectively represent the repulsive forces of the adjacent object j in the x - direction and the y - direction;
[0130] Combine the gravitational force and the repulsive force to obtain the resultant force Fi received by the agent i
[0131]
[0132] In the formula, represents the set of adjacent objects of the agent i; σ ij is the collision function, indicating the influence degree of each member in the adjacent objects on the repulsive force received by the agent i. The expression of σ ij is as follows:
[0133]
[0134] d ij is the Euclidean distance between the agent i and the adjacent object j; d mis the minimum collision distance between agent i and adjacent object j; d c is the sensor detection distance; d r is a constant, and its value range is (d m , d c ); The calculation formulas and expressions of parameters a, b, c, and d are as follows:
[0135]
[0136] The linear velocity v of agent i under the action of the resultant force i and the angular velocity ω i are calculated as follows:
[0137]
[0138] where, k u is the velocity control gain, and its value is the maximum velocity v max ; ||P i - P ig || represents the Euclidean distance between the position of agent i and the target position; k ω is the angular velocity control constant, and its value is k ω = 2.5; ψ i represents the heading angle value of agent i, represents the resultant force 's direction angle, and its value is represents 's derivative value with respect to time, 's expression is as follows:
[0139]
[0140] Step 4.2, generate the decision-making action of the agent based on the Actor network: The Actor network receives the state input of the agent, extracts features from the state through a multi-layer neural network, and finally outputs the action that the agent should execute; The Actor network enhances the exploration ability by introducing random noise to ensure effective exploration of the agent in an unknown environment; As the training progresses, the Actor network continuously optimizes the policy parameters;
[0141] The input of the Actor network is the current state s of the agent i , and the input layer passes the original state information to the hidden layer for feature extraction; Each layer in the hidden layer performs weighted processing on the input and processes the non-linear relationship through the activation function ReLu. The Actor network extracts the key features in the state layer by layer and finally generates the action A i of the agent:
[0142]
[0143] Among them, are the trainable parameters of the Actor network; at each time step Δt, the agent will interact with the environment according to the action a output by the Actor network i and update the parameters of the Actor network according to the feedback of the environment to optimize the policy and improve the action efficiency of the agent.
[0144] Furthermore, in step 5, the update of the agent state is as follows:
[0145] The agent interacts with the environment according to the current state s i and the action A generated by the policy network i , and the environment will feedback a new state s t+1 and a reward r t ; the update process of the state s t+1 ensures that the agent can adjust its behavior policy through the feedback received from the environment and provides data support for subsequent experience replay.
[0146] Furthermore, in step 6, the HER algorithm is used to generate excellent experiences as follows:
[0147] When the agent fails to complete the goal, the actually reached state is redefined as a new goal, so that the experience trajectory originally regarded as a failure can still provide experience for policy learning; in each training cycle, the agent starts from the initial state s1 and executes the policy π, samples the trajectory according to the current policy and stores the policy sampling trajectory at time t in the experience pool D in the format of (s t , A t , r t , s t+1 , g); when the agent fails to complete the goal g, the HER algorithm randomly selects the state st' from the current trajectory as the alternative goal g' and recalculates the reward r t ', forms a new sampling trajectory {s t ', A t , r t ', s t+1 , g'}; stores the new sampling trajectory in the excellent experience pool D'.
[0148] The counterexample module generates poor-quality experiences in two ways:
[0149] One way is to directly extract poor-quality experiences from the collision or failure behaviors of the agent: once a collision or failure behavior is detected, immediately store the final state of the trajectory as poor-quality experience in the poor-quality experience pool D”;
[0150] Another way is to generate inferior experiences through reverse operations on excellent historical experiences; during the trajectory sampling process, monitor whether the behavior of the intelligent agent shows collision or failure behavior: randomly sample some high-quality experiences at time t from the excellent experience pool D' (s t , A t , r t , s t+1 , g), and generate inferior experiences by reversely adjusting s t and A t ; calculate the reversely generated s new+1 and r new , and store the generated inferior experiences (s new , A new , r new , s new+1 ) into the inferior experience pool D”; the counterexample module provides more diverse training data for the policy learning of the intelligent agent, further improving the robustness and adaptability of the algorithm.
[0151] Furthermore, in step 7, the judgment on whether the multi-agent enters the training stage is as follows:
[0152] When the intelligent agent explores the environment, it needs to store experiences first for subsequent training and update; the number of experiences in the regular experience pool D, the excellent experience pool D', and the inferior experience pool D” are |D|, |D'|, and |D”| respectively; add |D|, |D'|, and |D”| to obtain the total number of experiences stored in the three experience pools; compare the total number of experiences with the sampling number η. If |D| + |D'| + |D”| ≥ Q, the condition for entering the training stage is met, and go to step 8; if |D| + |D'| + |D”| < Q, go to step 4 to continue generating the number of experiences.
[0153] Furthermore, in step 8, the setting process of the sampling strategy and experience replay is as follows:
[0154] Step 8.1, perform experience sampling:
[0155] The intelligent agent will sample experiences from the regular experience pool D, the excellent experience pool D', and the inferior experience pool D” to update the parameters of the Actor-Critic network; set the sampling probability as p0 = 1 / 8, and the number of samples sampled from D, D', and D” are n0, n1, and n2 respectively:
[0156]
[0157] Randomly sample from the three experience pools and store the sampled experience data into the sample set for calculating the target Q value;
[0158] Step 8.2, calculate the target Q value:
[0159] According to the sampled empirical samples Use the target Critic network to calculate the target Q-value, which combines the current immediate reward r i and the expectation of future rewards to help the agent evaluate the long-term return of performing a specific action in the current state;
[0160]
[0161] Step 8.3, update the Critic network to optimize the agent's Q-value estimation:
[0162] The Critic network calculates the error between the currently estimated Q-value and the target Q-value to adjust the Critic network parameters; by minimizing the error, the Critic network can learn more accurate Q-values, providing an accurate reference for the policy update of the Actor network; the update method of the Critic network is shown in the following formula:
[0163]
[0164] Loss function The smaller it is, the more accurate the Critic network's estimation of the Q-value; the update of the Critic network can learn more accurate Q-values;
[0165] Step 8.4, update the Actor network:
[0166] The goal of the Actor network is to find the action policy that maximizes the Q-value of the Critic network; the update formula of the Actor network is as follows:
[0167]
[0168] The Actor network maximizes the Q-value estimated by the Critic network by minimizing the policy gradient loss such that the generated action can maximize the Q-value estimated by the Critic network.
[0169] Step 8.5, determine whether to update the target network:
[0170] The target network parameters are updated only once every fixed time interval ξ, which can reduce the oscillation during the policy update process and make the estimation of the target Q-value more stable; when Δt % ξ = 0 is satisfied, it turns to Step 8.7; when Δt % ξ = 0 is not satisfied, the current target network parameters are continued to be used without updating the Actor-Critic network parameters;
[0171] Step 8.6, update the target network in a soft update manner to maintain the stability of the target network:
[0172] Target network parameters and The update method is as follows:
[0173]
[0174] Among them, is the parameter of the Critic network; is the parameter of the Actor network; τ is the update coefficient that gradually increases from 0 to 1, which makes the parameters of the Actor and Critic networks update slowly, improving the stability of neural network training;
[0175] Step 8.7, iteration termination:
[0176] The training of the neural network continues until the Actor and Critic networks of the agent have reached the expected performance standard through the training data; when the training terminates, the agent system saves the parameters of the Actor and Critic networks at the termination and uses these parameters in practical applications.
[0177] The beneficial effects of the present invention are as follows:
[0178] (1) By using the HER algorithm, the present invention takes the actual state reached by the agent during replay as a new target, enabling the experience trajectories that were originally regarded as failures to still provide useful information for policy learning, and storing the originally regarded as failed experiences in the excellent experience pool to form an excellent experience pool; excellent experiences usually contain the optimal decisions made by the agent in a specific environment, which helps the agent quickly learn and imitate these successful strategies.
[0179] (2) The present invention directly extracts negative experiences from the collision or failure behaviors of the agent, and at the same time generates inferior experiences by reversing historical high-quality experiences to form an inferior experience pool; the samples in the inferior experience pool help to enhance the learning performance of the agent in a complex environment, enabling it to better avoid failure behaviors and optimize the decision-making process;
[0180] (3) The MADDPG algorithm involving the mixing of three types of experiences in the present invention divides the experience pool into three types of experience pools: a conventional experience pool, an excellent experience pool, and an inferior experience pool, effectively improving the utilization rate and diversity of experience samples, and avoiding the problem of too high a proportion of meaningless experiences in the traditional MADDPG algorithm;
[0181] (4) The present invention determines the sampling ratio from different experience pools according to the set sampling probability; samples are taken from the excellent experience pool and the inferior experience pool respectively, and the remaining samples are taken from the conventional experience pool; the sampling strategy of mixing three types of experiences can ensure that the agent learns from various types of experiences, enabling it to evenly obtain policy information from different experience pools during the learning process. Description of the Drawings
[0182] Figure 1 Flowchart of the multi-agent motion control method for mixing three types of experiences;
[0183] Figure 2 Schematic diagram of multi-agent motion control in a two-dimensional space;
[0184] Figure 3 Framework diagram of the MADDPG algorithm for mixing three types of experiences;
[0185] Figure 4 Graph of arrival rate test data when the number of agents is 5, 6, 7, and 8 respectively. Detailed implementation method
[0186] As Figure 2 shown, agent i is deployed at the initial position P i (represented by the blue circular frame), and its target position is P ig (represented by the blue-filled circle). During the movement process, the agent must move within the limited space range, pass through various obstacle areas (represented by black circles), and must not collide or conflict with other moving obstacles (such as other agents j and k). The obstacles move irregularly, and each agent can sense the potential obstacle threats ahead through on-board light detection and ranging or other sensors.
[0187] Figure 2 Shows the scene layout of the multi-agent motion planning model, where the circles with colored borders represent agents with a radius of r i different agents, the arrows on them indicate the movement directions, the pure-color-filled circles are the destinations of the corresponding agents, and the solid-color lines connecting the two circles represent the trajectories of the agents. The circles with black borders represent obstacles with a radius of r obs moving obstacles, and the dotted lines starting from the obstacles are regarded as their random trajectories. Define the collision distance between the agent and the obstacle as D io = r i + r obs . The agent can finally reach the target position safely through continuous sensing, obstacle avoidance, and path planning in the environment, thus completing the autonomous motion task in a complex environment.
[0188] In actual training, for the cases of 5 to 8 agents and different numbers of obstacles, the 3E-MADDPG and MADDPG algorithms were respectively tested. Under each experimental scenario, 1000 episodes of tests were conducted, and the positions of the agents, targets, and obstacles in each episode were randomly set. The test episode ends under any of the following conditions: all agents reach the target, all agents collide, or the episode time reaches the upper limit. The experimental index is the average arrival rate of the agents, and the specific data is shown in Table 1, and the relevant curves are as Figure 4 shown.
[0189] Test data for agents 5, 6, 7, and 8
[0190]
[0191] Analysis of the chart shows that in an environment with five agents, when the number of obstacles is 5, the arrival rate of the improved MADDPG exceeds 95% and reaches 97.69%, while the arrival rate of the MADDPG algorithm is only 86.80%. When the number of obstacles is 10, the arrival rate of MADDPG drops to 71.94%, indicating that the algorithm is difficult to effectively control agents in a complex environment. However, the improved MADDPG algorithm can still maintain an arrival rate of 79.55%. As the number of obstacles increases, the performance of the improved MADDPG decreases relatively slowly, while that of MADDPG decreases rapidly, especially when the number of obstacles is greater than 8, and the performance of MADDPG drops sharply. This shows that the improved MADDPG is more adaptable to environmental complexity..
[0192] From the overall results, the average arrival rate of the improved MADDPG algorithm at all obstacle numbers is 88.61%, which is about 9.93% higher than the average arrival rate of MADDPG at 78.67%. Especially when the number of obstacles is large, such as 10 obstacles, the arrival rate of the improved MADDPG algorithm can still remain at about 79.55%, showing good robustness.
[0193] Regarding the test results for different numbers of agents, when there are 10 obstacles in the environment, the improved MADDPG can still maintain an arrival rate of 80.53% when planning for 6 agents, while MADDPG only has 73.05%; when the number of agents becomes 7, the arrival rate of the improved MADDPG is 78.60%, remaining around 80%, while the arrival rate of MADDPG has dropped below 70%; when the number of agents continues to increase to 9, the arrival rate of the improved MADDPG is 73.71%, still higher than that of MADDPG. As the number of agents increases from 5 to 8, the arrival rates of both algorithms decrease, but the decreasing trend of the improved MADDPG is relatively gentle, maintaining a high arrival rate. Even in the case of 8 agents, the improved MADDPG still maintains an arrival rate of over 73.71%. This shows that the improved MADDPG has stronger scalability in multi-agent cooperation tasks. As the number of agents increases, the complexity of the system will increase significantly, including factors such as communication between agents and path coordination. The improved MADDPG can better handle such multi-agent cooperation tasks, while MADDPG seems rather inadequate when dealing with multi-agent tasks.
[0194] In summary, the improved MADDPG algorithm demonstrates stronger adaptability and robustness in test environments with multiple agents and different numbers of obstacles. Whether in scenarios with fewer or more obstacles, the arrival rate of the improved MADDPG is always significantly higher than that of MADDPG. Especially in high-complexity environments such as those with 10 obstacles, the improved MADDPG algorithm shows obvious advantages. This indicates that the improved MADDPG algorithm can better plan paths in complex and changing environments, effectively improving the task completion ability of agents.
Claims
1. A MADDPG multi-agent motion control method that mixes three types of experiences, characterized in that, The technical solution of the control method includes the following steps: Step 1: Set the environment model and the agent motion model of the agent. Step 2: Construct the multi-agent joint state model. Step 3: Construct the Actor-Critic network of the MADDPG algorithm, and initialize the parameters of the Actor-Critic network and the experience pool. Step 4: Determine the way for the agent to select actions. Step 5: Update the agent state. Step 6: Use the HER algorithm to generate excellent experiences and use the counterexample module to generate poor experiences. Step 7: Determine whether the multi-agent enters the training stage. Step 8: Set the sampling strategy and experience replay.
2. The multi-agent motion control method according to claim 1, wherein In Step 1, the environment model is set as follows: Each agent in the multi-agent system is a circular body with a radius of r for the agent i ; The external environment of the multi-agent system contains a series of dynamic obstacles. During flight, the agent needs to avoid circular obstacles with a radius of r obs to safely reach the target position; The collision distance D io between the agent and the obstacle is i r obs + r io ; When the Euclidean distance d between the agent and the obstacle satisfies d ≤ D io , it means that a collision has occurred to the agent When the obstacle is in a moving state, the obstacle motion model is the same as the agent motion model. The agent motion model is set as follows: The agent motion model simulates the motion of an agent in a two-dimensional plane, flying at a constant altitude and turning with inertial coordination; the agent obtains its own data through an inertial measurement device, and the update of the agent's motion state at each time step t is determined by the acceleration velocity position heading angle and angular velocity determined by: Angular velocity is calculated as follows: wherein, is the heading angle of the agent at time t and is the rate of change, i.e., the angular velocity; is the velocity in the x direction; is the velocity in the y direction; the range of the heading angle is [0, 2π]; The agent obtains acceleration through an inertial measurement device and velocity Obtain the current position based on the navigation data Acceleration is determined by the resultant force acting on the agent and the mass m u of the agent, and the acceleration is expressed as: and respectively represent the accelerations in the x - direction and y - direction; and respectively represent the resultant forces in the x - direction and y - direction; The speed of the agent at time t is determined by the speed at the previous time t-1 and the acceleration ; similarly, the position information is also determined by the position information at the previous time and the current speed. The formulas for speed and information are as follows: and respectively represent the positions in the x - direction and y - direction; δt represents the time interval from the previous moment t - 1 to the current moment t; To ensure the physicality of the agent system, the maximum acceleration a of the agent is set max and the maximum speed v max to prevent the motion parameters of the agent from exceeding the actual constraints.
3. The multi-agent motion control method according to claim 1, characterized in that, In Step 2, the process of constructing the multi-agent joint state model is as follows: Step 2.1, Use the Markov decision model for multi-agent path planning: The Markov decision model is described by a five-tuple <N, S, A, C, R>. Among them, N = {1, 2, …, n}, representing a set of n agents; n represents the number of agents; S = s1 × s2 × s i × … × s n , S represents the joint state space, and s i is the state of agent i. S is composed of the state combinations of all agents. Because the current state of an agent is not only related to the previous state of the agent itself but also closely related to the states of other agents; A = A1 × A2 × A i × … × A n , A represents the joint action space, and A i is the action of agent i. Each agent selects an action based on the current state and the cooperation or competition information from other agents; C is the state transition model, and the state transition model is: S × A × S → [0, 1], indicating the probability of the multi-agent transferring to the next joint state after taking a joint action in the current joint state; R = r1 × r2 × r i × … × r n , R is the joint reward function, and r i represents the reward value obtained by agent i under a specific state-action pair; In the Markov decision model, the optimal policy J(θ i ) of the agent is to maximize the joint cumulative reward of the multi-agent system, and J(θ i ) is achieved through the expected reward: Among them, θ i is the policy parameter of agent i; G i is the expected cumulative return of agent i under the current policy; γ t is the discount factor at time t, and the discount factor represents the importance that the agent attaches to future rewards; is the immediate reward of agent i at time t; Step 2.2, Construct the multi-agent joint state space: The sensor of each agent can detect the detection information in 7 different directions to obtain the detected distance value; when detecting other agents or obstacles, the sensor returns the straight-line distance between the agent and the detected object; if no object is detected, the sensor returns the default value of the maximum detection distance; through the sensor, the agent can obtain multi-dimensional information about itself, the environment, and the target point, so as to make decisions in complex task scenarios. Agent state information S i including its own state information s u , sensor detection information s r and environmental information s e ; that is, S i =(s u , s r , s e ); s u = [v i , ψ i Equation (6) Among them, the self-state information s u includes the current speed v of the agent i and the heading angle ψ i ; the sensor detection information s r is composed of the detection information returned by the sensor, and respectively represent the distance detection values in 7 directions detected by the sensor; the target position P ig =(P igx , P igy ), P igx and P igy respectively represent the positions of P ig in the x-direction and y-direction; the relative position between the agent i and the target The relative distance between the agent i and other agents represents the heading angle difference between the agent i and the target; represents the heading angle difference between the agent i and other agents, and the range of s e is [0, π]; Normalize the state information S of the agent i and map it to the range of [0, 1]; S i The normalization process is as follows: First, normalization of the self-state information s u : Among them, v i represents the velocity of agent i; ψ i represents the heading angle of agent i; π represents the set of strategies selected by all agents during the entire motion process; Secondly, the normalization of the sensor detection information s r : where d R represents the maximum detection distance of the sensor of agent i; Finally, the standardization of environmental information s e : Among them, L max is the boundary length of the environment; The final agent state space S i is a 15-dimensional vector that contains all the information related to multi-agent cooperation and competition tasks, ensuring that the agent can effectively learn and adapt to changes in different task scenarios; Step 2.3, Construct the multi-agent joint action space: By controlling the acceleration and angular velocity of the agent, the motion attitude and flight speed of the agent can be controlled; the action space A of the agent i is as follows: A i = [a i , ω i Formula (12) where a i is the acceleration of agent i; ω i is the angular velocity of agent i, and a i ∈[-a max , a max , ω i ∈[-ω max , ω max , a max represents the maximum acceleration of agent i, and ω max represents the maximum angular velocity of agent i; To ensure that the agent moves under physical constraints, the velocity v of agent i i and the heading angle ψ i both need to satisfy the constraints that v i ∈[0, v max , ψ i ∈[0, ψ max , where v max represents the maximum velocity of the agent, and ψ max represents the maximum heading angle of the agent; Action space A i describes the motion adjustment ability of the agent at any moment. Due to the complex cooperation and competition relationships among agents, the joint action space A = A1 × A2 × A i ×…× A n is crucial for the overall motion control of the agent system; Step 2.4, Construct the multi-agent objective function: The goal of path planning is to ensure that the agent reaches the target position in the shortest time and path without collisions, so as to achieve the optimal motion strategy π for all agents to reach the target * : Among them, P i represents the position of agent i; P ig represents the target position; T i represents the time when agent i reaches the target position; L i represents the total path length for agent i to reach the target position; the optimal motion strategy π * , that is, the strategy that minimizes the total time value for all agents to reach the target and the total path length; D io is the collision distance; means that at any moment, agent i cannot collide with obstacles or other agents; j represents an obstacle or other agent, that is, an adjacent object; represents the position of agent i at time t; represents the position of j at time t; Step 2.5, Construct the multi-agent reward function: Based on the framework of deep reinforcement learning, the reward function also sets a sparse reward mechanism to guide the agent to avoid obstacles and other agents; the reward function is divided into sparse rewards and dense rewards. When the agent achieves the goal without a collision, i.e., ρ ig <d goal , a positive reward r arrive : r arrive = C arrive + σ t Equation (16) Among them, C arrive is the reward constant for the agent to successfully reach the target; σ t is the time-related dense reward, used to encourage the agent to reach the target quickly, σ t = -W(ΔT - t min ), W is the penalty parameter, ΔT is the consumption time, t min = [||P0, P ig || - d goal / v max , t min is the shortest time consumed by the agent at the maximum speed v max , P0 is the initial position of the agent, P ig is the target position, d goal is the distance threshold for the agent to reach the target; When the agents collide, give a negative reward r collision : r collision = C collision When min(d1, d2, …, d7) < d collision at this time, Equation (17) C collision =-C arrive Equation (18) Among them, C collision represents the penalty for the agent to collide; d collision is the distance threshold for collision; When the agent does not reach the target and does not collide, 4 non-sparse reward functions are set: r4 = C t Equation (22) r1 represents the distance difference of agent i relative to the target from time t-1 to t. If r1>0, it means the agent is approaching the target, and a positive reward is given; otherwise, a penalty is imposed. r2 represents the heading angle difference of agent i relative to the target from time t-1 to t. If r2>0, it means the agent is adjusting the angle towards the target, and a positive reward is given; otherwise, a penalty is imposed. Among them, represents the relative position between agent i and the target at time t; represents the heading angle difference between agent i and the target at time t; r3 represents the agent collision warning. When an obstacle is detected within the sensor detection range, r3 = 0. When an obstacle is detected, a penalty is given to warn the agent to stay away from the obstacle. C t Represents the time penalty at each moment, hoping that the intelligent agent can reach the target point as soon as possible; Reward function r when the agent moves normally normal : Among them, represents the normalization of relevant numerical values; μ1 + μ2 + μ3 + μ4 = μ = 1, where μ represents the contribution rate to the reward function.
4. The multi-agent motion control method according to claim 1, wherein In Step 3, the process of constructing the Actor-Critic network is as follows: In the MADDPG algorithm, an independent Actor policy network and a Critic evaluation network are used on each agent; the Actor network generates the optimal policy for the agent, and the Critic network evaluates the value function of each agent's action; the Actor network optimizes the policy through the policy gradient method according to the feedback provided by the Critic network to maximize the long-term cumulative return; when there are n agents in the agent system, there are n pairs of independent Actor-Critic networks. Step 3.1, construct the Actor policy network: The Actor network outputs an action a based on the state s of the agent i ; The update of the Actor network is carried out by the policy gradient method, aiming to continuously optimize the policy parameters θ i , so that the action a generated by the agent in the state s i can obtain the maximum long-term reward; The gradient of the expected cumulative return of agent i i under the state s i for the action a generated Among them, is the Q value evaluated by the Critic network, representing the expected long-term return of agent i in the joint state s and all agent actions A1, …, A n under; represents the gradient of the Q value with respect to the action of agent i; μ i (A i ∣s i ) represents the policy function μ i , taking the input state s i and outputting the action A i ; represents the gradient of the policy function μ i with respect to θ i ; θ i are the policy parameters of agent i, used to generate the policy function μ i ; In the joint state, the multi-agent system learns from the experiences in the experience pool and continuously optimizes its policy using the Actor network to improve the coordination ability and task completion efficiency of the entire system; Step 3.2, construct the Critic evaluation network: In the MADDPG algorithm, the Critic network Q i (s, A|φ i ) evaluates the action A of each agent i and outputs the long-term return estimate of the agent in state s i ; the Critic network combines the joint state S and joint action A of all agents and calculates the Q value based on this, that is, evaluates the long-term return of the current action in the global state; the update of the Critic network is based on the Temporal Difference error. By minimizing this error, the Critic network can gradually learn the long-term return of the agent performing different actions in each state; Loss function of the Critic network Represents the mean squared error between the current Q-value prediction and the target Q-value, and the formula is as follows: Among them, y i represents the target Q value, which combines the immediate reward and the future discounted reward; γ represents the discount factor; is the immediate reward obtained by agent i at time t; represents the Q-value estimation of the target network for the state s' at the next time step and the actions A'1, A' i …, A' n ; A' i = μ' i (s i ) represents the action of agent i at the next time step, which is determined by its policy function μ' i (s i ).
5. The multi-agent motion control method according to claim 1, characterized in that In Step 3, the parameters and the experience pool initialization method are as follows: Step 3-1, set the number of training steps and the iteration period: Set the total number of training steps \(M = 15000\). During the training process, the simulation step size is set to \(\Delta t=0.2s\), and the maximum movement time of each agent within a single training episode is \(t\) max \(= 300\times\Delta t = 60s\) to ensure that the agents can fully explore the environment within a reasonable time range; Step 3-2, set the sampling probability and related hyperparameters: Set the sampling number η = 64, and the Actor-Critic network update frequency ξ = 100; to ensure that the agents learn from different types of experiences, it is necessary to set the sampling probability p0 = 0.1, and the sampling probability p0 controls the sampling ratio from the excellent experience pool D' and the poor experience pool D”; Generate a random number p, p ∈ [0, 1]; if p ≤ p0, then select the next action to execute according to the artificial potential field method; if p > p0, then select the action according to the policy network; Both the Actor network and the Critic network are fully connected neural networks, both including an input layer, two hidden layers, and an output layer; the activation function of the hidden layer of the Actor network uses ReLU; the first hidden layer of the Actor network contains 64 nodes, and the second hidden layer contains 32 nodes; the output layer of the Actor network is 1 node, representing the action selection without human intervention, and the activation function used is the tanh function, and the learning rate of the Actor network is 0.001; The number of nodes in the two hidden layers of the Critic network is 64 each, and the output layer is 1 node, outputting the Q value, using the linear activation function y = x + b, where b is the bias parameter, and the learning rate of the Critic network is 0.0001; Both the Actor network and the Critic network use the Adam optimizer, and set the discount factor γ = 0.97 to balance the influence of the current reward and the future reward; Step 3-3, initialize the noise model: To increase the exploration ability of the agent, exploration noise needs to be added to each agent; the Ornstein-Uhlenbeck (OU) noise model is used to generate random perturbations when generating actions for each agent, thereby enhancing the exploration ability; the OU noise dx t The expression is as follows: dx t = θ(μ - x t )dt + ρdW Equation (26) where x t represents the current value of the OU noise; θ represents the noise mean reversion rate; v represents the long-term mean of the noise; ρ represents the noise volatility intensity; dW represents the differential of the noise Brownian motion; Step 3-4, initialize the experience pool: In the 3E-MADDPG algorithm, the learning of the agents depends on three experience pools to store different types of experiences. The regular experience pool D stores the regular experiences collected by the agents when exploring the environment; the excellent experience pool stores the high-quality experiences generated by the HER algorithm and the artificial potential field method; the poor experience pool stores the failed experiences during the training process, and the poor experiences generated by the counterexample model; For the regular experience pool D: store the regular experiences collected by the agent when exploring the environment; each sample is a five-tuple (s t , A t , r t , s t+1 , g), where s t , A t , r t , s t+1 and g represent the state of the agent at the current moment, the action generated by the agent according to the policy network in the current state, the immediate reward obtained by the agent from the environment after executing the action A t , the next state of the agent after executing the action A t , and the task goal, respectively; For the excellent experience pool D': store the excellent experiences generated by the HER algorithm and the artificial potential field method; the excellent experiences are used to handle the sparse reward problem and the goal achievement problem in complex environments; initialize the policy parameters θ of the Actor network i When setting, use the excellent experience pool D' generated by HER; For the inferior experience pool D": Store the failure experiences encountered by the agent during training. The failure experiences include the agent colliding and deviating from the path; the counterexample model helps the agent recognize and avoid repeating the same mistakes through the inferior experiences; initialize the policy parameter θ of the Actor network i When setting the excellent experience pool D" generated by HER; The capacity size of the regular experience replay buffer is D = 1000000. Since D' and D” are auxiliary experience pools, their capacities are both set to 50000, and training starts after the experience pool capacity reaches 5000.
6. The multi-agent motion control method according to claim 1, wherein In Step 4, the determination process of the action selection method of the agent is as follows: Determine the action selection method of the agent according to the result of Step 3; If p ≤ p0, then select the next action to execute according to the artificial potential field method; If p > p0, then select the action according to the policy network; Step 4.1, the agent generates a decision-making action according to the artificial potential field method: it is necessary to obtain the position Pi of the agent i at time t i Gravitational force received And repulsive force Let Pi i = [Pi ix , Pi iy represent the position of agent i, and Pi ig = [Pi igx , Pi igy be the target position. Then the normalized gravitational force component received by the agent is expressed as: and respectively represent the gravitational forces in the x - direction and the y - direction; The expression of the normalized repulsive force component of agent i affected by adjacent object j is as follows: The position of the adjacent object j is P j = [P jx , P jy ; P jx and P jy respectively represent the positions of P j in the x - direction and the y - direction; and respectively represent the repulsive forces of the adjacent object j in the x - direction and the y - direction; Combine the gravitational force and the repulsive force to obtain the resultant force acting on agent i wherein, represents the set of adjacent objects of agent i; σ ij is a collision function, indicating the influence degree of each member in the adjacent objects on the repulsive force received by agent i, σ ij has the following expression: d ij is the Euclidean distance between agent i and adjacent object j; d m is the minimum collision distance between agent i and adjacent object j; d c is the detection distance of the sensor; d r is a constant, and its value range is (d m , d c ); The calculation formulas and expressions of parameters a, b, c, and d are as follows: The linear velocity v of agent i under the combined force i and the angular velocity ω i are calculated as follows: where k u is the speed control gain, with a value of the maximum speed v max ; ||P i -P ig || represents the Euclidean distance between the position of agent i and the target position; k ω is the angular velocity control constant, with a value of k ω = 2.5; ψ i represents the heading angle value of agent i, represents the resultant force of the direction angle, with a value of represents the derivative value with respect to time, The expression of is as follows: Step 4.2, generating the decision-making actions of the agent based on the Actor network: The Actor network receives the state input of the agent, extracts features from the state through a multi-layer neural network, and finally outputs the actions that the agent should execute; the Actor network enhances the exploration ability by introducing random noise to ensure effective exploration of the agent in an unknown environment; as the training progresses, the Actor network continuously optimizes the policy parameters; The input of the Actor network is the current state s of the agent i , and the input layer extracts features by passing the original state information to the hidden layer; each layer in the hidden layer will perform weighted processing on the input and handle the non-linear relationship through the activation function ReLu. The Actor network extracts the key features in the state layer by layer and finally generates the action A of the agent i : Among them, are the trainable parameters of the Actor network; at each time step Δt, the agent interacts with the environment according to the action a output by the Actor network i and updates the parameters of the Actor network according to the feedback of the environment to optimize the policy and improve the action efficiency of the agent.
7. The multi-agent motion control method according to claim 1, characterized in that In Step 5, the update of the agent state is as follows: The agent acts according to the current state s i and the action A generated by the policy network i , and interacts with the environment, which will feedback a new state s t+1 and a reward r t ; The state s t+1 The update process ensures that the agent can adjust its behavior strategy based on the feedback received from the environment and provides data support for subsequent experience replay.
8. The multi-agent motion control method according to claim 1, characterized in that, In Step 6, using the HER algorithm to generate excellent experiences is as follows: When the agent fails to achieve the goal, the actually reached state is redefined as the new goal, so that the experience trajectory that was originally regarded as a failure can still provide experience for policy learning; in each training cycle, the agent starts from the initial state s1 and executes the policy π, sampling the trajectory according to the current policy and storing the policy sampling trajectory at time t in the experience pool D in the format of (s t , A t , r t , s t+1 , g); when the agent fails to achieve the goal g, the HER algorithm randomly selects a state s t ' from the current trajectory as the substitute goal g', and recalculates the reward r t ' to form a new sampling trajectory {s t ', A t , r t ', s t+1 , g'}; store the new sampling trajectory in the excellent experience pool D'. The counterexample module generates poor experiences in two ways: One way is to directly extract inferior experiences from the collision or failure behaviors of the agent: once a collision or failure behavior is detected, immediately store the final state of the trajectory as an inferior experience in the inferior experience pool D”; Another way is to generate inferior experiences through reverse operations on excellent historical experiences; during the trajectory sampling process, monitor whether the behavior of the intelligent agent exhibits collision or failure behaviors: randomly sample some high-quality experiences at time t from the excellent experience pool D' (s t , A t , r t , s t+1 , g), and generate inferior experiences by reversely adjusting s t and A t ; calculate the reversely generated s new+1 and r new , and store the generated inferior experiences (s new , A new , r new , s new+1 ) into the inferior experience pool D”; the counterexample module provides more diverse training data for the policy learning of the intelligent agent, further enhancing the robustness and adaptability of the algorithm.
9. The multi-agent motion control method according to claim 1, characterized in that In Step 7, the judgment of whether the multi-agent enters the training stage is as follows: When exploring the environment, the agent needs to store experiences first for subsequent training and update; the number of experiences in the regular experience pool D, the excellent experience pool D', and the poor experience pool D” are |D|, |D'|, and |D”| respectively; adding |D|, |D'|, and |D”| together gives the total number of experiences stored in the three experience pools; comparing the total number of experiences with the sampling number η, if ||D| + |D|'| + |D”| ≥ Q, the condition for entering the training stage is met, and it transfers to Step 8; if |D| + |D'| + |D”| < Q, it transfers to Step 4 to continue generating the number of experiences.
10. The multi-agent motion control method according to claim 1, wherein In Step 8, the setting process of the sampling strategy and experience replay is as follows: Step 8.1, performing experience sampling: The agent will sample experiences from the regular experience pool D, the excellent experience pool D', and the poor experience pool D” to update the parameters of the Actor-Critic network; setting the sampling probability as p0 = 1 / 8, the number of samples sampled from D, D', and D” are n0, n1, and n2 respectively: n0 = N - n1 - n2 n1 = [N × p0] Equation (29) n2 = [N × p0] Randomly sample from three experience pools and store the sampled experience data in the sample set For calculating the target Q value; Step 8.2, calculating the target Q value: According to the sampled empirical samples Use the target Critic network to calculate the target Q-value, which combines the current immediate reward r i and the expectation of future rewards to help the agent evaluate the long-term return of performing a specific action in the current state; Step 8.3, updating the Critic network to optimize the Q value estimation of the agent: The Critic network calculates the error between the currently estimated Q value and the target Q value to adjust the parameters of the Critic network; by minimizing the error, the Critic network can learn a more accurate Q value, providing an accurate reference basis for the policy update of the Actor network; the update method of the Critic network is shown in the following formula: Loss function The smaller it is, the more accurate the Critic network's estimation of the Q value; the update of the Critic network can learn a more accurate Q value; Step 8.4, updating the Actor network: The goal of the Actor network is to find the action policy that maximizes the Q value of the Critic network; the update formula of the Actor network is as follows: The Actor network minimizes the policy gradient loss to make the generated actions maximize the Q-value estimated by the Critic network. Step 8.5, judging whether to update the target network: The target network parameters are updated only once every fixed time interval ξ, which can reduce the oscillation during the policy update process and make the estimation of the target Q value more stable; when Δt % ξ = 0 is satisfied, it transfers to Step 8.7; when Δt % ξ = 0 is not satisfied, continue to use the current target network parameters without updating the Actor-Critic network parameters; Step 8.6, updating the target network in a soft update manner to maintain the stability of the target network: Target network parameters and are updated as follows: Among them, parameters of the Critic network; are the parameters of the Actor network; τ is the update coefficient that gradually increases from 0 to 1, so as to slowly update the parameters of the Actor and Critic networks and improve the stability of neural network training; Step 8.7, iteration termination: The training of the neural network continues until the Actor and Critic networks of the agent have achieved the expected performance standards with the training data; when the training terminates, the agent system saves the parameters of the Actor and Critic networks at the termination and uses these parameters in practical applications.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle motion planning method based on artificial potential field method and MADDPG
CN112947562A
Mixed-experience multi-agent reinforcement learning motion planning method
CN113341958A
New energy vehicle ecological driving method based on heterogeneous multi-agent deep reinforcement learning
CN115495997A
Group intelligent learning method fusing thought of see-after-notice
CN115660052A
Method and apparatus for generating multi-drone network cooperative operation plan based on reinforcement learning
US20230297859A1
Cited By
Aircraft intelligent control method based on dynamic teaching deep reinforcement learning
CN122085645A