Multi-agent path planning method based on QMIX algorithm

By combining A* and QMIX algorithms, the multi-agent path planning model is optimized, and the problem of path planning flexibility and global target priority balance in dynamic environments is solved, achieving more efficient resource allocation and task completion.

CN120087875APending Publication Date: 2025-06-03HENAN UNIVERSITY

Patent Information

Application Number
CN202510250502.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing multi-agent path planning methods lack flexibility and dynamic balance of global target priorities in complex multi-task scenarios with dynamic changes, resulting in uneven resource allocation, task hunger and local optimality problems.

Method used

By combining the A* algorithm with the QMIX algorithm, a multi-agent path planning model is established, a post-event experience replay mechanism is used to optimize task allocation, and a noise mechanism is introduced to improve the diversity of strategy exploration, achieving a dynamic balance between global and local exploration.

Benefits of technology

It improves the adaptability of agents to complex environments and path planning efficiency, avoids uneven resource allocation and task hunger, and optimizes the overall effect of agents' collaboration and path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087875A_ABST
    Figure CN120087875A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-agent path planning method based on a QMIX algorithm, and relates to the technical field of machine learning, and the method comprises the steps: building a multi-agent path planning model based on the QMIX algorithm in combination with an A * algorithm, and training the multi-agent path planning model in combination with an after-event experience playback mechanism to obtain a trained multi-agent path planning model; and based on the trained multi-agent path planning model, realizing multi-agent path planning in a dynamically changing complex multi-task scene. The technical problem that an existing method lacks flexibility and dynamic balance of global target priority in a dynamically changing complex multi-task scene is solved, an innovative solution is provided for multi-agent path planning, efficient combination of a global strategy and local path optimization is achieved, the overall efficiency of path planning is improved, and the path planning efficiency is improved. And the method has universality and practicability and is worthy of popularization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and particularly to a multi-agent path planning method. Background Art

[0002] Multi-agent systems have become an important technical means for solving problems such as automated production, logistics transportation, and complex task scheduling. Especially in scenarios such as warehousing and manufacturing. Multiple agents often need to cooperate to complete tasks such as cargo handling and path planning. However, with the complication of the environment, problems such as dynamic obstacles, path conflicts between agents, and dynamic target adjustment make it difficult for traditional single planning algorithms to adapt to the complex and changeable environment.

[0003] Currently, the A* algorithm, as a classic path planning algorithm, is widely used for global path planning in static maps. However, the A* algorithm can only provide a static globally optimal path and is difficult to cope with the dynamically changing environment in real time. For example, when an agent encounters a suddenly emerging obstacle, the A* algorithm cannot quickly adjust the planned path, so there are certain limitations.

[0004] To address the complex cooperation problems in multi-agent systems, in recent years, path planning algorithms based on reinforcement learning have gradually received attention. Among them, the QMIX algorithm, as a distributed reinforcement learning algorithm, can effectively solve the partially observable problems in multi-agent systems, support each agent to make independent decisions and cooperate to complete global tasks. However, the existing QMIX algorithm still has deficiencies in the dynamically changing complex environment. Especially in the path planning task, how to combine global information with local dynamic changes is an issue that the existing technology urgently needs to solve.

[0005] The patent with the application number 202210868106.8 discloses a multi-agent unknown environment search and rescue method and system based on deep reinforcement learning. First, a Markov decision model for multi-agent unknown environment search and rescue is obtained; then, based on the Markov decision model, the QMIX algorithm is used to obtain the actions of each agent; each agent plans the optimal path from the current point to the next state target point using the A* algorithm based on the above actions; the process of continuously cycling the QMIX algorithm to determine actions and the A* algorithm to plan paths is repeated until the multi-agent unknown environment search and rescue result is output when a preset termination condition is reached. The above patent effectively improves the efficiency of multi-agent search and rescue in unknown environments. However, the above patent lacks flexibility in a dynamic environment and lacks dynamic balance of global target priorities in complex multi-task scenarios, which is likely to lead to uneven resource allocation, task starvation, and being trapped in local optima. Summary of the Invention

[0006] Aiming at the technical problems that existing multi-agent path planning methods lack flexibility and dynamic balance of global goal priorities in dynamically changing complex multi-task scenarios, the present invention proposes a multi-agent path planning method based on the QMIX algorithm. By combining the A* algorithm with the QMIX algorithm, the adaptability of agents to complex environments and the path planning efficiency are improved, the dynamic balance between global and local exploration is achieved, the situations of uneven resource allocation, task starvation, and falling into local optima are avoided, and the overall effect of agent cooperation and path planning is optimized.

[0007] To achieve the above object, the technical solution of the present invention is realized as follows:

[0008] A multi-agent path planning method based on the QMIX algorithm includes the following steps:

[0009] Based on the QMIX algorithm and combined with the A* algorithm, establish a multi-agent path planning model, and train the multi-agent path planning model in combination with the experience replay mechanism after the event to obtain a trained multi-agent path planning model;

[0010] Based on the trained multi-agent path planning model, realize multi-agent path planning in a dynamically changing complex multi-task scenario.

[0011] Preferably, the step of establishing a multi-agent path planning model based on the QMIX algorithm and combined with the A* algorithm includes:

[0012] S1. Model and transform the path planning problem of the agent into a distributed partially observable Markov decision model;

[0013] S2. Based on the distributed partially observable Markov decision model, use the QMIX algorithm to learn the global cooperation strategy of multi-agents, comprehensively analyze the global state information and task requirements, and optimize task allocation and overall behavior decision-making;

[0014] S3. Based on task allocation, determine the action of each agent according to the principle of maximizing the joint Q value in the QMIX algorithm; when the agent encounters an obstacle or conflicts with other agents, use the A* algorithm to plan the path from the current node to the local target node;

[0015] S4. Loop and execute steps S2 and S3 until the task is completed.

[0016] Preferably, the method of modeling and transforming the path planning problem of the agent into a distributed partially observable Markov decision model is:

[0017] S1.1. According to the physical characteristics and functional requirements of the agent, define the observations and actions of the agent, and model the global state of the agent;

[0018] The local observation information of agent i at time t includes the position vector of agent i velocity vector target position vector g i and obstacle information Then the local observation information of agent i at time t

[0019] wherein, the position vector of agent i the velocity vector of agent i the target position vector of agent i the obstacle information perceived by agent i includes the positions of the nearest dynamic and static obstacles; a Cartesian coordinate system is established with the lower left corner as the origin, represents the position of agent i in the x-axis direction, represents the position of agent i in the y-axis direction, represents the velocity component of agent i in the x-axis direction, represents the velocity component of agent i in the y-axis direction, represents the position of the target position vector of agent i in the x-axis direction, represents the position of the target position vector of agent i in the y-axis direction;

[0020] The action selected by agent i at time t is a discrete action of up, down, left, right or stop, and the moving distance of agent i depends on the time step and the applied acceleration;

[0021] The positions and velocities of all agents, the current state of the environment, and the relevant parameters of the task are used as the global state s;

[0022] S1.2. Based on the defined observations and actions of the agents, combined with the global state of the agents, design a structured reward function to obtain a distributed partially observable Markov decision model.

[0023] Preferably, the expression of the structured reward function is:

[0024]

[0025] where, K i ∈(0,1], i = 1, 2, 3, 4, 5, K 1 is the weight parameter of the reward when agent i reaches the target point of, K 2 is the penalty when agent i collides with or approaches an obstacle of, K 3 is the path length penalty The weight parameter, K 4 is the penalty for the invalid movement or delay of agent i The weight parameter, K 5 is the reward for agent i to maintain a safe distance from the dynamic obstacle The weight parameter.

[0026] Preferably, the expression for the reward when agent i reaches the target point is:

[0027]

[0028] The expression for the penalty when agent i collides with or approaches an obstacle is:

[0029]

[0030] where min_range represents the minimum distance from agent i to the nearest obstacle or other agent detected by agent i; the expression for the path length penalty is:

[0031]

[0032] where λ 3 represents the path penalty constant, and L path represents the path length of agent i;

[0033] The expression for the penalty for the invalid movement or delay of agent i is:

[0034]

[0035] where λ 4 represents the penalty constant for invalid movement or unnecessary waiting time, and t idle represents the unnecessary waiting time or invalid movement time of agent i;

[0036] The expression for the reward for agent i to maintain a safe distance from the dynamic obstacle is:

[0037]

[0038] where d represents the current distance between agent i and the obstacle, and d safe represents the predefined safe distance, and λ 5 、λ 6 and α are adjustable parameters.

[0039] Preferably, the distributed partially observable Markov decision model is represented by the seven-tuple where I = {1, 2, …, N} represents the set of agents, and N is the number of agents; S represents the global state space, and A iDenote the action space of agent \(i\), and the action \(a\) selected by agent \(i\) i ∈A i , at time \(t\), combine the actions of all agents to form the joint action \(a=(a\) 1 ,a 2 , ……,a N ) at time \(t\); \(O\) i is the local observation space generated by the observation function \(O\) i (s,a i ); \(P\) is the state transition probability; \(R\) represents the reward obtained when all agents take the joint action \(a\) under the global state information \(s\); \(\gamma\in[0,1)\) represents the discount factor used to control the weight of future rewards.

[0040] Preferably, the method of using the QMIX algorithm to determine the action of each agent according to the principle of maximizing the joint Q-value in the QMIX algorithm is as follows:

[0041] The QMIX network consists of a mixing network and the local Q-network of each agent; among them, the structure of the local Q-network includes an input layer - the first multi-layer perceptron - a gated recurrent unit - the second multi-layer perceptron connected in sequence, and a noise mechanism is introduced into the weights of the second multi-layer perceptron of the local Q-network;

[0042] The agent receives local observation information and previous actions through the input layer of the local Q-network and sends them to the first multi-layer perceptron. The first multi-layer perceptron performs feature fusion on the input data to generate a feature vector. The gated recurrent unit combines the historical trajectory information of the agent and the feature vector to update the trajectory representation of the agent. The second multi-layer perceptron serves as the output layer and maps the updated trajectory representation of the agent to the local Q-value of the agent; the agent selects an action based on the local Q-value by executing the \(\epsilon\)-greedy policy;

[0043] Since a noise mechanism is introduced into the weights of the second multi-layer perceptron of the local Q-network, during each forward propagation, the noise weight parameter and random noise will act dynamically on the weights of the second multi-layer perceptron, making the weights of the second multi-layer perceptron have the characteristics of randomization, and enhancing the exploration ability of the agent in the decision-making process and the diversity of action selection;

[0044] The mixing network receives the local Q-values of all agents and the global state as inputs, and generates the joint Q-value \(Q\) tot through a feedforward neural network based on a non-linear weighted mixing method, tot and ensures that the joint Q-value \(Q\) satisfies the monotonicity constraint tot such that maximizing the joint Q-value \(Q\) i is equivalent to maximizing each local Q-value \(Q\)

[0045] Preferably, when the agent encounters an obstacle or conflicts with other agents, the method for using the A* algorithm to plan the path from the current node to the local target node is as follows:

[0046] If the distance between the agent and the obstacle or other agents is less than the safety threshold d safe , trigger local path replanning: taking the current position as the starting point and the local target node as the end point, generate a set of candidate actions for obstacle avoidance paths in combination with dynamic environment information The subsequent actions of the agent are only selected from the set of candidate actions A for obstacle avoidance paths safe .

[0047] Preferably, training the multi-agent path planning model in combination with the experience replay mechanism afterwards includes the following steps:

[0048] Ⅰ. Initialize parameters and define the initial experience replay pool: Initializing parameters includes initializing the scenario and agents, initializing the network weight parameters of the QMIX algorithm, setting the total number of iteration rounds and the network parameter update frequency; Ⅱ. Path planning and action execution; Ⅲ. Experience storage and network update: After the agent executes an action, observe the new state feedback by the environment and the immediate reward obtained, record the samples of state transition, and store them in the experience replay pool; Calculate the temporal error of the samples in the experience replay pool and set the priority of the samples. When the capacity of the replay pool exceeds the threshold, remove the samples with the lowest priority; Minimize the weighted loss function and update the network weight parameters and task allocation strategies; Ⅳ. Task completion and evaluation: Repeat steps Ⅱ and Ⅲ until all agents reach the target position or reach the maximum number of iteration rounds. Perform performance evaluation on the multi-agent path planning model that has completed updating the network weight parameters and task allocation strategies. If the evaluation result meets the preset goal, the multi-agent path planning model that has completed updating the network weight parameters and task allocation strategies is the trained multi-agent path planning model.

[0049] Preferably, the calculation formula for the temporal error is:

[0050] δ = |Q tot (s,a;θ)-(R+γmax a′ (Q tot (s′,a′;θ - )))|

[0051] Set the sample priority p = |δ| + ∈, where ∈ is a constant to prevent zero probability, and calculate the weighted loss function value of the batch of samples drawn according to the probability , where α ∈ [0,1] is a variable to control uniformity;

[0052] The expression of the weighted loss function is:

[0053]

[0054] Among them, represents the expectation of the mean square error of multiple state-action pairs, is the experience replay pool, w i =(N·P(i)) -β is the importance sampling weight, β is a variable controlling the strength of bias correction, y = r + γQ tot (s′, argmax a′ Q tot (s′, a′; θ); θ - ) represents the target Q value at the current time step, Q tot (s, a; θ) represents the estimation of the joint action a in the global state s when the network parameter is θ, Q tot (s′, a′; θ - ) represents that when the network parameter is θ - , the estimation of the new joint action a′ in the new state s′.

[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0056] By introducing a noise mechanism into the QMIX algorithm, the present invention enhances the diversity of policy exploration, improves the decision-making robustness and adaptability of the multi-agent system. In a dynamic environment, the local Q network with a noise mechanism enables agents to better explore effective paths in the face of complex and uncertain factors and avoid falling into local optima. The present invention realizes optimization in terms of resource allocation and path conflict, improving the response speed and efficiency of multi-agent path planning.

[0057] By introducing a post hoc experience replay mechanism to optimize the priority task allocation, the present invention achieves a dynamic balance between global and local exploration, improves the adaptability and decision-making quality of agents in complex environments, strengthens the learning ability and path planning effect of agents. Post hoc experience replay allows agents to repeatedly utilize key path data in multiple iterations, thereby accelerating the convergence of the multi-agent path planning model and improving the learning efficiency in a dynamic environment. It also enables agents to summarize experience from historical decisions, optimize the execution path of current tasks, effectively avoid resource waste, and improve stability and decision-making accuracy.

[0058] The present invention provides an innovative solution for multi-agent path planning, realizes the efficient combination of global strategy and local path optimization, improves the overall efficiency of path planning, and has generality and practicality, worthy of popularization. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0060] Figure 1 This is the flowchart of the present invention.

[0061] Figure 2 This is a schematic diagram of the network structure of the QMIX algorithm introducing the noise mechanism of the present invention; among them, Figure 2 -(a) is a schematic diagram of the structure of the mixing network, Figure 2 -(b) is a schematic diagram of the structure of the QMIX network, Figure 2 -(c) is a schematic diagram of the structure of the local Q network

[0062] Figure 3 This is the flowchart of training the multi-agent path planning model of the present invention. Specific embodiments

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0064] As Figure 1 shown, a multi-agent path planning method based on the QMIX algorithm includes the following steps:

[0065] Based on the QMIX algorithm and combined with the A* algorithm, establish a multi-agent path planning model, and train the multi-agent path planning model in combination with the experience replay mechanism after the event to obtain a trained multi-agent path planning model.

[0066] Based on the trained multi-agent path planning model, realize multi-agent path planning in a dynamically changing complex multi-task scenario.

[0067] Furthermore, the establishment of the multi-agent path planning model based on the QMIX algorithm and combined with the A* algorithm includes the following steps:

[0068] S1. Model and transform the path planning problem of the agent into a distributed partially observable Markov decision model.

[0069] S1.1. Define the observations and actions of the agent according to its physical characteristics and functional requirements, and model the global state of the agent.

[0070] Specifically, the local observation information of agent i at time t includes the position vector of agent i velocity vector target position vector g i and obstacle information Then the local observation information of agent i at time t

[0071] where, the position vector of agent i the velocity vector of agent i the target position vector of agent i the obstacle information perceived by agent i includes the positions of the nearest dynamic and static obstacles. Establish a Cartesian coordinate system with the lower left corner as the origin. represents the position of agent i in the x-axis direction, represents the position of agent i in the y-axis direction, represents the velocity component of agent i in the x-axis direction, represents the velocity component of agent i in the y-axis direction, represents the position of the target position vector of agent i in the x-axis direction, represents the position of the target position vector of agent i in the y-axis direction.

[0072] The action selected by agent i at time t is a discrete action: up, down, left, right, or stop. The moving distance of agent i depends on the time step and the applied acceleration.

[0073] The positions and velocities of all agents, the current state of the environment (including the distribution of obstacles and the dynamic change characteristics and update frequency of the environment), and the relevant parameters of the task (including the target position, task priority, and task type) are used as the global state s.

[0074] S1.2. Based on the defined observations and actions of the agent, combined with the global state of the agent, design a structured reward function to obtain a distributed partially observable Markov decision model.

[0075] The reward when agent i reaches the target point is:

[0076]

[0077] The penalty when agent i collides with or approaches an obstacle is:

[0078]

[0079] Among them, min_range represents the minimum distance from agent i to the nearest obstacle or other agent detected by agent i.

[0080] To encourage agent i to choose the shortest path, the path length penalty is set as:

[0081]

[0082] Among them, λ 3 represents the path penalty constant, and L path represents the path length of agent i. The greater the path length, the greater the penalty, thus encouraging agent i to choose a shorter path.

[0083] The penalty for invalid movement or delay of agent i is:

[0084]

[0085] Among them, λ 4 represents the penalty constant for invalid movement or unnecessary waiting time, and t idle represents the unnecessary waiting time or invalid movement time of agent i. The longer the unnecessary waiting time or invalid movement time, the greater the penalty.

[0086] The reward for agent i to maintain a safe distance from dynamic obstacles is:

[0087]

[0088] Among them, d represents the current distance between agent i and the obstacle, and d safe represents the predefined safe distance, and λ 5 , λ 6 and α are adjustable parameters.

[0089] In summary, the reward obtained by agent i at time t is:

[0090]

[0091] Among them, K i ∈(0,1], i = 1, 2, 3, 4, 5, and K 1 is the weight parameter of the reward when agent i reaches the target point, K 2 is the weight parameter of the penalty when agent i collides with or approaches an obstacle, K 3 is the weight parameter of the path length penalty, K 4 is the weight parameter of the penalty for invalid movement or delay of agent i, and K 5 is the weight parameter of the reward for agent i to maintain a safe distance from dynamic obstacles.

[0092] The distributed partially observable Markov decision model is obtained by integrating the observations, actions, global states, and reward functions of the comprehensive agent.

[0093] The distributed partially observable Markov decision model is represented by a seven-tuple where \(I = \{1, 2, \ldots, N\}\) represents the set of agents, and \(N\) is the number of agents; \(S\) represents the global state space, which is the set of all global states \(s\in S\); \(A\) i represents the action space of agent \(i\), and the action \(a\) selected by agent \(i\) i \(\in A\) i , and at time \(t\), the actions of all agents are combined to form the joint action \(a=(a\) 1 , a\) 2 , \(\ldots\), a\) N ) at time \(t\); \(O\) i is the local observation space generated by the observation function \(O\) i (s, a\) i ); the state transition probability \(P\) represents the probability of transitioning to a new state \(s'\) given the joint action \(a\) under the global state \(s\); \(R\) represents the reward obtained when all agents take the joint action \(a\) under the global state information \(s\); \(\gamma\in[0, 1)\) represents the discount factor used to control the weight of future rewards.

[0094] S2. Based on the distributed partially observable Markov decision model, use the QMIX algorithm to learn the global cooperation strategy of multi-agent systems, comprehensively analyze the global state information and task requirements, and optimize the task allocation and overall behavior decision-making to ensure the optimality of the global path planning.

[0095] S3. Based on the task allocation, determine the action of each agent according to the principle of maximizing the joint Q-value in the QMIX algorithm, providing a unified long-term guidance for the overall goal; when an agent encounters an obstacle or conflicts with other agents, use the A* algorithm to plan the path from the current node to the local target node, quickly respond to obstacles or dynamic changes in the environment through local observation information, optimize the local path decision-making, and flexibly handle local obstacles and dynamic changes to ensure the efficient combination of local path decision-making and global path planning.

[0096] Furthermore, the network structure of the QMIX algorithm is as Figure 2 shown. The QMIX network consists of a mixing network and the local Q-network of each agent. Among them, the structure of the local Q-network includes an input layer - the first multi-layer perceptron - a gated recurrent unit - the second multi-layer perceptron. To enhance the diversity of policy exploration, a noise mechanism is introduced into the weights of the second multi-layer perceptron of the local Q-network.

[0097] Furthermore, asFigure 2 -(c) and Figure 2 As shown in -(b), the agent generates the local Q-value of the agent through the local Q-network Among them, represents the generated trajectory representation of agent i, represents the action of agent i at time t.

[0098] Specifically, at time t, the input layer of the local Q-network receives local observation information and the previous action and sends them into the first multi-layer perceptron MLP. The first multi-layer perceptron MLP performs feature fusion on the input data to generate a feature vector Then, the gated recurrent unit GRU combines the historical trajectory information and the output of the first multi-layer perceptron MLP to update the trajectory representation The second multi-layer perceptron MLP serves as the output layer and maps the updated trajectory representation to the local Q-value of the agent The agent executes the ∈-greedy policy to select actions based on the local Q-value

[0099] The expression for calculating the noise weight is as follows:

[0100] W noisy = W + σ W ·∈ W

[0101] Among them, W is the standard weight; σ W is the learnable noise weight parameter and will be dynamically updated according to the gradient information; ∈ W is the random noise sampled from the standard normal distribution.

[0102] Since a noise mechanism is introduced into the weights of the second multi-layer perceptron of the local Q-network, during each forward propagation, the noise weight parameter σ W and the random noise ∈ W will act dynamically on the weights of the second multi-layer perceptron MLP, making the weights of the second multi-layer perceptron MLP have the characteristics of randomization, thus effectively improving the exploration ability of the agent in the decision-making process and the diversity of action selection.

[0103] Furthermore, as shown in Figure 2 -(a) and Figure 2 -(b), the hybrid network receives the local Q-values of all agents and the global state s as inputs, and generates the joint Q-value Q through the feedforward neural network based on the non-linear weighted mixing method tot = W 2 ·σ(W 1 ·[Q 1 ,Q2 , …, Q N Τ +b 1 ) + b 2 Perform global task optimization, where W 1 and W 2 are non - negative weight matrices generated by the hyper - network, σ(·) is a monotonically increasing activation function, b 1 and b 2 are bias vectors generated by the hyper - network, and ensure that the joint Q - value Q tot satisfies the monotonicity constraint such that maximizing the joint Q - value Q tot is equivalent to maximizing each local Q - value Q i .

[0104] S4. Loop and execute step S2 and step S3 until the task is completed.

[0105] Furthermore, when training the multi - agent path - planning model, introduce experience replay, enabling the agents to summarize effective coping strategies from historical data, continuously optimize the local decisions of the agents, thereby achieving global path optimization and improving the adaptability and flexibility to dynamic environmental changes, as Figure 3 shown, including the following steps:[[]]

[0106] Ⅰ. Initialize parameters and define the initial experience replay pool: ① Initialize the scenario and agents: Convert the actual environment into a grid representation, define dynamic and static obstacles and their update rules; Set the initial positions and target positions of each agent, and set the task priorities according to the task requirements; Assign an initial state to each agent, including position, speed, historical trajectory, target point, etc. ② Initialize the network weight parameters of the improved QMIX algorithm, including all trainable parameters of the local Q - network and the mixing network; Set the total number of iteration rounds and the network parameter update frequency. ③ Define the initial experience replay pool for storing historical states, actions, rewards, and next states, etc.

[0107] Ⅱ. Path planning and action execution: Obtain the current global state s and the local observations collected by each agent The agents calculate the Q - values of each action through the local Q - network Adopt the ∈ - greedy strategy to select actions The mixing network fuses the local Q - values of all agents and the global state s to generate the joint Q - value Q tot , satisfying the monotonicity constraint Optimize the global strategy.

[0108] If the distance between the agent and an obstacle or other agents is less than the safety threshold d safe ​, trigger local path replanning. Taking the current position as the starting point and the local target node as the ending point, generate a set of candidate actions for obstacle avoidance paths in combination with dynamic environment information The subsequent actions of the agent are only selected from the set of candidate actions A for obstacle avoidance paths safe to ensure safety and path feasibility.

[0109] III. Experience Storage and Network Update: After the agent executes an action , observe the new state s′ feedback by the environment and the immediate reward Record the sample (s, a, R, s′) of the state transition and store it in the experience replay pool

[0110] Calculate the temporal difference error (TD error) of the samples in the experience replay pool and set the priority of the samples. When the capacity of the experience replay pool exceeds the threshold Remove the sample with the lowest priority;

[0111] The formula for calculating the temporal difference error is:

[0112] δ = |Q tot (s,a;θ) - (R + γmax a′ (Q tot (s′,a′;θ - )))|

[0113] Set the priority p of the sample = |δ| + ∈, where ∈ is a very small constant to prevent zero probability, and calculate the weighted loss function value of the batch of samples drawn according to the probability , where α ∈ [0,1] is a variable to control uniformity.

[0114] Minimize the weighted loss function:

[0115]

[0116] Among them, represents the expectation of the mean square error of multiple state-action pairs, D is the experience replay pool, w i = (N·P(i)) -β is the importance sampling weight, β is a variable to control the strength of bias correction, y = r + γQ tot (s′,argmax a′ Q tot (s′,a′;θ - ) represents the target Q value at the current time step, Q tot (x,a;θ) represents the estimation of the joint action a in the global state s when the network parameter is θ, Q tot (x′,a′;θ - ) represents that the network parameter is θ -At this time, the evaluation value of the new joint action a' in the new state x'. Synchronously update the target network parameter θ to θ every other time step - .

[0117] Update the network weight parameters and the task assignment strategy.

[0118] IV. Task Completion and Evaluation: Repeat steps II and III until all agents reach the target position or the maximum number of iterations is reached. Evaluate the performance of the multi-agent path planning model that has completed updating the network weight parameters and the task assignment strategy. If the evaluation result reaches the preset target, the multi-agent path planning model that has completed updating the network weight parameters and the task assignment strategy is the trained multi-agent path planning model

[0119] When all agents reach the target position, the task is considered completed; if the task fails (such as an agent failing to complete the task within the specified time, an agent colliding with an obstacle, or agents colliding with each other), record the relevant data for subsequent analysis and improvement, and then retrain. Evaluate the performance of the multi-agent path planning model through indicators such as the completion rate, average path length, calculation time, and conflict rate.

[0120] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi-agent path planning method based on QMIX algorithm, characterized in that: The following steps are involved: A multi-agent path planning model is established based on the QMIX algorithm and combined with the A* algorithm. The multi-agent path planning model is trained with the post-experience playback mechanism to obtain a trained multi-agent path planning model. Based on the trained multi-agent path planning model, multi-agent path planning is realized in dynamically changing complex multi-task scenarios.

2. The multi-agent path planning method based on the QMIX algorithm according to claim 1 is characterized in that: The steps of establishing a multi-agent path planning model based on the QMIX algorithm and in combination with the A* algorithm include: S1. Transform the path planning problem modeling of the intelligent agent into a distributed partially observable Markov decision model; S2. Based on the distributed partially observable Markov decision model, the QMIX algorithm is used to learn the global collaboration strategy of multiple agents, comprehensively analyze the global state information and task requirements, and optimize the task allocation and overall behavior decision; S3, based on task allocation, determine the action of each agent according to the principle of maximizing the joint Q value in the QMIX algorithm; when the agent encounters an obstacle or conflicts with other agents, use the A* algorithm to plan the path from the current node to the local target node; S4, loop through steps S2 and S3 until the task is completed.

3. The multi-agent path planning method based on the QMIX algorithm according to claim 2 is characterized in that: The method of converting the path planning problem modeling of the intelligent agent into a distributed partially observable Markov decision model is: S1.

1. Define the observation and action of the agent according to the physical characteristics and functional requirements of the agent, and model the global state of the agent; Local observation information of agent i at time t Contains the position vector of agent i Velocity Vector Target position vector g i and obstacle information Then the local observation information of agent i at time t is Among them, the position vector of agent i is The velocity vector of agent i The target position vector of agent i Obstacle information perceived by agent i Including the position of the nearest dynamic and static obstacles; establish a Cartesian coordinate system with the lower left corner as the origin, represents the position of agent i in the x-axis direction, represents the position of agent i on the y-axis, represents the velocity component of agent i in the x-axis direction, represents the velocity component of agent i in the y-axis direction, represents the position of the target position vector of agent i in the x-axis direction, Represents the position of the target position vector of agent i in the y-axis direction; The action chosen by agent i at time t For the discrete actions up, down, left, right, or stop, the moving distance of agent i depends on the time step and the applied acceleration; The positions and velocities of all agents, the current state of the environment, and the relevant parameters of the task are taken as the global state s; S1.

2. Based on the observations and actions of the defined intelligent agent and combined with the global state of the intelligent agent, a structured reward function is designed to obtain a distributed partially observable Markov decision model.

4. The multi-agent path planning method based on the QMIX algorithm according to claim 3 is characterized in that: The expression of the structured reward function is: Among them, K i ∈(0,1], i=1,2,3,4,5, K1 is the reward when agent i reaches the target point The weight parameter K2 is the penalty for agent i when it collides with or approaches an obstacle. The weight parameter, K3 is the path length penalty The weight parameter K4 is the penalty for invalid movement or delay of agent i The weight parameter K5 is the reward for agent i to maintain a safe distance from dynamic obstacles. The weight parameter of .

5. The multi-agent path planning method based on the QMIX algorithm according to claim 4 is characterized in that: The expression of the reward when agent i reaches the target point is: The penalty for agent i when it collides with or approaches an obstacle is expressed as: Among them, min_range represents the minimum distance from agent i to the nearest obstacle or other agents detected by agent i; The expression of path length penalty is: Among them, λ3 represents the path penalty constant, L path represents the path length of agent i; The penalty for ineffective movement or delay of agent i is expressed as: Where λ4 represents the penalty constant for invalid movement or unnecessary waiting time, t idle represents the unnecessary waiting time or invalid movement time of agent i; The expression of the reward for agent i to maintain a safe distance from a dynamic obstacle is: Where d represents the current distance between agent i and the obstacle, d safe represents the predefined safety distance, λ5, λ6 and α are adjustable parameters.

6. The multi-agent path planning method based on the QMIX algorithm according to any one of claims 2 to 5, characterized in that: The distributed partially observable Markov decision model uses seven-tuple Indicates, where I = {1, 2, ..., N} represents the set of agents, N is the number of agents; S represents the global state space, A i represents the action space of agent i, and the action a selected by agent i i ∈A i , at time t, combine the actions of all agents to form a joint action a=(a1, a2, ..., a N );O i The observation function O i (s,a i ) is the local observation space generated; P is the state transition probability; R represents the reward obtained when all agents take a joint action a under the global state information s; γ∈[0,1) represents the discount factor used to control the weight of future rewards.

7. The multi-agent path planning method based on the QMIX algorithm according to claim 2 is characterized in that: The method for determining the action of each agent based on the principle of maximizing the joint Q value in the QMIX algorithm is: The QMIX network in the QMIX algorithm consists of a hybrid network and a local Q network of each agent; wherein the structure of the local Q network includes an input layer-a first multilayer perceptron-a gated recurrent unit-a second multilayer perceptron connected in sequence, and a noise mechanism is introduced into the weights of the second multilayer perceptron of the local Q network; The agent receives local observation information and previous actions through the input layer of the local Q network and sends them to the first multi-layer perceptron. The first multi-layer perceptron performs feature fusion on the input data to generate a feature vector. The gated recurrent unit combines the historical trajectory information of the agent and the feature vector to update the trajectory representation of the agent. The second multi-layer perceptron serves as the output layer and maps the updated trajectory representation of the agent to the local Q value of the agent. The agent selects actions based on the local Q value using the ∈-greedy strategy. Since a noise mechanism is introduced into the weights of the second multilayer perceptron of the local Q network, the noise weight parameters and random noise will dynamically act on the weights of the second multilayer perceptron each time it propagates forward, making the weights of the second multilayer perceptron have random characteristics, thereby improving the exploration ability and diversity of action selection of the intelligent agent in the decision-making process; The hybrid network receives the local Q values ​​and global states of all agents as input, and generates a joint Q value Q through a feedforward neural network based on a nonlinear weighted hybrid method. tot , and ensure that the joint Q value Q tot Satisfy the monotonicity constraint Maximize the joint Q value Q tot This is equivalent to maximizing each local Q value Q i .

8. The multi-agent path planning method based on the QMIX algorithm according to claim 7, characterized in that: When the agent encounters an obstacle or conflicts with other agents, the method of using the A* algorithm to plan the path from the current node to the local target node is: If the distance between the agent and the obstacle or other agents is less than the safety threshold d safe , triggering local path replanning: taking the current position as the starting point and the local target node as the end point, combining the dynamic environment information to generate a set of candidate actions for the obstacle avoidance path The subsequent actions of the agent are only selected from the candidate action set A of the obstacle avoidance path. safe Select from.

9. The multi-agent path planning method based on the QMIX algorithm according to any one of claims 1, 2 or 7, characterized in that: The training of the multi-agent path planning model in combination with the post-experience playback mechanism includes the following steps: Ⅰ. Initialize parameters and define the initial experience replay pool: Initialize parameters include initializing scenes and agents, initializing network weight parameters of the QMIX algorithm, and setting the total number of iterations and the frequency of network parameter updates; Ⅱ. Path planning and action execution; Ⅲ. Experience storage and network update: After the agent performs an action, observe the new state of the environment feedback and the immediate reward, record the state transition samples, and store them in the experience replay pool; Calculate the timing error of samples in the experience replay pool and set the priority of the samples. When the capacity of the replay pool exceeds the threshold, remove the samples with the lowest priority; minimize the weighted loss function, update the network weight parameters and task allocation strategy; IV. Task completion and evaluation: Repeat steps II and III until all agents reach the target position or reach the maximum number of iterations, and perform performance evaluation on the multi-agent path planning model that has completed the update of the network weight parameters and task allocation strategy. If the evaluation result meets the preset goal, the multi-agent path planning model that has completed the update of the network weight parameters and task allocation strategy is a trained multi-agent path planning model.

10. The multi-agent path planning method based on the QMIX algorithm according to claim 1, characterized in that: The calculation formula of the timing error is: δ=|Q tot (s,a;θ)-(R+γmax a′ (Q tot (s′,a′;θ - )))| Set the sample priority p = |δ| + ∈, ∈ is a constant to prevent zero probability, and calculate the probability The weighted loss function value of the extracted batch samples, α∈[0,1] is the variable controlling uniformity; The expression of the weighted loss function is: in, represents the expected mean square error of multiple state-action pairs, is the experience replay pool, w i =(N·P(i)) -β is the importance sampling weight, β is the variable controlling the bias correction strength, y = r + γQ tot (s′,argmax a′ Q tot (s′,a′;θ);θ - ) represents the target Q value of the current time step, Q tot (s,a;θ) represents the estimation of joint action a in global state s when the network parameter is θ, Q tot (s′,a′;θ - ) indicates that the network parameters are θ - When , the valuation of the new joint action a′ in the new state s′.

Citation Information

Patent Citations

  • Multi-agent unknown environment search and rescue method and system based on deep reinforcement learning

    CN115330029A

Cited By

  • Heterogeneous agent path planning method

    CN120430486A

  • A path planning method for heterogeneous intelligent agents

    CN120430486B

  • DRL-based AUV path planning and obstacle avoidance method in partially observable environment

    CN121115021A

  • AUV path planning and obstacle avoidance method based on DRL in partially observable environment

    CN121115021B

  • Control method and system under combined transportation scene of steelmaking travelling crane and trolley

    CN121500840A