Multi-agent collision-free path planning method based on fusion DQN algorithm

By constructing a two-dimensional grid map and four-channel CNN feature extraction in multi-agent path planning, and combining the A* algorithm and the BC model trained by behavior cloning, the weights of the DQN model are dynamically adjusted, which solves the problems of high computational complexity and conflict in traditional algorithms in multi-agent scenarios, and achieves efficient and collision-free path planning.

CN120949778APending Publication Date: 2025-11-14CHANGZHOU UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511123106.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional path planning algorithms have high computational complexity and are prone to conflicts in multi-agent scenarios, making it difficult to guarantee path optimality. Furthermore, the traditional DQN algorithm suffers from slow convergence, low sample efficiency, and inability to effectively handle cooperation/competition relationships between agents.

Method used

A multi-agent collision-free path planning method based on the fusion DQN algorithm is adopted. By constructing a two-dimensional grid map and extracting features from a four-channel CNN, the A* algorithm is introduced to generate expert data. The BC model is trained by combining behavior cloning, and the fusion weights of the BC model and the DQN model are dynamically adjusted through an adaptive fusion framework. Finally, the CBS algorithm is used for conflict resolution.

Benefits of technology

It significantly improves the efficiency and quality of multi-agent path planning, ensures collision-free characteristics, enhances planning efficiency and path quality, reduces computation time and energy consumption, and strengthens the collaborative and competitive processing capabilities among agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120949778A_ABST
    Figure CN120949778A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of agent path planning, in particular to a multi-agent collision-free path planning method based on a fusion DQN algorithm. The multi-agent collision-free path planning method comprises the following steps: firstly, constructing a two-dimensional grid map as an environment, and carrying out feature extraction by utilizing a CNN (Convolutional Neural Network); then, behavior clone learning is carried out through the expert model to obtain a BC model; the core innovation lies in that a BC model and a CNNDQN model are fused, an adaptive strategy learning framework is constructed, and intelligent dynamic combination of expert experience and reinforcement learning exploration is realized by adopting uncertainty estimation, antagonistic knowledge distillation and performance perception sampling technologies; and finally, further processing an initial path output by the fusion model by a CBS algorithm, and completing multi-agent collision-free path planning. According to the method, the accuracy and efficiency of path planning are optimized through a mixed learning strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent agent path planning technology, and in particular to a multi-agent collision-free path planning method based on the fusion DQN algorithm. Background Technology

[0002] In warehousing and logistics environments, collision-free path planning for multiple agents is crucial for improving efficiency and reducing costs. However, traditional path planning algorithms often face challenges such as long computation times and poor planning results when dealing with collision-free path planning for multiple agents. This not only leads to a waste of time and energy but may also cause system congestion, a surge in energy consumption, and even the risk of agent collisions due to path conflicts or invalid moves. Especially when the number of agents increases or the map size expands, the computational complexity often increases explosively, and it is difficult to guarantee the optimality of the final path.

[0003] With the development of artificial intelligence, reinforcement learning (RL) and deep learning (DL) methods have been introduced into the field of path planning. Among them, the Deep Q-Network (DQN) algorithm combines the decision-making ability of reinforcement learning with the perception ability of deep learning, effectively alleviating the "curse of dimensionality" problem in complex environments. It has shown good results and has been applied in single-agent path planning. However, the traditional DQN algorithm faces significant challenges in multi-agent scenarios: slow convergence speed and low sample efficiency in the early stages of training, and its architecture design is difficult to directly adapt to the cooperative and competitive relationships between agents. Summary of the Invention

[0004] The technical problem to be solved by this invention is: in order to address the issues of high computational complexity, easy conflict generation, and difficulty in guaranteeing path optimality in multi-agent scenarios by traditional path planning algorithms in the background art; and the problems of slow convergence, low sample efficiency, and inability to effectively handle cooperation / competition relationships between agents by traditional DQN algorithms, this invention provides a multi-agent collision-free path planning method based on a fusion DQN algorithm.

[0005] The technical solution adopted by this invention to solve its technical problem is: a multi-agent collision-free path planning method based on the fusion DQN algorithm, comprising the following steps: S1. Construct a two-dimensional grid map for path planning. The two-dimensional grid map includes a starting grid, an ending grid, obstacle grids, and feasible grids, and the starting grid and the ending grid are connected. S2. Based on the two-dimensional grid map, a CNN neural network is constructed to extract four-channel features from the two-dimensional grid map. The four channels represent the starting position, the ending position, the obstacle area, and the feasible area, respectively. S3. Introduce expert data generated by the A* algorithm, and train the BC model through behavior cloning to learn expert strategies; S4. The BC model and the CNNDQN model are dynamically fused through an adaptive fusion framework, which includes an uncertainty estimator, a fusion weight network and a discriminator network. The model decision confidence is quantified by the Monte Carlo Dropout method, and the fusion weights of the BC model and the DQN model are dynamically adjusted. S5. Use the fused DQN model to perform preliminary path planning for multiple agents, pass the preliminary path to the CBS algorithm for conflict resolution, and finally generate a collision-free multi-agent path.

[0006] Environmental modeling using 2D grid maps enables a digital representation of complex warehouse environments; four-channel CNN feature extraction allows agents to comprehensively perceive key environmental elements (start point, end point, obstacles, and feasible areas); the introduction of A* expert policies and the establishment of BC models through behavior cloning provide high-quality initial strategies for reinforcement learning; an innovative adaptive fusion framework effectively solves the problem of slow convergence in the early stages of traditional DQN training through a dynamic weight adjustment mechanism; finally, the CBS algorithm is combined for conflict resolution, ensuring the collision-free characteristics of multi-agent paths; this phased processing method significantly improves planning efficiency and path quality.

[0007] According to an embodiment of the present invention, the CNN neural network in step S2 includes multiple convolutional layers, batch normalization layers, nonlinear activation functions and pooling layers, used to extract spatial features of a two-dimensional grid map, and output a feature map of fixed size through adaptive pooling; the extracted feature map is flattened and then enters a DQN decision module consisting of two fully connected layers, and a Dropout mechanism is introduced, the output of which is the Q value corresponding to the action space, used to guide the agent to select the optimal action.

[0008] By combining convolutional layers and batch normalization layers, efficient extraction of map spatial features is achieved; nonlinear activation functions enhance the network's expressive power; the use of pooling layers and adaptive pooling preserves key feature information while achieving uniform feature map size; the introduction of the Dropout mechanism effectively prevents overfitting; this CNN neural network enables the agent to accurately understand the environmental structure, providing a reliable feature representation basis for subsequent decision-making.

[0009] According to one embodiment of the present invention, in step S2, the extracted features are optimized using a reward function, including arrival reward, penalty reward, dwell penalty, and distance reward, calculated using the following formula: ; In the formula, The Manhattan distance between the agent's previous position and the destination. This represents the Manhattan distance between the current location of the agent and its destination.

[0010] The arrival reward (+5.0) reinforces goal-oriented behavior, the collision penalty (-0.5) avoids dangerous actions, the lingering penalty suppresses ineffective wandering, and the dynamic reward based on Manhattan distance (±0.01) enables fine-grained path optimization. The multi-level reward pattern significantly accelerates model convergence and enables the agent to learn a movement strategy that is both efficient and safe.

[0011] According to one embodiment of the present invention, in the initial training phase of the adaptive fusion framework dynamic fusion in step S4, the exploration rate ε is optimized. The initial exploration rate is set to 0.3, and gradually decreases to a minimum value of 0.01 as the number of training rounds increases; the decrease formula is: , In the formula, The exploration rate after decay. To achieve the minimum exploration rate value, Based on the current exploration rate, is the attenuation coefficient, and t indicates starting from the t-th episode.

[0012] The initial high exploration rate (0.3) ensures sufficient exploration in the early stages of training; as training progresses, the exploration rate gradually decreases to the minimum value (0.01), making the strategy tend to stabilize; this avoids the inefficiency of completely random exploration and prevents the algorithm from getting trapped in local optima too early, which is an important guarantee for the algorithm to continuously optimize.

[0013] According to an embodiment of the present invention, the adaptive fusion framework in step S4 achieves dynamic fusion through the following steps: S41. Output the Q-value of the current state using both the BC model and the DQN model. and , S42. Using two independent uncertainty estimators and The prediction variances of the two models are calculated through multiple forward propagations to quantify the decision confidence. Specifically, the uncertainties of the BC model and the DQN model in the current state are as follows: ; ; In the two formulas above, Let Q be the Q-value of the agent in the BC model at time t in state s. Let Q be the Q-value of the agent in the DQN model at time t in state s; S43. Based on the Q-value from step S41 and the uncertainty information from step S42, the weights of the BC model and the DQN model are dynamically allocated through a fusion weight network to generate the fused Q-value; the model weights are specifically as follows: .

[0014] By employing parallel inference with both BC and DQN models, the algorithm fully leverages expert knowledge and self-learning capabilities. The uncertainty estimation achieved through the Monte Carlo Dropout method provides an objective basis for model weight allocation. Dynamic weight calculation based on Q-values ​​and uncertainties ensures that greater decision weights are assigned when the model has high confidence, significantly improving the robustness and adaptability of the algorithm.

[0015] According to one embodiment of the present invention, a priority experience replay strategy is further included between step S4 and step S5, specifically: priority experience replay is implemented using a SumTree structure based on TD error; three performance-aware sub-buffers are maintained to store BC model advantage samples, DQN model advantage samples and high uncertainty samples respectively; an 80%-20% mixed sampling strategy is adopted, with 80% of the samples coming from priority experience replay and 20% of the samples being uniformly sampled from the three performance-aware sub-buffers.

[0016] The SumTree structure based on TD error ensures priority learning of important experiences. The separate storage and sampling of three special types of samples specifically enhances the learning of key scenarios. The 80%-20% mixed sampling strategy balances the relationship between focused learning and comprehensive learning, significantly improving sample utilization efficiency.

[0017] According to one embodiment of the present invention, the optimization of the DQN model in step S5 is achieved through a triple loss function, including: S51, TD error loss based on fused Q-value; where... Total loss function of the fusion module The calculation formula is: In the formula, The target weights are calculated based on performance-uncertainty guidance. This represents the uncertainty regularization strength coefficient. and These are the predicted uncertainties for BC and DQN, respectively. and The true uncertainty is calculated using Monte Carlo Dropout; S52 and KL divergence constraint loss are used to maintain consistency with the output distribution of the BC model; S53, Adversarial loss, improves the robustness of DQN output through the discriminator network; where the total loss function of the discriminator... The calculation formula is: ; In the formula, For the discriminator network, The binary cross-entropy loss function is... The output quality of the DQN model can be improved through adversarial training. The specific calculation formula is as follows: In the formula, Let be the regularization mildness coefficient of the KL divergence. For the Softmax function; Therefore, the total loss function of the DQN model The calculation formula is: ; In the formula, For time-series difference loss based on fusion Q-value; As a constraint for knowledge distillation The regularization strength coefficient to counteract loss.

[0018] The TD error loss ensures the accuracy of Q-value learning; the KL divergence constraint maintains consistency with the expert policy; the adversarial loss improves the robustness of decision-making; this composite loss function design retains the exploratory ability of reinforcement learning and inherits the reliability of the expert policy.

[0019] According to one embodiment of the present invention, the KL divergence constraint loss in step S52 adopts a Softmax variant with temperature coefficient adjustment, specifically: applying a temperature parameter T, T>1, to the output Q values ​​of the BC model and the DQN model to soften the probability distribution and enhance the knowledge transfer effect; dynamically adjusting the temperature coefficient, setting T=5 in the early stage of training to smooth the distribution difference, and gradually reducing it to T=1 in the later stage to strengthen the strategy focus; wherein, the modified KL divergence loss function is: ; In the formula, For the Softmax function, and The outputs are the Q-values ​​for the BC model and the DQN model, respectively.

[0020] The higher temperature parameter (T=5) in the early stage of training smoothed out the differences in policy distribution and reduced the learning difficulty; as training progressed, the temperature was gradually reduced, which allowed the policies to gradually focus; this effectively alleviated the problem of policy distribution mismatch and made knowledge transfer more stable and efficient.

[0021] According to an embodiment of the present invention, a dynamic priority mechanism is introduced in the CBS short-term conflict resolution stage in step S5. Specifically, the path planning priority is dynamically adjusted according to the Manhattan distance between the agent's current position and the target point, with higher priority for closer distances. A path replanning strategy is adopted for high-priority agents, while a waiting or local detour strategy is adopted for low-priority agents. The feasibility of the path is verified through a time window conflict detection algorithm to ensure a deadlock-free solution.

[0022] Prioritization based on Manhattan distance aligns with path planning intuition; differentiated processing strategies (replanning / waiting / detour) achieve rational resource allocation; time window detection ensures the feasibility of the solution; and significantly improves the coordination efficiency of multi-agent systems.

[0023] According to one embodiment of the present invention, each square of the two-dimensional grid map is the same size and evenly distributed, and there is at least one feasible path between the starting square and the ending square in the map.

[0024] The uniformly sized grid simplifies environment representation and computation, the clear definition of the four types of grids standardizes environment modeling, and the connectivity requirement ensures the solvability of the problem; these specifications provide clear input standards for algorithm implementation and are the basic guarantee for the reliable operation of the system.

[0025] The beneficial effects of this invention are: 1. A convolutional neural network (CNN) is used to replace the traditional fully connected structure. The two-dimensional grid map is encoded into four input channels, representing the start point, end point, obstacles, and feasible region, respectively. This enables the agent to fully perceive the environmental structure and improves the ability to extract map features. Batch normalization is introduced to accelerate training convergence and stabilize network performance. Pooling layers are introduced to enhance the robustness of feature representation. At the same time, the Dropout regularization mechanism is combined to effectively suppress the risk of overfitting and improve the model's generalization ability in complex environments, thereby achieving better path planning results. 2. Dynamically and adaptively fusing the BC model and the CNNDQN model accelerates the convergence speed and avoids useless exploration by the agent; 3. An adaptive fusion strategy is adopted to dynamically integrate the output Q-values ​​of the BC model and the CNNDQN model: A dedicated fusion network is trained to calculate the optimal weight allocation of the two models in real time. This network takes the Q-value outputs of the two models and their corresponding uncertainty estimates as input features, and outputs the dynamic weight coefficients of the BC model, thus obtaining the fused Q-value. In the early stages of training, when the BC model exhibits high decision certainty and good performance, the fusion network adaptively increases the weight of the BC model, making full use of the mature experience of the BC model to avoid ineffective exploration. As the CNNDQN model is continuously optimized and its performance improves during training, the fusion network can intelligently identify which model is more advantageous in the current state and dynamically adjust the weight allocation, so that the CNNDQN model gradually takes the lead. This adaptive fusion mechanism based on performance awareness and uncertainty estimation not only ensures the stability in the early stages of training, but also enables the CNNDQN model to further explore path optimization based on the BC model in the later stages. 4. The loss function was improved in multiple layers: First, a KL divergence regularization term was introduced to ensure that the output distribution of the CNNDQN model is consistent with that of the BC model, thus ensuring the effective inheritance of expert knowledge during training. Second, an adversarial loss component was added to guide the CNNDQN model to generate more robust decision outputs using the discriminator network, thereby improving the model's generalization ability. At the same time, importance sampling weights were used to weight the standard DQN loss, effectively mitigating the distribution bias problem introduced by the priority experience replay mechanism. In addition, uncertainty regularization loss was integrated to improve the model's ability to accurately estimate decision confidence. Through this multi-constraint composite loss function design, the model maintains its continuous learning ability while effectively preventing the forgetting of expert knowledge, achieving an organic unity of knowledge preservation and performance optimization. Attached Figure Description

[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0027] Figure 1 This is a flowchart of the present invention.

[0028] Figure 2 This is a starting grid map of size 40×44 and containing 20 intelligent agents in a specific embodiment of the present invention.

[0029] Figure 3 This is a grid map of size 40×44, representing the arrival of 20 intelligent agents at the destination, according to a specific embodiment of the present invention.

[0030] Figure 4 This is a schematic diagram of the specific process of the present invention.

[0031] Figure 5 This is a network structure diagram of the CNNDQN model in this invention.

[0032] Figure 6 This is a network structure diagram of the UncertaintyEstimator in this invention.

[0033] Figure 7 This is a network structure diagram of the Discriminator in this invention. Detailed Implementation

[0034] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0035] like Figure 1 As shown, a multi-agent collision-free path planning method based on the fusion DQN algorithm includes the following steps: Step 1: Construct a two-dimensional grid map for path planning. The two-dimensional grid map contains a start grid, an end grid, obstacle grids, and feasible grids, and the start grid and the end grid are connected. Step 2: Based on the two-dimensional grid map, construct a CNN neural network to extract four-channel features from the two-dimensional grid map. The four channels represent the starting position, ending position, obstacle area, and feasible area, respectively. Step 3: Introduce expert data generated by the A* algorithm, train the BC model through behavior cloning, and learn expert strategies; Step 4: Dynamically fuse the BC model and the CNNDQN model using an adaptive fusion framework. The adaptive fusion framework includes an uncertainty estimator, a fusion weight network, and a discriminator network. The model decision confidence is quantified using the Monte Carlo Dropout method, and the fusion weights of the BC model and the DQN model are dynamically adjusted. Step 5: Use the fused DQN model to perform preliminary path planning for the multi-agent team, pass the preliminary path to the CBS algorithm for conflict resolution, and finally generate a collision-free multi-agent path.

[0036] like Figure 2 and Figure 3 As shown, each cell in the two-dimensional grid map is the same size and evenly distributed, and there is at least one feasible path between the starting cell and the ending cell. Specifically, each cell in the two-dimensional grid map is the same size and evenly distributed, gray cells represent obstacles, and white cells represent passable areas. Figure 2 The initial positions of the 20 agents are shown. Figure 3 The corresponding target location is displayed.

[0037] Among them, the path planning of agent 1: starting point( Figure 2 ): Position (19, 12), End Point ( Figure 3 ): Position (30, 14). Agent 1 needs to move from the upper left through multiple shelf aisles, avoiding obstacles, and to the lower right. The adaptive fusion model of this invention generates an initial path for it, with a path length of approximately 35 steps.

[0038] Path planning for Agent 12: Starting point ( Figure 2 : Position (25, 3), Endpoint ( Figure 3): Position (37, 23). Agent 12 needs to make long-distance diagonal movements, bypassing multiple shelf obstacles. The fusion model finds a near-optimal path by combining the experience knowledge of BC and the exploration capabilities of DQN. When 20 agents plan paths simultaneously, the initial paths generated by the path planning method in this embodiment may have conflicts. These initial paths are then fed into the CBS algorithm for conflict resolution, ultimately achieving 100% collision-free multi-agent cooperative navigation.

[0039] like Figure 4 As shown, the method for constructing a preliminary DQN path planning model first includes: combining deep neural networks with the Q-learning algorithm, using the neural network as a function approximation to replace the Q-value table, and calculating the value function. The neural network serves as the carrier of the state-action value function, and the state-action value function is approximated by iteratively updating the f-parameter θ of the neural network, defined as: ,in, Let represent an approximate substitution function, where the Q-value is replaced by the output of a neural network, s represents the current state, and a represents the action.

[0040] The state function, action function, and reward function of the DQN path planning model are initialized. Each cell in the grid map corresponds to a state. The agent interacts with the environment to obtain the current state s. Then, the BC model and the CNNDQN model are used to select action a according to the set fusion weights. There is a probability ε of selecting the optimal action and a probability of 1-ε of selecting a random action to enhance the exploration ability. After execution, a new state is obtained. The agent's state-behavior sequence (st, at, rt, st+1) generated during exploration is stored in the experience pool along with a reward value r. Action a includes four different actions: up, down, left, and right, with each step moving only one unit distance (one square). The extracted features are then optimized using a reward function, with the reward value calculated as follows: ; In the formula, The Manhattan distance between the agent's previous position and the destination. This represents the Manhattan distance between the current location of the agent and its destination.

[0041] In other words, when the agent reaches the target point, it receives a high reward of +5.0; when it collides with an obstacle or moves illegally, it receives a reward of -0.5; when the agent stays in place or takes a longer route, it receives a larger penalty for being stuck. The reward is based on the Manhattan distance between the agent's previous step and the current distance to the destination. When the distance is shortened (closer to the destination), it receives a reward of +0.01; when the distance is lengthened (farther from the destination), it receives a reward of -0.01.

[0042] Secondly, combining CNN neural networks with Q-learning algorithms, such as... Figure 5 As shown, a CNN neural network is used to extract features from a two-dimensional grid map, where the map is encoded as a four-channel image: the first channel represents the starting position, the second channel represents the ending position, the third channel marks obstacle regions, and the fourth channel represents feasible regions. This four-channel input is processed through multiple convolutional layers, batch normalization layers, and nonlinear activation functions to extract spatial features. Pooling layers are used to progressively reduce the spatial resolution to enhance the receptive field, and finally, adaptive pooling outputs a fixed-size feature map. The extracted feature map is flattened and then fed into a DQN decision module consisting of two fully connected layers, which incorporates a Dropout mechanism to improve generalization ability. The output is the Q-value corresponding to the action space, used to guide the agent in selecting the optimal action. The parameters of each layer of the entire network are initialized using the Kaiming method to ensure training stability and convergence speed under the ReLU activation function.

[0043] Then, the model trained by the behavior clone is subjected to reinforcement learning fusion training. The expert model and the untrained DQN model are dynamically fused through a multi-level adaptive fusion framework. This framework solves problems such as the policy distribution gap, the blind spot of dynamic weight allocation, the lack of decision confidence quantification, and the mismatch of heterogeneous learning progress in the fusion of BC model and DQN model through the cascade mechanism of uncertainty perception, adaptive weight calculation, and adversarial optimization.

[0044] Specifically, the adaptive fusion framework achieves dynamic fusion through the following steps: outputting the Q-value of the current state through the BC model and the DQN model respectively, and outputting... and ;like Figure 6 As shown, two independent uncertainty estimators are used. and The prediction variances of the two models are calculated through multiple forward propagations to quantify the decision confidence. A smaller variance indicates higher confidence and model reliability, thus increasing the proportion of that model. Conversely, a larger variance indicates lower confidence and model instability, thus decreasing the proportion of that model. Specifically, the uncertainties of the BC model and the DQN model in the current state are as follows: ; ; In the two formulas above, Let Q be the Q-value of the agent in the BC model at time t in state s. Let Q be the Q-value of the agent in the DQN model at time t in state s; based on the Q-value in step S41 and the uncertainty information in step S42, the weights of the BC model and the DQN model are dynamically allocated through a fusion weight network to generate the fused Q-value; the specific model weights are as follows: .

[0045] In the initial training phase of the adaptive fusion framework's dynamic fusion, the exploration rate ε is optimized. The initial exploration rate is set to 0.3, and it gradually decays to a minimum value of 0.01 as the number of training epochs increases. The decay formula is: , In the formula, The exploration rate after decay. To achieve the minimum exploration rate value, Based on the current exploration rate, is the attenuation coefficient, and t indicates starting from the t-th episode.

[0046] A multi-dimensional performance-aware adaptive priority experience replay strategy can construct a two-layer experience management architecture that integrates TD error-driven and scenario-specific learning. Specifically, the priority experience replay strategy maintains a global priority ranking based on TD errors through a SumTree structure, while introducing a three-dimensional performance-aware classification mechanism: dynamically identifying and independently storing feature experience samples from scenarios where the BC model excels (bc_better), scenarios where the DQN model excels (dqn_better), and decision-difficult scenarios (uncertain_cases) filtered based on an uncertainty threshold. The buffer implements an adaptive hybrid sampling strategy, employing an 80%-20% dynamic allocation mechanism: the main sampling channel selects high-value transfer experiences from the TD error priority queue based on importance weights, while the auxiliary sampling channel extracts scenario-specific samples evenly from the three performance-aware subdomains, ensuring that the model can obtain reinforcement signals from key state transitions and extract targeted strategy knowledge from successful cases of multi-model collaborative decision-making. This mechanism achieves intelligent classification and adaptive weight allocation of experience samples by real-time monitoring of the matching degree between reward feedback and model prediction, combined with uncertainty quantification indicators. Compared with the traditional priority experience replay mechanism based solely on TD error, this strategy can specifically retain and repeatedly learn typical scenarios of BC model success, DQN model success, and decision difficulty, avoiding the dilution of different types of key experiences in a unified buffer and improving the learning efficiency of the fusion model for their respective advantageous regions.

[0047] Finally, a multi-objective joint optimization loss function is constructed. The DQN model is trained through a triple loss mechanism: the standard TD error loss ensures the accuracy of Q-value learning, the KL divergence constraint loss maintains consistency with the probability distribution of the BC model, and the adversarial loss makes the DQN output closer to the characteristics of expert policies. Compared with the single learning objective of traditional DQN that relies solely on environmental reward feedback, this multi-constraint mechanism effectively avoids blind exploration in the early stages of training, maintains the continuous inheritance of expert knowledge while ensuring reinforcement learning capabilities, and improves the quality of policy representation through adversarial game mechanisms.

[0048] Among them, the total loss function of the fusion module The calculation formula is: In the formula, The target weights are calculated based on performance-uncertainty guidance. This represents the uncertainty regularization strength coefficient. and These are the predicted uncertainties for BC and DQN, respectively. and The true uncertainty is calculated using Monte Carlo Dropout; KL divergence constraint loss is used to maintain consistency with the output distribution of the BC model; the KL divergence constraint loss employs a temperature-coefficient-adjusted Softmax variant, specifically: applying a temperature parameter T, T>1, to the output Q values ​​of both the BC and DQN models to soften the probability distribution and enhance knowledge transfer; the temperature coefficient is dynamically adjusted, initially set to T=5 to smooth distribution differences, and gradually reduced to T=1 later to strengthen policy focus; the modified KL divergence loss function is: ; In the formula, For the Softmax function, and Output the Q-values ​​for the BC model and the DQN model, respectively; like Figure 7 As shown, adversarial loss is used to improve the robustness of the DQN output through the discriminator network; where the total loss function of the discriminator is... The calculation formula is: ; In the formula, For the discriminator network, The binary cross-entropy loss function is... The output quality of the DQN model can be improved through adversarial training. The specific calculation formula is as follows: In the formula, Let be the regularization mildness coefficient of the KL divergence. For the Softmax function; Therefore, the total loss function of the DQN model The calculation formula is: ; In the formula, For time-series difference loss based on fusion Q-value; As a constraint for knowledge distillation The regularization strength coefficient to counteract loss.

[0049] As a preferred approach, the CBS short-term conflict resolution phase introduces a dynamic priority mechanism, specifically: the path planning priority is dynamically adjusted based on the Manhattan distance between the agent's current position and the target point, with higher priority for closer distances; a path replanning strategy is adopted for high-priority agents, while a waiting or partial detour strategy is adopted for low-priority agents; and the feasibility of the path is verified through a time window conflict detection algorithm to ensure a deadlock-free solution.

[0050] This embodiment of the multi-agent collision-free path planning method based on the fused DQN algorithm first learns the expert decision-making strategy of the A* algorithm through behavior cloning technology to establish a BC baseline model to obtain high-quality initial policy knowledge. Then, an adaptive fusion module is designed to fuse the BC model and the CNNDQN model. This module includes an uncertainty estimator, a fusion weight network, a discriminator network, and a priority experience replay buffer. The uncertainty estimator adopts the Monte Carlo Dropout method, calculates the prediction variance through multiple forward propagation sampling, and quantifies the decision confidence of the BC model and the CNNDQN model in the current state. The fusion weight network receives the Q-value vector and uncertainty information as input, and learns the optimal weight allocation strategy related to the state through a multi-layer neural network. The discriminator network is trained adversarially with the CNNDQN model based on the GAN framework, and distinguishes between BC output and DQN output through a binary classification task, prompting the CNNDQN model to learn the expert strategy. The priority experience replay buffer adopts a priority sampling mechanism based on TD error, and classifies and stores experience samples according to BC advantage, DQN advantage, and high uncertainty scenarios. During training, the fusion module dynamically calculates the fusion weights of the BC model and the CNNDQN model based on the uncertainty assessment of the current state and historical performance statistics. When a model exhibits low uncertainty and good decision-making performance in a specific scenario, its weight contribution in the final decision is increased accordingly. By jointly optimizing a multi-objective function that includes reinforcement learning loss, knowledge distillation loss, and adversarial loss, the BC model and the CNNDQN model are effectively combined, ultimately obtaining a multi-agent path planning model with good convergence performance and decision quality.

[0051] Table 1 lists the training parameter configurations for behavior cloning using the A* expert model in this embodiment of the invention. These parameters were experimentally optimized for training the agent's behavior cloning model in a warehouse environment (40×44 grid map).

[0052] Table 1 Behavior Cloning Parameter Settings

[0053] After the behavior clone model in Table 1 is trained, it serves as the BC baseline model for fusion reinforcement learning in Table 2. Through an adaptive fusion mechanism, it is combined with the DQN model to further improve path planning performance.

[0054] Table 2 lists the training parameter configurations of the adaptive fusion model in this embodiment of the invention. These parameters have been experimentally optimized and are used to train a path planning model of 20 agents in a warehouse environment (40×44 grid map).

[0055] Table 2. Parameter Settings for Fusion Reinforcement Learning

[0056] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A multi-agent collision-free path planning method based on a fusion DQN algorithm, characterized in that, Includes the following steps: S1. Construct a two-dimensional grid map for path planning. The two-dimensional grid map includes a starting grid, an ending grid, obstacle grids, and feasible grids, and the starting grid and the ending grid are connected. S2. Based on the two-dimensional grid map, a CNN neural network is constructed to extract four-channel features from the two-dimensional grid map. The four channels represent the starting position, the ending position, the obstacle area, and the feasible area, respectively. S3. Introduce expert data generated by the A* algorithm, and train the BC model through behavior cloning to learn expert strategies; S4. The BC model and the CNNDQN model are dynamically fused through an adaptive fusion framework, which includes an uncertainty estimator, a fusion weight network and a discriminator network. The model decision confidence is quantified by the Monte Carlo Dropout method, and the fusion weights of the BC model and the DQN model are dynamically adjusted. S5. Use the fused DQN model to perform preliminary path planning for multiple agents, pass the preliminary path to the CBS algorithm for conflict resolution, and finally generate a collision-free multi-agent path.

2. The multi-agent collision-free path planning method based on the fused DQN algorithm according to claim 1, characterized in that: The CNN neural network in step S2 includes multiple convolutional layers, batch normalization layers, nonlinear activation functions, and pooling layers, which are used to extract spatial features of the two-dimensional grid map and output a fixed-size feature map through adaptive pooling. The extracted feature map is flattened and then enters the DQN decision module consisting of two fully connected layers, and the Dropout mechanism is introduced. The output is the Q value corresponding to the action space, which is used to guide the agent to select the optimal action.

3. The multi-agent collision-free path planning method based on the fused DQN algorithm according to claim 1, characterized in that: In step S2, the extracted features are optimized using a reward function, including arrival reward, penalty reward, dwell penalty, and distance reward. The calculation formula is as follows: ; In the formula, The Manhattan distance between the agent's previous position and the destination. This represents the Manhattan distance between the current location of the agent and its destination.

4. The multi-agent collision-free path planning method based on the fused DQN algorithm according to claim 1, characterized in that: In step S4, during the initial training phase of the adaptive fusion framework's dynamic fusion, the exploration rate ε is optimized. The initial exploration rate is set to 0.3, and it gradually decays to a minimum value of 0.01 as the number of training epochs increases. The decay formula is: , In the formula, The exploration rate after decay. To achieve the minimum exploration rate value, Based on the current exploration rate, is the attenuation coefficient, and t indicates starting from the t-th episode.

5. The multi-agent collision-free path planning method based on the fused DQN algorithm according to claim 1, characterized in that: The adaptive fusion framework in step S4 achieves dynamic fusion through the following steps: S41. Output the Q-value of the current state using both the BC model and the DQN model. and , S42. Using two independent uncertainty estimators and The prediction variances of the two models are calculated through multiple forward propagations to quantify the decision confidence. Specifically, the uncertainties of the BC model and the DQN model in the current state are as follows: ; ; In the two formulas above, Let Q be the Q-value of the agent in the BC model at time t in state s. Let Q be the Q-value of the agent in the DQN model at time t in state s; S43. Based on the Q-value from step S41 and the uncertainty information from step S42, the weights of the BC model and the DQN model are dynamically allocated through a fusion weight network to generate the fused Q-value; the model weights are specifically as follows: 。 6. The multi-agent collision-free path planning method based on the fused DQN algorithm according to claim 1, characterized in that: Between steps S4 and S5, a priority experience replay strategy is also included, specifically: a SumTree structure based on TD error is used to implement priority experience replay; three performance-aware sub-buffers are maintained to store BC model advantage samples, DQN model advantage samples and high uncertainty samples respectively; an 80%-20% mixed sampling strategy is adopted, with 80% of the samples coming from priority experience replay and 20% of the samples being uniformly sampled from the three performance-aware sub-buffers.

7. The multi-agent collision-free path planning method based on the fused DQN algorithm according to claim 1, characterized in that: The optimization of the DQN model in step S5 is achieved through a triple loss function, including: S51, TD error loss based on fused Q-value; where... Total loss function of the fusion module The calculation formula is: In the formula, The target weights are calculated based on performance-uncertainty guidance. This represents the uncertainty regularization strength coefficient. and These are the predicted uncertainties for BC and DQN, respectively. and The true uncertainty is calculated using Monte Carlo Dropout; S52 and KL divergence constraint loss are used to maintain consistency with the output distribution of the BC model; S53, Adversarial loss, improves the robustness of DQN output through the discriminator network; where the total loss function of the discriminator... The calculation formula is: ; In the formula, For the discriminator network, The binary cross-entropy loss function is... The output quality of the DQN model can be improved through adversarial training. The specific calculation formula is as follows: In the formula, Let be the regularization mildness coefficient of the KL divergence. For the Softmax function; Therefore, the total loss function of the DQN model The calculation formula is: ; In the formula, For time-series difference loss based on fusion Q-value; As a constraint for knowledge distillation The regularization strength coefficient to counteract loss.

8. The multi-agent collision-free path planning method based on the fused DQN algorithm according to claim 7, characterized in that: In step S52, the KL divergence constraint loss employs a temperature-coefficient-adjusted Softmax variant. Specifically, a temperature parameter T (T>1) is applied to the output Q values ​​of the BC and DQN models to soften the probability distribution and enhance knowledge transfer. The temperature coefficient is dynamically adjusted, initially set to T=5 to smooth distribution differences, and then gradually reduced to T=1 to strengthen policy focus. The modified KL divergence loss function is as follows: ; In the formula, For the Softmax function, and The outputs are the Q-values ​​for the BC model and the DQN model, respectively.

9. The multi-agent collision-free path planning method based on the fused DQN algorithm according to claim 1, characterized in that: In step S5, the CBS short-term conflict resolution phase introduces a dynamic priority mechanism, which is as follows: the path planning priority is dynamically adjusted according to the Manhattan distance between the agent's current position and the target point, with higher priority for closer distances; a path replanning strategy is adopted for high-priority agents, while a waiting or local detour strategy is adopted for low-priority agents; and the feasibility of the path is verified through a time window conflict detection algorithm to ensure a deadlock-free solution.

10. The multi-agent collision-free path planning method based on the fused DQN algorithm according to claim 1, characterized in that: Each cell in the two-dimensional grid map is the same size and evenly distributed, and there is at least one feasible path between the starting cell and the ending cell in the map.

Citation Information

Cited By

  • Intelligent agent learning training method and device, computer equipment and storage medium

    CN121390197A

  • Agent learning training method and device, computer equipment and storage medium

    CN121390197B

  • Low-altitude traffic flow collaborative awareness and conflict prediction method and system based on multi-modal large model

    CN121725677A

  • Low-altitude traffic flow cooperative perception and conflict prediction method and system based on multi-modal large model

    CN121725677B