Multi-agent dynamic task allocation and collaborative path-finding system for label-free distributed deep reinforcement learning
Through label-free distributed deep reinforcement learning and inter-agent state exchange mechanism, adaptive task allocation and collaborative pathfinding of multi-agent systems are realized, which solves the adaptive problems of task allocation and path planning in dynamic environments, and improves system efficiency and resource utilization.
Patent Information
- Application Number
- CN202510437220.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multi-agent system lacks adaptability in task allocation and path planning in dynamic environments, relies on manual annotation data, and task allocation and path planning are independently processed, resulting in inefficiency of the system.
Label-free distributed deep reinforcement learning and state exchange mechanism between agents are adopted, and through multi-source heterogeneous information fusion and adaptive matching, decentralized task allocation and collaborative pathfinding between agents are realized, reducing dependence on manual labeled data, and improving system adaptability and resource utilization.
Improve task allocation efficiency, enhance system adaptability, optimize energy efficiency, reduce communication overhead, improve coordination efficiency, reduce path conflicts, and improve system throughput.
Smart Images

Figure CN120373826A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and multi-agent systems, and in particular to a multi-agent dynamic task allocation and collaborative pathfinding system for unlabeled distributed deep reinforcement learning. Background Art
[0002] With the rapid development of artificial intelligence technology, multi-agent systems are increasingly widely used in fields such as logistics distribution, UAV formations, intelligent warehousing, and autonomous vehicle scheduling. A multi-agent system is a complex system composed of multiple intelligent agents with autonomous decision-making capabilities. These intelligent agents can perceive the environment, interact with the environment and other intelligent agents, and make decisions based on the information obtained. In actual application scenarios, multi-agent systems face challenges such as dynamically changing tasks, limited resources, and complex and ever-changing environments, and effective cooperation among intelligent agents is required to achieve common goals.
[0003] In terms of multi-agent dynamic task allocation, traditional methods are mainly divided into three categories: centralized methods, distributed methods, and hybrid methods. The centralized method uniformly plans the task allocation and path planning of all intelligent agents by a central controller, and can obtain the global optimal solution. However, as the number of intelligent agents and task complexity increase, the computational complexity increases exponentially, making it difficult to meet the real-time requirements. The distributed method allows intelligent agents to make decisions based on local information, with small communication overhead and strong scalability, but it is difficult to guarantee global optimality. The hybrid method combines the advantages of the first two methods, but still faces challenges in adaptability in dynamic environments.
[0004] At the same time, the multi-agent collaborative pathfinding problem usually involves how to make multiple intelligent agents reach their respective target points from their starting points, while avoiding collisions and optimizing the overall performance. Traditional methods include rule-based methods (such as priority planning, spatio-temporal reservation, etc.), search-based methods (such as A*, RRT, etc.), and optimization-based methods (such as mixed integer linear programming, etc.). These methods perform well in static environments and predefined task scenarios, but lack adaptability in dynamic environments and scenarios with frequently changing tasks.
[0005] In recent years, deep reinforcement learning technology has shown great potential in multi-agent systems, and can learn from experience and adapt to complex environments. However, most existing multi-agent systems based on deep reinforcement learning rely on manually labeled data and a central controller, require a large amount of labeled data during the learning process, and the cooperation mechanism between intelligent agents is often predefined, lacking self-adaptability. In addition, traditional methods usually treat task allocation and path planning as independent problems, ignoring the mutual influence between the two, resulting in a reduction in the overall efficiency of the system.
[0006] Most existing methods are one-way frameworks that separate the task allocation and path planning processes, lacking comprehensive consideration of the agent's state (such as battery level, load capacity), task characteristics (such as priority, time window), and environmental dynamic changes. In such one-way frameworks, the upper limit of system performance is often limited by the quality of the initial task allocation. Therefore, there is an urgent need to introduce a multi-agent dynamic task allocation and collaborative pathfinding system based on unlabeled distributed deep reinforcement learning. Through the state exchange and adaptive collaboration mechanism among agents, the optimal matching of tasks and agents and path planning can be achieved, thereby improving the overall efficiency and adaptability of the system while reducing the dependence on manually labeled data.
[0007] The information disclosed in this background section is only intended to enhance the overall understanding of the present invention and should not be regarded as an admission or any form of implication that this information constitutes the prior art already known to those of ordinary skill in the art. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to overcome the above technical defects and provide a multi-agent dynamic task allocation and collaborative pathfinding system based on unlabeled distributed deep reinforcement learning. Agents can autonomously evaluate the matching degree with each task based on their own state and the information exchanged with other agents, realizing decentralized dynamic task allocation and collaborative pathfinding without the intervention of manually labeled data and a central controller. By applying distributed deep reinforcement learning technology and the state exchange mechanism among agents, the system can quickly adapt to environmental changes and task dynamics, improving resource utilization and task completion efficiency.
[0009] To solve the above problems, the technical solution of the present invention is a multi-agent dynamic task allocation and collaborative pathfinding system based on unlabeled distributed deep reinforcement learning, including an environment module, a DHC model, a training module, and a task allocation module;
[0010] The method of using the system includes the following steps:
[0011] Step 1: Receive the state information of each agent and environmental perception data in a multi-agent system based on distributed deep reinforcement learning;
[0012] The agents include heterogeneous agents with different performance parameters;
[0013] The state information includes the battery state, load capacity, current position, and target information of the agent;
[0014] The environmental perception data includes local observation information and state data exchanged among agents;
[0015] Each agent is equipped with a neural network decision model;
[0016] Step 2: Extract the feature representations of the environmental perception data and the agent state information, and perform multi-source heterogeneous information fusion through an attention mechanism to obtain state-task matching features;
[0017] Step 3: Transmit the state-task matching features to the multi-agent network in real time, and adopt a hierarchical scheduling and state exchange mechanism based on task priorities to achieve task assignment negotiation among agents, dynamically detect new task types and perform incremental learning, and realize the online update and iteration of the model to assist the multi-agent in optimizing the task assignment strategy and path planning actions;
[0018] Based on the iterated model, achieve cooperative obstacle avoidance and path optimization in the multi-agent environment.
[0019] Furthermore, in Step 2, adaptive multi-task learning is adopted, and the formula of the loss function is:
[0020] L(θ t ,θ s )=∑L tas k(G t (V m ,θ t ),y t )+β∑L state (G s (V m ,θ s ),y s );
[0021] Where the state-task matching feature V m represents the agent-task matching space; y t is the task completion priority feature, and y m is the agent state feature; G t and G s represent the prediction results of the task assignment network and the state evaluation network, L task represents the task assignment loss, L state represents the state evaluation loss, θ t and θ s correspond to the internal parameters of the task assignment network and the state evaluation network respectively, and β is a trade-off hyperparameter.
[0022] Furthermore, Step 3 specifically includes the following steps: In the dynamic change process of the task flow, the model calculates the matching score using the agent state vector and the task feature vector, and the formula is:
[0023] Match(αi,tj)=min(μd,1)·Comp(αi,tj)+max(1-μd,0)·Dist(αi,tj)
[0024] Match is the agent αi With t j If the matching score of the task is higher than the set threshold and is better than other agents, assign the task to agent α i ; d is the number of tasks completed in each assignment cycle, and the max and min operators keep the weight between (0, 1); μ is the dynamic balance parameter used to balance the task compatibility and distance factors; Comp is the compatibility score between the agent and the task; Dist is the normalized distance score from the agent to the task;
[0025] When the agent updates its decision model according to the newly assigned task, the change of the model parameters further updates the state information of the agent. The agent updates the strategy and action according to the immediate state information and task information. The formula is:
[0026]
[0027] Q(s,α) is the value function of action α in state s, E π is the expectation function, R t+1 is the reward obtained by the agent at time t + 1, s · is the next state after executing action α, and γ is the discount factor with a value range from 0 to 1.
[0028] Furthermore, in the path planning process of the agent, the task completion efficiency is composed of the task completion time and the energy consumption ratio. The formula is:
[0029] η = ω·Φ(p,t)+(1 - ω)·ε(p,e)
[0030] ω is the balance coefficient, p represents the path decision, t represents the time factor, e represents the energy factor, Φ represents the time efficiency function, and ε represents the energy efficiency function;
[0031] In the construction of the decision - making cost function for path planning and task execution, based on the current state of the agent, the distribution of environmental obstacles, and the positions of other agents, the total decision - making cost function can be constructed as:
[0032] Θ = w c ·C c (p)+wd·Cd(p)+w e ·C e (p,e)+K
[0033] C c 、C d and C e are the cost functions of collision avoidance, path distance, and energy consumption respectively, normalized to values between 0 and 1, w c 、w d and w eis the weight of three cost functions, and K measures the cost function of violating task constraints during the execution of the agent:
[0034] K = w p ·C p (p) + w t ·C t (p, t) + w o ·C o (p)
[0035] C p is the priority violation cost, representing the penalty for violating the task priority sorting, C t is the time window cost, representing the penalty for violating the task time window constraint, C o is the path overlap cost, representing the penalty for excessive overlap with the paths of other agents, w p 、w t and w o are the weights of the three cost functions.
[0036] Furthermore, in step one, the neural network decision model adopts an adaptive attention communication mechanism, expressed as:
[0037]
[0038] h i and h′i are the hidden state representations of agent i before and after communication, h i is the hidden state representation of agent j, A ij is the attention weight of agent i to the information of agent j, W ν is the value conversion matrix, a i is the communication adjustment coefficient of the agent, which is dynamically adjusted according to the current task complexity and agent state.
[0039] Furthermore, the distributed deep reinforcement learning adopts the asynchronous advantage actor-critic algorithm (A3C), and each agent maintains a value network and a policy network. The learning objective is:
[0040]
[0041]
[0042] θ s and θ ν are the parameters of the policy network and the value network respectively, Q(s, a) is the state-action value function, V(s, θ ν ) is the state value function, θ is the immediate reward, s′ is the next state, and γ is the discount factor.
[0043] Furthermore, the system also includes a progressive curriculum learning mechanism that dynamically adjusts the environmental complexity through the following formula:
[0044]
[0045] D t is the environmental complexity parameter at time t, including the number of tasks, time constraints, etc. D0 is the initial complexity, and D max is the maximum complexity, T is the total duration of the preset curriculum learning, and λ is a non-linear adjustment factor that controls the rate of difficulty increase.
[0046] Furthermore, the system also includes a prioritized experience replay mechanism that calculates the sample priority in the following way:
[0047] p i =(∫δ i ∣+ε) α
[0048] p i is the priority of sample i, δ i is the temporal difference error of this sample, ε is a small constant to prevent the priority from being zero, and α is the priority exponent parameter that controls the degree of priority allocation.
[0049] Furthermore, the system also includes a multi-agent coordination layer that realizes optimal resource allocation by calculating the cooperation coefficient between agents:
[0050]
[0051] C ij is the cooperation coefficient between agents i and j, d ij is the distance between the two agents, σ is the distance sensitivity parameter, ν i and ν j are the task vectors of the two agents, cos(ν i ,ν j ) calculates the cosine similarity of the two vectors, e i and e j are the energy levels of the two agents, and the cooperation coefficient is used to guide the task transfer and cooperation decision-making between agents.
[0052] Furthermore, the system also includes an adaptive reward shaping mechanism, and the reward function is defined as:
[0053] R = Rt as k + R energy + R co ll a b + R cons t ra i n t
[0054] R task For task completion rewards, related to task priority and completion time; R energy For energy efficiency rewards, related to energy consumption; R collab For collaboration rewards, encouraging effective collaboration among agents; R constraint For constraint rewards, punishing behaviors that violate system constraints, and the weights of each part are dynamically adjusted according to the current system state and task characteristics.
[0055] The advantages of the present invention compared with the existing technologies are as follows:
[0056] 1. The present invention improves task allocation efficiency: Through the state exchange and adaptive matching mechanism among agents, the optimal matching of tasks and agents is achieved. Compared with the traditional centralized allocation method, the task completion time is reduced and the resource utilization rate is improved.
[0057] 2. The present invention enhances system adaptability: The tagless deep reinforcement learning enables the system to adapt to environmental changes and new task types without manual data annotation. In the dynamic task scenario test, the system's adaptability to new task types is improved compared with the traditional supervised learning method.
[0058] 3. The present invention optimizes energy efficiency: By considering the power status of agents and task characteristics, the optimization of energy consumption is achieved. In the long-term operation test, compared with the system that does not consider energy factors, the average energy consumption of agents is reduced and the charging times are decreased.
[0059] 4. The present invention reduces communication overhead: The selective communication based on the attention mechanism significantly reduces the system communication requirements. Compared with the fully connected communication network, the communication data volume is reduced while maintaining similar system performance.
[0060] 5. The present invention improves collaboration efficiency: The multi-agent collaborative pathfinding mechanism effectively reduces the path conflicts and waiting times among agents. In the high-density agent scenario, the path conflicts are reduced and the system throughput is improved. Brief Description of the Drawings
[0061] Figure 1 is the overall architecture schematic diagram of the present invention.
[0062] Figure 2 is the schematic diagram of the intelligent scheduling module of the present invention.
[0063] Figure 3 is the schematic diagram of the neural network model structure of the present invention.
[0064] Figure 4 is the schematic diagram of the communication mechanism of the present invention.
[0065] Figure 5It is a schematic diagram of the training workflow of the present invention.
[0066] Figure 6 It is a system workflow diagram of the present invention. Specific embodiments
[0067] In order to make the content of the present invention easier to be clearly understood, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0068] As Figures 1-6 shown, for the problems of dynamic task allocation and collaborative pathfinding in a multi-agent system, taking the theory of unlabeled distributed deep reinforcement learning and multi-agent state exchange as the main theoretical basis, for the adaptive learning problem of agent state perception and task matching, through the state exchange mechanism between multi-agents, multi-source heterogeneous information fusion technology, and adaptive learning technology of dynamic task allocation, a multi-agent dynamic task allocation and collaborative pathfinding system and method based on unlabeled distributed deep reinforcement learning are proposed.
[0069] The overall architecture of the system includes five core modules: The environment module is responsible for environment configuration, agent definition, map generation, and reward function design;
[0070] The model module includes the DHC model and its network structure, communication mechanism, experience replay, and action selection;
[0071] The training module is responsible for the collaborative work of the trainer, worker nodes, global buffer, and evaluator;
[0072] The tool module provides visualization and configuration management functions;
[0073] The task allocation module realizes task priority management, power management, task distribution optimization, and dynamic matching.
[0074] The method includes the following steps:
[0075] Step 1: The system receives the state information and environment perception data of each agent in the multi-agent system based on distributed deep reinforcement learning; the agents include heterogeneous agents with different performance parameters, the state information includes the power state, load capacity, current position, and target information of the agent, and the environment perception data includes local observation information and state data exchanged between agents; each agent is equipped with a neural network decision model.
[0076] In this embodiment, the objects are heterogeneous agents with different performance parameters in a multi-agent environment, including types such as high-load-capacity agents, high-speed moving agents, and long-endurance agents, which are used to complete task requirements with different characteristics. Basic preparation stage: 1) Data preparation: Each agent captures surrounding environment data through a multi-channel sensor, including information such as the distribution of static obstacles, the positions of other agents, and the positions of tasks to be assigned, and performs necessary preprocessing; 2) Basic model loading: Each agent loads a neural network decision model suitable for its own characteristics to ensure basic perception and decision-making capabilities.
[0077] Step 2: Extract the feature representations of the environmental perception data and agent state information, and perform multi-source heterogeneous information fusion through an attention mechanism to obtain state-task matching features.
[0078] In this embodiment, multi-source heterogeneous information fusion is adopted: Each agent transmits the environment observation data and its own state information after feature extraction to the multi-source information fusion module, and uses the attention mechanism to perform weighted fusion on information from different sources to obtain the state-task matching feature V m . The specific description is as follows:
[0079] Based on the theory of multi-agent system modeling, each agent has independent decision-making capabilities and can adjust its behavior according to its own state. By analyzing and judging the states and task characteristics of multiple agents, the deep relationship between agent state data and task characteristics can be effectively summarized. First, as Figure 2 shown, the core innovations of the present invention include: task priority management (ensuring the priority completion of key tasks, optimizing resource allocation, and preventing task delays), power dynamic management (charging route planning, power warning mechanism, and ensuring service continuity), task distribution optimization (load balancing scheduling, preventing traffic congestion, and improving resource utilization), and unlabeled dynamic matching (state information exchange, optimal matching selection, and adaptive task allocation). Joint modeling of multi-source heterogeneous information (environmental observations, agent states, task characteristics, etc.):
[0080] V m = ψ(O1, O2,..., O n , S1, S2,..., S m , T1, T2,..., T k )
[0081] where ψ is the feature fusion function, and V m represents the state-task matching feature, O i represents the environmental observation feature of the i-th agent, and T iDenote the state features of the $i$-th agent (including power, load, location, etc.), and denote the characteristic features of the $j$-th task (including location, priority, deadline, etc.). The state-task matching feature $V$ m is input into the subsequent task allocation network to achieve the optimal matching between the agent and the task.
[0082] Under the adaptive learning framework of the multi-agent system, in order to better utilize unlabeled data for online learning, an adaptive multi-task learning method is adopted to jointly optimize task allocation and state evaluation through shared representation learning. The state-task matching feature $V$ m obtained is used as the shared representation to simultaneously optimize the two related tasks of task allocation and state evaluation. At this time, the loss function of the adaptive multi-task learning can be obtained:
[0083] $L(\theta$ t , $\theta$ s ) = $\sum L$ task $(G$ t $(V$ m , $\theta$ t ), $y$ t ) + $\beta\sum L$ state $(G$ s $(V$ m , $\theta$ s ), $y$ s )
[0084] where $G$ t and $G$ s respectively represent the prediction results of the task allocation network and the state evaluation network, $L$ task represents the task allocation loss, $L$ state represents the state evaluation loss, $\theta$ t and $\theta$ s respectively correspond to the internal parameters of the task allocation network and the state evaluation network, and $\beta$ is a trade-off hyperparameter used to balance the importance of the two tasks. In this way, the system can simultaneously optimize the task allocation ability and the state evaluation ability of the agent in the case of unlabeled data, and achieve unsupervised collaborative learning.
[0085] Step 3: Transmit the state-task matching feature to the multi-agent network in real time, and adopt a hierarchical scheduling and state exchange mechanism based on task priority to achieve task allocation negotiation among agents, dynamically detect new task types and perform incremental learning, and realize the online update and iteration of the model to assist the multi-agent in optimizing the task allocation strategy and path planning actions.
[0086] In this embodiment, hierarchical scheduling and state exchange based on task priorities: According to the real-time states of each agent and their matching degrees with tasks, hierarchical scheduling is performed according to task priorities. At the same time, through the state exchange mechanism, key information is shared among agents to negotiate task allocation schemes. Dynamic task detection and incremental learning: The system can detect newly added task types and update the model through incremental learning to adapt to environmental changes. Specifically as follows:
[0087] Design a multi-agent collaborative online learning method to achieve the adaptive allocation of multiple agents for dynamic tasks and reduce human intervention. For example Figure 3 As shown, the neural network model structure adopted by the present invention includes: an observation input layer, a CNN encoder for feature extraction, multiple residual blocks for deep feature learning, GRU units for processing temporal information, a flattening layer for converting data dimensions, and finally value evaluation and action decision are output through an advantage network architecture. This network structure can effectively process complex state information in a multi-agent environment, and through the matching mechanism of agent state information and task characteristics, and at the same time based on the state exchange network between multi-agent systems, the attention mechanism is used to selectively share key information to achieve decentralized task allocation decisions; at the same time, each agent continuously optimizes its own decision-making model through shared experience and collaborative learning.
[0088] In the dynamic change process of the task flow, the system calculates the matching score using the agent state vector and the task feature vector to achieve dynamic task allocation. The calculation formula of the matching score is as follows:
[0089] Match(a i ,t j )=min(μd,1)·Comp(a i ,t j )+max(1-μd,0)·Dist(a i ,t j )
[0090] In the formula, Match is the matching score between agent a i and task t j . If it is higher than the set threshold and better than other agents, the task is assigned to agent a i ; d is the number of tasks completed in each allocation cycle, and the max and min operators keep the weight between (0, 1); μ is a dynamic balance parameter used to balance task compatibility and distance factors; Comp is the compatibility score between the agent and the task; Dist is the normalized distance score from the agent to the task, reflecting the spatial proximity.
[0091] When the agent updates its decision-making model according to the newly assigned task, the action policy is further optimized through deep reinforcement learning, and the formula is as follows:
[0092] Q(s,a) = E π [R t+1 + γ·max a' Q(s',a') | S t = s, A t = a]
[0093] Where Q(s,a) is the value function of action a in state s, and E π is the expectation function, R t+1 is the reward obtained by the agent at time t + 1, s' is the next state after executing action a, and γ is the discount factor with a value range from 0 to 1. The agent continuously adjusts its behavior according to environmental changes and task requirements. At the same time, its decision-making model is also continuously iteratively updated through online learning to adapt to environmental changes.
[0094] As Figure 4 shown, the multi-agent state exchange network realizes selective information sharing based on the attention mechanism, demonstrating the communication mechanism between multi-agents. Each agent i generates a query vector q i , and calculates the attention weight a j with the key vector k ij of other agent j, selectively focuses on the most relevant agent information, and updates its own state representation accordingly:
[0095]
[0096] Where h i and h i ' are the hidden state representations of agent i before and after communication respectively, h j is the hidden state representation of agent j, A ij is the attention weight of agent i to the information of agent j, W v is the value conversion matrix, and a i is the communication adjustment coefficient of agent i, which is dynamically adjusted according to the current task complexity and agent state.
[0097] The advantages of multi-agent collaborative learning are reflected in: 1) Agents can autonomously evaluate the matching degree based on their own states and task characteristics to achieve decentralized task allocation; 2) Through the state exchange network, agents can efficiently share key information and reduce communication overhead; 3) The system can dynamically detect new task types and perform incremental learning without manual annotation of data; 4) The framework based on deep reinforcement learning enables the system to continuously optimize from experience, improving task completion efficiency and resource utilization.
[0098] Study the impact on online learning by exploring the correlation between the reward signal and task characteristics in multi-agent reinforcement learning; at the same time, establish an efficient cooperation mechanism in a dynamic task environment through the joint optimization of task allocation and path planning to maximize the overall performance of the system.
[0099] In a multi-agent dynamic task allocation and collaborative pathfinding system for unlabeled distributed deep reinforcement learning in one embodiment, it also includes, based on the iterated model, to achieve collaborative obstacle avoidance and path optimization in a multi-agent environment. The specific steps of this process are as follows:
[0100] As Figure 5 shown, the training workflow of the present invention includes three key modules that operate collaboratively: the agent training execution process (environment initialization, interaction loop, end-of-episode processing, model update, state output), the data processing process (data reception, data preprocessing, priority update), and the training process (network initialization, training loop, periodic operations). The three modules form a closed-loop feedback mechanism by sending experiences, providing training batches, and updating model weights, ensuring the continuous optimization ability of the system. In a multi-agent environment, based on the state-task matching mechanism and multi-agent collaborative learning method, the joint optimization of dynamic task allocation and collaborative path planning is achieved to improve task completion efficiency and reduce resource consumption. In the process of multiple agents collaborating to complete tasks, achieving efficient task allocation and path planning is the core goal of this system. The goal of the system is to maximize the task completion efficiency η while minimizing the cost function Θ of path planning.
[0101] As Figure 6 shown, the system workflow of the present invention demonstrates the complete task allocation process from state evaluation (battery level, current location, task load, priority status), state information exchange (broadcasting its own state, receiving the states of other agents, constructing a state matrix) to optimal matching selection (calculating the matching degree, selecting the optimal task, dynamic priority adjustment), and finally entering the task execution and update stage (executing the assigned task, updating state information, dynamic path planning). The whole process forms a closed-loop feedback mechanism to ensure that the system can continuously adapt to environmental changes and optimize decisions.
[0102] In the path planning process of the agent, the task completion efficiency is composed of the task completion time and the energy consumption ratio, and its formula can be constructed as follows:
[0103] η = ω·Φ(p,t) + (1 - ω)·ε(p,e)
[0104] Among them, ω is the balance coefficient, p represents path decision-making, t represents the time factor, e represents the energy factor, Φ represents the time efficiency function, which is the time cost evaluation of the agent during the task execution, and ε represents the energy efficiency function, which is the energy consumption evaluation of the agent during the task execution. The system will dynamically adjust the value of ω according to the current power state of the agent. When the power is sufficient, it pays more attention to time efficiency, and when the power is insufficient, it pays more attention to energy efficiency.
[0105] In the construction of the path planning and task execution decision cost function, based on the current state of the agent, the distribution of environmental obstacles, and the positions of other agents, the total decision cost function Θ can be constructed as:
[0106] Θ = w c ·C c (p) + w d ·C d (p) + w e ·C e (p, e) + K
[0107] Among them, C c 、C d and C e are the cost functions of collision avoidance, path distance, and energy consumption respectively, normalized to values between 0 and 1, w c 、w d and w e are the weights of the three cost functions, and K measures the cost function of violating task constraints during the agent's execution process:
[0108] K = w p ·C p (p) + w t ·C t (p, t) + w o ·C o (p)
[0109] Among them, C p is the priority violation cost, representing the penalty for violating the task priority order, C t is the time window cost, representing the penalty for violating the task time window constraint, C o is the path overlap cost, representing the penalty for excessive overlap with the paths of other agents, w p 、w t and are w o the weights of the three cost functions.
[0110] During the multi-agent collaborative pathfinding process, each agent generates an optimal path through a deep reinforcement learning network based on its own state information and task characteristics, combined with the shared information from other agents, while ensuring the maximization of the overall system efficiency. Agents avoid path conflicts through a negotiation mechanism, achieve collaborative obstacle avoidance and path optimization, and improve the overall task completion rate and resource utilization efficiency of the system.
[0111] It should be understood that although the steps in the above flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the above flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0112] The above is only the preferred implementation mode of the multi-agent dynamic task allocation and collaborative pathfinding system and method of the tagless distributed deep reinforcement learning disclosed in the present invention, and is not used to limit the protection scope of the embodiments of this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of this specification shall be included in the protection scope of the embodiments of this specification.
Claims
1. A multi-agent dynamic task allocation and collaborative pathfinding system for tagless distributed deep reinforcement learning, characterized in that: It includes an environment module, a DHC model, a training module, and a task allocation module; The method for using the system includes the following steps: Step 1: Receive the state information and environmental perception data of each agent in the multi-agent system based on distributed deep reinforcement learning; The agents include heterogeneous agents with different performance parameters; The state information includes the power state, load capacity, current position, and target information of the agent; The environmental perception data includes local observation information and state data exchanged between agents; Each of the agents is equipped with a neural network decision model; Step 2: Extract the feature representations of the environmental perception data and agent state information, and perform multi-source heterogeneous information fusion through an attention mechanism to obtain state-task matching features; Step 3: Transmit the state-task matching features to the multi-agent network in real time, adopt a hierarchical scheduling and state exchange mechanism based on task priorities to achieve task allocation negotiation between agents, dynamically detect new task types and perform incremental learning, and realize online update and iteration of the model to assist the multi-agent in optimizing the task allocation strategy and path planning actions; Based on the iterated model, collaborative obstacle avoidance and path optimization in the multi-agent environment are realized.
2. The multi-agent dynamic task allocation and cooperative pathfinding system for tagless distributed deep reinforcement learning according to claim 1, characterized in that: In Step 2, adaptive multi-task learning is adopted, and the formula for the loss function is: L(θ t ,θ s ) = ∑L task (G t (V m ,θ t ), y t ) + β∑L state (G s (V m ,θ s ), y s ); Among them, the state-task matching feature V m represents the agent-task matching space; y t is the task completion priority feature, y m is the agent state feature; G t and G s represent the prediction results of the task assignment network and the state evaluation network, L task represents the task assignment loss, L state represents the state evaluation loss, θ t and θ s correspond to the internal parameters of the task assignment network and the state evaluation network respectively, and β is the trade-off hyperparameter.
3. The multi-agent dynamic task allocation and cooperative pathfinding system for label-free distributed deep reinforcement learning according to claim 1, characterized in that: Step 3 specifically includes the following steps: During the dynamic change process of the task flow, the model calculates the matching score using the agent state vector and task feature vector, and the formula is: Match(α i ,t j ) = min(μd, 1)·Comp(α i ,t j ) + max(1 - μd, 0)·Dist(α i ,t j ) Match is for agent α i and t j The matching score of the task. If it is higher than the set threshold and better than other agents, then assign this task to agent α i ; d is the number of tasks completed in each allocation cycle, and the max and min operators keep the weight between (0, 1); μ is the dynamic balance parameter, which is used to balance the task compatibility and distance factors; Comp is the compatibility score between the agent and the task; Dist is the normalized distance score from the agent to the task; When the agent updates its decision model according to the newly assigned task, the change in the model parameters further updates the state information of the agent, and the agent updates the strategy and actions according to the immediate state information and task information, and the formula is: Q(s,α) is the value function of action α in state s, and E π is the expected function, R t+1 is the reward obtained by the agent at time t+1, s· is the next state after executing action α, and γ is the discount factor with a value range from 0 to 1.
4. The multi-agent dynamic task allocation and cooperative pathfinding system for tagless distributed deep reinforcement learning according to claim 1, characterized in that: During the path planning process of the agent, the task completion efficiency consists of the task completion time and the energy consumption ratio, and the formula is: η = ω·Φ(p,t)+(1 - ω)·ε(p,e) ω is the balance coefficient, p represents the path decision, t represents the time factor, e represents the energy factor, Φ represents the time efficiency function, and ε represents the energy efficiency function; In the construction of the decision-making cost function for path planning and task execution, based on the current state of the agent, the distribution of environmental obstacles, and the positions of other agents, the total decision-making cost function can be constructed as: Θ = w c ·C c (p) + w d ·C d (p) + we·C e (p, e) + K C c 、C d and C e are the cost functions for collision avoidance, path distance, and energy consumption, respectively, normalized to values between 0 and 1, where w c 、w d and w e are the weights of the three cost functions, and K measures the cost function for violating task constraints during the execution of the agent: K = w p ·C p (p) + w t ·C t (p, t) + w o ·C o (p) C p is the priority violation cost, representing the penalty for violating the task priority ranking, C t is the time window cost, representing the penalty for violating the task time window constraint, C o is the path overlap cost, representing the penalty for excessive overlap with the paths of other agents, w p 、w t and w o are the weights of the three cost functions.
5. The multi-agent dynamic task allocation and cooperative pathfinding system for tagless distributed deep reinforcement learning according to claim 1, characterized in that: In Step 1, the neural network decision model adopts an adaptive attention communication mechanism, which is expressed as: h i and h' i are the hidden state representations of agent i before and after communication, respectively, where h i is the hidden state representation of agent j, and A ij is the attention weight of agent i to the information of agent j, and W ν is the value conversion matrix, and a i is the communication adjustment coefficient of the agent, which is dynamically adjusted according to the current task complexity and agent state.
6. The multi-agent dynamic task allocation and cooperative pathfinding system for tagless distributed deep reinforcement learning according to claim 1, characterized in that: The distributed deep reinforcement learning adopts the asynchronous advantage actor-critic algorithm (A3C), and each agent maintains a value network and a policy network, and the learning objective is: θ s and θ ν are the parameters of the policy network and the value network respectively, Q(s,a) is the state-action value function, V(s,θ ν ) is the state value function, θ is the immediate reward, s′ is the next state, and γ is the discount factor.
7. The multi-agent dynamic task allocation and cooperative pathfinding system for tagless distributed deep reinforcement learning according to claim 1, characterized in that: The system also includes a progressive curriculum learning mechanism, which dynamically adjusts the environmental complexity through the following formula: D t is the environmental complexity parameter at time t, including the number of tasks, time constraints, etc. D0 is the initial complexity, and D max is the maximum complexity, T is the preset total duration of course learning, and λ is a non-linear adjustment factor that controls the rate of difficulty increase.
8. The multi-agent dynamic task allocation and collaborative pathfinding system for tagless distributed deep reinforcement learning according to claim 1, characterized in that: The system also includes a prioritized experience replay mechanism, and the prioritized experience replay mechanism calculates the sample priority through the following method: p i = (∫δ i | + ε) α p i is the priority of the i-th sample, and δ i is the temporal difference error of the sample, ε is a small constant to prevent the priority from being zero, and α is the priority exponent parameter that controls the degree of priority allocation.
9. The multi-agent dynamic task allocation and cooperative pathfinding system for tagless distributed deep reinforcement learning according to claim 1, characterized in that: The system also includes a multi-agent coordination layer, which realizes optimal resource allocation by calculating the cooperation coefficient between agents: C ij is the cooperation coefficient between agents i and j, d ij is the distance between the two agents, σ is the distance sensitivity parameter, ν i and ν j are the task vectors of the two agents, cos(ν i , ν j ) calculates the cosine similarity of the two vectors, e i and e j are the energy levels of the two agents. The cooperation coefficient is used to guide the task transfer and cooperation decision-making between agents.
10. The multi-agent dynamic task allocation and cooperative pathfinding system for tagless distributed deep reinforcement learning according to claim 1, characterized in that: The system also includes an adaptive reward shaping mechanism, and the reward function is defined as: R = R task + R energy + R collab + R constraint R task For task completion rewards, related to task priority and completion time; R energy For energy efficiency rewards, related to energy consumption; R collab For collaboration rewards, encouraging effective collaboration among agents; R constraint For constraint rewards, punishing behaviors that violate system constraints, and the weights of each part are dynamically adjusted according to the current system state and task characteristics.
Citation Information
Cited By
Multi-agent task allocation method based on network coverage dynamic matching control
CN120935588A
Multi-agent task allocation method based on network coverage dynamic matching control
CN120935588B
Dynamic optimization method for multi-agent task allocation in distributed environment
CN121279742A
Distributed agent dynamic collaboration method and device, computer equipment, storage medium and computer program product
CN121441958A
Unmanned post task scheduling system and method fused with reinforcement learning
CN121504312A