A Multi-Agent Distributed Task Coordination Method and System
By constructing a constrained optimization model and a global optimization algorithm in a multi-agent system, the computational burden and inflexibility of centralized scheduling systems are solved, achieving efficient and flexible task allocation and resource scheduling, and improving the system's response speed and stability.
Patent Information
- Application Number
- CN202510616865.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-05-14
AI Technical Summary
Traditional centralized multi-agent systems suffer from excessive computational burden, communication latency, and insufficient flexibility when facing complex tasks and dynamic environments, resulting in inefficient task allocation and slow response speed.
By broadcasting a set of tasks to be assigned to the agent, generating a task expected duration matrix, constructing a constraint optimization model, and using a global constraint optimization algorithm for task allocation, the flexibility and efficiency of task allocation are ensured by combining the agent's load information and historical data.
It improves the flexibility and response speed of task allocation, reduces communication latency, enhances the flexibility and stability of the system, avoids resource waste and task delays, and achieves efficient task allocation and resource utilization.
Smart Images

Figure CN120542814B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a multi-agent distributed task coordination method and system. Background Technology
[0002] As a key link in the logistics industry, warehouse management no longer relies solely on manual operation and single automated equipment. With the demands of e-commerce, rapid turnover, and large-scale supply chains, warehouse management systems are gradually developing towards greater efficiency, flexibility, and intelligence.
[0003] Traditional warehousing systems typically rely on centralized scheduling systems for task allocation and resource allocation. Their core is a centralized control unit that manages and schedules all warehousing tasks and resources. While this approach worked well in early, small-scale warehousing environments, as the number of tasks and the types of resources increase, the computational and management burden on centralized scheduling systems rises significantly, leading to slower system response times and processing capacity becoming a bottleneck. Furthermore, warehousing systems require rapid responses to dynamic changes in tasks and unexpected events, but traditional centralized methods often lack the flexibility to adapt quickly to real-time environmental changes.
[0004] A multi-agent system (MAS) consists of multiple agents with autonomous perception, decision-making, and execution capabilities. Multi-agent task collaboration refers to the coordination and cooperation among agents within the system to jointly complete a specific task. Task allocation, as a key link in task collaboration, directly affects the operational efficiency and success rate of the entire multi-agent system.
[0005] Currently, multi-agent task allocation methods mainly fall into two categories: centralized control and traditional distributed control. Centralized control uses one or a few central nodes to handle task allocation, information synchronization, and agent coordination. It performs well in small-scale, relatively simple scenarios and can coordinate and optimize global objectives. However, as the number of agents increases and the environment becomes more dynamic, the computational load and communication latency of the central nodes increase dramatically, easily leading to decision-making bottlenecks and single points of failure. Traditional distributed task allocation methods rely on static or predefined rules. When faced with increasing task complexity and environmental uncertainty, they cannot effectively adapt to rapid changes in the task environment, resulting in low collaborative efficiency and even task failure.
[0006] Currently, the industry has not proposed a better technical solution to the above problems. Summary of the Invention
[0007] This application provides a multi-agent distributed task collaboration method, system, storage medium, computer program product, and electronic device to at least solve the problems of insufficient centralized scheduling flexibility, insufficient real-time adaptability, and low collaboration efficiency in current related technologies for multi-agent systems.
[0008] In a first aspect, embodiments of this application provide a multi-agent distributed task coordination method, comprising: broadcasting a set of tasks to be assigned to each agent, such that each agent estimates the expected completion time of each task to be assigned and generates a corresponding expected completion time matrix for the tasks to be assigned; receiving the expected completion time matrix for the tasks to be assigned and the estimated completion time of historically assigned tasks from each agent; wherein, if the agent is in an idle state, the estimated completion time of the queued tasks is 0; if the agent is in a running state, the estimated completion time of the historically assigned tasks is the sum of the expected completion times of at least one assigned task to be executed in the agent; constructing a constraint optimization model based on the expected completion time matrix for the tasks to be assigned and the estimated completion time of historically assigned tasks corresponding to each agent; the constraint optimization model includes task execution constraints, constraints on the number of tasks queued by the agent, constraints on the maximum load of the agent, and an objective function; solving the constraint optimization model using a global constraint optimization algorithm, and assigning each task to be assigned to a corresponding agent according to the optimal solution;
[0009] The objective function is expressed by the following formula:
[0010]
[0011] In the formula, M represents the total number of tasks in the task set to be assigned, and N represents the total number of agents; x ij X is a decision variable, representing whether to assign task i to agent j; if so, then X... ij =1, otherwise x ij =0; t ij This represents the time required for agent j to execute task i, which is determined based on the expected duration distribution of the task. It is the total estimated time of the assigned tasks for agent j;
[0012] The task execution constraints require that each task must be assigned to a single agent:
[0013]
[0014] The agent queue task quantity constraint is used to limit the upper limit of the number of tasks assigned to an agent, in order to avoid task backlog:
[0015]
[0016] In the formula, Q jIt is the threshold number of tasks assigned to agent j;
[0017] The maximum load constraint of the agent is used to ensure that the agent does not become overloaded while executing queued tasks and currently assigned tasks:
[0018]
[0019] In the formula, C j This represents the maximum workload that agent j can withstand.
[0020] Secondly, embodiments of this application provide a multi-agent distributed task collaboration system, comprising: a task set broadcasting unit, configured to broadcast a set of tasks to be assigned to each agent, enabling each agent to estimate the expected completion time of each task to be assigned and generate a corresponding expected completion time matrix for the tasks to be assigned; an estimated completion time receiving unit, configured to receive the expected completion time matrix for the tasks to be assigned and the estimated completion time of historically assigned tasks from each agent; wherein, if an agent is in an idle state, the estimated completion time of the queued tasks is 0; if an agent is in a running state, the estimated completion time of historically assigned tasks is the sum of the expected completion times of at least one assigned task to be executed in the agent; a constraint optimization modeling unit, configured to construct a constraint optimization model based on the expected completion time matrix of the tasks to be assigned and the estimated completion time of historically assigned tasks corresponding to each agent; the constraint optimization model includes task execution constraints, constraints on the number of tasks queued by the agent, constraints on the maximum load of the agent, and an objective function; and a global constraint solving unit, configured to solve the constraint optimization model using a global constraint optimization algorithm, and allocate each task to be assigned to the corresponding agent according to the optimal solution.
[0021] The objective function is expressed by the following formula:
[0022]
[0023] In the formula, M represents the total number of tasks in the task set to be assigned, and N represents the total number of agents; x ij X is a decision variable, representing whether to assign task i to agent j; if so, then X... ij =1, otherwise x ij =0; t ij This represents the time required for agent j to execute task i, which is determined based on the expected duration distribution of the task. It is the total estimated time of the assigned tasks for agent j;
[0024] The task execution constraints require that each task must be assigned to a single agent:
[0025]
[0026] The agent queue task quantity constraint is used to limit the upper limit of the number of tasks assigned to an agent, in order to avoid task backlog:
[0027]
[0028] In the formula, Q j It is the threshold number of tasks assigned to agent j;
[0029] The maximum load constraint of the agent is used to ensure that the agent does not become overloaded while executing queued tasks and currently assigned tasks:
[0030]
[0031] In the formula, C j This represents the maximum workload that agent j can withstand.
[0032] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the multi-agent distributed task coordination method of any embodiment of this application.
[0033] Fourthly, embodiments of this application provide a storage medium storing a computer program thereon, characterized in that, when the program is executed by a processor, it implements the steps of the multi-agent distributed task collaboration method of any embodiment of this application.
[0034] Fifthly, embodiments of this application provide a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the multi-agent distributed task coordination method of any embodiment of this application.
[0035] The multi-agent distributed task coordination method and system provided in this application can achieve at least the following technical effects:
[0036] (1) By broadcasting a set of tasks to be assigned and requiring each agent to estimate the expected completion time of each task, task assignment can take into account the load of each agent and the execution time of the task. By generating a task expected duration matrix and combining it with historical task data, the dynamic adjustment method based on state can more accurately evaluate the timing of each task assignment and the most suitable agent. This makes task assignment not only dependent on static rules or simple centralized control, but also based on real-time calculation and data feedback, which improves the flexibility of task assignment, reduces resource waste and unreasonable allocation in task assignment, and thus improves the response speed and efficiency of the entire warehousing system.
[0037] (2) Compared with the traditional centralized control method, each agent autonomously evaluates the task duration of each task to be assigned and shares the load information. Through distributed task collaboration, the computational burden of the central node is avoided, while communication latency is reduced, the system's response speed and flexibility are enhanced, and high task allocation efficiency is maintained when facing a large number of agents and dynamically changing environments.
[0038] (3) By constructing a constraint optimization model and using a global constraint optimization algorithm to solve it, we ensure that each task is assigned to the appropriate agent in the best way. While satisfying multiple constraints such as task execution time, agent load limit, and number of task queues, we maximize the utilization of resources and ensure that the agent can complete the task within a reasonable load range. This avoids task execution delay or failure due to overload and also avoids resource idleness, thus achieving global optimal scheduling.
[0039] This technical solution, by combining a global constraint optimization algorithm and a distributed collaborative mechanism of a multi-agent system, significantly improves the efficiency, stability, and reliability of the warehouse management system when dealing with complex task allocation and environmental changes. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 A flowchart illustrating an example of a multi-agent distributed task coordination method according to an embodiment of this application is shown.
[0042] Figure 2 A structural block diagram of an example of a multi-agent distributed task collaboration system according to an embodiment of this application is shown;
[0043] Figure 3 This is a schematic diagram of the structure of an embodiment of the electronic device of this application. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0045] The technical solutions in this application, including the collection, storage, use, processing, transmission, provision, and disclosure of users' personal information, comply with relevant laws and regulations and do not violate public order and good morals.
[0046] Figure 1 A flowchart illustrating an example of a multi-agent distributed task coordination method according to an embodiment of this application is shown.
[0047] Regarding the execution entity of the method in the embodiments of this application, it can be any controller or processor with computing or processing capabilities. Specifically, it can be implemented by a multi-agent task allocation and management platform. Through a distributed architecture and constrained optimization model, it significantly improves the response speed, computing efficiency, and global optimization capabilities of the warehouse management system, enhances the system's flexibility and robustness in complex dynamic environments, reduces operation and maintenance costs, and achieves efficient and reliable task allocation and resource scheduling.
[0048] In some examples, it can be integrated into an electronic device or terminal through software, hardware, or a combination of both, and the type of terminal or electronic device can be diverse, such as mobile phones, tablets, or desktop computers, etc.
[0049] like Figure 1 As shown, in step S110, a set of tasks to be assigned is broadcast to each agent, so that each agent estimates the expected completion time of each task to be assigned and generates a corresponding expected duration matrix of the tasks to be assigned.
[0050] In some implementations, the platform first identifies all tasks to be assigned and transmits all task information to each agent via network broadcast or multicast. These tasks may be warehouse operation tasks or other types of tasks, and may include corresponding task attribute information, such as task type, task resource requirements, task priority, and other relevant parameters.
[0051] For example, the intelligent agent is a warehouse intelligent agent, such as an Automated Guided Vehicle (AGV), an intelligent forklift, and a sorting robot. The task is a warehouse operation task, and the corresponding task types include any one of the following: receiving tasks, picking tasks, and inventory counting tasks. Thus, each task type has different standard execution times and resource consumption. Task resource requirements refer to the physical resources needed to complete the task, such as storage space, transportation equipment, and robotic equipment. This allows each intelligent agent to independently estimate the completion time of each task based on task attributes and its own state or performance, and generate a matrix of expected completion times for tasks to be assigned. This matrix contains the expected completion time for each task for which the intelligent agent is expected. For example, each intelligent agent, through its own computational model, combined with factors such as task attributes and dynamic changes in the environment, estimates the expected completion time of each task, and then returns the matrix of expected completion times for tasks to be assigned to the management platform.
[0052] Therefore, by broadcasting the set of tasks to be assigned to all agents through the central management platform, information transparency and sharing are ensured, enabling each agent to understand the task status within the system in a timely manner. Each agent can estimate the completion time of tasks in real time based on historical experience and its own execution capabilities, thereby improving the system's sensitivity to task execution and response speed.
[0053] In step S120, the expected duration matrix of the tasks to be assigned and the estimated duration of historically assigned tasks are received from each agent.
[0054] Here, if the agent is in an idle state, the estimated time for queuing tasks is... A value of 0 indicates that the agent currently has no tasks waiting to be executed. If the agent is in a running state, the estimated duration of historically assigned tasks is the sum of the expected completion times of at least one assigned task waiting to be executed in the agent, representing its current load status.
[0055] In this way, the management platform receives the expected duration matrix of tasks to be assigned from each agent and synchronizes the historical task allocation data of each agent, realizing a real-time feedback mechanism for task status. This helps the management platform understand the current workload of each agent, thereby flexibly adjusting the task allocation strategy to ensure that the task allocation can reasonably balance the load and avoid agents from being overloaded or idle.
[0056] In step S130, a constraint optimization model is constructed based on the expected duration matrix of the tasks to be assigned to each agent and the estimated duration of historical assigned tasks. The constraint optimization model includes task execution constraints, constraints on the number of tasks queued by the agent, constraints on the maximum load of the agent, and an objective function.
[0057] Here, based on the expected duration distribution of each agent and the estimated duration of historical tasks, a comprehensive constraint optimization model is constructed. The goal is to optimize task allocation so that the task execution time of the entire system is minimized, while satisfying the load and resource constraints of the agents.
[0058] Specifically, task execution constraints ensure that each task is appropriately assigned to an agent for execution. Queued task quantity constraints limit the number of tasks waiting to be executed by each agent, preventing an agent from having too many tasks and becoming overloaded. Maximum load constraints ensure that the workload of each agent (including the estimated duration of current tasks and historically assigned queued tasks) does not exceed its maximum capacity, preventing agent overload.
[0059] More specifically, the objective function is expressed by the following formula:
[0060]
[0061] In the formula, M represents the total number of tasks in the task set to be assigned, and N represents the total number of agents; x ij X is a decision variable, representing whether to assign task i to agent j; if so, then X... ij =1, otherwise x ij =0; t ij This represents the time required for agent j to execute task i, which is determined based on the expected duration distribution of the task. It is the total estimated time of the assigned tasks for agent j.
[0062] In the objective function of equation (1), the core objective is to minimize the total execution time of all tasks, thereby optimizing the efficiency of the entire system. The execution time of each task i includes not only the estimated duration t of the task itself. ij It is also necessary to consider the current load of the agent, i.e., the estimated time of the queuing tasks. This approach ensures that task allocation not only focuses on the execution of the current task but also takes into account the actual workload of the agent, effectively balancing the task allocation for each agent and preventing task delays caused by overload of a particular agent.
[0063] Task execution constraints require that each task must be assigned to a single agent:
[0064]
[0065] In the task execution constraints, it is ensured that each task i must be assigned to one and only one agent j, thus ensuring the uniqueness of task assignment and avoiding the situation where tasks are executed repeatedly or by multiple agents.
[0066] The agent queue task quantity constraint is used to limit the upper limit of the number of tasks assigned to an agent, in order to avoid task backlog:
[0067]
[0068] In the formula, Q j It is the threshold for the number of tasks assigned to agent j.
[0069] In the constraint on the number of queued tasks for agents, the number of queued tasks for each agent j is limited, i.e., a maximum of Q tasks can be assigned simultaneously. j This system limits the number of tasks in the queue to prevent agents from becoming overloaded due to processing too many tasks consecutively. Therefore, by limiting the number of tasks in the queue, the system can rationally schedule the tasks of each agent, avoiding task backlog and reducing task delays caused by agent overload.
[0070] The agent's maximum load constraint is used to ensure that the agent does not become overloaded while executing queued tasks and currently assigned tasks.
[0071]
[0072] In the formula, C j This represents the maximum workload that agent j can withstand.
[0073] According to equation (4), the maximum load constraint of the agent ensures that the workload of agent j will not exceed its maximum capacity C during the execution of the task. j This ensures that agents do not exceed their resource and capability limitations when executing tasks. By introducing maximum load constraints for agents, the system can limit the workload of each agent. When assigning new tasks to an agent, the total duration of its already executed tasks and queued tasks must be considered, thus ensuring that it does not exceed its maximum workload capacity. This effectively avoids the problem of agent overload and reduces task delays and execution failures caused by overload. Furthermore, reasonable load control for each agent improves the stability of the entire warehouse management system and the on-time completion rate of tasks, guaranteeing the quality of agent task execution.
[0074] Therefore, by introducing multiple constraints, the management platform can comprehensively consider factors such as task execution time, agent load, and task queue. It can not only meet the needs of each task, but also ensure that the load of each agent is within an acceptable range. It can optimize task allocation globally, reasonably schedule agent tasks, and avoid agents from being overloaded or having backlogged tasks, thereby improving task execution efficiency and the overall stability of the multi-agent system.
[0075] In step S140, the constraint optimization model is solved by the global constraint optimization algorithm, and each task to be assigned is assigned to the corresponding agent according to the optimal solution.
[0076] Specifically, in solving the constructed constraint optimization model using a global constraint optimization algorithm, the task allocation scheme is adjusted through multiple iterations to gradually approach the global optimum. The global optimization algorithm ensures that the task allocation scheme reaches its optimum under all constraints, thereby improving the overall task execution efficiency of the multi-agent system.
[0077] It should be noted that the types of global constraint optimization algorithms can be diverse. For example, in the task allocation problem, due to the decision variable x... ij Since the variables are binary (each task may or may not be assigned to an agent), the MILP (Mixed-Integer Linear Programming) algorithm, which has integer decision variables and linear constraints, can be used to solve the problem. Furthermore, given the large number and variety of tasks and agents in the warehouse environment, various heuristic algorithms (such as genetic algorithms and particle swarm optimization) can be employed under complex constraints. These algorithms support achieving lower computational time consumption even with high computational complexity, and are not limited here.
[0078] Through the embodiments of this application, by using a distributed task collaboration method combined with global constraint optimization, it is possible to respond to environmental changes in real time, effectively avoid the risk of single point of failure, optimize resource utilization and load balancing, and achieve flexibility in task allocation strategies.
[0079] Regarding the details of estimating the expected completion time of the task in step S110, in some examples of embodiments of this application, it can be implemented by a deep learning model based on a multilayer perceptron deployed in the agent.
[0080] In some implementations, the agent performs the following operations to estimate the expected completion time of the task to be assigned: parses the task type and resource requirements corresponding to the task to be assigned, and obtains the agent's device performance indicators to construct multivariate input features; then inputs the multivariate input features into a deep learning model based on a multilayer perceptron (MLP) to determine the corresponding expected completion time of the task.
[0081] Regarding the details of task type parsing, in some implementations, each task to be assigned includes task type information (e.g., handling, sorting, storage, etc.), and the agent needs to parse the execution characteristics of the task based on this task type. Different task types correspond to different operational complexities, required resources, and execution times. To enable the agent to adapt to different tasks, task types can be converted using One-Hot encoding.
[0082] Regarding the details of the task resource requirements analysis, the task resource requirements include the equipment (such as robots, sensors, etc.), storage space, power requirements, etc., such as the number of robots required, the sensor accuracy required for each task, and the required storage space or processing power.
[0083] Regarding the description of equipment performance indicators, each intelligent agent will regularly update its equipment performance indicators, such as robot movement speed, handling efficiency, sensor accuracy, battery capacity, and processor performance, etc., which reflect the processing capability of the intelligent agent when performing tasks and directly affect the task execution time.
[0084] By analyzing the type of task and resource requirements, the agent can accurately identify the basic needs of each task. By acquiring the agent's device performance indicators and matching them with the task resource requirements, the agent can dynamically adjust the task duration estimate based on its own capabilities, ensuring the efficiency and accuracy of task allocation.
[0085] A multidimensional input feature vector is constructed using the above information and then fed into an MLP-based deep learning model. The MLP model is then used to perform regression prediction of the task duration through a multi-layer neural network.
[0086] It should be noted that an MLP model can include multiple hidden layers, each containing multiple neurons. The number of neurons in each layer and the number of layers can be tuned through cross-validation to ensure the model's training effectiveness. Non-linear activation functions (e.g., ReLU activation function) are used between layers to improve the network's non-linear modeling capabilities. In particular, through the Multilayer Perceptron (MLP) model, the platform can accurately predict the execution time of each task based on the task type (receiving task, picking task, inventory counting task) and the agent's device performance.
[0087] Through the embodiments of this application, a multivariate prediction model based on MLP is deployed in an intelligent agent. The agent can accurately predict the execution time of a task based on the task type, resource requirements, and its own device performance. Because the MLP model has a strong nonlinear fitting capability, it can capture the relationships between complex features, thus making the task duration prediction more accurate. Furthermore, as the task environment dynamically changes (such as changes in task type and resource requirements), the MLP model can continuously optimize the task duration prediction by training and learning from new data, thereby adapting to tasks of different types and complexities. With the accumulation of data and feedback, the model can automatically adjust and optimize the duration prediction by continuously learning from the agent's task execution feedback, iteratively improving the accuracy of task duration prediction.
[0088] In some examples of embodiments of this application, the global constraint optimization algorithm adopts the cooperative particle swarm optimization algorithm. The solution process based on the global constraint optimization algorithm mainly includes the following operations:
[0089] Particle swarm initialization:
[0090] Each particle is defined as a task assignment scheme with dimensions M×N, where each dimension corresponds to the assignment relationship between each task and each agent; the particle position is defined using decision variables. This indicates whether task i is assigned to agent j in particle k.
[0091] Here, in the particle swarm optimization algorithm, the task allocation scheme is defined by the particle position matrix of each particle, and the position of each particle... Corresponding to the assignment relationship between task i and agent j, in the initial state, the particle position matrix and the corresponding task assignment scheme are random. By covering a wider range of possibilities in the search space, we can ensure that all possible task assignments are explored without any prior knowledge.
[0092] Under the premise of simultaneously satisfying task execution constraints, agent queue task quantity constraints, and agent maximum load constraints, the initial position of the particles is randomly generated; the particle velocity is then utilized. This represents the change in the assignment of task i to agent j in particle k. The initial velocity of the particle is zero. The velocity of the particle controls the change in the task assignment scheme. The particle updates its velocity and position in each iteration to optimize task scheduling.
[0093] Specifically, task execution constraints ensure that each task is assigned to one agent, and task queueing constraints ensure that the number of tasks assigned to each agent does not exceed a preset upper limit Q. jThe maximum load constraint of the agent ensures that the total load of each agent (the execution time of the tasks to be assigned plus the estimated time of the queued tasks) does not exceed the maximum load capacity C of that agent. j .
[0094] Particle swarm optimization based on neighborhood cooperative search:
[0095] Particles not only rely on their own optimal solutions, but also utilize the information-sharing mechanism of neighboring particles to accelerate the convergence of the global optimal solution through cooperative search.
[0096] The particle velocity update formula is:
[0097]
[0098] In the formula, Let represent the particle velocity of particle k in the t-th iteration, and let represent the change when task i is assigned to agent j; w(t) is the dynamic inertia weight, which determines the balance between global search and local exploitation of the particle, and gradually decreases with the number of iterations; c1(t) and c2(t) represent the individual optimal solution acceleration constant and the global optimal solution acceleration constant, respectively; r1 and r2 are random numbers in the range [0,1], used to increase the randomness of the search. It is the individual optimal solution of particle k in the t-th iteration, where task i is assigned to agent j. is the task allocation scheme of particle k in the global optimal solution, representing the optimal task allocation scheme among all particles; N(k) is the set of neighboring particles of particle k, and the particles exchange information through cooperative search; l represents the neighboring particles of particle k. Let represent the particle position of particle k in the t-th iteration, and let represent whether task i is assigned to agent j; α is the influence coefficient of cooperative search, which controls the strength of cooperation between particles.
[0099] In the particle velocity update formula shown in Equation (5), the particle velocity is used to represent the change in the task allocation scheme, controlling the direction and step size of the task allocation change. Specifically, the guiding role of particles between the current optimal solution and the global optimal solution is considered. At the same time, cooperative search is carried out through information transmission between neighboring particles, allowing the particle swarm to take into account neighborhood information during local search, thereby accelerating convergence.
[0100] Specifically, based on It provides an information-sharing mechanism among neighboring particles, allowing particles to adjust their speed and position based on the behavior of their neighbors. This enhances cooperation among particles, avoids getting trapped in local optima, ensures that the optimal solution is found in a larger search space, and also accelerates the convergence of the global optimal solution.
[0101] This effectively balances global search and local exploration, and integrates neighborhood collaborative search, enabling particles to find the global optimum faster during the search process and avoid getting trapped in local optima.
[0102] The particle position update formula is:
[0103]
[0104] In the formula, σ(·) represents the Sigmoid activation function, which is used to convert particle velocity into probability to determine whether task i is assigned to agent j; r ij It is a random number in the range [0,1], used to control whether task i is assigned to agent j.
[0105] In equation (6), the particle update is based on the Sigmoid activation function, which converts velocity into a probability value to determine whether a task is assigned to agent j. This allows the particle to adjust its task assignment in each iteration, ensuring the diversity of the particle search. Thus, it ensures that the task assignment decision is based on the particle's current velocity, while also enabling the particle to explore appropriately during the search process, improving the randomness of task assignment and enhancing the algorithm's global search capability.
[0106] After the particle position is updated, it is checked whether the task allocation scheme corresponding to the updated particle simultaneously satisfies the task execution constraint, the agent queue task number constraint, and the agent maximum load constraint. If satisfied, the iteration stop condition is checked; if not satisfied, the particle velocity is reinitialized and adjusted to zero.
[0107] Here, after each particle position update, a constraint verification is performed on the task allocation scheme represented by the particle. Only when the task execution constraints, the number of tasks queued by the agent, and the maximum load constraint of the agent are satisfied will the particle continue to participate in subsequent iterations. Thus, through the constraint verification mechanism, it is ensured that the updates of the particle after each iteration meet all the constraint requirements, thereby avoiding the generation of invalid solutions and guaranteeing the rationality and feasibility of the final solution.
[0108] The update formula for dynamic inertia weight is:
[0109]
[0110] In the formula, w max and w min These are the maximum and minimum values of the inertia weight, T. max It represents the maximum number of iterations.
[0111] Referring to equation (7), the balance between global and local searches of particles is controlled by the inertia weight w(t). The inertia weight gradually decreases with the increase of the number of iterations, so that in the early stage of the search, the inertia weight is large, and the particles have a strong global search ability; while in the later stage, the inertia weight gradually decreases, making the particles more concentrated in the local search. Thus, it helps the particle swarm avoid premature convergence in the early stage of the search and improves the accuracy of the search in the later stage.
[0112] The update formula for the adaptive speedup constant is:
[0113]
[0114] In the formula, c min,1 and c max,1 c represents the minimum and maximum acceleration constants of the individual optimal solution, respectively; min,2 and c max,2 These represent the minimum and maximum acceleration constants of the global optimal solution, respectively.
[0115] It should be noted that c1(t) is the optimal solution for controlling the particle for each individual. The acceleration constant of the degree of dependence. Because particles are more dependent on the global optimum in the early stages of the search space. Extensive exploration is conducted, while in the later stages, more reliance is placed on individual optimal solutions. To perform fine local optimization, the acceleration constant c1(t) should decrease as the number of iterations increases, so as to guide the particles to gradually turn to local search and enhance the local optimization capability.
[0116] Furthermore, c2(t) controls the particle's approach to the global optimal solution. The degree of dependence on the global optimum should be adjusted in the early stages of the search. In the later stages, particles should rely more on the global optimum for extensive exploration, while in the later stages, particles gradually shift towards their own optimal solutions. Therefore, c2(t) should be larger in the early stages to enhance global search capability, and gradually decreased in the later stages to enhance local search capability.
[0117] Thus, in the initial stage, the particle relies heavily on the global optimum with a large c2(t), enabling it to explore a wider solution space and avoid getting trapped in local optima. As the number of iterations increases, the particle gradually reduces its dependence on the global optimum and relies more on its own optimal solution for a refined local search, thereby accelerating the convergence process and improving the accuracy of the solution.
[0118] In the iteration stopping condition judgment, when the number of iterations reaches the maximum number of iterations T, maxWhen the preset fitness convergence condition is met, the iteration is terminated, and the particle with the highest fitness in the current particle swarm is returned as the optimal solution. Based on the optimal solution, each task to be assigned is assigned to the corresponding agent.
[0119] Specifically, the algorithm terminates when the maximum number of iterations is reached or the fitness changes of all particles satisfy the convergence condition. This ensures that the algorithm stops at an appropriate time, avoids excessive iteration and waste of computational resources, and returns the particle with the highest fitness as the optimal solution. Finally, the task is assigned to the agent.
[0120] Through the embodiments of this application, a cooperative particle swarm optimization algorithm is employed. By dynamically adjusting the inertia weight and adaptive acceleration constant, particles can effectively balance global exploration and local optimization during the search process and flexibly respond to different search stages, thus improving efficiency and accuracy. Furthermore, the cooperative search among particles improves the global convergence speed and avoids local optima, enhancing the algorithm's global search capability. Finally, the task allocation scheme is determined by the particle positions of the optimal solution, ensuring that tasks can be efficiently allocated to appropriate agents under constraints.
[0121] In some examples of embodiments of this application, the fitness function of the cooperative particle swarm optimization algorithm is expressed by the following formula:
[0122] F = F1 + λF2, Equation (10)
[0123]
[0124] In the formula, F represents the fitness function, F1 represents the task execution time optimization term, F2 represents the task timeout penalty term, and λ represents the fitness balance factor; t i,limit This represents the maximum time limit for task i, which is determined based on a pre-defined mapping relationship between task type and / or task priority and time limit.
[0125] Here, the fitness function F, used to evaluate the quality of each particle's solution, comprises two parts: the total task execution time optimization term F1 and the task timeout penalty term F2. Its aim is to simultaneously minimize both the total task execution time and the timeout penalty. A balance factor λ controls the relative importance of these two objectives, allowing adjustment of their relative weights during the optimization process. This enables the algorithm to comprehensively consider task execution time and time constraints during task scheduling, flexibly balancing execution efficiency and time constraints to meet the dynamic adjustment needs of different scenarios.
[0126] In Equation (11), the total task execution time optimization term F1 refers to the objective function of the constrained optimization model, calculates the total time required for task execution, including the execution time of the task itself and the time for agents to queue tasks, and calculates the total execution time of all tasks on all agents. This can minimize the total task allocation time and improve the overall working efficiency of the multi-agent system. In particular, by considering the impact of agent queuing tasks, it can effectively avoid uneven agent load or overload caused by task allocation, and ensure the high efficiency of task scheduling.
[0127] In equation (12), the task timeout penalty term F2 is used to evaluate the penalty for handling tasks exceeding the maximum time limit by calculating the difference between the task execution time and the task's maximum time limit. Specifically, it calculates whether all tasks time out and penalizes tasks that time out. The penalty for timeout tasks is max(0,t). ij -T i,limit The penalty is the difference between the actual execution time of the task and the maximum time limit of the task. The larger the difference, the greater the penalty, thus minimizing the chance of task timeouts. Therefore, the penalty term effectively prevents task timeouts, and by penalizing timed-out tasks, it motivates the algorithm to find task allocation schemes that meet the time limit requirements.
[0128] Through the embodiments of this application, by introducing the particle swarm optimization algorithm and combining the dual optimization objectives of the total execution time of multi-agent tasks and the time limit penalty of each task, the time and time limit requirements of task execution can be efficiently balanced, so that the task timeout problem can be effectively avoided during task scheduling, while maximizing the efficiency of task execution.
[0129] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of combined actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Secondly, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application. In the above embodiments, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0130] Figure 2 A structural block diagram of an example of a multi-agent distributed task collaboration system according to an embodiment of this application is shown.
[0131] like Figure 2As shown, the multi-agent distributed task collaboration system 200 includes a task set broadcasting unit 210, an estimated duration receiving unit 220, a constraint optimization modeling unit 230, and a global constraint solving unit 240.
[0132] The task set broadcasting unit 210 is used to broadcast the set of tasks to be assigned to each agent, so that the agents can estimate the expected completion time of each task to be assigned and generate a corresponding expected duration matrix of the tasks to be assigned.
[0133] The estimated duration receiving unit 220 is used to receive the expected duration matrix of the tasks to be assigned and the estimated duration of the historical assigned tasks from each of the intelligent agents; wherein, if the intelligent agent is in an idle state, the estimated duration of the queued tasks is 0; if the intelligent agent is in a running state, the estimated duration of the historical assigned tasks is the sum of the expected completion times of the tasks corresponding to at least one assigned task to be executed in the intelligent agent.
[0134] The constraint optimization modeling unit 230 is used to construct a constraint optimization model based on the expected duration matrix of the tasks to be assigned to each agent and the estimated duration of historical assigned tasks. The constraint optimization model includes task execution constraints, constraints on the number of tasks queued by the agent, constraints on the maximum load of the agent, and an objective function.
[0135] The global constraint solving unit 240 is used to solve the constraint optimization model through a global constraint optimization algorithm, and to assign each of the tasks to be assigned to the corresponding agents according to the optimal solution.
[0136] The objective function is expressed by the following formula:
[0137]
[0138] In the formula, M represents the total number of tasks in the task set to be assigned, and N represents the total number of agents; x ij X is a decision variable, representing whether to assign task i to agent j; if so, then X... ij =1, otherwise x ij =0; t ij This represents the time required for agent j to execute task i, which is determined based on the expected duration distribution of the task. It is the total estimated time of the assigned tasks for agent j;
[0139] The task execution constraints require that each task must be assigned to a single agent:
[0140]
[0141] The agent queue task quantity constraint is used to limit the upper limit of the number of tasks assigned to an agent, in order to avoid task backlog:
[0142]
[0143] In the formula, Q j It is the threshold number of tasks assigned to agent j;
[0144] The maximum load constraint of the agent is used to ensure that the agent does not become overloaded while executing queued tasks and currently assigned tasks:
[0145]
[0146] In the formula, C j This represents the maximum workload that agent j can withstand.
[0147] In some embodiments, this application provides a non-volatile computer-readable storage medium storing one or more programs including execution instructions. The execution instructions can be read and executed by an electronic device (including but not limited to a computer, server, or network device) to perform the steps of any of the multi-agent distributed task cooperation methods described above.
[0148] In some embodiments, this application also provides a computer program product, the computer program product including a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the steps of any of the above-described multi-agent distributed task cooperation methods.
[0149] In some embodiments, this application also provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of a multi-agent distributed task coordination method.
[0150] Figure 3 This is a schematic diagram of the hardware structure of an electronic device for executing a multi-agent distributed task coordination method according to another embodiment of this application, as shown below. Figure 3 As shown, the device includes:
[0151] One or more processors 310 and memory 320, Figure 3 Take the 310 processor as an example.
[0152] The device for implementing the multi-agent distributed task coordination method may further include: an input device 330 and an output device 340.
[0153] The processor 310, memory 320, input device 330, and output device 340 can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.
[0154] The memory 320, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the multi-agent distributed task collaboration method in the embodiments of this application. The processor 310 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 320, thereby realizing the multi-agent distributed task collaboration method in the above-described method embodiments.
[0155] The memory 320 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 320 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 320 may optionally include memory remotely located relative to the processor 310, and these remote memories can be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0156] Input device 330 can receive input digital or character information and generate signals related to user settings and function control of the electronic device. Output device 340 may include display devices such as a display screen.
[0157] The one or more modules are stored in the memory 320, and when executed by the one or more processors 310, they execute the multi-agent distributed task collaboration method in any of the above method embodiments.
[0158] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.
[0159] The electronic devices in this application embodiments exist in various forms, including but not limited to:
[0160] (1) Mobile communication devices: These devices are characterized by their mobile communication capabilities and primarily aim to provide voice and data communication. These terminals include smartphones, multimedia phones, feature phones, and low-end phones.
[0161] (2) Ultra-mobile personal computer devices: These devices fall under the category of personal computers, possessing computing and processing capabilities, and generally also have mobile internet access features. These terminals include: PDAs, MIDs, and UMPCs, etc.
[0162] (3) Portable entertainment devices: These devices can display and play multimedia content. This category includes audio and video players, handheld game consoles, e-book readers, as well as smart toys and portable car navigation devices.
[0163] (4) Other airborne electronic devices with data interaction capabilities, such as vehicle-mounted systems installed on vehicles.
[0164] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0165] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0166] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A multi-agent distributed task coordination method, the method comprising: Broadcast the set of tasks to be assigned to each agent, so that each agent can estimate the expected completion time of each task and generate a corresponding expected duration matrix of the tasks to be assigned. Each of the aforementioned agents receives a matrix of expected durations for tasks to be assigned and an estimated duration for historically assigned tasks; wherein, if the agent is in an idle state, the estimated duration for queued tasks is 0; if the agent is in a running state, the estimated duration for historically assigned tasks is the sum of the expected completion times of at least one assigned task to be executed in the agent. A constraint optimization model is constructed based on the expected duration matrix of the tasks to be assigned to each agent and the estimated duration of historical assigned tasks. The constraint optimization model includes task execution constraints, constraints on the number of tasks queued by the agent, constraints on the maximum load of the agent, and an objective function. The constraint optimization model is solved using a global constraint optimization algorithm, and each of the tasks to be assigned is assigned to the corresponding agents based on the optimal solution. The objective function is expressed by the following formula: In the formula, M represents the total number of tasks in the task set to be assigned, and N represents the total number of agents; x ij X is a decision variable, representing whether to assign task i to agent j; if so, then X... ij =1, otherwise x ij =0; t ij This represents the time required for agent j to execute task i, which is determined based on the expected duration distribution of the task. It is the total estimated time of the assigned tasks for agent j; The task execution constraints require that each task must be assigned to a single agent: The agent queue task quantity constraint is used to limit the upper limit of the number of tasks assigned to an agent, in order to avoid task backlog: In the formula, Q j It is the threshold number of tasks assigned to agent j; The maximum load constraint of the agent is used to ensure that the agent does not become overloaded while executing queued tasks and currently assigned tasks: In the formula, C j This represents the maximum workload that agent j can withstand.
2. The method according to claim 1, wherein, The agent deploys a deep learning model based on a multilayer perceptron. The agent performs the following operations to estimate the expected completion time of the task to be assigned: The task type and resource requirements corresponding to the task to be assigned are analyzed, and the device performance indicators of the agent are obtained to construct multivariate input features; The multivariate input features are input into the deep learning model based on a multilayer perceptron to determine the expected completion time of the corresponding task.
3. The method according to claim 1 or 2, wherein, The intelligent agent is a warehouse intelligent agent, and the task is a warehouse operation task. The corresponding task type includes any one of the following: receiving task, picking task, and inventory counting task.
4. The method according to claim 1, wherein, The global constraint optimization algorithm adopts the cooperative particle swarm optimization algorithm. The step of solving the constraint optimization model using a global constraint optimization algorithm and assigning each of the tasks to be assigned to the corresponding agents based on the optimal solution includes: Particle swarm initialization: Each particle is defined as a task assignment scheme with dimensions M×N, where each dimension corresponds to the assignment relationship between each task and each agent; the particle position is defined using decision variables. It indicates whether task i is assigned to agent j in particle k; Under the premise of simultaneously satisfying task execution constraints, agent queue task quantity constraints, and agent maximum load constraints, the initial position of the particles is randomly generated; the particle velocity is then utilized. This represents the change in the assignment of task i to agent j in particle k. The initial velocity of the particle is zero. The velocity of the particle will control the change in the task assignment scheme. The particle updates its velocity and position in each iteration to optimize task scheduling. Particle swarm optimization based on neighborhood cooperative search: Particles not only rely on their own optimal solutions, but also utilize the information-sharing mechanism of neighboring particles to accelerate the convergence of the global optimal solution through cooperative search; The particle velocity update formula is: In the formula, Let represent the particle velocity of particle k in the t-th iteration, and let represent the change when task i is assigned to agent j; w(t) is the dynamic inertia weight, which determines the balance between global search and local exploitation of the particle, and gradually decreases with the number of iterations; c1(t) and c2(t) represent the individual optimal solution acceleration constant and the global optimal solution acceleration constant, respectively; r1 and r2 are random numbers in the range [0, 1], used to increase the randomness of the search. It is the individual optimal solution of particle k in the t-th iteration, where task i is assigned to agent j. is the task allocation scheme of particle k in the global optimal solution, representing the optimal task allocation scheme among all particles; N(k) is the set of neighboring particles of particle k, and the particles exchange information through cooperative search; l represents the neighboring particles of particle k. Let represent the particle position of particle k in the t-th iteration, and let represent whether task i is assigned to agent j; α is the influence coefficient of cooperative search, which controls the strength of cooperation between particles. The particle position update formula is: In the formula, σ(·) represents the Sigmoid activation function, which is used to convert particle velocity into probability to determine whether task i is assigned to agent j; r ij It is a random number in the range [0, 1], used to control whether task i is assigned to agent j; After the particle position is updated, it is checked whether the task allocation scheme corresponding to the updated particle simultaneously satisfies the task execution constraint, the agent queue task number constraint, and the agent maximum load constraint. If satisfied, the iteration stop condition is checked; if not satisfied, the particle velocity is reinitialized and adjusted to zero. The update formula for the dynamic inertia weight is as follows: In the formula, w max and w min These are the maximum and minimum values of the inertia weight, T. max It is the maximum number of iterations; The update formula for the adaptive acceleration constant is as follows: In the formula, c min,1 and c max,1 c represents the minimum and maximum acceleration constants of the individual optimal solution, respectively; min,2 and c max,2 These represent the minimum and maximum speedup constants of the global optimal solution, respectively. In the iteration stopping condition judgment, when the number of iterations reaches the maximum number of iterations T, max When the preset fitness convergence condition is met, the iteration is terminated, and the particle with the highest fitness in the current particle swarm is returned as the optimal solution. Based on the optimal solution, each of the tasks to be assigned is assigned to the corresponding agent.
5. The method according to claim 4, wherein, The fitness function of the cooperative particle swarm optimization algorithm is expressed by the following formula: F = F1 + λF2, In the formula, F represents the fitness function, F1 represents the total task execution time optimization term, F2 represents the task timeout penalty term, and λ represents the fitness balance factor. t i,limit This represents the maximum time limit for task i, which is determined based on a pre-defined mapping relationship between task type and / or task priority and time limit.
6. A multi-agent distributed task coordination system, comprising: The task set broadcasting unit is used to broadcast the set of tasks to be assigned to each agent, so that the agents can estimate the expected completion time of each task to be assigned and generate a corresponding expected time matrix of the tasks to be assigned. The estimated duration receiving unit is used to receive the expected duration matrix of the tasks to be assigned and the estimated duration of the historical assigned tasks from each of the intelligent agents; wherein, if the intelligent agent is in an idle state, the estimated duration of the queued tasks is 0; if the intelligent agent is in a running state, the estimated duration of the historical assigned tasks is the sum of the expected completion times of the tasks corresponding to at least one assigned task to be executed in the intelligent agent. The constraint optimization modeling unit is used to construct a constraint optimization model based on the expected duration matrix of the tasks to be assigned to each agent and the estimated duration of historical assigned tasks. The constraint optimization model includes task execution constraints, constraints on the number of tasks queued by the agent, constraints on the maximum load of the agent, and an objective function. The global constraint solving unit is used to solve the constraint optimization model through a global constraint optimization algorithm, and to assign each of the tasks to be assigned to the corresponding agents according to the optimal solution; The objective function is expressed by the following formula: In the formula, M represents the total number of tasks in the task set to be assigned, and N represents the total number of agents; x ij X is a decision variable, representing whether to assign task i to agent j; if so, then X... ij =1, otherwise x ij =0; t ij This represents the time required for agent j to execute task i, which is determined based on the expected duration distribution of the task. It is the total estimated time of the assigned tasks for agent j; The task execution constraints require that each task must be assigned to a single agent: The agent queue task quantity constraint is used to limit the upper limit of the number of tasks assigned to an agent, in order to avoid task backlog: In the formula, Q j It is the threshold number of tasks assigned to agent j; The maximum load constraint of the agent is used to ensure that the agent does not become overloaded while executing queued tasks and currently assigned tasks: In the formula, C j This represents the maximum workload that agent j can withstand.
Citation Information
Patent Citations
Task allocation method and device
CN117455193A
Autonomous agent task priority scheduling
US20220335714A1