Multi-agent task allocation method based on MDPSO algorithm
Through the matrix-based discrete particle swarm optimization algorithm (MDPSO), the multi-constraint and strong coupling problems of target allocation in multi-agent systems are solved, efficient and rapid task allocation is achieved, and the effectiveness of multi-agent collaborative combat is improved.
Patent Information
- Application Number
- CN202510637106.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-26
AI Technical Summary
The existing technology is difficult to efficiently solve the complex multi-constrained and strongly coupled multi-objective optimization problems of target allocation in multi-agent systems, affecting the effectiveness of collaborative combat of multi-agents.
The discrete particle swarm optimization algorithm (MDPSO) based on matrix form is used to design particle swarm initialization, fitness calculation, update and movement processes through particle swarm movement and mutation methods, find the optimal task allocation scheme, and meet the multi-objective and multi-attribute constraints.
It realizes efficient and rapid finding the optimal task allocation plan in multi-agent systems, avoids local optimality, and improves the effectiveness of multi-agent collaborative combat.
Smart Images

Figure CN120540808A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multi-agent task allocation. When multi-agent systems are operating with a known environment and a known target, it employs dual-matrix real-number encoding to directly represent the task sequence assigned to each agent, facilitating practical operational applications. Furthermore, its particle movement design can partially address target constraints, while its mutation design can avoid falling into local optima. Background Art
[0002] In recent years, with the coordinated development of multiple disciplines, including electronics, computer science, control theory, communications, and artificial intelligence, multi-agent systems, such as cruise missiles and drones, have moved from the laboratory to the battlefield. The problem of multi-agent collaborative task allocation has gradually attracted the attention of researchers both domestically and internationally, and a large number of research efforts have yielded fruitful results.
[0003] During collaborative combat, multi-agent systems must formulate firepower coordination strategies based on multiple target information. The quality of target allocation closely impacts the effectiveness of multi-agent collaborative combat, striving to maximize kill gains with minimal losses. Therefore, designing an efficient target allocation algorithm is a crucial issue in multi-agent collaborative combat.
[0004] The multi-agent collaborative task allocation problem is a complex multi-objective optimization problem with multiple constraints and strong coupling. Particle swarm optimization (PSO) is a key technology for solving task allocation problems. This paper proposes a matrix-based discrete particle swarm optimization (MDPSO) algorithm. This algorithm aims to find the optimal task allocation solution for a multi-agent system by using particle movement and mutation within the PSO algorithm, while ensuring that the multi-agent system meets multiple objective and multi-attribute constraints and optimization goals. Summary of the Invention
[0005] The purpose of this invention is to provide a multi-agent task allocation method based on the MDPSO algorithm, which can provide each agent with an autonomous task allocation goal. Figure 5 shown.
[0006] A multi-agent task allocation method based on the MDPSO algorithm includes the following steps:
[0007] Step 1: Initialize the particle population. Set the population size and the total number of iterations k max , initialize the particle swarm and randomly generate the particle position matrix Z according to the constraints;
[0008] Step 2: Fitness value calculation. Generate the task sequence matrix W based on the position matrix Z, and then calculate the fitness value J of each particle based on the constraints and fitness function;
[0009] Step 3: According to the minimum fitness value Jmin Select the optimal particle p in the current cycle of the population Best and the global optimal particle g Best ;
[0010] Step 4: According to the probability, all particles in the population are mutated to Z1 with probability β;
[0011] Step 5: Update the optimal particle p in this cycle Best and the global optimal particle g Best ;
[0012] Step 6: Update the velocity matrix V according to the probability α t ;
[0013] Step 6: All particles Z1 are divided into two groups according to the velocity sequence V. t to move;
[0014] Step 7: Update the particle position Z1 to Z2 and calculate the optimal particle fitness value at the current position;
[0015] Step 8: Update the optimal particle p in this cycle Best and the global optimal particle g Best ;
[0016] Step 9: Output the optimal solution until the loop condition is met, otherwise return to step 4. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Multi-agent task allocation scheme structure diagram;
[0018] Figure 2 Schematic diagram of particle swarm initialization;
[0019] Figure 3 Particle matrix movement pattern diagram;
[0020] Figure 4a ) Schematic diagram of swapping two columns of a particle matrix
[0021] Figure 4b ) Schematic diagram of swapping two rows of a particle matrix
[0022] Figure 5 Flowchart of multi-agent task allocation method based on MDPSO algorithm DETAILED DESCRIPTION
[0023] The present invention will be described in further detail below with reference to the accompanying drawings.
[0024] Combine Figure 1 , detailing the mathematical model of the multi-agent task allocation problem. Assume that the enemy target position and battlefield environment have been obtained in advance. U={U1,U2,...,U n} represents a set of n agents; T = {T1, T2, ..., T m} represents a set of m enemy targets.
[0025] Assume that the attribute is Indicates that U i represents the i-th agent, Represents the agent U i The two-dimensional coordinates of Represents the agent U i The initial energy corresponds to the range, Represents the agent U i The initial bomb load, Represents the agent U i Target T j The probability of damage, τ represents the time required to perform tasks such as identification, attack and damage assessment on a single target, Represents the agent U i The value of itself.
[0026] Assume that the target T j Attributes by Indicates that T j represents the jth target, Indicates the target T j The two-dimensional coordinates of Indicates the target T j For agent U i degree of threat.
[0027] This paper requires that multiple agents are assigned a target execution sequence after task allocation, namely:
[0028]
[0029] In formula (1), Represents the agent U i The distribution result, Represents the agent U i Execution target T j , and eventually every U i Return to the nearest base is required after completing the full sequence of objectives.
[0030] The mathematical model of the fitness function is:
[0031]
[0032] In formula (2), λ1 and λ2 represent the benefit effect coefficients of a single agent, λ3 and λ4 represent the penalty scales for a single agent exceeding the constraints, and λ5 and λ6 represent the collaborative allocation effect coefficients. i Indicates the threat cost;i Indicates the attack benefit; Indicates the remaining ammunition penalty; Indicates timeout penalty; L i represents the fuel consumption penalty; Λ represents the time penalty. The following is an explanation of the above parameters:
[0033] 1) Threat Cost C i The expression is:
[0034]
[0035] In formula (3), the agent U i The value of The expression is Represents the agent U i Intrinsic value, Indicates the type of ammunition r value, υ indicates the ammunition load. ij Represents the agent U i Target T j The collaborative decision-making amount of agent U i The task sequence contains the target T j , then x ij =1, otherwise x ij =0. Represents the agent U i Target T j Probability of hitting.
[0036] 2) Attack benefit Η i The expression is:
[0037]
[0038] In formula (4), is the target T j The value of It is the intelligent agent U i Target T j probability of killing.
[0039] 3) Remaining ammunition penalty The expression is:
[0040]
[0041] In formula (5), the agent U i The initial ammunition load is ω kij Represented by integers {0, 1}, if ω ijk =1, indicating that the agent U i From the target T j Fly to target T k , otherwise ωijk =0, indicating that the agent U i Stay at T j Place.
[0042] 4) Timeout penalty The expression is:
[0043]
[0044] In formula (6), represents the upper limit of the time required for multiple agents to perform tasks, Represents the agent U i Cruise speed, Represents the agent U i The number of targets in the task sequence.
[0045] 5) Fuel consumption penalty L i The expression is:
[0046]
[0047] Mode, Represents the agent U i The final energy corresponds to the range, ξ represents the remaining energy coefficient, which is generally 0.5. Represents the agent U i Initial energy corresponds to range.
[0048] 6) The expression of the time cost Λ is:
[0049]
[0050] The present invention uses the time cost Λ to characterize the time efficiency and coordination ability of multi-agent task allocation, that is, the ratio of the time required by the agent with the longest task time to the time required by the agent with the shortest task time.
[0051] Combine Figure 2 , detail the particle swarm initialization process.
[0052] The multi-agent task allocation scenario described in this patent has three constraints, namely task duration constraint, task coordination constraint and capability constraint. The capability constraint includes the agent payload constraint and fuel constraint. The task duration constraint is the timeout penalty in the mathematical model of the multi-agent task allocation problem. In the mathematical model of multi-agent task allocation problem, the fuel consumption penalty L is solved. i Solved in.
[0053] The task coordination constraint means that each target can only be assigned to one agent, and each agent must complete at least one task target. The payload constraint is the agent's own capability indicator, which is expressed as:
[0054]
[0055] The particle position is represented by a double matrix, Z = X + Y, and the n × m dimensional target allocation matrix X is a row full rank matrix, x ij ∈{0, 1}, requiring that each column has only one non-zero element 1.
[0056] The sequence matrix Y is also an n×m dimensional matrix, where the element y ij ∈(0,1), represents the priority of the execution target, and the task with the highest priority is executed first. The number of decimal places in the sequence matrix Y is determined by the number of targets, that is, T ID =y×10 a , where a represents the number of target points. The decimal value is equal to the target point number. For example, if the target to be searched is T2, and the intelligent agent UAV2 performs the task of the target point, then x 2,2 =1,y 2,2 =0.2, z 2,2 =1.2. At this time x i,2 =y i,2 =z i,2 =0, and i≠2.
[0057] Regarding the problem of how to execute the particle swarm initialization, the following method is used to handle it. Assuming that there are m intelligent agents and n task targets, first randomly extract m targets and evenly distribute them to each intelligent agent, and then extract k targets from the remaining targets, each time extracting The targets are assigned to an agent without duplication, and are drawn at most m times. If there are still targets that have not been assigned, the loop phase begins, where one target is drawn each time and given to the agent that does not meet the most task constraints, until all targets are assigned and the loop ends.
[0058] Combine Figure 3 , detailing the particle movement method.
[0059] This patent proposes a discrete particle swarm optimization algorithm based on matrix form:
[0060]
[0061] In the above formula, V t It represents the moving speed of the particle swarm when the agent group performs the tth cycle, P represents the optimal task allocation scheme of the agent in the current cycle, that is, the local optimal particle position, G represents the global optimal task allocation scheme of the agent, that is, the global optimal particle position, Z t Indicates the position of the entire population of particles in the tth cycle. t This means that there exists a 1×m moving sequence pc such that Z tThe matrix P is obtained by the pc movement rule. Similarly, there is another sequence gc such that Z t The matrix G is obtained by the gc movement rule. α represents the probability of non-zero movement in each column. Assuming that the probability of movement is equal, the probability follows a uniform distribution. Symbol Indicates that two movement sequences are combined into one movement sequence, and finally form a speed sequence V t Then pc and gc are expressed as:
[0062]
[0063] The symbol ⊙ represents the particle population position Z t Execution speed sequence V t , let the velocity sequence V t Expressed as: V t ={k1,k2,...,k i ,...,k m}, then Z t The update method is expressed as:
[0064]
[0065] Combined with Figure 4, the particle mutation method is described in detail. The mathematical model of the matrix discrete particle swarm (MDPSO) algorithm after adding the mutation form is as follows:
[0066]
[0067] In formula (13), Represents Z αt Perform row and column transformation according to probability β, and then get the new particle position Z βt .
[0068] Mutations are divided into pairs Z t Swap rows and columns:
[0069] Z t Row transformation (wc): represents the exchange of any two rows in the matrix Z. The physical meaning is that the two agents execute each other's task sequence, as shown in Figure 4(a);
[0070] Z t Column transformation (rc): It means that any two columns in the matrix Z are interchanged. The physical meaning is that two agents exchange a task, as shown in Figure 4(b).
[0071] β represents Z αt The probability of swapping rows and columns, assuming that Υ is a random number uniformly distributed between [0,1]. In order to converge quickly, in the kth iteration calculation, β is updated in the form of formula (14):
[0072]
[0073] In formula (14), N represents the total number of cycles, and n represents the current number of cycles. When β≤Y, there is no need to αt Perform any operation; when β<Y, then Z αt Perform row and column swap operations. In order to express the probability of row and column swapping, let β = {β1, β2}, where β1 and β2 are independent of each other and represent the probability of row swapping and column swapping, respectively.
[0074] Before each iteration, a certain probability β is used to choose whether Z needs to be changed t The mutation operation is performed by applying t This is achieved by positioning each particle in .
Claims
1. A multi-agent task allocation method based on the MDPSO algorithm, which includes the following steps: constructing a particle description in the form of a double matrix, encoding the target number and priority level into the matrix; designing a fitness function based on the target threat cost, energy consumption cost, time cost, attack benefit, etc., and obtaining the historical optimal solution p of the particle individual Best and the historical optimal solution g of the particle population Best ; According to the matrix characteristics, design the particle movement and mutation methods; according to the optimal solution p of the particle swarm in the current cycle Best and the historical optimal solution g of the particle population Best , update the particle swarm, and after reaching the preset number of iterations, obtain the multi-agent task allocation plan for multiple objectives.
2. A multi-agent task allocation method based on MDPSO algorithm according to claim 1, characterized in that: The problem is described as follows: Taking the air-to-ground battlefield scenario as an example, in a two-dimensional map, our air reconnaissance agent is assigned the task of attacking the enemy target. Assuming that the enemy target position and battlefield environment have been obtained in advance, U = {U1, U2, ..., U n } represents the set of n agents on our side; T={T1,T2,...,T m } represents a set of m enemy targets; assuming that the attributes of the agent Ui are Indicates that U i represents the i-th agent; Represents the agent U i The two-dimensional coordinates of Represents the agent U i The initial energy corresponds to the range, Represents the agent U i The initial bomb load, Represents the agent U i Target T j The probability of damage, τ represents the necessary time to perform the terminal tasks such as identification, attack and damage assessment on a single target, Represents the agent U i Cruise speed, Represents the agent U i The mission targets studied in this paper are divided into multiple types of fixed enemy targets such as bunkers, caves, and bridges. The mission includes identification, attack, and damage assessment, so the mission time cannot be ignored. Assuming that the target T j Attributes by Indicates that T j represents the jth target, Indicates the target T j The two-dimensional coordinates of Indicates the target T j For agent U i the degree of threat; This patent requires that agents start from multiple different bases and traverse all enemy targets as a whole. A single agent must correspond to at least one enemy target. The result of multi-agent task allocation is: assign a target execution sequence to each agent, expressed as: In formula (1), Represents the agent U i The distribution result, Represents the agent U i Execution target T j , and eventually every U i Return to the nearest base is required after completing the full sequence of objectives.
3. The multi-agent task allocation method based on the MDPSO algorithm according to claim 1 is characterized in that: The fitness function is: In formula (2), λ1 and λ2 represent the benefit effect coefficients of a single agent, λ3 and λ4 represent the penalty scales for a single agent exceeding the constraints, and λ5 and λ6 represent the collaborative allocation effect coefficients; C i represents the threat cost, Η i Indicates the attack benefit, Indicates the remaining ammunition penalty. Indicates the timeout penalty, L i represents the energy consumption penalty, and Λ represents the time cost; the above parameters are explained below: 1) Threat Cost C i The expression is: In formula (3), the value of the agent is The expression is Represents the agent U i Intrinsic value, represents the type of ammunition r value, υ represents the ammunition load; x ij Indicates U i Target T j The collaborative decision-making amount of agent U i The task sequence contains the target T j , then x ij =1, otherwise x ij =0; Represents the agent U i Target T j Probability of hitting; 2) Attack benefit Η i The expression is: In formula (4), is the target T j The value of It is the intelligent agent U i Target T j The probability of killing; 3) Remaining ammunition penalty The expression is: In formula (5), the agent U i The initial ammunition load is ω kij Represented by integers {0, 1}; if ω ijk =1, indicating that the agent U i From the target T j Fly to target T k Otherwise ijk =0, indicating that the agent U i Stay at T j Department; 4) Timeout penalty The expression is: In formula (6), represents the upper limit of the time required for multiple agents to perform tasks, Represents the agent U i Cruise speed, Represents the agent U i The number of targets in the mission sequence; 5) Energy consumption penalty L i The expression is: In formula (7), Represents the agent U i The final energy corresponds to the range, ξ represents the remaining energy coefficient, which is generally 0.
5. Represents the agent U i Initial energy corresponds to range; 6) The expression of the time cost Λ is: The present invention uses the time cost Λ to characterize the time efficiency and coordination ability of multi-agent task allocation, that is, the ratio of the time required by the agent with the longest task time to the time required by the agent with the shortest task time.
4. The multi-agent task allocation method based on the MDPSO algorithm according to claim 1 is characterized in that: The double matrix is generated as follows: The multi-agent task allocation scenario described in this patent has three constraints, namely task duration constraint, task coordination constraint and capability constraint; the capability constraint includes the payload constraint and energy constraint of the agent, and the task duration constraint is the timeout penalty in the mathematical model of the multi-agent task allocation problem. The energy penalty L in the mathematical model of the multi-agent task allocation problem is solved in i Solved in The task coordination constraint means that each target can only be assigned to one agent, and each agent must complete at least one task target. The payload constraint is an indicator of the agent's own capability, expressed as: The particle position is represented by a double matrix, Z = X + Y, and the n × m dimensional target allocation matrix X is a row full rank matrix, x ij ∈{0, 1}, requiring each column to have only one non-zero element 1; The sequence matrix Y is also an n×m dimensional matrix, where the element y ij ∈(0,1), represents the priority of the execution target, and the task with the highest priority is executed first; the number of decimal places in the sequence matrix Y is determined by the number of targets, that is, T ID =y×10 a , where a represents the target number; the decimal value is equal to the target point number. For example, if the target to be searched is T2, and the agent U2 performs the task of the target point, then x 2,2 =1,y 2,2 =0.2, z 2,2 =1.2; at this time x i,2 =y i,2 =z i,2 =0, and i≠2; As for how to execute the above formula, we use the following method to deal with it. Assume that there are n agents and m task targets (m≥n). First, randomly select n targets and evenly distribute them to each agent. Then, we extract k targets from the remaining targets. Each time we extract The targets are assigned to an agent without duplication, and are drawn at most m times. If there are still targets that have not been assigned, the loop phase begins, where one target is drawn each time and given to the agent that does not meet the most task constraints, until all targets are assigned and the loop ends.
5. The multi-agent task allocation method based on the MDPSO algorithm according to claim 1 is characterized in that: The particles move as follows: This patent proposes a discrete particle swarm optimization algorithm based on matrix form as follows: In the above formula, V t It represents the moving speed of the particle swarm when the agent group performs the tth cycle, P represents the optimal task allocation scheme of the agent in the current cycle, that is, the local optimal particle position, G represents the global optimal task allocation scheme of the agent, that is, the global optimal particle position, Z t Indicates the position of the entire population particle at the tth cycle, PZ t This means that there exists a 1×m moving sequence pc such that Z t The matrix P is obtained by the pc movement rule; similarly, there is another sequence gc such that Z t The matrix G is obtained by the gc movement rule. α represents the probability of non-zero numbers in each column moving. Assuming that they move with equal probability, their probability obeys uniform distribution. The symbol Indicates that two movement sequences are combined into one movement sequence, and finally form a speed sequence V t , then pc and gc are expressed as follows: The symbol ⊙ represents the particle population position Z t Execution speed sequence V t , let the velocity sequence V t Expressed as: V t ={k1,k2,...,k i ,...,k m }, then Z t The updating method is formula (12).
6. The multi-agent task allocation method based on the MDPSO algorithm according to claim 1 is characterized in that: The particle mutation method is as follows; The mathematical model of the matrix discrete particle swarm optimization (MDPSO) algorithm after adding mutation form is as follows: in, Represents Z αt Perform row and column transformation according to probability β, and then get the new particle position Z βt ; Mutations are divided into pairs Z t Swap rows and columns: Z t Row transformation (wc): represents the exchange of any two rows in the matrix Z. The physical meaning is that the two agents execute each other's task sequence, as shown in Figure 4(a); Z t Column transformation (rc): It means that any two columns in the matrix Z are interchanged. The physical meaning is that two agents exchange a task, as shown in Figure 4(b). β represents Z αt The probability of swapping rows and columns, assuming that Υ is a random number uniformly distributed between [0,1]. In order to converge quickly, in the kth iteration calculation, β is updated in the following form: In the above formula, N represents the total number of cycles, and n represents the current number of cycles. When β≤Y, there is no need to adjust Z. αt Perform any operation; when β<Y, then Z αt Perform row and column swap operations. To express the probability of row and column swapping, let β = {β1, β2}, where β1 and β2 are independent of each other and represent the probability of row and column swapping, respectively. Before each iteration, a certain probability β is used to choose whether Z needs to be changed t , the mutation operation is performed by applying t This is achieved by positioning each particle in .
7. The multi-agent task allocation method based on the MDPSO algorithm according to claim 1 is characterized in that: The steps of task allocation are as follows: Step 1: Initialize the particle population, set the population size, and the total number of iterations k max , initialize the particle swarm and randomly generate the particle position matrix Z according to the constraints; Step 2: Fitness value calculation, generate the task sequence matrix W according to the position matrix Z, and then calculate the fitness value J of each particle according to the constraint conditions and fitness function; Step 3: According to the minimum fitness value J min Select the optimal particle p in the current cycle of the population Best and the global optimal particle g Best ; Step 4: According to the probability, all particles in the population are mutated to Z1 with probability β; Step 5: Update the optimal particle p in this cycle Best and the global optimal particle g Best ; Step 6: Update the velocity matrix V according to the probability α t ; Step 6: All particles Z1 are divided into two groups according to the velocity sequence V. t to move; Step 7: Update the particle position Z1 to Z2 and calculate the optimal particle fitness value at the current position; Step 8: Update the optimal particle p in this cycle Best and the global optimal particle g Best ; Step 9: Output the optimal solution until the loop condition is met, otherwise return to step 4.