Multi-target multi-element distribution decision-making method based on multiple agents

Through the goal allocation strategy of multi-agent reinforced learning training, the problem of difficulty in solving set screening in traditional air defense deployment is solved, and the rapid screening of Pareto solution sets in air defense deployment is realized, and the decision-making efficiency of the air defense system is improved.

CN120337548AActive Publication Date: 2025-07-18HARBIN INST OF TECH

Patent Information

Application Number
CN202510425638.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Traditional multi-objective optimization algorithms cannot effectively screen out the optimal solution that meets actual needs in air defense deployment, and cannot take into account multiple scenario elements, resulting in the actual performance of the solution set being unable to be guaranteed.

Method used

The multi-objective and multi-factor allocation decision-making method based on multi-agents is adopted, and the optimal deployment solution is selected through air defense scenario modeling, agent action and observation space construction, reward and punishment function design, and MADDPG reinforcement learning.

Benefits of technology

It has achieved the rapid selection of the optimal air defense deployment plan in complex and changeable real scenarios, ensuring that the system achieves optimal performance in various key indicators, and improving the efficiency of the air defense system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337548A_ABST
    Figure CN120337548A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-target multi-element distribution decision-making method based on multiple agents. The invention relates to the technical field of optimal solution set decision making, and the problem is solved by innovatively using a target allocation strategy obtained based on reinforcement learning training for the situation that deployment solution sets are difficult to screen in traditional air defense deployment multi-target optimization. Abstracting the problem into a plurality of objective functions to be optimized so as to solve a deployment solution set of the problem by utilizing a multi-objective optimization algorithm; for the evaluation and screening problem of a deployment solution set, a relatively real air defense battlefield environment is abstracted on the basis of a deployment model, a multi-agent reinforcement learning algorithm is utilized to train agents to obtain a dynamic target allocation strategy, and the strategy is utilized to realize the evaluation and screening of a deployment scheme. It can be known through simulation analysis that the adopted target allocation strategy air defense deployment Pareto solution set rapid screening method based on reinforcement learning can achieve rapid screening of the air defense deployment Pareto solution set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optimizing solution set decision-making, and is a multi-objective and multi-factor allocation decision-making method based on multi-agent. Background Art

[0002] For complex large-scale multi-objective optimization problems, in order to ensure that each objective function is not ignored by the algorithm during optimization, resulting in the solution set falling into an extreme situation, a multi-objective optimization algorithm is generally used to optimize and solve it. However, the deployment solution obtained by the multi-objective optimization algorithm is not a single optimal solution, but a solution set composed of all dominant solutions. Therefore, for practical applications, after obtaining the Pareto solution set, it is still necessary to screen out the individual solutions that best meet the actual needs from the solution set according to the actual environment, that is, the decision-making problem in the Pareto solution set. However, for decision-making problems, existing traditional methods such as the weighted method, the ideal point method, and the constraint method cannot reproduce the actual scenario, the actual performance of the selected optimal solution cannot be guaranteed, and all actual scenario elements cannot be taken into account during the screening process. Therefore, an algorithm that can quickly screen the solution set according to the actual environment, taking into account multiple factors comprehensively, to obtain the optimal solution is urgently needed. Summary of the Invention

[0003] In view of the problem that the traditional optimization decision-making method cannot reproduce the real scenario and the actual performance of the selected optimal solution cannot be guaranteed, the present invention provides a multi-objective and multi-factor allocation decision-making method based on multi-agent, and the present invention provides the following technical solutions:

[0004] A multi-objective and multi-factor allocation decision-making method based on multi-agent, the method comprising the following steps:

[0005] Step 1: Perform air defense scenario modeling to simulate an air attack scenario;

[0006] Step 2: Build an agent action and observation space, and observe the space matrix;

[0007] Step 3: Establish a reward and punishment function, so that after each action is executed by the agent, a reward value will be fed back to the agent according to the action;

[0008] Step 4: Design an air defense environment process to lock the protected target;

[0009] Step 5: Perform MADDPG reinforcement learning to generate a strategy, optimize the air defense strategy, and perform multi-objective optimization to solve and deploy the solution set.

[0010] Preferably, the specific content of step 1 is:

[0011] Suppose there are three protected targets A, B, and C, where A and B are important protected targets with higher protection weights than C; the three protected targets are located in an area densely distributed with protected objects such as population gathering areas and factories under attack. This area is regarded as a circular protected object with a radius of 125 km. The center of the circular area is taken as the coordinate origin, the due north direction is the positive y-axis direction, and the due east direction is the positive x-axis direction; the incoming direction is due north, that is, from the positive y-axis direction, and the incoming mode is set to fly horizontally at a certain height. Each unit randomly locks a strike target and flies towards it at a constant speed.

[0012] Preferably, the specific steps of step 2 are as follows:

[0013] There are N enemies above the current air defense. Define the action space that the air defense unit can only take as:

[0014] action = discrete(a1, a2,..., a N ) T (1)

[0015] Among them, discrete means that this action space is a discrete space, that is, under this space, the action a i is a binary function with values of 0 or 1;

[0016] When its value is 1, it means that the air defense unit agent will turn its aiming target to the i-th air raid unit of the enemy at this moment; on the contrary, 0 means not to turn to this unit. At the same moment, an air defense unit can only aim at one enemy air raid unit;

[0017] For the information input to the agent at each moment, a observation space matrix is constructed, which is a matrix of size 5×N:

[0018]

[0019] The first row parameter d i ′ in the observation space matrix represents the relative distance between the air defense unit agent and the i-th air raid unit. The calculation method is:

[0020]

[0021] Among them, d i is the Euclidean distance between the air defense unit and the air raid unit, and r is the effective combat range of the air defense unit; when d i ′ < 1, it means that the enemy is within the strike range of the air defense unit agent;

[0022] The second row parameter Δθ in the observation space matrix irepresents the change value of the entry angle of the i-th air raid unit relative to the air defense unit agent; let the current time be t and the time interval be Δt, then the calculation method of the entry angle change value is as follows:

[0023]

[0024] In the formula, θ is the entry angle, which is the angle between the line connecting the air raid unit and the air defense unit agent and the x-axis direction; when Δθ i = 0, it means that the target of the i-th air raid unit at this time is the current air defense unit agent;

[0025] Determine the third row parameter m in the space matrix i represents how many air defense unit agents have targeted and locked the i-th air raid unit; this parameter is generated by the communication between air defense unit agents. If the targeting target of the current air defense unit agent is j, then m j takes a negative number, indicating that the current target has been targeted by this agent;

[0026] Determine the fourth row parameter s in the space matrix i represents the survival status of the i-th air raid unit; if the air raid unit is still alive, its value is set to 1, and if the air raid unit has been shot down, its value is set to -1;

[0027] Determine the fifth row parameter Δv in the space matrix i represents the value of firepower transfer; assume that the targeting target of the agent at time (t - Δt) is the j-th air raid unit with a value of v j ; at time t, the agent transfers the targeting target to the i-th air raid unit with a value of v i ; then the calculation method of the firepower transfer value in this case is:

[0028]

[0029] That is, the transfer value is equal to the ratio of the planned transfer target value to the current relative distance minus the ratio of the previous targeting target value to the previous relative distance, which is used to evaluate the rationality of the air defense unit agent's transfer of the targeting target.

[0030] Preferably, step 3 is specifically as follows:

[0031] Step 3.1: Determine the hit reward. When the agent successfully shoots down the i-th air raid unit, give it a reward:

[0032] r = 30v i (6)

[0033] That is, 30 times the value of the i-th air raid unit;

[0034] Step 3.2: Determine the firepower transfer reward. When the agent attempts to transfer from the state of aiming at the $i$-th air raid unit to aiming at the $j$-th air raid unit, give it a reward or punishment:

[0035] $r = 1.5\Delta v$ j -0.6 (7)

[0036] Step 3.3: Determine the penalty for ineffective aiming. When the agent attempts to aim at an air raid unit that has already been shot down, give it a penalty of $r = -1$;

[0037] Step 3.4: Aiming maintenance reward. When the aiming target of the agent at the current moment is the same as that at the previous moment, that is, when the actions at adjacent moments are the same:

[0038] $a$ t $= a$ t+Δt (8)

[0039] Give the agent a reward of $r = 0.3$, so as to encourage the agent to focus on aiming at the air raid unit with higher threat for a long time;

[0040] Step 3.5: Process end reward. When the process ends with all air raid units being shot down by the air defense unit agents, give all air defense unit agents a reward of $r = 10$. When the process ends with one or more defended targets being destroyed by the air raid units, give it a penalty:

[0041] $r = -50n$ (9)

[0042] where $n$ is the number of defended targets destroyed at the end of the process, and it is 1.

[0043] Preferably, the specific content of step 4 is as follows:

[0044] Add 6 air defense units; and a total of 8 enemy air raid units are added, which lock three defended targets and three long-range air defense units respectively, and the rest lock two important defended targets A and B among the defended targets;

[0045] Before the start of each round of the process, it is necessary to initialize the attributes of the agents and air raid units in the environment; for the air defense unit agents, the initialization operations mainly include: clearing the radar data, initializing the fire interval state, and initializing the firepower transfer state; for the enemy air raid units, the initialization operations mainly include: selecting the strike target of the air raid unit, randomly generating the starting position, and initializing the survival state. To ensure that the strategies generated by the agents have a certain universality and prevent overfitting, the $x$-axis coordinate in the initialization of the starting position of the air raid unit is randomly generated.

[0046] Preferably, the specific content of step 5 is as follows:

[0047] After the state space, reward function, and environment process design are completed, the MADDPG reinforcement learning algorithm can be used to train the air defense unit agents;

[0048] Among them, for the construction of the Actor and Critic neural networks, the A-C network adopts a multi-input neural network design: the observation space matrix is divided into 5 vectors of size 1×8 and respectively input into 5 fully connected layers. After passing through one or more hidden layers, the output neurons are concatenated together and then input into the next hidden layer;

[0049] By using the Gumbel-Softmax method to sample from the discrete distribution to make it differentiable, the Gumbel-Softmax sampling is written as:

[0050]

[0051] Among them, the parameter τ is called the temperature parameter of the distribution, and this parameter is always greater than 0, and g i is a noise sampled from Gumbel(0,1), called the reparameterization factor:

[0052] g i = -log(-logu), u~Uniform(0,1) (11).

[0053] Preferably, the specific step 6 is as follows:

[0054] Input the deployment coordinates of the long-range air defense units and medium-range air defense units in each deployment plan generated by the multi-objective optimization algorithm into the air defense battlefield environment, use the trained agent strategy to simulate the air defense battlefield, obtain the shooting down rates under different deployment plans, so as to evaluate the performance of the overall potential deployment plan and screen out the best individuals among them.

[0055] A multi-objective multi-element allocation decision-making system based on multi-agents, the system includes:

[0056] A simulation module, which conducts air defense scenario modeling and simulates an air raid scenario;

[0057] A construction module, which constructs the agent action and observation space and observes the space matrix;

[0058] A reward and punishment module, which establishes a reward and punishment function so that the agent will be given a reward value according to the action feedback after each action;

[0059] A design module, which designs the air defense environment process and locks the protected target;

[0060] An optimization module, which performs MADDPG reinforcement learning to generate a policy, optimize the air defense policy, and perform multi-objective optimization to solve and deploy a solution set.

[0061] A computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement a multi-objective and multi-factor allocation decision-making method based on multi-agent.

[0062] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements a multi-objective and multi-factor allocation decision-making method based on multi-agent.

[0063] The present invention has the following beneficial effects:

[0064] Compared with the prior art, the present invention:

[0065] For the multi-objective optimization solution set of a certain real scenario, the present invention first needs to abstract the complex real system into a simulation model of the interaction between multiple agents. In this process, the multi-factor requirements of the real scenario (such as resource allocation, time constraints, space limitations, etc.) are analyzed in detail and incorporated into the system design. Each agent represents a sub-part of the system and interacts with other agents through rules and algorithms. To ensure that the multi-agent system can effectively coordinate and make decisions, the present invention needs to carefully set the reward function for each agent, which not only considers the performance of a single agent but also comprehensively considers the mutual influence between multiple agents and the overall system objectives.

[0066] Next, the multi-agent system is trained with the aim of optimizing its behavioral policy through multiple simulations and feedback, and during the continuous adjustment process, ensuring that the system can meet the expected performance standards when facing a complex environment. During the training process, the present invention will continuously observe the performance of the agents in the simulation environment and evaluate their decision-making effects and system efficiency in different situations.

[0067] Finally, the present invention conducts inference tests on each individual solution in the multi-objective optimization solution set to further verify its actual performance. At this stage, through multiple rounds of simulation experiments, the present invention can select the solution that best meets the target requirements from the optimization solution set and use it as the optimal decision. In this way, the present invention can find a relatively ideal solution in a complex and changeable real scenario, ensuring that the system can achieve the best performance in all key indicators.

[0068] The present invention solves the problem that it is difficult to screen the deployment solution set in the traditional air defense deployment multi-objective optimization, and innovatively uses the target allocation strategy obtained by reinforcement learning training to solve the problem. The problem is abstracted into multiple objective functions to be optimized, and the deployment solution set is solved by using the multi-objective optimization algorithm; for the evaluation and screening of the deployment solution set, it is proposed to abstract a more realistic air defense battlefield environment based on the deployment model and use the multi-agent reinforcement learning algorithm to train the agent to obtain a dynamic target allocation strategy, and the strategy is used to realize the evaluation and screening of the deployment plan.

[0069] Through simulation analysis, it can be seen that the adopted reinforcement learning-based target allocation strategy air defense deployment Pareto solution set rapid screening method can realize the rapid screening of air defense deployment Pareto solution set. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0071] Figure 1 Shown is a flow chart for establishing a multi-agent simulation model of the present invention;

[0072] Figure 2 Shown is a schematic diagram of the air attack unit entry angle of the present invention;

[0073] Figure 3 Shown is a schematic diagram of the air strike unit action flow chart of the present invention;

[0074] Figure 4 Shown is a flow chart of the air defense unit firing determination of the present invention;

[0075] Figure 5 Shown is a schematic diagram of the AC network design of the present invention;

[0076] Figure 6 Shown is a schematic diagram of a scatter plot of reward values of the present invention;

[0077] Figure 7 Shown is a schematic diagram of a scatter plot of the number of knockdowns of the present invention;

[0078] Figure 8 Shown is a scatter plot of the number of shots of the present invention;

[0079] Figure 9 Shown is a scatter diagram of the number of firepower transfers of the present invention;

[0080] Figure 10 Shown as the evaluation of the kill rate of the present invention;

[0081] Figure 11 Shown as the test histogram of the present invention;

[0082] Figure 12 Shown as the evaluation of the optimal deployment plan of the present invention. Specific embodiments

[0083] The following combines specific embodiments to elaborate on the present invention in detail. Specific Embodiment 1:

[0085] According to Figures 1 to 12 As shown, the specific optimization technical solution adopted by the present invention to solve the above technical problems is: The present invention relates to a multi-objective and multi-element allocation decision-making method based on multi-agent.

[0086] A multi-objective and multi-element allocation decision-making method based on multi-agent, the method includes the following steps: The method includes the following steps:

[0087] Step 1: Conduct air defense scenario modeling to simulate an air attack scenario;

[0088] Step 2: Build the agent action and observation space and observe the space matrix;

[0089] Step 3: Establish a reward and punishment function so that the agent will be given a reward value according to the action feedback to the agent after each action is executed;

[0090] Step 4: Design the air defense environment process to lock the protected target;

[0091] Step 5: Conduct MADDPG reinforcement learning to generate a strategy, optimize the air defense strategy, and conduct multi-objective optimization to solve the deployment solution set.

[0092] For the situation where it is difficult to screen the deployment solution set in the traditional air defense deployment multi-objective optimization, the present invention innovatively uses the target allocation strategy obtained through reinforcement learning training to solve this problem. By abstracting this problem into multiple objective functions to be optimized, the multi-objective optimization algorithm is used to solve its deployment solution set; for the evaluation and screening problem of the deployment solution set, a relatively realistic air defense battlefield environment is abstracted on the basis of the deployment model, and the multi-agent reinforcement learning algorithm is used to train the agent to obtain a dynamic target allocation strategy, and this strategy is used to realize the evaluation and screening of the deployment plan.

[0093] Through simulation analysis, it can be known that the adopted method for quickly screening the Pareto solution set of air defense deployment based on the reinforcement learning-based target allocation strategy can realize the quick screening of the Pareto solution set of air defense deployment. Specific Embodiment 2:

[0095] The difference between the second embodiment and the first embodiment of this application is only that:

[0096] The specific content of step 1 is as follows:

[0097] Suppose there are three protected targets A, B, and C, where A and B are important protected targets with a higher protection weight than C; the three protected targets are located in an area densely distributed with protected objects in the population gathering area and the factory under attack. This area is regarded as a circular protected object with a radius of 125 km. The center of the circular area is used as the coordinate origin, the due north direction is the positive direction of the y-axis, and the due east direction is the positive direction of the x-axis; the incoming direction is due north, that is, from the positive direction of the y-axis, and the incoming method is set to fly horizontally at a certain height. Each unit randomly locks a strike target and flies towards it at a constant speed. Specific embodiment three:

[0099] The difference between the third embodiment and the second embodiment of this application is only that:

[0100] The specific content of step 2 is as follows:

[0101] There are N enemies above the current air defense. Define the action space that the air defense unit can only perform as:

[0102] action = discrete(a1, a2,..., a N ) T (1)

[0103] Among them, discrete indicates that this action space is a discrete space, that is, under this space, the action a i is a binary function with values of 0 or 1;

[0104] When its value is 1, it means that the air defense unit agent will turn its aiming target to the i-th air raid unit of the enemy at this moment; on the contrary, 0 means not to turn to this unit. At the same moment, an air defense unit can only aim at one enemy air raid unit;

[0105] For the information input to the agent at each moment, a observation space matrix is constructed, which is a matrix of size 5×N:

[0106]

[0107] The first row parameter d in the observation space matrix i ′ represents the relative distance between the air defense unit agent and the i-th air raid unit, and the calculation method is:

[0108]

[0109] Among them, d iis the Euclidean distance between the air defense unit and the air raid unit, and r is the effective combat range of the air defense unit; when d i ′ < 1, it means that the enemy is within the strike range of the air defense unit agent;

[0110] The second row parameter Δθ in the observation space matrix i represents the change value of the entry angle of the i-th air raid unit relative to the air defense unit agent; let the current time be t and the time interval be Δt, then the calculation method of the entry angle change value is:

[0111]

[0112] In the formula, θ is the entry angle, which is the angle between the line connecting the air raid unit and the air defense unit agent and the x-axis direction; when Δθ i = 0, it means that the strike target of the i-th air raid unit at this time is the current air defense unit agent;

[0113] Determine the third row parameter m in the space matrix i represents how many air defense unit agents the i-th air raid unit has been targeted and locked by; this parameter is generated by the communication between air defense unit agents. If the current air defense unit agent's targeting target is j, then m j takes a negative number, indicating that the current target has been targeted by this agent;

[0114] Determine the fourth row parameter s in the space matrix i represents the survival status of the i-th air raid unit; if the air raid unit is still alive, its value is set to 1, and if the air raid unit has been shot down, its value is set to -1;

[0115] Determine the fifth row parameter Δv in the space matrix i represents the firepower transfer value; assume that at the time (t - Δt), the targeting target of the agent is the j-th air raid unit with a value of v j ; at time t, the agent transfers the targeting target to the i-th air raid unit with a value of v i ; then the calculation method of the firepower transfer value in this case is:

[0116]

[0117] That is, the transfer value is equal to the ratio of the planned transfer target value to the current relative distance minus the ratio of the previous targeting target value to the previous relative distance, which is used to evaluate the rationality of the air defense unit agent's transfer of the targeting target. Specific Embodiment Four:

[0119] The difference between Embodiment Four and Embodiment Three of this application is only that:

[0120] The specific content of step 3 is:

[0121] Step 3.1: Determine the hit reward. When the agent successfully shoots down the i-th air raid unit, give it a reward:

[0122] r = 30v i (6)

[0123] That is, 30 times the value of the i-th air raid unit;

[0124] Step 3.2: Determine the firepower transfer reward. When the agent attempts to transfer from the state of aiming at the i-th air raid unit to aiming at the j-th air raid unit, give it a reward or punishment:

[0125] r = 1.5Δv j -0.6 (7)

[0126] Step 3.3: Determine the invalid aiming penalty. When the agent attempts to aim at an air raid unit that has already been shot down, give it a penalty of r = -1;

[0127] Step 3.4: Aiming hold reward. When the aiming target of the agent at the current moment is the same as that at the previous moment, that is, when the actions at adjacent moments are the same:

[0128] a t = a t+Δt (8)

[0129] Give the agent a reward of r = 0.3 to encourage the agent to focus on aiming at the air raid unit with a higher threat for a long time;

[0130] Step 3.5: Process end reward. When the process ends with all air raid units being shot down by the air defense unit agent, give all air defense unit agents a reward of r = 10. When the process ends with one or more defended targets being destroyed by the air raid units, give it a penalty:

[0131] r = -50n (9)

[0132] Where n is the number of defended targets destroyed at the end of the process, which is 1. Specific Embodiment Five:

[0134] The difference between Embodiment Five and Embodiment Four of the present invention is only that:

[0135] The specific content of Step 4 is as follows:

[0136] Add 6 air defense units; and a total of 8 enemy air raid units are added, which lock three defended targets and three long-range air defense units respectively, and the remaining two lock two important defended targets A and B among the defended targets;

[0137] Before the start of each round of the process, it is necessary to initialize the attributes of the agents and air raid units in the environment; for the air defense unit agents, the initialization operations mainly include: clearing the radar data, initializing the fire interval state, and initializing the fire transfer state; for the enemy air raid units, the initialization operations mainly include: selecting a strike target for the air raid unit, randomly generating the starting position, and initializing the survival state. To ensure the generality of the strategies generated by the agents and prevent overfitting, the x-axis coordinate in the initialization of the starting position of the air raid unit is randomly generated. Specific Embodiment Six:

[0139] The difference between Embodiment Six and Embodiment Five of the present invention lies only in:

[0140] Step 5 is specifically:

[0141] After completing the design of the state space, reward function, and environmental process, the MADDPG reinforcement learning algorithm can be used to train the air defense unit agents;

[0142] Among them, for the construction of the Actor and Critic neural networks, the A-C network adopts a multi-input neural network design: the observation space matrix is divided into 5 vectors of size 1×8, which are respectively input into 5 fully connected layers. After passing through one or more hidden layers, the output neurons are concatenated together and then input into the next hidden layer;

[0143] By using the Gumbel-Softmax method to sample from a discrete distribution to make it differentiable, the Gumbel-Softmax sampling is written as:

[0144]

[0145] Among them, the parameter τ is called the temperature parameter of the distribution, and this parameter is always greater than 0, and g i is a noise sampled from Gumbel(0,1), which is called the reparameterization factor:

[0146] g i =-log(-logu), u~Uniform(0,1) (11). Specific Embodiment Seven:

[0148] The difference between Embodiment Seven and Embodiment Six of the present invention lies only in:

[0149] Step 6 is specifically:

[0150] Input the deployment coordinates of the long-range air defense units and medium-range air defense units in each deployment plan generated by the multi-objective optimization algorithm into the air defense battlefield environment, and use the intelligent agent strategy obtained through training to simulate the air defense battlefield, so as to obtain the shooting down rate under different deployment plans, thereby evaluating the performance of the overall potential deployment plan and screening out the best individuals among them. Specific Embodiment VIII:

[0152] The difference between Embodiment VIII and Embodiment VII of the present invention lies only in:

[0153] The present invention provides a multi-objective multi-element allocation decision-making system based on multi-intelligent agents, and the system includes:

[0154] A simulation module that models the air defense scenario and simulates the air raid scenario;

[0155] A construction module that constructs the intelligent agent action and observation space and observes the space matrix;

[0156] A reward and punishment module that establishes a reward and punishment function so that after each action is executed by the intelligent agent, a reward value will be given to the intelligent agent according to the action feedback;

[0157] A design module that designs the air defense environment process and locks the protected target;

[0158] An optimization module that performs MADDPG reinforcement learning to generate a strategy, optimizes the air defense strategy, and performs multi-objective optimization to solve the deployment solution set. Specific Embodiment IX:

[0160] The difference between Embodiment IX and Embodiment VIII of the present invention lies only in:

[0161] The present invention provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement a multi-objective multi-element allocation decision-making method based on multi-intelligent agents. Specific Embodiment X:

[0163] The difference between Embodiment X and Embodiment IX of the present invention lies only in:

[0164] The present invention provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements a multi-objective multi-element allocation decision-making method based on multi-intelligent agents. Specific Embodiment XI:

[0166] The difference between Embodiment XI and Embodiment X of the present invention lies only in:

[0167] Taking the air defense deployment scenario as an example, the present invention will specifically demonstrate the implementation process of the multi-objective and multi-factor allocation decision-making method based on multi-agent. In this scenario, multiple agents represent different air defense units, which need to optimize the deployment of air defense forces and resource allocation according to real-time intelligence and enemy threats under limited resource conditions to maximize the defense effect of the system. Through this example, we can specifically elaborate on the application of the above method and show how to use the cooperation and optimization of agents to improve the system efficiency in complex decision-making processes.

[0168] Step 1: Air defense scenario modeling

[0169] Suppose there are currently three defense targets A, B, and C, where A and B are important defense targets with higher defense weights than C. In this scenario, the three defense targets are located in an area densely distributed with defense objects such as population gathering areas and factories that are more likely to be attacked. This area can be regarded as a circular defense object with a radius of 125 km. If the center of the circular area is taken as the coordinate origin, the positive y-axis direction is due north, and the positive x-axis direction is due east. Assume that the incoming direction is due north, that is, from the positive y-axis direction. The incoming mode is set to level flight at a certain altitude, and each unit randomly locks a strike target and flies towards it at a constant speed.

[0170] Step 2: Construction of agent action and observation spaces

[0171] Suppose there are N enemies above the current air defense. Then the action space that can be defined for the air defense unit is:

[0172] action = discrete(a1, a2,..., a N ) T (1)

[0173] Among them, discrete indicates that this action space is a discrete space, that is, under this space, the action a i is a binary function with values of 0 or 1. When its value is 1, it means that at this moment, the air defense unit agent turns its aiming target to the i-th enemy air raid unit; on the contrary, 0 means not turning to this unit. At the same moment, an air defense unit can only aim at one enemy air raid unit.

[0174] Since in the real air defense environment, the air defense system cannot have global information about the air defense environment and can only obtain partial observation information based on the operation of observation systems such as radars. Therefore, for the information input to the agent at each moment, a observation space matrix is constructed, which is a matrix of size 5×N:

[0175]

[0176] The first row of parameters d in the observation space matrixi d' represents the relative distance between the air defense unit agent and the i-th air raid unit, and the calculation method is as follows:

[0177]

[0178] In the formula, d i is the Euclidean distance between the air defense unit and the air raid unit, and r is the effective combat range of the air defense unit. When d i ' < 1, it means that the enemy is within the strike range of the air defense unit agent.

[0179] The second row parameter Δθ in the observation space matrix i represents the change value of the entry angle of the i-th air raid unit relative to the air defense unit agent. Let the current time be t and the time interval be Δt, then the calculation method of the entry angle change value is as follows:

[0180]

[0181] In the formula, θ is the entry angle, which is the angle between the line connecting the air raid unit and the air defense unit agent and the x-axis direction. When Δθ i = 0, it means that the strike target of the i-th air raid unit at this time is the current air defense unit agent. The longer the duration of this value being zero, the higher the risk factor of the current air defense unit agent.

[0182] The third row parameter m in the observation space matrix i represents how many air defense unit agents the i-th air raid unit has been targeted and locked by. This parameter is generated by the communication between air defense unit agents. If the targeting target of the current air defense unit agent is j, then m j takes a negative number, indicating that the current target has been targeted by this agent.

[0183] The fourth row parameter s in the observation space matrix i represents the survival status of the i-th air raid unit. If the air raid unit is still alive, its value is set to 1; if the air raid unit has been shot down, its value is set to -1.

[0184] The fifth row parameter Δv in the observation space matrix i represents the value of firepower transfer. Suppose the targeting target of the agent at (t - Δt) time is the j-th air raid unit, and the value is v j ; at time t, the agent transfers the targeting target to the i-th air raid unit, and the value is v i . Then the calculation method of the firepower transfer value in this case is as follows:

[0185]

[0186] That is, the transfer value is equal to the ratio of the planned transfer target value to the current relative distance minus the ratio of the aiming target value at the previous moment to the relative distance at the previous moment. This value is used to evaluate the rationality of the air defense unit agent's transfer and aiming at the target.

[0187] Step 3: Design of the reward and punishment function

[0188] Based on the characteristics of intensive air defense battlefield intelligence and rapid scenario changes, the reward and punishment function needs to be designed as a dense reward and punishment function, that is, after each action executed by the agent, a reward value will be given to the agent according to the action feedback. The reward and punishment function of the agent mainly includes the following items:

[0189] (1) Hit reward

[0190] When the agent successfully shoots down the i-th air raid unit, it is given the following reward:

[0191] r = 30v i (6)

[0192] That is, 30 times the value of the i-th air raid unit.

[0193] (2) Firepower transfer reward

[0194] When the agent attempts to transfer from the state of aiming at the i-th air raid unit to aiming at the j-th air raid unit, it is given a reward or punishment:

[0195] r = 1.5Δv j -0.6 (7)

[0196] This reward and punishment function consists of two parts: the first part is the reward part, that is, the 1.5Δv j item. The higher the firepower transfer value generated by the firepower transfer, the higher the reward value; the second part is the penalty item, that is, the -0.6 item, which represents the basic penalty for time waste caused by the existence of firepower transfer time during firepower transfer. To sum up, only when a firepower transfer action with a high enough value is generated will the system give the agent a positive reward feedback, thus preventing the agent from randomly aiming aimlessly.

[0197] (3) Penalty for ineffective aiming

[0198] When the agent attempts to aim at an air raid unit that has been shot down, it is given a penalty of r = -1.

[0199] (4) Reward for maintaining aiming

[0200] When the aiming target of the agent at the current moment is the same as that at the previous moment, that is, when the actions at adjacent moments are the same:

[0201] a t = a t+Δt(8)

[0202] Give the agent a reward of r = 0.3 to encourage the agent to focus on aiming at the air raid units with a higher threat level for a long time.

[0203] (5) End-of-process reward

[0204] When the process ends with all air raid units being shot down by the air defense unit agents, give all air defense unit agents a reward of r = 10. When the process ends with one or more defended targets being destroyed by the air raid units, give a penalty:

[0205] r = -50n (9)

[0206] In the formula, n is the number of defended targets destroyed at the end of the process, usually 1.

[0207] Step 4: Design of the air defense environment process

[0208] Due to the relatively short effective combat radius and short interval of the short- and medium-range air defense units, their resistance to air raid units is extremely poor, but their ability to strike air raid units on the near bombing line of unmanned aerial vehicles is extremely strong. Therefore, after the air raid units actually enter the formation, for the sake of certain simplification considerations, a total of 6 air defense units are added to this environment; and a total of 8 enemy air raid units are added, which respectively lock three defended targets and three long-range air defense units, and the remaining ones lock two important defended targets A and B among the defended targets.

[0209] Before the start of each round of the process, it is necessary to initialize the attributes of the agents and air raid units in the environment. For the air defense unit agents, the initialization operations mainly include: clearing the radar data, initializing the fire interval state, and initializing the fire transfer state; for the enemy air raid units, the initialization operations mainly include: selecting the strike target of the air raid unit, randomly generating the starting position, and initializing the survival state. Among them, to ensure that the strategies generated by the agents have a certain degree of universality and prevent overfitting, the x-axis coordinate in the initialization of the starting position of the air raid unit is randomly generated.

[0210] In each round of the process, the running step size is set to 1 second, and three actions are performed in each step, which are respectively:

[0211] (1) Actions of enemy air raid units

[0212] Each air raid unit first determines whether it has been shot down. If not, it updates the distance between itself and the target according to the speed. After the update, it determines whether the distance is less than or equal to the bombing line. If it is less than or equal to the bombing line, the process ends.

[0213] (2) Actions of air defense units

[0214] Each air defense unit agent generates actions based on the observation space matrix provided by the radar. First, a fire transfer determination is made. If the selected enemy target in the action is the same as the previous moment, the target is maintained; otherwise, it is returned that the fire has been transferred. Then, a firing determination is started. If it is not in the fire interval or fire transfer state at this time, and the enemy target is alive and within the shooting range, a hit determination is made; if the enemy is not alive, an invalid capture is returned.

[0215] (3) Step size settlement

[0216] After all enemy air raid units and air defense unit agents have completed their actions in a step, the environment will perform the settlement work. First, based on the information returned by the air defense unit agent during the action, the reward value feedback to each agent is calculated through the reward function formula given in step three; then, a process termination determination is made based on the bomb dropping and shot-down situations of the air raid units.

[0217] Step 5: MADDPG reinforcement learning algorithm generates strategies

[0218] After completing the design of the state space, reward function, and environmental process, the MADDPG reinforcement learning algorithm can be used to train the air defense unit agent.

[0219] Among them, for the construction of the Actor and Critic neural networks, since the observation space matrix of the air defense environment in this training is as shown in formula (2), which is a matrix of size 5×8. And generally, a fully connected network can only input a vector, and a convolutional neural network can input a matrix but this environment does not require the overall matrix recognition ability. Therefore, the A-C network in this environment adopts a multi-input neural network design. The method is as follows: The observation space matrix is divided into 5 vectors of size 1×8 and respectively input into 5 fully connected layers. After passing through one or more hidden layers, the output neurons are concatenated together and then input into the next hidden layer.

[0220] In addition, to solve the problem that the action space in the air defense environment is discrete, resulting in the non-differentiability of the agent's actions with respect to the policy parameters, sampling can be performed from the discrete distribution by using the Gumbel-Softmax method to make it differentiable. The Gumbel-Softmax sampling can be written as:

[0221]

[0222] In the formula, the parameter τ is called the temperature parameter of the distribution, and this parameter is always greater than 0. By adjusting this parameter, the approximation degree of the generated sampling distribution to the discrete distribution can be controlled. The larger τ is, the closer the sampling distribution is to the uniform distribution; the smaller τ is, the closer the sampling is to the one-hot result; g iA noise sampled from Gumbel(0,1), called the reparameterization factor:

[0223] g i = -log(-logu), u ~ Uniform(0,1) (11)

[0224] Step 6: Solve the deployment solution set using a multi-objective optimization algorithm

[0225] According to the air defense environment and actual requirements, design the corresponding objective function and constraints to form an optimization problem, and use a variety of multi-objective optimization algorithms to solve this optimization problem to obtain the solution set of deployment plans. Typical multi-objective optimization algorithms such as NSGA-III and AGE-MOEA2 algorithms.

[0226] The NSGA-III algorithm is based on the NSGA-II algorithm, introducing the concept of reference points, and guiding the distribution of solutions through uniformly distributed reference points. Each objective has a corresponding reference point in a standardized space, enabling the algorithm to obtain better uniformity when dealing with multiple objectives, so that the algorithm can solve optimization problems with a larger number of objective functions.

[0227] The AGE-MOEA2 algorithm uses a simple heuristic method to simulate the shape of the Pareto front, and calculates the evaluation score by introducing Newton's iterative method for root finding and geodesics to measure the distance of non-dominated solutions on the surface manifold, thus better ensuring the diversity and authenticity of the solution set.

[0228] Step 7: Evaluate and screen the deployment plans

[0229] Input the deployment coordinates of the long-range air defense units and medium-range air defense units in each deployment plan generated by the multi-objective optimization algorithm into the air defense battlefield environment, use the intelligent agent strategy obtained through training to simulate the air defense battlefield, and obtain the shooting down rate under different deployment plans, so as to evaluate the performance of the overall potential deployment plans and screen the best individuals among them.

[0230] Set the hyperparameters as follows, and obtain the target allocation strategy table 1 through training:

[0231] Table 1. Allocation strategy table

[0232]

[0233]

[0234] Among them, the deployment positions of the air defense unit agents are generated after regularization according to the general intuitive ideas in reality, and the deployment positions are shown in Table 2 below:

[0235] Table 2. Deployment position table

[0236]

[0237] Take the total number of training Episodes as Episodes = 7000. Let the reward value in each Episode be the sum of the reward values obtained by all agents in each step. Then, a scatter plot of the reward value versus the training Episode can be obtained as shown in Figure 6 shown. After approximately 700 Episodes, the reward value exceeds zero for the first time, indicating that there has been a situation where the air defense unit agent wins; after approximately 5000 Episodes, the curve of the reward value greater than 0 becomes denser and denser, and the efficiency of the agent in obtaining rewards has been improved.

[0238] Figure 7 This is a scatter plot of the number of enemy air raid units shot down by the air defense unit agent in each round of Episode. It can be seen that as the Episode increases, the density of the scatter points with a lower number of shots down gradually decreases, and the density of the scatter points with a higher number of shots down gradually becomes denser. This shows that the ability of each agent to cooperate with each other to fight against enemy air raid units is continuously improving, and after 1000 rounds of Episode, the situation of shooting down 0 aircraft no longer occurs.

[0239] Figure 8 This is a scatter plot of the number of firing times. It can be seen that it has a high correlation with the number of shots down, and its final convergence range is slightly higher than the number of shots down because the hit rate of the air defense unit is not 100%. It can be observed that there are occasional situations where the number of firing times is too high, which indicates that it is impossible to rule out the situation where the air defense unit fails to hit after multiple rounds of shooting. For such accidental events, some anti-aircraft guns can be considered for deployment as a firepower supplement and the final defense measure.

[0240] Figure 9 This is a scatter plot of the number of firepower transfer times. It can be seen that under the guidance of the reward and punishment function, the total number of firepower transfer times per round of Episode decreases rapidly and finally converges to around 75.

[0241] During the training process, after every 100 Episodes of training are completed, an evaluation is carried out. The evaluation will repeat 100 processes. The air defense environment remains unchanged, while the agents will select actions to execute in a decentralized manner. After 100 processes are completed, the average enemy air raid unit shooting down rate per round is calculated as the evaluation value. As shown in Figure 10 shown is the change situation of the evaluation value in 7000 Episodes. It can be seen that it has been on an upward trend since the start of training and converges to around 0.95 after approximately 5000 Episodes.

[0242] The AGE-MOEA2 multi-objective optimization algorithm is used to solve the optimal deployment problem. The population size is 100. The set of deployment plans is obtained and placed in the air defense environment. The strategies obtained through training are used to execute the process 100 times in a decentralized manner for each deployment plan, and the average kill rate is obtained. The kill rate histogram of the set of deployment plans can be obtained as Figure 11 shown, and the average kill rate is 0.2524. It can be seen that in the set of deployment plans, except for a small number of individuals with a lower kill rate distribution, most individuals are concentrated in the range of 0.25 - 0.35.

[0243] Among them, the best deployment plan is the 22nd individual, and the test kill rate is 0.3325. The deployment diagram is as Figure 12 shown, and the expected goal is achieved.

[0244] Through simulation analysis, it can be known that the method for quickly screening the Pareto solution set of air defense deployment using the target assignment strategy based on reinforcement learning can achieve the quick screening of the Pareto solution set of air defense deployment.

[0245] The above is only a preferred implementation manner of a multi-objective and multi-factor allocation decision method based on multi-agent. The protection scope of a multi-objective and multi-factor allocation decision method based on multi-agent is not limited to the above embodiments. All technical solutions within this idea belong to the protection scope of the present invention. It should be pointed out that for those skilled in the art, several improvements and changes made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A multi-objective and multi-factor allocation decision-making method based on multi-agent, characterized in that: The method includes the following steps: Step 1: Conduct air defense scenario modeling to simulate air attack scenarios; Step 2: Build the agent's action and observation space, and observe the space matrix; Step 3: Establish a reward and punishment function so that the agent will be given a reward value according to the action feedback after each action is executed; Step 4: Design the air defense environment process to lock the protected targets; Step 5: Perform MADDPG reinforcement learning to generate strategies, optimize the air defense strategy, and perform multi-objective optimization to solve and deploy the solution set.

2. The method according to claim 1, characterized in that: Specifically, Step 1 is as follows: Set three protected targets A, B, and C, where A and B are important protected targets with higher protection weights than C; the three protected targets are located in an area densely distributed with protected objects such as population gathering areas and factories under attack. This area is regarded as a circular protected object with a radius of 125 km. The center of the circular area is used as the coordinate origin, the due north direction is the positive direction of the y-axis, and the due east direction is the positive direction of the x-axis.

3. The method according to claim 2, characterized in that: Specifically, Step 2 is as follows: There are N enemies above the current air defense. Define the action space that the air defense unit can only take as: action = discrete(a1, a2,..., a N ) T (1) Among them, "discrete" indicates that the action space is a discrete space, that is, under this space, the action a i is a binary function with values of 0 or 1; For the information input to the agent at each moment, choose to construct an observation space matrix, which is a matrix of size 5×N: Observe the first row of parameters d in the observation space matrix i ' represents the relative distance between the air defense unit agent and the i-th air raid unit, and the calculation method is as follows: Observe the second row parameter Δθ in the observation space matrix i Represents the change value of the entry angle of the i-th air raid unit relative to the air defense unit agent; Let the current time be t and the time interval be Δt, then the calculation method of the entry angle change value is as follows: Determine the third row parameter m in the spatial matrix i Represents how many anti-aircraft unit agents are currently aiming and locking on the i-th air raid unit; this parameter is generated by the communication between anti-aircraft unit agents. If the current aiming target of the anti-aircraft unit agent is j, then m j Takes a negative number, representing that the current target has been aimed at by this agent; Determine the fourth row parameter s in the spatial matrix i Represents the survival status of the i-th air raid unit; if the air raid unit is still alive, its value is set to 1, and if the air raid unit has been shot down, its value is set to -1; Determine the fifth row parameter Δv in the spatial matrix i Represents the value of firepower transfer; assume that at time t - Δt, the aiming target of the agent is the j-th air raid unit with a value of v j ; at time t, the agent transfers the aiming target to the i-th air raid unit with a value of v i ; then the calculation method of the firepower transfer value in this case is as follows: That is, the transfer value is equal to the ratio of the planned transfer target value to the current relative distance minus the ratio of the aiming target value at the previous moment to the relative distance at the previous moment, which is used to evaluate the rationality of the air defense unit agent's transfer and aiming at the target.

4. The method according to claim 3, characterized in that: Specifically, Step 3 is as follows: Step 3.1: Determine the hit reward. When the agent successfully shoots down the i-th air raid unit, give it a reward: r=30v i (6) That is, 30 times the value of the i-th air raid unit; Step 3.2: Determine the firepower transfer reward. When the agent tries to transfer from the state of aiming at the i-th air raid unit to aiming at the j-th air raid unit, give it a reward or punishment: r = 1.5Δv j -0.6 (7) Step 3.3: Determine the invalid aiming penalty. When the agent tries to aim at an air raid unit that has been shot down, give it a penalty of r=-1; Step 3.4: Aiming hold reward. When the agent's aiming target at the current moment is the same as that at the previous moment, that is, when the actions at adjacent moments are the same: a t = a t+Δt (8) Give the agent a reward of r = 0.3 to encourage the agent to focus on aiming at the air raid unit with a higher threat for a long time; Step 3.5: Process end reward. When the process ends with the air defense unit agent shooting down all air raid units, give all air defense unit agents a reward of r = 10. When the process ends with one or more protected targets being destroyed by air raid units, give it a penalty: r=-50n (9) Among them, n is the number of protected targets destroyed at the end of the process, which is 1.

5. The method according to claim 4, characterized in that: Specifically, Step 4 is as follows: Add 6 air defense units; and a total of 8 enemy air raid units are added, which lock three protected targets and three long-range air defense units respectively, and the rest lock two important protected targets A and B among the protected targets; Before the start of each round of the process, it is necessary to initialize the attributes of the agents and air raid units in the environment.

6. The method according to claim 5, characterized in that: Specifically, Step 5 is as follows: For the construction of the Actor and Critic neural networks, the A-C network adopts a multi-input neural network design: the observation space matrix is divided into 5 vectors of size 1×8 and respectively input into 5 fully connected layers. After passing through one or more hidden layers, the output neurons are concatenated together and then input into the next hidden layer; By using the Gumbel-Softmax method to sample from the discrete distribution to make it differentiable, the Gumbel-Softmax sampling is written as: Among them, the parameter τ is called the temperature parameter of the distribution, and this parameter is always greater than 0. g i is a noise sampled from Gumbel(0,1), which is called the reparameterization factor: g i = -log(-log u), where u ~ Uniform(0,1) (11).

7. The method according to claim 5, characterized in that: Specifically, step 6 is as follows: The deployment coordinates of the remote air defense units and medium-range air defense units in each deployment plan generated by the multi-objective optimization algorithm are input into the air defense battlefield environment. The intelligent agent strategy obtained through training is used to simulate the air defense battlefield, and the hit rates under different deployment plans are obtained, so as to evaluate the performance of the overall potential deployment plan and screen the best individuals among them.

8. A multi-objective and multi-factor allocation decision-making system based on multi-agent, characterized in that: The system includes: A simulation module that models air defense scenarios and simulates air raid scenarios; A construction module that constructs the intelligent agent action and observation space and observes the space matrix; A reward and punishment module that establishes a reward and punishment function so that the intelligent agent will be given a reward value according to the action after each action is executed; A design module that designs the air defense environment process and locks the protected target; An optimization module that performs MADDPG reinforcement learning to generate a strategy, optimize the air defense strategy, and perform multi-objective optimization to solve the deployment solution set.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method as claimed in claims 1-7.

10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, the method as claimed in claims 1-7 is implemented.

Citation Information

Patent Citations

  • Dynamic air multi-target distribution and strike method based on multi-agent reinforcement learning

    CN116956705A

  • Unmanned aerial vehicle cooperative air combat decision-making method based on GRU-MAPPO deep reinforcement learning

    CN119129413A

  • Cooperative multi-goal, multi-agent, multi-stage reinforcement learning

    US20200160168A1

  • Multi-agent reinforcement learning scheduling method and system and electronic device

    WO2020181896A1

Cited By

  • Green way site selection optimization method and device based on graph neural network and reinforcement learning

    CN121189686A