A multi-agent based multi-objective multi-factor allocation decision method

By using multi-agent system air defense scenario modeling and reinforcement learning, the problem of multi-objective optimization that traditional methods cannot take into account multiple factors has been solved. This enables the rapid selection of the optimal air defense deployment strategy in complex environments, thereby improving the efficiency of air defense decision-making.

CN120337548BActive Publication Date: 2026-04-24HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2025-04-07
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing multi-objective optimization algorithms cannot take into account multiple factors in the real environment when selecting the optimal solution, resulting in the selected solution set failing to guarantee actual performance.

Method used

A multi-agent, multi-objective, multi-factor allocation decision-making method is adopted. Through air defense scenario modeling, agent action and observation space construction, reward and punishment function design, and MADDPG reinforcement learning, an optimized air defense strategy is generated to select the Pareto solution set.

Benefits of technology

It enables the rapid selection of the optimal solution in complex environments, ensuring that the system achieves optimal performance in various key indicators and improving the decision-making efficiency of air defense deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337548B_ABST
    Figure CN120337548B_ABST
Patent Text Reader

Abstract

The application is a multi-agent-based multi-target multi-element allocation decision method. The application relates to the technical field of optimization solution set decision, and innovatively uses a target allocation strategy obtained based on reinforcement learning training to solve the problem that a traditional air defense deployment multi-target optimization is difficult to screen a deployment solution set. The problem is abstracted into multiple target functions to be optimized, so that a multi-target optimization algorithm is used to solve the deployment solution set. In view of the evaluation and screening problem of the deployment solution set, a relatively real air defense battlefield environment is abstracted on the basis of a deployment model, a multi-agent reinforcement learning algorithm is used to train the agent to obtain a dynamic target allocation strategy, and the strategy is used to realize the evaluation and screening of the deployment scheme. Through simulation analysis, it can be known that the air defense deployment Pareto solution set rapid screening method using the target allocation strategy based on reinforcement learning can realize the rapid screening of the air defense deployment Pareto solution set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optimization solution set decision technology, and is a multi-agent, multi-objective, multi-factor allocation decision method. Background Technology

[0002] For complex, large-scale multi-objective optimization problems, multi-objective optimization algorithms are generally used to ensure that each objective function is not ignored by the algorithm during optimization, thus preventing the solution set from falling into extreme cases. However, the deployment scheme obtained by multi-objective optimization algorithms is not a single optimal solution, but a solution set composed of all dominant solutions. Therefore, for practical applications, after obtaining the Pareto solution set, it is still necessary to select the individual solutions that best meet the actual needs from the solution set based on the actual environment, i.e., the decision problem within the Pareto solution set. However, for decision problems, existing traditional methods such as weighted methods, ideal point methods, and constraint methods cannot reproduce the actual scenario, and the actual performance of the selected optimal solution cannot be guaranteed. Furthermore, the selection process cannot take into account all factors of the actual scenario. Therefore, an algorithm that can quickly select the optimal solution from the solution set based on the actual environment and considering multiple factors is urgently needed. Summary of the Invention

[0003] This invention addresses the problems of traditional optimization decision-making methods failing to reproduce real-world scenarios and the inability to guarantee the actual performance of the selected optimal solution. This invention provides a multi-agent, multi-objective, multi-factor allocation decision-making method, and offers the following technical solutions:

[0004] A multi-agent, multi-objective, multi-factor allocation decision-making method, the method comprising the following steps:

[0005] Step 1: Conduct air defense scenario modeling to simulate aerial attack scenarios;

[0006] Step 2: Construct the action and observation space of the intelligent agent, and observe the spatial matrix;

[0007] Step 3: Establish a reward and punishment function so that the agent receives a reward value after each action.

[0008] Step 4: Design the air defense environment process and identify the targets to be protected;

[0009] Step 5: Perform MADDPG reinforcement learning to generate strategies, optimize air defense strategies, and perform multi-objective optimization to solve for the deployment solution set.

[0010] Preferably, step 1 specifically comprises:

[0011] Three targets, A, B, and C, are designated for defense. A and B are considered important targets with a higher defense weight than C. These three targets are located in an area densely populated with population centers and factories, which are considered targets of attack. This area is regarded as a circular target with a radius of 125 km. The center of the circular area is taken as the origin of the coordinate system, with due north as the positive y-axis and due east as the positive x-axis. The direction of attack is due north, i.e., from the positive y-axis. The attack mode is assumed to be level flight at a certain altitude. Each unit randomly locks onto a target and flies towards it at a constant speed.

[0012] Preferably, step 2 specifically comprises:

[0013] There are currently N enemy targets above the air defense system. The action space that an air defense unit can only perform is defined as:

[0014] action = discrete(a1, a2, ..., a N ) T (1)

[0015] Here, "discrete" indicates that the action space is a discrete space, that is, action a is performed in this space. i It is a binary function that takes the value 0 or 1;

[0016] When its value is 1, it means that the air defense unit agent will turn its aiming target towards the i-th enemy air attack unit at this moment; conversely, 0 means that it will not turn to that unit. At the same time, an air defense unit can only aim at one enemy air attack unit.

[0017] For the information input to the agent at each time step, an observation space matrix is ​​constructed, which is a matrix of size 5×N:

[0018]

[0019] Observe the first row of parameters d in the spatial matrix i ′ represents the relative distance between the air defense unit agent and the i-th air attack unit, calculated as follows:

[0020]

[0021] Where, d i Let d be the Euclidean distance between the air defense unit and the air attack unit, and r be the effective combat range of the air defense unit; when d i When ′ < 1, it means that the enemy is already within the attack range of the air defense unit's intelligent agent;

[0022] Observe the second row of parameters Δθ in the spatial matrix iLet represent the change in the entry angle of the i-th air attack unit relative to the air defense unit agent; assuming the current time is t and the timing interval is Δt, the method for calculating the change in the entry angle is as follows:

[0023]

[0024] In the formula, θ is the entry angle, which is the angle between the line connecting the air attack unit and the air defense unit intelligent agent and the x-axis direction; when Δθ i When =0, it means that the target of the i-th air attack unit is the current air defense unit agent;

[0025] Determine the parameter m in the third row of the space matrix i This represents the number of air defense agents currently targeting and locking onto the i-th air attack unit; this parameter is generated through communication between air defense agents. If the current target of the air defense agent is j, then m... j A negative number indicates that the target has been targeted by this agent.

[0026] Determine the parameter s in the fourth row of the space matrix i This represents the survival status of the i-th air strike unit; if the air strike unit is still alive, its value is set to 1, and if the air strike unit has been shot down, its value is set to -1.

[0027] Determine the parameter Δv in the fifth row of the space matrix. i Represents the value of firepower transfer; let the target of the agent at time (t-Δt) be the j-th air attack unit, and its value be v. j At time t, the agent shifts its target to the i-th air attack unit, with a value of v. i Therefore, the method for calculating the value of firepower transfer in this situation is as follows:

[0028]

[0029] In other words, the transfer value is equal to the ratio of the planned transfer target value to the current relative distance, minus the ratio of the target value at the previous moment to the relative distance at the previous moment. It is used to evaluate the rationality of the air defense unit's intelligent agent transferring the target.

[0030] Preferably, step 3 specifically comprises:

[0031] Step 3.1: Determine the hit reward. When the agent successfully shoots down the i-th air attack unit, it is given a reward:

[0032] r = 30v i (6)

[0033] That is, 30 times the value of the i-th air strike unit;

[0034] Step 3.2: Determine the reward for firepower transfer. When an agent attempts to transfer its firepower from targeting the i-th air attack unit to targeting the j-th air attack unit, it is given a reward or penalty:

[0035] r = 1.5Δv j -0.6 (7)

[0036] Step 3.3: Determine the penalty for invalid aiming. When an agent attempts to aim at an air attack unit that has been shot down, it is given a penalty of r = -1.

[0037] Step 3.4: Target retention reward: When the agent's target is the same in the current and previous time steps, i.e., when actions in adjacent time steps are identical:

[0038] a t =a t+Δt (8)

[0039] The agent is given a reward of r=0.3 to encourage it to focus on targeting high-threat air units for extended periods.

[0040] Step 3.5: Process End Reward. When the process ends with the air defense unit agents shooting down all air attack units, all air defense unit agents are given a reward of r=10. If the process ends with one or more protected targets being destroyed by air attack units, a penalty is imposed.

[0041] r = -50n (9)

[0042] Where n is the number of protected targets destroyed at the end of the process, which is 1.

[0043] Preferably, step 4 specifically comprises:

[0044] Six air defense units were added; while eight enemy air attack units were added, locking onto three protected targets and three long-range air defense units respectively, and the rest locking onto two important protected targets A and B among the protected targets;

[0045] Before each round of the process begins, the attributes of the agents and air attack units in the environment need to be initialized. For the air defense unit agents, the initialization operations mainly include: clearing radar data, initializing the fire interval state, and initializing the fire transfer state. For the enemy air attack units, the initialization operations mainly include: the air attack units selecting targets, randomly generating the starting position, and initializing the survival state. In order to ensure that the strategy for generating agents has a certain degree of universality and to prevent overfitting, the x-axis coordinate in the initial position initialization of the air attack units is randomly generated.

[0046] Preferably, step 5 specifically comprises:

[0047] After completing the design of the state space, reward and punishment function, and environmental process, the MADDPG reinforcement learning algorithm can be used to train the air defense unit agent.

[0048] For the construction of the Actor and Critic neural networks, the AC network adopts a multi-input neural network design: the observation space matrix is ​​divided into 5 vectors of size 1×8, which are respectively input into 5 fully connected layers. After passing through one or more hidden layers, the output neurons are spliced ​​together and then input into the next hidden layer.

[0049] By sampling from a discrete distribution using the Gumbel-Softmax method to make it differentiable, Gumbel-Softmax sampling can be written as:

[0050]

[0051] Wherein, the parameter τ is called the temperature parameter of the distribution, and this parameter is always greater than 0, g i The noise sampled from Gumbel(0,1) is called the reparameter factor:

[0052] g i =-log(-logu),u~Uniform(0,1) (11).

[0053] Preferably, step 6 specifically comprises:

[0054] The deployment coordinates of long-range and medium-range air defense units in each deployment scheme generated by the multi-objective optimization algorithm are input into the air defense battlefield environment. The air defense battlefield is simulated using the intelligent agent strategy obtained through training. The kill rate under different deployment schemes is obtained, thereby evaluating the overall performance of potential deployment schemes and selecting the best individual among them.

[0055] A multi-agent, multi-objective, multi-factor allocation decision-making system, the system comprising:

[0056] The simulation module performs air defense scenario modeling and simulates air attack scenarios.

[0057] A construction module is used to construct the action and observation space of the intelligent agent and to observe the spatial matrix.

[0058] The reward and punishment module establishes a reward and punishment function so that after the agent performs each action, a reward value is fed back to the agent based on the action.

[0059] The design module designs the air defense environment process and locks onto the protected targets.

[0060] The optimization module uses MADDPG reinforcement learning to generate strategies, optimizes air defense strategies, and performs multi-objective optimization to solve for and deploy solution sets.

[0061] A computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement a multi-agent, multi-objective, multi-factor allocation decision-making method.

[0062] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a multi-agent, multi-objective, multi-factor allocation decision-making method.

[0063] The present invention has the following beneficial effects:

[0064] Compared with the prior art, the present invention:

[0065] This invention addresses multi-objective optimization solutions for a specific real-world scenario. First, it abstracts the complex real-world system into a simulation model of interactions between multiple agents. During this process, the various requirements of the real-world scenario (such as resource allocation, time constraints, and space limitations) are analyzed in detail and integrated into the system design. Each agent represents a sub-part of the system and interacts with other agents through rules and algorithms. To ensure effective coordination and decision-making within the multi-agent system, this invention carefully sets the reward function for each agent. These functions consider not only the performance of individual agents but also the mutual influence between multiple agents and the overall system objective.

[0066] Next, the multi-agent system is trained to optimize its behavioral strategies through multiple simulations and feedback, ensuring that it achieves the desired performance standards in complex environments through continuous adjustments. During training, this invention continuously observes the agents' performance in the simulation environment, evaluating their decision-making effectiveness and system efficiency in different scenarios.

[0067] Finally, this invention will conduct reasoning tests on each individual solution in the multi-objective optimization solution set to further verify its actual performance. At this stage, through multiple rounds of simulation experiments, this invention can select the solution that best meets the objective requirements from the optimization solution set and use it as the optimal decision. In this way, this invention can find a relatively ideal solution in complex and ever-changing real-world scenarios, ensuring that the system achieves optimal performance across all key indicators.

[0068] This invention addresses the challenge of selecting deployment solutions in traditional multi-objective optimization for air defense deployments by innovatively employing a target allocation strategy obtained through reinforcement learning training. The problem is abstracted into multiple objective functions to be optimized, allowing a multi-objective optimization algorithm to solve for the deployment solution set. For the evaluation and selection of the deployment solution set, a more realistic air defense battlefield environment is abstracted based on the deployment model. A dynamic target allocation strategy is obtained by training agents using a multi-agent reinforcement learning algorithm, and this strategy is then used to evaluate and select deployment schemes.

[0069] Simulation analysis shows that the reinforcement learning-based target allocation strategy for rapid selection of Pareto solutions for air defense deployment can achieve rapid selection of Pareto solutions for air defense deployment. Attached Figure Description

[0070] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0071] Figure 1 The flowchart shown is for establishing the multi-agent simulation model of the present invention.

[0072] Figure 2 The diagram shows the entry angle of the air raid unit according to the present invention.

[0073] Figure 3 The diagram shown is a schematic flowchart of the air raid unit operation of the present invention.

[0074] Figure 4 The flowchart shown is a flowchart of the air defense unit firing determination process of the present invention.

[0075] Figure 5 The diagram shown is a schematic representation of the AC network design of this invention.

[0076] Figure 6 The diagram shown is a scatter plot of the reward values ​​of this invention.

[0077] Figure 7 The diagram shown is a scatter plot of the number of kills according to the present invention.

[0078] Figure 8 The output is shown as a scatter plot of the firing count of this invention.

[0079] Figure 9 The diagram shows the number of fire transfers according to the present invention.

[0080] Figure 10 The displayed value represents the estimated kill rate of this invention.

[0081] Figure 11 The results are shown as the test histogram of this invention.

[0082] Figure 12 This is shown as the optimal deployment scheme for evaluating the present invention. Detailed Implementation

[0083] The present invention will be described in detail below with reference to specific embodiments. Specific Implementation Example 1:

[0085] according to Figures 1 to 12 As shown, the specific optimization technical solution adopted by the present invention to solve the above-mentioned technical problems is: The present invention relates to a multi-agent, multi-objective, multi-factor allocation decision-making method.

[0086] A multi-agent, multi-objective, multi-factor allocation decision-making method, the method comprising the following steps:

[0087] Step 1: Conduct air defense scenario modeling to simulate aerial attack scenarios;

[0088] Step 2: Construct the action and observation space of the intelligent agent, and observe the spatial matrix;

[0089] Step 3: Establish a reward and punishment function so that the agent receives a reward value after each action.

[0090] Step 4: Design the air defense environment process and identify the targets to be protected;

[0091] Step 5: Perform MADDPG reinforcement learning to generate strategies, optimize air defense strategies, and perform multi-objective optimization to solve for the deployment solution set.

[0092] This invention addresses the challenge of selecting deployment solutions in traditional multi-objective optimization for air defense deployments by innovatively employing a target allocation strategy obtained through reinforcement learning training. The problem is abstracted into multiple objective functions to be optimized, allowing a multi-objective optimization algorithm to solve for the deployment solution set. For the evaluation and selection of the deployment solution set, a more realistic air defense battlefield environment is abstracted based on the deployment model. A dynamic target allocation strategy is obtained by training agents using a multi-agent reinforcement learning algorithm, and this strategy is then used to evaluate and select deployment schemes.

[0093] Simulation analysis shows that the reinforcement learning-based target allocation strategy for rapid selection of Pareto solutions for air defense deployment can achieve rapid selection of Pareto solutions for air defense deployment. Specific Implementation Example 2:

[0095] The only difference between Embodiment 2 and Embodiment 1 of this application is that:

[0096] Step 1 specifically involves:

[0097] Three targets, A, B, and C, are designated for defense. A and B are considered important targets with a higher defense weight than C. These three targets are located in an area densely populated with population centers and factories, which are considered targets of attack. This area is regarded as a circular target with a radius of 125 km. The center of the circular area is taken as the origin of the coordinate system, with due north as the positive y-axis and due east as the positive x-axis. The direction of attack is due north, i.e., from the positive y-axis. The attack mode is assumed to be level flight at a certain altitude. Each unit randomly locks onto a target and flies towards it at a constant speed. Specific Implementation Example 3:

[0099] The only difference between Embodiment 3 and Embodiment 2 of this application is that:

[0100] Step 2 specifically involves:

[0101] There are currently N enemy targets above the air defense system. The action space that an air defense unit can only perform is defined as:

[0102] action = discrete(a1, a2, ..., a N ) T (1)

[0103] Here, "discrete" indicates that the action space is a discrete space, that is, action a is performed in this space. i It is a binary function that takes the value 0 or 1;

[0104] When its value is 1, it means that the air defense unit agent will turn its aiming target towards the i-th enemy air attack unit at this moment; conversely, 0 means that it will not turn to that unit. At the same time, an air defense unit can only aim at one enemy air attack unit.

[0105] For the information input to the agent at each time step, an observation space matrix is ​​constructed, which is a matrix of size 5×N:

[0106]

[0107] Observe the first row of parameters d in the spatial matrix i ′ represents the relative distance between the air defense unit agent and the i-th air attack unit, calculated as follows:

[0108]

[0109] Where, d iLet d be the Euclidean distance between the air defense unit and the air attack unit, and r be the effective combat range of the air defense unit; when d i When ′ < 1, it means that the enemy is already within the attack range of the air defense unit's intelligent agent;

[0110] Observe the second row of parameters Δθ in the spatial matrix i Let represent the change in the entry angle of the i-th air attack unit relative to the air defense unit agent; assuming the current time is t and the timing interval is Δt, the method for calculating the change in the entry angle is as follows:

[0111]

[0112] In the formula, θ is the entry angle, which is the angle between the line connecting the air attack unit and the air defense unit intelligent agent and the x-axis direction; when Δθ i When =0, it means that the target of the i-th air attack unit is the current air defense unit agent;

[0113] Determine the parameter m in the third row of the space matrix i This represents the number of air defense agents currently targeting and locking onto the i-th air attack unit; this parameter is generated through communication between air defense agents. If the current target of the air defense agent is j, then m... j A negative number indicates that the target has been targeted by this agent.

[0114] Determine the parameter s in the fourth row of the space matrix i This represents the survival status of the i-th air strike unit; if the air strike unit is still alive, its value is set to 1, and if the air strike unit has been shot down, its value is set to -1.

[0115] Determine the parameter Δv in the fifth row of the space matrix. i Represents the value of firepower transfer; let the target of the agent at time (t-Δt) be the j-th air attack unit, and its value be v. j At time t, the agent shifts its target to the i-th air attack unit, with a value of v. i Therefore, the method for calculating the value of firepower transfer in this situation is as follows:

[0116]

[0117] In other words, the transfer value is equal to the ratio of the planned transfer target value to the current relative distance, minus the ratio of the target value at the previous moment to the relative distance at the previous moment. It is used to evaluate the rationality of the air defense unit's intelligent agent transferring the target. Specific Implementation Example 4:

[0119] The only difference between Embodiment 4 and Embodiment 3 of this application is that:

[0120] Step 3 specifically involves:

[0121] Step 3.1: Determine the hit reward. When the agent successfully shoots down the i-th air attack unit, it is given a reward:

[0122] r = 30v i (6)

[0123] That is, 30 times the value of the i-th air strike unit;

[0124] Step 3.2: Determine the reward for firepower transfer. When an agent attempts to transfer its firepower from targeting the i-th air attack unit to targeting the j-th air attack unit, it is given a reward or penalty:

[0125] r = 1.5Δv j -0.6 (7)

[0126] Step 3.3: Determine the penalty for invalid aiming. When an agent attempts to aim at an air attack unit that has been shot down, it is given a penalty of r = -1.

[0127] Step 3.4: Target retention reward: When the agent's target is the same in the current and previous time steps, i.e., when actions in adjacent time steps are identical:

[0128] a t =a t+Δt (8)

[0129] The agent is given a reward of r=0.3 to encourage it to focus on targeting high-threat air units for extended periods.

[0130] Step 3.5: Process End Reward. When the process ends with the air defense unit agents shooting down all air attack units, all air defense unit agents are given a reward of r=10. If the process ends with one or more protected targets being destroyed by air attack units, a penalty is imposed.

[0131] r = -50n (9)

[0132] Where n is the number of protected targets destroyed at the end of the process, which is 1. Specific Implementation Example 5:

[0134] The difference between Embodiment 5 and Embodiment 4 of the present invention lies only in:

[0135] Step 4 specifically involves:

[0136] Six air defense units were added; while eight enemy air attack units were added, locking onto three protected targets and three long-range air defense units respectively, and the rest locking onto two important protected targets A and B among the protected targets;

[0137] Before each round of the process begins, the attributes of the agents and air attack units in the environment need to be initialized. For the air defense unit agents, the initialization operations mainly include: clearing radar data, initializing the fire interval state, and initializing the fire transfer state. For the enemy air attack units, the initialization operations mainly include: the air attack units selecting targets, randomly generating the starting position, and initializing the survival state. In order to ensure that the strategy for generating agents has a certain degree of universality and to prevent overfitting, the x-axis coordinate in the initial position initialization of the air attack units is randomly generated. Specific Implementation Example Six:

[0139] The difference between Embodiment Six and Embodiment Five of the present invention lies only in:

[0140] Step 5 specifically involves:

[0141] After completing the design of the state space, reward and punishment function, and environmental process, the MADDPG reinforcement learning algorithm can be used to train the air defense unit agent.

[0142] For the construction of the Actor and Critic neural networks, the AC network adopts a multi-input neural network design: the observation space matrix is ​​divided into 5 vectors of size 1×8, which are respectively input into 5 fully connected layers. After passing through one or more hidden layers, the output neurons are spliced ​​together and then input into the next hidden layer.

[0143] By sampling from a discrete distribution using the Gumbel-Softmax method to make it differentiable, Gumbel-Softmax sampling can be written as:

[0144]

[0145] Wherein, the parameter τ is called the temperature parameter of the distribution, and this parameter is always greater than 0, g i The noise sampled from Gumbel(0,1) is called the reparameter factor:

[0146] g i =-log(-logu),u~Uniform(0,1) (11). Specific Implementation Example 7:

[0148] The difference between Embodiment Seven and Embodiment Six of the present invention lies only in:

[0149] Step 6 specifically involves:

[0150] The deployment coordinates of long-range and medium-range air defense units in each deployment scheme generated by the multi-objective optimization algorithm are input into the air defense battlefield environment. The air defense battlefield is simulated using the intelligent agent strategy obtained through training. The kill rate under different deployment schemes is obtained, thereby evaluating the overall performance of potential deployment schemes and selecting the best individual among them. Specific Implementation Example 8:

[0152] The difference between Embodiment 8 and Embodiment 7 of the present invention lies only in:

[0153] This invention provides a multi-agent, multi-objective, multi-factor allocation decision-making system, the system comprising:

[0154] The simulation module performs air defense scenario modeling and simulates air attack scenarios.

[0155] A construction module is used to construct the action and observation space of the intelligent agent and to observe the spatial matrix.

[0156] The reward and punishment module establishes a reward and punishment function so that after the agent performs each action, a reward value is fed back to the agent based on the action.

[0157] The design module designs the air defense environment process and locks onto the protected targets.

[0158] The optimization module uses MADDPG reinforcement learning to generate strategies, optimizes air defense strategies, and performs multi-objective optimization to solve for and deploy solution sets. Specific Implementation Example Nine:

[0160] The difference between Embodiment Nine and Embodiment Eight of the present invention lies only in:

[0161] The present invention provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement a multi-agent, multi-objective, multi-factor allocation decision-making method. Specific Implementation Example 10:

[0163] The only difference between Embodiment 10 and Embodiment 9 of the present invention is that:

[0164] The present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement a multi-agent, multi-objective, multi-factor allocation decision-making method. Specific Implementation Example Eleven:

[0166] The only difference between Embodiment Eleven and Embodiment Ten of this invention is that:

[0167] Taking an air defense deployment scenario as an example, this invention will specifically demonstrate the implementation process of a multi-agent, multi-target, multi-factor allocation decision-making method. In this scenario, multiple agents represent different air defense units and, under limited resource conditions, need to optimize the deployment of air defense forces and resource allocation based on real-time intelligence and enemy threats to maximize the system's defensive effectiveness. Through this example, we can concretely illustrate the application of the above method and demonstrate how to leverage agent collaboration and optimization to improve system efficiency in complex decision-making processes.

[0168] Step 1: Air Defense Scenario Modeling

[0169] Suppose there are three protected targets, A, B, and C, where A and B are important targets with a higher protection weight than C. In this scenario, the three targets are located in an area densely populated with high-risk protected objects such as densely populated areas and factories. This area can be considered as a circular protected object with a radius of 125km. If we take the center of the circular area as the origin, with due north as the positive y-axis and due east as the positive x-axis, and assume the incoming attack direction is due north, i.e., from the positive y-axis, and the attack method is assumed to be level flight at a certain altitude, with each unit randomly locking onto a target and flying towards it at a constant speed.

[0170] Step 2: Building the Action and Observation Space for the Intelligent Agent

[0171] Assuming there are N enemy targets above the current air defense, the action space that an air defense unit can only perform can be defined as follows:

[0172] action = discrete(a1, a2, ..., a N ) T (1)

[0173] Here, "discrete" indicates that the action space is a discrete space, that is, action a is performed in this space. i This is a binary function, taking values ​​of 0 or 1. When its value is 1, it means that the air defense unit agent will turn its aim at the i-th enemy air attack unit at this moment; conversely, 0 means it will not turn its aim at that unit. At any given moment, an air defense unit can only target one enemy air attack unit.

[0174] In a real air defense environment, air defense systems cannot possess global information about the entire environment; they can only obtain partial observation information based on the operation of radar and other observation systems. Therefore, for the information input to the agent at each moment, an observation space matrix is ​​constructed, which is a 5×N matrix:

[0175]

[0176] Observe the first row of parameters d in the spatial matrixi ′ represents the relative distance between the air defense unit agent and the i-th air attack unit, calculated as follows:

[0177]

[0178] In the formula, d i Let d be the Euclidean distance between the air defense unit and the air attack unit, and r be the effective combat range of the air defense unit. i When ′ < 1, it means that the enemy is already within the attack range of the air defense unit's intelligent agent.

[0179] Observe the second row of parameters Δθ in the spatial matrix i This represents the change in the entry angle of the i-th air attack unit relative to the air defense unit agent. Let the current time be t, and the timing interval be Δt. The method for calculating the change in the entry angle is as follows:

[0180]

[0181] In the formula, θ is the entry angle, which is the angle between the line connecting the air attack unit and the air defense unit intelligent agent and the x-axis direction. When Δθ i When the value is 0, it means that the target of the i-th air attack unit is the current air defense unit agent. The longer this value remains zero, the higher the danger level of the current air defense unit agent.

[0182] Observe the parameter m in the third row of the spatial matrix i This represents the number of air defense agents currently targeting and locking onto the i-th air attack unit. This parameter is generated through communication between air defense agents; if the current target of the air defense agent is j, then m... j A negative number indicates that the target has been targeted by this agent.

[0183] Observe the parameter s in the fourth row of the spatial matrix i This represents the survival status of the i-th air strike unit. If the air strike unit is still alive, its value is set to 1; if the air strike unit has been shot down, its value is set to -1.

[0184] Observe the fifth row parameter Δv in the spatial matrix i This represents the value of firepower transfer. Let the target of the agent at time (t-Δt) be the j-th air attack unit, and its value be v. j At time t, the agent shifts its target to the i-th air attack unit, with a value of v. i Therefore, the method for calculating the value of firepower transfer in this situation is as follows:

[0185]

[0186] The transfer value is equal to the ratio of the planned transfer target value to the current relative distance, minus the ratio of the target value at the previous moment to the relative distance at the previous moment. This value is used to evaluate the rationality of the air defense unit's intelligent agent transferring the target.

[0187] Step 3: Design of reward and punishment functions

[0188] Given the characteristics of dense intelligence and rapidly changing scenarios in air defense battlefields, the reward and punishment function needs to be designed as a dense reward and punishment function, meaning that after each action is performed by the agent, a reward value is given to the agent based on the action. The agent's reward and punishment function mainly includes the following:

[0189] (1) Hit reward

[0190] When the agent successfully shoots down the i-th air attack unit, it is rewarded:

[0191] r = 30v i (6)

[0192] That is, 30 times the value of the i-th air strike unit.

[0193] (2) Firepower Transfer Reward

[0194] When an agent attempts to transition from targeting the i-th air unit to targeting the j-th air unit, it is given a reward or penalty:

[0195] r = 1.5Δv j -0.6 (7)

[0196] This reward and penalty function consists of two parts: the first part is the reward, which is 1.5Δv. j The reward for a firepower transfer is higher when the value of the transfer is greater. The penalty is -0.6, representing the basic penalty for wasted time due to the firepower transfer. In summary, the system only provides a positive reward to the agent when a firepower transfer action of sufficient value is performed, thus preventing the agent from engaging in aimless and random targeting.

[0197] (3) Penalty for Invalid Aiming

[0198] When an agent attempts to target an air strike unit that has already been shot down, it is penalized with a penalty of r = -1.

[0199] (4) Aim to maintain the reward

[0200] When the agent's target is the same in the current moment and the previous moment, that is, when the actions in adjacent moments are the same:

[0201] a t =a t+Δt(8)

[0202] The agent is given a reward of r=0.3 to encourage it to focus on targeting high-threat air units for extended periods.

[0203] (5) Process completion reward

[0204] When the process ends with the air defense unit agents shooting down all air attack units, all air defense unit agents are awarded a reward of r=10. When the process ends with one or more protected targets being destroyed by air attack units, a penalty is imposed:

[0205] r = -50n (9)

[0206] In the formula, n is the number of protected targets destroyed at the end of the process, which is usually 1.

[0207] Step 4: Air Defense Environment Process Design

[0208] Due to their relatively short effective combat radius and short intervals, short- and medium-range air defense units are extremely vulnerable to air attack units, but highly effective against air attack units near the drone's bombing line. Therefore, after the air attack units actually enter the formation, for simplification, a total of 6 air defense units are included in this scenario; while 8 enemy air attack units are included, locking onto three protected targets and three long-range air defense units, with the remaining units locking onto two important protected targets, A and B.

[0209] Before each round of the process begins, the attributes of the agents and air attack units in the environment need to be initialized. For air defense unit agents, the initialization operations mainly include: clearing radar data, initializing fire interval status, and initializing fire transfer status; for enemy air attack units, the initialization operations mainly include: air attack units selecting targets, randomly generating starting positions, and initializing survival status. To ensure the agent generation strategy has a certain degree of universality and to prevent overfitting, the x-axis coordinate in the initial position initialization of air attack units is randomly generated.

[0210] In each round of the process, the step size is set to 1 second, and three actions are performed in each step:

[0211] (1) Enemy air strike unit actions

[0212] Each air strike unit first determines whether it has been shot down. If not, it updates the distance between itself and the target according to its speed. After updating, it checks whether the distance is less than or equal to the bombing line. If it is less than or equal to the bombing line, the process ends.

[0213] (2) Air defense unit operations

[0214] Each air defense unit agent generates actions based on the observation space matrix provided by the radar. First, a fire transfer determination is performed. If the selected enemy target is the same as the one from the previous moment, the action returns to target holding; otherwise, it returns to fire transfer. Next, a firing determination is performed. If there is no fire interval or fire transfer state at this time, and the enemy target is alive and within firing range, a hit determination is performed; if the enemy is not alive, it returns to invalid acquisition.

[0215] (3) Step-size calculation

[0216] After all enemy air attack units and air defense unit agents complete their actions within a single step, the environment will perform a settlement process. First, based on the information returned by the air defense unit agents during their actions, the reward value fed back to each agent is calculated using the reward and penalty function formula given in step three. Then, the process termination is determined based on the bombing situation and the number of units shot down by the air attack units.

[0217] Step 5: MADDPG Reinforcement Learning Algorithm Generation Strategy

[0218] After completing the design of the state space, reward and punishment functions, and environmental processes, the MADDPG reinforcement learning algorithm can be used to train the air defense unit agent.

[0219] In the construction of the Actor and Critic neural networks, the observation space matrix of the air defense environment in this training is a 5×8 matrix as shown in formula (2). While a fully connected network can only input one vector, and a convolutional neural network can input a matrix, this environment does not require the ability to recognize the entire matrix. Therefore, the AC network in this environment adopts a multi-input neural network design. The method is as follows: the observation space matrix is ​​divided into five 1×8 vectors, which are then input into five fully connected layers. After passing through one or more hidden layers, the output neurons are concatenated together and then input into the next hidden layer.

[0220] Furthermore, to address the issue of discretized action spaces in air defense environments leading to non-differentiable agent actions with respect to policy parameters, the Gumbel-Softmax method can be used to sample from a discrete distribution, making it differentiable. Gumbel-Softmax sampling can be written as:

[0221]

[0222] In the formula, the parameter τ is called the temperature parameter of the distribution, and this parameter is always greater than 0. Adjusting this parameter controls the approximation of the generated sampling distribution to the discrete distribution; the larger τ is, the closer the sampling distribution is to a uniform distribution; the smaller τ is, the closer the sampling is to the isolated heat result; g iThe noise sampled from Gumbel(0,1) is called the reparameter factor:

[0223] g i =-log(-logu),u~Uniform(0,1) (11)

[0224] Step Six: Solve and deploy the solution set using a multi-objective optimization algorithm.

[0225] Based on the air defense environment and actual needs, an optimization problem is designed with corresponding objective functions and constraints. Various multi-objective optimization algorithms are then used to solve this problem to obtain a solution set for deployment schemes. Typical multi-objective optimization algorithms include NSGA-III and AGE-MOEA2 algorithms.

[0226] The NSGA-III algorithm, building upon the NSGA-II algorithm, introduces the concept of reference points. These uniformly distributed reference points guide the distribution of solutions. Each objective has a corresponding reference point in a standardized space, enabling the algorithm to achieve better uniformity when dealing with multiple objectives. This allows the algorithm to solve optimization problems with a larger number of objective functions.

[0227] The AGE-MOEA2 algorithm uses a simple heuristic to simulate the shape of the Pareto front, and calculates the evaluation score by introducing Newton's iteration method for root search and geodesics to measure the distance of non-dominated solutions on the surface manifold, thereby better ensuring the diversity and authenticity of the solution set.

[0228] Step 7: Deployment Plan Evaluation and Screening

[0229] The deployment coordinates of long-range and medium-range air defense units in each deployment scheme generated by the multi-objective optimization algorithm are input into the air defense battlefield environment. The air defense battlefield is simulated using the intelligent agent strategy obtained through training. The kill rate under different deployment schemes is obtained, thereby evaluating the overall performance of potential deployment schemes and selecting the best individual among them.

[0230] With the hyperparameters set as follows, target allocation strategy table 1 is obtained after training:

[0231] Table 1. Allocation Strategy Table

[0232]

[0233]

[0234] The deployment locations of the air defense unit agents are generated through a rule-based approach based on common practical intuition, as shown in Table 2 below:

[0235] Table 2. Deployment Location Table

[0236]

[0237] Let the total number of training episodes be Episodes = 7000. Let the reward value in each episode be the sum of the rewards obtained by all agents in each step. Then, a scatter plot of the reward value as a function of the training episodes can be obtained as follows: Figure 6 As shown, after about 700 episodes, the reward value exceeded zero for the first time, indicating that the air defense unit agent had achieved victory; after about 5000 episodes, the reward value curve with a value greater than 0 became increasingly dense, indicating that the agent's efficiency in obtaining rewards had improved.

[0238] Figure 7 This is a scatter plot showing the change in the number of enemy air attack units shot down by the air defense agents in each episode. It can be seen that as the number of episodes increases, the density of scatter points with lower kill counts gradually decreases, while the density of scatter points with higher kill counts gradually increases. This indicates that the agents' ability to cooperate in resisting enemy air attack units is continuously improving, and after 1000 episodes, there are no longer any instances of zero kills.

[0239] Figure 8 The graph shows a scatter plot of the number of shots fired. It's evident that the number of shots fired is highly correlated with the number of kills, with the final convergence range being slightly higher than the kill count. This is because the hit rate of air defense units is not 100%. Observations reveal occasional instances where the number of shots fired is excessively high, indicating that situations where multiple volleys of fire from air defense units fail to hit the target cannot be ruled out. For such occasional events, deploying some anti-aircraft guns as supplementary firepower and a last resort defense measure could be considered.

[0240] Figure 9 The plot shows the scatter plot of the number of firepower transfers. It can be seen that, guided by the reward / penalty function, the total number of firepower transfers in each episode decreases rapidly and eventually converges to around 75.

[0241] During training, an evaluation is performed after each training session of 100 episodes. This evaluation is repeated 100 times, with the air defense environment remaining constant, and the agent decentralizedly selecting actions to execute. After 100 rounds, the average enemy air unit kill rate per round is calculated as the evaluation value. Figure 10 The figure shows the changes in the evaluation value over 7000 episodes. It can be seen that it has been on an upward trend since the start of training, and converges to about 0.95 after about 5000 episodes.

[0242] The AGE-MOEA2 multi-objective optimization algorithm is used to solve the optimal deployment problem. A population size of 100 is used to obtain a set of deployment schemes, which are then placed into an air defense environment. The trained strategy is then applied to each deployment scheme 100 times in a decentralized manner, and the average kill rate is obtained. The kill rate histogram of the deployment scheme set is then generated as follows: Figure 11 As shown, the average kill rate is 0.2524. It can be seen that in the deployment scheme set, apart from a small number of individuals distributed at lower kill rates, most individuals are concentrated in the range of 0.25-0.35.

[0243] The optimal deployment plan was for individual #22, with a kill rate of 0.3325. The deployment diagram is shown below. Figure 12 As shown, the expected goals have been achieved.

[0244] Simulation analysis shows that the reinforcement learning-based target allocation strategy for rapid selection of Pareto solutions for air defense deployment can achieve rapid selection of Pareto solutions for air defense deployment.

[0245] The above description is merely a preferred embodiment of a multi-agent, multi-objective, multi-factor allocation decision-making method. The scope of protection for such a method is not limited to the above embodiments; all technical solutions falling within this framework are within the scope of protection of this invention. It should be noted that for those skilled in the art, any improvements and variations made without departing from the principles of this invention should also be considered within the scope of protection of this invention.

Claims

1. A multi-agent, multi-objective, multi-factor allocation decision-making method, characterized by: The method includes the following steps: Step 1: Conduct air defense scenario modeling to simulate aerial attack scenarios; Step 2: Construct the action and observation space of the intelligent agent, and observe the spatial matrix; Step 2 specifically involves: Currently, there are above air defenses For an enemy, the action space that an air defense unit can only initiate is defined as follows: (1) Here, "discrete" indicates that the action space is a discrete space, meaning that actions occur within this space. It is a binary function that takes the value 0 or 1; For the information input to the agent at each time step, an observation space matrix is ​​constructed, which is a matrix of size ... The matrix: (2) Observe the parameters in the first row of the space matrix Representing the air defense unit intelligent agent and the first The relative distance between air attack units is calculated as follows: (3) Observe the second row of parameters in the space matrix Representing the The change in the angle of entry of each air attack unit relative to the air defense unit agent; let the current time be... The timing interval is The method for calculating the change value of the entry angle is as follows: (4) Determine the parameters of the third row in the space matrix Representing the The number of air defense agents currently targeting and locking onto a given air attack unit; this parameter is generated through communication between air defense agents, and the current target of the air defense agents is... ,but A negative number indicates that the target has been targeted by this agent. Determine the parameters of the fourth row in the space matrix Representing the The survival status of each air strike unit; if the air strike unit is still alive, its value is set to 1; if the air strike unit has been shot down, its value is set to -1. Determine the parameters of the fifth row in the space matrix. Represents the value of firepower transfer; set at The target of the agent at any given moment is the first... One air strike unit, valued at ;exist The agent will shift its target to the next moment. One air strike unit, valued at Therefore, the method for calculating the value of firepower transfer in this situation is as follows: (5) That is, the transfer value is equal to the ratio of the planned transfer target value to the current relative distance, minus the ratio of the target value at the previous moment to the relative distance at the previous moment, and is used to evaluate the rationality of the air defense unit's intelligent agent transferring the target. Step 3: Establish a reward and punishment function so that the agent receives a reward value after each action. Step 4: Design the air defense environment process and identify the targets to be protected; Step 5: Perform MADDPG reinforcement learning to generate strategies, optimize air defense strategies, and perform multi-objective optimization to solve for the deployment solution set; Step 5 specifically involves: For the construction of the Actor and Critic neural networks, the AC network adopts a multi-input neural network design: the observation space matrix is ​​divided into 5... The vectors of size are input into the five fully connected layers respectively. After passing through one or more hidden layers, the output neurons are concatenated together and then input into the next hidden layer. By sampling from a discrete distribution using the Gumbel-Softmax method to make it differentiable, Gumbel-Softmax sampling can be written as: (10) Among them, parameters This is called the temperature parameter of the distribution, and this parameter is always greater than 0. For a sample The noise is called the reparameter factor: (11)。 2. The method according to claim 1, characterized in that: Step 1 specifically involves: Three defense targets were set. A , B , C ,in A , B As an important target to be protected, its protection weight is relatively high. C Higher; the three protected targets are located in an area densely populated with attacked factories and other protected objects. This area is considered a circular protected object with a radius of 125km. The center of the circular area is taken as the origin of the coordinate system, with due north as the coordinate axis. The positive direction of the axis is east. Positive direction of the axis.

3. The method according to claim 2, characterized in that: Step 3 specifically involves: Step 3.1: Determine the hit reward. When the agent successfully shoots down the first... When an air strike unit is attacked, a reward will be given: (6) That is, 30 times the first The value of an air strike unit; Step 3.2: Determine the reward for transferring firepower when the agent attempts to shift firepower from the first target. The status of the air strike unit has been transferred to target the first air strike unit. When an air strike unit is attacked, it should be rewarded or punished accordingly. (7) Step 3.3: Determine the penalty for invalid targeting. When an agent attempts to target an air strike unit that has already been shot down, issue it a penalty. Punishment; Step 3.4: Target retention reward: When the agent's target is the same in the current and previous time steps, i.e., when actions in adjacent time steps are identical: (8) Give the agent a The rewards encourage agents to focus on targeting high-threat air units for extended periods. Step 3.5: Process Completion Reward. When the process ends with the air defense unit agents shooting down all air attack units, all air defense unit agents will be awarded a reward. The reward is given when the process ends with one or more protected targets being destroyed by air attack units, and the penalty is given when the process ends with the target being destroyed by air attack units. (9) in, The number of protected targets destroyed at the end of the process, which is 1.

4. The method according to claim 3, characterized in that: Step 4 specifically involves: Six air defense units were added; while eight enemy air attack units were added, one locking onto three defended targets and three long-range air defense units, and the remaining units locking onto two key defended targets. A , B ; Before each round of the process begins, the attributes of the agents and air attack units in the environment need to be initialized.

5. The method according to claim 4, characterized in that: The method further includes step 6, which specifically involves: The deployment coordinates of long-range and medium-range air defense units in each deployment scheme generated by the multi-objective optimization algorithm are input into the air defense battlefield environment. The air defense battlefield is simulated using the intelligent agent strategy obtained through training. The kill rate under different deployment schemes is obtained, thereby evaluating the overall performance of potential deployment schemes and selecting the best individual among them.

6. A multi-agent, multi-objective, multi-factor allocation decision-making system, wherein the system operates based on the multi-agent, multi-objective, multi-factor allocation decision-making method of claim 1, characterized in that: The system includes: The simulation module performs air defense scenario modeling and simulates air attack scenarios. A construction module is used to construct the action and observation space of the intelligent agent and to observe the spatial matrix. The reward and punishment module establishes a reward and punishment function so that after the agent performs each action, a reward value is fed back to the agent based on the action. The design module designs the air defense environment process and locks onto the protected targets. The optimization module uses MADDPG reinforcement learning to generate strategies, optimizes air defense strategies, and performs multi-objective optimization to solve for and deploy solution sets.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method as claimed in any one of claims 1-5.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Dynamic air multi-target distribution and strike method based on multi-agent reinforcement learning

    CN116956705A

  • Unmanned aerial vehicle cooperative air combat decision-making method based on GRU-MAPPO deep reinforcement learning

    CN119129413A