An aircraft intelligent many-to-many game target assignment method, electronic equipment and storage medium

By training an intelligent multi-to-multi game target allocation method for aircraft using the OW-QMIX algorithm, the problem of poor target allocation strategies in multi-aircraft combat is solved, achieving more efficient and robust target allocation and improving the performance of the algorithm in rapidly changing environments.

CN120805633BActive Publication Date: 2026-05-01HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2024-12-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

In multi-aircraft versus multi-player game adversarial processes, existing algorithms struggle to adapt to the rapid changes in the adversarial format, resulting in poor target allocation strategies.

Method used

The OW-QMIX algorithm is used to train distributed agents. Combined with the OW-QMIX agent structure, a target allocation method for intelligent multi-to-multi game is designed for aircraft. By constructing a relative motion model of aircraft game adversarial and the state space of multiple agents, fuel consumption and reward values ​​are calculated to optimize the target allocation strategy.

Benefits of technology

It improves the target allocation strategy's performance metrics and training speed, enhances the algorithm's robustness, and outperforms the QMIX algorithm, especially in cases where some aircraft fail.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805633B_ABST
    Figure CN120805633B_ABST
Patent Text Reader

Abstract

The application discloses an aircraft intelligent many-to-many game target allocation method, an electronic device and a storage medium, and belongs to the technical field of aircraft control. In order to realize the many-to-many fast game target allocation of an aircraft, an aircraft game confrontation relative motion model is established; a training environment of the intelligent many-to-many game confrontation of the aircraft, a state space of a plurality of intelligent agents and an action space of the intelligent agents are designed; observation values of the aircraft in the training environment are collected as the state space of the plurality of intelligent agents, input numbers of the aircraft are used as the action space of the intelligent agents, fuel consumption of the aircraft and reward values obtained by the aircraft are calculated; an OW-QMIX intelligent agent structure is constructed, then the state space of the plurality of intelligent agents, the action space of the intelligent agents and the reward values obtained are input into the OW-QMIX intelligent agent structure, and an aircraft intelligent many-to-many game target allocation strategy training result is output, and then simulation verification is carried out. The application realizes the many-to-many fast game target allocation of the aircraft.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of aircraft control technology, specifically relating to an intelligent many-to-many game target allocation method, electronic equipment, and storage medium for aircraft. Background Technology

[0002] In multi-aircraft versus multi-target game adversarial processes, the spatial distribution of adversarial targets typically covers a large spatial area, thus requiring aircraft to assign tasks to these targets. Optimizing the target assignment process yields better game strategies. However, due to the rapidly changing nature of multi-aircraft versus multi-target adversarial processes, conventional algorithms struggle to adapt to the rapid shifts in the game's adversarial dynamics. Therefore, research is needed on the problem of rapid target assignment in multi-aircraft versus multi-target game adversarial processes. Summary of the Invention

[0003] The problem this invention aims to solve is to achieve rapid target allocation in many-to-many games for aircraft, and proposes an intelligent many-to-many game target allocation method, electronic device, and storage medium for aircraft.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A method for target allocation in intelligent many-to-many game theory involving aircraft includes the following steps:

[0006] S1. Establish a game-theoretic relative motion model for aircraft;

[0007] S2. Based on the relative motion model of the aircraft game adversarial competition obtained in step S1, design the training environment, state space of multiple agents, and action space of agents for intelligent multi-to-multi game adversarial competition of aircraft. Collect the observation values ​​of aircraft in the training environment as the state space of multiple agents, and the input number of aircraft as the action space of agents. Calculate the fuel consumption and reward value of aircraft.

[0008] S3. Construct the OW-QMIX agent structure, and then input the state space, action space and reward value of the multi-agent obtained in step S2 into the OW-QMIX agent structure, and output the training results of the aircraft intelligent multi-to-multi game objective allocation strategy.

[0009] S4. Simulate and verify the training results of the target allocation strategy for intelligent multi-to-multi game of aircraft obtained in step S3.

[0010] Furthermore, the specific implementation method of step S1 includes the following steps:

[0011] S1.1. Construct the three-degree-of-freedom kinematic equations for the aircraft's flight outside the atmosphere, expressed as follows:

[0012]

[0013] in, Let V be the rate of change of the aircraft's velocity relative to the x-axis, and let V be the displacement vector and velocity vector of the aircraft in the ballistic frame relative to the inertial frame. The velocity change rate of the aircraft relative to the y-axis. Let θ be the rate of change of the aircraft's velocity relative to the z-axis, θ be the trajectory angle, and ψ be the velocity of the aircraft relative to the z-axis. P For heading angle;

[0014]

[0015] in, pitch angle The rate of change of velocity, Yaw angle ψ m The rate of change of velocity, For roll angle γ m rate of change of velocity, ω xB Let ω be the angular velocity of the spacecraft relative to the x-axis. yB Let ω be the angular velocity of the aircraft relative to the y-axis. zB The angular velocity of the aircraft relative to the z-axis;

[0016] S1.2. Construct the dynamic equations for the aircraft's flight outside the atmosphere, the expression of which is:

[0017]

[0018] Where m is the mass of the aircraft, r is the geocentric radius vector of the aircraft, V is the displacement and velocity vector of the aircraft in the ballistic frame relative to the inertial frame, Ω is the angular velocity of the ballistic frame relative to the inertial frame, P is the thrust of the game aircraft's engine, F is the aerodynamic force and disturbance force other than the thrust, H is the angular momentum of the game aircraft's center of mass, and M P M is the torque generated by the thrust eccentricity of the aircraft engine, M is the aerodynamic torque and disturbance torque, and ω is the angular velocity of the aircraft system relative to the inertial frame.

[0019] S1.3. Set the position of the spacecraft at time T to r. M =[x M ,y M ,z M ] T The position of the target aircraft at time T is r. T =[x T ,y T ,z T ] T Construct a motion position parameter model for the aircraft and the target aircraft, with the following expression:

[0020] The expression for the relative position Δr is:

[0021]

[0022] Based on the relative distance, the expression for the relative velocity v is obtained by differentiation:

[0023]

[0024] After obtaining the relative speed, the corresponding distance proximity rate is obtained. The expression is:

[0025]

[0026] Angle of view q g and line of sight deflection q f The expression is:

[0027]

[0028] in, q f = (-π, π];

[0029] Since the target uses proportional guidance, the expression for the aircraft-target line-of-sight angle rotation rate is:

[0030]

[0031] in, For the aircraft-target line-of-sight tilt angle rotation, The deflection rate of the aircraft-target line of sight.

[0032] Furthermore, the specific implementation method of step S2 includes the following steps:

[0033] S2.1. Based on the relative motion model of aircraft game adversarial competition obtained in step S1, design a training environment for intelligent multi-to-multi game adversarial competition of aircraft.

[0034] Initialization settings are performed, including randomly generating multiple aircraft and game targets, and randomly determining the speed and direction of the aircraft. At the start of the simulation, each aircraft conducts a detection and acquires the relative position and speed information of the game targets within its field of view. The initial position range of the aircraft is set to 0 to 1 km in the x-axis direction, while the target's x-axis position is 10 to 11 km, the y-axis position is -1 to 12 km, and the z-axis position is -10 to 10 km. The initial speed value range is 5000 to 5500 m / s, and the speed tilt and deflection angles range from -10° to 10°.

[0035] S2.2. The expression for the state space S of a multi-agent system is defined as follows:

[0036]

[0037] In the formula, Δx1, Δy1, and Δz1 represent the relative positions of the three axes of the first game's aircraft. Let be the rate of change of the relative distance between the aircraft in the first game;

[0038] The expression for the action space A of the intelligent agent is defined as follows:

[0039] A = [num] (10)

[0040] Wherein, the action space of the agent is the target allocation result of the corresponding game aircraft, and num is the corresponding number of the game aircraft;

[0041] S2.3. Input the observation values ​​of each aircraft into the agent network. The agent network outputs the corresponding number of the aircraft in the game. The aircraft group simulates the relative motion model of the aircraft game confrontation. During the aircraft game confrontation, the fuel consumption of the aircraft is calculated, and the reward obtained by the allocation algorithm is calculated in combination with the result of the game confrontation.

[0042] The expression for the aircraft's fuel consumption is:

[0043]

[0044] Where Fuel is the fuel consumption index, n is the total number of aircraft, i is any one of n, N represents the total overload of aircraft maneuvers, and t represents the total duration of aircraft maneuvers;

[0045] The expression for the reward obtained by the allocation algorithm is:

[0046]

[0047] Among them, Fuel i Let be the fuel consumption index of the i-th aircraft, and let be the sum of the game competition reward value and the fuel consumption reward value.

[0048] Furthermore, the specific implementation method of step S3 includes the following steps:

[0049] S3.1. Input the state space, action space and reward value of the multi-agent obtained in step S2 into the OW-QMIX agent structure. Randomly select the first step of input as S0. Use the dynamic equation in step S1 to simulate the flight trajectory. After completing a time step of 0.1s, randomly obtain the state space S′ of the next moment. Build the data into a quadruple (S,A,S′,r) and store it in the experience pool D. Set the experience pool loop stopping condition to a maximum number of rounds of 100,000.

[0050] Then set the current network parameter to θ. now The updated network parameters are θ. new The value function Q of the i-th single agent is obtained by forward propagation of the single agent network. i The expression for ′ is:

[0051] Q i ′=q i (S i ′,A i ;θ now (13)

[0052] S3.2. To address the value allocation problem in multi-agent learning, a Value Decomposition Network (VDN) is used for multi-agent training. The network weights of the OW-QMIX agent structure are set to Q(S,A;θ), the experience pool to D, and the learning rate to α. The joint action value function is set to Q. tot (S,A), the expression is:

[0053]

[0054] Among them, Q i Let i be the value function of the i-th single agent;

[0055] S3.3. When multiple agents need to extract distributed policies, VDN guarantees that operating on the globally maximum independent variable and operating on the individually maximum independent variable on the agent network yields the same result, resulting in the expression:

[0056]

[0057] Each individual agent collects information through its agent network, aggregates it, and inputs it into a hybrid network for computation. The hybrid network consists of a supernetwork that uses the absolute value function as its heuristic function, ultimately calculating Q. tot (S,A);

[0058] S3.4. Set the TD target y TD,i With TD error δ TD,i The expression is:

[0059] yTD,i =reward + γQ i ′ (16)

[0060] δ TD,i =Q tot -y TD (17)

[0061] Set up backpropagation to obtain the gradient g. i The expression is:

[0062]

[0063] The expression for setting and updating network parameters is:

[0064] θ new ←θ now +α·δ TD,i ·g i (19)

[0065] S3.5. Repeat steps S3.1-S3.4 until the training results of the intelligent many-to-many game objective allocation strategy for the aircraft are output.

[0066] Furthermore, the specific implementation method of step S4 includes the following steps:

[0067] S4.1. Determine the number of multiple aircraft in the simulation to be 16 and the number of targets to be 12. During environment initialization, the parameters of the aircraft and targets will be randomly selected for initialization.

[0068] S4.2. Set the hyperparameters of the OW-QMIX network. The hardware environment used for simulation is: CPU i5-11400F, GPU GTX1060. Under normal working conditions, the aircraft does not have partial or complete failure issues. The simulation is then performed for verification.

[0069] An electronic device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the intelligent many-to-many game target allocation method for aircraft.

[0070] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the aforementioned intelligent many-to-many game target allocation method for aircraft.

[0071] The beneficial effects of this invention are:

[0072] The present invention discloses an intelligent multi-to-multi game target allocation method for aircraft. It designs a distributed aircraft target allocation method based on optimization weighted hybrid Q learning. Considering the limited observations between aircraft and different communication topologies, the OW-QMIX algorithm is used to train the distributed agents, and related multi-agent algorithm simulations are performed.

[0073] The target allocation method for intelligent many-to-many game of aircraft described in this invention, after testing, shows that the OW-QMIX algorithm improves the target completion rate by 5.12% compared to the QMIX algorithm. The training speed of the OW-QMIX algorithm is also faster than that of the QMIX network. Therefore, under the premise of using the same hyperparameters, the OW-QMIX algorithm has stronger computation speed and target completion rate than the QMIX algorithm. Furthermore, due to the characteristics of offline training and online computation, the OW-QMIX algorithm has stronger robustness compared to the traditional greedy strategy. Attached Figure Description

[0074] Figure 1 This is a flowchart of a target allocation method for intelligent many-to-many game in aircraft, as described in this invention.

[0075] Figure 2 This is a schematic diagram illustrating the overall application process of the OW-QMIX algorithm of this invention;

[0076] Figure 3 This is a diagram of the intelligent agent structure of the QMIX algorithm of this invention;

[0077] Figure 4 This is a design diagram of the training environment for this invention;

[0078] Figure 5 This is the training process curve of the algorithm in the normal working environment of this invention;

[0079] Figure 6 These are the algorithm test results for the normal working environment of this invention;

[0080] Figure 7 This is a training process curve for part of the failure environment algorithm of this invention;

[0081] Figure 8 These are the test results of the failure environment algorithm of this invention. Detailed Implementation

[0082] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention; that is, the described specific embodiments are merely a part of the embodiments of the invention, and not all of them. The components of the specific embodiments of the invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations, and the invention may also have other embodiments.

[0083] Therefore, the following detailed description of specific embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected specific embodiments of the invention. All other specific embodiments obtained by those skilled in the art based on these specific embodiments without inventive effort are within the scope of protection of this invention.

[0084] To further understand the invention's content, features, and effects, the following specific embodiments are provided, along with accompanying drawings. Figure 1 - Appendix Figure 8 Detailed explanation is as follows: Specific implementation method one:

[0086] A method for target allocation in intelligent many-to-many game theory involving aircraft includes the following steps:

[0087] S1. Establish a game-theoretic relative motion model for aircraft;

[0088] Furthermore, the specific implementation method of step S1 includes the following steps:

[0089] S1.1. Construct the three-degree-of-freedom kinematic equations for the aircraft's flight outside the atmosphere, expressed as follows:

[0090]

[0091] in, Let V be the rate of change of the aircraft's velocity relative to the x-axis, and let V be the displacement vector and velocity vector of the aircraft in the ballistic frame relative to the inertial frame. The velocity change rate of the aircraft relative to the y-axis. Let θ be the rate of change of the aircraft's velocity relative to the z-axis, θ be the trajectory angle, and ψ be the velocity of the aircraft relative to the z-axis. P For heading angle;

[0092]

[0093] in, pitch angle The rate of change of velocity, Yaw angle ψm The rate of change of velocity, For roll angle γ m rate of change of velocity, ω xB Let ω be the angular velocity of the spacecraft relative to the x-axis. yB Let ω be the angular velocity of the aircraft relative to the y-axis. zB The angular velocity of the aircraft relative to the z-axis;

[0094] S1.2. Construct the dynamic equations for the aircraft's flight outside the atmosphere, the expression of which is:

[0095]

[0096] Where m is the mass of the aircraft, r is the geocentric radius vector of the aircraft, V is the displacement and velocity vector of the aircraft in the ballistic frame relative to the inertial frame, Ω is the angular velocity of the ballistic frame relative to the inertial frame, P is the thrust of the game aircraft's engine, F is the aerodynamic force and disturbance force other than the thrust, H is the angular momentum of the game aircraft's center of mass, and M P M is the torque generated by the thrust eccentricity of the aircraft engine, M is the aerodynamic torque and disturbance torque, and ω is the angular velocity of the aircraft system relative to the inertial frame.

[0097] S1.3. Set the position of the spacecraft at time T to r. M =[x M ,y M ,z M ] T The position of the target aircraft at time T is r. T =[x T ,y T ,z T ] T Construct a motion position parameter model for the aircraft and the target aircraft, with the following expression:

[0098] The expression for the relative position Δr is:

[0099]

[0100] Based on the relative distance, the expression for the relative velocity v is obtained by differentiation:

[0101]

[0102] After obtaining the relative speed, the corresponding distance proximity rate is obtained. The expression is:

[0103]

[0104] Angle of view q g and line of sight deflection q f The expression is:

[0105]

[0106] in, q f = (-π, π];

[0107] Since the target uses proportional guidance, the expression for the aircraft-target line-of-sight angle rotation rate is:

[0108]

[0109] in, For the aircraft-target line-of-sight tilt angle rotation, For the aircraft-target line-of-sight deflection angle rotation rate;

[0110] S2. Based on the relative motion model of the aircraft game adversarial competition obtained in step S1, design the training environment, state space of multiple agents, and action space of agents for intelligent multi-to-multi game adversarial competition of aircraft. Collect the observation values ​​of aircraft in the training environment as the state space of multiple agents, and the input number of aircraft as the action space of agents. Calculate the fuel consumption and reward value of aircraft.

[0111] Furthermore, the specific implementation method of step S2 includes the following steps:

[0112] S2.1. Based on the relative motion model of aircraft game adversarial competition obtained in step S1, design a training environment for intelligent multi-to-multi game adversarial competition of aircraft.

[0113] Initialization settings are performed, including randomly generating multiple aircraft and game targets, and randomly determining the speed and direction of the aircraft. At the start of the simulation, each aircraft conducts a detection and acquires the relative position and speed information of the game targets within its field of view. The initial position range of the aircraft is set to 0 to 1 km in the x-axis direction, while the target's x-axis position is 10 to 11 km, the y-axis position is -1 to 12 km, and the z-axis position is -10 to 10 km. The initial speed value range is 5000 to 5500 m / s, and the speed tilt and deflection angles range from -10° to 10°.

[0114] S2.2. The expression for the state space S of a multi-agent system is defined as follows:

[0115]

[0116] In the formula, Δx1, Δy1, and Δz1 represent the relative positions of the three axes of the first game's aircraft. Let be the rate of change of the relative distance between the aircraft in the first game;

[0117] The expression for the action space A of the intelligent agent is defined as follows:

[0118] A = [num] (10)

[0119] Wherein, the action space of the agent is the target allocation result of the corresponding game aircraft, and num is the corresponding number of the game aircraft;

[0120] S2.3. Input the observation values ​​of each aircraft into the agent network. The agent network outputs the corresponding number of the aircraft in the game. The aircraft group simulates the relative motion model of the aircraft game confrontation. During the aircraft game confrontation, the fuel consumption of the aircraft is calculated, and the reward obtained by the allocation algorithm is calculated in combination with the result of the game confrontation.

[0121] The expression for the aircraft's fuel consumption is:

[0122]

[0123] Where Fuel is the fuel consumption index, n is the total number of aircraft, i is any one of n, N represents the total overload of aircraft maneuvers, and t represents the total duration of aircraft maneuvers;

[0124] The expression for the reward obtained by the allocation algorithm is:

[0125]

[0126] Among them, Fuel i Let be the fuel consumption index of the i-th aircraft, and let be the sum of the game competition reward value and the fuel consumption reward value.

[0127] Furthermore, the reward is a normalized reward, which includes game-playing rewards and fuel consumption rewards. The final reward is normalized according to the number of aircraft to obtain the final reward value.

[0128] S3. Construct the OW-QMIX agent structure, and then input the state space, action space and reward value of the multi-agent obtained in step S2 into the OW-QMIX agent structure, and output the training results of the aircraft intelligent multi-to-multi game objective allocation strategy.

[0129] Furthermore, the specific implementation method of step S3 includes the following steps:

[0130] S3.1. Input the state space, action space and reward value of the multi-agent obtained in step S2 into the OW-QMIX agent structure. Randomly select the first step of input as S0. Use the dynamic equation in step S1 to simulate the flight trajectory. After completing a time step of 0.1s, randomly obtain the state space S′ of the next moment. Build the data into a quadruple (S,A,S′,r) and store it in the experience pool D. Set the experience pool loop stopping condition to a maximum number of rounds of 100,000.

[0131] Then set the current network parameter to θ. now The updated network parameters are θ. new The value function Q of the i-th single agent is obtained by forward propagation of the single agent network. i The expression for ′ is:

[0132] Q i ′=q i (S i ′,A i ;θ now (13)

[0133] S3.2. To address the value allocation problem in multi-agent learning, a Value Decomposition Network (VDN) is used for multi-agent training. The network weights of the OW-QMIX agent structure are set to Q(S,A;θ), the experience pool to D, and the learning rate to α. The joint action value function is set to Q. tot (S,A), the expression is:

[0134]

[0135] Among them, Q i Let i be the value function of the i-th single agent;

[0136] S3.3. When multiple agents need to extract distributed policies, VDN guarantees that operating on the globally maximum independent variable and operating on the individually maximum independent variable on the agent network yields the same result, resulting in the expression:

[0137]

[0138] Each individual agent collects information through its agent network, aggregates it, and inputs it into a hybrid network for computation. The hybrid network consists of a supernetwork that uses the absolute value function as its heuristic function, ultimately calculating Q. tot (S,A);

[0139] S3.4. Set the TD target y TD,i With TD error δ TD,i The expression is:

[0140] yTD,i =reward + γQ i ′ (16)

[0141] δ TD,i =Q tot -y TD (17)

[0142] Set up backpropagation to obtain the gradient g. i The expression is:

[0143]

[0144] The expression for setting and updating network parameters is:

[0145] θ new ←θ now +α·δ TD,i ·g i (19)

[0146] S3.5. Repeat steps S3.1-S3.4 until the training results of the intelligent many-to-many game objective allocation strategy for the aircraft are output.

[0147] S4. Simulate and verify the training results of the target allocation strategy for intelligent multi-to-multi game of aircraft obtained in step S3.

[0148] Furthermore, the specific implementation method of step S4 includes the following steps:

[0149] S4.1. Determine the number of multiple aircraft in the simulation to be 16 and the number of targets to be 12. During environment initialization, the parameters of the aircraft and targets will be randomly selected for initialization.

[0150] S4.2. Set the hyperparameters of the OW-QMIX network. The hardware environment used for simulation is: CPU i5-11400F, GPU GTX1060. Under normal working conditions, the aircraft does not have partial or complete failure issues. The simulation is then performed for verification.

[0151] Furthermore, the hyperparameter settings for the OW-QMIX network are shown in Table 1.

[0152] Table 1 Hyperparameter Settings

[0153]

[0154]

[0155] In the multi-vehicle-to-multi-target assignment task, the number of multi-vehicles in the simulation is determined to be 16, and the number of targets is 12. During environment initialization, parameter perturbations of the aircraft and targets are randomly selected for initialization. The initial parameters of 16 aircraft in one training scenario are given in Table 2 below:

[0156] Table 2 Initial parameters of the aircraft

[0157]

[0158] Meanwhile, examples of the initial parameter settings for the target are shown in Table 3.

[0159] Table 3 Initial Parameters of the Target

[0160]

[0161] The simulation results and analysis under normal working conditions are as follows:

[0162] After setting the hyperparameters, the simulation hardware environment was: CPU i5-11400F, GPU GTX1060. Under normal operating conditions, the aircraft does not experience partial or complete failure. The training processes for the QMIX and OW-QMIX algorithms are as follows: Figure 5 As shown in the figure. The horizontal axis represents the number of training sessions, and the vertical axis represents the total reward value obtained by the agent. Taking the initial environment as an example, the initial conditions of the environment were changed multiple times, including the initial position and initial velocity, and the number of game targets and the number of aircraft were varied between 12 and 16. Finally, the completion scores of each algorithm were normalized and listed in Table 4. The normalized reward mentioned earlier is the completion score obtained by normalizing the interception result and fuel consumption.

[0163] Table 4. Algorithm Indicator Completion Statistics

[0164]

[0165] The index distributions of the three algorithms are as follows Figure 6 As shown.

[0166] The simulation results and analysis for the two failure scenarios are as follows:

[0167] This simulation exhibits an issue where some aircraft fail individually. The failure mechanism involves randomly numbered aircraft, and failed aircraft cannot participate in the mission. The training processes for the QMIX and OW-QMIX algorithms are as follows: Figure 7 As shown in Table 5, the horizontal axis represents the number of training sessions, and the vertical axis represents the total reward value obtained by the agent. After multiple simulations, the normalized performance metrics of each algorithm are listed below:

[0168] Table 5. Algorithm Indicator Completion Statistics

[0169]

[0170] The index distributions of the three algorithms are as follows Figure 8 As shown.

[0171] In summary, after Monte Carlo shooting tests, in the case of partial failures among multiple aircraft, the distributed multi-aircraft many-to-many game target allocation algorithm based on the OW-QMIX algorithm outperforms the QMIX algorithm in convergence speed. Similar to Simulation 1, it can be seen that the OW-QMIX algorithm has better convergence speed and index completion rate than both the QMIX and greedy algorithms. Overall, the OW-QMIX algorithm has a significant advantage. Specific Implementation Method Two:

[0173] An electronic device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the target allocation method for intelligent many-to-many games of an aircraft as described in Specific Embodiment 1.

[0174] The computer device of the present invention may include a processor and a memory, such as a microcontroller containing a central processing unit. Furthermore, the processor executes the computer program stored in the memory to implement the steps of the aforementioned intelligent many-to-many game target allocation method for aircraft.

[0175] The processor referred to can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0176] The memory may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function (such as sound playback, image playback, etc.); the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart media cards (SMC), secure digital cards (SD cards), flash cards, at least one disk storage device, flash memory device, or other volatile solid-state storage devices. Specific implementation method three:

[0178] A computer-readable storage medium storing a computer program thereon, characterized in that, when the computer program is executed by a processor, it implements the target allocation method for intelligent many-to-many game of aircraft as described in Specific Embodiment 1.

[0179] The computer-readable storage medium of the present invention can be any form of storage medium that can be read by the processor of a computer device, including but not limited to non-volatile memory, volatile memory, ferroelectric memory, etc. The computer-readable storage medium stores a computer program. When the processor of the computer device reads and executes the computer program stored in the memory, the steps of the above-mentioned intelligent many-to-many game target allocation method for aircraft can be implemented.

[0180] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0181] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0182] Although this application has been described above with reference to specific embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of this application. In particular, as long as there is no structural conflict, the features in the specific embodiments disclosed in this application can be combined with each other in any way. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, this application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A target allocation method for intelligent many-to-many game in aircraft, characterized in that, Includes the following steps: S1. Establish a game-theoretic relative motion model for aircraft; S2. Based on the relative motion model of aircraft game adversarial competition obtained in step S1, design the training environment, state space of multiple agents, and action space of agents for intelligent multi-to-multi game adversarial competition of aircraft. Collect the observation values ​​of aircraft in the training environment as the state space of multiple agents, and the input number of aircraft as the action space of agents. Calculate the fuel consumption and reward value of aircraft. The specific implementation method of step S2 includes the following steps: S2.

1. Based on the relative motion model of aircraft game adversarial competition obtained in step S1, design a training environment for intelligent many-to-many game adversarial competition of aircraft. Initialization settings are performed, including randomly generating multiple aircraft and game targets, and randomly determining the speed and direction of the aircraft. At the start of the simulation, each aircraft conducts a detection and acquires the relative position and speed information of the game targets within its field of view. The initial position range of the aircraft is set to 0~1km in the x-axis direction, while the target's x-axis position is 10~11km, y-axis position is -1~12km, and z-axis position is -10~10km. The initial speed value range is 5000~5500m / s, and the speed tilt angle and deflection angle range is -10°~10°. S2.

2. Define the state space of the multi-agent system. The expression is: ; In the formula, The three axes represent the relative positions of the first game's aircraft. Let be the rate of change of the relative distance between the aircraft in the first game; Define the action space of the intelligent agent The expression is: ; The agent's action space consists of the corresponding game-theoretic target assignment results. The corresponding number for the game's aircraft; S2.

3. Input the observations of each aircraft into the agent network. The agent network outputs the corresponding number of the aircraft in the game. The aircraft group simulates the relative motion model of the aircraft game confrontation. During the aircraft game confrontation, the fuel consumption of the aircraft is calculated, and the reward obtained by the allocation algorithm is calculated in combination with the result of the game confrontation. The expression for the aircraft's fuel consumption is: ; Where Fuel is the fuel consumption index, n is the total number of aircraft, i is any one of n, N represents the total overload of aircraft maneuvers, and t represents the total duration of aircraft maneuvers; The reward obtained by the allocation algorithm The expression is: ; in, Let be the fuel consumption index of the i-th aircraft, and let be the sum of the game competition reward value and the fuel consumption reward value. S3. Construct the OW-QMIX agent structure, and then input the state space, action space and reward value of the multi-agent obtained in step S2 into the OW-QMIX agent structure, and output the training results of the aircraft intelligent multi-to-multi game objective allocation strategy. The specific implementation method of step S3 includes the following steps: S3.

1. Input the state space, action space, and reward value of the multi-agent obtained in step S2 into the OW-QMIX agent structure, and randomly select the first step as the input. The flight trajectory is simulated using the dynamic equations in step S1. After completing a time step of 0.1s, the state space of the next moment is randomly obtained. Form the data into quadruples Store in experience pool D, and set the experience pool loop to stop at a maximum of 100,000 rounds; Then set the parameters of the current network as follows: The updated network parameters are as follows: The value function of the i-th single agent is obtained by forward propagation of the single agent network. The expression is: ; S3.

2. To address the value allocation problem in multi-agent learning, a Value Decomposition Network (VDN) is used for multi-agent training. The network weights of the OW-QMIX agent structure are set to Q(S,A;θ), the experience pool to D, and the learning rate to... Set the joint action value function as The expression is: ; in, Let i be the value function of the i-th single agent; S3.

3. When multiple agents need to extract distributed policies, VDN guarantees that operating on the globally maximum independent variable and operating on the individually maximum independent variable on the agent network yields the same result, resulting in the expression: ; Each individual agent collects information through its agent network, aggregates it, and inputs it into a hybrid network for computation. The hybrid network consists of a supernetwork that uses the absolute value function as its heuristic function, ultimately calculating... ; S3.

4. Set TD target With TD error The expression is: ; ; Set up backpropagation to obtain the gradient. The expression is: ; The expression for setting and updating network parameters is: ; S3.

5. Repeat steps S3.1-S3.4 until the training results of the target allocation strategy for the intelligent multi-to-multi game of the aircraft are output; S4. Simulate and verify the training results of the intelligent many-to-many game objective allocation strategy for aircraft obtained in step S3.

2. The target allocation method for intelligent many-to-many game in aircraft according to claim 1, characterized in that, The specific implementation method of step S1 includes the following steps: S1.

1. Construct the three-degree-of-freedom kinematic equations for the aircraft's flight outside the atmosphere, expressed as follows: ; in, Let V be the rate of change of the aircraft's velocity relative to the x-axis, and let V be the displacement vector and velocity vector of the aircraft in the ballistic frame relative to the inertial frame. The velocity change rate of the aircraft relative to the y-axis. Let z be the rate of change of the aircraft's velocity relative to the z-axis. For the track angle, For heading angle; ; in, pitch angle The rate of change of velocity, Yaw angle The rate of change of velocity, For roll angle The rate of change of velocity, Let be the angular velocity of the aircraft relative to the x-axis. Let be the angular velocity of the aircraft relative to the y-axis. The angular velocity of the aircraft relative to the z-axis; S1.

2. Construct the dynamic equations for the aircraft's flight outside the atmosphere, expressed as follows: ; in, For the mass of the aircraft, The geocentric radius of the spacecraft Let these be the displacement vector and velocity vector of the aircraft in the ballistic frame relative to the inertial frame. Let be the angular velocity of the ballistic frame relative to the launch frame. To compete for the thrust of aircraft engines, In addition to thrust, aerodynamic forces and disturbance forces, To determine the angular momentum of the center of mass of the aircraft. The torque generated by the thrust eccentricity of the aircraft engine. For aerodynamic torque and disturbance torque, The angular velocity of the aircraft's own system relative to the inertial frame; S1.

3. Set the position of the spacecraft at time T as follows: The position of the target aircraft at time T is Construct a motion position parameter model for the aircraft and the target aircraft, with the following expression: relative position The expression is: ; The relative velocity can be obtained by differentiating from the relative distance. The expression is: ; After obtaining the relative speed, the corresponding distance proximity rate is obtained. The expression is: ; Angle of view and line of sight angle The expression is: ; in, , ; Since the target uses proportional guidance, the expression for the aircraft-target line-of-sight angle rotation rate is: ; in, For the aircraft-target line-of-sight tilt angle rotation, The deflection rate of the aircraft-target line of sight.

3. The target allocation method for intelligent many-to-many game in aircraft according to claim 2, characterized in that, The specific implementation method of step S4 includes the following steps: S4.

1. The number of aircraft in the simulation is set to 16, and the number of targets is set to 12. During environment initialization, parameter perturbations for aircraft and targets will be randomly selected. S4.

2. Set the hyperparameters of the OW-QMIX network. The hardware environment used for simulation is: CPU i5-11400F, GPU GTX1060. Under normal working conditions, the aircraft does not have partial or complete failure issues. The simulation is then performed for verification.

4. An electronic device, characterized in that, It includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the intelligent many-to-many game target allocation method for aircraft as described in any one of claims 1-3.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the target allocation method for intelligent many-to-many games of aircraft as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Collaborative battle method and device of intelligent agent

    CN113893539A

  • BC-QMIX off-line and on-line multi-agent behavior decision modeling method oriented to military force game confrontation

    CN115964898A