Multi-unmanned aerial vehicle cooperative game autonomous decision-making method based on deep reinforcement learning
Through the autonomous decision-making method of collaborative game of multi-UAVs based on deep reinforcement learning, the problem of insufficient flexibility and adaptability of traditional technologies in complex battlefield environments is solved, and efficient collaboration and adaptive control of drones in multi-target and multi-task environments is achieved.
Patent Information
- Application Number
- CN202510030285.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-23
AI Technical Summary
The traditional stand-alone game confrontation model cannot meet the multi-objective and multi-task needs in complex battlefield environments, and lacks sufficient flexibility and adaptability in the multi-drone collaborative game control.
The multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning is adopted. By obtaining the aircraft's feature information, inputting the preset situation evaluation model, obtaining relative situation information, and reconstructing the target observation value through a multi-level adaptive feature fusion framework, and finally inputting it into the decision model trained by deep reinforcement learning to control the aircraft.
It has realized that drones independently learn and adjust their flight strategies in complex air combat environments, dynamically optimize maneuver paths, and maintain coordination and formation with other drones, achieving efficient tactical coordination and adaptive response to enemy threats.
Smart Images

Figure CN120029053A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aircraft control technology, and in particular to a multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning. Background Art
[0002] With the rapid development of drone technology, the traditional single-machine game confrontation mode can no longer meet the multi-target and multi-task requirements in complex battlefield environments. Compared with a single drone, the coordinated confrontation of multiple drones can effectively improve combat efficiency, increase tactical flexibility, and have strong robustness. Through collaborative cooperation, drone swarms can not only improve overall confrontation capabilities, but also realize the allocation and execution of complex tasks. However, with the increase in the number of drones and the complexity of tasks, how to achieve collaborative game control of multiple drones has become a major challenge. In the drone game confrontation environment, each drone not only needs to consider its own flight status, but also must make collaborative decisions with other drones to deal with enemy threats and tactical changes.
[0003] Traditional adversarial decision-making methods, such as those based on game theory and optimization, can guarantee the mission execution of drones to a certain extent, but they often lack sufficient flexibility and adaptability when faced with highly dynamic and complex battlefield environments. Summary of the invention
[0004] Based on this, it is necessary to propose an aircraft control method and related equipment to address the problem that the existing technology lacks sufficient flexibility and adaptability in aircraft (UAV) game confrontation.
[0005] In a first aspect, a multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning is provided, the method comprising:
[0006] Acquiring characteristic information of the aircraft, and inputting the characteristic information into a preset situation assessment model to obtain relative situation information;
[0007] Reconstructing the relative situation information through a context-aware multi-level adaptive feature fusion framework to obtain a target observation value;
[0008] The target observation value is input into a preset decision model to obtain decision data for the feature information, and the aircraft is controlled according to the decision data.
[0009] Optionally, before the step of inputting the feature information into a preset situation assessment model to obtain relative situation information, the step further includes:
[0010] The lateral motion, longitudinal motion and vertical motion of the aircraft are modeled to obtain a degree of freedom model of the aircraft and a dynamic model of the aircraft. The mathematical representation of the degree of freedom model is:
[0011]
[0012] Among them, x, y, z are the coordinate values of the drone's position, p = [x, y, z], v is the drone's velocity vector, and v represents the rate of change of position in three directions, μ, θ, Represent the roll angle, track angle and heading angle respectively, is the projection of v on the xoy plane is the projection of v on the xoy plane;
[0013] The mathematical representation of the dynamic model of the aircraft is:
[0014]
[0015] Where g is the acceleration due to gravity, n x is the overload in the speed direction, n z is the normal overload, μ is the roll angle, and the basic control parameter n x , n z and μ can be expressed as a control input vector u = [n x , n z , μ], n x Used to control the speed of the drone, n z and μ are used to control the direction of the velocity vector.
[0016] Optionally, before the step of inputting the target observation value into a preset decision model, the method further includes:
[0017] Obtaining an initial observation value output by the multi-level adaptive feature fusion framework, and inputting the initial observation value into an Actor-Critic network to obtain a decision action for the initial observation value;
[0018] Controlling the aircraft to execute the decision action, and acquiring status data of the aircraft after executing the decision action;
[0019] The reward and punishment values are calculated based on the state data and the preset reward function, and the Actor-Critic network is trained by the gradient descent method according to the reward and punishment values until the reward and punishment values converge to obtain the trained Actor-Critic network, and a preset decision model is obtained according to the Actor-Critic network.
[0020] Optionally, the preset reward function includes: a relative situation reward function, a battlefield departure penalty function, a defeat reward function, a global reward function and a comprehensive reward function;
[0021] Among them, the mathematical representation of the relative situation reward function is:
[0022]
[0023] Among them, X i,j Represents the situation function of UAV i relative to UAV j. For UAV i With UAV j The distance between is the maximum sensing distance of the sensor, γ S is the reward coefficient of relative situation reward;
[0024] The mathematical representation of the battlefield departure penalty function is:
[0025]
[0026] Among them, Zone high is the maximum value of the battlefield boundary, Zone low is the minimum value of the battlefield boundary, γ p The reward coefficient for the penalty of crossing the boundary;
[0027] The mathematical representation of the defeat reward function is:
[0028]
[0029] The mathematical representation of the global reward function is:
[0030]
[0031] Among them, Number blue and Number red are the number of our and enemy aircraft remaining at the end of each round, γ w is the reward coefficient of the global reward;
[0032] The comprehensive reward function is composed of a relative situation reward function, a battlefield exit penalty function, a defeat reward function, and a global reward function. The mathematical representation of the comprehensive reward function is:
[0033]
[0034] Optionally, the multi-level adaptive feature fusion framework includes a bottom-level GCN network, a middle-level GAT network, and a high-level multi-head attention network;
[0035] Among them, the mathematical representation of the underlying GCN network is:
[0036]
[0037] The mathematical representation of the middle-level GAT network is:
[0038]
[0039] The mathematical representation of the high-level multi-head attention network is:
[0040] Q=H'W Q , K=H'W K , V=H'W V .
[0041] Optionally, the step of reconstructing the relative situation information through a context-aware multi-level adaptive feature fusion framework to obtain a target observation value includes:
[0042] Inputting the relative situation information into the underlying GCN network in a multi-level adaptive feature fusion framework based on context awareness to obtain initial feature data;
[0043] Inputting the initial feature data into a middle-layer GAT network in a multi-level adaptive feature fusion framework based on context perception to obtain fused feature data;
[0044] Inputting the fused feature data into a high-level multi-head attention network in a context-aware multi-level adaptive feature fusion framework to obtain output vector data;
[0045] Combine the global information of the current moment and the historical moment to construct an embedding vector based on contextual information;
[0046] The relative situation information is reconstructed according to the output vector data and the embedded vector to obtain a target observation value.
[0047] In a second aspect, a multi-UAV collaborative game autonomous decision-making device based on deep reinforcement learning is provided, the device comprising:
[0048] A data acquisition module is used to obtain characteristic information of the aircraft and input the characteristic information into a preset situation assessment model to obtain relative situation information;
[0049] A reconstruction module, used to reconstruct the relative situation information through a multi-level adaptive feature fusion framework based on context awareness to obtain a target observation value;
[0050] The decision module is used to input the target observation value into a preset decision model to obtain decision data for the feature information and control the aircraft according to the decision data.
[0051] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of the above-mentioned multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning are implemented.
[0052] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning are implemented.
[0053] The present application obtains the characteristic information of the aircraft and inputs the characteristic information into a preset situation assessment model to obtain relative situation information; reconstructs the relative situation information through a multi-level adaptive feature fusion framework based on context perception to obtain target observation values; inputs the target observation values into a preset decision model to obtain decision data for the characteristic information, and controls the aircraft according to the decision data. The preset decision model is a decision model trained by deep reinforcement learning, which enables aircraft such as drones to autonomously learn how to adjust flight strategies in complex air combat environments and dynamically optimize maneuvering paths while maintaining coordination and formation with other drones, so as to achieve efficient tactical coordination and adaptive response to enemy threats. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0055] in:
[0056] Figure 1 It is a flowchart of a multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning in one embodiment;
[0057] Figure 2 A multi-level adaptive feature fusion framework diagram based on context perception in a multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning in one embodiment;
[0058] Figure 3A framework detail diagram of a context-aware multi-level adaptive feature fusion framework in a multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning in one embodiment;
[0059] Figure 4 It is a structural block diagram of a multi-UAV collaborative game autonomous decision-making device based on deep reinforcement learning in one embodiment;
[0060] Figure 5 is a structural block diagram of a computer device in one embodiment;
[0061] Figure 6 It is a structural block diagram of a computer device in another embodiment. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] The present invention is described in detail below through specific embodiments.
[0064] See also Figure 1 As shown, Figure 1 A flowchart of a multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning provided by an embodiment of the present invention includes the following steps:
[0065] S101, acquiring characteristic information of an aircraft, and inputting the characteristic information into a preset situation assessment model to obtain relative situation information;
[0066] Exemplarily, the aircraft may be understood as a flying device such as a drone.
[0067] S102, reconstructing the relative situation information through a multi-level adaptive feature fusion framework based on context perception to obtain a target observation value;
[0068] S103: Input the target observation value into a preset decision model to obtain decision data for the feature information, and control the aircraft according to the decision data.
[0069] Relative situation information is obtained by acquiring characteristic information of an aircraft and inputting the characteristic information into a preset situation assessment model; the relative situation information is reconstructed through a multi-level adaptive feature fusion framework based on context perception to obtain a target observation value; the target observation value is input into a preset decision model to obtain decision data for the characteristic information, and the aircraft is controlled according to the decision data. The preset decision model is a decision model trained by deep reinforcement learning, which enables aircraft such as drones to autonomously learn how to adjust flight strategies in complex air combat environments, dynamically optimize maneuvering paths, and maintain coordination and formation with other drones, so as to achieve efficient tactical coordination and adaptive response to enemy threats.
[0070] In a possible embodiment, before the step of inputting the feature information into a preset situation assessment model to obtain relative situation information, the method further includes:
[0071] The lateral motion, longitudinal motion and vertical motion of the aircraft are modeled to obtain a degree of freedom model of the aircraft and a dynamic model of the aircraft. The mathematical representation of the degree of freedom model is:
[0072]
[0073] Among them, x, y, z are the coordinate values of the drone's position, p = [x, y, z], v is the drone's velocity vector, and v represents the rate of change of position in three directions, μ, θ, Represent the roll angle, track angle and heading angle respectively, is the projection of v on the xoy plane is the projection of v on the xoy plane;
[0074] The mathematical representation of the dynamic model of the aircraft is:
[0075]
[0076] Where g is the acceleration due to gravity, n x is the overload in the speed direction, n z is the normal overload, μ is the roll angle, and the basic control parameter n x , n z
[0077] and μ can be expressed as a control input vector u = [n x , n z , μ], n x Used to control the speed of the drone, n z and μ are used to control the direction of the velocity vector.
[0078] For example, since the maneuvering decision of an aircraft such as a drone mainly considers the relative position and relative speed relationship between drones, the lateral motion, longitudinal motion and vertical motion of the drone are modeled as a three-degree-of-freedom model of the drone as a motion model of the drone.
[0079] When the speed direction of the drone coincides with the fuselage axis, the motion model of the drone is:
[0080]
[0081] Among them, x, y, z are the coordinate values of the drone's position, p = [x, y, z], v is the drone's velocity vector, and v represents the rate of change of position in three directions, μ, θ, Represent the roll angle, track angle and heading angle respectively, is the projection of v on the xoy plane.
[0082] The dynamic model of the UAV is:
[0083]
[0084] Where g is the acceleration due to gravity, n x is the overload in the speed direction, n z is the normal overload, μ is the roll angle, and the basic control parameter n x , n z
[0085] and μ can be expressed as a control input vector u = [n x , n z , μ], n x Used to control the speed of the drone, n z and μ are used to control the direction of the velocity vector.
[0086] In a possible embodiment, before the step of inputting the target observation value into a preset decision model, the method further includes:
[0087] Obtaining an initial observation value output by the multi-level adaptive feature fusion framework, and inputting the initial observation value into an Actor-Critic network to obtain a decision action for the initial observation value;
[0088] Controlling the aircraft to execute the decision action, and acquiring status data of the aircraft after executing the decision action;
[0089] The reward and punishment values are calculated based on the state data and the preset reward function, and the Actor-Critic network is trained by the gradient descent method according to the reward and punishment values until the reward and punishment values converge, so as to obtain the trained Actor-Critic network, and the preset decision model is obtained according to the Actor-Critic network.
[0090] For example, the gradient descent algorithm is also commonly referred to as the batch gradient descent algorithm. Batch gradient descent uses the entire training set for each learning, so these calculations are redundant because the exact same sample set is used each time. But its advantage is that each update will proceed in the right direction, and it is guaranteed to converge to the extreme point in the end.
[0091] In a possible embodiment, the preset reward function includes: a relative situation reward function, a battlefield departure penalty function, a defeat reward function, a global reward function and a comprehensive reward function;
[0092] Among them, the mathematical representation of the relative situation reward function is:
[0093]
[0094] Among them, X i,j Represents the situation function of UAV i relative to UAV j. For UAV i With UAV j The distance between is the maximum sensing distance of the sensor, γ s is the reward coefficient of relative situation reward;
[0095] The mathematical representation of the battlefield departure penalty function is:
[0096]
[0097] Among them, Zone high is the maximum value of the battlefield boundary, Zone low is the minimum value of the battlefield boundary, γ p The reward coefficient for the penalty of crossing the boundary;
[0098] The mathematical representation of the defeat reward function is:
[0099]
[0100] The mathematical representation of the global reward function is:
[0101]
[0102] Among them, Number blue and Number redare the number of our and enemy aircraft remaining at the end of each round, γ w is the reward coefficient of the global reward;
[0103] The comprehensive reward function is composed of a relative situation reward function, a battlefield exit penalty function, a defeat reward function, and a global reward function. The mathematical representation of the comprehensive reward function is:
[0104]
[0105] For example, a reward function is constructed based on the task requirements to guide the UAV to learn the optimal maneuver strategy. Therefore, a reward function suitable for the multi-UAV collaborative air combat task is designed to guide the learning process. Specifically:
[0106] Relative situation reward function, the situation function value matrix is calculated once at each step. For each drone at each moment, calculate the relative situation value of the drone relative to other enemy drones within its observation range. Multiply the relative situation function by the reward factor as the relative situation reward value. The relative situation reward function of the drone is:
[0107]
[0108] Among them, X i,j Represents the situation function of UAV i relative to UAV j. For UAV i With UAV j The distance between is the maximum sensing distance of the sensor, γ s is the reward coefficient of the relative situation reward.
[0109] The battlefield penalty function is determined once at each time step based on the coordinate value. The out-of-bounds penalty can guide the UAV to make maneuver decisions within a reasonable air combat area. The battlefield penalty function can be expressed by the following formula:
[0110]
[0111] Among them, Zone high is the maximum value of the battlefield boundary, Zone low is the minimum value of the battlefield boundary, γ p The reward coefficient for out-of-bounds punishment.
[0112] The defeat reward function is to reward our drone whenever it defeats the enemy drone, and punish it otherwise. The defeat reward function can be expressed by the following formula:
[0113]
[0114] The global reward function is determined by the group victory or defeat of the drone air combat. The reward is determined at the end of each round. Based on the number of drones remaining on both sides at the end of the confrontation, all our drones are given the same global reward or penalty. This can guide drones to learn group cooperative maneuvering strategies and ultimately win the air combat. The calculation formula of the global reward function can be expressed as follows:
[0115]
[0116] Number blue and Number red are the number of remaining friendly and enemy drones at the end of each round, γ w is the reward coefficient of the global reward.
[0117] The comprehensive reward function consists of relative situation reward, battlefield exit penalty, defeat reward and global reward, which is expressed by the formula of the comprehensive reward function:
[0118]
[0119] The reward coefficients involved in the reward are assigned to γ s =0.9,γ p =20.0,γ d =50.0,γ w =100.0.
[0120] In a possible implementation, the multi-level adaptive feature fusion framework includes a bottom-level GCN network, a middle-level GAT network, and a high-level multi-head attention network;
[0121] Among them, the mathematical representation of the underlying GCN network is:
[0122]
[0123] The mathematical representation of the middle-level GAT network is:
[0124]
[0125] The mathematical representation of the high-level multi-head attention network is:
[0126] Q=H′W Q , K=H′W K , V=H′W V .
[0127] In a possible implementation, the step of reconstructing the relative situation information through a context-aware multi-level adaptive feature fusion framework to obtain a target observation value includes:
[0128] Inputting the relative situation information into the underlying GCN network in a multi-level adaptive feature fusion framework based on context awareness to obtain initial feature data;
[0129] Inputting the initial feature data into a middle-layer GAT network in a multi-level adaptive feature fusion framework based on context perception to obtain fused feature data;
[0130] Inputting the fused feature data into a high-level multi-head attention network in a context-aware multi-level adaptive feature fusion framework to obtain output vector data;
[0131] Combine the global information of the current moment and the historical moment to construct an embedding vector based on contextual information;
[0132] The relative situation information is reconstructed according to the output vector data and the embedded vector to obtain a target observation value.
[0133] For example, Figure 2-Figure 3 As shown in the figure, first, the drone passes the state information (relative situation information) captured by the sensor to the underlying GCN network. The underlying GCN network is responsible for processing the drone sensor data, aggregating the information of each node and its local neighbors, and generating feature representations of the nodes. It captures the basic connections and shared features between nodes. It provides a higher level of abstract representation for the middle-level GAT. The calculation formula of the underlying GCN network is as follows:
[0134]
[0135] Among them, H (l) Represents the node feature matrix of the lth layer, which reduces the training parameters and computational complexity while ensuring effective data processing. This paper adopts a single-layer GCN network. The initial input H (0) =[S 1 , S 2 , …, S N ] T , is the adjacency matrix with self-connection A is the adjacency matrix, which is used to represent the connection relationship between nodes in the graph. N is the identity matrix, which is used to contain the characteristics of each node itself. yes The degree matrix, W (l) is the weight matrix of the lth layer, and σ is the activation function ReLU function.
[0136] Based on the features (initial feature data) extracted by the underlying GCN network, the adaptive feature fusion module adaptively fuses the feature information of different nodes in the middle-layer GAT network according to the changes in the battlefield situation, thereby flexibly learning the complex relationship between nodes and dynamically allocating attention according to the importance of the nodes. Effective dimensionality reduction of group situation information is achieved. The calculation formula of the middle layer of the framework is as follows:
[0137]
[0138] Where W is the weight matrix used to transform the feature vector of each node into a higher-dimensional space. β is the parameter vector used to calculate the attention coefficient. || represents the concatenation operation, and LeakyReLU is the activation function used to increase the nonlinearity of the network. represents the neighbor set of node i. The final node feature can be expressed as:
[0139]
[0140] Where K is the number of attention heads, and W k is the attention coefficient and weight matrix of the kth attention head.
[0141] The high-level framework uses multi-head attention to stabilize the self-attention method to integrate the information of all drone networks and construct a global task vector representing the overall task. The high-level calculation formula of the framework is as follows
[0142] Q=H'W Q , K=H'W K , V=H'W V
[0143] Where H′ is the input feature matrix, W Q , W K , W V are the weight matrices corresponding to query, key, and value, respectively.
[0144]
[0145] The output calculated by each head is the attention weight, A is the attention score, d k is the dimension of the key vector, and the division operation is to scale the result of the dot product, which helps the stability of training. Finally, the outputs of all heads are concatenated and fused through a final linear layer to obtain the output vector.
[0146] W′ is another weight matrix used to fuse the outputs of all heads into the final output vector.
[0147] Since the situational relationship between the air combat environment and the UAV is highly time-varying, relying solely on the state information at a certain moment is not enough to describe the entire battlefield. In order to better capture the situational relationship based on time series, the global information of the current moment and the historical moment is combined to construct an embedding vector based on contextual information to reconstruct the original observation value, thereby better guiding the UAV to make air combat maneuver decisions. The generation of the message is calculated by the following formula:
[0148]
[0149] in, is the characteristic of edge ij, t ij It is the time when the edge occurs. It is the state at the previous moment;
[0150] The initial input is h′ i is the initial feature of node i, f init is the initialization function. MSG Is the function used to generate the message. represents the neighbors of node i, f AGG Is an aggregate function.
[0151]
[0152] in, is the message received at the current moment, f update is the state update function. The encoding based on time features is obtained by the following formula
[0153] t' i =TE(t i )
[0154] Among them, t i is the timestamp, TE is the time encoding function, expressed as a cosine function. The final output observation value is obtained by the following formula
[0155]
[0156] Among them, f output is the output function, and the function f is composed of fully connected layers.
[0157] In one embodiment, Figure 4 As shown, a multi-UAV collaborative game autonomous decision-making device based on deep reinforcement learning is provided, and the device includes:
[0158] The data acquisition module 201 is used to obtain characteristic information of the aircraft and input the characteristic information into a preset situation assessment model to obtain relative situation information;
[0159] A reconstruction module 202, configured to reconstruct the relative situation information through a multi-level adaptive feature fusion framework based on context awareness to obtain a target observation value;
[0160] The decision module 203 is used to input the target observation value into a preset decision model to obtain decision data for the feature information and control the aircraft according to the decision data.
[0161] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning.
[0162] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning.
[0163] In one embodiment, a computer device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: acquiring characteristic information of an aircraft, and inputting the characteristic information into a preset situation assessment model to obtain relative situation information; reconstructing the relative situation information through a multi-level adaptive feature fusion framework based on context awareness to obtain a target observation value; inputting the target observation value into a preset decision model to obtain decision data for the characteristic information, and controlling the aircraft according to the decision data.
[0164] In one embodiment, a computer-readable storage medium is proposed, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: acquiring characteristic information of an aircraft, and inputting the characteristic information into a preset situation assessment model to obtain relative situation information; reconstructing the relative situation information through a multi-level adaptive feature fusion framework based on context awareness to obtain a target observation value; inputting the target observation value into a preset decision model to obtain decision data for the characteristic information, and controlling the aircraft according to the decision data.
[0165] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0166] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0167] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0168] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning, characterized in that: The method comprises: Acquiring characteristic information of the aircraft, and inputting the characteristic information into a preset situation assessment model to obtain relative situation information; Reconstructing the relative situation information through a context-aware multi-level adaptive feature fusion framework to obtain a target observation value; The target observation value is input into a preset decision model to obtain decision data for the feature information, and the aircraft is controlled according to the decision data.
2. The multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning according to claim 1 is characterized in that: Before the step of inputting the characteristic information into a preset situation assessment model to obtain relative situation information, the step further includes: The lateral motion, longitudinal motion and vertical motion of the aircraft are modeled to obtain a degree of freedom model of the aircraft and a dynamic model of the aircraft. The mathematical representation of the degree of freedom model is: Among them, x, y, z are the coordinate values of the drone's position, p = [x, y, z], v is the drone's velocity vector, and v represents the rate of change of position in three directions, μ, θ, Represent the roll angle, track angle and heading angle respectively, is the projection of v on the xoy plane; The mathematical representation of the dynamic model of the aircraft is: Where g is the acceleration due to gravity, n x is the overload in the speed direction, n z is the normal overload, μ is the roll angle, and the basic control parameter n x , n z and μ can be expressed as a control input vector u = [n x , n z , μ], n x Used to control the speed of the drone, n z and μ are used to control the direction of the velocity vector.
3. The multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning according to claim 1 is characterized in that: Before the step of inputting the target observation value into a preset decision model, the method further includes: Obtaining an initial observation value output by the multi-level adaptive feature fusion framework, and inputting the initial observation value into an Actor-Critic network to obtain a decision action for the initial observation value; Controlling the aircraft to execute the decision action, and acquiring status data of the aircraft after executing the decision action; The reward and punishment values are calculated based on the state data and the preset reward function, and the Actor-Critic network is trained by the gradient descent method according to the reward and punishment values until the reward and punishment values converge, so as to obtain the trained Actor-Critic network, and the preset decision model is obtained according to the Actor-Critic network.
4. The multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning according to claim 3 is characterized in that: The preset reward functions include: relative situation reward function, battlefield departure penalty function, defeat reward function, global reward function and comprehensive reward function; Among them, the mathematical representation of the relative situation reward function is: Among them, X i,j Represents the situation function of UAV i relative to UAV j. For UAV i With UAV j The distance between is the maximum sensing distance of the sensor, γ S is the reward coefficient of relative situation reward; The mathematical representation of the battlefield departure penalty function is: Among them, Zone high is the maximum value of the battlefield boundary, Zone low is the minimum value of the battlefield boundary, γ p The reward coefficient for the penalty of crossing the boundary; The mathematical representation of the defeat reward function is: The mathematical representation of the global reward function is: Among them, Number blue and Number red are the number of our and enemy aircraft remaining at the end of each round, γ w is the reward coefficient of the global reward; The comprehensive reward function is composed of a relative situation reward function, a battlefield exit penalty function, a defeat reward function, and a global reward function. The mathematical representation of the comprehensive reward function is:
5. The multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning according to claim 1 is characterized in that: The multi-level adaptive feature fusion framework includes a bottom-level GCN network, a middle-level GAT network, and a high-level multi-head attention network; Among them, the mathematical representation of the underlying GCN network is: The mathematical representation of the middle-level GAT network is: The mathematical representation of the high-level multi-head attention network is: Q=H′W Q ,K=H′W K ,V=H′W V 。 6. The multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning according to claim 5 is characterized in that: The step of reconstructing the relative situation information through a context-aware multi-level adaptive feature fusion framework to obtain a target observation value comprises: Inputting the relative situation information into the underlying GCN network in a multi-level adaptive feature fusion framework based on context awareness to obtain initial feature data; Inputting the initial feature data into a middle-layer GAT network in a multi-level adaptive feature fusion framework based on context perception to obtain fused feature data; Inputting the fused feature data into a high-level multi-head attention network in a context-aware multi-level adaptive feature fusion framework to obtain output vector data; Combine the global information of the current moment and the historical moment to construct an embedding vector based on contextual information; The relative situation information is reconstructed according to the output vector data and the embedded vector to obtain a target observation value.
7. A multi-UAV collaborative game autonomous decision-making device based on deep reinforcement learning, characterized in that: The device comprises: A data acquisition module is used to obtain characteristic information of the aircraft and input the characteristic information into a preset situation assessment model to obtain relative situation information; A reconstruction module, used to reconstruct the relative situation information through a multi-level adaptive feature fusion framework based on context awareness to obtain a target observation value; The decision module is used to input the target observation value into a preset decision model to obtain decision data for the feature information and control the aircraft according to the decision data.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the multi-UAV collaborative game autonomous decision-making method based on deep reinforcement learning are implemented as described in any one of claims 1 to 6.
Citation Information
Cited By
Target matching method and device for unmanned aerial vehicle cluster, terminal and storage medium
CN120335496A