Unmanned aerial vehicle short-range air combat maneuvering autonomous decision method, system, device and terminal

By improving state representation and introducing node clustering methods, the complexity and computational efficiency issues of autonomous decision-making in close-range UAV air combat were resolved, achieving efficient autonomous learning and decision-making and enhancing the UAV's maneuverability in complex situations.

CN116400718BActive Publication Date: 2025-12-23PLA AIR FORCE AVIATION UNIVERSITY
View PDF -1 Cites 0 Cited by

Patent Information

Application Number
CN202310362825.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2025-12-23
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

Autonomous decision-making in close-range air combat by UAVs is a complex problem. Existing algorithms have difficulty converging under complex situations and require a large amount of computation. The accumulation of a large amount of experience in the experience pool leads to low efficiency, and traditional methods are difficult to effectively handle air combat experience.

Method used

We adopt an autonomous decision-making method for UAV close-range air combat maneuvers based on node clustering and deep deterministic policy gradient (DDPG). By improving the state representation and introducing fuzzy clustering, we eliminate similar experiences and retain typical experiences, and design an autonomous learning and decision-making algorithm.

Benefits of technology

It improves the algorithm efficiency of UAVs' autonomous decision-making in close-range air combat, reduces the demand for storage and computing resources, and enhances the speed and accuracy of autonomous learning and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116400718B_ABST
    Figure CN116400718B_ABST
Patent Text Reader

Abstract

The present application belongs to the field of unmanned aerial vehicle maneuver decision technology, and discloses a kind of unmanned aerial vehicle short-range air combat maneuver autonomous decision method, system, equipment and terminal, establish short-range air combat situation, missile attack area and the expression of missile allowed launch condition is achieved, construct unmanned aerial vehicle movement dynamics model and carry out action representation in reinforcement learning;According to node and fuzzy clustering method, similar experience in experience pool is clustered and excluded, to ensure that the experience retained is typical;Design air combat maneuver autonomous learning, autonomous decision algorithm based on node clustering and DDPG, realize the autonomous decision of short-range air combat maneuver.The present application is used to exclude similar experience, and retain the most typical experience.The present application effectively solves the problem of large amount of experience pool data and low algorithm efficiency when making decision on complex tasks.The present application improves the short-range air combat problem description model, changes the state expression method according to the characteristics of short-range air combat, and better reflects the characteristics of short-range air combat.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of unmanned aerial vehicle (UAV) maneuver decision-making technology, and particularly relates to a method, system, device and terminal for autonomous decision-making in close-range air combat of UAVs. Background Technology

[0002] Currently, the autonomous decision-making problem of drones in close-range air combat remains a challenging and cutting-edge research hotspot, and some countries are conducting technological exploration and experimental verification.

[0003] The challenges of air combat maneuver decision-making lie in several aspects. First, from an air combat perspective, the nonlinear changes in state and the abrupt shifts in strategy necessitate powerful neural networks to reflect these nonlinear changes. Second, the learning of decision-making neural networks requires a large number of supervised samples to achieve correct decisions under various states, but air combat experience is difficult to acquire, requiring novel intelligent methods to address this. Third, the maneuver decision-making problem itself is complex; air combat states change with factors such as weapons, targets, environment, the aircraft itself, and target characteristics, making it a highly complex decision-making problem. Furthermore, air combat experience needs to be acquired with a holistic and process-oriented perspective, rather than based on short-term gains and losses. These characteristics of the air combat maneuver decision-making problem pose insurmountable obstacles to traditional mathematical methods, urgently requiring novel solutions.

[0004] The emergence and development of deep reinforcement learning has undoubtedly provided a powerful tool for solving the problem of air combat maneuver decision-making. It has both the "trial-and-error" characteristics of reinforcement learning and the powerful mapping ability of neural networks, and can realize the ideal self-learning mode of learning and improving at the same time and trying and learning at the same time.

[0005] Some researchers have begun to try to handle the air combat maneuver decision problem based on different deep reinforcement learning algorithms. Several scholars have used the DQN method for autonomous learning and decision-making in maneuvering, such as Liu Pin and Yang Qiming [1,2]. Hu Dongyuan [3] improved the DQN algorithm by replacing the policy network in DQN with a perceptual context layer and a value fitting layer and applied it to maneuver decision-making in beyond-visual-range air combat. Li Yongfeng [4] proposed a multi-step dual-network deep reinforcement learning algorithm (MS-DDQN) based on the maneuver decision-making requirements of close-range air combat and successfully realized autonomous learning and decision-making in close-range air combat. However, DQN and its improved algorithms only solve the decision problem in continuous state space. Its action space adopts the form of a maneuver action library. The limited maneuver action library is difficult to reflect the maneuver actions in actual air combat.

[0006] The Deep Deterministic Policy Gradient (DDPG) method integrates the ideas of DQN and PPG, and can learn and make decisions autonomously in continuous state space and action space. Yang[5] established an air combat decision training framework based on DDPG and used optimization algorithm to generate prior air combat maneuver values ​​to improve the learning efficiency of DDPG algorithm. However, its aircraft model is simplified, and the simulation verification is relatively simple. Jing Xianyong[6] improved the experience playback strategy, ε-Greedy, and experience pool concentration adjustment of DDPG algorithm, and successfully applied it to close-range air combat maneuver decision. However, the efficiency of the improved algorithm still needs to be further improved when facing equal and disadvantageous situations. Kong7 studied an air combat strategy generation method based on multi-agent deep deterministic policy gradient (MADDPG) algorithm for computer force generation simulation verification, but it also lacks detailed discussion and sufficient simulation verification.

[0007] -------------------------

[0008] 1 Liu P,Ma YA Deep reinforcement learning based intelligent decision method for UCAV air combat.Singapore:Springer; 2017.p.274-286.https: / / doi.org / 10.1007 / 978-981-10-6463-0_24.

[0009] 2 Yang Q, Zhang J, Shi G, Wu Y. Maneuver decision of UCAV in short-range air combat based on deep reinforcement learning. IEEE Access 2019; 8:363-368. https: / / doi.org / 10.1109 / ACCESS.2019.2961426.

[0010] 3D.Hu,R.Yang,J.Zuo,Z.Zhang,J.Wu and Y.Wang,"Application of DeepReinforcement Learning in Maneuver Planning of Beyond-Visual-Range AirCombat,"in IEEE Access,vol.9,pp.32282-32297,2021,doi:10.1109 / ACCESS.2021.3060426.

[0011] 4 Y.-f.Li,J.-p.Shi,W.Jiang et al.,Autonomous maneuver decision-makingfor a UCAV in short-range aerial combat based on an MS-DDQN algorithm,DefenceTechnology,https: / / doi.org / 10.1016 / j.dt.2021.09.014

[0012] 5 Yang Q,Zhu Y,Zhang J,Qiao S,Liu J.UCAV air combat autonomousmaneuver decision based on DDPG algorithm.In:IEEE 15th internationalconference on control and automation;2019.p.37e42.https: / / doi.org / 10.1109 / ICCA.2019.8899703.

[0013] 6 Jing.Xianyong,M.Hou,G.Wu,Z.Ma and Z.Tao,"Research on ManeuveringDecision Algorithm Based on Improved Deep Deterministic Policy Gradient,"inIEEE Access,vol.10,pp.92426-92445,2022,doi:10.1109 / ACCESS.2022.3202918.

[0014] 7W.KONG, D.ZHOU and Z.YANG, "Air Combat Strategies Generation of CGFBased on MADDPG and Reward Shaping," 2020International Conference on ComputerVision, Image and Deep Learning(CVIDL), 2020, pp.651-655, doi:10.1109 / CVIDL51233.2020.000-7.

[0015] However, the DDPG algorithm still faces difficulties in convergence, especially when the optimal strategy is complex. The autonomous learning process is lengthy and computationally intensive. Because the maneuver strategies in these situations are complex and difficult to obtain, there is a significant amount of trial and error, resulting in a large amount of air combat experience. As the algorithm runs, a large amount of experience accumulates in the experience pool, requiring not only significant memory but also causing a lengthy experience replay process, severely impacting the algorithm's efficiency.

[0016] Based on the above analysis, the problems and shortcomings of the existing technology are as follows:

[0017] (1) Air combat maneuver decision-making is quite complex. The air combat state changes with factors such as weapons, targets, environment, aircraft itself, and target characteristics, and there are serious nonlinear phenomena. Furthermore, due to the special nature of air combat, experience is difficult to acquire. In addition, air combat maneuver decision-making requires a global and process-oriented perspective, rather than making trade-offs based on short-term gains and losses. These characteristics pose a great challenge to traditional mathematical methods, and new solutions are urgently needed.

[0018] (2) When making maneuver decisions based on the DDPG method, the algorithm requires a lot of trial and error during operation, especially when the optimal strategy is relatively complex. This results in a large amount of experience accumulating in the experience pool, which not only requires a large amount of running memory, but also causes the experience replay process to take a long time, which seriously affects the efficiency of the algorithm.

[0019] (3) The concentration suppression mechanism of the immune algorithm has a certain ability to exclude similar experience, but it also has a large degree of randomness and is sensitive to the initial value, resulting in a long convergence time of the algorithm. A better method is needed to handle similar experience in the experience pool. Summary of the Invention

[0020] To address the problems existing in the prior art, this invention provides a method, system, device, and terminal for autonomous decision-making in close-range air combat of unmanned aerial vehicles (UAVs), and particularly relates to a method, system, device, and terminal for autonomous decision-making in close-range air combat of UAVs based on node clustering and DDPG.

[0021] This invention is implemented as follows: an autonomous decision-making method for close-range air combat maneuvers of unmanned aerial vehicles (UAVs). The autonomous decision-making method for close-range air combat maneuvers of UAVs includes: proposing a new problem description model for close-range air combat, changing the state expression method according to the characteristics of close-range air combat to reflect the characteristics of close-range air combat of UAVs; proposing a node fuzzy clustering method, and processing the experience pool based on this method to exclude similar experiences and retain the most representative experiences; designing an autonomous learning and autonomous decision-making algorithm flow for close-range air combat maneuvers, and providing pseudocode for the main steps of the algorithm.

[0022] Furthermore, the autonomous decision-making method for close-range air combat maneuvers of unmanned aerial vehicles includes the following steps:

[0023] Step 1, close-range air combat problem modeling: establish expressions for close-range air combat situation, missile attack zone and missile launch conditions, construct UAV motion dynamics model and perform action representation in reinforcement learning;

[0024] Step 2, similar experience elimination: Using the ideas of grid partitioning and concentration suppression, the state space nodes are divided, and similar experiences in the experience pool are clustered and eliminated based on the nodes and fuzzy clustering methods to ensure that the retained experiences are typical;

[0025] Step 3, autonomous learning and decision-making for air combat maneuvers: Design an autonomous learning and decision-making algorithm for air combat maneuvers based on node clustering and DDPG, including the algorithm flow and pseudocode of the main steps, to realize autonomous decision-making for close-range air combat maneuvers.

[0026] Furthermore, the close-range air combat situation description in step one includes:

[0027] According to the conventions of reinforcement learning, the current close-range air combat situation is represented by a vector S, which includes all factors that affect the course of the battle, including the relative spatial positions of both sides and their motion states.

[0028] Among them, the spatial relative position variables of the two sides include distance, azimuth and pitch angles between them, azimuth angle υ1 and pitch angle μ of U2 relative to U1, and azimuth angle υ2 and pitch angle -μ of U1 relative to U2; the motion state factors and variables of the two sides include speed magnitude and track inclination angle, and OXYZ is the inertial ground coordinate system.

[0029] Adding the UAV's flight altitude H to the current state, and calculating the target aircraft's flight altitude using the pitch angle μ and relative distance d, the current state is represented as:

[0030] S=[d,υ1,υ2,μ,v1,v2,θ1,θ2,H]

[0031] Furthermore, the expressions for the missile attack zone and the missile launch permit conditions in step one include:

[0032] For a typical infrared short-range missile, the basic conditions for launch include:

[0033] (1) The attacking aircraft U1 enters a certain area around the target U2, denoted by D;

[0034] (2) The target U2 is located within a certain area around the attacking aircraft U1, denoted by D′;

[0035] (3) The distance between the two sides d∈[d max ,d min ], d min d max These are the minimum and maximum launch distances for the missile;

[0036] Among them, D, D′, d min d max The size of the missile is related to the aircraft's motion state, target parameters, missile performance, and environmental parameters, and changes dynamically.

[0037] Furthermore, the construction of the UAV motion dynamics model in step one includes:

[0038] (1) Constructing a kinematic model of a particle

[0039] A three-degree-of-freedom kinematic model of a UAV is established in a geographic coordinate system, where OXYZ is the inertial ground coordinate system, and the position of U1 in the coordinate system is represented by [x,y,z]; the angle between the velocity vector V1 and the OXY plane is represented by θ, which is the trajectory inclination angle; the angle between U1 and the OZX plane is represented by ψ, which is the trajectory deflection angle. The equations of motion for the particle are:

[0040]

[0041] (2) Constructing an aerodynamic calculation model

[0042] When the sideslip angle is 0, there is only one angle of attack, denoted by α, between the aircraft coordinate system and the velocity coordinate system. (OXYZ) p Let (OXYZ) be the aircraft coordinate system. v Let OX be the velocity coordinate system. p OX v The included angle is the angle of attack α.

[0043] When the sideslip angle is 0, the transformation matrix from the aircraft coordinate system to the velocity coordinate system simplifies to:

[0044]

[0045] In the velocity coordinate system (OXYZ) v In the middle, the lift force F on the UAV L The formula for calculating air resistance f is as follows:

[0046]

[0047] In the formula, ρ is the air density, S is the equivalent wing area of ​​the UCAV; C F (V,α), C x (V, α) are the lift and drag coefficients, respectively.

[0048] (3) Constructing a dynamic model

[0049] If the speed roll angle is γ v The trajectory deflection angle is ψ, the trajectory inclination angle is θ, and the transformation matrix from the ground coordinate system to the velocity coordinate system is:

[0050]

[0051] If F L The expressions for f in the ground coordinate system OXYZ are F′=[F′ x F′ y F′ z ] T f′=[f′ x f′ y f′ z ] T ,but:

[0052]

[0053]

[0054] When the engine mounting angle is 0, the thrust F T The expression F′ in OXYZ T for:

[0055]

[0056] The three-degree-of-freedom dynamic model is then:

[0057]

[0058] In the formula, m is the mass and g is the acceleration due to gravity.

[0059] Furthermore, the actions in step one include: when the UAV uses BTT control, during maneuvering flight, the UAV controls the angle of attack α and the roll angle γ. v Engine thrust F TThe change generates the required overload to achieve the maneuver objective; therefore, the action in reinforcement learning is represented as:

[0060] Action=[α,γ v ,F T ].

[0061] Furthermore, the similarity exclusion algorithm designed in step two, utilizing grid partitioning and concentration suppression, includes:

[0062] (1) Determine the gridding scale, and based on the scale, blur each dimension of each set of data in the data space to its nearest neighbor node; a set of data in the data space represents an experience. Assume that a set of standardized data contains u dimensions x i =[x i1 ,x i2 ,...,x iu ], for x i any dimension x iu The fuzziness scale is set using d. u If expressed as [d1, d2, ..., d], then the scale vector is d = [d1, d2, ..., d]. u ]; Scale vector d = [d1, d2, ..., d u Once set, d will no longer change.

[0063] (2) For a set of data x i u dimensions x iu Using x iu Divisible by d u equals p u The remainder is e. u ,but:

[0064] x iu =p u ·d u +e u ;

[0065] According to the following formula, x iu Perform node blurring:

[0066]

[0067] Then we get x i The node fuzzy clustering result x′ i =[x′ i1 ,x′ i2 ,...,x′ iu ]. Calculate x i 、x′ i Euclidean distance between them and r i express.

[0068] (3) with x′ j Represents another set of data x j After the node fuzzy clustering, the result of step (2) may be x i x j The fuzzy clustering results of the nodes in the two sets of data are the same, that is...

[0069] x′ i =x′ j

[0070] Then we can consider x i x j If two sets of data (based on experience) are similar, one of them can be eliminated. Typically, a set of data x... i With node x′ i The smaller the Euclidean distance, the closer the data can be considered to be to node x′. i If the data set is larger, then this set of data will better reflect the characteristics of the node data set. Therefore, the data set to be retained can be determined based on Euclidean distance. Let r... j Represents array x j With its node x′ j Euclidean distance, such as r j >r i Then keep x i Otherwise, keep x. j .

[0071] In reinforcement learning, each set of data includes the following factors:

[0072] {S t A t ,r t+1 ,S t+1}

[0073] In the experience of reinforcement learning, r t+1 S t+1 The values ​​are all determined by S t A t The decision is therefore made only for S. t A t Perform approximate node clustering; if S in the two empirical values... t A t After approximate clustering of nodes, they all tend to converge to the same node, that is:

[0074] {S 1t A 1t ,r 1,t+1 ,S 1,t+1}——>{S′ 1t ,A′ 1t ,r 1,t+1 ,S 1,t+1}

[0075] {S 2t A 2t ,r 2,t+1 ,S 2,t+1}——>{S′ 1t ,A′ 1t ,r 2,t+1 ,S 2,t+1}

[0076] Then, one set of empirical methods can be eliminated based on the minimum Euclidean distance method described above.

[0077] Furthermore, the design of the autonomous decision-making learning algorithm for air combat maneuvers in step three includes:

[0078] (1) Decision exploration stage

[0079] 1) UAV1 policy network NN(θ) π ) Receive the current state S from the environment t To determine the action 'a' to be performed in the current state. t =[n y γ F T ];

[0080] 2) UAV1 performs action a t The target aircraft, UAV2, also performed an action a2. t Change the state of the environment to S t+1 Based on the changes in state, UAV1 receives a reward r. t ;

[0081] 3) Gain one experience point {S} t ,a t ,r t ,S t+1} and store the experience in an experience database;

[0082] 4) Termination condition judgment.

[0083] During the decision exploration phase, the Actor_target network NN(θ) π- ), Critic_online network NN(θ) Q ), and Critic_target network NN (θ Q- It does not participate in the algorithm process; the judgment condition is set according to the number of episodes, and the judgment is made according to the set condition; when the set condition is met, the training and learning phase is started.

[0084] (2) Training and learning phase

[0085] 1) Based on the experience playback strategy, the experience database E DPerform experience replay to obtain N pieces of experience data for training and learning;

[0086] 2) For any replay experience {S} t ,a t ,r t ,S t+1}, UAV1 currently evaluates the network NN(θ) Q According to [S] t ,a t The current state action value Q(S) is obtained. t ,a t |θ Q );

[0087] 3) UAV1 target policy network NN(θ) π- Based on experience, S t+1 Determine the optimal action under the given conditions. Will As input, according to the target evaluation network NN(θ) Q- ) Result in state action value And obtain empirical evaluation of the network's expected value.

[0088] 4) Construct the Critic_online network NN(θ) Q The loss function L(θ) Q ):

[0089]

[0090] Based on L(θ) Q ) and error gradient backpropagation algorithm to update NN(θ) Q ), where the gradient is represented as:

[0091]

[0092] 5) Update the current policy network NN(θ) using the error gradient. π ), NN(θ) π The loss function is L(θ) π L(θ) represents the sum of squares of the differences between the network's current output value and the target value; π Regarding θ π The gradient is equivalent to Q(s,a|θ) Q For θ π The expected gradient is obtained, and the expected value is estimated unbiasedly using mini-batch training set data according to the Monte Carlo method.

[0093]

[0094] Since a=π(s|θ π Therefore, substituting into the unbiased estimation formula, we get:

[0095]

[0096] Update the parameters in the direction that the Q value increases.

[0097] 6) After a certain number of iteration intervals, the target network NN(θ) is processed. π- ), NN(θ) Q- Update;

[0098]

[0099]

[0100] in, and θ represents the target policy network parameters and the target evaluation network parameters, respectively; π and θ Q These represent the current policy network parameters and the target evaluation network parameters, respectively; τ is the learning rate, 0 < τ < 1, which represents the update speed of the target network, and the update method is called Soft Target Updates.

[0101] Another objective of this invention is to provide an autonomous decision-making system for close-range air combat maneuvers of a UAV that applies the aforementioned autonomous decision-making method for close-range air combat maneuvers. The autonomous decision-making system for close-range air combat maneuvers of a UAV includes:

[0102] The close-range air combat problem modeling module is used to describe the mathematical expressions of close-range air combat situation, attack zone and attack conditions, and to construct the UAV motion dynamics model and perform action representation in reinforcement learning.

[0103] The similar experience exclusion module is used to propose and design an exclusion algorithm based on the ideas of grid partitioning and concentration suppression to cluster and exclude similar experiences in the experience pool, retaining the most representative experience data;

[0104] The autonomous decision-making learning module for air combat maneuvers is used to design an autonomous learning algorithm for air combat maneuvers based on node clustering and DDPG, and to achieve autonomous decision-making for air combat maneuvers through decision exploration and learning training phases.

[0105] Another object of the present invention is to provide a computer device, the computer device including a memory and a processor, the memory storing a computer program, and when the computer program is executed by the processor, causing the processor to perform the steps of the aforementioned UAV close-range air combat maneuver autonomous decision-making method.

[0106] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the aforementioned UAV close-range air combat maneuver autonomous decision-making method.

[0107] Another objective of this invention is to provide an information data processing terminal for implementing the aforementioned UAV close-range air combat maneuver autonomous decision-making system.

[0108] Based on the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by this invention are as follows:

[0109] First, this invention addresses the problem that the DDPG algorithm generates a large amount of air combat experience during autonomous learning in both balanced and disadvantageous situations. This experience consumes significant storage and computational resources during replay, leading to long learning times and high computational demands. To resolve this, this invention proposes a fuzzy clustering method to process the experience pool, eliminating similar experiences and retaining only the most representative ones. Simultaneously, this invention borrows from grid clustering to propose a node clustering similarity elimination method for removing similar experiences. This invention also improves the close-range air combat problem description model, changing the state representation method to better reflect the characteristics of close-range air combat.

[0110] Second, current DDPG-based close-range air combat maneuver autonomous learning and decision-making algorithms require a large amount of storage and computing resources. This invention proposes a method based on fuzzy clustering to retain typical experiences while reducing the size of the experience pool; it proposes a node clustering method based on the purpose of the algorithm clustering, which can effectively retain typical experiences and also fit the DDPG algorithm framework; and it proposes a new form of air combat state expression that can better reflect the state factors in close-range air combat.

[0111] Third, as supplementary evidence of the inventive step of the claims of this invention, it is also reflected in the following important aspects:

[0112] Autonomous learning and decision-making in close-range air combat maneuvers is a challenging frontier research problem. This invention realizes autonomous learning and continuous decision-making of UAVs in close-range air combat under continuous states based on the Deep Deterministic Policy Gradient (DDPG) algorithm. To address the low efficiency of the DDPG algorithm under complex situations, a node clustering algorithm is proposed, which retains typical experience while effectively reducing the number of experiences in the experience pool, thus significantly improving the algorithm efficiency. The algorithm of this invention can be used to autonomously train UAV maneuver decision-making units in close-range air combat, and has significant engineering application value. Attached Figure Description

[0113] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0114] Figure 1 This is a flowchart of the autonomous decision-making method for close-range air combat maneuvers of unmanned aerial vehicles provided in an embodiment of the present invention;

[0115] Figure 2 This is a spatial relative position diagram of both sides in close-range air combat provided by an embodiment of the present invention;

[0116] Figure 3 This is a motion state diagram of the two opposing sides in close-range air combat provided in an embodiment of the present invention;

[0117] Figure 4 This is a schematic diagram illustrating the achievement of infrared short-range missile attack conditions provided in an embodiment of the present invention;

[0118] Figure 5 This is a diagram of the aircraft coordinate system and velocity coordinate system and their geometric relationship provided in an embodiment of the present invention;

[0119] Figure 6 This is a schematic diagram of the node approximate clustering method provided in an embodiment of the present invention;

[0120] Figure 7 This is a schematic diagram of the overall framework of DDPG provided in an embodiment of the present invention;

[0121] Figure 8A This is a visual representation of the empirical distribution before clustering, using [d,υ1,υ2] as coordinate values, provided by an embodiment of the present invention.

[0122] Figure 8B This is an intuitive diagram showing the empirical distribution after clustering, using [d,υ1,υ2] as coordinate values, provided by an embodiment of the present invention.

[0123] Figure 8C This is a visual representation of the empirical distribution after excluding vector moments, using [d,υ1,υ2] as coordinate values, provided by an embodiment of the present invention.

[0124] Figure 9A This is a schematic diagram of the three-dimensional distribution of TS with [υ1,υ2,α] as coordinate values, provided by an embodiment of the present invention.

[0125] Figure 9B This is a schematic diagram of the three-dimensional distribution of TS1 with [υ1,υ2,α] as coordinate values ​​provided in an embodiment of the present invention.

[0126] Figure 9C This is a schematic diagram of the three-dimensional distribution of TS2 with [υ1,υ2,α] as coordinate values ​​provided in an embodiment of the present invention;

[0127] Figure 10 This is the advantageous situation provided by the embodiments of the present invention under P Su P Fl P Oth The change curve;

[0128] Figure 11 is a diagram illustrating some typical processes of algorithm emergence under advantageous conditions provided in the embodiments of the present invention;

[0129] Figure 12 This is a schematic diagram of the maneuvering process generated by making maneuvering decisions based on an advantage under the guidance of a trained policy network, as provided in an embodiment of the present invention.

[0130] Figure 13A The U1 provided in this embodiment of the invention is based on NN(θ) π The engine thrust variation curve resulting from the decision;

[0131] Figure 13B The U1 provided in this embodiment of the invention is based on NN(θ) π The curve showing the change in angle of attack determined by the decision;

[0132] Figure 13C The U1 provided in this embodiment of the invention is based on NN(θ) π The resulting speed roll angle variation curve;

[0133] Figure 13D This is a speed change curve of U1 during the maneuvering process provided in an embodiment of the present invention;

[0134] Figure 14 This is the equilibrium state P provided in the embodiments of the present invention. Su P Fl P Oth The change curve;

[0135] Figure 15 This is a diagram illustrating the maneuvering process generated by a balanced power maneuvering decision based on a trained policy network, as provided in an embodiment of the present invention.

[0136] Figure 16A The U1 provided in this embodiment of the invention is based on NN(θ) π The curve showing the change in angle of attack determined by the decision;

[0137] Figure 16B The U1 provided in this embodiment of the invention is based on NN(θ) π The engine thrust variation curve resulting from the decision;

[0138] Figure 16C The U1 provided in this embodiment of the invention is based on NN(θ) π The resulting speed roll angle variation curve;

[0139] Figure 16D This is a schematic diagram of the normal overload change of U1 during the maneuvering process provided in an embodiment of the present invention;

[0140] Figure 16E This is a speed change curve of U1 during the maneuvering process provided in an embodiment of the present invention;

[0141] Figure 17 This is the disadvantageous situation P provided in the embodiments of the present invention. Su P Fl P Oth The change curve;

[0142] Figure 18 This is a diagram illustrating the maneuvering process generated by a trained policy network under adverse conditions, as provided in an embodiment of the present invention.

[0143] Figure 19A The U1 provided in this embodiment of the invention is based on NN(θ) π The engine thrust variation curve resulting from the decision;

[0144] Figure 19B The U1 provided in this embodiment of the invention is based on NN(θ) π The curve showing the change in angle of attack determined by the decision;

[0145] Figure 19C The U1 provided in this embodiment of the invention is based on NN(θ) π The curve showing the change in engine speed roll angle determined by the decision;

[0146] Figure 19D This is a graph showing the normal overload variation of U1 during the maneuvering process, provided in an embodiment of the present invention.

[0147] Figure 19E This is a speed change curve of U1 during the maneuvering process provided in an embodiment of the present invention. Detailed Implementation

[0148] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0149] To address the problems existing in the prior art, this invention provides a method, system, device, and terminal for autonomous decision-making in close-range air combat of unmanned aerial vehicles (UAVs). The invention will be described in detail below with reference to the accompanying drawings.

[0150] This invention addresses the issues of the long learning time and high computational cost of the DDPG algorithm's autonomous learning process for air combat maneuvers under advantageous and disadvantageous situations. It proposes a fuzzy clustering method to exclude similar experiences from the experience pool, retaining only the most representative experiences. However, since the purpose of clustering is classification, this invention borrows the idea of ​​grid partitioning to propose a node clustering method for excluding similar experiences. Simultaneously, this invention improves the problem description model for close-range air combat, proposing a new state representation method based on the characteristics of close-range air combat, which better reflects the features of close-range air combat.

[0151] like Figure 1 As shown, the UAV close-range air combat maneuver autonomous learning decision-making method provided in this embodiment of the invention includes the following steps:

[0152] S101, Close-range air combat problem modeling: Establish expressions for close-range air combat situation, missile attack zone and missile launch conditions, construct UAV motion dynamics model and perform action representation in reinforcement learning;

[0153] S102, Similar Experience Exclusion: Using the ideas of grid partitioning and concentration suppression, the state space nodes are divided, and similar experiences in the experience pool are clustered and excluded based on the nodes and fuzzy clustering methods to ensure that the retained experiences are typical.

[0154] S103, Autonomous Learning and Decision Making for Air Combat Maneuvers: Design an autonomous learning and decision-making algorithm for air combat maneuvers based on node clustering and DDPG, including the algorithm flow and pseudocode of the main steps, to achieve autonomous decision-making for close-range air combat maneuvers.

[0155] As a preferred embodiment of the present invention, the UAV close-range air combat maneuver autonomous decision-making method further includes:

[0156] 1. Modeling of close-range air combat problems

[0157] In current close-range air combat, infrared short-range missiles are the primary weapons. Therefore, this invention's analysis is based on the premise that UAVs use infrared short-range missiles in air combat. Both sides maneuver to bring the opponent into their missile's attack zone while simultaneously avoiding entering the opponent's missile's attack zone.

[0158] 1.1 Close-range air combat situation description

[0159] Accurately describing the problem state is a prerequisite for effective problem-solving. Based on existing analysis, this invention improves the state representation method according to the characteristics of close-range air combat, making it more reflective of these characteristics. Following the representational conventions of reinforcement learning, this invention uses a vector S to represent the current close-range air combat situation. S should include all factors affecting the combat process, including the relative spatial positions of both sides and their motion states.

[0160] The spatial relative positions of the two parties mainly include the following variables: relative distance, mutual azimuth and elevation angles, including the azimuth angle υ1 and elevation angle μ of U2 relative to U1, and the azimuth angle υ2 and elevation angle -μ of U1 relative to U2. The meanings of each variable are as follows: Figure 2 As shown.

[0161] Figure 2 In the diagram, d represents the relative distance between the two drones U1 and U2; υ1 represents the angle between the projection of the velocity vector V1 of drone U1 onto the horizontal plane and the projection of the target line U1U2 onto the horizontal plane; υ2 represents the angle between the projection of the velocity vector V2 of drone U2 onto the horizontal plane and the projection of the target line U1U2 onto the horizontal plane; and μ represents the angle between the target line U1U2 and the horizontal plane.

[0162] The motion states of both sides in a battle mainly include the following factors and variables: speed magnitude, trajectory angle, and the definitions of each variable are as follows: Figure 3 As shown, OXYZ is an inertial ground coordinate system.

[0163] Figure 3 In the diagram, v1 and v2 represent the speeds of U1 and U2, respectively, and are scalars; θ1 represents the trajectory tilt angle of U1; and θ2 represents the trajectory tilt angle of U2.

[0164] In fact, environmental factors also affect the aircraft's maneuver command decisions. For example, changes in air density due to changes in altitude will alter the aircraft's available overload, which in turn will limit the decision-making range. Therefore, the aircraft's altitude H should also be included in the current state. The target aircraft's altitude can be calculated using the pitch angle μ and relative distance d, so it does not need to be listed separately.

[0165] Based on the above analysis, the current state can be represented as:

[0166] S=[d,υ1,υ2,μ,v1,v2,θ1,θ2,H]

[0167] 1.2 Representation of the attack zone and fulfillment of attack conditions

[0168] like Figure 4 As shown, for a typical infrared short-range missile, the following three basic conditions must be met to constitute a launch condition:

[0169] (1) The attack aircraft (U1) enters a certain area around the target (U2), which is denoted by D;

[0170] (2) The target (U2) is located within a certain area around the attacking aircraft (U1), and this area is denoted by D′;

[0171] (3) Relative distance d∈[d min ,d max ], d min d max These are the minimum and maximum launch distances for the missile;

[0172] Among them, D, D′, d min d max The size of the missile is related to the aircraft's motion state, target parameters, missile performance, and environmental parameters, and it changes dynamically.

[0173] 1.3 Dynamics Model of Unmanned Aerial Vehicles

[0174] In aerial combat, unmanned aerial vehicles (UAVs) use maneuvering to create advantageous combat positions, evade enemy attacks, and create conditions for friendly missile launches. According to modern energy-based air combat theory, the energy loss and replenishment of UAVs during maneuvers must be considered, thus necessitating the influence of aerodynamics and engine thrust on the maneuvering process. Therefore, a point mass motion model, a dynamics model, and an aerodynamic model will be established separately.

[0175] 1.3.1 Kinematic Model of a Particle

[0176] A three-degree-of-freedom kinematic model of a UAV is established in a geographic coordinate system, but it is usually simplified to some extent:

[0177] (1) The change of the UAV centroid is not considered;

[0178] (2) Treat the ground coordinate system as an inertial coordinate system and ignore the effects of the Earth's rotation and revolution;

[0179] (3) Ignore the curvature of the Earth.

[0180] Where OXYZ is the inertial ground coordinate system, and the position of U1 in the coordinate system is represented by [x,y,z]. The angle between the velocity vector V1 and the OXY plane is represented by θ, which is the trajectory inclination angle; the angle between U1 and the OZX plane is represented by ψ, which is the trajectory deflection angle. The equation of motion of the particle can be expressed as:

[0181]

[0182] 1.3.2 Aerodynamic Calculation Model

[0183] During the maneuvering flight of an unmanned aerial vehicle (UAV), the UAV adjusts its flight attitude based on maneuver commands to control the control surfaces, generating different maneuver overloads. Conversely, maneuver overloads also affect the UAV's subsequent flight status and attitude. Therefore, aerodynamic calculations are essential to reflect this fundamental change pattern.

[0184] Meanwhile, to simplify the model, it is assumed that the UAV's sideslip path control is good, and the existence of sideslip angle is not considered. When the sideslip angle is not considered, there is only one angle of attack between the aircraft coordinate system and the velocity coordinate system, denoted by α, as follows: Figure 5 As shown. (OXYZ) p The coordinate system is (OXYZ). v Represents the velocity coordinate system, OX p OX v The included angle is the angle of attack α.

[0185] Since the sideslip angle is assumed to be 0, the transformation matrix from the aircraft coordinate system to the velocity coordinate system can be simplified to:

[0186]

[0187] In the velocity coordinate system (OXYZ) v In the middle, the lift force F on the UAV L The formula for calculating air resistance f is as follows:

[0188]

[0189] In the formula, ρ is the air density, S is the equivalent wing area of ​​the UCAV; C F (V,α), C x (V, α) are the lift and drag coefficients, respectively.

[0190] 1.3.3 Dynamic Model

[0191] If the speed roll angle is γ v The trajectory deflection angle is ψ, the trajectory inclination angle is θ, and the transformation matrix from the ground coordinate system to the velocity coordinate system is:

[0192]

[0193] If F L The expressions for f in the ground coordinate system OXYZ are F′=[F′ x F′ y F′ z ] T f′=[f′ x f′ y f′ z ] T ,but:

[0194]

[0195]

[0196] Similarly, assuming the engine mounting angle is 0, the thrust F can be obtained. TThe expression F′ in OXYZ T :

[0197]

[0198] This leads to a three-degree-of-freedom dynamic model:

[0199]

[0200] Where m is mass and g is gravitational acceleration.

[0201] 1.4 Action Representation

[0202] Assuming the UAV uses BTT (Browser-to-Trip) control, and for the sake of simplicity, the control process is simplified. It is also assumed that the UAV's control characteristics are good and that it can adjust its flight attitude within permissible ranges. During maneuvering flight, the UAV can control its angle of attack α and roll angle γ. v Engine thrust F T The change generates the required overload to achieve the maneuver objective. Therefore, the action in reinforcement learning can be represented as:

[0203] Action=[α,γ v ,F T ]

[0204] 2. Design of an autonomous learning algorithm for air combat maneuvering based on clustering methods and DDPG

[0205] 2.1 Introduction to the DDPG Algorithm

[0206] The DDPG algorithm has many natural applicability issues for air combat maneuver decision-making problems. First, it retains the trial-and-error exploration and learning characteristics of reinforcement learning algorithms, which can solve the problem of scarce and difficult-to-acquire air combat experience. Second, DDPG uses value networks and policy networks, which can search in continuous state and action spaces, and air combat decision-making problems are typical continuous decision-making problems. Third, DDPG uses neural networks to implement value functions and policy functions, which have good nonlinear mapping characteristics for complex state and action spaces.

[0207] The DDPG algorithm employs four neural networks: two identical Actor policy networks, namely the Actor_online network and the Actor_target network, each using NN(θ) π ), NN(θ) π- ) represents two identical Critic evaluation networks, namely the Critic_online network and the Critic_target network, denoted by NN(θ). Q ), NN(θ) Q-) represents. Set NN(θ) π- ), NN(θ) Q- The primary purpose of a neural network (NN) is to maintain the stability of the algorithm when generating training datasets. π ) and NN(θ Q The network is primarily used for training and is periodically updated. Q ), NN(θ) Q- ) Network parameters.

[0208] θ π Represents NN(θ) π The network parameters, which function as a deterministic policy function π(·), take the current state s as input and output the agent's action, i.e.:

[0209] Action=π(s|θ π )

[0210] Using the network NN(θ) Q ) represents the value function Q(s,Action|θ) Q ). θ Q This represents the network parameters, whose inputs are the current state s and the agent's action Action, and whose output is the state-action value. The detailed principles of the algorithm will be explained in later chapters.

[0211] However, DDPG has encountered some difficulties in its attempts to solve the air combat decision-making problem, including the problem of an exploding experience replay database. Since each air combat trial and error process generates a large amount of experience, and when the maneuver strategy is complex, the required trial and error process is even more extensive, resulting in a massive air combat experience database. While theoretically, more experience is more beneficial for network learning, storing all of this air combat experience in the experience pool would have a disastrous impact on the algorithm's operation, significantly increasing computational load and time. Furthermore, many of these air combat experiences are similar and can be discarded; only representative air combat experiences need to be retained. This necessitates a method to group similar experiences into categories, using a typical experience as the representative of that category. Therefore, this invention proposes a node clustering method.

[0212] 2.2 Clustering Method Based on Node-Fuzzy Approach

[0213] Since the primary objective of this invention is to exclude similar experiences, it only borrows the ideas of grid partitioning and density calculation, without performing clustering, but only exclusion. Based on this objective, and drawing on the idea of ​​gridding, the following exclusion algorithm is designed:

[0214] (1) Determine the gridding scale, and based on the scale, blur each dimension of each set of data in the data space to its nearest neighbor node; a set of data in the data space represents an experience. Assume that a set of standardized data contains u dimensions x i =[x i1 ,x i2 ,...,x iu ], for x i any dimension x iu The fuzziness scale is set using d. u If expressed as [d1, d2, ..., d], then the scale vector is d = [d1, d2, ..., d]. u ]; Scale vector d = [d1, d2, ..., d u Once set, d will no longer change.

[0215] (2) For a set of data x i u dimensions x iu Using x iu Divisible by d u equals p u The remainder is e. u ,but:

[0216] x iu =p u ·d u +e u ;

[0217] According to the following formula, x iu Perform node blurring:

[0218]

[0219] Then we get x i The node fuzzy clustering result x′ i =[x′ i1 ,x′ i2 ,...,x′ iu ]. Calculate x i 、x′ i Euclidean distance between them and r i express.

[0220] (3) with x′ j Represents another set of data x j After the node fuzzy clustering, the result of step (2) may be x i x j The fuzzy clustering results of the nodes in the two sets of data are the same, that is...

[0221] x′ i =x′ j

[0222] Then we can consider x i x j If two sets of data (empirical) are similar, one of them can be eliminated.

[0223] 2.3 Definition and Application of Node Membership

[0224] like Figure 6 As shown, the idea behind the node clustering algorithm can be briefly illustrated below.

[0225] The plane in the diagram represents a state space, and a series of nodes can be generated in the state space based on the scale vector [d1, d2]. [x 11 ,x 12 ]、[x 21 ,x 22 ]、[x 31 ,x 32 ] represents any 3 states. After node clustering, the result is always [x′]. 11 ,x′ 12 The graph shows that these three states are similar. But which of these three sets of data should be kept? As can be seen from the graph, the three states are similar to the node state [x′]. 11 ,x′ 12 The Euclidean distance of a state x is different from that of a given state x. i With node state x′ i The smaller the Euclidean distance, the closer the data can be to the node state x′. i If the state is true, then it better reflects the characteristics of the node state. Therefore, the state to be retained can be determined based on Euclidean distance. The node membership degree of a state is defined as:

[0226] R i,node =1 / ||[x i1 ,x i2 ]-[x i1,node ,x i2,node ]||

[0227] Among them, [x i1 ,x i2 [x] represents any state. i1,node ,x i2,node ] is [x i1 ,x i2 The node state after node clustering. If X i =[x i1 ,x i2 ], with X i ∈[x i1,node ,x i2,node ] represents X i The node state after approximate clustering is [xi1,node ,x i2,node ], then we have:

[0228] X i ←minR i,node

[0229] X∈[x i1,node ,x i2,node ]

[0230] The advantage of using this approach is that it reduces the algorithm's memory requirements, as it only needs to store the typical states involved in the database, instead of storing the useless grid parts in the data space.

[0231] In reinforcement learning, each set of data includes the following factors:

[0232] {S t A t ,r t+1 ,S t+1}

[0233] In the experience of reinforcement learning, r t+1 S t+1 The values ​​are all determined by S t A t The decision is therefore made only for S. t A t Perform approximate node clustering; if S in the two empirical values... t A t After approximate clustering of nodes, they all tend to converge to the same node, that is:

[0234] {S 1t A 1t ,r 1,t+1 ,S 1,t+1}——>{S′ 1t ,A′ 1t ,r 1,t+1 ,S 1,t+1}

[0235] {S 2t A 2t ,r 2,t+1 ,S 2,t+1}——>{S′ 1t ,A′ 1t ,r 2,t+1 ,S 2,t+1}

[0236] Then, one set of empirical methods can be eliminated based on the minimum Euclidean distance method described above.

[0237] 3. Design of an autonomous learning algorithm for air combat maneuvering based on node clustering and DDPG

[0238] 3.1 Algorithm Flow and Principle

[0239] DDPG overall framework as follows Figure 7 As shown, the algorithm can be divided into two phases: Phase 1 – Decision Exploration Phase and Phase 2 – Learning and Training Phase. Its basic process is as follows:

[0240] During the decision-making exploration phase:

[0241] 1) UAV1 policy network NN(θ) π ) Receive the current state S from the environment t To determine the action 'a' to be performed in the current state. t =[n y γ F T ];

[0242] 2) UAV1 performs action a t The target aircraft, UAV2, also performed an action a2. t Change the state of the environment to S t+1 Based on the changes in state, UAV1 receives a reward r. t ;

[0243] 3) Gain one experience point {S} t ,a t ,r t ,S t+1} and store the experience in an experience database;

[0244] 4) Termination condition judgment.

[0245] It is evident that in the Stage I decision exploration phase, the Actor_target network NN(θ) π- ), Critic_online network NN(θ) Q ), and Critic_target network NN (θ Q- It does not participate in the algorithm process;

[0246] In the DDPG algorithm, Phase II is not performed in every episode. Instead, it is triggered based on predefined conditions. The Phase II algorithm only starts when these conditions are met. Generally, the condition can be set based on the number of episodes. During the Phase II training phase:

[0247] (1) First, the experience database E is analyzed according to the experience playback strategy. D Perform experience replay to obtain N pieces of experience data for training and learning;

[0248] (2) For any replay experience {S} t ,at ,r t ,S t+1}, UAV1 currently evaluates the network NN(θ) Q According to [S] t ,a t The current state action value Q(S) is obtained. t ,a t |θ Q );

[0249] (3) UAV1 target policy network NN(θ) π- Based on experience, S t+1 Determine the optimal action under the given conditions. Will As input, according to the target evaluation network NN(θ) Q- ) Result in state action value And obtain empirical evaluation of the network's expected value.

[0250] (4) Construct the Critic_online network NN(θ) Q The loss function L(θ) Q ):

[0251]

[0252] Based on L(θ) Q ) and error gradient backpropagation algorithm to update NN(θ) Q ), where the gradient is represented as:

[0253]

[0254] (5) Similarly, update the current policy network NN(θ) using the error gradient. π ), NN(θ) π The loss function is L(θ) π L(θ) represents the sum of squares of the differences between the network's current output value and the target value; π Regarding θ π The gradient is equivalent to Q(s,a|θ) Q For θ π The expected gradient is obtained, and the expected value is estimated unbiasedly using mini-batch training set data according to the Monte Carlo method.

[0255]

[0256] Since a=π(s|θ π Therefore, substituting into the unbiased estimation formula, we get:

[0257]

[0258] Update the parameters in the direction that the Q value increases.

[0259] (6) After a certain number of iteration intervals, the target network NN(θ) is processed. π- ), NN(θ) Q- Update;

[0260]

[0261]

[0262] in, and θ represents the target policy network parameters and the target evaluation network parameters, respectively; π and θ Q These represent the current policy network parameters and the target evaluation network parameters, respectively. τ is the learning rate, generally 0 < τ < 1, representing the update speed of the target network. This update method is also called Soft Target Updates, which can make the algorithm run more stably.

[0263] 3.2 Algorithm Pseudocode

[0264] 3.2.1 Main program pseudocode (see Table 1)

[0265] Table 1. Main Program Pseudocode

[0266]

[0267]

[0268] 3.2.2 Node Clustering - Similarity Exclusion Pseudocode (see Table 2)

[0269] For database E D E D1 The basic idea behind clustering and similarity exclusion is as follows:

[0270] (1) Perform node clustering on the temporary database D1, save the similarity R of each experience, and exclude similar experiences within D1 based on the similarity.

[0271] (2) Compare the experiences in D1 and D pairwise, and exclude the experiences in D1 based on similarity;

[0272] (3) Merge D1 after similarity exclusion with D to obtain a new experience replay library D.

[0273] The basic methods and steps for node clustering have been described in section 3.5.3 and will not be repeated here. The pseudocode for similarity exclusion is designed as follows:

[0274] Table 2 Node Clustering - Similarity Exclusion Pseudocode

[0275]

[0276] 3.2.3 Pseudocode of the Network Learning Algorithm in the DDPG Algorithm

[0277] The pseudocode for the network learning algorithm in the DDPG algorithm is shown in Table 3.

[0278] Table 3. Pseudocode of network learning algorithms based on DDPG principle

[0279]

[0280]

[0281] 3.3 Improved Reward Function Design

[0282] Air combat decision-making is constrained by various factors such as aircraft performance, weapon attack conditions, and environment. It is a complex nonlinear decision-making problem. Since the mechanisms of many of these factors are not yet fully understood, this invention adopts a model-free approach to design reward functions.

[0283] The current air combat situation is divided into three undisputed scenarios: attack conditions met, attack conditions not met and the aircraft is not attacked, and the aircraft is attacked by the opponent. The influence of flight control and the environment is also represented by reward functions, namely, decision actions are infeasible, and the situation exceeds the environmental constraints. The reward functions for the five scenarios are as follows:

[0284] (1) When the attack conditions are met, the unit should receive a large positive reward, and according to the energy air combat theory, the higher the unit's survival rate, the better. Therefore, the reward function is defined as follows:

[0285] r t =100·(v1 / v2)

[0286] Where v1 and v2 are the speeds of the local machine and the target machine, respectively, in m / s.

[0287] (2) When the attack conditions are not met and no attack is launched, it means that the action has not achieved a valid result and should receive a negative reward value. The reward function is defined as follows:

[0288] r t =-1

[0289] (3) When the attack conditions are not met and the target is attacked, the target should receive a larger negative reward (penalty). The higher the target's survival speed, the stronger its ability to evade missile attacks, and the negative reward should decrease accordingly. The higher the target's survival speed, the more difficult it is to evade missile attacks, so the negative reward increases accordingly. The reward function is defined as follows:

[0290] r t = -50·(v2 / v1)

[0291] (4) When the UAV is unable to execute the decision command; if the decision action exceeds the current capability of the aircraft, the decision result is incorrect and a larger negative reward should be given. The reward function is defined as follows:

[0292] r t =-20

[0293] (5) When the current state exceeds the limit; this state is not currently of concern and should be given a negative reward value to avoid this state, therefore it is defined as follows:

[0294] r t =-10

[0295] The UAV close-range air combat maneuver autonomous decision-making system provided in this embodiment of the invention includes:

[0296] The close-range air combat problem modeling module is used to describe the close-range air combat situation, attack zone and attack achievement conditions, construct the UAV motion dynamics model and perform action representation in reinforcement learning;

[0297] The similar experience exclusion module is used to design node clustering and similarity exclusion algorithms using the ideas of grid partitioning and concentration suppression calculation. It clusters and excludes similar experiences in the experience pool to improve algorithm efficiency.

[0298] The autonomous decision-making learning module for air combat maneuvers is used to design an autonomous learning algorithm for air combat maneuvers based on node clustering and DDPG, and to realize autonomous learning decision-making for air combat maneuvers through decision exploration and learning training phases.

[0299] The embodiments of the present invention have achieved some positive results during the research and development or use process, and have indeed great advantages compared with the prior art. The following content describes them in conjunction with the data, charts and other information of the experimental process.

[0300] Due to computational limitations, this embodiment reduces the simulation difficulty by limiting both the target and the aircraft to maneuvering within a single plane, without vertical maneuvers, and at a flight altitude of 2000m. It also assumes the target's velocity remains constant and it undergoes circular motion, with a target velocity of 220m / s and the aircraft's initial velocity also being 220m / s.

[0301] Considering the complexity of the missile attack zone and target radiation characteristics, a simplified missile launch zone is adopted in this embodiment. D and D′ represent two fixed cones representing the missile detection performance and aircraft radiation characteristics, respectively. In this embodiment, the apex angle of D is set to 160°, and the apex angle of D′ is set to 100°. Combined with... Figure 4 It can be seen that launch condition 1 can be described as: the angle between the target line U1U2 and the velocity vector V2 is greater than 120°, and launch condition 2 can be described as: the angle between the target line U1U2 and the velocity vector V1 is less than or equal to 50°. Let d... min d max The values ​​are fixed at 500m and 6km respectively.

[0302] 1. Validation of the clustering algorithm

[0303] To observe the effectiveness of the node clustering algorithm, we take s t The three variable values ​​are used as three coordinate values, and a point is generated in the coordinate system based on this set of coordinates. This method can intuitively present the distribution of experiences in the experience database. Furthermore, for ease of comparison, a visual representation of the experience database before and after clustering is provided.

[0304] First, [d,υ1,υ2] are selected as coordinate values ​​for 3D representation. TS is a temporary experience base generated after 500 air combats by the algorithm, containing 40,189 experiences. The experience base obtained after clustering nodes in TS is denoted as TS1, with the total number reduced to 14,081, taking 3.502 seconds. The experience base obtained after vector moment similarity elimination in TS is denoted as TS2, with the total number reduced to 12,744, taking 40.716 seconds. The time taken by the vector moment method is much longer than that of the method of this invention. This is because vector moment similarity elimination requires calculating the vector moments between each pair of nodes, which is computationally intensive. In contrast, in the method of this invention, each experience only needs to calculate the Euclidean distance with the corresponding node once and then compare them, thus greatly reducing the computational load.

[0305] After normalizing [d,υ1,υ2] of each experience according to a unified rule, the three-dimensional distributions of TS, TS1, and TS2 are as follows: Figures 8A to 8C As shown.

[0306] exist Figures 8A to 8C In the diagram, each point represents a piece of experience. Figure 8A This provides a visual representation of the database experience prior to clustering. Figure 8B This is a visual demonstration of the database experience after node clustering according to the present invention. Figure 8CThis is a visual representation of the empirical distribution after similarity exclusion using the vector moments method. A direct comparison of the empirical distributions shows that after node clustering or vector moments similarity exclusion, the empirical data in densely populated regions is significantly reduced, such as the Zs2 region. Figure 8B , Figure 8C The color of the medium region becomes significantly lighter, indicating a decrease in the density of experience; at the same time, the number of experiences in sparse regions, such as region Zs1, does not decrease significantly, indicating that the algorithm retains the experience in sparse regions. This proves that both methods can effectively eliminate similar experiences.

[0307] contrast Figure 8B , Figure 8C From the Zs3 and Zs4 regions, it can be seen that... Figure 8B The graphs in the original database are closer to the distribution shape of the original database and are more uniformly distributed. They are better at preserving sparse samples. Therefore, it can be considered that the similarity exclusion method of the present invention is better able to maintain the diversity of experience, which is beneficial to subsequent network learning and helps to improve the robustness of the network.

[0308] To verify the generality of the conclusion, [υ1,υ2,α] was again selected as the coordinate values ​​for a three-dimensional representation. After normalizing [υ1,υ2,α] of each experience according to a unified rule, the three-dimensional distributions of TS, TS1, and TS2 are as follows: Figures 9A to 9C As shown.

[0309] It can be seen that, overall, Figure 9B The empirical distribution map, which is closer to the original data, can be considered to better preserve the distribution characteristics of the original database; comparing the Zn1 regions of the three maps, it can be seen that... Figure 9B and Figure 9A Almost identical, and Figure 9C and Figure 9B Compared to the previous method, this method is much sparser, indicating that some experience has been excluded from this region. Therefore, the node clustering method in this invention has a better ability to preserve sparse experience, better maintains the diversity of experience, and is more conducive to the subsequent learning and training of the network.

[0310] 2. Verification of Algorithm Effectiveness

[0311] To verify the algorithm's autonomous learning and decision-making capabilities, different initial conditions were set for verification, including advantageous and disadvantageous situations.

[0312] 2.1 Verification of Algorithm Effectiveness under Advantageous Situation

[0313] In the OXYZ coordinate system, set the heading angle of U1 to 60°, the heading angle of U2 to -20°, the initial coordinates of U1 to (0, 0), and the initial coordinates of U2 to (6000, 0). The initial distance between the two aircraft is 6km. At this time, the azimuth angle of the target relative to the aircraft is υ1 = 60°, and the azimuth angle of the aircraft relative to the target is υ2 = 160°. Set the maximum number of air combat rounds EP = 5000, and the clustering exclusion frequency f. c =100 episodes, network learning frequency f T =200, network soft update frequency f u =400.

[0314] Within a certain period, the algorithm performed N... Cy In a round-based air combat scenario, the number of rounds to win is N. Su The winning ratio is defined as:

[0315] p Su =N Su / N Cy

[0316] Similarly, if the number of rounds of failure is N Fl The number of rounds for other situations (such as loss of control, exceeding boundaries, etc.) is N. Oth The failure rate and the rates for other scenarios can be defined as follows:

[0317] p Fl =N Fl / N Cy

[0318] p Oth =N Oth / N Cy

[0319] The algorithm spent approximately 37 minutes during the 5000 rounds of exploration and learning. Su p Fl p Oth The change curve is as follows Figure 10 As shown in the figure, in the early stages of operation, the algorithm initially avoids other situations through exploration. However, since the correct strategy has not yet been found, the probability of failure increases. As the learning process progresses, once the algorithm finds the optimal strategy, the success rate gradually increases, while the failure rate and other ratios decrease rapidly.

[0320] During the learning process, the algorithm continuously explores strategies, and typical emerging processes include... Figures 11A to 11E As shown, the changes in the machine's maneuvering process indicate that, overall, the machine conducts more exploration and experimentation in the early stages of algorithm operation, resulting in many typical processes, and gradually finds the best winning strategy in the later stages.

[0321] Based on the trained policy network NN(θ) π Under the same initial conditions, the maneuver decisions were made, and the results are as follows, where the maneuver processes of U1 and U2 are as follows: Figure 12 As shown, it can be seen that U1 achieved the conditions for attacking U2 by performing a left turn maneuver with high overload.

[0322] Figures 13A-13C In this process, U1 is based on NN(θ) π The curves showing the changes in engine thrust, angle of attack, and roll angle determined by the decision. Figure 13D This is the speed change curve during the maneuver.

[0323] Simulation results show that, based on NN(θ) π The drone made correct decisions regarding its actions, ensuring both flight control of the drone itself and meeting the conditions for attack by its own missiles.

[0324] 2.2 Verification of Algorithm Effectiveness under Balanced Potential

[0325] In the OXYZ coordinate system, with U1 heading angle set to 90° and U2 heading angle set to -90°, this situation can be considered a disadvantageous situation in close-range air combat. During the algorithm's exploration and learning process, p... Su p Fl p Oth The change curve is as follows Figure 14 As shown.

[0326] After 150,000 runs, the algorithm learned the correct response strategy, increasing its win rate to 100%. This process took approximately 131 hours. It is evident that the exploration and learning of air combat strategies in a balanced situation is a relatively long process, and the computational time and workload consumed by the algorithm far exceed those in a dominant situation. This reflects the complexity of maneuver strategies in a balanced situation.

[0327] from Figure 15 It can be clearly seen that the agent first made a right turn to avoid the opponent's missile attack zone. Then, when the target was about to expose its tail, it made a sharp turn around to ensure that the agent did not lose control while bringing the target into the launch range of its own missile, thus achieving the launch conditions of its own missile. This reflects the algorithm's ability to learn complex strategies under various constraints.

[0328] Figures 16A-16C In this process, U1 is based on NN(θ) π The curves showing the changes in angle of attack, engine thrust, and speed roll angle determined by the decision. Figure 16D The curve shows the change in normal overload during the U1 maneuver. Figure 16EThis is the curve showing the change in velocity of U1 during the maneuver.

[0329] 2.3 Disadvantages

[0330] In the OXYZ coordinate system, with U1 heading angle set to 90° and U2 heading angle to -130°, this situation can be considered a disadvantageous one in close-range air combat. During the algorithm's exploration and learning process, p... Su p Fl p Oth The change curve is as follows Figure 17 As shown.

[0331] After approximately 110,000 runs, the algorithm learned the correct response strategy, achieving a 100% win rate. The 120,000 learning runs took about 116 hours. This demonstrates that the exploration and learning of air combat strategies under disadvantageous circumstances also involves a relatively long process, reflecting the complexity of maneuver decision-making problems in such situations.

[0332] However, compared with the original algorithm (V_DDPG) that uses vector moments for similarity exclusion, the algorithm of this invention shows advantages under both equal and disadvantageous conditions. See Table 4 for a detailed comparison.

[0333] As can be seen from the comparison in Table 4, regardless of whether the situation is balanced or disadvantageous, the algorithm of this invention takes less time and fewer episodes than existing algorithms, which can be considered to have higher autonomous learning efficiency.

[0334] Table 4 Algorithm Comparison Results

[0335]

[0336] Based on the trained NN(θ) π Make maneuver decisions under this situation. Figure 18 This provides a visual demonstration of the maneuvering process. The diagram clearly shows that U1 continuously made left turns, maintaining its position in the forward hemisphere of the target to prevent the opponent from launching. Simultaneously, it maintained a distance of at least 2000 meters from the target (the simulation is programmed to terminate when the distance is less than 2000 meters). When U1 entered the rear hemisphere of the target, it performed a high-G right turn, pointing its nose at the target. Since the distance was now relatively short, the missile's launch conditions were met when the target entered the missile's angular range. Throughout the entire process, U1 maintained control and all parameters remained within acceptable limits, further demonstrating the algorithm's ability to learn complex strategies under various constraints.

[0337] Figures 19A-19C In this process, U1 is based on NN(θ) πThe curves showing the changes in engine thrust, angle of attack, and roll angle determined by the decision. Figure 19D This is a graph showing the change in normal overload during the maneuver. Figure 19E This is the speed change curve during the maneuver.

[0338] It should be noted that embodiments of the present invention can be implemented in hardware, software, or a combination of both. The hardware portion can be implemented using dedicated logic; the software portion can be stored in memory and executed by a suitable instruction execution system, such as a microprocessor or dedicated-design hardware. Those skilled in the art will understand that the above-described devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, on a carrier medium such as a disk, CD, or DVD-ROM, a programmable memory such as read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuitry such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field-programmable gate arrays, programmable logic devices, etc., or by software executed by various types of processors, or by a combination of the above-described hardware circuitry and software, such as firmware.

[0339] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for autonomous decision-making in close-range air combat maneuvers for unmanned aerial vehicles, characterized in that: include: Establish expressions for close-range air combat situation, missile attack zone, and missile launch permission conditions; construct a UAV motion dynamics model and perform action representation in reinforcement learning. Similar experiences in the experience pool are clustered and excluded based on node clustering and fuzzy clustering methods to ensure that the retained experiences are typical; an autonomous learning and decision-making algorithm for air combat maneuvers based on node clustering and DDPG is designed to realize autonomous decision-making for close-range air combat maneuvers. The autonomous decision-making method for close-range air combat maneuvers of unmanned aerial vehicles includes the following steps: Step 1, close-range air combat problem modeling: establish expressions for close-range air combat situation, missile attack zone and missile launch conditions, construct UAV motion dynamics model and perform action representation in reinforcement learning; Step 2, similar experience elimination: Using the ideas of grid partitioning and concentration suppression, the state space nodes are divided, and similar experiences in the experience pool are clustered and eliminated based on the nodes and fuzzy clustering methods to ensure that the retained experiences are typical; Step 3, autonomous learning and decision-making for air combat maneuvers: Design an autonomous learning and decision-making algorithm for air combat maneuvers based on node clustering and DDPG, including the algorithm flow and pseudocode of the main steps, to realize autonomous decision-making for close-range air combat maneuvers; The construction of the UAV motion dynamics model in step one includes: (1) Constructing a kinematic model of a particle A three-degree-of-freedom kinematic model of a UAV is established in a geographic coordinate system, where OXYZ is the inertial ground coordinate system, and the position of U1 in the coordinate system is represented by [x,y,z]; the angle between the velocity vector V1 and the OXY plane is represented by θ, which is the trajectory inclination angle; the angle between U1 and the OZX plane is represented by ψ, which is the trajectory deflection angle. The equations of motion for the particle are: (2) Constructing an aerodynamic calculation model When the sideslip angle is 0, there is only one angle of attack between the aircraft coordinate system and the velocity coordinate system, denoted by α; (OXYZ) p Let (OXYZ) be the aircraft coordinate system. v Let OX be the velocity coordinate system. p OX v The included angle is the angle of attack α; When the sideslip angle is 0, the transformation matrix from the aircraft coordinate system to the velocity coordinate system simplifies to: In the velocity coordinate system (OXYZ) v In the middle, the lift force F on the UAV L The formula for calculating air resistance f is as follows: In the formula, ρ is the air density, S is the equivalent wing area of ​​the UCAV; C F (V,α), C x (V, α) are the lift and drag coefficients, respectively; (3) Constructing a dynamic model If the roll angle is γ, the deflection angle is ψ, and the inclination angle is θ, the transformation matrix from the ground coordinate system to the velocity coordinate system is: If F L The expressions for f in the ground coordinate system OXYZ are respectively F′=[F′ x F′ y F′ z ] T f′=[f′ x f′ y f′ z ] T ,but: When the engine mounting angle is 0, the thrust F T The expression F′ in OXYZ T for: The three-degree-of-freedom dynamic model is then: In the formula, m is the mass and g is the acceleration due to gravity; The actions in step one include: when the UAV uses BTT control, during maneuvering flight, the UAV controls the angle of attack α, speed roll angle γ, and engine thrust F. T The change generates the required overload to achieve the maneuver objective; therefore, the action in reinforcement learning is represented as: Action=[α,γ,F T ]; Step two involves designing a similar node clustering exclusion algorithm using grid partitioning and concentration suppression, including: (1) Determine the gridding scale and fuzzify each set of data in the data space according to the scale; when any set of standardized data contains u dimensions x i =[x i1 ,x i2 ,...,x iu ], for x i any dimension x iu The fuzziness scale is set using d. u If expressed as [d1, d2, ..., d], then the scale vector is d = [d1, d2, ..., d]. u ]; Scale vector d = [d1, d2, ..., d u Once set, for any set of data x in the data space i d no longer changes; (2) For any set of data x i u dimensions x iu Using x iu Divisible by d u equals p u The remainder is e. u ,but: x iu =p u ·d u +e u ; According to the following formula, x iu Blur: Then we get x i The node fuzzy clustering result x′ i =[x′ i1 ,x′ i2 ,...,x′ iu ]; Calculate x i 、x′ i Euclidean distance between them and r i express; (3) with x′ j Represents another set of data x j After the node fuzzy clustering, the result of step (2) may be x i x j The fuzzy clustering results of the nodes in the two sets of data are the same, that is... x′ i =x′ j Then we can consider x i x j If two sets of data are similar, one of them can be eliminated; typically, a set of data x i With node x′ i The smaller the Euclidean distance, the closer the data can be considered to be to node x′. i If the data set is accurate, then this set of data will better reflect the characteristics of the node data set. Therefore, the data set to be retained can be determined based on the Euclidean distance; with r j Represents array x j With its node x′ j Euclidean distance, such as r j >r i Then keep x i Otherwise, keep x. j ; Step two, which excludes similar experiences, also includes: Define a typical state set, and for any given state, cluster the node states according to their membership degree; In reinforcement learning, each set of data includes the following factors: {S t ,A t ,r t+1 ,S t+1 }; In the experience of reinforcement learning, r t+1 S t+1 The values ​​are all determined by S t A t The decision is therefore made only for S. t A t Perform approximate clustering of nodes; if S in the two empirical values... t A t After approximate clustering of nodes, they all tend to converge to the same node, that is: {S 1t ,A 1t ,r 1,t+1 ,S 1,t+1 }——>{S′ 1t ,A′ 1t ,r 1,t+1 ,S 1,t+1 }; {S 2t ,A 2t ,r 2,t+1 ,S 2,t+1 }——>{S′ 1t ,A′ 1t ,r 2,t+1 ,S 2,t+1 }; Then, one set of empirical methods can be eliminated based on the minimum Euclidean distance method described above.

2. The autonomous decision-making method for close-range air combat maneuvers of unmanned aerial vehicles as described in claim 1, characterized in that, The close-range air combat situation description in Step One includes: According to the conventions of reinforcement learning, the current close-range air combat situation is represented by vector S, which includes all factors that affect the course of the battle, including the relative spatial positions of both sides and their motion states. Among them, the spatial relative position variables of the two sides include distance, azimuth and pitch angles between them, azimuth angle υ1 and pitch angle μ of U2 relative to U1, and azimuth angle υ2 and pitch angle -μ of U1 relative to U2; the motion state factors and variables of the two sides include speed magnitude and track inclination angle, and OXYZ is the inertial ground coordinate system; Adding the UAV's flight altitude H to the current state, and calculating the target aircraft's flight altitude using the pitch angle μ and relative distance d, the current state is represented as: S=[d,υ1,υ2,μ,v1,v2,θ1,θ2,H] The expressions for the missile attack zone and the missile launch permit conditions in Step One include: For a typical infrared short-range missile, the basic conditions for launch include: (1) The attacking aircraft U1 enters a certain area around the target U2, denoted by D; (2) The target U2 is located within a certain area around the attacking aircraft U1, denoted by D′; (3) The distance between the two sides d∈[d max ,d min ], d min d max These are the minimum and maximum launch distances for the missile; Among them, D, D′, d min d max The size of the missile is related to the aircraft's motion state, target parameters, missile performance, and environmental parameters, and changes dynamically.

3. The autonomous decision-making method for close-range air combat maneuvers of unmanned aerial vehicles as described in claim 1, characterized in that, The autonomous decision-making learning algorithm for air combat maneuvers in step three includes: (1) Decision exploration stage 1) UAV1 policy network NN(θ) π ) Receive the current state S from the environment t To determine the action 'a' to be performed in the current state. t =[n y γ F T ]; 2) UAV1 performs action a t The target aircraft, UAV2, also performed an action a2. t Change the state of the environment to S t+1 Based on the changes in state, UAV1 receives a reward r. t ; 3) Gain one experience point {S} t ,a t ,r t ,S t+1 } and store the experience in an experience database; 4) Termination condition judgment; During the decision exploration phase, the Actor_target network NN(θ) π- ), Critic_online network NN(θ) Q ), and Critic_target network NN (θ) Q- It does not participate in the algorithm process; the judgment condition is set according to the number of episodes, and the judgment is made according to the set condition; when the set condition is met, the training and learning phase is started. (2) Training and learning phase 1) Based on the experience playback strategy, the experience database E D Perform experience replay to obtain N pieces of experience data for training and learning; 2) For any replay experience {S} t ,a t ,r t ,S t+1 }, UAV1 currently evaluates the network NN(θ) Q According to [S] t ,a t The current state action value Q(S) is obtained. t ,a t |θ Q ); 3) UAV1 target policy network NN(θ) π- Based on experience, S t+1 Determine the optimal action under the given conditions. Will As input, according to the target evaluation network NN(θ) Q- ) Result in state action value And obtain empirical evaluation of the network's expected value. 4) Construct the Critic_online network NN(θ) Q The loss function L(θ) Q ): Based on L(θ) Q ) and error gradient backpropagation algorithm to update NN(θ) Q ), where the gradient is represented as: 5) Update the current policy network NN(θ) using the error gradient. π ), NN(θ) π The loss function is L(θ) π L(θ) represents the sum of squares of the differences between the network's current output value and the target value; π Regarding θ π The gradient is equivalent to Q(s,a|θ) Q For θ π The expected gradient is obtained, and the expected value is estimated unbiasedly using mini-batch training set data according to the Monte Carlo method. Since a=π(s|θ π Therefore, substituting into the unbiased estimation formula, we get: Update the parameters in the direction that the Q value increases; 6) After a certain number of iteration intervals, the target network NN(θ) is processed. π- ), NN(θ) Q- Update; in, and θ represents the target policy network parameters and the target evaluation network parameters, respectively; π and θ Q These represent the current policy network parameters and the target evaluation network parameters, respectively; τ is the learning rate, 0 < τ < 1, which represents the update speed of the target network, and the update method is called Soft Target Updates.

4. A UAV close-range air combat maneuver autonomous decision-making system applying the UAV close-range air combat maneuver autonomous decision-making method as described in any one of claims 1 to 3, characterized in that, The autonomous decision-making system for close-range air combat maneuvering of unmanned aerial vehicles includes: The close-range air combat problem modeling module is used to describe the mathematical expressions of close-range air combat situation, attack zone and attack conditions, and to construct the UAV motion dynamics model and perform action representation in reinforcement learning. The similar experience exclusion module is used to propose and design an exclusion algorithm based on the ideas of grid partitioning and concentration suppression to cluster and exclude similar experiences in the experience pool, retaining the most representative experience data; The autonomous decision-making learning module for air combat maneuvers is used to design an autonomous learning algorithm for air combat maneuvers based on node clustering and DDPG, and to achieve autonomous decision-making for air combat maneuvers through decision exploration and learning training phases.

5. A computer device, characterized in that, The computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, it causes the processor to perform the steps of the autonomous decision-making method for close-range air combat maneuvers of the UAV as described in any one of claims 1 to 3.

6. A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the autonomous decision-making method for close-range air combat maneuvers of an unmanned aerial vehicle as described in any one of claims 1 to 3.

7. An information data processing terminal, characterized in that, The information data processing terminal is used to implement the UAV close-range air combat maneuver autonomous decision-making system as described in claim 4.