2-to-2 intelligent air combat real-time maneuvering decision-making method
By adopting a target allocation-situation advantage-nonlinear model predictive control architecture, the 2v2 air combat problem is reduced to a parallel 1v1 optimal control subproblem, which solves the problems of real-time performance, safety and portability in 2v2 air combat, and achieves efficient maneuver decision-making and safety assurance.
Patent Information
- Application Number
- CN202511826856.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2045-12-05
AI Technical Summary
Existing technologies struggle to balance real-time performance, security, and portability in 2v2 beyond-visual-range air combat, especially in highly dynamic, strongly coupled, and resource-constrained scenarios. There is a lack of a unified dynamics-tactical coupling model, and existing decision-making methods are insufficient in terms of real-time performance and scalability.
The target assignment-situation advantage-nonlinear model predictive control (NMPC) architecture is adopted to reduce the 2v2 air combat problem to two parallel 1v1 optimal control subproblems. By using an adaptive air combat advantage function and a real-time target assignment strategy, and by utilizing control parameterization and constraint transcription techniques, the real-time performance and safety of maneuver decisions are ensured.
It significantly improves the real-time performance and solution efficiency of online decision-making, has good versatility and portability, ensures flight safety through inherent hard constraints, effectively solves the dimensionality curse problem in multi-aircraft collaboration, and enhances tactical adaptability to dynamic battlefield environments.
Smart Images

Figure CN121613744A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of autonomous control and decision-making technology for unmanned combat aerial vehicles (UCAVs), specifically involving a real-time decision-making method for 2v2 beyond-visual-range air combat of UAVs that combines target allocation strategy, adaptive situation assessment and nonlinear model predictive control (NMPC). Background Technology
[0002] Modern air combat has evolved from traditional close-range dogfights to beyond-visual-range (BVR) engagements, where medium- and long-range missiles are the primary means of attack. As a core subset of BVR air combat, medium-range 2v2 air combat not only requires individual aircraft to possess the ability to detect and fire first, but also necessitates integrated solutions for cooperative target allocation, real-time maneuver decision-making, and inter-aircraft coupling constraints. Existing research primarily follows three technical routes:
[0003] (1) Decision tree method: Relying on pre-encoded tactical rules or expert knowledge base, maneuver selection is achieved through IF-THEN logic. This type of method is robust under the conditions of fixed mission scenarios and controllable environmental disturbances, but it has the following inherent defects: Poor portability: When the performance parameters of the aircraft platform and radar / missile change, the rule base needs to be reconstructed manually, and the maintenance cost is extremely high; Multi-aircraft expansion bottleneck: In 2v2 air combat, the enemy and friendly situations are highly coupled, and the rule combination explosion leads to insufficient completeness, making it difficult to cover complex interaction scenarios.
[0004] (2) Differential game and game theory model: Air combat is abstracted into a two-person zero-sum differential game, or incomplete information is handled through matrix game. Its advantage is that it can give analytical or semi-analytical optimal strategy boundaries. However, the modeling assumptions are strict: the constraints of airframe dynamics, radar / missile envelope nonlinearity and environmental disturbances are generally ignored, and it is only applicable to idealized pursuit-escape simplified scenarios; poor scalability: when it is expanded from 1v1 to 2v2, the state-action space expands exponentially, and the real-time performance of solving the HJI equation or Nash equilibrium deteriorates sharply.
[0005] (3) Data-driven reinforcement learning: In recent years, deep multi-agent reinforcement learning (MARL) has been introduced into UCAV decision-making. By acquiring policies through large-scale simulation interaction, it can get rid of the dependence on expert rules to a certain extent. However, there are the following key bottlenecks: huge data and training overhead: typical MARL algorithms require tens of millions of simulation rounds to converge. When the aircraft model, radar / missile parameters or battlefield environment drift, retraining is required, resulting in high migration costs; insufficient security and interpretability: the network policy lacks explicit handling of hard constraints such as overload and roll angle change rate, and is prone to violating the flight envelope at key boundaries; weak generalization of adversary policies: the difference between training and testing distributions leads to a significant decrease in policy performance when facing unknown adversaries or new tactics.
[0006] The technical problems with the above methods are: (1) Most existing studies focus on 1v1 or single-platform decision-making and lack a unified dynamics-tactical coupling model for 2v2 beyond-visual-range air combat; (2) Differential game, evolutionary algorithm and deep reinforcement learning methods have poor real-time performance in 2v2 scenarios; (3) Existing rule or learning-based decision-making generally treats overload, roll angle and its rate of change constraints as ex-post verification items rather than endogenous constraints of optimization variables.
[0007] In summary, existing technologies still struggle to achieve a balance between "real-time performance, security, and portability," especially in the highly dynamic, tightly coupled, and resource-constrained typical scenario of 2v2 medium-range air combat. There is an urgent need for a new decision-making framework that can integrate model prediction, target allocation, and constraint protection. Summary of the Invention
[0008] This invention constructs a real-time decision-making method for UAVs suitable for 2v2 beyond-visual-range air combat through a "target allocation-situation advantage-constraint transcription-NMPC" architecture, reducing the complex 2v2 game to two parallel 1v1 optimal control subproblems. Within the framework of nonlinear model predictive control (NMPC), control parameterization and constraint transcription techniques are used to transform the maneuverability limit into a low-complexity nonlinear programming problem that can be solved in real time. Through an adaptive air combat advantage function and a real-time target allocation strategy, it ensures the continuous output of maneuver decision commands that satisfy the aircraft's constraints, are trackable, and approach the global optimum in a dynamic battlefield environment.
[0009] To solve the above problems, the technical solution adopted by the present invention is as follows:
[0010] A real-time maneuver decision-making method for 2-on-2 intelligent air combat includes the following steps:
[0011] Step S1: Establish a three-degree-of-freedom kinematic model of a UAV mass, and use the kinematic model as the NMPC prediction model;
[0012] Step S2: Obtain real-time status information of enemy and friendly UAVs, and construct a comprehensive air combat situation advantage function. The comprehensive air combat situation advantage function includes five sub-items: azimuth advantage, entry angle advantage, range advantage, altitude advantage, and energy advantage.
[0013] Step S3: Design adaptive weighting rules based on enemy-friendly distance, missile attack range, and angle conditions to dynamically adjust the weights of the advantages of each sub-item in Step S2.
[0014] Step S4: Introduce the target allocation model to maximize our overall air combat situation. By solving the target allocation matrix, the 2v2 air combat problem is reduced in dimension and transformed into two parallel 1v1 air combat subproblems.
[0015] Step S5: The weighted comprehensive air combat situation advantage function is used as the optimization objective function of the nonlinear model predictive control. Under the premise of satisfying the performance constraints of the UAV, the optimal control problem for air combat maneuver decision-making is established.
[0016] Step S6: Discretize the optimal control problem using the control parameterization method, and use the constraint transcription method to transform the state inequality constraints into differentiable integral equality constraints to construct a nonlinear programming problem.
[0017] Step S7: Solve the nonlinear programming problem in each control cycle to generate real-time maneuvering decision commands to drive the UAV to execute.
[0018] Further, the three-degree-of-freedom kinematic model of the UAV established in step S1 includes: selecting the UAV's state vector and control vector; the state vector includes flight speed, yaw angle, pitch angle, and position coordinates in a three-dimensional inertial frame; the control vector includes tangential overload, normal overload, and roll angle; establishing guidance equations based on the principles of UAV dynamics to describe the relationship between the UAV's position and velocity vectors and the control variables, and discretizing the guidance equations using the fourth-order Runge-Kutta method to obtain discrete difference equations for predicting the state at the next moment.
[0019] Furthermore, the construction rules for each sub-item of the comprehensive air combat situation advantage function described in step S2 are as follows:
[0020] The range advantage function is constructed to be a negative exponential function of the enemy-friendly distance when the enemy-friendly distance is greater than the maximum attack distance, and a constant when the enemy-friendly distance is less than the maximum attack distance, in order to drive the UAV into the effective attack range;
[0021] The azimuth dominance function and the entry angle dominance function are constructed as linearly decreasing functions of the ratio of angle to pi, used to drive the angle toward zero.
[0022] The height advantage function is constructed as a sigmoid function with respect to relative height within a set height range, used to enhance height advantage within the set height range;
[0023] The energy advantage function is constructed as a sigmoid function of the ratio of enemy and friendly maneuvering energy when the ratio is within a set range, and a constant when the ratio exceeds the set range. It is used to dynamically reflect the energy advantage when the ratio is in the transition region, where maneuvering energy combines speed and altitude potential energy.
[0024] Furthermore, the adaptive weighting rules in step S3 specifically include: when the distance between the enemy and friendly forces is greater than the missile's maximum attack range and exceeds a set safety threshold, the weighting coefficient of the distance advantage is set to the maximum, the angle weight is ignored, and the UAV is driven to accelerate and approach the enemy aircraft; when the distance between the enemy and friendly forces is less than the missile's maximum attack range and greater than the set safety threshold, the weighting coefficients of altitude advantage and energy advantage are increased, enabling the UAV to gain energy advantage through climbing; when the distance between the enemy and friendly forces is close to or enters the missile's maximum attack range, and the enemy's entry angle is within the threat range, the weighting coefficients of azimuth advantage and entry angle advantage are set as the main weights to ensure tactical positioning and personal safety; when both sides are in the attack zone and the attack angle conditions are met, the altitude weight is reduced, and the high angle weight is maintained to maintain the attack posture.
[0025] Furthermore, the target allocation model in step S4 takes maximizing the sum of our overall air combat strike situation and air combat process situation as its objective function; the air combat strike situation is obtained by calculating the difference between the probability that our missile attack zone covers the position of the enemy UAV and the probability that the enemy missile attack zone covers the position of our UAV; the air combat process situation is constructed as a function of our aircraft's azimuth angle, the enemy aircraft's entry angle, and the distance between us and the enemy. This function is characterized by calculating the normalized ratio of pi to the difference between the azimuth angle and the entry angle, and combining it with a negative exponential decay term when the distance between us and the enemy exceeds the maximum attack distance, to represent the attack geometry of both sides at the current distance.
[0026] Furthermore, the solution process of the target allocation model includes: defining a target allocation matrix with rows corresponding to the number of friendly drones and columns corresponding to the number of enemy drones; constructing an optimization problem to find the optimal allocation matrix element values such that the sum of the products of the target allocation matrix elements and the corresponding objective function values is maximized; setting constraints to restrict the sum of the elements in each row of the allocation matrix to be equal to 1, so as to ensure that each friendly drone locks onto only one enemy target at any given time.
[0027] Furthermore, the UAV performance constraints in step S5 include state constraints and control constraints; the state constraints limit the minimum and maximum values of flight speed and flight altitude; the control constraints limit the minimum and maximum values of normal overload, tangential overload, roll angle and their respective rates of change.
[0028] Furthermore, the control parameterization method described in step S6 specifically includes: discretizing the continuous air combat decision optimization time domain into several time nodes; approximating the control quantity using a piecewise constantization method, setting the control quantity as a constant within a certain time period, thereby reducing the dimensionality of the optimization variables; and setting the number of segments according to the response characteristics of the underlying controller, so that the discretized control commands can be effectively tracked by the underlying flight control system.
[0029] Furthermore, step S6, which involves using the constraint transcription method to transform state inequality constraints into differentiable integral equality constraints, specifically includes: uniformly representing the upper and lower bound inequality constraints of all state variables as standard form state inequality constraint functions; introducing two greater than zero slack variables to construct a smoothing penalty function containing slack variables and relating to the state inequality constraint function values; defining an integral constraint function, which is the integral of the smoothing penalty function over the entire optimization time domain; and replacing the multiple discrete-point state inequality constraints in the original nonlinear programming problem with equality constraints or single-point constraints whose integral constraint function values are less than or equal to zero, thereby reducing the solution complexity of the nonlinear programming problem.
[0030] Further, step S7 specifically includes: at the initial moment of each control cycle, using the currently observed state information, solving the nonlinear programming problem constructed in step S6 to obtain the optimal control quantity sequence in the future prediction time domain; selecting the control quantity of the first control cycle in the optimal control quantity sequence as the current maneuver decision command; and inputting the maneuver decision command into the UAV flight control system.
[0031] Compared with the prior art, the beneficial effects of the present invention are:
[0032] (1) Significantly improves the real-time performance and solution efficiency of online decision-making. Compared with traditional differential game theory or conventional nonlinear model predictive control, this invention introduces control parameterization and constraint transcription techniques. By transforming the originally discrete and numerous state inequality constraints (such as speed, altitude, and overload limits) into smooth and differentiable integral equality constraints, the complexity and constraint scale of the nonlinear programming problem are greatly reduced. This enables complex air combat maneuver decision-making problems to converge quickly with limited computing resources, meeting the millisecond-level real-time requirements of air combat scenarios.
[0033] (2) No large-scale data training is required, and it has good versatility and portability. Compared with decision-making methods based on reinforcement learning, this invention does not require massive amounts of simulation data for training, avoiding the high transfer training costs caused by changes in fighter jet models and radar / missile parameters. This invention constructs an optimization problem based on a dynamic model and situation function. When platform parameters change, only the model parameters need to be updated, which has stronger engineering versatility.
[0034] (3) Intrinsic hard constraints ensure flight safety. Unlike methods based on expert rules or deep learning that often treat constraints such as overload and roll angle as post-hoc verification, this invention directly embeds the maneuverability limits of the UAV (such as maximum overload and roll angle change rate) as intrinsic hard constraints into the NMPC optimization model. Through rolling time-domain optimization, it is ensured that every guidance command generated strictly falls within the flight envelope, effectively avoiding flight accidents caused by maneuver commands exceeding physical limits, and significantly improving the safety of autonomous air combat.
[0035] (4) Effectively solves the curse of dimensionality problem in multi-aircraft collaboration. This invention adopts a hierarchical architecture of "target allocation-situational advantage-NMPC", which reduces the high-dimensional coupled 2v2 game problem into two parallel 1v1 optimal control subproblems through the target allocation model. This design not only ensures the optimization of the overall strike situation and process situation of the formation, but also avoids the exponential expansion of the state space caused by directly solving multi-aircraft joint operations, and realizes an effective combination of collaborative decision-making and single-aircraft maneuver.
[0036] (5) Enhanced tactical adaptability to dynamic battlefield environments. This invention designs an adaptive weighting rule based on enemy-friendly distance and missile attack zone status. This rule can dynamically adjust the weights of advantageous terms such as azimuth, distance, and energy in the objective function according to changes in the battlefield situation (such as whether the no-escape zone has been entered, the advantage or disadvantage of the angle, etc.). This enables UAVs to flexibly switch between different tactical intentions such as "long-range positioning," "medium-range attack," and "close-range dogfighting," just like human pilots, thus improving the intelligence level of decision-making.
[0037] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, embodiments of the present invention are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0038] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a framework diagram of the technical solution of the present invention;
[0040] Figure 2 This is a diagram illustrating the changing targets of a 2v2 attack by a drone according to an embodiment of the present invention.
[0041] Figure 3 This is a flight trajectory diagram of a UAV during a 2v2 beyond-visual-range air combat scenario according to an embodiment of the present invention, at a time of 122 seconds.
[0042] Figure 4 This is a flight trajectory diagram of a UAV during a 2v2 beyond-visual-range air combat exercise at 144 seconds, according to an embodiment of the present invention.
[0043] Figure 5 This is a weight allocation diagram of the air combat advantage function according to an embodiment of the present invention;
[0044] Figure 6 This is a command tracking diagram of a six-degree-of-freedom model of a UAV according to an embodiment of the present invention;
[0045] Figure 7 This is a graph showing the overload, roll angle, and rate of change of a UAV according to an embodiment of the present invention.
[0046] Figure 8 This invention demonstrates that, compared to traditional nonlinear model predictive control algorithms, it introduces constraint transcription technology in control parameterization to smooth nonlinear constraints, thereby improving the solution speed of the optimal problem. A comparison of the solution time of the air combat optimal problem in two scenarios shows that the method adopted by this invention has a shorter solution time and can better meet the real-time requirements of air combat maneuver decision-making. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.
[0048] This invention proposes a real-time decision-making method for 2v2 UAV air combat based on nonlinear model predictive control (NMPC). This method, through a "target assignment-situation advantage-constraint transcription-NMPC" architecture, reduces the complex 2v2 game into two parallel 1v1 optimal control subproblems. The specific implementation steps are as follows:
[0049] 1. Establish a three-degree-of-freedom kinematic model of a UAV.
[0050] When studying aircraft maneuver decision-making and trajectory generation, the general mass kinematics model is selected. After obtaining guidance commands such as reference overload of the aircraft through the decision algorithm, the actual aircraft maneuver control is carried out through the flight control law.
[0051] The kinematic model of the aircraft mass is selected as follows:
[0052] (1.1)
[0053] Where: UAV mass state vector By flight speed Yaw angle Pitch angle UAV location coordinates Composition; control vector Tangential overload Overload and roll angle constitute.
[0054] Based on the discretization of the maneuver decision optimization problem using the three-degree-of-freedom dynamic guidance equations of the aircraft, the above-mentioned kinematic model of the aircraft can be expressed as follows:
[0055] (1.2)
[0056] in, For the spacecraft at the initial moment state, Let be the aircraft state vector at time t. Let be the control vector at time t.
[0057] In the subsequent NMPC solution, the model is discretized using the fourth-order Runge-Kutta method to obtain discrete difference equations for predicting the state at the next time step.
[0058] 2. Construct a comprehensive air combat situational advantage function
[0059] Considering that the subsequent target allocation method can transform the multi-aircraft air combat problem into multiple local single-aircraft air combat problems, we establish the following single-aircraft air combat comprehensive advantage function based on the real-time status information of enemy and friendly UAVs:
[0060] (1.3)
[0061] In the formula, R is the value of the comprehensive air combat situation advantage function; These are five sub-items: azimuth advantage, angle of entry advantage, altitude advantage, distance advantage, and energy advantage; These are the weight coefficients corresponding to each sub-advantage, and all are greater than 0, satisfying the condition that... .
[0062] The specific design of each advantage function is given below:
[0063] Define the real-time distance between the two machines as The maximum attack range of drone weapons is By optimizing the range advantage function, the distance to enemy aircraft can be reduced, allowing them to gradually enter the attack range of our weapons.
[0064] The range advantage function is constructed as a negative exponential function of the enemy-friendly distance when the enemy-friendly distance is greater than the maximum attack range, and a constant when the enemy-friendly distance is less than the maximum attack range. This function is used to drive the UAV into its effective attack range. Specifically:
[0065] (1.4)
[0066] Define the azimuth angle as the angle between the line connecting the velocity vector of our aircraft and the line connecting our position to the position vector of the enemy. This reduces our azimuth angle, gradually bringing enemy aircraft within our weapon's attack angle. The azimuth advantage function is constructed as a linearly decreasing function of the ratio of angle to pi, used to drive the angle towards zero, specifically:
[0067] (1.5)
[0068] Define the angle between the enemy aircraft's velocity vector and the line connecting our aircraft's position to the enemy's position vector as the angle of entry. Reducing the enemy's angle of entry can disrupt the enemy's weapon's angle of attack. The angle of entry advantage function is constructed as a linearly decreasing function of the ratio of angle to pi, used to drive the angle towards zero, specifically:
[0069] (1.6)
[0070] Define the altitude of our aircraft relative to the enemy aircraft as Increasing the relative altitude of our aircraft can extend the missile attack range of our aircraft, thereby enabling the weapons to strike ahead of time. The altitude advantage function is constructed as a sigmoid function with respect to relative altitude within a set altitude range, used to enhance altitude advantage within that range, specifically:
[0071] (1.7)
[0072] in Let represent the relative altitude between the two drones at the current moment. In aerial combat, increasing altitude can increase the attack range of our drone's weapons, so increasing the distance between our drone and the enemy is advantageous. However, excessive relative altitude will cause our drone to ignore relative angles, so the relative altitude should be within a certain range for optimal performance. The piecewise form of formula (1.7) is an approximate engineering implementation of the characteristics of the Sigmoid function, aiming to define the advantage of increasing altitude within a certain range (e.g., relative altitude from 0 to 2000m).
[0073] Define the maneuvering energy of our UAVs during air combat as Considering both maneuverability and flight speed and flight altitude Factors include maximizing the maneuverability advantage of friendly drones during air combat to enable them to attack the enemy drones before they can. Therefore, an energy advantage function is constructed as a sigmoid function of the ratio of friendly to enemy maneuverability energy within a set range, and a constant when the ratio exceeds the set range. This function dynamically reflects the energy advantage when the ratio is in a transitional region. Specifically:
[0074] (1.8)
[0075] The maneuvering energy during aerial combat between the two drones is defined as a function of speed and altitude in the above equation. The energy advantage function is designed based on the ratio of the maneuvering energy of the two drones as follows:
[0076] (1.9)
[0077] Where E is the specific energy. For the maneuvering energy of our drones, Provides mobility energy for enemy drones.
[0078] 3. Design adaptive weight rules
[0079] In aerial combat, the combat situation is constantly changing. Fixed advantage function weights cannot adapt to all combat scenarios. The combat situation between UAVs changes rapidly, and using fixed optimization function weight coefficients would lead to overly simplistic optimal maneuver decisions. Therefore, this invention establishes adaptive rules for the weight coefficients in the advantage function, aiming to create a dynamic advantage function that changes with the combat situation, thereby better guiding our aircraft in making maneuver decisions. Specifically:
[0080] For weighting coefficients The specific weighting rules are as follows:
[0081] 1) Long-range pursuit phase: At the current moment, neither drone has entered its respective weapon's attack range and the distance between the two drones is... (That is, greater than the set threshold), meaning the distance between the two drones is still far from the attack range, it is too early to consider the angles of both drones during the air combat. Our drone should accelerate at full speed to catch up with the enemy drone, ignoring the angle weight. The weight of the air combat advantage function should be set to... .
[0082] 2) Mid-range positioning phase: At the current moment, neither of the two machines has entered the attack range and the distance is... When the distance between the two drones approaches again, but is still greater than the attack range of both drones' weapons, our drones will increase their altitude to gain a greater weapon attack range. At the same time, our drones will enhance their energy advantage, enabling them to have greater maneuverability and respond quickly once both drones enter their weapon attack range.
[0083] The weights of the air combat advantage function should be set to .
[0084] 3) Close-range defense phase: At the current moment, neither of the two machines has entered the attack range and the distance is... When the distance between the two drones approaches again, but remains greater than the attack range of both drones' weapons, although they haven't yet entered the attack range of either side's weapons, the distance isn't too far either. To ensure the safety of our drones, we should ensure that the enemy's angle of entry is not too large. If the enemy's angle of entry meets the requirements... (Significant threat) To ensure safety, the enemy's entry angle should be limited, and the weights of the air combat advantage function should be set to... .
[0085] 4) Close-range attack preparation phase: At the current moment, neither of the two machines has entered the attack range and the distance is... (Close attack range), meaning when the two aircraft are close to each other within the attack range, it is crucial to maximize the angular advantage of our drones during air combat and simultaneously reduce the distance between the two enemy aircraft.
[0086] Furthermore, when our drone weapons have a greater attack range than the enemy's (a significant advantage), the weight of the air combat advantage function should be set to...
[0087] ,
[0088] Furthermore, when the attack range of our drone weapons is smaller than that of the enemy, the weight of the air combat advantage function should be set to...
[0089] .
[0090] 5) Attack Phase: When both sides enter the attack zone, but our side has a greater angular advantage (meeting the angular condition), the weight of the air combat advantage function should be set to... And in other cases, settings .
[0091] 4. Target allocation in multi-aircraft air combat
[0092] After designing the beyond-visual-range air combat advantage function and weighting coefficients, a target allocation model is introduced to transform the 2v2 air combat into two simultaneous 1v1 air combats. The input to the target allocation is the state of the enemy and friendly UAVs, and the output is the strike targets of our two UAVs. The allocation is based on maximizing the air combat situation objective function, which can be divided into air combat strike situation and air combat process situation.
[0093] (1) Definition of air combat strike situation and process situation
[0094] Regarding our drones, air combat strike posture Defined as our air combat strike posture and the enemy's posture of attacking our side The difference:
[0095] (2.1)
[0096] The respective air combat strike postures of both sides' drones (i.e., the probability of missile attack zone coverage) are represented as follows:
[0097] (2.2)
[0098] (2.3)
[0099] in, and These are the coordinates of our side and the enemy's positions, respectively. and These are the attack zones for our and the enemy's air combat weapons, respectively.
[0100] Air combat situation It is constructed as a function of our aircraft's azimuth, the enemy aircraft's approach angle, and the distance between us and the enemy. This function characterizes the attack geometry of both sides at the current distance by calculating the normalized ratio of pi to the difference between the azimuth and the approach angle, and incorporating a negative exponential decay term when the distance between us and the enemy exceeds the maximum attack range. Specifically:
[0101] (2.4)
[0102] During the air combat, our drones against enemy drones Objective function for air combat target allocation It is the sum of these two states, which can be expressed as:
[0103] (2.5)
[0104] In 2v2 air combat, after our two drones receive the real-time status of the enemy's two drones, they use a target allocation model to select the enemy drone with the greater attack posture for each of our drones to make maneuver decisions, so as to optimize the overall attack posture of our drones, that is, to have a greater advantage in the battle.
[0105] (2) Target assignment matrix and solution
[0106] Define the target assignment matrix :
[0107] (2.6)
[0108] In the formula, and These represent the total number of friendly and enemy drones in a 2v2 scenario. , Choose 0 or 1; choosing 1 represents our drone. The target of the attack was enemy drones. If the value is 0, it represents a drone. The target was not an enemy drone. .
[0109] To optimize our attack posture after target allocation, we solve the following target allocation problem to obtain the target allocation matrix at the current moment. :
[0110] (2.7)
[0111] In formula (2.7), the constraint conditions require each row All The sum must equal 1, which means that at any given moment each of our drones has exactly one target, while an enemy drone can be targeted by multiple of our drones simultaneously. The result is based on The value of , under the constraints, will When they are the same, the relatively larger corresponding Corresponding The value is assigned to 1 and filled into the matrix, thus completing the target allocation.
[0112] In the next step, only the rolling time domain maneuver trajectory optimization will be performed on the "one-to-one" confrontation relationship (which aircraft is attacking which enemy aircraft) that has been determined by the target allocation model in this step, and the target will not be reselected or switched.
[0113] 5. Maneuver decision optimization based on nonlinear model predictive control
[0114] This section establishes the optimal control problem for a single UAV in air combat within a nonlinear model predictive control framework, and then performs real-time decision optimization using a control parameterization method. The established optimization problem uses a constructed weighted air combat advantage function as the objective function, and considers UAV performance constraints (including state constraints and control constraints). The optimized maneuver decision yields the desired control quantities from the UAV's three-degree-of-freedom kinematic equations, including normal overload, tangential overload, and roll angle. These desired control quantities are then fed to the UAV's six-degree-of-freedom flight controller. The flight controller calculates the actual throttle and rudder deflection of the UAV based on the desired control quantities, thus obtaining the UAV's state at the next moment.
[0115] During UAV air combat flight, it is necessary to constantly consider necessary state constraints such as speed and altitude, as well as control input limitations such as overload and roll angle. Specifically, we consider the following constraints:
[0116] (3.1)
[0117] Based on the established UAV guidance model, single-aircraft air combat advantage function, target allocation algorithm, and UAV performance constraints, we can obtain our UAV's... Air combat maneuver decision-making optimal control problem :
[0118] (3.2)
[0119] In the formula, Optimize the time domain for air combat decision-making. and For our drones State and control variables, For enemy drones The state variables, The value is determined by solving the current The problem at that moment (2.7) was decided. Due to the enemy aircraft's status... Only The time is known, in the future The inside is unknown, if If the value is relatively small, it can be assumed that its value remains approximately constant, or mature methods such as polynomial fitting and intention prediction can be used to predict the enemy aircraft's status, providing input for calculating the air combat advantage value.
[0120] In formula (3.2) This refers to the objective function that needs to be optimized at this specific moment. In other words, it refers to the states of both sides. After substituting the values, the weights of the fuzzy rules are calculated to obtain the air combat advantage function. The calculation rules are detailed in formula (1.3).
[0121] For the air combat maneuver decision optimization problem (3.2), we solve it within a nonlinear model predictive control framework to achieve the maneuver decision-making effect of "situation prediction and feedback execution," and to some extent reduce the problem of prediction errors in the future time domain regarding the enemy UAV's state. The optimal control sequence is solved in each control cycle using a rolling time-domain optimization strategy. However, in actual guidance processes, only the following is executed. Optimal control quantity within time By utilizing state feedback and periodic online re-optimization to achieve closed-loop control, the system ensures rapid response to changes in the battlefield situation while meeting constraints, thereby continuously enhancing air combat superiority. Among these, For execution segment, This is the segment for predicting air combat advantages.
[0122] 6. Discretization and Constrained Transcription Solving
[0123] The air combat maneuver decision optimization problem (3.2) is time-domain continuous and has complex guidance model constraints, UAV state and control constraints, making it difficult to solve online. Here, a control parameterization method is used to simplify the problem, transforming it into a low-complexity nonlinear programming problem.
[0124] First, the continuous time domain is discretized to obtain a discrete step size. Below Approximate time-point sequence :
[0125] (3.3)
[0126] After discretizing time, the continuous-time maneuver decision control quantity can be approximated by parameter discretization, that is...
[0127] (3.4)
[0128] Based on the UAV guidance model (1.2), the following discrete difference equation can be obtained using the fourth-order Runge-Kutta method:
[0129] (3.5)
[0130] State and control constraints can be approximated as corresponding constraints at each discrete point. Taking overload constraints as an example, they can be approximated as follows:
[0131] (3.6)
[0132] Due to the high real-time requirements of air combat maneuver decision-making and the short optimization time domain, further optimization of the tracking control hysteresis and computation time of the underlying controller is needed, resulting in piecewise constantized control variables:
[0133] (3.7)
[0134] In the formula, Let be the number of segments, which is an integer and such that... It is also an integer.
[0135] After discretizing the state constraints such as velocity in the optimal maneuver optimization problem of UAV, as shown in Equation (3.7), the number of each constraint is reduced by the number of discrete constraints. It is determined that the number of nonlinear constraints in the underlying nonlinear programming problem will be large, resulting in a slow solution speed. Therefore, the state inequality constraints are processed using the constraint transcription method, transforming them into one-dimensional state integral inequality constraints. First, the state inequality constraints in the optimization problem (3.2) are uniformly expressed as...
[0136] (3.8)
[0137] In the formula,
[0138] (3.9)
[0139] By introducing slack variables, constructing a smoothing penalty function, and transforming it into an integral constraint, the specific implementation is as follows:
[0140] (3.10)
[0141] In the formula, As slack variables, The smoothing penalty function is defined as follows:
[0142] (3.11)
[0143] Through the above processing, the optimal control optimization problem for maneuver decision-making in beyond-visual-range air combat (3.2) can be transformed into the following low-complexity nonlinear programming problem that can be solved in real time. :
[0144] (3.12)
[0145] Thus, we have transformed the multi-aircraft beyond-visual-range air combat maneuver decision-making problem into multiple nonlinear programming problems of single-aircraft air combat maneuver decision-making that can be solved in parallel. .
[0146] 7. Rolling Solving and Execution
[0147] Within each control cycle, the above nonlinear programming problem is solved using solvers such as Fmincon to obtain the optimal control sequence. Only the control input from the first control cycle in the sequence is input into the UAV flight control system as the current maneuver decision command. At the next moment, the UAV state is updated. Repeat steps 1-7 above to achieve closed-loop rolling optimization.
[0148] 8. Simulation Verification and Conclusion
[0149] To verify the feasibility of the 2v2 air combat decision-making method based on the NMPC algorithm and target allocation model described in this invention, further explanation is provided below with simulation examples. The simulation was performed in the Matlab 2022b environment, using the Fmincon solver to solve the discretized nonlinear programming problem online. It is assumed that both friendly and enemy UAVs are of the same type, i.e., all flight parameters are identical. The thresholds for various maneuver performance constraints of the UAVs are determined based on the specific aircraft model's dynamics and the control module's performance. The aircraft control and state constraints adopted in this patent are as follows:
[0150] (4.1)
[0151] The simulation parameters are set as follows:
[0152] (4.2)
[0153] A 2v2 beyond-visual-range (BVR) air combat is being conducted between red and blue drones. Considering that the two blue drones employ a pursuit strategy, chasing Red 1 and Red 2 respectively, the blue drones are making maneuvering decisions to close the distance to the enemy drones as quickly as possible. The initial states of the four drones on both sides are as follows:
[0154] (4.3)
[0155] The target change diagram obtained by the two Red Team drones according to the target allocation model is shown below. Figure 2 As shown, the trajectories of the red and blue drones in a 2v2 medium-range air combat are as follows: Figure 3 and Figure 4 As shown, during the air combat, based on the drone target allocation model, our two drones were assigned target drones with a greater air combat advantage. Before 122 seconds, the two Red drones, based on the target allocation model, comprehensively considered air combat distance and altitude, assigning them targets with a relatively greater air combat advantage. After obtaining the targets, the two Red drones optimized their decisions based on the air combat advantage function to gain a greater air combat advantage until victory. At 122 seconds, Red 1 defeated Blue 1 first. Figure 3 As shown, Red 1 and Red 2 then locked onto Blue 2, and finally, at 144 seconds, Red 2 defeated Blue 2 first. Figure 4 As shown, they achieved final victory in the air battle.
[0156] Taking Red 1 as an example, the changes in the weight values of the advantage function during air combat are as follows: Figure 5 As shown, in UAV air combat, the initial advantage lies in altitude and energy to increase weapon attack range. As enemy and friendly UAVs approach, the angular advantage is increased to bring the enemy UAV into the friendly UAV's attack range. Once the angular advantage is secured, the distance is quickly closed to achieve a preemptive strike. The output commands of the decision module during the Red Force UAV's flight and the actual tracking effect of the target drone are shown in the figure. Figure 6 As shown, the tracking trend was consistent and the tracking error was small throughout the entire flight phase. Overload tracking fluctuated slightly around 40 seconds, but quickly resumed tracking to the actual overload. Speed and roll angle were tracked well throughout the entire process, with roll angle tracking error not exceeding 5 degrees and speed tracking error not exceeding 5 m / s. During the overall flight phase, the aircraft's overload, roll angle, and rate of change were as follows: Figure 7 As shown, all performance constraints of the UAV are met: maximum overload is 2.3, maximum overload change rate is 1.4, maximum roll angle is 47 degrees, and maximum absolute value of roll angle change rate is 16.
[0157] The algorithm proposed in this invention can guarantee the trackability of guidance commands and the satisfaction of flight constraints. This invention takes into account the maneuverability of actual UAVs, constraining not only the UAV's control variables but also the rate of change of those variables. Furthermore, it solves the problem within a framework of nonlinear model predictive control, resulting in an air combat trajectory that satisfies the UAV's performance constraints, as shown in the simulation scenario described above. Figure 3 and Figure 4 The optimized control variables and their rates of change are all within the set constraints. The optimized real-time commands are loaded into the UAV's six-DOF Simulink model for tracking and verification. The tracking and verification process is as follows: Figure 6 As shown, the overall tracking trend is consistent and the tracking error is small. While there are local deviations, they can be quickly corrected later. Compared to traditional nonlinear model predictive control algorithms, this invention introduces constraint transcription technology in control parameterization to smooth nonlinear constraints, thereby improving the solution speed of the optimal problem. The solution time for the air combat optimal problem in the two scenarios is compared as follows: Figure 8 As shown, the method used in this invention has a shorter solution time and can better meet the real-time requirements of air combat maneuver decision-making.
[0158] This invention proposes an online optimization algorithm for maneuver decision-making in multi-aircraft beyond-visual-range (BVR) air combat. First, the complex air combat problem is decomposed into multiple single-aircraft air combat adversarial problems using a target allocation method. Then, a weighted adaptive air combat advantage function is designed under time-varying missile attack ranges, and considering the actual performance constraints of the UAVs, a rolling optimization problem for UAV air combat maneuver decision-making is established within a nonlinear model predictive control framework. Finally, a control parameterization method is used to obtain an online solvable nonlinear programming problem for maneuver decision-making. The proposed maneuver decision-making optimization method exhibits good flight safety, engineering feasibility (guided commands can be tracked), decision optimality, real-time performance, portability, and scalability.
[0159] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A 2 vs 2 intelligent air combat real-time maneuver decision method, characterized in that, The method comprises the following steps: Step S1, a three-degree-of-freedom particle kinematics model of a UAV is established, and the kinematics model is taken as an NMPC prediction model; Step S2, real-time state information of enemy and own UAVs is acquired, and a comprehensive air combat situation advantage function is constructed, the comprehensive air combat situation advantage function comprising five sub-items of azimuth angle advantage, entering angle advantage, distance advantage, height advantage and energy advantage; Step S3, an adaptive weight rule is designed according to the distance between the enemy and the own UAV, the missile attack distance and the angle condition, and the weight of each sub-item advantage in step S2 is dynamically adjusted; Step S4, a target allocation model is introduced, the target allocation model taking maximizing the overall air combat situation of the own side as a target, and the 2v2 air combat problem is reduced in dimension into two parallel 1v1 air combat sub-problems through solving a target allocation matrix; Step S5, the weighted comprehensive air combat situation advantage function is taken as an optimization target function of nonlinear model prediction control, and an optimal control problem of air combat maneuvering decision is established under the premise of satisfying the UAV performance constraint; Step S6, a control parameterization method is adopted to discretize the optimal control problem, and a state inequality constraint is converted into a derivable integral equation constraint by using a constraint transcription method, and a nonlinear programming problem is constructed; Step S7, the nonlinear programming problem is solved in each control period, and real-time maneuvering decision instructions are generated to drive the UAV to execute.
2. The 2 vs 2 intelligent air combat real-time maneuver decision method according to claim 1, characterized in that, The three-degree-of-freedom particle kinematics model established in step S1 comprises: selecting a UAV particle state vector and a control amount vector; the state vector comprises flight speed, yaw angle, pitch angle and position coordinates in a three-dimensional inertial system; the control amount vector comprises tangential overload, normal overload and roll angle; a guidance equation is established based on UAV dynamics principles, the guidance equation describes the change relationship between the position and velocity vector of the UAV and the control amount, and the guidance equation is discretized by using a fourth-order Runge-Kutta method, and a discrete difference equation for predicting the state at the next time is obtained.
3. The 2 vs. 2 intelligent air combat real-time maneuver decision method according to claim 1, characterized in that, The construction rule of each sub-item of the comprehensive air combat situation advantage function in step S2 is as follows: The distance advantage function is constructed as a negative exponential function about the distance between the enemy and the own UAV when the distance is greater than the maximum attack distance, and as a constant when the distance is less than the maximum attack distance, for driving the UAV to enter the effective attack distance; The azimuth angle advantage function and the entering angle advantage function are constructed as linearly decreasing functions about the ratio of the angle to the circular constant, for driving the angle to tend to zero; The height advantage function is constructed as a Sigmoid function about the relative height within a set height range, for improving the height advantage within the set height range; The energy advantage function is constructed as a Sigmoid function about the ratio of the maneuvering energy of the enemy to the own UAV when the ratio is within a set interval, and as a constant when the ratio is out of the set interval, for dynamically reflecting the energy advantage when the ratio is in the transition region, wherein the maneuvering energy comprehensively considers the speed and height potential energy.
4. The 2 vs. 2 intelligent air combat real-time maneuver decision method according to claim 1, characterized in that, The adaptive weight rule in step S3 specifically includes: when the distance between the enemy and the UAV is greater than the maximum attack distance of the missile and exceeds a set safety threshold, setting the weight coefficient of the distance advantage to be the maximum, ignoring the angle weight, and driving the UAV to accelerate to approach the enemy; when the distance between the enemy and the UAV is less than the maximum attack distance of the missile and greater than the set safety threshold, increasing the weight coefficients of the height advantage and the energy advantage, so that the UAV obtains the energy advantage by climbing; when the distance between the enemy and the UAV is close to or within the maximum attack distance of the missile, and the enemy enters the threat range, setting the weight coefficients of the azimuth angle advantage and the entering angle advantage as the main weights, to ensure the tactical occupation and safety; when both sides are in the attack area and meet the attack angle condition, reducing the height weight and maintaining the high angle weight to maintain the attack posture.
5. The 2 vs 2 intelligent air combat real-time maneuver decision method according to claim 1, characterized in that, The target allocation model in step S4 takes the sum of the air combat attack posture and the air combat process posture of the UAV as the objective function; the air combat attack posture is obtained by calculating the difference between the probability that the missile attack area of the UAV covers the position of the enemy UAV and the probability that the missile attack area of the enemy covers the position of the UAV; the air combat process posture is constructed as a function of the azimuth angle of the UAV, the entering angle of the enemy and the distance between the enemy and the UAV, which is constructed by calculating the normalized ratio of the difference between the azimuth angle and the entering angle and the circular constant, and combining the negative exponential decay term when the distance between the enemy and the UAV exceeds the maximum attack distance, to represent the attack geometric posture of the enemy and the UAV at the current distance.
6. The 2 vs. 2 intelligent air combat real-time maneuver decision method according to claim 5, characterized in that, The solving process of the target allocation model includes: defining a target allocation matrix with the number of rows corresponding to the number of UAVs of the UAV and the number of columns corresponding to the number of UAVs of the enemy; constructing an optimization problem to find the optimal element value of the allocation matrix, so that the cumulative sum of the product of the element of the allocation matrix and the corresponding target function value is the maximum; setting the constraint condition to limit the sum of each row element of the allocation matrix to be equal to 1, to ensure that each UAV of the UAV locks only one enemy target at each time.
7. The 2 vs. 2 intelligent air combat real-time maneuver decision method according to claim 1, characterized in that, The UAV performance constraint in step S5 includes state constraints and control constraints; the state constraints limit the minimum and maximum values of the flight speed and the flight height; the control constraints limit the minimum and maximum values of the normal overload, the tangential overload, the roll angle and their respective change rates.
8. The 2 vs. 2 intelligent air combat real-time maneuver decision method according to claim 1, characterized in that, The control parameterization method in step S6 specifically includes: discretizing the continuous air combat decision optimization time domain into a plurality of time nodes; using the piecewise constant approximation method to approximate the control quantity, and setting the control quantity as a constant within a certain time period, thereby reducing the dimension of the optimization variable; setting the number of segments according to the response characteristics of the bottom controller, so that the discrete control instruction can be effectively tracked by the bottom flight control system.
9. The 2 vs. 2 intelligent air combat real-time maneuver decision method according to claim 1, characterized in that, The state inequality constraint is converted into a derivable integral equation constraint by using the constraint transcription method in step S6, and the conversion specifically includes: representing the upper and lower bound inequality constraints of all state variables as state inequality constraint functions in a standard form; introducing two slack variables greater than zero, and constructing a smooth penalty function containing the slack variables and related to the state inequality constraint function values; defining an integral constraint function, which is the integral of the smooth penalty function over the entire optimization time domain; replacing the multiple discrete point state inequality constraints in the original nonlinear programming problem with an equality constraint or a single point constraint that the integral constraint function value is less than or equal to zero, thereby reducing the solution complexity of the nonlinear programming problem.
10. The 2 vs. 2 intelligent air combat real-time maneuver decision method according to claim 1, characterized in that, Step S7 specifically includes: at the initial time of each control period, using the currently observed state information, solving the nonlinear programming problem constructed in step S6 to obtain an optimal control quantity sequence in a future prediction time domain; selecting the control quantity of the first control period in the optimal control quantity sequence as the current maneuver decision instruction; and inputting the maneuver decision instruction into the unmanned aerial vehicle flight control system.
Citation Information
Patent Citations
Hesitant fuzzy and dynamic deep reinforcement learning combined unmanned aerial vehicle maneuvering decision-making method
CN111666631A
Unmanned aerial vehicle air combat maneuver decision-making method for simulating Harris hawk intelligent predation optimization
CN113741500A
Close-range air combat maneuver decision-making method and system based on target maneuver intention prediction
CN115993835A
Aircraft over-the-horizon air combat maneuver decision-making method based on nonlinear predictive control
CN116861645A
Close-range air combat maneuver decision optimization method based on control parameterization
CN117192982A