Fighter bottom layer flight control method based on deep reinforcement learning

By applying deep reinforcement learning methods in the underlying flight control of fighter aircraft, combined with PPO algorithm and GAE method, the problem of poor flight stability of air combat agents in high dynamic environments is solved, and more efficient training and more stable control effects are achieved.

CN120233791AActive Publication Date: 2025-07-01POLIXIR TECH LTD

Patent Information

Application Number
CN202411807483.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-07-01
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

Air combat agents are prone to multiple stalls and multiple jitter problems in high dynamic environments, resulting in poor flight stability. Traditional deep reinforcement learning algorithms converge slowly in large-scale distributed interactive environments, making it difficult to meet the needs of air combat applications.

Method used

The fighter's underlying flight control method based on deep reinforcement learning is adopted. Through the deep neural network model, the fighter's underlying control system is refined to optimize the underlying flight control module, reduce the training of basic decisions for the underlying aircraft control, and accelerate the training process using PPO algorithm and generalized advantage estimation (GAE) method.

Benefits of technology

It significantly improves the flight control stability and convergence speed of the agent in a high dynamic environment, reduces stall and jitter phenomena, and improves training efficiency and fault tolerance, adaptability and robustness of the control system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233791A_ABST
    Figure CN120233791A_ABST
Patent Text Reader

Abstract

The invention relates to a fighter bottom layer flight control method based on deep reinforcement learning. The method comprises the following steps: S1, designing a target updating mode and a termination condition function of flight control according to a simulation environment; s2, constructing a state space, an action space and a reward function controlled by the fighter according to the simulation environment; s3, building a deep neural network model based on a PPO algorithm; and S4, training a fighter bottom layer flight control model based on a PPO algorithm. According to the control method disclosed by the invention, the intelligent agent only needs to concentrate on advanced task planning when completing high-level task decisions such as battle and tracking, so that the training of basic decisions for bottom-layer aircraft control is reduced, the algorithm is easier to converge, and the learning cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a control method, in particular to a bottom-layer flight control method for a fighter based on deep reinforcement learning, belonging to the technical field of fighter flight control. Background Art

[0002] As a complex electromechanical system, a fighter relies on a flight control system (flight control system) to achieve the dynamic conversion between bottom-layer control commands (such as ailerons, elevators, rudders, and throttles) and flight trajectories, so as to ensure the stability and maneuverability of the fighter in a changeable and high-load air combat environment. Since the response of the flight control system of a fighter to input commands is limited by multiple factors such as flight dynamics characteristics, aerodynamic layout, and the physiological endurance of pilots, the operation difficulty is extremely high, and pilots need to undergo long-term training to master it. This also leads to a high dependence of traditional fighter control operations on manual operations, posing a severe challenge to the man-machine cooperation efficiency in within-visual-range (WVR) or beyond-visual-range (BVR) air combat missions.

[0003] With the rapid development of artificial intelligence (AI) and deep learning technologies, the application potential of intelligent agents in air combat has been widely studied. Especially in the scenarios of unmanned aerial vehicles and manned-unmanned cooperative combat, the application of AI intelligent agents can significantly reduce the burden on pilots and improve combat efficiency. In practical applications, intelligent agents can be used to take over the bottom-layer flight control of a fighter, enabling pilots to concentrate on higher-level tactical decisions; or perform multi-aircraft cooperative tasks in cluster flight, such as reconnaissance, escort, and luring. These potential applications have greatly promoted the research progress of air combat intelligent agents. However, due to the high-precision simulation in the air combat simulation environment (including the flight dynamics characteristics of a fighter, the reaction model of the flight control system, etc.) which consumes a large amount of computing resources, it is difficult for the deep reinforcement learning training of intelligent agents to iterate quickly, resulting in limited algorithm convergence speed.

[0004] Current air combat intelligent agent training methods usually divide the tasks of intelligent agents into two modules: high-level decision-making and bottom-layer control, in order to simplify the training process and improve learning efficiency. In this framework, the high-level module focuses on air combat strategies and tactical choices, and the bottom-layer module is responsible for basic flight control. However, this modular method faces many technical challenges in the implementation process. In particular, due to the high-dynamic flight characteristics of a fighter (such as high speed, instantaneous acceleration, and turning, etc.), intelligent agents are prone to problems such as multiple stalls and multiple jitters in bottom-layer flight control, seriously affecting flight stability, and further causing the high-level decision-making module to be difficult to focus on high-level tasks. In addition, traditional deep reinforcement learning algorithms are difficult to achieve fast convergence when facing large-scale distributed interaction sampling requirements, and the final performance of the algorithms often fails to meet the requirements of air combat applications.

[0005] To solve this problem, existing research usually optimizes the flight control process through two technical measures: one is to optimize the reward mechanism of the agent so that it can quickly obtain effective feedback after completing specific actions to accelerate the learning process; the other is to divide the learning process of the agent into multiple stages, first complete basic actions such as takeoff and cruise, and then gradually learn more complex air combat maneuver actions. However, this phased learning method still faces the problem of poor convergence when encountering dynamic and difficult flight tasks (such as high-maneuver actions of fighter jets under high-speed and high-overload conditions), resulting in poor control stability of the agent in complex environments. Therefore, there is an urgent need for a new solution to solve this technical problem. Summary of the Invention

[0006] The present invention precisely aims at the technical problems existing in the prior art and provides a fighter aircraft bottom-layer flight control method based on deep reinforcement learning. This solution solves the problems of multi-stall and multi-jitter that are prone to occur in the air combat agent in a high-dynamic environment. This method targets speed and angle through a deep reinforcement learning model to finely control the bottom-layer control system of the fighter aircraft, so as to reduce the stall and jitter problems caused by complex environmental changes in the bottom-layer control of the air combat agent. By optimizing the bottom-layer flight control module, the control method of the present invention enables the agent to focus only on high-level task planning when completing high-level task decisions such as combat and tracking, thereby reducing the training of basic decisions for bottom-layer aircraft control, making the algorithm easier to converge, and reducing the learning cost.

[0007] To achieve the above object, the technical solution of the present invention is as follows. A fighter aircraft bottom-layer flight control method based on deep reinforcement learning, the method includes the following steps:

[0008] S1. Design the target update method and termination condition function of the flight control according to the simulation environment;

[0009] S2. Construct the state space, action space, and reward function of the fighter aircraft control according to the simulation environment;

[0010] S3. Build a deep neural network model based on the PPO algorithm;

[0011] S4. Train the fighter aircraft bottom-layer flight control model based on the PPO algorithm.

[0012] Among them, in step S1, when designing the target update method and termination condition function of the flight control according to the simulation environment, the specific steps are as follows:

[0013] S11. Define the origin of the body coordinate system as the center of mass of the aircraft, and build the kinematic equation and dynamic equation of the fighter aircraft trajectory motion:

[0014]

[0015] Wherein, X, Y, and Z are the positions of the aircraft in the ground coordinate system, u, v, and w are the components of the aircraft's airspeed in the body coordinate system, θ, φ are the pitch angle, yaw angle, and roll angle of the aircraft,

[0016]

[0017] Wherein, p, q, and r are the roll angular velocity, pitch angular velocity, and yaw angular velocity of the aircraft, F x 、F y 、F z are the components of the resultant force in the body coordinate system, g is the acceleration due to gravity, m is the mass of the aircraft,

[0018] S12. Design the target update method for flight control training,

[0019] This method is used to initialize the state parameters related to the target and set the new flight target state when resetting the flight target in the flight control system, so as to ensure that the new target setting is not affected by the previous state. The specific steps are as follows:

[0020] (1) Reset the state flag variable related to the flight target, and randomly generate a three-dimensional deviation vector for speed, pitch angle, and yaw angle, with the range of each component being [-1, 1],

[0021] (2) Receive the state information corresponding to the current fighter aircraft, store it in the current state vector, decode the generated deviation vector, and convert it into the actual target change range,

[0022] (3) Add the decoded deviation value to the current state value to generate a new target state, and constrain the upper and lower limits of the target value to ensure that the set target value is within the controllable range,

[0023] (4) Calculate the difference between the newly set target state and the current state to obtain the deviation state, which will ultimately be part of the state space,

[0024] S13. Design the termination condition function for flight control training,

[0025] Through the termination condition setting of the present invention, the system can effectively monitor the achievement state during the flight control training process and provide timely reward or punishment signals, as follows:

[0026] (1) Achievement condition check: According to the difference between the current state and the target, judge whether the set speed and angle targets are achieved, compare the difference between the current state and the target state with the threshold value. If the difference is less than the target threshold value, mark the target as achieved. If all speed and angle targets are achieved, mark this unit as the "all achieved" state,

[0027] (2) Achievement Marking and Recording: Based on the achievement status of each unit, update the internal marking variables and counters to record the number of steps achieved each time, the number of achievement times, and the consecutive steps of target achievement. If the target is continuously achieved, update the counter; if the target is not continuously achieved, record the number of failure times.

[0028] (3) Termination Condition Judgment: If the target is continuously achieved 5 times, mark the task as completed and trigger a reward.

[0029] If the current flight control steps reach the preset step limit, trigger the "timeout" state; if the current unit does not meet the target state (such as too low speed, too low or too high flight altitude, etc.), the task is judged to fail. Whether it is timeout, failure or success, the termination flag is always non-terminated.

[0030] Among them, in step S2, construct the state space, action space, and reward function of the aircraft control according to the simulation environment. The specific steps are as follows:

[0031] S21. Define the flight control state space of the fighter aircraft.

[0032] When training flight control in the simulation environment, the flight posture of the fighter aircraft can be described by its own attitude, the attitude relationship with the target, the speed relationship, etc. The specific description is as follows:

[0033]

[0034] Among them, V x,y,z is the three-way speed of the fighter aircraft in the north-east-earth coordinate system; is the pitch angle, yaw angle, and roll angle of the fighter aircraft; ω p,q,r is the roll angular velocity, pitch angular velocity, and yaw angular velocity of the fighter aircraft; H is the altitude of the fighter aircraft; is the difference between the target and the current speed scalar; is the difference between the target and the current speed pitch angle and yaw angle.

[0035] S22. Define the flight control action space of the fighter aircraft:

[0036] A = [W aileron , W elevator , W rudder , δ t

[0037] Among them, W aileron controls the roll of the aircraft by the fighter aircraft's control stick, W elevator controls the pitch of the aircraft by the fighter aircraft's control stick, W rudder controls the yaw of the aircraft by the fighter aircraft's control stick, δ T is the displacement of the fighter aircraft's throttle lever. ​

[0038] S23. Define the fighter flight control reward function:

[0039] The total reward function is composed of a weighted sum of the altitude reward function, roll reward function, angle reward function, speed reward function, time-consuming reward function, and safety reward function: r = ∑ω j r j , where r j is all the above reward items, and ω j is the weight coefficient of the response reward item.

[0040] Among them, the altitude, safety, and time-consuming reward functions are all event rewards, and the main reward functions are as follows:

[0041] The speed reward function is measured by the matching degree between the current flight speed and the target speed. The specific formula is as follows:

[0042] R v = exp(-w v Δv)

[0043] where Δv is the weighted sum of speed differences, including three parts: the difference between the current speed scalar and the target speed scalar; the difference between the current speed tilt angle and the target speed tilt angle; the difference between the current speed yaw angle and the target speed yaw angle. w v is the speed reward weight coefficient, which is used to control the punishment intensity

[0044] This reward function comprehensively evaluates the matching degree between the current speed of the fighter and the target speed, including the differences in speed magnitude, tilt angle, and yaw angle. Compared with the consideration of a single speed index in the prior art, this method greatly improves the accuracy of speed control, enabling the intelligent agent to approach the target flight state more efficiently under complex flight conditions. Through this reward function, the accurate control ability of the fighter for the speed state can be significantly improved, making it closer to the target flight trajectory.

[0045] The core objective of the angle reward is to detect whether the flight attitude is abnormal and give punishment or reward based on the degree of abnormality. The specific formula is as follows:

[0046]

[0047] where w θ is the angle reward weight coefficient. When the absolute value of the roll angle θ is greater than 90 degrees, it indicates that the aircraft's attitude is abnormal and punishment needs to be imposed on it. Otherwise, it is normal and the reward is 0. The angle reward detects whether the flight attitude exceeds the safety threshold, and real-time identification and intervention are carried out on potential dangerous states. This punishment mechanism further strengthens the safety during the flight process. Compared with the prior art, it more comprehensively avoids the occurrence of dangerous flight states and ensures the stability of the control model in a high-dynamic environment.

[0048] The roll reward term aims to suppress excessive changes in roll angular velocity, thereby reducing flight instability caused by excessive movements. Its formula is as follows:

[0049]

[0050] where w roll is the roll reward weight coefficient. If the roll angular velocity σ roll and the roll angular velocity σ' at the previous moment roll have inconsistent signs, it indicates that the aircraft attitude has abnormal jitter. At this time, a penalty is given. By penalizing the drastic change in roll angular velocity, the interference of action jitter on flight attitude can be effectively avoided.

[0051] Real-time detection of attitude jitter is added to the roll reward. By judging the change trend of roll angular velocity, abnormal jitter is punished in a timely manner. Such a design effectively reduces the attitude jitter problem during flight, significantly improving the control stability and execution accuracy. Compared with traditional methods that ignore the details of roll dynamics, the present invention pays more attention to the dynamic change process, enhancing the sensitivity and correction ability of the model to attitude abnormalities.

[0052] The present invention realizes the upgrade from sparse reward to multi-dimensional immediate reward in the design of the reward function, significantly improving the training efficiency and convergence speed of the model. Through the penalty design of the roll reward and angle reward functions, the model can detect and correct flight attitude abnormalities in real time, thereby enhancing the control stability and safety. At the same time, through comprehensive consideration of multiple aspects such as speed, attitude, and safety, the technical solution has much better robustness and adaptability in complex air combat environments than traditional methods, fully reflecting its innovation and practical value.

[0053] Among them, in step S3, the policy network of the PPO algorithm is used to select actions, and the policy is optimized through the reward feedback of the action-state pair. The steps are as follows:

[0054] The PPO algorithm model consists of an experience pool and two neural networks, including a policy Actor network and a value Critic network. The parameters of the two networks are θ and φ respectively.

[0055] S31. The experience pool is used to store the experience samples generated by the interaction between the UAV and the environment during the training process. These experience samples are stored in the buffer for subsequent training use, which can improve the independence, sample efficiency, generalization performance, and stability of the training.

[0056] S32. The input of the policy Actor network is the state S t , and the output is the action probability distribution π(a t |s t) The optimization goal is that the better the action probability distribution output by the policy Actor network can increase the value of the policy objective, the better the policy effect. The PPO algorithm limits the update amplitude by introducing a clipping mechanism to ensure training stability. The objective function of policy optimization is defined as a clipped loss function:

[0057]

[0058] Among them is the policy probability ratio, measuring the change ratio between the old and new policies; is the advantage function, used to evaluate the relative advantages and disadvantages of the current action; ε is the clipping hyperparameter, controlling the policy update amplitude to avoid drastic policy changes.

[0059] S33. The input of the value Critic network is the state, and the output is the estimated state value V(s). The Critic network is optimized by minimizing the value error, enabling the model to more accurately evaluate the value of the current state. The mean squared error loss function is used to define the loss function of the Critic network:

[0060]

[0061] Among them, V φ (s t ) is the estimated value of the state S t by the Critic network; R t is the actual return value.

[0062] S34. The advantage function calculation uses the Generalized Advantage Estimation (GAE) method to calculate the advantage function to evaluate the advantages and disadvantages of each action relative to the current policy. The expression of GAE is:

[0063]

[0064] Among them, δ t = r t + γV(s t+1 ) - V(s t ) represents the TD error, γ is the discount factor, and λ is the decay factor.

[0065] By introducing the discount factor and decay factor, the GAE method can balance the relationship between long-term rewards and immediate feedback, effectively reducing the volatility and instability of advantage estimation in traditional methods. This smooth estimation mechanism makes the model more robust during training and can adapt to dynamic changes in complex environments faster.

[0066] By flexibly adjusting the temporal difference error, the importance of short-term and long-term rewards can be dynamically balanced in different scenarios. This feature is particularly suitable for high-dynamic scenarios in fighter flight control, such as high-g turns or complex flight missions, and significantly improves the adaptability and performance of the model in multi-task switching compared to traditional advantage estimation methods.

[0067] In addition, GAE can effectively alleviate the problem of insufficient sample utilization. In existing technologies, there is usually a high dependence on sampled samples, while GAE improves the utilization efficiency of samples through more efficient advantage evaluation, thereby reducing the demand for a large amount of computing resources during training.

[0068] Among them, in step S4, training the underlying flight control model of the fighter based on the PPO algorithm is as follows:

[0069] The specific steps for training the network are as follows:

[0070] S41. Set network training parameters: Randomly initialize the parameters of the policy Actor network and the value Critic network, and initialize the experience buffer pool for storing the experience samples generated during training.

[0071] S42. Initialize the fighter state, initialize the state parameters related to the target, and set a new flight target state.

[0072] S43. Input the current state into the policy Actor network to generate corresponding actions; execute the actions in the current state, obtain the new state and rewards, record the relevant training data, and store it in the experience buffer pool.

[0073] S44. Check the data volume in the experience buffer pool: When the amount of experience data is greater than the set mini-batch size, start training. Randomly sample a batch of samples from the experience buffer pool for updating the policy Actor network and the Critic network.

[0074] S45. Optimize the parameters of the Actor network through the clipping loss function defined by PPO; optimize the parameters of the Critic network through the mean squared error loss function; use the Adam optimizer for gradient update.

[0075] S46. Policy update: Check the policy update condition. If the condition is met, perform the clipped update of the policy; if not, keep the current policy unchanged.

[0076] S47. Check whether the end-of-episode condition is reached: If it has been reached, execute step (47); otherwise, return to step (43) to continue a new round of state-action interaction.

[0077] S48. Training termination condition: Determine whether the algorithm model has reached the preset maximum number of training rounds. If it has, end the training and save the main policy model after the training ends. If not, return to step (42) to continue the training.

[0078] First, in the training initialization stage (S41, S42) of the present invention, by randomly initializing network parameters and designing an efficient experience cache pool, the diversity and utilization efficiency of training samples are significantly improved. Compared with the prior art that relies on a single initialization method, this mechanism effectively enhances the generalization ability of the model, enabling it to adapt to more complex flight environments.

[0079] Secondly, in the state and action interaction link (S43, S44), the present invention strengthens the use of the experience cache pool. By dynamically sampling small batches of data and optimizing the network training process, the efficient utilization of samples is achieved. At the same time, the step-by-step optimized policy network and value network (Actor and Critic networks) are adopted to ensure that the model can be quickly updated in each iteration, significantly improving the convergence speed.

[0080] Furthermore, through the combination of the PPO clipped loss function (S45) and the Adam optimizer, the stability and efficiency of policy optimization are improved. Compared with the problems of too fast or too slow policy updates that are prone to occur in traditional methods, the present invention introduces a clipping mechanism to avoid drastic fluctuations in the policy network, ensuring a more stable training process.

[0081] In the policy update and episode check link (S46, S47), by flexibly setting the update conditions and episode end conditions, the present invention can dynamically adjust the training process, avoiding resource waste caused by overtraining or performance degradation caused by insufficient training. At the same time, the added multi-round loop mechanism provides guarantee for the long-term optimization of the model, enabling it to continuously improve performance in complex task scenarios.

[0082] Finally, in the setting of the termination condition (S48), the present invention adopts a clear training round limit and model saving mechanism. This design avoids the occurrence of ineffective training and ensures that the optimal policy model can be output at the end of the training. Compared with the termination conditions of the prior art that are too simple or lack optimization, it shows higher efficiency and reliability.

[0083] In summary, by improving the training steps, the present invention significantly improves the convergence speed, adaptability, and robustness of the PPO algorithm in the fighter aircraft's underlying flight control model, enabling it to exhibit excellent control capabilities in high-dynamic complex environments.

[0084] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the described fighter aircraft low-level flight control method based on deep reinforcement learning.

[0085] A computer-readable storage medium stores computer instructions thereon. When the computer instructions are executed by a processor, they implement the described fighter aircraft low-level flight control method based on deep reinforcement learning.

[0086] The fighter aircraft low-level flight control method based on deep reinforcement learning proposed by the present invention significantly improves the flight control stability and convergence speed of the intelligent agent in a high-dynamic environment. By performing refined control on the fighter aircraft low-level flight control system, especially in a high-dynamic and high-load air combat environment, the method of the present invention exhibits many advantages:

[0087] (1) Improve flight control stability,

[0088] In traditional methods, when the intelligent agent executes flight tasks under high-speed and high-overload conditions, problems such as stall and jitter often occur due to the complex environment. By taking speed and angle as control targets and combining an improved reward mechanism, the present invention enables the intelligent agent to more precisely regulate the attitude and speed of the fighter aircraft. This method effectively reduces the multiple stall and jitter phenomena caused by sudden attitude changes, greatly improving the flight control stability and ensuring that the intelligent agent can maintain stable operation under high-speed turning and dynamic condition changes.

[0089] (2) Optimize the training efficiency of the intelligent agent,

[0090] Traditional deep reinforcement learning methods are difficult to meet the requirements of air combat applications in a large-scale distributed interaction environment due to their slow convergence speed. The present invention adopts the PPO algorithm, optimizes the training process of the policy network through the experience pool mechanism and the clipped loss function, improves the sample utilization rate of the training process, and ensures the independence and stability of the training. At the same time, the Generalized Advantage Estimation (GAE) method is used to accelerate the calculation of the advantage function, enabling the intelligent agent to obtain training feedback faster, ultimately achieving more efficient learning convergence and reducing the training cost.

[0091] (3) Enhance the fault tolerance of the control system,

[0092] By setting the target update method and the termination condition function, the present invention increases the robustness of the control system. For example, the system can continuously monitor the difference between the flight state and the target state of the fighter aircraft and trigger corresponding reward or punishment signals in a timely manner to prompt the intelligent agent to correct the deviation as soon as possible. This function effectively enhances the fault tolerance of the intelligent agent in a complex and dynamic environment, enabling it to adapt to the changing air combat environment and ensuring the safety during the control process.

[0093] (4) The modular design improves the system adaptability.

[0094] The underlying control method of the present invention can operate independently of the high-level tactical decision-making module. The two achieve flight control and tactical selection respectively through separate designs. In this way, even in the case of changing tactical requirements, the underlying flight control module can still maintain the stability and accuracy of its control. At the same time, this method can be reused in different fighter simulation environments and has good adaptability and generality. Description of the Drawings

[0095] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0096] Figure 2 It is a schematic diagram of the specific process of the training network. Detailed Implementation Manner

[0097] To deepen the understanding of the present invention, the following will make a detailed description of this embodiment in conjunction with the drawings.

[0098] Embodiment 1: Refer to Figure 1 、 Figure 2 , a fighter underlying flight control method based on deep reinforcement learning, the method includes the following steps:

[0099] S1. Design the target update method and termination condition function of the flight control according to the simulation environment;

[0100] S2. Construct the state space, action space and reward function of the fighter control according to the simulation environment;

[0101] S3. Build a deep neural network model based on the PPO algorithm;

[0102] S4. Train the fighter underlying flight control model based on the PPO algorithm.

[0103] Among them, in step S1, the target update method and termination condition function of the flight control are designed according to the simulation environment, and the specific steps are as follows:

[0104] S11. Define the origin of the body coordinate system as the center of mass of the aircraft, and build the kinematic equation and dynamic equation of the fighter's trajectory movement:

[0105]

[0106] Among them, X, Y, and Z are the positions of the aircraft in the ground coordinate system, u, v, and w are the components of the aircraft airspeed on the body coordinate system, θ, φ are the pitch angle, yaw angle, and roll angle of the aircraft,

[0107]

[0108] where p, q, and r are the roll angular velocity, pitch angular velocity, and yaw angular velocity of the aircraft, F x , F y , F z are the components of the resultant force in the body coordinate system, g is the acceleration due to gravity, and m is the mass of the aircraft,

[0109] S12. Design the target update method for flight control training,

[0110] This method is used to initialize the state parameters related to the target and set the new flight target state when resetting the flight target in the flight control system, so as to ensure that the new target setting is not affected by the previous state. The specific steps are as follows:

[0111] (1) Reset the state flag variable related to the flight target and randomly generate a three-dimensional deviation vector for velocity, pitch angle, and yaw angle. The range of each component is [-1, 1],

[0112] (2) Receive the state information corresponding to the current fighter and store it in the current state vector. Decode the generated deviation vector and convert it into the actual target change range,

[0113] (3) Add the decoded deviation value to the current state value to generate a new target state, and constrain the upper and lower limits of the target value to ensure that the set target value is within the controllable range,

[0114] (4) Calculate the difference between the newly set target state and the current state to obtain the deviation state, which will ultimately be part of the state space,

[0115] S13. Design the termination condition function for flight control training,

[0116] Through the termination condition setting of the present invention, the system can effectively monitor the achievement state during the flight control training process and provide timely reward or punishment signals, as follows:

[0117] (1) Achievement condition check: Based on the difference between the current state and the target, determine whether the set speed and angle targets are achieved. Compare the difference between the current state and the target state with the threshold. If the difference is less than the target threshold, mark the target as achieved. If all speed and angle targets are achieved, mark the unit as the "fully achieved" state,

[0118] (2) Achievement marking and recording: Based on the achievement state of each unit, update the internal marking variable and counter to record the number of steps achieved each time, the number of achievement times, and the continuous number of steps of target achievement. If the target is continuously achieved, update the counter; if the target is not continuously achieved, record the number of failure times of achievement,

[0119] (3) Termination condition judgment: If the target is continuously achieved 5 times, mark the task as completed and trigger a reward.

[0120] If the current flight control step reaches the preset step upper limit, trigger the "timeout" state; if the current unit does not meet the target state (such as too low speed, too low or too high flight altitude, etc.), the task is determined to fail. Whether it is timeout, failure or success, the termination flag is always non-terminated.

[0121] Among them, in step S2, construct the state space, action space and reward function for aircraft control according to the simulation environment. The specific steps are as follows:

[0122] S21. Define the flight control state space of the fighter plane.

[0123] When training flight control in the simulation environment, the flight posture of the fighter plane can be described by its own attitude, the relative posture relationship with the target, the speed relationship, etc. The specific description is as follows:

[0124]

[0125] Among them, V x,y,z is the three-way speed of the fighter plane in the north-east-earth coordinate system; is the pitch angle, yaw angle, roll angle of the fighter plane; ω p,q,r is the roll angular velocity, pitch angular velocity, yaw angular velocity of the fighter plane; H is the altitude of the fighter plane; is the difference between the target and the current speed scalar; is the difference between the target and the pitch angle and yaw angle of the current speed,

[0126] S22. Define the flight control action space of the fighter plane:

[0127] A = [W aileron , W elevator , W rudder , δ t

[0128] Among them, W aileron is the roll of the fighter plane controlled by the joystick, W elevator is the pitch of the fighter plane controlled by the joystick, W rudder is the yaw of the fighter plane controlled by the joystick, δ T is the displacement of the throttle lever of the fighter plane,

[0129] S23. Define the flight control reward function of the fighter plane:

[0130] The total reward function is composed of a weighted sum of the altitude reward function, roll reward function, angle reward function, speed reward function, time-consuming reward function, and safety reward function: r = ∑ω j r​j , where r j is all the above reward items, and ω j is the weight coefficient of the response reward item.

[0131] Among them, in step S3, the policy network using the PPO algorithm is used to select actions, and the policy is optimized through the reward feedback of the action-state pair. The steps are as follows:

[0132] The PPO algorithm model consists of an experience pool and two neural networks, including a policy Actor network and a value Critic network. The parameters of the two networks are θ and φ respectively.

[0133] S31. The experience pool is used to store the experience samples generated by the interaction between the drone and the environment during the training process. These experience samples are stored in the buffer for subsequent training use, which can improve the independence, sample efficiency, generalization performance and stability of the training.

[0134] S32. The input of the policy Actor network is the state S t , and the output is the action probability distribution π(a t |s t ). The optimization goal is that the better the action probability distribution output by the policy Actor network can increase the value of the policy objective, the better the policy effect. The PPO algorithm limits the update amplitude by introducing a clipping mechanism to ensure the training stability. The objective function of the policy optimization is defined as the clipped loss function:

[0135]

[0136] where is the policy probability ratio, which measures the change ratio between the old and new policies; is the advantage function, which is used to evaluate the relative pros and cons of the current action; ε is the clipping hyperparameter, which controls the policy update amplitude to avoid drastic policy changes.

[0137] S33. The input of the value Critic network is the state, and the output is the estimated state value V(s). The Critic network is optimized by minimizing the value error, so that the model can more accurately evaluate the value of the current state. The mean squared error loss function is used to define the loss function of the Critic network:

[0138]

[0139] where, V φ (s t ) is the estimated value of the state S t by the Critic network; R t is the actual return value.

[0140] S34. The advantage function calculation uses the Generalized Advantage Estimation (GAE) method to calculate the advantage function to evaluate the quality of each action relative to the current policy. The expression of GAE is:

[0141]

[0142] where, δ t = r t + γV(s t+1 ) - V(s t ) represents the TD error, γ is the discount factor, and λ is the decay factor.

[0143] Among them, in step S4, train the underlying flight control model of the fighter based on the PPO algorithm. The steps are as follows:

[0144] See Figure 2 , the specific steps of training the network are as follows:

[0145] S41. Set the network training parameters: Randomly initialize the parameters of the policy Actor network and the value Critic network, and initialize the experience buffer pool for storing the experience samples generated during the training process.

[0146] S42. Initialize the fighter state, initialize the state parameters related to the target, and set a new flight target state.

[0147] S43. Input the current state into the policy Actor network to generate the corresponding action; execute the action in the current state, obtain the new state and reward, record the relevant training data, and store it in the experience buffer pool.

[0148] S44. Check the data volume of the experience buffer pool: When the experience data volume is greater than the set mini - batch size, start training. Randomly sample a batch of samples from the experience buffer pool for updating the policy Actor network and the Critic network.

[0149] S45. Optimize the parameters of the Actor network through the clipping loss function defined by PPO; optimize the parameters of the Critic network through the mean - squared error loss function; use the Adam optimizer for gradient update.

[0150] S46. Policy update: Check the policy update condition. If the condition is met, perform the clipped update of the policy; if not, keep the current policy unchanged.

[0151] S47. Check whether the end - of - episode condition is reached: If it has been reached, execute step (47); otherwise, return to step (43) to continue a new round of state - action interaction.

[0152] S48. Training termination condition: Determine whether the algorithm model has reached the preset maximum number of training rounds. If it has, end the training and save the main policy model after the training ends. If not, return to step (42) to continue the training.

[0153] It should be noted that the above embodiments are not used to limit the protection scope of the present invention. Equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. A fighter jet low-level flight control method based on deep reinforcement learning, characterized in that: The method comprises the following steps: S1. Design the target update method and termination condition function of the flight control according to the simulation environment; S2, construct the state space, action space and reward function of fighter control according to the simulation environment; S3. Build a deep neural network model based on the PP0 algorithm; S4. Train the fighter jet underlying flight control model based on the PP0 algorithm.

2. The fighter jet low-level flight control method based on deep reinforcement learning according to claim 1 is characterized in that: In step S1, the target update method and termination condition function of the flight control are designed according to the simulation environment. The specific steps are as follows: S11. Define the origin of the body coordinate system as the center of mass of the aircraft, and build the kinematic equation and dynamic equation of the fighter's flight path motion: in, is the position of the aircraft in the ground coordinate system, u, υ, w are the velocity components of the aircraft airspeed in the body coordinate system, are the pitch angle, yaw angle, and roll angle of the aircraft, in, is the component of the aircraft in the body coordinate system, p, q, r are the roll angular velocity, pitch angular velocity, and yaw angular velocity of the aircraft, and F x 、F y 、F z is the component of the resultant force in the body coordinate system, g is the acceleration of gravity, m is the mass of the aircraft, S12. Design a target update method for flight control training. The specific steps are as follows: (1) Reset the state flag variables related to the flight target and randomly generate a three-dimensional deviation vector for speed, pitch angle and yaw angle, with each component ranging from [-1, 1], (2) Receive the state information corresponding to the current fighter, store it in the current state vector, decode the generated deviation vector, and convert it into the actual target change range, (3) Add the decoded deviation value to the current state value to generate a new target state, and set upper and lower limits on the target value to ensure that the set target value is within a controllable range. (4) Calculate the difference between the new target state and the current state to obtain the deviation state, which will eventually be used as part of the state space. S13. Design the termination condition function of flight control training. The details are as follows: (1) Achievement condition check: Based on the difference between the current state and the target, determine whether the set speed and angle targets are achieved. Compare the difference between the current state and the target state with the threshold. If the difference is less than the target threshold, mark the target as achieved. If all speed and angle targets are achieved, mark the unit as "all achieved". (2) Achievement marking and recording: Based on the achievement status of each unit, the internal marking variables and counters are updated to record the number of steps achieved each time, the number of times the goal has been achieved, and the number of consecutive steps to achieve the goal. If the goal has been achieved continuously, the counter is updated; if the goal has not been achieved continuously, the number of failed steps to achieve the goal is recorded. (3) Termination condition judgment: If the goal is achieved five times in a row, the task is marked as completed and the reward is triggered. If the current flight control step number reaches the preset step limit, the "timeout" state is triggered; if the current unit does not meet the target state (including too low speed, too low or too high flight altitude), the mission is judged to have failed. Regardless of timeout, failure or success, the termination flag will always be not terminated.

3. The fighter jet low-level flight control method based on deep reinforcement learning according to claim 2 is characterized in that: In step S2, the state space, action space and reward function of the aircraft control are constructed according to the simulation environment. The specific steps are as follows: S21. Define the fighter flight control state space. When training flight control in a simulation environment, the flight situation of a fighter jet can be described by its own posture, its situation relationship with the target, and its speed relationship, as shown below: Among them, V x,y,z is the three-dimensional speed of the fighter in the north-east coordinate system; are the pitch angle, yaw angle and roll angle of the fighter; ω p,q,r are the fighter's roll angular velocity, pitch angular velocity, and yaw angular velocity; H is the fighter's altitude; is the difference between the target and current speed scalars; is the difference between the target and current speed pitch angle and yaw angle, S22. Define the fighter flight control action space: A=[W aileron ,W elevator ,W rudder ,δ t ] Among them, W aileron The fighter's joystick controls the aircraft's roll, W elevator The fighter's joystick controls the aircraft's pitch. rudder The fighter's joystick controls the aircraft's yaw, δ T is the fighter throttle lever displacement, S23. Define the fighter flight control reward function: The total reward function is composed of the weighted height reward function, roll reward function, angle reward function, speed reward function, time reward function, and safety reward function: r = ∑ω j r j , where r j For all the above bonus items, ω j is the weight coefficient of the response reward item.

4. The fighter jet low-level flight control method based on deep reinforcement learning according to claim 3 is characterized in that: In step S3, the policy network of the PPO algorithm is used to select actions, and the policy is optimized by reward feedback for action-state pairs. The steps are as follows: The PPO algorithm model consists of an experience pool and two neural networks, including a strategy Actor network and a value Critic network. The parameters of the two networks are θ and φ respectively. S31, the experience pool is used to store the experience samples generated by the interaction between the drone and the environment during the training process, and store these experience samples in the buffer for subsequent training. S32, the input of the strategy Actor network is state S t , the output is the action probability distribution π(a t |s t ), the optimization goal is that the more the action probability distribution output by the policy Actor network can increase the value of the policy goal, the better the policy effect. The PPO algorithm limits the update amplitude by introducing a clipping mechanism to ensure training stability. The objective function of policy optimization is defined as the clipped loss function: in is the strategy probability ratio, which measures the change ratio between the new and old strategies; is the advantage function, which is used to evaluate the relative merits of the current action; ε is the clipping hyperparameter, which controls the strategy update amplitude to avoid drastic changes in strategy. S33. The input of the value critic network is the state, and the output is the estimated state value V(s). The critic network is optimized by minimizing the value error, so that the model can more accurately evaluate the value of the current state. The mean square error loss function is used to define the loss function of the critic network: Among them, V φ (s t ) is the Critic network for state S t The estimated value of R t is the actual return value, S34. Advantage function calculation is to use the generalized advantage estimation (GAE) method to calculate the advantage function To evaluate the pros and cons of each action relative to the current strategy, the expression of GAE is: Among them, δ t =r t +γV(s t+1 )-V(s t ) represents the TD error, γ is the discount factor, and λ is the attenuation factor.

5. The fighter jet low-level flight control method based on deep reinforcement learning according to claim 2 is characterized in that: In step S4, the fighter jet underlying flight control model based on the PPO algorithm is trained as follows: The specific steps for training the network are as follows: S41. Set network training parameters: randomly initialize the parameters of the strategy Actor network and the value Critic network, initialize the experience buffer pool to store the experience samples generated during the training process, and set the optimizer's learning rate, adaptive gradient clipping threshold and gradient penalty coefficient according to the task scenario to improve training stability. S42, initialize the fighter state, initialize the state parameters related to the target and set the new flight target state, and design a dynamic target adjustment strategy so that the target state can be dynamically adjusted according to the complexity of the environment during training to improve the generalization ability of the model. S43, input the current state into the strategy Actor network, generate the corresponding action; execute the action under the current state, obtain the new state and reward, record the relevant training data, and store it in the experience cache pool, while adding state transfer constraint rules to ensure the physical rationality of state transfer, S44, check the data volume of the experience buffer pool: when the amount of experience data is greater than the set small batch size, start training, randomly sample a batch of samples from the experience buffer pool, and use them to update the policy Actor network and Critic network. At the same time, use the priority experience replay mechanism to assign a higher sampling probability to samples related to higher rewards or scarce states. S45, optimize the parameters of the Actor network through the clipping loss function defined by PPO; optimize the parameters of the Critic network through the mean square error loss function; use the Adam optimizer for gradient update, and add the gradient entropy regularization term to prevent the model from falling into the local optimal solution, S46, strategy update: Check the strategy update conditions. If the conditions are met, the strategy is updated; if not, the current strategy is kept unchanged and strategy diversity constraints are introduced to avoid the problem of a single behavior pattern during the strategy update process. S47, check whether the round end condition is met: if it is met, execute step (47); otherwise, return to step (43) to continue a new round of state-action interaction. At the same time, increase the flight behavior diversity reward in this step to encourage the agent to explore more effective strategies. S48, training termination condition: determine whether the algorithm model has reached the preset maximum number of training rounds. If so, terminate the training and save the main strategy model after the training. If not, return to step (42) to continue training.

6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the fighter jet underlying flight control method based on deep reinforcement learning as described in any one of claims 1 to 5 above.

7. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by the processor, the underlying flight control method for a fighter jet based on deep reinforcement learning as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Multi-machine collaborative air combat planning method and system based on deep reinforcement learning

    CN112861442A

  • Unmanned aerial vehicle trajectory planning method based on continuous action dominant function learning

    CN116700327A

  • Air combat maneuvering strategy generation method based on example strategy constraint

    CN116796505A

  • Intelligent agent behavior learning method and system generated based on virtual simulation scene

    CN118485132A

  • Unmanned aerial vehicle autonomous navigation method based on improved TD3 algorithm

    CN118963407A

Cited By

  • Operation speed safety control method based on reinforcement learning hybrid control strategy

    CN119758838A

  • Traffic light control method and device for low-altitude takeoff and landing field

    CN120431773A

  • Adversarial reinforcement learning training method for robust control of fixed-wing aircraft

    CN121115529A