A deep reinforcement learning aircraft control method based on LLM guidance

By introducing a deep reinforcement learning method based on LLM guidance in aircraft control, a local text knowledge base is constructed and back-questioning strategy guidance is provided, the problems of low flexibility, low adaptability and uncertain exploration direction in existing aircraft control methods are solved, and more efficient and flexible aircraft control is achieved.

CN118034368BActive Publication Date: 2025-05-13SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410109451.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-05-13
Estimated Expiration
2044-01-25

AI Technical Summary

Technical Problem

The existing aircraft control methods have problems such as low control flexibility, inability to complete complex maneuvering actions, low adaptability, low robustness, and uncertain direction of agent learning and exploration in deep reinforcement learning training.

Method used

The deep reinforcement learning aircraft control method based on LLM guidance is adopted. By designing the state space and action space of the six-degree of freedom aircraft agent, a local text knowledge base for flight control is constructed, and the flight controller model is trained based on the LLM and local text knowledge base. The LLM is used to receive the information of the local text knowledge base and the state action text information generated by the interaction between the agent and the environment during the training process, and a backward questioning strategy is used to guide the flight control instructions of the agent.

Benefits of technology

It improves the flexibility and adaptability of aircraft control, can complete complex maneuvering actions, and improves robustness and training speed, solving the problem of uncertainty in the learning and exploration direction of agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118034368B_ABST
    Figure CN118034368B_ABST
Patent Text Reader

Abstract

The present invention discloses a deep reinforcement learning aircraft control method based on LLM guidance, which relates to the field of aircraft control technology, and includes the following steps: designing the state space and action space of a six-degree-of-freedom aircraft agent; constructing a local text knowledge base for flight control; designing a reward function to train a flight controller model; using LLM to receive information from the local text knowledge base and state-action text information generated by the interaction between the agent and the environment during the training process, and using a backward questioning strategy to guide the flight control instructions of the agent, while combining the guidance action with the flight control instruction to generate a new control instruction; repeating the above interaction process until the model converges. The present invention solves the problems of low control flexibility, inability to complete complex maneuvers, low adaptability, low robustness, and uncertainty in the learning and exploration direction of the agent in deep reinforcement learning training in existing aircraft control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of aircraft control technology, and in particular to a deep reinforcement learning aircraft control method based on LLM guidance. Background Art

[0002] As a complex information processing system, the flight controller is the bridge that connects tactical decision-making strategies with actual actions, and is the basis of intelligent air combat game decision-making. People have done a lot of research on flight controllers, which are mainly divided into model-based flight control methods and model-free flight control methods.

[0003] Model-based flight control methods include nonlinear model predictive control, linear quadratic Gaussian model control, backstepping control, sliding mode control, gain scheduling control, nonlinear neural network adaptive control, etc. Although model-based control methods can solve flight control problems to a certain extent and achieve corresponding control goals, model-based flight control systems are heavily dependent on the accuracy of the mathematical model assumed by the system. In fact, it is difficult to achieve perfect mathematical modeling for flight control systems, and there is even a lack of understanding, which is bound to affect the control effect.

[0004] Model-free flight control methods include fuzzy system control methods, data-driven control methods, neural network control methods, deep reinforcement learning methods, etc. Compared with model-based control methods, model-free control methods do not need to rely on accurate mathematical models, and they are adaptive and more robust when facing uncertain interference factors.

[0005] At present, aircraft control mainly uses deep reinforcement learning to achieve end-to-end six-degree-of-freedom control of the aircraft. However, there are problems such as low control flexibility, inability to complete complex maneuvers, low adaptability, low robustness, and uncertainty in the learning and exploration direction of the intelligent agent during deep reinforcement learning training. Summary of the invention

[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides a deep reinforcement learning aircraft control method based on LLM guidance, which solves the problems of low control flexibility, inability to complete complex maneuvers, low adaptability, low robustness, and uncertain learning and exploration direction of the intelligent agent in deep reinforcement learning training.

[0007] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a deep reinforcement learning aircraft control method based on LLM guidance, comprising the following steps:

[0008] S1: Design the state space and action space of the 6-DOF aircraft agent;

[0009] S2: Construct a local textual knowledge base of flight control based on the parameters of flight control tasks, state space and action space;

[0010] S3: Design a reward function based on the flight control task and train the flight controller model based on the LLM and local text knowledge base;

[0011] S4: Use LLM to receive information from the local text knowledge base and the state-action text information generated by the agent's interaction with the environment during training, and use the backward questioning strategy to guide the agent's flight control instructions, while combining the guidance actions and flight control instructions to generate new control instructions;

[0012] S5: Based on the new control instruction, determine whether the model has reached convergence. If yes, end the training and enter step S6. If no, return to step S4 and repeat the above interaction process.

[0013] S6: Test the trained model and complete the deep reinforcement learning aircraft control based on LLM guidance.

[0014] The beneficial effect of the above scheme is: the present invention takes into account the uncertainty problem of deep reinforcement learning exploration direction and the large motion space of six-degree-of-freedom aircraft, introduces LLM based on local knowledge base reasoning for action guidance, guides the control direction of the aircraft, improves the training speed, and solves the problems of low control flexibility, inability to complete complex maneuvers, low adaptability, low robustness, and uncertainty in the learning and exploration direction of the intelligent agent in deep reinforcement learning training.

[0015] Furthermore, the state s in the S1 state space t Obtained from the flight control environment, including the aircraft's flight status and the status of the aircraft's tracking target signal;

[0016] The aircraft flight status includes pitch angle, roll angle, yaw angle, altitude and speed;

[0017] The state of the aircraft tracking the target signal includes a roll angle error, a yaw angle error, an altitude error and a speed error.

[0018] The beneficial effect of the above further scheme is that in the design of the intelligent agent state space, the flight state of the aircraft and the state of the aircraft tracking target signal are taken into account at the same time for model training, which can not only realize direct control of the height and speed of the aircraft, but also increase the control of the aircraft attitude, and solve the problems of low flight control flexibility, inability to complete complex maneuvers, low adaptability and low robustness.

[0019] Furthermore, action a in the action space S1 tIncluding aileron control commands, throttle commands, elevator control commands and rudder control commands.

[0020] The beneficial effect of the above further solution is: by setting the above action instructions, the aircraft can be controlled to improve the flexibility of flight control.

[0021] Furthermore, the local text knowledge base in S2 includes flight targets and control instructions;

[0022] The flight targets include turning left, turning right, rolling left, rolling right, ascending altitude, descending altitude, accelerating and decelerating;

[0023] The control instructions include pressing the joystick right, pressing the joystick left, holding the joystick steady, pulling the joystick, pushing the joystick, refueling and reducing the fuel.

[0024] The beneficial effect of the above further scheme is that the present invention textualizes flight driving by constructing a local text knowledge base, and converts flight control instructions and flight status into natural language understandable to LLM.

[0025] Furthermore, the reward function R in S3 is:

[0026]

[0027] Among them, R φ is the roll angle error reward, is the yaw angle error reward, R alt is the height error reward, Reward for speed error;

[0028]

[0029]

[0030]

[0031]

[0032] Among them, Δ φ is the roll angle error, is the yaw angle error, Δ alt is the height error, is the speed error, σ φ is the roll angle error reward variance, is the yaw angle error reward variance, σ alt is the height error reward variance, is the speed error reward variance.

[0033] The beneficial effect of the above further scheme is: the reward function is a very important part in reinforcement learning. The present invention introduces roll angle error, yaw angle error, altitude error and speed error into training. In order to control the flight by controlling the error values ​​and tracking the target signal, the present application uses potential-based reward shaping (PBRS) technology to design a potential function consistent with the Gaussian function form, so that the overall effect of the flight controller is optimized, rather than being able to perform only a certain attitude control.

[0034] Furthermore, during the S4 training process, the agent interacts with the environment to generate state-action text information, which includes the following steps:

[0035] S4-1: Initialize the aircraft state, policy network actor, evaluation network critic and experience replay pool replaybuffer;

[0036] S4-2: Based on the initialization result, use the strategy network actor according to the current state s t Send action a t , and determine whether the current signal reaches the target signal. If yes, proceed to step S4-3. If not, the current round of tasks ends, and the advantage value is calculated and stored in the experience replay pool replay buffer as state action text information;

[0037] S4-3: Re-randomly generate the target signal {s t+1 ,a t ,R}, and store it in the experience replay pool replay buffer as state action text information, where s t+1 is the state at the next moment, and R is the reward function;

[0038] S4-4: Based on the maximum number of steps set for each round of tasks, continue to track the next target signal and re-randomly generate the target signal {s t+1 ,a t ,R} as the current signal and return to step S4-2.

[0039] The beneficial effect of the above further scheme is: the present invention adopts the PPO algorithm to train the flight controller, tracks the target signal, checks whether the target signal is reached within the specified time, and if the target signal is reached, the target signal will be randomly reinitialized, and the task is not completed at this time; if the target signal is not reached, the task is completed, and the target signal and advantage value generated in each round are stored in the experience replay pool. The experience replay pool is mainly used to store historical data. By using historical data for training, it is helpful to improve sample efficiency and reduce training instability caused by time series data.

[0040] Furthermore, the randomness of the target signal in S4-2 is set to gradient ascent.

[0041] The beneficial effect of the above further scheme is: considering that the action space of the six-degree-of-freedom aircraft is too large, the present invention sets the randomness of the target signal to gradient ascent, thereby improving the training speed and stability of the flight controller.

[0042] Furthermore, the advantage value in S4-2 for:

[0043]

[0044] Among them, G t is the action-state value function, and V is the value function.

[0045] The beneficial effect of the above further scheme is: through the above formula, using the Bellman equation, the advantage value at the end of each round is obtained and stored in the experience replay pool.

[0046] Furthermore, the termination signal Done of the task completion in S4-2 is:

[0047] Done=Done signal ∨Done G ∨Done Altitude ∨Done step

[0048] Among them, Done signal Is the termination signal of the target signal reached, Done G Is the termination signal of whether the maximum overload has been reached, Done Altitude Is the termination signal of whether the ceiling is reached, Done step is the termination signal of whether the maximum number of steps is exceeded, and ∨ is an OR operation;

[0049]

[0050]

[0051] Among them, G is overload, G max is the maximum overload;

[0052]

[0053] Among them, alt is the height, Altitude max is the ceiling;

[0054]

[0055] Among them, step is the number of steps, step max is the maximum number of steps.

[0056] The beneficial effect of the above further scheme is: through the above technical scheme, the termination conditions for the end of the mission are established to prevent the aircraft from endlessly tracking the next target signal. In addition, considering extreme flight conditions encountered during the flight, such as exceeding the maximum overload and ceiling, the mission will be terminated. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 The figure is a flow chart of a deep reinforcement learning aircraft control method based on LLM guidance.

[0058] Figure 2 A block diagram of deep reinforcement learning aircraft control based on LLM guidance. DETAILED DESCRIPTION

[0059] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0060] like Figure 1 As shown, a deep reinforcement learning aircraft control method based on LLM guidance includes the following steps:

[0061] S1: Design the state space and action space of the 6-DOF aircraft agent;

[0062] S2: Construct a local textual knowledge base of flight control based on the parameters of flight control tasks, state space and action space;

[0063] S3: Design a reward function based on the flight control task and train the flight controller model based on the LLM and local text knowledge base;

[0064] S4: Use LLM to receive information from the local text knowledge base and the state-action text information generated by the agent's interaction with the environment during training, and use the backward questioning strategy to guide the agent's flight control instructions, while combining the guidance actions and flight control instructions to generate new control instructions;

[0065] S5: Based on the new control instruction, determine whether the model has reached convergence. If yes, end the training and enter step S6. If no, return to step S4 and repeat the above interaction process.

[0066] S6: Test the trained model and complete the deep reinforcement learning aircraft control based on LLM guidance.

[0067] State s in the S1 state space t Obtained from the flight control environment, including the aircraft's flight status and the status of the aircraft's tracking target signal;

[0068] The aircraft flight status includes pitch angle, roll angle, yaw angle, altitude and speed;

[0069] The state of the aircraft tracking the target signal includes a roll angle error, a yaw angle error, an altitude error and a speed error.

[0070] Action a in S1 action space t It includes aileron control command, throttle command, elevator control command and rudder control command. Aileron, elevator and rudder can control the roll angle, pitch angle and yaw angle of the aircraft respectively.

[0071] The present invention textualizes flight driving by constructing a local text knowledge base, and uses LLM to conduct a backward questioning strategy based on the local text knowledge base in the early stage of training, thinking about the core of the problem step by step from a logical perspective, and judging whether the flight control actions made by the agent in the deep reinforcement learning training process are correct. A key to this technology is how to convert flight control instructions and flight status into natural language understandable to LLM, and at the same time convert LLM's text guidance information into digital information that can guide flight control instructions. The present invention tracks target signals to achieve control of the aircraft, so in this part, the corresponding text description of the target signal and other related states tracked by the aircraft and the flight control instructions made by the agent is constructed.

[0072] The local text knowledge base in S2 includes flight targets and control instructions;

[0073] The flight targets include turning left, turning right, rolling left, rolling right, ascending altitude, descending altitude, accelerating and decelerating;

[0074] The control instructions include pressing the joystick right, pressing the joystick left, holding the joystick steady, pulling the joystick, pushing the joystick, refueling and reducing the fuel.

[0075] In one embodiment of the present invention, the local text knowledge base can be described as:

[0076] Flight target: {turn left, turn right, roll left, roll right, ascend altitude, descend altitude, accelerate, decelerate};

[0077] Control instructions: {Press the joystick to the right, press the joystick to the left, hold the joystick steady, pull the joystick, push the joystick, add fuel, reduce fuel}.

[0078] In this embodiment, the current flight target of the aircraft is 1 goal , the control command of the aircraft is 1 command When f LLM (1 goal ,1 command) According to the flight target and control instructions, it will answer the following question: Should this operation be performed? LLM will answer "yes" or "no", and such answers can be easily converted into "1" or "0". Clear answers will be more conducive to guiding the flight control actions, and then combined with the control instructions of the intelligent body to obtain the guided actions, such as Figure 2 As shown in Figure 2, action guidance in the early stage of training can effectively guide the direction of the agent's exploration and speed up the training.

[0079] The reward function R in S3 is:

[0080]

[0081] Among them, R φ is the roll angle error reward. The smaller the absolute value of the roll angle error is, the greater the reward value is. is the yaw angle error reward, R alt is the height error reward, Reward for speed error;

[0082]

[0083]

[0084]

[0085]

[0086] Among them, Δ φ is the roll angle error, is the yaw angle error, Δ alt is the height error, is the speed error, σ φ is the roll angle error reward variance, is the yaw angle error reward variance, σ alt is the height error reward variance, is the speed error reward variance.

[0087] During the S4 training process, the agent interacts with the environment to generate state-action text information, including the following steps:

[0088] S4-1: Initialize the aircraft state, policy network actor, evaluation network critic and experience replay pool replaybuffer;

[0089] In one embodiment of the present invention, actor is a policy network, which is used to generate policies, input states, and output action instructions; critic is an evaluation network, which is used to evaluate the value of policies, input states, and output state values; replay buffer is an experience replay pool, whose purpose is to store historical data. By using historical data for training, it helps to improve sample efficiency and reduce training instability caused by time series data.

[0090] S4-2: Based on the initialization result, use the strategy network actor according to the current state s t Send action a t , and determine whether the current signal reaches the target signal. If yes, proceed to step S4-3. If not, the current round of tasks ends, and the advantage value is calculated and stored in the experience replay pool replay buffer as state action text information;

[0091] S4-3: Re-randomly generate the target signal {s t+1 ,a t ,R}, and store it in the experience replay pool replay buffer as state action text information, where s t+1 is the state at the next moment, and R is the reward function;

[0092] S4-4: Based on the maximum number of steps set for each round of tasks, continue to track the next target signal and re-randomly generate the target signal {s t+1 ,a t ,R} as the current signal and return to step S4-2.

[0093] In this embodiment, when the flight controller interacts with the environment to generate enough data, the actor and critic are optimized.

[0094] The randomness of the target signal in S4-2 is set to gradient ascent.

[0095] S4-2 medium advantage value for:

[0096]

[0097] Among them, G t is the action-state value function, and V is the value function.

[0098] The termination signal Done of the task in S4-2 is:

[0099] Done=Done signal ∨Done G ∨Done Altitude ∨Done step

[0100] Among them, Done signal Is the termination signal of the target signal reached, Done G Is the termination signal of whether the maximum overload has been reached, Done Altitude Is the termination signal of whether the ceiling is reached, Done step is the termination signal of whether the maximum number of steps is exceeded, and ∨ is an OR operation;

[0101]

[0102]

[0103] Among them, G is overload, G max is the maximum overload;

[0104]

[0105] Among them, alt is the height, Altitude max is the ceiling;

[0106]

[0107] Among them, step is the number of steps, step max is the maximum number of steps.

[0108] The present invention introduces roll angle error, yaw angle error, height error and speed error into training, and inputs corresponding errors into the flight controller by tracking the target roll angle, target yaw angle, target height and target speed, and outputs the commands of aileron, elevator, rudder and throttle to directly control the aircraft, which can not only realize the direct control of the height and speed of the aircraft, but also increase the control of the aircraft attitude, solve the problems of low flexibility of flight control, inability to complete complex maneuvers, low adaptability and low robustness. At the same time, considering the uncertainty of the exploration direction of deep reinforcement learning and the large action space of six-degree-of-freedom aircraft, LLM based on local knowledge base reasoning is introduced for action guidance, guiding the control direction of the aircraft, improving the training speed, and the flight controller finally obtained can be deployed offline, and has adaptability and strong robustness. LLM action guidance based on local knowledge base has the characteristics of being portable and scalable, and can guide the intelligent agent to explore in different control directions according to different tasks and scenarios, while accelerating the training speed of intelligence.

[0109] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the invention.

Claims

1. A deep reinforcement learning aircraft control method based on LLM guidance, characterized in that: The following steps are involved: S1: Design the state space and action space of the 6-DOF aircraft agent; S2: Construct a local textual knowledge base of flight control based on the parameters of flight control tasks, state space and action space; S3: Design a reward function based on the flight control task and train the flight controller model based on the LLM and local text knowledge base; S4: Use LLM to receive information from the local text knowledge base and the state-action text information generated by the agent's interaction with the environment during training, and use the backward questioning strategy to guide the agent's flight control instructions, while combining the guidance actions and flight control instructions to generate new control instructions; S5: Based on the new control instruction, determine whether the model has reached convergence. If yes, end the training and enter step S6. If no, return to step S4 and repeat the above interaction process. S6: Test the trained model and complete the deep reinforcement learning aircraft control based on LLM guidance; The reward function R in S3 is: Among them, R φ is the roll angle error reward, is the yaw angle error reward, R alt is the height error reward, Reward for speed error; Among them, Δ φ is the roll angle error, is the yaw angle error, Δ alt is the height error, is the speed error, σ φ is the roll angle error reward variance, is the yaw angle error reward variance, σ alt is the height error reward variance, is the speed error reward variance.

2. The deep reinforcement learning aircraft control method based on LLM guidance according to claim 1 is characterized in that: The state s in the S1 state space t Obtained from the flight control environment, including the aircraft's flight status and the status of the aircraft's tracking target signal; The aircraft flight status includes pitch angle, roll angle, yaw angle, altitude and speed; The state of the aircraft tracking the target signal includes a roll angle error, a yaw angle error, an altitude error and a speed error.

3. The LLM-guided deep reinforcement learning aircraft control method according to claim 1, characterized in that: Action a in the S1 action space t Including aileron control commands, throttle commands, elevator control commands and rudder control commands.

4. The LLM-guided deep reinforcement learning aircraft control method according to claim 1, characterized in that: The local text knowledge base in S2 includes flight targets and control instructions; The flight targets include turning left, turning right, rolling left, rolling right, ascending altitude, descending altitude, accelerating and decelerating; The control instructions include pressing the joystick right, pressing the joystick left, holding the joystick steady, pulling the joystick, pushing the joystick, refueling and reducing the fuel.

5. The LLM-guided deep reinforcement learning aircraft control method according to claim 1, characterized in that: The S4 training process involves the agent interacting with the environment to generate state-action text information, including the following steps: S4-1: Initialize the aircraft state, policy network actor, evaluation network critic and experience replay pool replay buffer; S4-2: Based on the initialization result, use the strategy network actor according to the current state s t Send action a t , and determine whether the current signal reaches the target signal. If yes, proceed to step S4-3. If not, the current round of tasks ends, and the advantage value is calculated and stored in the experience replay pool replay buffer as state action text information; S4-3: Re-randomly generate the target signal {s t+1 ,a t ,R}, and store it in the experience replay pool replay buffer as state action text information, where s t+1 is the state at the next moment, and R is the reward function; S4-4: Based on the maximum number of steps set for each round of tasks, continue to track the next target signal and re-randomly generate the target signal {s t+1 ,a t ,R} as the current signal and return to step S4-2.

6. The LLM-guided deep reinforcement learning aircraft control method according to claim 5, characterized in that: The randomness of the target signal in S4-2 is set to gradient ascent.

7. The LLM-guided deep reinforcement learning aircraft control method according to claim 5, characterized in that: The advantage value in S4-2 for: Among them, G t is the action-state value function, and V is the value function.

8. The LLM-guided deep reinforcement learning aircraft control method according to claim 5, characterized in that: The termination signal Done of the task in S4-2 is: Done=Done signal ∨Done G ∨Done Altitude ∨Done step Among them, Done signal Is the termination signal of the target signal reached, Done G Is the termination signal of whether the maximum overload has been reached, Done Altitude Is the termination signal of whether the ceiling is reached, Done step is the termination signal of whether the maximum number of steps is exceeded, and ∨ is an OR operation; Among them, G is overload, G max is the maximum overload; Among them, alt is the height, Altitude max is the ceiling; Among them, step is the number of steps, step max is the maximum number of steps.