A robust control method for morphing aircraft based on deep reinforcement learning

By using deep reinforcement learning, a CAD model and aerodynamic data of a mutated aircraft were established. The DQN algorithm was used for control, which solved the problems of autonomous deformation and stability of the mutated aircraft in complex environments, and improved the aerodynamic performance and control effect of the aircraft.

CN116560384BActive Publication Date: 2025-11-18TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310276318.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-21
Publication Date
2025-11-18
Estimated Expiration
2043-03-21

AI Technical Summary

Technical Problem

Variant aircraft struggle to achieve autonomous deformation control in complex environments. Traditional control methods cannot effectively characterize flight states, and the complexity and stability of the system are difficult to guarantee.

Method used

A robust control method based on deep reinforcement learning is adopted. The CAD model and aerodynamic data of the aircraft are established using CATIA and FLUENT software. The DQN algorithm is used for deformation and motion control, and an intelligent agent is constructed for decision-making.

Benefits of technology

Stable deformation control of morphing aircraft in complex environments has been achieved, adapting to multimodal and strong nonlinear characteristics, improving the aerodynamic performance and stability of the aircraft, and simplifying the design of control strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116560384B_ABST
    Figure CN116560384B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on deep reinforcement learning's variant aircraft robust control method, comprising the following steps: S1, the CAD model of variant aircraft is established, then simulation obtains aircraft aerodynamic data, according to the data obtained, the kinematics and dynamics equation of aircraft is solved, complete the physical model building of variant aircraft movement;S2, using DQN algorithm carries out the deep reinforcement learning of variant aircraft deformation and motion control, trains value function network;S3, according to the value function network of training intelligent agent is constructed, and the control of variant aircraft is made reasonable decision by the intelligent agent;The application can carry out deformation design, analysis and control to variant aircraft;In the reasonable case of task design, make aircraft can make reasonable deformation, adapt to the characteristics of system multimode, strong nonlinear and strong coupling brought by aircraft deformation process, optimal control strategy can be selected under complex control, ensure flight stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of mutated aircraft, and in particular to a robust control method for mutated aircraft based on deep reinforcement learning. Background Technology

[0002] With the development of flight technology, morphing aircraft have received increasing attention due to their significant military value. The concept of morphing aircraft was first proposed by German scientists during World War II, with the American Bell X-5 being the first variable-sweep wing aircraft. Later, aircraft such as the F-111 and F-14 adopted this technology; variable sweep angles can improve the aerodynamic characteristics of an aircraft. In recent years, morphing mechanisms such as folding wings, retractable wings, and variable noses have also been applied to aircraft. These variable devices can effectively improve the aerodynamic characteristics of aircraft, enabling rapid roll and lift maneuvers that are difficult for traditional aircraft to achieve, expanding the aircraft's flight envelope, and increasing its maximum speed.

[0003] Compared to traditional aircraft, morphing aircraft offer numerous advantages: First, they improve aerodynamic characteristics and reduce energy consumption through active deformation; second, they enhance control capabilities by using active deformation to assist maneuvering; and finally, they can adapt to various flight environments and missions by changing their configuration, expanding their application range. These advantages make morphing aircraft a promising foundation for breakthroughs in future high-performance aircraft development, possessing immense development potential and practical value in both military and civilian applications.

[0004] However, morphing aircraft are difficult to apply in large quantities in reality, for three main reasons:

[0005] I. High cost: Variant aircraft require transformation, which places high demands on hardware. The variable modules need to be stably connected to the airframe, and fragile materials cannot be used.

[0006] II. Greater Weight: To achieve transformation, morphing aircraft require the use of robust materials, which are often quite heavy. Furthermore, compared to traditional aircraft, morphing aircraft require the operation of systems such as motors, which adds to their weight.

[0007] III. System complexity: Large systems are difficult to plan using control algorithms.

[0008] Regarding the third point mentioned above, current research on mutator aircraft mainly focuses on mutation design and control. Currently, Q-learning-based switching control for mutator aircraft has been proposed; and reinforcement learning-based adaptive mutation strategies and flight control methods for mutator aircraft have also been proposed.

[0009] However, the existing technology has the following problems:

[0010] 1. This switching control method switches to the internal reinforcement learning controller when the aircraft's altitude tracking error is small. The Q-Learning method cannot fully characterize the aircraft's flight state.

[0011] 2. This method studies a variable-airfoil morphing aircraft. Using dive, cruise, and climb as three states and the sweep angle as the action, the optimal strategy for morphing is explored using Q-learning. However, this control method limits the morphing aircraft to only one configuration in each situation, which contradicts the expanded flight envelope proposed in morphing aircraft design, thus limiting its applicability.

[0012] While deformability enhances aircraft performance, it also presents new requirements and challenges for modeling and control: morphing aircraft need autonomous deformation capabilities in complex battlefield environments; the deformation process results in multimodal, highly nonlinear, and strongly coupled characteristics; and aircraft are highly susceptible to various internal and external disturbances during deformable flight, making flight stability difficult to guarantee. Traditional control system design methods for fixed-shape aircraft are no longer sufficient to meet the needs of morphing aircraft. Summary of the Invention

[0013] The purpose of this invention is to solve the modeling and deformation control problems of morphing aircraft and to provide a robust control method for morphing aircraft based on deep reinforcement learning.

[0014] To achieve the above objectives, the present invention adopts the following technical solution:

[0015] A robust control method for variant aircraft based on deep reinforcement learning includes the following steps:

[0016] S1. Establish a CAD model of the mutated aircraft, then simulate to obtain aerodynamic data of the aircraft. Based on the obtained data, solve the kinematic and dynamic equations of the aircraft to complete the physical model of the mutated aircraft's motion.

[0017] S2. Use the DQN (Deep Q-Learning) algorithm to perform deep reinforcement learning for the deformation and motion control of the morphing aircraft, and train the value function network;

[0018] S3. Construct an intelligent agent based on the trained value function network, and make reasonable decisions on the control of the variant aircraft through the intelligent agent.

[0019] In some embodiments of the present invention, in step S1, a CAD model of the variant aircraft is established using CATIA software, and then the aerodynamic data of the aircraft is obtained through simulation using FLUENT software.

[0020] In some embodiments of the present invention, in step S1, during modeling, the ailerons and engine of the morphing aircraft are ignored and regarded as a model that balances forces and maintains the direction of the aircraft's velocity. The wings and tail of the morphing aircraft have deformable parts. When the variable parts on the wings deform, they are equal in size and opposite in direction. When the variable parts on the tail deform, they are equal in size and in the same direction. Different actions are formed by the combination of wing deformation, horizontal tail deformation, and velocity change.

[0021] In some embodiments of the present invention, the deformable angles of the wings of the morphing aircraft are [-18°, -13.5°, -9°, -4.5°, 0°, 4.5°, 9°, 13.5°, 18°], and the deformable angles of the tail fin of the morphing aircraft are [0°, 3°, 6°, 9°, 12°, 15°]. There are a total of 30 different combinations of deformation angles.

[0022] In some embodiments of the present invention, in step S1, it is assumed that the velocity direction of the aircraft is always the same as the x-axis direction, that is, the trajectory coordinate system and the body coordinate system are in the same direction. According to different combinations of deformation angles, the FLUENT module is used to measure the force on the aircraft in the xyz axes and the sum of the torques of the three axes about the center of mass at various speeds.

[0023] The estimated moment of inertia J of the aircraft is as follows:

[0024]

[0025] Where W is the mass of the aircraft, and L a The total length of the aircraft Let be the dimensionless radius of gyration along the y-axis of the body coordinate system;

[0026] The aircraft's torque M, angular acceleration α, and angular velocity ω are:

[0027] M = J × a (2)

[0028]

[0029]

[0030] Where t is time;

[0031] The transformation matrices for the angles through which the original coordinate system rotates about the x and y axes are respectively:

[0032]

[0033]

[0034] in,

[0035]

[0036] When an aircraft is in flight, it is equivalent to first deforming the wings to rotate the aircraft around the x-axis, and then deforming the tail fins to change the aircraft's direction; the coordinates after rotation are: The coordinates before rotation are: If the aircraft first rotates by ξ radians around the x-axis and then by η radians around the y-axis, then:

[0037]

[0038] The projection [x',y',z'] of the x-axis direction in the track coordinate system onto the ground coordinate system is:

[0039]

[0040] The track deviation angle χ and the inclination angle γ are:

[0041]

[0042]

[0043] The velocity V of the aircraft in the ground coordinate system Ox g y g z g The three velocity components are:

[0044]

[0045]

[0046]

[0047] Integrating the three velocity components respectively, the position of the aircraft is determined, and a physical model of the motion of the variant aircraft is constructed based on the obtained data.

[0048] In some embodiments of the present invention, in step S2, when training the trajectory tracking of the variant aircraft, an encoder method is used to reduce the action space of the value function network, thereby accelerating the convergence speed of the deep learning network during training.

[0049] In some embodiments of the present invention, in step S2, the deep reinforcement learning includes: obtaining the target trajectory within one cycle as the task; the state of the aircraft is 14-dimensional, including the ratio of the target position to the current position of the aircraft, the aircraft's deflection angle, the aircraft's tilt angle, the aircraft's angular velocity, the aircraft's wing pose, the aircraft's tail pose, the aircraft's current coordinates, and the target coordinates of the aircraft in the next second; the aircraft performs three specific actions within one flight cycle, two of which are selected by the flight task through an encoder from 26 different actions formed by the combination of wing deformation, horizontal tail deformation, and speed change, and then combined with the aircraft's non-deformation actions to form the three actions; preferably, the actions of wing deformation and tail deformation are opposite; thereby, the constructed aircraft physical model is operated to obtain path data.

[0050] In some embodiments of the present invention, the task period is set to 30s and the step size of the action is 1s. Through different flight tasks, the encoder is used to reduce the action space.

[0051] In some embodiments of the present invention, the deep reinforcement learning includes: placing the acquired path into the corresponding position in the state for reinforcement learning training; in each episode, first resetting the environment so that the aircraft appears at the initial position of the path, and then selecting the action with the largest value according to the value function network; randomly selecting an action with a probability of (1-greedy), and before reaching a predetermined number of episodes, the greedy value increases linearly with a value less than 1, and then reaches 1; if the state after taking the action does not cause the aircraft to accelerate too much, interacting with the environment, and storing the array of state-action-reward-next state into a register; otherwise, randomly selecting one from the remaining actions, interacting with the environment, and storing it; when the amount stored in the register exceeds a predetermined amount, training the value function network with a set update frequency, learning rate, and reward discount; until the predetermined number of episodes is completed.

[0052] In some embodiments of the present invention, an encoder is completed based on the relationship between the generated route and actions, through which it is possible to determine what combination of actions the aircraft uses to fly out of a given trajectory.

[0053] The present invention has the following beneficial effects:

[0054] This invention overcomes the drawback of the difficulty in modeling morphing aircraft through aerodynamic parameter analysis. It designs a framework to analyze and control the motion of morphing aircraft even without aerodynamic data. Through this framework, the deformation design, analysis and control of morphing aircraft can be carried out, adapting to the multimodal, strong nonlinear and strong coupling characteristics of the system brought about by the deformation process. It can select the optimal control strategy under complex control to ensure flight stability.

[0055] Controlling aircraft using deep reinforcement learning methods allows you to input the factors to be considered into a deep learning network and assign a reward value, which is simpler than designing too many equations or inequalities using traditional optimization methods.

[0056] In this embodiment of the invention, when controlling the aircraft to perform trajectory tracking, an encoder is used to reduce the size of the value function network, which accelerates the convergence speed of the deep learning network during training and achieves a relatively ideal trajectory tracking effect. Using an encoder to reduce the motion space helps to accelerate training convergence while avoiding abnormal actions that could cause the aircraft to stall. Under reasonable mission design, the aircraft can reasonably deform. Attached Figure Description

[0057] Figure 1 This is a schematic diagram of the variant aircraft modeling in an embodiment of the present invention;

[0058] Figures 2a to 2c These are schematic diagrams showing different deformations of the wings of a morphing aircraft;

[0059] Figures 2d to 2f These are schematic diagrams showing different deformations of the tail fin of a morphing aircraft;

[0060] Figure 3 This is a schematic diagram of the coordinate system of the variant aircraft in an embodiment of the present invention;

[0061] Figure 4 This is a line graph showing the lift iteration process of a FLUENT simulated aircraft.

[0062] Figure 5 This is a schematic diagram of the reinforcement learning process in an embodiment of the present invention;

[0063] Figure 6a This is the convergence curve of reinforcement learning for Task 1 in Example 1;

[0064] Figure 6b This is a schematic diagram of the target trajectory of Task 1 in Example 1 and the trajectory of the aircraft controlled by the trained intelligent agent;

[0065] Figure 7a This is the convergence curve of reinforcement learning for Task 2 in Example 1;

[0066] Figure 7b This is a schematic diagram of the target trajectory of Task 2 in Example 1 and the trajectory of the aircraft controlled by the trained intelligent agent.

[0067] Figure 8a This is the convergence curve of the reinforcement learning process in the first 30 seconds of Task 3 in Example 1;

[0068] Figure 8bThis is a schematic diagram of the target trajectory in the first 30 seconds of Task 3 in Example 1 and the trajectory of the aircraft controlled by the intelligent agent after training.

[0069] Figure 8c This is the convergence curve of reinforcement learning in the last 30 seconds of Task 3 in Example 1;

[0070] Figure 8d This is a schematic diagram of the target trajectory 30 seconds after Task 3 in Example 1 and the trajectory of the aircraft controlled by the intelligent agent after training;

[0071] Figure 9 This is a flowchart of the steps of the robust control method for variant aircraft based on deep reinforcement learning in an embodiment of the present invention;

[0072] Figure 10 This is an overall block diagram of the robust control method for variant aircraft based on deep reinforcement learning in this embodiment of the invention. Detailed Implementation

[0073] The present invention will be further described below with reference to the accompanying drawings and preferred embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other.

[0074] It should be noted that the directional terms such as left, right, up, down, top, and bottom used in this embodiment are only relative concepts or are based on the normal use of the product, and should not be considered as restrictive.

[0075] This invention addresses the morphing control problem of morphing aircraft. In studying the control of morphing aircraft, the lift and drag coefficients change with deformation, making it difficult to analyze performance using aerodynamic parameters alone; modeling and subsequent analysis are necessary. This invention employs a deep reinforcement learning-based control method for the design, analysis, and control of morphing aircraft morphing, selecting the optimal control strategy for the complex control problem of morphing aircraft.

[0076] The modeling of the variant aircraft involved creating a CAD model using CATIA software, followed by simulation using FLUENT software to obtain aerodynamic data. Based on the obtained data, the kinematic and dynamic equations of the aircraft were solved to establish the aircraft model.

[0077] The aircraft control employs the Deep Q-Learning (DQN) algorithm, a classic reinforcement learning algorithm. Reinforcement learning, also known as reward learning, evaluation learning, or enhancement learning, is one of the paradigms and methodologies of machine learning. In the reinforcement learning framework, an agent explores the environment and receives rewards based on its state and actions. The agent's learning task is to find a policy that selects actions that maximize rewards in the long run. This learning task requires not only selecting actions that yield the highest reward for the current state but also considering long-term consequences to select actions that maximize global rewards for the agent. Currently, reinforcement learning algorithms are mainly divided into three types: value-based reinforcement learning, policy-based reinforcement learning, and the actor-critic framework. The most representative algorithms in the three reinforcement learning frameworks are DQN, TRPO (Trust Region Policy Optimization), and DDPG (Deep Deterministic Policy Gradient). The value function is a prediction of the expected, cumulative, discounted, and future reward, measuring the goodness or badness of each state or state-action pair. Deep reinforcement learning is a combination of deep learning and reinforcement learning, using neural networks to better represent policies or value functions. Currently, deep reinforcement learning is widely used in fields such as fault diagnosis and autonomous driving, and is a relatively complete algorithm. However, its disadvantages include difficulty in convergence and high computational cost, requiring optimization for industrial applications.

[0078] The following embodiments of the present invention address the complexities of analyzing mutated aircraft by proposing a complete framework for data analysis based on CATIA-FLUENT and using Python, as follows: Figure 10 As shown, based on a specific CAD model and the corresponding simulation data, this control is more realistic.

[0079] The following embodiments of the present invention also address the difficulty of morphing control for variant aircraft by designing a deep reinforcement learning control method with a step size of 1 second. During the control process, the task is designed with a 30-second cycle, and the action space is reduced using an encoder by varying the flight tasks. This accelerates the convergence speed of the deep network.

[0080] The following embodiments of the present invention propose a robust control method for variant aircraft based on deep reinforcement learning, such as... Figure 9 As shown, the process includes the following steps: S1. Establish a CAD model of the variator aircraft, then simulate to obtain aerodynamic data of the aircraft. Based on the obtained data, solve the kinematic and dynamic equations of the aircraft to complete the physical model of the variator aircraft's motion; S2. Use the DQN (Deep Q-Learning) algorithm to perform deep reinforcement learning for the deformation and motion control of the variator aircraft, and train the value function network; S3. Construct an intelligent agent based on the trained value function network, and make reasonable decisions on the control of the variator aircraft through the intelligent agent.

[0081] In a specific embodiment, in step S1, it is assumed that the velocity direction of the aircraft is always the same as the x-axis direction, that is, the trajectory coordinate system and the body coordinate system are in the same direction. Based on different combinations of deformation angles, the FLUENT module is used to measure the forces on the aircraft in the x, y, and z axes, and the sum of the torques of the three axes about the center of mass at various speeds.

[0082] In a specific embodiment, in step S2, when training the trajectory tracking of the variant aircraft, an encoder method is used to reduce the action space of the value function network, thereby accelerating the convergence speed of the deep learning network during training.

[0083] In a specific embodiment, step S2, the deep reinforcement learning includes: obtaining the target trajectory within one cycle as the task; the aircraft's state is 14-dimensional, including the ratio of the target's position to the current position, the aircraft's deflection angle, tilt angle, angular velocity, wing pose, tail pose, current coordinates, and target coordinates for the next second; the aircraft performs three specific actions within one flight cycle, two of which are selected by the flight task through an encoder from 26 different actions formed by the combination of wing deformation, horizontal tail deformation, and speed change, and then combined with the aircraft's non-deformation actions to form the three actions; preferably, actions with opposite wing and tail deformation are selected; thereby, the constructed aircraft physical model is manipulated to obtain path data.

[0084] In a specific embodiment, the deep reinforcement learning includes: placing the obtained path into the corresponding position in the state for reinforcement learning training; in each episode, first resetting the environment so that the aircraft appears at the initial position of the path, and then selecting the action with the largest value according to the value function network; randomly selecting an action with a probability of (1-greedy), and before reaching a predetermined number of episodes, the greedy value increases linearly with a value less than 1, and then reaches 1; if the state after taking the action does not cause the aircraft to accelerate too much, interacting with the environment, and storing the array of state-action-reward-next state into a register; otherwise, randomly selecting one from the remaining actions, interacting with the environment, and storing it; when the amount stored in the register exceeds a predetermined amount, training the value function network with a set update frequency, learning rate, and reward discount; until the predetermined number of episodes is completed.

[0085] The following describes the method of an embodiment of the present invention:

[0086] 1. Modeling and Simulation

[0087] This invention designs and builds a variant aircraft in which ailerons and engines are ignored and considered as modules that balance forces and maintain the aircraft's velocity direction. The CAD model is as follows: Figure 1 As shown.

[0088] Figure 1 The dotted areas on the wing and tail are deformable parts. When the deformable parts on the wing deform, they deform to the same size but in opposite directions. The deformable parts on the tail deform to the same size but in the same direction. The wing's deformation pattern is as follows: Figure 2a , Figure 2b and Figure 2c As shown, the deformation of the tail fin is as follows: Figure 2d , Figure 2e and Figure 2f As shown, different maneuvers are created by combining the deformation of the wings, the deformation of the horizontal stabilizer, and changes in speed.

[0089] The deformable angles of the left wing are specified as [-18°, -13.5°, -9°, -4.5°, 0°, 4.5°, 9°, 13.5°, 18°], and the deformable angles of the tail wing are specified as [0°, 3°, 6°, 9°, 12°, 15°]. There are a total of 30 different combinations of deformation angles (since the aircraft is completely symmetrical, the cases where the wing deformations are opposites are simply the opposites of the data in the y-direction of the aircraft).

[0090] ANSYS is a large-scale general-purpose finite element analysis (FEA) software. Using its FLUENT module for finite element analysis is a common method in fluid mechanics. In this embodiment of the invention, it is assumed that the aircraft's velocity is always in the same direction as the x-axis. That is, the trajectory coordinate system and the body coordinate system are in the same direction. The engine can cancel out the components of lift, gravity, and drag on the aircraft, and only the aircraft's torque is considered. For the convenience of calculating the experimental data in this embodiment of the invention, the trajectory coordinate system differs from the common case; the x-axis direction remains unchanged, while the y-axis and z-axis are reversed. The coordinate system is as follows: Figure 3 As shown.

[0091] Thirty STP files were input into ANSYS. The FLUENT module was used to measure the forces, moments, and torques along the three axes of the aircraft at velocities in the range [170, 180, 190, 200, 210, 220, 230, 240, 250] m / s. With the wing angle at 0°, the tail angle at 9°, and the velocity at 190 m / s, the FLUENT simulation of the aircraft's lift iteration process is as follows: Figure 4 As shown, the horizontal axis represents the number of iterations, and the vertical axis represents the lift (Lift / N).

[0092] The method for estimating the aircraft's moment of inertia J is as follows:

[0093]

[0094] Where W is the mass of the aircraft, and L a The total length of the aircraft The dimensionless radius of gyration along the y-axis of the body coordinate system is preferably... Take 0.3.

[0095] The formulas for moment of inertia J, torque M, angular acceleration α, angular velocity ω, and time t are as follows:

[0096]

[0097] The transformation matrices for the angles through which the original coordinate system rotates about the x and y axes are respectively:

[0098]

[0099]

[0100] in,

[0101]

[0102] During flight, this embodiment of the invention equates this process to first deforming the wings to rotate the aircraft around the x-axis, and then deforming the tail fin to change the aircraft's direction. In this process, let the rotated coordinates be... The coordinates before rotation are If the aircraft first rotates by ξ radians around the x-axis and then by η radians around the y-axis, then:

[0103]

[0104] The projection [x',y',z'] of the x-axis direction in the track coordinate system onto the ground coordinate system is:

[0105]

[0106] The solutions for the track deviation angle χ and the inclination angle γ are as follows:

[0107]

[0108]

[0109] Based on the flight path deflection and tilt angles, the aircraft's velocity V in the ground coordinate system Ox can be calculated. g y g z g The three velocity components:

[0110]

[0111]

[0112]

[0113] The position of the aircraft can be obtained by integrating the three velocity components. Based on the obtained data, a physical model of the aircraft's motion can be built.

[0114] 2. Strengthen control

[0115] Starting from t=0, assuming the aircraft flies for 30 seconds as one cycle, it performs three specific actions within one cycle. The task is the target trajectory within one cycle.

[0116] Within one cycle, the state is 14-dimensional, composed of the ratio of the target position to the current position (2-dimensional), the aircraft's yaw angle, tilt angle, angular velocity (2-dimensional), wing pose, tail pose, current coordinates (3-dimensional), and target coordinates for the next second (3-dimensional). The action is 3-dimensional, composed of two actions selected from 26 actions (see appendix) by the encoder, combined with action 0 (no aircraft deformation). The reward is the distance between the post-action coordinates and the target trajectory coordinates at that moment, divided by 100. The value function network is a 14*240*160*3 fully connected network with ReLU activation. When the agent selects an action that might cause excessive acceleration or stall, the action selector will avoid these actions. The overall diagram of reinforcement learning is shown below. Figure 5 .

[0117] The following is Example 1

[0118] First, the constructed aircraft physical model is manipulated to obtain path data. Two non-zero actions from the left side of the appendix are selected and combined with the zero action to form a three-dimensional action combination. In this combination, one action is taken per second for 30 seconds, and the aircraft's path for 30 seconds can be calculated on the physical model. In this embodiment, the starting point of the path is (0,0,0). Some noise is added to the path for training purposes; in this embodiment, only simple rounding is used to add noise. It should be noted that when selecting the above actions, choosing actions with opposite wing and tail deformations can make the aircraft's flight more consistent; otherwise, discrepancies with reality may occur.

[0119] Then, the obtained path is placed into the corresponding position in the state, and reinforcement learning training begins. In each episode, the environment is first reset so that the aircraft appears in the initial position (the same starting point as the obtained path). Then, the action with the largest value is selected according to the value function network. An action is randomly selected with a probability of (1-greedy). The greedy value increases linearly from 0.8 to 0.98 before 12,000 episodes, and then sets to 1. The purpose of setting the greedy value is to make the agent explore actions with smaller value functions, avoiding local optima. Then, if the state after taking the action does not cause the aircraft to accelerate too much, it interacts with the environment and stores the array of state-action-reward-next state into the buffer register; otherwise, a random action is selected from the remaining actions, interacts with the environment, and is stored in the buffer. When the buffer contains more than 2,000 entries, the value function network is trained with an update frequency of 100, a learning rate of 0.006, and a reward discount of 0.9. Training stops after 15,000 episodes are completed, and the agent can make reasonable decisions based on the value function network.

[0120] Finally, based on the relationship between the generated route and actions, an encoder is completed, which can then determine, given a route, the combination of actions by which the aircraft flies out of the trajectory.

[0121] Experimental results

[0122] The experiment was designed to track the trajectory of an aircraft along three different paths. Task 1 involved tracking the trajectory of the aircraft during its cruise, climb, and descent flight for 60 seconds. The experimental results are as follows: Figure 6a and Figure 6b As shown, where, Figure 6a It is the convergence curve of reinforcement learning. Figure 6a The horizontal axis represents the number of episodes explored, and the vertical axis represents the total reward. Figure 6b The points in the diagram represent the target trajectory and the trajectory of the aircraft controlled by the trained agent, respectively, and it can be seen that the consistency is good.

[0123] Mission 2 involves the aircraft performing cruise, turning, and climb maneuvers, followed by a 30-second trajectory tracking task. The experimental results are as follows: Figure 7a and Figure 7b As shown, where, Figure 7a It is the convergence curve of reinforcement learning. Figure 7a The horizontal axis represents the number of episodes explored, and the vertical axis represents the total reward. Figure 7b It is a trajectory tracking curve.

[0124] Task 3 involves the aircraft performing one cycle (30s) of each of the two tasks described above, for a total flight time of 60s, tracking its trajectory. The first 30s are for Task 1, and the experimental results are as follows: Figure 8a and Figure 8b As shown, where, Figure 8a It is the convergence curve of reinforcement learning. Figure 8a The horizontal axis represents the number of episodes explored, and the vertical axis represents the total reward. Figure 8b This is the trajectory tracking curve. The last 30 seconds were for Task 2. The experimental results are as follows: Figure 8c and Figure 8d As shown, where, Figure 8c It is the convergence curve of reinforcement learning. Figure 8c The horizontal axis represents the number of episodes explored, and the vertical axis represents the total reward. Figure 8d It is a trajectory tracking curve.

[0125] In the total reward-episodes curves of the three experiments above, it can be seen that when the greedy value increases to 1, the convergence of the total reward is very good; even with fluctuations, the reward remains within -100, thus the trajectory is within an acceptable range. In the trajectory tracking curves of the three experiments above, the aircraft's flight trajectory and the target trajectory almost overlap, achieving a relatively ideal trajectory tracking effect. The results from the three tasks show that the reinforcement network has good convergence, and the piecewise simplified task enables the aircraft to perform trajectory tracking well. Experimental results indicate that the DQN network has generalization ability and can adapt well to different deformation methods and trajectories. Furthermore, adding an encoder at the action selection point of the reinforcement network significantly reduces the network size, making it easier for the network to converge. Research shows that the aerodynamic performance of this variant aircraft is superior to that of the traditional aircraft, achieving better path tracking and stably performing rapid ascent and sharp turn maneuvers.

[0126] The embodiments of the present invention have the following beneficial effects:

[0127] 1. Typically, the kinematic and dynamic equations of an aircraft are directly given using aerodynamic characteristics. This invention overcomes the difficulty of modeling morphing aircraft through aerodynamic parameter analysis, providing a framework for analyzing the motion of morphing aircraft even without aerodynamic data. This framework enables the design, analysis, and control of morphing aircraft, adapting to the multimodal, highly nonlinear, and strongly coupled characteristics of the system during aircraft deformation. It allows for the selection of optimal control strategies under complex control conditions, ensuring flight stability.

[0128] 2. When controlling the aircraft for trajectory tracking, an encoder was used to reduce the size of the value function network, which accelerated the convergence speed of the deep learning network during training and achieved a more ideal trajectory tracking effect. Using an encoder to reduce the motion space helps to accelerate training convergence while preventing the aircraft from making abnormal movements that could lead to stall.

[0129] With proper mission design, the aircraft can reasonably transform.

[0130] 3. Controlling aircraft through reinforcement learning allows the factors to be considered to be input into a deep learning network and a reward value to be given, which is simpler than designing too many equations or inequalities using traditional optimization methods.

[0131] The embodiments of the present invention also have the following characteristics:

[0132] 1. An encoder that reduces the motion space based on different tasks;

[0133] 2. Modeling methods and simulation data generation methods for variant aircraft;

[0134] 3. Deep reinforcement learning control method for variant aircraft under 30 different configurations.

[0135] appendix

[0136]

[0137]

[0138] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, several equivalent substitutions or obvious modifications can be made without departing from the concept of the present invention, and all such modifications, achieving the same performance or purpose, should be considered within the scope of protection of the present invention.

Claims

1. A method for robust control of a morphing aircraft based on deep reinforcement learning, characterized in that, Comprise the following steps: S1, the CAD model of the morphing aircraft is established, and then the aerodynamic data of the aircraft is simulated, and the kinematics and dynamics equations of the aircraft are solved according to the obtained data, and the physical model of the motion of the morphing aircraft is completed; S2, the DQN algorithm is used for deep reinforcement learning of morphing and motion control of the morphing aircraft, and the value function network is trained; wherein, according to the relationship between the generated route and the action, an encoder is completed, through which it can be determined that the aircraft flies out of the given trajectory by combining the actions, and when the trajectory tracking training of the morphing aircraft is performed, the method of using the encoder reduces the action space of the value function network, so that the convergence speed of the deep learning network during training is accelerated; wherein, the aircraft performs a specific three actions in a cycle, wherein two actions are selected from a plurality of different actions formed by the combination of the deformation of the wing, the deformation of the horizontal tail and the change of the speed, and the third action is the action of the aircraft without deformation; S3, the intelligent agent is constructed according to the trained value function network, and reasonable decisions are made for the control of the morphing aircraft through the intelligent agent.

2. The method of claim 1, wherein, In step S1, the CAD model of the morphing aircraft is established by CATIA software, and then the aerodynamic data of the aircraft is simulated by FLUENT software.

3. The method of claim 2, wherein, In step S1, during modeling, the aileron and engine of the morphing aircraft are ignored, and the model is regarded as a balanced force and keeps the direction of the aircraft speed, while the wing and tail of the morphing aircraft have deformable parts, the deformable parts on the wing are equal in size and opposite in direction when deformed, and the deformable parts on the tail are equal in size and same in direction when deformed; wherein, different actions are combined from the deformation of the wing, the deformation of the horizontal tail and the change of the speed.

4. The method of claim 3, wherein, The deformable angle of the wing of the morphing aircraft is [-18°, -13.5°, -9°, -4.5°, 0°, 4.5°, 9°, 13.5°, 18°], and the deformable angle of the tail of the morphing aircraft is [0°, 3°, 6°, 9°, 12°, 15°], and there are 30 kinds of different deformation angle combinations.

5. The method of claim 3, wherein, In step S1, it is assumed that the direction of the speed of the aircraft is always the same as the direction of the x-axis, that is, the track coordinate system is the same as the direction of the body coordinate system, and according to different deformation angle combinations, the force on the aircraft in the x, y and z axial directions and the total torque of the three axial directions on the center of mass are measured by the FLUENT module when the speed is multiple; The estimation of the moment of inertia J of the aircraft is: where W is the aircraft mass, L a is the overall aircraft length, is the dimensionless radius of gyration in the y-axis direction of the body coordinate system; The moment M, angular acceleration a and angular velocity ω of the aircraft are: M=J×a (2) Where t is time; The conversion matrix of the original coordinate system rotating around the x and y axes by an angle is respectively: Where, When the aircraft is flying, it is equivalent to first deforming the wing to make the aircraft rotate around the x-axis, and then deforming the tail to realize the turning of the aircraft; the coordinates after rotation are: [x q ,y q ,z q ], and the coordinates before rotation are: [x p ,y p ,z p ]; if the aircraft rotates ξ radian around the x-axis first, and then rotates η radian around the y-axis, there are: The projection of the x-axis direction in the track coordinate system on the ground coordinate system is [x', y', z']: The track angle χ and the inclination angle γ are: The three velocity components of the aircraft V in the ground coordinate system Ox g y g z g are: Integrate the three speed components respectively to determine the position of the aircraft, and build the physical model of the motion of the morphing aircraft according to the obtained data.

6. The method of claim 3, wherein, In step S2, the deep reinforcement learning includes: taking the target trajectory in a cycle as a task, the state of the aircraft being 14-dimensional, including the ratio of the target position to the current position of the aircraft, the deflection angle of the aircraft, the tilt angle of the aircraft, the angular velocity of the aircraft, the wing pose of the aircraft, the tail pose of the aircraft, the current coordinates of the aircraft, and the target coordinates of the aircraft in the next second; the aircraft performing specific three actions in a cycle of flight, wherein two actions are selected from 26 different actions formed by combining the deformation of the wing, the deformation of the horizontal tail, and the change of the speed, and the third action is the action of the aircraft without deformation.

7. The method of claim 6, wherein, The action in which the wing deformation and the tail deformation are opposite is selected from the 26 different actions.

8. The method of claim 6, wherein, The cycle of the task is set to 30 seconds, the step length of the action is 1 second, and the action space is reduced by the encoder through different flight tasks.

9. The method of claim 3, wherein, The deep reinforcement learning includes: placing the obtained path in the corresponding position in the state, and performing reinforcement learning training; in each exploration, first reset the environment to make the aircraft appear at the initial position of the path, then select the action with the maximum value according to the value function network; there is a probability of 1-greedy to randomly select an action, and the greedy value is linearly increased to less than 1 before reaching the predetermined number of explorations, and then takes 1; if the state after taking the action does not make the acceleration of the aircraft too large, interact with the environment, and store the array of state-action-reward-next state in the register; otherwise, randomly select an action from the remaining actions, and store it after interacting with the environment; when the content stored in the register exceeds the predetermined amount, train the value function network at a set update frequency, learning rate and reward discount; until the predetermined number of explorations is completed.

Citation Information

Patent Citations

  • Dynamics modeling and analyzing method for aerospace vehicle

    CN103593524A

  • Flight coordination control system for morphing aircraft

    CN114637313A

  • Variant aircraft intelligent decision and control based on DQN-PID framework

    CN117755477A