Adaptive flight control method and system based on deep learning and reinforcement learning
Through an adaptive flight control method combining deep learning and enhanced learning, the adaptability and robustness of the aircraft in complex environments is solved, intelligent and accurate flight control is achieved, and the safety and reliability of the flight is ensured.
Patent Information
- Application Number
- CN202510396252.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When facing a complex and changing flight environment, existing aircraft control systems are not adaptable and robust, making it difficult to cope with sudden meteorological changes and aircraft performance degradation, resulting in degradation or failure of control system performance.
Using a combination of deep learning and enhanced learning, we use the method of training the flight stage prediction model and building a Q table, combining ε-greedy strategy, dynamically adjust the control strategy, and introduce a fault-tolerant mechanism to ensure the adaptive control of the aircraft in complex environments.
It improves the targeted and efficient flight control, enhances adaptability and robustness, ensures flight safety and reliability, and can deal with unknown or rapidly changing environments in real time.
Smart Images

Figure CN120255347A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of automatic control of aircraft, and particularly to an adaptive flight control method and system based on deep learning and reinforcement learning. Background Art
[0002] In traditional aircraft control systems, flight control strategies are usually implemented based on fixed mathematical models and pre-set control rules. These systems often rely on an accurate understanding and modeling of aircraft dynamics, as well as an accurate prediction of the flight environment. However, during actual flight, an aircraft may encounter various unpredictable complex situations, such as sudden meteorological changes, changes in flight load, and performance degradation of the aircraft itself. These can all lead to a decline in the performance of traditional control systems based on fixed models, or even failure.
[0003] With the development of aviation technology, higher requirements are put forward for the flight control systems of aircraft, including stronger adaptability, autonomy, and safety. To address these challenges, in recent years, artificial intelligence technologies, especially machine learning and pattern recognition technologies, have been increasingly applied in aircraft control systems. These technologies can provide more flexible and intelligent control strategies, and continuously optimize the flight control process by learning historical data and real-time feedback.
[0004] Nevertheless, existing artificial-intelligence-based flight control systems still have some limitations. For example, they may require a large amount of training data, have high requirements for computing resources, or their adaptability and robustness still need to be improved when facing unknown or rapidly changing environments. In addition, how to effectively combine artificial intelligence technologies with the real-time control requirements of aircraft while ensuring the safety and reliability of the system is also a current research hotspot and difficulty. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention provides an adaptive flight control method and system based on deep learning and reinforcement learning. By combining deep learning and reinforcement learning technologies, the method realizes adaptive control of the aircraft in various flight stages and complex environments, and at the same time introduces a fault tolerance mechanism and a human-machine interaction interface to improve the comprehensive performance and application value of the system.
[0006] The present invention is implemented through the following technical solutions: In a first aspect, the present application provides an adaptive flight control method based on deep learning and reinforcement learning, including the following steps: Step 1: According to the flight states of the aircraft in each flight stage, as well as environmental data and corresponding meteorological conditions, train a flight stage prediction model. The trained flight stage prediction model outputs the optimal flight stage of the aircraft under the meteorological conditions. Step 2: Construct a Q-table with the flight state and environmental parameters as states and control commands as actions. Based on the current action of the aircraft and in combination with the ε-greedy strategy, determine the action with the highest Q-value in the current state, use the action with the highest Q-value at the current moment as the action for the next moment, and determine the optimal control strategy according to the action with the highest Q-value at each moment. Step 3: Control the flight of the aircraft in the optimal flight phase under meteorological conditions according to the optimal control strategy.
[0007] Preferably, the flight phases described in Step 1 include takeoff, cruise, descent, or landing; The flight state includes the three-dimensional velocity, altitude, attitude, acceleration, and / or angular velocity of the aircraft; The environmental data includes atmospheric pressure, temperature, humidity, wind speed, and / or wind direction.
[0008] The meteorological conditions include turbulence, thunderstorm clear sky, cloudy, rainy, or snowy days.
[0009] Preferably, the training method of the flight phase prediction model is as follows: Use the flight state and environmental data as training data, use the flight phase and meteorological conditions as labels, construct a data set, and divide the data set into a training data set and a test data set; Train a deep learning model using the training data set and evaluate the performance of the trained deep learning model using the test data set; Evaluate the performance of the deep learning model according to the accuracy, recall rate, and F1 score of the deep learning model. When the performance of the deep learning model does not meet the requirements, optimize the architecture or hyperparameters of the deep learning model.
[0010] Preferably, the determination method of the optimal control strategy in Step 2 is as follows: Based on the current action of the aircraft and in combination with the ε-greedy strategy, select the action with the highest Q-value in the current state; Control the aircraft at the next moment according to the action with the highest Q-value, and obtain the state and corresponding reward at the next moment; Update the corresponding Q-value in the Q-table according to the state and corresponding reward at the next moment; Repeat the above process, iteratively update the Q-table until the set stop condition is met, and determine the optimal control strategy of the aircraft in this flight phase according to the action with the largest Q-value at each moment.
[0011] Preferably, the update method of the Q-value is as follows:
[0012] Wherein, s t ist The state at a moment a t is t the action taken at a moment α is the learning factor γ is the discount factor
[0013] Preferably, it further includes step 4: Obtain the state data of each subsystem of the aircraft during flight, perform real-time fault diagnosis on the aircraft according to the state data, and start the backup control strategy when the aircraft fails
[0014] Preferably, the flight phase prediction model in step 1 is a feedforward neural network, a convolutional neural network, a recurrent neural network, a generative adversarial network or a deep neural network
[0015] In a second aspect, the present application provides an adaptive flight control system based on deep learning and reinforcement learning, including: A prediction module, configured to train a flight phase prediction model according to the flight state data of the aircraft in each flight phase, as well as environmental data and corresponding meteorological conditions, and the trained flight phase prediction model outputs the optimal flight phase of the aircraft under the meteorological conditions A reinforcement learning module, configured to construct a Q-table with the flight state and environmental parameters as the state and the control instruction as the action, determine the action with the highest Q value in the current state according to the current action of the aircraft and in combination with the ε-greedy strategy, use the action with the highest Q value at the current moment as the action for the next moment, and determine the optimal control strategy according to the action with the highest Q value at each moment
[0016] A control module, configured to control the flight of the aircraft in the optimal flight phase under the meteorological conditions according to the optimal control strategy
[0017] In a third aspect, the present application provides an electronic device, including: A memory, configured to store a computer program A processor, configured to implement the steps of the adaptive flight control method based on deep learning and reinforcement learning when executing the computer program
[0018] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the adaptive flight control method based on deep learning and reinforcement learning are implemented
[0019] Compared with the prior art, the present invention has the following beneficial technical effects: An adaptive flight control method based on deep learning and reinforcement learning provided by this application. First, a flight phase prediction model is trained through deep learning technology. This method can accurately predict the optimal flight phase of the aircraft under specific meteorological conditions, which not only improves the pertinence of flight but also greatly enhances flight efficiency. Second, a Q-table is constructed using reinforcement learning and combined with the ε-greedy strategy to dynamically determine the optimal control strategy, enabling the aircraft to adjust control commands in real-time according to flight states and environmental parameters. This adaptability greatly enhances the robustness of the flight control system. In addition, this method comprehensively considers various flight factors, provides a more intelligent and accurate flight control solution for the aircraft, effectively copes with the complex and changeable flight environment, and ensures flight safety and reliability. Brief Description of the Drawings
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other relevant drawings can also be obtained based on these drawings.
[0021] Figure 1 It is a flowchart of the adaptive flight control method based on deep learning and reinforcement learning of the present invention.
[0022] Figure 2 It is a flowchart of the adaptive flight control method of Embodiment 4 of the present invention. Detailed Embodiments
[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. Usually, the components of the embodiments of this application described and shown in the drawings here can be arranged and designed in various different configurations.
[0024] Therefore, the detailed description of the embodiments of this application provided in the drawings below is not intended to limit the scope of the claimed application, but merely represents the selected embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0025] An adaptive flight control method based on deep learning and reinforcement learning includes the following steps: Step 1: According to the flight state data of the aircraft in each flight phase, as well as the environmental data and corresponding meteorological conditions, train the flight phase prediction model. The trained flight phase prediction model outputs the optimal flight phase of the aircraft under the meteorological conditions. Step 2: Construct a Q-table with the flight state and environmental parameters as the state and the control command as the action. According to the current action of the aircraft and combined with the ε-greedy strategy, determine the action with the highest Q value in the current state. Take the action with the highest Q value at the current moment as the action for the next moment, and determine the optimal control strategy according to the action with the highest Q value at each moment.
[0026] Step 3: Control the flight of the aircraft in the optimal flight phase under the meteorological conditions according to the optimal control strategy.
[0027] This adaptive flight control method based on deep learning and reinforcement learning first trains the flight phase prediction model through deep learning, which can accurately output the optimal flight phase of the aircraft under specific meteorological conditions, improving the pertinence and efficiency of flight. Then, it uses reinforcement learning to construct a Q-table and combines the ε-greedy strategy to determine the optimal control strategy, enabling the aircraft to dynamically adjust the control command according to the real-time state, enhancing the adaptability and robustness of the flight. This method comprehensively considers the flight state, environmental parameters, and meteorological conditions, providing a more intelligent and accurate flight control scheme for the aircraft.
[0028] Embodiment 1 An adaptive flight control method based on deep learning and reinforcement learning includes the following steps: Step 1: Obtain the flight state data of the aircraft in each flight phase, as well as the environmental data and corresponding meteorological conditions.
[0029] The flight phases include, but are not limited to, takeoff, cruise, descent, and landing.
[0030] The flight state data includes, but is not limited to, the three-dimensional velocity, altitude, attitude (pitch angle, roll angle, yaw angle), acceleration, angular velocity, etc. of the aircraft.
[0031] The environmental data includes, but is not limited to, atmospheric pressure, temperature, humidity, wind speed, wind direction, etc.
[0032] The meteorological conditions include, but are not limited to, turbulence, thunderstorm clear sky, cloudy, rainy, snowy.
[0033] Step 2: Use the flight state data and environmental data as training data, and the flight phase and meteorological conditions as label data to train the deep learning model. The trained deep learning model predicts the optimal flight phase of the aircraft under the meteorological conditions.
[0034] During the training process of the deep learning model, the collected dataset is divided into a training set and a test set. The training set is used to train the deep learning model, and the test set is used to evaluate the performance of the trained deep learning model. Metrics such as the accuracy, recall rate, and F1 score of the model are calculated. If the model performance is not good, the model architecture or hyperparameters are adjusted.
[0035] For the trained deep learning model, new flight state data and environmental data are input into the trained model, and the model will output the optimal flight phase of the aircraft under meteorological conditions.
[0036] The deep learning model is a feedforward neural network, a convolutional neural network, a recurrent neural network, a generative adversarial network, or a deep neural network.
[0037] Step 3: Construct a Q-table with the flight state and environmental parameters as the state and the control command as the action. According to the current action of the aircraft and combined with the ε-greedy strategy, determine the action with the highest Q-value in the current state, take the action with the highest Q-value at the current moment as the action for the next moment, and determine the optimal control strategy based on the action with the highest Q-value at each moment.
[0038] S3.1: In the process of using the Q-learning algorithm to obtain the optimal control strategy of the aircraft, first, a Q-table needs to be created. The state of the Q-table is composed of the flight state data of the aircraft (such as three-dimensional velocity, altitude, attitude, acceleration, etc.) and environmental data (such as atmospheric pressure, temperature, humidity, wind speed, wind direction, etc.), and the action is the control input of the aircraft. For example, the deflection angles of each rudder surface and the thrust setting of the engine.
[0039] S3.2: Initialize the Q-table by initializing all Q-values to 0 or a small random value.
[0040] S3.3: According to the current action of the aircraft and combined with the ε-greedy strategy, select the action with the highest Q-value in the current state.
[0041] ε represents the exploration factor, and its value range is 0 to 1. If ε is close to 1, the aircraft will choose to explore the environment more, that is, randomly select actions; if ε is close to 0, the aircraft will make more use of the environment, that is, directly select the action with the highest Q-value in the current state.
[0042] The aircraft uses the ε-greedy strategy to select actions. That is, randomly select an action with a probability of ε, which can explore new control strategies, and select the action with the highest Q-value in the current state with a probability of 1 - ε. As the training progresses, the value of ε will gradually decrease, making the algorithm more and more dependent on the learned experience.
[0043] S3.4. Control the aircraft at the next moment according to the action with the highest Q value, and obtain the state and corresponding reward at the next moment; The aircraft enters a new state according to the currently selected action and simultaneously obtains a reward value. The setting of this reward value is closely centered around the goal of the optimal control strategy. For example, when the aircraft develops in the direction of minimizing the flight path and minimizing the flight time, such as getting closer to the target path and shortening the flight time, it will receive a positive reward; conversely, if it deviates from the target path and extends the flight time, it will receive a negative reward. In terms of energy consumption optimization, if the aircraft reduces energy consumption, it will receive a positive reward, and an increase in energy consumption is a negative reward. For the stability and safety of the aircraft, if it can maintain stable flight and no potential hazards are detected, a positive reward is given; once unstable conditions occur or potential hazards are detected, a negative reward is given.
[0044] S3.5. Update the corresponding Q value in the Q-table according to the state and corresponding reward at the next moment. The update method of the Q value is as follows:
[0045] where, s t is the state at time t, a t is t the action taken at time α is the learning factor, and its value range is 0 to 1, γ is the discount factor, and its value range is 0 to 1.
[0046] S3.6. Repeat steps S3.3 - S3.5 to iteratively update the Q-table until the set stop condition is met, and determine the optimal control strategy of the aircraft in this flight phase according to the action with the largest Q value at each moment.
[0047] Optionally, when the set stop condition is met, this stop condition can be reaching a pre-set number of iterations or the change in the Q value being less than a certain set threshold.
[0048] When all possible actions have been executed in the state at a given moment, an optimal action can be selected according to the feedback reward information of the environment to enter the next state. The Q value of each state is also continuously updated as the exploration progresses until the Q values of all states tend to be stable, which means that after multiple iterations, the optimization goal is achieved. At this time, select the action with the largest Q value in each state and combine these actions to form the optimal control strategy.
[0049] Step 4. According to the flight phase under the meteorological conditions predicted by the deep learning model in step 2, use the optimal control strategy obtained by the Q-learning algorithm in step 3 to adjust the flight state of the aircraft in real time.
[0050] According to the flight phase and corresponding meteorological conditions predicted by the deep learning model in step 2, and based on the optimal control strategy obtained by the Q-learning algorithm in step 3, the flight state of the aircraft is adjusted in real time. For example, the deflection angles of the control surfaces (ailerons, elevators, rudders) are precisely adjusted according to the current flight conditions to change the flight direction and attitude of the aircraft; the thrust setting of the engine also changes in real time, and is reasonably adjusted according to factors such as flight speed, altitude, and load to ensure that the aircraft has sufficient power.
[0051] Step 5: Obtain the state data of each subsystem of the aircraft during flight, perform real-time fault diagnosis on the aircraft according to the state data, and start the backup control strategy when the aircraft fails.
[0052] During the flight of the aircraft, the states of key subsystems of the aircraft are monitored in real time, including sensors, actuators, power systems, flight control systems, etc. The state data of each subsystem is detected and diagnosed in real time. When an abnormality is detected in a subsystem, the source of the fault can be quickly located and fault information can be provided so that the control system of the aircraft can take corresponding countermeasures. At the same time, when the main control strategy fails, it can immediately take over the control of the aircraft and execute preset safety procedures, such as automatic return and emergency landing.
[0053] The adaptive flight control method of the present invention: First, analyze the real-time data of the aircraft in a changing environment through a deep learning algorithm to achieve rapid environment recognition and dynamic adjustment of the control strategy, ensuring the performance and safety of the aircraft in the face of unknown or rapidly changing flight conditions; Second, through the fault tolerance and safety mechanism, this mechanism can immediately start the backup control strategy when an abnormality in the key system is detected, ensuring the safe and stable flight of the aircraft under various challenges, and playing a decisive role in improving the autonomous control level and application value of the aircraft.
[0054] Embodiment 2 Corresponding to the above-mentioned adaptive flight control method based on deep learning and reinforcement learning, the present application also provides an adaptive flight control system based on deep learning and reinforcement learning, which may include: A prediction module, configured to train a flight phase prediction model according to the flight state data of the aircraft in each flight phase, as well as environmental data and corresponding meteorological conditions, and the trained flight phase prediction model outputs the optimal flight phase of the aircraft under the meteorological conditions. The reinforcement learning module is used to construct a Q-table with the flight state and environmental parameters as the state and the control instruction as the action. According to the current action of the aircraft and combined with the ε-greedy strategy, it determines the action with the highest Q value in the current state, takes the action with the highest Q value at the current moment as the action for the next moment, and determines the optimal control strategy according to the action with the highest Q value at each moment.
[0055] The control module is used to control the flight of the aircraft in the optimal flight phase under meteorological conditions according to the optimal control strategy.
[0056] Embodiment 3 An adaptive flight control device based on deep learning and reinforcement learning, comprising: The data acquisition module is used to obtain the flight state data of the aircraft in each flight phase, as well as the environmental data and the corresponding meteorological conditions.
[0057] The data acquisition module includes various types of sensors on the aircraft and a data preprocessing unit. The sensors are used to collect the data of the aircraft and send it to the data preprocessing unit.
[0058] The aircraft is equipped with a high-precision sensor kit, including but not limited to: three-axis accelerometer, three-axis gyroscope, magnetometer, airspeed sensor, static and dynamic pressure sensors, GPS receiver, meteorological sensors, etc., for real-time collection of the flight state and external environmental data of the aircraft.
[0059] The data preprocessing unit performs noise filtering, data fusion, time synchronization, and outlier rejection on the sensor data to ensure the accuracy and availability of the data.
[0060] The deep learning module is used to predict the optimal flight phase under the current meteorological conditions according to the flight state data and environmental data of the aircraft.
[0061] The real-time flight state data and environmental data of the aircraft are input into the trained deep learning model, and the deep learning model outputs the current flight phase and environmental conditions of the aircraft.
[0062] Flight phases: takeoff and climb, cruise, descent and approach, landing roll; Environmental conditions: clear, cloudy, rain and snow, thunderstorm, night flight, turbulence, etc.
[0063] The reinforcement learning module adopts the Q-learning algorithm to automatically adjust the mapping relationship between the aircraft control input and output through interaction with the environment to obtain the optimal control strategy.
[0064] The detection and diagnosis module is used to obtain the state data of each subsystem of the aircraft during flight, perform real-time fault diagnosis on the aircraft according to the state data, and start the backup control strategy when the aircraft fails.
[0065] The detection and diagnosis module includes a system health monitoring unit, a fault detection and diagnosis unit, and a backup control strategy unit.
[0066] System health monitoring unit: Real-time monitors the status of each subsystem of the aircraft, including sensors, actuators, power systems, etc., to ensure that all critical systems are in normal working condition; Fault detection and diagnosis unit: Provides fault diagnosis information when an anomaly is detected, quickly locates the source of the problem, such as sensor failures, actuator malfunctions, power system anomalies, etc.; Backup control strategy unit: Takes over the control of the aircraft when the main control strategy fails, ensuring that the aircraft can return safely or land, such as automatically switching to backup sensor data or executing an emergency landing procedure.
[0067] The human-machine interaction module is used to display the real-time flight status of the aircraft and system decisions.
[0068] The human-machine interaction interface provides an integrated pilot operation interface, including a high-resolution touch screen, a voice recognition system, physical buttons, and a joystick, for presenting the real-time flight status of the aircraft and system decisions. The interface design is intuitive and can clearly display flight parameters, system-recommended control strategies, environment recognition results, and any system warnings or fault information.
[0069] The human-machine interaction module includes a status display unit, a system decision display unit, and a manual control unit.
[0070] Status display unit: Displays the real-time flight status of the aircraft in a graphical interface, including key parameters such as speed, altitude, attitude, etc., as well as the system-recommended flight path and control strategy; System decision display unit: Displays the system-recommended control strategies, including flight path planning, rudder surface deflection angle, engine thrust, etc., as well as the current flight environment and phase; Manual control unit: Allows the pilot to manually adjust the flight status of the aircraft when necessary, including direct control through the joystick or buttons, and issuing commands through the voice recognition system.
[0071] This application integrates the above-mentioned modules into a complete adaptive flight control device and conducts preliminary tests in a high-fidelity flight simulator. After completing the simulator tests, the adaptive flight control system will conduct further flight tests on actual aircraft to verify its performance in a real flight environment. The design of this adaptive flight control system takes into account the specific requirements of different types of aircraft, such as fixed-wing aircraft, helicopters, multi-rotor aircraft, etc., and can be customized and optimized according to the flight characteristics of different aircraft. The application of the system is not limited to the civil aviation field, but also applicable to military aircraft. Especially when performing complex tasks or flying in harsh environments, it can significantly improve the survival ability and mission success rate of the aircraft.
[0072] The adaptive flight control system of the present invention can achieve adaptive control of the aircraft in various flight phases and complex environments, while ensuring flight safety and system reliability. The design of this system fully considers the actual needs of aircraft control and the latest progress of artificial intelligence technology, and has high innovation and practicality.
[0073] The adaptive flight control system of the present invention provides innovative solutions to two key technical problems: the flight adaptability of the aircraft in complex environments and the safety and reliability of the flight control system. For the problem of complex environment adaptability, the system uses deep learning algorithms to identify changing flight conditions in real time and continuously optimizes control strategies through reinforcement learning algorithms to achieve effective adaptation to unknown or rapidly changing flight environments. At the same time, in order to improve the safety and reliability of the system, this system designs an advanced fault tolerance mechanism. When sensor failures, actuator failures or other key system anomalies are detected, it can immediately activate backup control strategies to ensure the stable flight and safe landing of the aircraft in the face of various challenges. The implementation of these technologies significantly enhances the autonomous control ability of the aircraft, broadens its application scope, and increases the survival probability in extreme or unexpected situations.
[0074] Embodiment 4 A control method for an adaptive flight control device based on deep learning and reinforcement learning as described in Embodiment 3 includes the following steps: S1. Data acquisition stage: The sensor suite carried by the aircraft is calibrated before flight to ensure the accuracy of data acquisition. In a typical flight mission, the sensors collect data at a sampling frequency of 1 kHz, including three-axis acceleration (±5g), three-axis angular velocity (±1200° / s), three-axis magnetic field (±2000 μT), and airspeed (±150 m / s).
[0075] Pre-flight preparation: Conduct a comprehensive inspection of the aircraft to ensure that all sensors and actuators are working properly. Sensor calibration parameter setting: The calibration coefficient of the accelerometer is set to 0.01 g / LSB, and the gyroscope zero bias stability is 0.01 ° / s.
[0076] S2. Data preprocessing: The collected data is filtered for noise through a band-pass filter to remove high-frequency noise and low-frequency drift. The cut-off frequencies of the filter are set to 5 Hz and 100 Hz.
[0077] Data synchronization processing to ensure the consistency of data from different sensors in terms of timestamps.
[0078] Use a synchronization module to synchronize the timestamps of sensor data to ensure that the data synchronization error is less than 1 millisecond. The collected flight state data includes: three-dimensional velocity (Vx, Vy, Vz), altitude (H), attitude angles (Pitch, Roll, Yaw), and acceleration (Ax, Ay, Az).
[0079] S3. Deep learning for environment recognition: Use CNN to predict the optimal flight phase of the aircraft under the current meteorological conditions.
[0080] S4. Optimization of reinforcement learning control strategy: The exploration rate of the Q-learning algorithm is initially set to 1.0 and gradually reduced to 0.1 as the training progresses to balance the relationship between exploration and exploitation. The optimal control strategy is optimized through 1000 iterations in the simulation environment. Each iteration includes 1000 time steps, and the flight path error of the aircraft is reduced from an initial average of 10 meters to an average of 2 meters.
[0081] In a high-fidelity flight simulator, use the Q-learning algorithm to conduct offline training on the control strategy. The training period is 5 days, and a total of 100 flight hours are simulated. The flight path error of the aircraft in the simulation environment using the trained control strategy is reduced by 50%.
[0082] S5. Real-time adaptive control: The real-time control algorithm calculates the deflection angle of the control surface according to the current flight state and the optimized control strategy. For example, the elevator deflection angle is adjusted between -10° and +10° in steps of 0.1°. The engine thrust setting is adjusted in real time according to the flight speed and altitude, and the thrust change range is between 0% and 100% in steps of 1%.
[0083] The real-time control algorithm dynamically adjusts the deflection angle of the control surface and the engine thrust according to the current flight state and the optimized control strategy. The adjustment accuracy of the control surface deflection angle reaches 0.01°, and the adjustment accuracy of the engine thrust reaches 0.5%.
[0084] S6, Fault Tolerance and Safety Mechanism: The system health monitoring unit monitors the sensor data in real time and sets thresholds to trigger fault alarms.
[0085] For example, the abnormal threshold for accelerometer data is set at ±6g. During an actual flight test, when a sensor failed, the system successfully switched to the backup sensor within 0.5 seconds and maintained the stable flight of the aircraft.
[0086] S7, Human-Machine Interface: The pilot operation interface displays the real-time flight status of the aircraft at a refresh rate of 60 frames per second, including the three-dimensional flight path, speed vector, and attitude angle. During a flight demonstration, the pilot manually adjusted the flight path through the touch screen interface, and the system responded to the pilot's input within 1 second and adjusted the flight control parameters.
[0087] The pilot performed a simulated emergency landing operation through the human-machine interface. The system response time was less than 2 seconds, and the emergency landing procedure was correctly executed.
[0088] S8, The adaptive flight control system was tested in a high-fidelity flight simulator for 100 hours, simulating various flight environments and fault conditions. During the flight test on the actual aircraft, the system achieved autonomous control in 95% of the flight missions, and successfully maintained the stability of the aircraft when encountering unforeseen turbulence.
[0089] The system integration test was carried out in the actual flight environment, including stages such as takeoff, cruise, maneuvering flight, and landing. The test results showed that the system achieved adaptive control in 98% of the flight missions, significantly improving the automation level of flight. During a long-endurance reconnaissance mission, the adaptive flight control system enabled the aircraft to autonomously complete a flight mission of up to 8 hours without manual intervention. The system automatically generated a flight report, including the flight path, history of control parameter adjustment, and system performance evaluation. The report showed that the system successfully guided the aircraft to avoid thunderstorm areas under complex meteorological conditions, ensuring the smooth completion of the flight mission. The application of this system has significantly improved the autonomy level of the aircraft, reduced the workload of the pilot, and increased the success rate of mission execution.
[0090] The adaptive flight control system of the present invention demonstrates its effectiveness and reliability in actual flight missions. The design of the system meets the requirements of modern aircraft for intelligent and automated control, and provides strong technical support for the future development of aircraft.
[0091] Embodiment 5 An aircraft integrates the above-mentioned adaptive flight control system based on deep learning and reinforcement learning in its avionics system.
[0092] Integrate the adaptive flight control system with the avionics system of the aircraft to ensure that all sensors, actuators, and control algorithms can work seamlessly together. The system underwent preliminary functional tests at the ground control center by simulating the flight states of the aircraft. The test parameters included control surface deflection, engine response, sensor feedback, etc.
[0093] Calibrate the system parameters, including adjusting the sensitivity of sensors, optimizing the response time of actuators, and adjusting the parameters of control algorithms. Through ground tests, the optimal control parameters were determined. For example, the step angle of control surface deflection is 0.05°, and the minimum step of engine thrust adjustment is 0.1%.
[0094] Train pilots on the use of the adaptive flight control system in a flight simulator, simulating various flight environments and emergency situations. The pilots became familiar with the operation interface and control logic of the system through the simulator, improving their trust in the system and operational proficiency.
[0095] In actual flight tests, the aircraft was tested at different flight phases (such as takeoff, climb, cruise, descent, landing). The flight tests included control performance tests under normal flight conditions and fault tolerance tests under abnormal conditions.
[0096] Collect data from flight tests, including flight trajectories, control parameters, system status, pilot operations, etc. Use data analysis tools to conduct in-depth analysis of flight data to evaluate the control effect of the system and the operation performance of the pilots.
[0097] Based on the results of flight data analysis, evaluate the adaptive control performance of the system, including flight accuracy, response speed, stability, etc. By comparing flight test data and simulation data, the performance of the system in the actual flight environment was verified.
[0098] Based on the results of flight tests and data analysis, conduct iterative optimization of the system, including algorithm optimization, parameter adjustment, improvement of fault tolerance mechanisms, etc. The optimized system demonstrated better control performance and higher reliability in subsequent flight tests.
[0099] Apply the optimized adaptive flight control system to actual flight missions, such as border patrol, disaster monitoring, cargo transportation, etc. In actual applications, the system demonstrated excellent adaptive control capabilities, effectively coping with changing flight environments and complex mission requirements.
[0100] Collect feedback from pilots and ground maintenance personnel on the use of the system, including operation convenience, system stability, performance satisfaction, etc. Based on user feedback, conduct further user experience optimization and function upgrades for the system.
[0101] The adaptive flight control system of the present invention has been successfully applied and verified on actual aircraft. Steps such as the system's integration testing, flight testing, data analysis, performance evaluation, iterative optimization, and practical application have comprehensively demonstrated the excellent performance and high reliability of the system in actual flight missions.
[0102] It should be noted that in several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each module is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The modules described as separate components may or may not be physically separated. The components shown as modules can be one physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0103] In addition, in each embodiment of the present invention, each module can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0104] An electronic device provided by an embodiment of the present application includes a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the steps of the adaptive flight control method based on deep learning and reinforcement learning described in any of the above embodiments are implemented.
[0105] Another electronic device provided by an embodiment of the present application may further include: an input port connected to the processor, for transmitting multi-modal data collected by an external acquisition device to the processor; and a display unit connected to the processor, for displaying the processing result of the processor to the outside; a communication module connected to the processor, for realizing the communication between the electronic device and the outside. The display unit can be a display panel, a laser scanning display, etc.; the communication methods adopted by the communication module include but are not limited to Mobile High-Definition Link technology (HML), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), wireless connection (including Wireless Fidelity technology (WiFi), Bluetooth communication technology, Low-Energy Bluetooth communication technology, communication technology based on IEEE802.11s).
[0106] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, it implements the steps of the adaptive flight control method based on deep learning and reinforcement learning described in any of the foregoing embodiments.
[0107] For the description of the relevant parts in the adaptive flight control system, electronic device, and computer-readable storage medium based on deep learning and reinforcement learning provided by the embodiments of the present application, please refer to the detailed description of the corresponding parts in the adaptive flight control method based on deep learning and reinforcement learning provided by the embodiments of the present application, which will not be elaborated here. In addition, the parts of the above technical solutions provided by the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.
[0108] The above content is only to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any modification made on the basis of the technical solution according to the technical idea proposed by the present invention falls within the protection scope of the claims of the present invention.
Claims
1. An adaptive flight control method based on deep learning and reinforcement learning, characterized in that, It includes the following steps: Step 1: Train a flight phase prediction model based on the flight states of the aircraft in each flight phase, environmental data, and corresponding meteorological conditions. The trained flight phase prediction model outputs the optimal flight phase of the aircraft under the meteorological conditions; Step 2: Construct a Q-table with flight states and environmental parameters as states and control commands as actions. According to the current action of the aircraft and in combination with the ε-greedy strategy, determine the action with the highest Q-value in the current state. Take the action with the highest Q-value at the current moment as the action for the next moment, and determine the optimal control strategy based on the action with the highest Q-value at each moment; Step 3: Control the flight of the aircraft in the optimal flight phase under the meteorological conditions according to the optimal control strategy.
2. The adaptive flight control method based on deep learning and reinforcement learning according to claim 1, wherein The flight phases described in Step 1 include takeoff, cruise, descent, or landing; The flight states include the three-dimensional velocity, altitude, attitude, acceleration, and / or angular velocity of the aircraft; The environmental data includes atmospheric pressure, temperature, humidity, wind speed, and / or wind direction; The meteorological conditions include turbulence, thunderstorm clear sky, cloudy, rainy, or snowy days.
3. An adaptive flight control method based on deep learning and reinforcement learning according to claim 1, characterized in that The training method of the flight phase prediction model is as follows: Use the flight states and environmental data as training data, and the flight phases and meteorological conditions as labels to construct a dataset. Divide the dataset into a training dataset and a test dataset; Train a deep learning model using the training dataset and evaluate the performance of the trained deep learning model using the test dataset; Evaluate the performance of the deep learning model based on the accuracy, recall rate, and F1 score of the deep learning model. When the performance of the deep learning model does not meet the requirements, optimize the architecture or hyperparameters of the deep learning model.
4. An adaptive flight control method based on deep learning and reinforcement learning according to claim 1, characterized in that, The determination method of the optimal control strategy described in Step 2 is as follows: According to the current action of the aircraft and in combination with the ε-greedy strategy, select the action with the highest Q-value in the current state; Control the aircraft at the next moment according to the action with the highest Q-value to obtain the state and corresponding reward at the next moment; Update the corresponding Q-value in the Q-table according to the state and corresponding reward at the next moment; Repeat the above process, iteratively update the Q-table until the set stop condition is met, and determine the optimal control strategy of the aircraft in this flight phase according to the action with the largest Q-value at each moment.
5. The adaptive flight control method based on deep learning and reinforcement learning according to claim 4, characterized in that, The update method of the Q-value is as follows: Among them, s t is t the state at a moment, a t is t the action taken at a moment, α is the learning factor, γ is the discount factor.
6. An adaptive flight control method based on deep learning and reinforcement learning according to claim 1, characterized in that, It also includes Step 4: Obtain the state data of each subsystem of the aircraft during flight, perform real-time fault diagnosis on the aircraft according to the state data, and start the backup control strategy when the aircraft fails.
7. An adaptive flight control method based on deep learning and reinforcement learning according to claim 1, characterized in that The flight phase prediction model described in Step 1 is a feedforward neural network, convolutional neural network, recurrent neural network, generative adversarial network, or deep neural network.
8. An adaptive flight control system based on deep learning and reinforcement learning, characterized in that, It includes: A prediction module for training a flight phase prediction model according to the flight state data of the aircraft in each flight phase, environmental data, and corresponding meteorological conditions. The trained flight phase prediction model outputs the optimal flight phase of the aircraft under the meteorological conditions; The reinforcement learning module is used to construct a Q-table with the flight state and environmental parameters as the state and the control instruction as the action, determine the action with the highest Q-value in the current state according to the current action of the aircraft and in combination with the ε-greedy strategy, take the action with the highest Q-value at the current moment as the action for the next moment, and determine the optimal control strategy according to the action with the highest Q-value at each moment; The control module is used to control the flight of the aircraft in the optimal flight phase under meteorological conditions according to the optimal control strategy.
9. An electronic device, characterized in that, It includes: A memory for storing computer programs; A processor for implementing the steps of the adaptive flight control method based on deep learning and reinforcement learning according to any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps of the adaptive flight control method based on deep learning and reinforcement learning according to any one of claims 1-7.