A method, system and device for controlling the landing process of an aircraft in the presence of crosswind disturbances
By sharing control pre-training and course reinforcement learning, complex wind field disturbances are identified and processed in real time and graded, generating residual incremental control commands. This solves the problem of insufficient response to sudden disturbances during aircraft landing and achieves stable flight attitude and track tracking.
Patent Information
- Application Number
- CN202610421637.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-01
- Publication Date
- 2026-06-19
AI Technical Summary
Existing aircraft landing control methods are insufficient in response to sudden disturbances and struggle to cope with complex dynamic wind fields, resulting in an inability to maintain stable flight attitude and good track tracking capabilities.
By sharing the pre-training phase of control and the fine-tuning phase of course reinforcement learning, complex wind field disturbances such as crosswinds, gusts, and wind shears are identified and processed in real time and at different levels. A wind disturbance curriculum containing multiple training phases is constructed, and the control strategy is constrained by the composite reward function of KL divergence regularization term to generate residual incremental control commands for the aircraft control surfaces and throttle.
It achieves stable flight attitude and good trajectory tracking capability of aircraft under complex wind fields, avoids the overcorrection and correction lag problems in traditional methods, and improves anti-disturbance capability and adaptability.
Smart Images

Figure CN122239791A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of fixed-wing aircraft landing control technology, specifically to a method, system, and device for controlling the aircraft landing process in response to crosswind disturbances. Background Technology
[0002] The landing phase is one of the most complex, sensitive, and environmentally vulnerable phases of an aircraft's flight mission. Especially under conditions of crosswinds, wind shear, and turbulence, the aircraft's attitude control, heading alignment, descent trajectory maintenance, and touchdown stability are significantly affected. In aviation, crosswinds refer to winds that form an angle with the runway direction. When the crosswind component is significant, it directly impacts the aircraft's sideslip angle, roll angle, and heading maintenance. During the low-altitude, low-speed landing phase, aerodynamic stability is significantly reduced, and any lateral disturbance can cause the aircraft to deviate from the glide path or lose runway alignment. Crosswinds cause the aircraft to drift towards the windward side, requiring real-time corrections via rudder, ailerons, elevators, and other actuators. However, sudden changes in wind direction and speed often exceed the response range of traditional control systems, leading to attitude oscillations, increased sideslip angles, and even failure to land safely. Wind shear is also a common and strong disturbance during landing, typically manifesting as sudden changes in wind speed or direction over a short distance, causing sudden changes in lift, abnormal descent rate, or attitude instability. Improper handling of wind shear can cause an aircraft to suddenly lose altitude or experience abnormal pitch changes, increasing the difficulty of operation for the pilot and the automatic control system.
[0003] In existing aircraft landing control technologies, in order to improve the system's adaptability in complex aerodynamic environments, the existing control methods proposed by researchers can be divided into three categories: (1) A large number of model-based control methods have been proposed, such as Model Reference Control (MRC), Gain-Scheduled Control, and Linear Parameter Variation Control (LPV-Control). These methods rely on pre-established aerodynamic models and wind field models of the aircraft and adapt to different flight states by dynamically adjusting control parameters. However, due to the significant influence of terrain, temperature gradient, and random turbulence on the actual wind field, model errors are difficult to avoid. At the same time, LPV, MRC, and gain-scheduled control methods often require a large number of parameter calibrations in engineering implementation and are highly sensitive to model accuracy. Once the model does not match the real environment, it is easy to cause a decline in control performance. (2) A variety of nonlinear control methods have been proposed, including Dynamic Inversion Control, Adaptive Backstepping Control, and Sliding Mode Control. Control), adaptive second-order sliding mode control, and robust control methods combined with μ synthesis, etc.; these methods can handle the dynamic changes of aircraft under high angle of attack and strong nonlinear conditions, and improve the anti-disturbance capability within a certain range; however, nonlinear control generally has problems such as complex design, difficult parameter tuning, and high dependence on model structure; at the same time, since the landing phase involves multiple target coupling quantities such as altitude, pitch, roll, heading, and airspeed, a single nonlinear control method often cannot simultaneously take into account multiple control requirements, and is prone to control surface saturation or coupling conflict under strong wind interference; (3) learning-based control methods have been proposed, such as imitation learning based on teaching data, behavior cloning, and deep reinforcement learning based on interactive experiments; imitation learning relies on the pilot's demonstration trajectory and can reproduce human control logic within a limited range, but since flight tests cannot cover extreme crosswinds and sudden wind shear, its generalization ability is limited. While reinforcement learning methods have the advantage of not relying on explicit models, landing tasks are characterized by long durations, strong state coupling, and high risks, resulting in slow convergence of the learning process and extreme sensitivity to wind field distribution. Furthermore, under conditions of uneven wind direction distribution, random wind speed drift, or no wind field observed, reinforcement learning methods are prone to policy degradation or trajectory instability.
[0004] In summary, current aircraft landing process control methods are insufficient in response to sudden disturbances. Their parameters are usually fixed with design conditions, making it difficult to cope with complex dynamic wind fields, resulting in the inability to maintain stable flight attitude and good track tracking capabilities. Summary of the Invention
[0005] To address the shortcomings of existing technologies in responding to sudden disturbances and coping with complex dynamic wind fields, this invention proposes an aircraft landing process control method, system, and device oriented towards crosswind disturbances. By sharing a control pre-training phase and a course reinforcement learning fine-tuning phase, it achieves real-time identification, hierarchical processing, and control compensation for complex wind field disturbances such as crosswinds, gusts, and wind shear. This enables the aircraft to maintain a stable flight attitude and good track tracking capability even when the near-surface wind field is highly random and changes rapidly, thereby solving the problems existing in the prior art.
[0006] A control method for aircraft landing process oriented towards crosswind disturbances includes the following steps: Real-time acquisition of aircraft flight status data and environmental wind field data; Flight status data and environmental wind field data are input into a pre-trained shared control model to obtain basic control prior outputs; wherein, the shared control model is trained by fusing expert policies from multiple flight control sub-tasks; A wind disturbance curriculum is constructed, comprising multiple training stages. Each training stage corresponds to a different set of wind disturbance parameters, and the wind disturbance intensity of each training stage increases progressively according to a preset rule. The basic control prior output is used as the initial control strategy, and training begins from the stage with the lowest wind disturbance intensity in the wind disturbance curriculum. During each training process, the deviation between the current control strategy and the policy distribution of the shared control model is constrained by a composite reward function containing a KL divergence regularization term to update the control strategy. After the performance of the control strategy in the current training stage reaches a preset threshold, the training automatically switches to the next scenario with higher wind disturbance intensity until a final control strategy capable of stable landing under the highest wind disturbance intensity is generated. Based on the final control strategy, residual incremental control commands are generated for the aircraft control surfaces and throttle to complete landing control against crosswind disturbances.
[0007] Furthermore, the training process of the shared control model includes the following steps: The aircraft landing control task is decomposed into four independent sub-tasks: attitude control, speed control, altitude control, and heading control. Based on the near-end policy optimization algorithm, the algorithm is trained in an independent simulation environment for each sub-task to obtain four sub-task expert policies. Collect state-action data of various expert strategies during the training process, and perform knowledge distillation and fusion on the state-action data through behavior cloning to generate a unified shared control model.
[0008] Furthermore, the wind disturbance parameters include at least the crosswind duration, wind duration, and maximum wind speed.
[0009] Furthermore, the composite reward function containing the KL divergence regularization term is expressed as: ; ; Among them, As the current strategy, For pre-trained expert strategies, Regularization weight hyperparameter; trajectory tracking reward Used to penalize 3D yaw deviation, terminal landing reward A sparse reward structure is adopted, and a reward is given all at once when all landing conditions are met; The reward function for the attitude control subtask; The KL divergence between the current policy and the policy distribution of the pre-trained shared control model; This represents the cumulative reward obtained from completing a task.
[0010] Furthermore, the final control strategy is to adjust the aileron deflection angle, elevator deflection angle, rudder deflection angle, and throttle opening of the aircraft at the current moment. This adjustment will be added to the actual control input at the previous moment to obtain the actual control command to be executed at the current moment.
[0011] The present invention also includes an aircraft landing process control system oriented towards crosswind disturbances, comprising: The acquisition module is used to acquire real-time flight status data and environmental wind field data of the aircraft; The prior output module is used to input flight state data and environmental wind field data into a pre-trained shared control model to obtain basic control prior output; wherein, the shared control model is trained by fusing expert policies of multiple flight control sub-tasks; The training module is used to construct a wind disturbance curriculum containing multiple training stages. Each training stage corresponds to a different set of wind disturbance parameters, and the wind disturbance intensity of each training stage increases progressively according to a preset rule. The basic control prior output is used as the initial control strategy, and training starts from the stage with the lowest wind disturbance intensity in the wind disturbance curriculum. During each training process, the deviation between the current control strategy and the policy distribution of the shared control model is constrained by a composite reward function containing a KL divergence regularization term to update the control strategy. After the performance of the control strategy in the current training stage reaches a preset threshold, the module automatically switches to the next scenario with higher wind disturbance intensity to continue training until a final control strategy capable of stable landing under the highest wind disturbance intensity is generated. The control module is used to generate residual incremental control commands for the aircraft's control surfaces and throttle based on the final control strategy, in order to complete landing control against crosswind disturbances.
[0012] The present invention also includes a computer device for controlling the aircraft landing process in response to crosswind disturbances, comprising: a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the aircraft landing process control method in response to crosswind disturbances.
[0013] The present invention also includes a readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, are used to perform the steps of the aircraft landing process control method for crosswind disturbances.
[0014] This invention provides a control method for aircraft landing process oriented towards crosswind disturbances, which has the following beneficial effects: This invention obtains a "shared control model" by integrating expert strategies from multiple flight control sub-tasks. This eliminates the need for subsequent learning in complex environments to begin with disordered random actions, effectively avoiding high trial-and-error costs and the risk of losing control, thus improving the overall stability of training. By designing a "wind disturbance curriculum mechanism," the intensity and variation patterns of various crosswind disturbances are organized into training sequences from easy to difficult, allowing the control strategy to gradually adapt to complex wind field conditions in a natural and progressive manner. By constructing a landing control architecture oriented towards crosswind disturbances, this invention achieves real-time identification, hierarchical processing, and control compensation for complex wind field disturbances such as crosswinds, gusts, and wind shear. This enables the aircraft to maintain a stable flight attitude and good track tracking capability even in near-surface wind fields with strong randomness and rapid changes, significantly improving the aircraft's anti-disturbance and adaptability during landing in complex dynamic wind fields, and avoiding problems such as over-correction and correction lag that occur in traditional methods when wind direction changes abruptly. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of the overall architecture of the training framework in an embodiment of the present invention; Figure 2 This is a schematic diagram comparing direct action output and residual action output in an embodiment of the present invention; wherein, Figure 2 (a) in the diagram is a schematic diagram of the non-residual motion output result; Figure 2 (b) in the diagram is a schematic diagram of the residual motion output result; Figure 3 This is a schematic diagram comparing the reward curves of the baseline method in an embodiment of the present invention; wherein, Figure 3 (a), (b), and (c) in the figure represent the convergence curves of the present invention and the traditional reinforcement learning algorithms PPO and DDPG under the scenarios of strong crosswind ingress, constant wind field, and random wind field, respectively. Figure 4 This is a schematic diagram of the baseline method trajectory visualization results in an embodiment of the present invention; Figure 5 This is a schematic diagram of an ablation experiment using the shared control module in an embodiment of the present invention; wherein, Figure 5 (a) in the figure is a schematic diagram of the experimental results in a constant wind field mission. Figure 5 (b) in the diagram is a schematic diagram of the experimental results in a strong crosswind sudden change scenario; Figure 6 This is a visual schematic diagram of the landing process in an embodiment of the present invention; Figure 6 (a) in the diagram represents the start of the entire task. Figure 6 (b) and (c) in the diagram represent schematic diagrams of crosswinds with random direction and magnitude triggered in the environment. Figure 6 (d) and (e) in the diagram represent a successful landing near the runway by adjusting and controlling the elevator, rudder, throttle, and ailerons despite wind interference. Figure 6 (f) in the diagram represents the aircraft finally making contact with the ground and entering the taxiing phase; Figure 7 This is a schematic flowchart of an aircraft landing process control method oriented towards crosswind disturbance in an embodiment of the present invention. Detailed Implementation
[0016] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0017] In aviation, crosswinds refer to winds that form an angle with the runway direction. When the crosswind component is significant, it directly affects the aircraft's sideslip angle, roll angle, and heading maintenance. Aircraft landing is not a single-objective control task, but rather a coordinated control process involving multiple dynamic state variables, such as: ① Maintain the correct heading and align it with the runway centerline.
[0018] ② Maintain a stable rate of decline and altitude variation.
[0019] ③ Adjust the pitch to achieve smooth leveling and grounding.
[0020] ④ Control the roll to prevent the wings from tilting and touching down.
[0021] ⑥ Maintain a reasonable airspeed to prevent stall or overshoot.
[0022] These control objectives are closely coupled. For example, to counteract lateral drift caused by crosswinds, yaw correction is required via the rudder, but this action simultaneously triggers changes in sideslip angle and roll. Adjusting the throttle to control the rate of descent affects airspeed stability, thus impacting pitch control effectiveness. The actions between the actuators are often not completely independent. When one control surface becomes saturated or lags in response, other actuators must compensate, which can easily lead to control conflicts or system oscillations. Most existing control methods employ multi-loop independent designs with insufficient coordination between loops, making them highly susceptible to "mutual interference" in complex wind fields. For instance, yaw control loop corrections may negatively impact the roll control loop, causing continuous yaw; altitude and rate of descent control loops may experience pitch-up / pitch-down oscillations due to throttle correction lag. Traditional control structures struggle to handle multiple coupled objectives simultaneously, especially under significant external disturbances, making them prone to inconsistencies, slow responses, or command conflicts. Moreover, the wind field in the actual airport environment has significant uncertainty and randomness, which manifests as: continuous changes in wind speed, multiple deflections of wind direction, different wind field intensities at different altitudes (vertical wind shear), and unpredictable disturbances such as sudden gusts and turbulence. Traditional control methods usually rely on accurate aerodynamic models, wind field estimation, or linearized control models. These models can achieve good control performance in the laboratory or ideal environment, but the following problems may occur in the actual random wind field: (1) Model mismatch leads to performance degradation: When the actual wind field or aerodynamic parameters are inconsistent with the design model, the controller output may not be able to effectively resist disturbances, resulting in deviations in descent rate, attitude, or heading. (2) Difficulty in adapting to different aircraft types and different environments: The structure, weight, and aerodynamic characteristics of each aircraft type are different, requiring remodeling, parameter tuning, or even redesign of the control law, which greatly reduces the adaptability and engineering practicality of the system. (3) The actual wind field distribution is complex and lacks predictability: Some traditional control algorithms assume that the wind field is stable or slowly changing, while the actual flight environment is usually much more complex than the model, which makes it difficult for the control system to maintain consistent performance under multiple conditions.
[0023] Based on this, this invention proposes a comprehensive intelligent control method capable of maintaining stability under complex wind fields, strong random disturbances, and multi-objective coupling conditions. This method consists of two stages, such as... Figure 1 As shown, the shared control pre-training stage and the course reinforcement learning fine-tuning stage are respectively, with the goal of jointly solving the sample efficiency bottleneck under multi-objective control and the policy generalization problem in wind disturbance environment.
[0024] (1) First stage: Pre-training and online data collection process based on multi-expert models.
[0025] To address the coupling relationship among the four key objectives of "attitude control, speed control, altitude control, and heading control" in landing missions, this invention breaks down the landing process into four independent basic control tasks and constructs corresponding sub-task training environments for each. Each sub-task environment is built based on the aircraft's dynamic characteristics. Its inputs include necessary flight states such as current attitude angle, speed, altitude, heading deviation, and crosswind information, while its outputs are basic control variables such as ailerons, elevators, rudder, and throttle.
[0026] In each sub-task environment, this invention uses a step-by-step, test-and-train approach to enable the control strategy to independently learn how to achieve stable control under a single objective. For example, the attitude control sub-task is responsible for maintaining stable roll and pitch angles, the altitude control sub-task is responsible for keeping the aircraft near a designated descent altitude, the speed control sub-task ensures a stable and controllable airspeed is reached before touchdown, and the heading control sub-task ensures the aircraft's nose remains aligned with the runway direction. Through this structured training method, four sub-task expert models with stable control capabilities in their respective tasks can be obtained.
[0027] Subsequently, this invention unifies and summarizes the large amount of flight control data generated during the training process of the four types of sub-tasks, and integrates them into a "shared control model" using imitation learning. This shared control model is not directly used to execute the complete landing mission, but serves as the "basic control capability library" for the entire system. Its function is equivalent to establishing a stable and reliable flight action baseline for the model, so that subsequent learning in complex environments does not have to start from disordered random actions, thereby effectively avoiding problems such as high trial-and-error costs and easy loss of control, and improving the overall stability of training.
[0028] (2) Second Stage: Supervised Policy Training Based on Disturbance Curriculum for Landing Mission Optimization. After completing the pre-training of the multi-expert model and online data collection, this invention enters the second stage, aiming to achieve stable control of the complete landing mission under crosswind disturbance conditions. Through the policy distillation method, this invention uses the expert model trained in the first stage as the teacher model to provide supervision for the student policy model in the second stage. This process adopts a "wind disturbance curriculum mechanism," which organizes crosswind disturbances of different intensities and changing patterns into training sequences from easy to difficult, enabling the control strategy to gradually adapt to complex wind field conditions.
[0029] In the initial training phase, wind speeds are low, crosswinds change slowly, and the flight environment is relatively stable, making it easy for the system to maintain aircraft attitude and glide path stability while controlling basic maneuvers. As training progresses, the instructor model gradually guides the student model to cope with stronger winds, more frequent gusts, and more complex wind direction changes, enabling the student model to learn how to maintain trajectory and attitude stability under more challenging wind conditions. This course-based training method effectively prevents the strategy from going out of control when first exposed to strong winds, ensuring a smooth learning process.
[0030] During strategy distillation, the teacher model provides action prediction and strategy guidance, enabling the student model to gradually improve its performance in complex environments while ensuring consistency between the student and teacher strategies. To further optimize control quality, this invention designs a composite evaluation mechanism that comprehensively considers indicators such as 3D trajectory error, attitude deviation, velocity deviation, and whether soft landing requirements are met, guiding the training process to consistently improve landing quality and stability.
[0031] To avoid control instability caused by abrupt policy updates under strong disturbances, this invention introduces a distributed constraint mechanism to limit the update magnitude of the student policy in each training iteration, ensuring necessary consistency between the control behavior and the shared control model, and improving the safety and stability of the training process. Through course training with progressively increasing disturbance difficulty, a unified intelligent landing control strategy capable of adapting to both structured and stochastic wind fields is ultimately obtained, ensuring a high landing success rate and trajectory stability even under conditions of high wind speeds and strong gusts.
[0032] Furthermore, to improve the smoothness of the control strategy application, this invention introduces "residual motion design" in motion modeling. The output of the control strategy not only provides the final control command but also the incremental adjustment based on the existing control values. Specifically, the strategy output is the incremental change in ailerons, elevators, rudders, and throttle, rather than the absolute values of these control surfaces. This "gradual fine-tuning" approach is more in line with the habits of pilots in actual operation, effectively reducing over-controlling caused by sudden changes in wind disturbance, achieving smoother control surface adjustments, and thus making the aircraft's attitude changes more stable. This method also reduces high-frequency oscillations in the control signal, improves numerical stability during training, and reduces the risk of aircraft instability caused by drastic changes in motion. Experimental results show that in high-dynamic wind disturbance scenarios, the strategy using residual motion design can significantly reduce the jitter of motion values, prevent instability caused by sudden changes in control values, and thus improve the feasibility of deploying the intelligent control system.
[0033] Based on the above two stages, such as Figure 7 As shown, the method proposed in this invention includes the following steps: S1. First, expert policy learning for sub-tasks is performed under no-demonstration conditions. To address the training instability and policy convergence difficulties caused by multi-objective control optimization in landing missions, this invention divides the entire control task into four basic sub-tasks: altitude control, velocity control, heading control, and attitude control. Each sub-task corresponds to an independent sub-environment and sub-objective, where expert policies are trained separately and subsequently fused through behavior cloning. Unlike traditional imitation learning methods that rely on high-quality human demonstrations, this invention employs a no-demonstration reinforcement learning approach to train the expert policy for each sub-task. Specifically, based on the Proximal Policy Optimization (PPO) algorithm, the agent obtains state-action-reward triples through direct interaction with the sub-task environment and uses this interaction data to update the policy model. Each sub-task has its own independent reward function design.
[0034] For altitude control, the agent is guided to maintain the flight altitude at the target altitude. The reward function for the vicinity is defined as follows: ; in This indicates the deviation between the current altitude and the target altitude. This is a normalized parameter for altitude deviation. For airspeed control: considering both vertical descent speed and angle of attack, the reward function is as follows:
[0035] ; in For the desired rate of descent, For the current aircraft descent rate, For the target angle of attack, For the current angle of attack, , For the corresponding normalization parameters.
[0036] For heading control: adjust the flight path angle Try to align with the direction of the runway. The ratio of the aircraft's lateral and longitudinal velocities represents the reward function. ; in, These are the normalized parameters for heading control.
[0037] For attitude control: Requires roll angle Yaw angle Keep it close to 0, pitch angle Define the reward function within a reasonable range: ; in Penalty for angles outside the pitch range. For the corresponding normalization parameters.
[0038] Each subtask uses the PPO's pruning objective function for policy updates, with the specific policy loss function as follows: ; The probability ratio is This indicates the state of the new and old strategies. Select action The probability ratio; It is an advantage estimate, used to measure the relative advantage of taking a certain action in a given state; This refers to the pruning range, controlling the magnitude of policy updates; combined with value function regression and entropy regularization, the complete loss function for the expert manipulation task is: ; in and These are the weights of the value function loss term and the regularization term. The value function loss term is: Measure the current value function With target value The difference between them; the strategy entropy is: ,in Is the strategy in the state? Select action The probability of . Among them, i This indicates the expert's ID number. Through the above-described structured subtask modeling and demonstration-free reinforcement learning training, this invention obtains four stable and goal-focused subtask expert policies. These expert models will subsequently be used for behavioral cloning distillation of the shared control model, providing high-quality control priors for the initialization of the main policy and laying the foundation for the overall method's high sample efficiency and stable training.
[0039] S2. After obtaining the shared control model as a pre-trained policy, this invention further fine-tunes it through Curriculum Reinforcement Learning (CRL) to improve the robustness and generalization ability of the policy under complex wind disturbance conditions.
[0040] The core idea of the course mechanism is to control the difficulty of the task in the early stage of training, and gradually increase the complexity of the perturbation so that the strategy can gradually adapt to the high uncertainty environment, thus avoiding training instability or even strategy collapse due to exposure to severe perturbation at the beginning.
[0041] Table 1 Wind Disturbance Course Setup This invention constructs a wind disturbance schedule, as shown in Table 1. The disturbance parameters in the schedule include: crosswind duration, wind duration, and maximum wind speed. During training, the schedule is smoothly arranged in stages, with wind disturbance conditions gradually transitioning from "easy" to "difficult". This disturbance scheduling mechanism effectively constructs a task distribution sequence from simple to difficult, promoting the strategy to acquire basic robustness in low-complexity environments and gradually migrate to high-dynamic environments.
[0042] To ensure the stability of policy training and the density of reward signals, this invention employs the following composite reward function during the fine-tuning phase: ; Among them, trajectory tracking rewards Used to penalize 3D yaw deviation, terminal landing reward A sparse reward structure is adopted, awarding a reward all at once when all landing conditions are met. In the optimization objective, a KL regularization term is introduced to limit the policy from deviating excessively from the pre-trained policy distribution during fine-tuning, thereby improving policy stability. The overall optimization objective is:
[0043] ; in As the current strategy, For pre-trained expert strategies, For regularization weight hyperparameters, This represents the cumulative reward obtained from a single task. The introduction of the KL regularization term can significantly alleviate the collapse divergence problem of the policy in highly dynamic wind fields, and improve the convergence stability and policy security during the high-disturbance phase.
[0044] Furthermore, during fine-tuning, the task is no longer segmented into stages; instead, the entire flight process is modeled as a unified trajectory tracking task. The training objective is to ensure that the flight trajectory continuously closely follows the target path, avoiding the optimization challenges caused by sparse stage rewards in traditional methods. As the course progresses with increasingly severe wind disturbances, the strategy gradually acquires the ability to cope with various disturbance types, ultimately resulting in a robust control strategy.
[0045] To comprehensively evaluate the effectiveness and generalization ability of the training framework proposed in this invention, large-scale simulation experiments were conducted on the high-fidelity six-degree-of-freedom flight simulation platform JSBSim. JSBSim supports flexibly configurable aerodynamic modeling and flight control interfaces, and is suitable for various fixed-wing aircraft control tasks.
[0046] The experiment used three typical aircraft models: F16, Cessna 172P, and Airbus A320. As shown in Table 2, three representative wind disturbance scenarios were designed, covering structural disturbances and high-dynamic disturbances respectively:
[0047] (1) Strong Crosswind Approach: Apply crosswind pulses in the middle and late stages of landing to simulate sudden wind disturbances; (2) Constant Crosswind Approach: Apply constant wind speed throughout the landing process to test stable control capability; (3) Stochastic Wind Field: Use fractional Ornstein-Uhlenbeck process to generate wind speed changes to achieve high randomness and high frequency disturbance conditions.
[0048] First, we verify the performance of the method of this invention in three highly random wind fields (here, landing success rate is used as the primary metric) and compare it with baseline methods. The baseline methods include: traditional linear flight control strategy PID controller, behavior cloning (BC), and mainstream deep reinforcement learning algorithms DDPG, PPO, and SAC.
[0049] Table 2 Comparison of task success rates using baseline methods The experimental results are shown in Table 2. The method of this invention achieved the highest landing success rate across all missions, especially under strong disturbance conditions (such as stochastic wind fields), where its success rate significantly outperformed existing mainstream methods. Under structured wind fields, SAC performed slightly better than PPO and DDPG, but its performance fluctuated significantly with increased crosswind suddenness and coupled dynamics. The method of this invention achieved a success rate of 88%–91% on Cessna 172P, F16, and A320, demonstrating strong generalization and wind field adaptability.
[0050] Convergence curves compared to baseline deep reinforcement learning algorithms are shown below. Figure 3 As shown, Figure 3Images (a), (b), and (c) show the convergence curves of this invention and traditional reinforcement learning algorithms PPO and DDPG under strong crosswind approach, constant wind field, and random wind field scenarios, respectively. It can be seen that traditional reinforcement learning methods such as PPO and DDPG generally suffer from slow convergence speed and large training fluctuations in wind-disturbed environments. Especially in the early stages, limited by the lack of prior initialization and high-risk exploration, their policies are prone to frequently falling into an out-of-control state, leading to reward collapse and affecting the overall gradient quality and policy stability. In contrast, the method of this invention exhibits faster convergence speed and smoother reward growth curves in all three wind fields. First, the shared control module provides stable priors for low-level flight control through distillation of expert sub-policies, providing reasonable basic control responses in the early stages of training, significantly reducing the risk of round collapse caused by trial and error from scratch. Second, the curriculum reinforcement mechanism reduces the difficulty of policy adaptation to high-complexity wind fields in the early stages of training through progressive scheduling from low-disturbance tasks to high-disturbance tasks, alleviating the distribution shift problem of early exploration of extreme samples.
[0051] like Figure 4 The trajectory visualization results further corroborate this process: in the initial training phase, neither DDPG nor PPO achieved effective control, and the flight trajectories generally exhibited significant yaw, unstable speed, and early loss of control and crashes. In contrast, the method of this invention, due to its behavior cloning initialization strategy, can generate feasible action outputs early on. Although a precise landing has not yet been achieved, the overall trajectory can maintain basic attitude and heading control. As training progresses to 100,000 and 500,000 steps, SC+CRL gradually learns to actively correct deviations under wind disturbances, the flight trajectory becomes smoother, and ultimately exhibits highly stable and generalizable landing control behavior. In contrast, even at the end of convergence, PPO and DDPG still exhibit sensitivity to the direction of disturbances and cross-scenario instability, indicating a lack of wind disturbance robustness modeling capability.
[0052] To evaluate the generalization ability of the SC+CRL algorithm in unknown strong random wind fields, this invention designs a test set with progressively increasing perturbation amplitude, controls the wind speed range from 0 to 50 ft / s, and sets a test interval of 60 ft / s beyond the training distribution range to test the out-of-distribution generalization ability of the strategy.
[0053] Table 3 Comparison of task generalization capabilities of baseline methods The experimental results are shown in Table 3. The traditional behavior cloning method exhibits a sharp performance drop after wind disturbances exceed the training trajectory, demonstrating severe overfitting. The success rates of DDPG and PPO significantly decrease after wind speeds exceed 30 ft / s, exhibiting policy collapse or output failure characteristics. In contrast, the SC+CRL method maintains stable performance under all wind disturbance intensities, especially in tasks where no wind field size was encountered during training, significantly outperforming the strongest baseline, SAC (13.7%). Furthermore, this invention conducted an ablation experiment to investigate the role of the curriculum mechanism: when curriculum scheduling was removed and training was conducted directly in a high-disturbance environment using only shared control, the policy success rate in wind field conditions not encountered during training decreased to 17.1%, verifying the role of the curriculum structure in modeling the disturbance distribution structure and generalization transfer.
[0054] This invention further verifies the effect of expert policy shared control through ablation experiments: To quantify the role of the shared control module in sample efficiency and policy stability, a controlled variable experiment was designed, comparing the training process initialized with distillation using 0, 1, 2, and 4 expert sub-policies. The average reward value curve for the first 200,000 steps was statistically analyzed, and the results are as follows: Figure 5 As shown in the figure. Experimental results show that when an expert-shared control pre-trained model is added, the early performance of the policy is higher and the convergence is faster. Without expert initialization, the policy needs to explore from scratch, resulting in frequent round failures in the initial stage and slow growth in average reward. Using only a single subtask expert can bring some stability improvement, but policy inconsistency still exists. In control theory, this mechanism is equivalent to providing a set of heuristic initial conditions in the policy space, effectively reducing the effective sampling area of the state-action space during the exploration and trial-and-error phase, while providing control balancing capabilities across subtasks, making it suitable for multi-objective coupled tasks such as fixed-wing aircraft landing.
[0055] Finally, the stability of various attitude angles during the flight and landing process was tested and visualized. A landing demo was then created using the Flight Gear visualization tool, demonstrating that the method of this invention can maintain a stable glide path and attitude even under wind interference. Figure 6 As shown, the aircraft was able to maintain a stable attitude and heading throughout the entire descent, from the start of landing until it touched down, demonstrating the feasibility of the method from a visual perspective.
[0056] The key points of this invention are: (1) A smart control architecture for the entire landing process of an aircraft oriented towards crosswind disturbances is constructed. This architecture is based on the characteristics of crosswinds, gusts and wind shear during landing, and sets up core modules including perception and information processing, wind field disturbance identification, task hierarchical decision-making and execution command coordination, so that the system can maintain stable control in the near-ground environment with extremely strong wind field randomness. This invention focuses on protecting the functional positioning, information flow mode and task logic relationship of each module in this system structure, so that the control system can make real-time and effective control responses to crosswind disturbances of different intensities and types without relying on high-precision aerodynamic models. This overall architecture cannot be achieved by existing linear control, model-based control and single nonlinear control structures. (2) A cooperative control mechanism suitable for multi-objective coupled processes in the landing phase is proposed. The landing task involves multiple control objectives such as pitch holding, roll stabilization, yaw alignment, descent speed control and track tracking. These objectives are significantly coupled and are prone to mutual interference when encountering crosswinds and wind shears. This invention establishes a coordination relationship between the task layer and the execution layer to achieve unified planning and instruction fusion of multiple control quantities, thereby avoiding the control surface conflict and loop confrontation problems common in traditional PID control, dynamic inversion control, sliding mode control and other methods under strong disturbances. This invention focuses on protecting the collaborative control logic, the division of labor among control quantities and the unified instruction fusion strategy, so that the system can maintain heading alignment, descent channel stability and attitude smoothness under random wind field changes. (3) A disturbance classification processing and robust control mechanism for random wind fields and sudden wind shear is constructed. This mechanism identifies the intensity and changing trend of crosswinds, gusts and wind shear in real time, classifies disturbances into different levels, and automatically switches the corresponding attitude stabilization strategy, wind drift compensation strategy and control authority allocation method according to the disturbance level. This method enables the system to maintain gentle and smooth control in light crosswinds, and to trigger enhanced attitude stabilization actions under moderate or strong sudden wind shear, thereby significantly improving the system's anti-disturbance capability. This disturbance classification and control mode switching mechanism does not rely on an accurate aerodynamic model, nor does it require complex gain scheduling logic. This is a significant innovation that distinguishes it from traditional model reference control, gain scheduling control, and learning-based policy generation methods. This invention focuses on protecting this disturbance identification method, its classification strategy, and its coupling relationship with the control strategy.
[0057] This invention constructs a landing control architecture oriented towards crosswind disturbances, achieving real-time identification, hierarchical processing, and control compensation for complex wind field disturbances such as crosswinds, gusts, and wind shear. This enables the aircraft to maintain stable flight attitude and good trajectory tracking capability even under conditions of strong randomness and rapid changes in near-surface wind fields. Compared with traditional linear control systems that rely on fixed parameters, this invention can automatically adjust the control focus according to the disturbance level, making the control commands more closely match the actual wind field conditions. This significantly improves the system's disturbance rejection and adaptability, avoiding problems such as over-correction and correction lag that occur in traditional methods when wind direction changes abruptly.
[0058] The multi-objective cooperative control mechanism proposed in this invention can coordinate multiple control variables such as pitch, roll, yaw, glide path, and airspeed. This mechanism solves the engineering problem of mutual interference and even control surface conflicts between different control loops in existing control methods, enabling the aircraft to maintain coordination among control objectives during landing and reducing attitude oscillations and control instability caused by loop conflict. Through the fusion and rational allocation of control commands, this invention significantly improves the aircraft's heading alignment and descent path maintenance capabilities during strong crosswind landings, resulting in a smoother landing process with less deviation.
[0059] The disturbance classification processing mechanism constructed in this invention enables the system to stably cope with various unknown wind field conditions without relying on accurate aerodynamic models or complex gain scheduling. Traditional model-based methods often experience a sharp decline in performance when aerodynamic parameters are inaccurate or environmental changes are significant. However, this invention, through real-time analysis of disturbance characteristics, enables the control system to proactively adapt to wind field changes, maintaining stable control performance under different aircraft types, airports, and wind field distributions, significantly improving the system's robustness and cross-scenario adaptability.
[0060] In summary, this invention achieves multi-objective stable coordination, rapid compensation for random disturbances, and cross-wind field adaptability, which are difficult to achieve with traditional control systems. It can significantly improve the landing safety and stability of aircraft under crosswind and wind shear conditions, and has outstanding engineering practical value and broad application potential.
[0061] Example 1: Landing control application of light trainer aircraft under moderate crosswind conditions The landing control system of this invention was deployed on a light fixed-wing trainer aircraft. The aircraft is equipped with basic weather detection devices, inertial navigation equipment, and conventional control surface actuators. The control structure proposed in this invention is integrated into the flight control computer, which identifies and determines the level of disturbance from the incoming crosswind by acquiring data such as airspeed, crosswind component, and attitude angle in real time. Under moderate crosswind conditions (approximately 8-10 m / s), the system automatically enhances roll stability and heading alignment capabilities, coordinating the three key tasks of glide path maintenance, attitude stabilization, and wind deflection compensation. During actual landing, the aircraft's roll angle deviation is significantly suppressed, heading alignment is achieved more quickly, the glide path is smoother, and the pilot can complete a safe landing without frequent manual corrections. This embodiment demonstrates that the invention possesses good crosswind resistance and engineering feasibility on typical training aircraft.
[0062] Example 2: Landing control of regional airliners under gust and wind shear conditions The control system of this invention was applied on a regional jet test platform. This aircraft has comprehensive meteorological data access capabilities, enabling real-time acquisition of gust and wind speed gradient changes. When the aircraft descended to near-ground altitude (approximately 60-30 meters), it encountered a sudden wind shear, with wind speed and direction changing rapidly within a short period. This invention can immediately identify the disturbance level and automatically switch to a control mode suitable for sudden disturbances, enhancing pitch attitude stability and yaw coordination, while adjusting throttle commands to suppress descent rate fluctuations. In actual testing, the system restored attitude balance shortly after the wind shear occurred, the roll deviation did not exceed safe limits, and the glide path remained continuous and stable, without any runway deviation trend. This embodiment demonstrates that this invention can effectively handle sudden, strong disturbances and improve landing safety under wind shear conditions.
[0063] Example 3: Cross-scenario verification of fixed-wing UAVs under different airport wind field conditions This invention was applied to a fixed-wing unmanned aerial vehicle (UAV) to verify its adaptability under different airport wind field conditions. Since the UAV itself lacks a precise aerodynamic model, the robustness of the control system is crucial. Tests were conducted at coastal, valley, and plain airports, each with significantly different crosswind characteristics, gust frequencies, and wind shear features. The disturbance identification and control coordination mechanism of this invention can automatically adjust wind yaw compensation and attitude stabilization based on real-time wind field characteristics. At coastal airports, the system effectively suppressed yaw caused by persistent crosswinds; at valley airports, the system quickly stabilized its attitude in the face of frequent gusts; and at plain airports with frequent wind shear, the descent path remained safe and continuous. Tests show that this invention achieves stable control performance at different types of airports, demonstrating high cross-scenario adaptability and good versatility.
[0064] Based on the same inventive concept, this invention also proposes an aircraft landing process control system oriented towards crosswind disturbances, comprising: The acquisition module is used to acquire real-time flight status data and environmental wind field data of the aircraft.
[0065] The prior output module is used to input flight state data and environmental wind field data into the pre-trained shared control model to obtain the basic control prior output; the shared control model is trained by fusing expert policies from multiple flight control sub-tasks.
[0066] The training module is used to construct a wind disturbance curriculum containing multiple training stages. Each training stage corresponds to a different set of wind disturbance parameters, and the wind disturbance intensity of each training stage increases progressively according to preset rules. The basic control prior output is used as the initial control strategy, and training starts from the stage with the lowest wind disturbance intensity in the wind disturbance curriculum. During each training process, the deviation between the current control strategy and the shared control model strategy distribution is constrained by a composite reward function containing a KL divergence regularization term to update the control strategy. After the performance of the control strategy in the current training stage reaches a preset threshold, the module automatically switches to the next scenario with higher wind disturbance intensity to continue training until a final control strategy capable of stable landing under the highest wind disturbance intensity is generated.
[0067] The control module is used to generate residual incremental control commands for the aircraft's control surfaces and throttle based on the final control strategy, in order to complete landing control against crosswind disturbances.
[0068] The present invention also proposes a computer device for controlling the aircraft landing process in response to crosswind disturbances, comprising: a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the aircraft landing process control method in response to crosswind disturbances.
[0069] The present invention also proposes a readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, are used to perform steps of an aircraft landing process control method oriented towards crosswind disturbances.
[0070] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A control method for aircraft landing process oriented towards crosswind disturbance, characterized in that, Includes the following steps: Real-time acquisition of aircraft flight status data and environmental wind field data; Flight status data and environmental wind field data are input into a pre-trained shared control model to obtain basic control prior outputs; wherein, the shared control model is trained by fusing expert policies from multiple flight control sub-tasks; A wind disturbance curriculum is constructed, comprising multiple training stages. Each training stage corresponds to a different set of wind disturbance parameters, and the wind disturbance intensity of each training stage increases progressively according to a preset rule. The basic control prior output is used as the initial control strategy, and training begins from the stage with the lowest wind disturbance intensity in the wind disturbance curriculum. During each training process, the deviation between the current control strategy and the policy distribution of the shared control model is constrained by a composite reward function containing a KL divergence regularization term to update the control strategy. After the performance of the control strategy in the current training stage reaches a preset threshold, the training automatically switches to the next scenario with higher wind disturbance intensity until a final control strategy capable of stable landing under the highest wind disturbance intensity is generated. Based on the final control strategy, residual incremental control commands are generated for the aircraft control surfaces and throttle to complete landing control against crosswind disturbances.
2. The aircraft landing process control method for crosswind disturbance as described in claim 1, characterized in that, The training process of the shared control model includes the following steps: The aircraft landing control task is decomposed into four independent sub-tasks: attitude control, speed control, altitude control, and heading control. Based on the near-end policy optimization algorithm, the algorithm is trained in an independent simulation environment for each sub-task to obtain four sub-task expert policies. Collect state-action data of various expert strategies during the training process, and perform knowledge distillation and fusion on the state-action data through behavior cloning to generate a unified shared control model.
3. The aircraft landing process control method for crosswind disturbance as described in claim 1, characterized in that, The composite reward function including the KL divergence regularization term is expressed as: ; ; Among them, As the current strategy, For pre-trained expert strategies, Regularization weight hyperparameter; trajectory tracking reward Used to penalize 3D yaw deviation, terminal landing reward A sparse reward structure is adopted, and a reward is given all at once when all landing conditions are met; The reward function for the attitude control subtask; The KL divergence between the current policy and the policy distribution of the pre-trained shared control model; This represents the cumulative reward received in a single task. Indicates the use of strategy Sampling trajectory Expected cumulative return This is a discount factor for future steps. This is the discount factor.
4. The aircraft landing process control method for crosswind disturbance as described in claim 1, characterized in that, The final control strategy involves adjusting the aileron deflection, elevator deflection, rudder deflection, and throttle opening of the aircraft at the current moment. This adjustment is added to the actual control input at the previous moment to obtain the actual control command to be executed at the current moment.
5. The aircraft landing process control method for crosswind disturbance according to claim 1, characterized in that, The wind disturbance parameters include at least the crosswind duration, wind duration, and maximum wind speed.
6. A control system for aircraft landing process oriented to crosswind disturbances, characterized in that, include: The acquisition module is used to acquire real-time flight status data and environmental wind field data of the aircraft; The prior output module is used to input flight state data and environmental wind field data into a pre-trained shared control model to obtain basic control prior output; wherein, the shared control model is trained by fusing expert policies of multiple flight control sub-tasks; The training module is used to construct a wind disturbance curriculum containing multiple training stages. Each training stage corresponds to a different set of wind disturbance parameters, and the wind disturbance intensity of each training stage increases progressively according to a preset rule. The basic control prior output is used as the initial control strategy, and training starts from the stage with the lowest wind disturbance intensity in the wind disturbance curriculum. During each training process, the deviation between the current control strategy and the policy distribution of the shared control model is constrained by a composite reward function containing a KL divergence regularization term to update the control strategy. After the performance of the control strategy in the current training stage reaches a preset threshold, the module automatically switches to the next scenario with higher wind disturbance intensity to continue training until a final control strategy capable of stable landing under the highest wind disturbance intensity is generated. The control module is used to generate residual incremental control commands for the aircraft's control surfaces and throttle based on the final control strategy, in order to complete landing control against crosswind disturbances.
7. A computer device for controlling the landing process of an aircraft in response to crosswind disturbances, characterized in that, include: A memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the aircraft landing process control method for crosswind disturbances as described in any one of claims 1-5.
8. A readable storage medium, characterized in that, The readable storage medium stores a computer program, which includes program instructions that, when executed by a processor, are used to perform the steps of the aircraft landing process control method for crosswind disturbances as described in any one of claims 1-5.