An autonomous take-off and landing guidance method based on an end-to-end large model
By employing an end-to-end large-scale model-based autonomous takeoff and landing guidance method, the problems of spatiotemporal alignment, four-dimensional modeling, and aerodynamic disturbances in complex environments for eVTOL automatic landing were solved, thereby improving safety and stability and ensuring efficient automatic landing of eVTOL in urban environments.
Patent Information
- Application Number
- CN202511377317.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing eVTOL automatic landing technology suffers from problems such as spatiotemporal alignment difficulties, insufficient four-dimensional environment modeling, lack of near-ground aerodynamic disturbance modeling, lack of cross-altitude guidance consistency, insufficient airborne real-time performance, and excessive computation and energy consumption in complex urban environments, resulting in insufficient safety and stability.
An autonomous takeoff and landing guidance method based on an end-to-end large model is adopted. Spatiotemporal alignment is achieved through hardware synchronization and spatial calibration. Transformer feature encoding and depth estimation are performed using a multi-task perception large model. Dynamic prediction is performed by combining fluid dynamics priors and a spatial RNN model to generate phased guidance and control commands, thereby enabling the automatic landing of the aircraft.
It improves the safety, stability and engineering availability of eVTOL automatic landing, solves complex technical challenges such as spatiotemporal alignment, four-dimensional environment modeling, near-ground aerodynamic disturbances and airborne real-time performance, and achieves efficient automatic landing in complex environments.
Smart Images

Figure CN120871894B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical fields of electric vertical takeoff and landing aircraft and path planning, and in particular to an autonomous takeoff and landing guidance method based on an end-to-end large model. Background Technology
[0002] With the accelerated application of eVTOL (electric Vertical Take-Off and Landing) in urban air traffic (UAM), the automatic landing phase has become a key bottleneck for overall aircraft safety and availability. Existing engineering approaches often employ a "multi-sensor fusion + segmented processing pipeline" paradigm: under the joint perception of GNSS / RTK, lidar, millimeter-wave radar, and visual sensors, target / landing zone detection, local / global mapping, trajectory planning, and flight control execution are completed sequentially. This paradigm achieves good results in open airports or well-marked test environments, but it reveals significant limitations in high-constraint scenarios such as complex urban canyons, rooftop helipads, and temporary emergency landing sites: reliance on expensive and heavy active sensors, sensitivity to environmental factors such as electromagnetic shielding, multipath propagation, and rain / fog, closed-loop lag due to module-level latency accumulation, and difficulty in stably completing terminal guidance under complex aerodynamic disturbances. Therefore, common technical solutions in this field still suffer from the following complex technical problems that urgently require improvement:
[0003] First, there are difficulties in spatiotemporal alignment and the accumulation of pipeline-level delays. Under conditions of high-speed descent, strong body attitude changes, and rolling shutter speeds, the timestamps of external participants are prone to drift when multiple cameras / sensors are used. The traditional serial pipeline of "detection → mapping → planning → control" introduces encoding / decoding and buffering at each stage, resulting in an "information time lag" between the perception results and the control commands, which is particularly dangerous in the near-ground phase.
[0004] Second, there is a lack of a unified four-dimensional (3D + time) environmental representation. Existing solutions are mostly based on instantaneous 3D point clouds or static grids, and use heuristic temporal smoothing to compensate for dynamism. They are difficult to express the dynamic flow of obstacles / landable areas that "occupy over time". They are insufficient in predicting dynamic objects such as moving pedestrians, vehicles, and wind-driven drifting debris, resulting in unreliable end-point obstacle avoidance and landing point selection.
[0005] Third, the strong coupling of aerodynamic / fluid disturbances in the near-ground segment has not been modeled. After eVTOLs descend and leave the upper atmosphere, they are significantly affected by ground effects, downwash wakes, crosswind shear, and turbulence around buildings, resulting in significant changes in lift and drag characteristics, lateral coupling, and attitude response. Existing plans often roughly equate these disturbances to external disturbances or safety margins, lacking a mechanism to jointly model fluid dynamics priors and visual environmental cognition for dynamic prediction of landing points and guided convergence.
[0006] Fourth, there is a lack of consistent guidance strategies across altitudes. Most existing solutions use a single cost function / fixed safety potential radius, which does not reflect the differences in altitude-related sensor reliability, parallax scale, ground effect risk, and landing surface texture resolution. This makes it difficult to achieve strategy switching and dynamic shrinkage of the safety radius in the "high → medium → near ground" phase.
[0007] Fifth, computational / energy consumption and airborne deployment constraints. In existing solutions, multi-model cascading is difficult to meet near-ground high-frequency closed-loop requirements when airborne computing power and thermal design are limited; the lack of multi-task collaboration to simultaneously complete Transformer feature encoding, semantic segmentation and depth estimation in a unified feature space leads to redundant computation and excessive power consumption.
[0008] Sixth, the landing point “prediction-guidance-execution” link is fragmented. Existing solutions often select the landing point based on rules or instantaneous geometric centers and then hand it over to a general MPC for convergence; there is a lack of a sequence model for “landing point evolution over time” to predict and generate phased guidance, which makes it prone to convergence oscillations or over-conservatism under dynamic obstacles, gusts and near-ground aerodynamic disturbances.
[0009] Seventh, the calibration and maintenance costs of the engineering implementation are high. Once the spatial calibration / time synchronization of multi-source hardware drifts, it requires downtime for maintenance; the lack of unified management of spatiotemporal alignment as algorithm input results in insufficient reliability and maintainability of the system in long-term operation.
[0010] Based on the above challenges, it is necessary to propose an eVTOL autonomous driving landing guidance method that can solve complex technical problems such as spatiotemporal alignment, four-dimensional environment modeling, near-ground aerodynamic coupling, cross-altitude guidance consistency and airborne real-time performance. Summary of the Invention
[0011] To address the shortcomings of the existing technologies, this invention provides an autonomous takeoff and landing guidance method based on an end-to-end large model. This method can solve complex technical challenges in spatiotemporal alignment, four-dimensional environment modeling, near-ground aerodynamic disturbances, cross-altitude guidance consistency, and airborne real-time performance for eVTOL automatic landing, thereby achieving a comprehensive improvement in safety, stability, and engineering usability.
[0012] The autonomous takeoff and landing guidance method based on an end-to-end large model provided by this invention includes:
[0013] Hardware synchronization and spatial calibration are performed on the simulated landing vehicle to achieve spatiotemporal alignment of the simulated landing vehicle, generate a unified time series of the vehicle hardware, and collect image data reflecting the landing environment of the vehicle in real time through multiple cameras of the vehicle. The image data is then deblurred to obtain pure visual landing environment image data with enhanced motion fuzzy control.
[0014] The pure visual landing environment image data is input into the end-to-end multi-task perception model. The end-to-end multi-task perception model performs Transformer feature encoding, semantic segmentation and depth estimation on the pure visual landing environment image data, and outputs landing environment analysis data including a three-dimensional BEV space map, a semantic segmentation map of the landing area and obstacles and image resolution depth estimation.
[0015] The landing environment analysis data, fluid dynamics prior data, and time series generated during spatiotemporal alignment are input into a spatial RNN model. The spatial RNN model dynamically predicts the landing point of the aircraft based on the landing environment analysis data, the fluid dynamics prior data, and the time series to generate landing phase guidance outputs adapted to different landing altitudes.
[0016] The landing phase guidance is acquired, and the landing of the aircraft is guided and controlled in stages according to the acquired landing phase guidance. Control commands are generated and sent to the aircraft's actuators for execution to guide the aircraft to land automatically.
[0017] Compared with the prior art, the beneficial effects of this invention are as follows:
[0018] This invention provides an autonomous takeoff and landing guidance method based on an end-to-end large model. The method includes: performing hardware synchronization and spatial calibration on a simulated landing vehicle to achieve spatiotemporal alignment, generating a unified time series of the vehicle hardware; acquiring image data reflecting the landing environment in real time through multiple cameras on the vehicle, and deblurring the image data to obtain pure visual landing environment image data with enhanced motion fuzzy control; inputting the pure visual landing environment image data into an end-to-end multi-task perception large model, whereby the end-to-end multi-task perception large model performs Transformer feature encoding, semantic segmentation, and depth estimation on the pure visual landing environment image data, and outputs a packet... This method incorporates landing environment analysis data, including a 3D BEV spatial map, semantic segmentation maps of the landing area and obstacles, and image resolution depth estimation. The landing environment analysis data, prior hydrodynamic data, and time series generated during spatiotemporal alignment are input into a spatial RNN model. This spatial RNN model dynamically predicts the aircraft's landing point based on the landing environment analysis data, prior hydrodynamic data, and the time series, generating landing phase guidance outputs adapted to different landing altitudes. The landing phase guidance is acquired, and based on this guidance, the aircraft's landing is controlled in stages. Control commands are generated and issued to the aircraft's actuators for execution, guiding the aircraft to automatic landing. This method can solve complex technical challenges in eVTOL automatic landing, such as spatiotemporal alignment, 4D environment modeling, near-ground aerodynamic disturbances, cross-altitude guidance consistency, and airborne real-time performance, achieving a comprehensive improvement in safety, stability, and engineering usability. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. Some specific embodiments of the invention will be described in detail below with reference to the accompanying drawings in an exemplary and non-limiting manner. The same reference numerals in the drawings designate the same or similar parts or components. It should be understood by those skilled in the art that these drawings are not necessarily drawn to scale. In the drawings:
[0020] Figure 1 This is a flowchart illustrating an autonomous takeoff and landing guidance method based on an end-to-end large model according to an embodiment of the present invention.
[0021] Figure 2 This is a schematic diagram of an architecture of an autonomous take-off and landing guidance system based on an end-to-end large model according to an embodiment of the present invention. Detailed Implementation
[0022] To enable those skilled in the art to better understand the present invention, the technical solutions of this embodiment will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0023] See Figures 1-2 This embodiment provides an autonomous takeoff and landing guidance method based on an end-to-end large model. This method can be used in an autonomous takeoff and landing guidance system based on an end-to-end large model. The method includes the following steps:
[0024] S101. Perform hardware synchronization and spatial calibration processing on the proposed landing vehicle to complete the spatiotemporal alignment of the proposed landing vehicle, generate a unified time series of the vehicle hardware, collect image data reflecting the landing environment of the vehicle in real time through multiple cameras of the vehicle, and perform deblurring processing on the image data to obtain pure visual landing environment image data with enhanced motion fuzzy control.
[0025] S102. Input the pure visual landing environment image data into the end-to-end multi-task perception big model. The end-to-end multi-task perception big model performs Transformer feature encoding, semantic segmentation and depth estimation on the pure visual landing environment image data, and outputs landing environment analysis data including a three-dimensional BEV space map, a semantic segmentation map of the landing area and obstacles and image resolution depth estimation.
[0026] S103. The landing environment analysis data, fluid dynamics prior data, and time series generated during spatiotemporal alignment are input into the spatial RNN model. The spatial RNN model dynamically predicts the landing point of the aircraft based on the landing environment analysis data, the fluid dynamics prior data, and the time series to generate landing phase guidance outputs adapted to different landing altitudes.
[0027] S104. Obtain the landing phase guidance, and perform phased guidance control on the landing of the aircraft based on the obtained landing phase guidance, generate control commands and send them to the aircraft's actuators for execution, so as to guide the aircraft to land automatically.
[0028] In this embodiment, hardware synchronization and spatial calibration are performed on the proposed landing vehicle to achieve spatiotemporal alignment and generate a unified time series. This solves the problem of perception and control link distortion caused by timestamp drift and extrinsic parameter mismatch during high-speed descent between multiple cameras and inertial navigation systems. It achieves strict time synchronization and integrated coordinate expression across sensors and modules, providing a unified time base for subsequent end-to-end inference and closed-loop control. Multiple cameras acquire landing environment images in real time and perform deblurring processing to obtain pure visual image data that enhances motion fuzzy control. This solves the problems of feature degradation and unstable depth estimation caused by large attitude changes, rolling shutter, and high-speed displacement, achieving robust visual input and improved distinguishability of the landing area under high dynamic, low-light, or rain and fog conditions. By feeding pure visual input into an end-to-end multi-task perception model, performing Transformer feature encoding, semantic segmentation, and depth estimation, and outputting 3D BEV, landing area / obstacle segmentation, and resolution depth, the problem of repetitive calculations and time delay accumulation from detection and mapping to planning segmented pipelines can be solved. This achieves integrated four-dimensional (3D + time) environmental representation and efficient perception output within a unified feature space. By inputting landing environment analysis data, fluid dynamics priors, and a unified time series into a spatial RNN model, dynamic prediction of landing points is performed, generating phased guidance adapted to different landing altitudes. This addresses the lack of consistent strategies across altitudes and the difficulty in explicitly handling near-ground disturbances such as ground effects, crosswinds, and wakes. It enables guidance curves that converge hierarchically from high to medium to near-ground and provides forward-looking predictions of landing point evolution over time. Based on the landing phase guidance, phased guidance control of the aircraft is implemented and actuators are issued. This solves the problem of the disconnect between landing point selection and trajectory control, achieving closed-loop coupling from prediction, guidance, and execution, and reducing landing delays caused by near-ground oscillations and excessive conservatism. In this embodiment, the end-to-end multi-task perception model, through multi-task processing of pure visual images, can solve the power consumption and real-time bottlenecks of multi-model cascading under limited airborne computing power, achieving real-time deployment and improved engineering usability in lightweight sensing configurations. Furthermore, by introducing fluid priors and using time-series data to drive the spatial RNN, the problem of the inability to uniformly model the combined effects of dynamic obstacles and aerodynamic disturbances can be solved, achieving joint suppression of occupancy evolution and wind field disturbances, thereby achieving comprehensive improvements in safety, stability, and landing success rate.
[0029] It should be noted that the execution entity of step S104 can be a controller, which acquires the landing phase guidance and performs phased guidance control for the aircraft's landing based on the acquired landing phase guidance. Additionally, the end-to-end multi-task perception large model in step S102 can be a RegNet large model.
[0030] Preferably, the hardware synchronization and spatial calibration process includes: synchronizing the hardware timestamps of the multiple cameras of the proposed landing vehicle and the inertial navigation system; and aligning the coordinate system of the multiple cameras of the proposed landing vehicle and the inertial navigation system to the coordinate system of the vehicle body through an external parameter calibration matrix.
[0031] In this embodiment, synchronizing the hardware timestamps of multiple cameras and the inertial navigation system resolves the timing inconsistency caused by sampling frequencies and clock drift of different devices, achieving strict registration of multi-source data at the same reference time and reducing time jitter in the perception-control link. Aligning the coordinate systems of multiple cameras and the inertial navigation system to the body coordinate system using an extrinsic parameter calibration matrix solves the problem of accumulated projection and attitude conversion errors in cross-coordinate system fusion, achieving consistent mapping of pixel, ray, and body coordinates, and ensuring the accuracy of BEV generation and landing geometric constraints. This embodiment, through an integrated calibration process of synchronization and extrinsic parameter alignment, addresses the issue of high calibration and maintenance costs during long-term operation, enabling reusable and traceable calibration results, and providing a stable geometric and temporal reference for subsequent end-to-end model training and online updates.
[0032] Preferably, the inertial navigation system includes an IMU and a GNSS; the IMU outputs high-frequency attitude data to construct a unified local time series as a reference for aligning multi-camera image data; the timestamps of each image data are mapped to the local time series, and interpolation calculations are performed between the high-frequency attitude data output by adjacent IMUs to obtain the camera attitude parameters at the corresponding time, thereby achieving multi-sensor time alignment in the body coordinate system; the GNSS provides absolute position data to align the body coordinate system to the global coordinate system and outputs a UTC time reference as a globally unified time reference for multiple sensors and the system level. Combined with the extrinsic calibration matrix, the absolute position, and the UTC time reference, the spatiotemporal alignment of the body coordinate system and the global coordinate system is completed, and the relative pose increment of the visual odometry is corrected for drift using GNSS.
[0033] In this embodiment, the IMU outputs high-frequency attitude data to construct a unified local time series, which serves as a reference benchmark for aligning multi-camera image data. The timestamps of each image data are mapped to the local time series, and interpolation calculations are performed between the high-frequency attitude data output by adjacent IMUs to obtain the camera attitude parameters at the corresponding time, thereby achieving multi-sensor time series alignment in the body coordinate system. GNSS provides absolute position data to align the body coordinate system to the global coordinate system and outputs a UTC time reference as a globally unified time reference for multiple sensors and the system. Combining the extrinsic calibration matrix, the absolute position, and the UTC time reference, the spatiotemporal alignment of the body coordinate system and the global coordinate system is completed. GNSS is used to correct the drift of the relative pose increment of the visual odometry. This can solve the problems of inconsistent perception data, cumulative positioning errors, and coordinate system conflicts caused by the misalignment of multiple sensors (vision, IMU, GNSS, etc.) in time and space, achieving a high-precision and highly consistent spatiotemporal reference unification. This provides reliable and synchronous multi-source input for subsequent end-to-end perception and decision-making, effectively suppressing visual odometry drift and improving the robustness and accuracy of the autonomous take-off and landing system in complex environments.
[0034] Preferably, the deblurring process for the image data includes: acquiring IMU data of the aircraft to be landed, and using the IMU data to deblur the image data reflecting the landing environment of the aircraft to obtain pure visual landing environment image data with enhanced motion fuzzy control.
[0035] In this embodiment, acquiring IMU data and using it to deblur image data reflecting the landing environment can solve the problem of feature extraction failure caused by motion blur under high-speed descent and high angular velocity conditions, achieving pixel-level deblurring and detail restoration based on the actual kinematics of the aircraft. Coupled with the IMU and the image under a unified time series, the artifact problem caused by neglecting attitude / acceleration in independent image restoration can be solved, achieving stable visibility of feature points, edges, and textures, improving the quality of semantic segmentation and depth estimation, and providing a foundation for BEV robustness.
[0036] Preferably, when using the IMU data to deblur image data reflecting the aircraft landing environment, the process includes: acquiring the angular velocity and linear acceleration parameters of the IMU under a unified time series, and converting the parameters into camera rotation matrix and motion amplitude estimates; embedding the camera rotation matrix and motion amplitude as prior information for attention weights into the image feature matching and recovery process, so as to assign higher matching confidence to key feature points in motion-blurred regions and reduce the weight in unstable regions, thereby enhancing the image feature matching capability under motion-blurred conditions and generating pure visual landing environment image data with enhanced motion-blurred control.
[0037] In this embodiment, the IMU angular velocity and linear acceleration are converted into camera rotation matrix and motion amplitude estimates, which solves the scale and direction uncertainty caused by inferring motion solely from images, achieving physically consistent motion priors. Embedding these priors as attention weights into the feature matching and recovery process solves the problem of imbalanced feature confidence allocation in blurred regions, achieving enhanced matching in key regions and reduced weighting in unstable regions, thus reducing mismatches and drift. This embodiment outputs pure visual image data that enhances motion blur control, addressing the degradation of subsequent BEV, segmentation, and depth under high dynamic conditions, and providing robust support for landable boundaries, obstacle edges, and normal / slope estimation.
[0038] Preferably, when generating landing phase guidance adapted to different landing altitudes, if the aircraft is at a landing altitude of a preset altitude distance, the spatial RNN model dynamically predicts the landing point of the aircraft based on the landing environment analysis data and the time series and hydrodynamic prior data to generate high-altitude landing phase guidance adapted to the preset altitude distance; when the spatial RNN model performs dynamic prediction, the adaptive weight scheduling method uses the landing environment analysis data as input to schedule the weights of the hydrodynamic prior data and the Occupancy Flow dynamic obstacle avoidance module, where the weight of the Occupancy Flow dynamic obstacle avoidance module is 0.
[0039] For example, the preset high-altitude distance can be more than 100 meters from the landing point. In this embodiment, the spatial RNN model uses landing environment analysis data, time series data, and fluid dynamics prior data as input to dynamically predict the landing point and generate high-altitude landing phase guidance adapted to the preset high-altitude distance. This allows for the formation of time-continuous landing point extrapolation and directional guidance during the high-altitude phase, providing stable priors for the subsequent mid / near-ground phases. Utilizing a unified time series (from spatiotemporal alignment) to drive temporal recursion in dynamic prediction ensures consistency between guidance calculation and the perception time axis, improving the temporal stability of commands during the high-altitude phase. Introducing fluid dynamics prior data into the dynamic prediction of the landing point during the high-altitude phase makes the high-altitude guidance more wind-resistant and the landing point convergence direction more accurate. The adaptive weighted scheduling method uses landing environment analysis data as input, weights the fluid dynamics priors and the Occupancy Flow dynamic obstacle avoidance module, and sets the latter's weight to 0 in the high-altitude phase. This enables phased risk management: in the high-altitude phase, emphasis is placed on large-scale wind fields and global geometry to avoid overfitting and guidance jitter in the early stages of distant obstacles. Setting the Occupancy Flow to zero in the high-altitude phase, emphasizing prior wind fields and global structure, allows for strategy decoupling based on altitude: no fine-grained dynamic obstacle avoidance is performed in the high-altitude phase, reducing false alarms and oscillations, while preserving obstacle avoidance sensitivity in the mid / near-ground phases. In this embodiment, disabling the Occupancy Flow in the high-altitude phase reduces the computational load, allowing for increased loop frequency and thermal margin in the far-altitude phase, reserving computational resources for the near-ground high-frequency closed loop.
[0040] Preferably, when generating the high-altitude landing guidance, the spatial RNN model calls a Transformer-based BEV global planning module to perform long-term path optimization on the 3D BEV spatial map under the unified time series, and performs global path fusion with the satellite map to form a target path and candidate landing points as global priors for the spatial RNN model to output the high-altitude landing guidance; during the generation of the high-altitude landing guidance, the spatial RNN model uses an adaptive weight scheduling method to schedule the weights of the hydrodynamic prior data to low weights within a preset range based on the landing environment analysis data.
[0041] In this embodiment, the Transformer-based BEV global planning module is invoked to perform long-term path optimization under a unified time series. This addresses the issues of insufficient time consistency and local optima in path planning during the transition and near-ground phases, enabling long-term rolling optimization of dynamic scenarios. Global path fusion with satellite maps to form the target path and candidate landing points serves as the global prior for the spatial RNN. This resolves the lack of global ground feature / no-fly zone awareness caused by relying solely on short-term airborne observations, achieving geographical consistency constraints for track and landing point selection.
[0042] In addition, by adjusting the weights of prior fluid dynamics data to a low weight within a preset range, the problem of excessive sensitivity to local aerodynamic disturbances and the tendency to introduce uncertain priors into global planning can be solved. This enables robust global decision-making based primarily on topology and macroscopic accessibility, while aerodynamic details are reinforced in the near-range / transition phase, thereby improving the repeatability and cross-scenario generalization of the global path.
[0043] Preferably, when generating landing phase guidance adapted to different landing altitudes, if the aircraft is at a landing altitude with a preset transition distance, and the preset transition distance is lower than the preset high-altitude distance, the spatial RNN model dynamically predicts the landing point of the aircraft based on the landing environment analysis data, the hydrodynamic prior data, and the time series, to generate transition landing phase guidance adapted to the landing altitude with the preset transition distance; during the dynamic prediction process of the spatial RNN model, the adaptive weight scheduling method assigns high weights to the hydrodynamic prior data within a preset interval, where the weight value of the high weight is greater than the weight value of the low weight, and assigns low weights to the Occupancy Flow dynamic obstacle avoidance module within a preset interval.
[0044] In this embodiment, when the aircraft is at a preset transition distance (below high altitude), landing environment analysis data, fluid dynamics priors, and time series data are jointly input into a spatial RNN to generate transition phase guidance. This solves the model discontinuity problem during the transition from far-field global planning to near-ground fine control, achieving a smooth transition considering ground effects and shear wind influences. For example, the preset transition distance can be a distance between 30 and 100 meters from the landing point. By introducing mechanistic priors at the transition altitude, this embodiment addresses the problem of traditional planning approximating aerodynamic disturbances only with a safety margin, enabling real-time re-estimation and risk-adaptive convergence of the landing point and path in response to wind field and turbulence.
[0045] Furthermore, adaptive weight scheduling assigns high weights to fluid dynamics priors and low weights to occupancy flow, enabling aerodynamic feedforward to stabilize downwash / crosswind-induced yaw and descent trends while suppressing overfitting to distant small targets, ensuring guidance continuity and computational concentration. In this embodiment, when the aircraft is at a landing altitude with a preset transition distance, increasing the weights of fluid dynamics priors and decreasing the weights of occupancy flow allows the maneuver bandwidth to be concentrated on mitigating macroscopic aerodynamic disturbances and the increasing ground effect, establishing a more stable velocity / attitude envelope before entering the near-ground phase. Through weight redistribution and dynamic prediction during the transition phase, a highly correlated strategy continuum is achieved: parameters and attention migrate monotonically with altitude, reducing manual intervention and maintenance costs.
[0046] Preferably, when generating the transition landing phase guidance, the spatial RNN model employs a Convolutional Long Short-Term Memory (ConvLSTM) submodule to perform spatiotemporal prediction of the wind field state under the unified time series. The wind field state includes a three-dimensional vector field of wind speed and direction. Furthermore, using the prior hydrodynamic data as constraints, continuity and momentum conservation constraints are applied to the predicted field to achieve regularization. Simultaneously, based on IMU time series calculations, turbulence intensity feedback indices including acceleration variance and angular velocity spectral density are used to perform online compensation of the predicted wind field, correcting the guidance output of the spatial RNN model to suppress jitter and overshoot, and generating transition landing phase guidance adapted to the landing altitude of the preset transition distance.
[0047] In this embodiment, a ConvLSTM submodule is used to perform spatiotemporal prediction of the three-dimensional vector field of wind speed / direction under a unified time series. This can solve the problem of short-term unpredictability caused by sudden gusts and bypass flow, enabling proactive modeling and short-term extrapolation of the wind field's temporal structure. Applying continuity and momentum conservation constraints to the predicted field using fluid dynamics priors can solve the problems of distortion and divergence in pure data-driven predictions, achieving physical feasibility and numerical stability in wind field estimation. Based on IMU time series calculation of turbulence intensity feedback and online compensation for the predicted wind field, the spatial RNN guided output is corrected to suppress jitter and overshoot, which can solve the trajectory oscillation problem caused by model-reality deviations, achieving steady-state transition of attitude / trajectory and improving the smoothness of the landing path.
[0048] Preferably, when generating landing phase guidance adapted to different landing altitudes, if the aircraft is at a landing altitude with a preset near-ground distance, and the preset near-ground distance is lower than the preset conversion distance, the spatial RNN model performs Occupancy Flow dynamic obstacle avoidance based on the landing environment analysis data, the time series, and the hydrodynamic prior data. Specifically, this includes: fusing a 3D BEV spatial map with depth estimation to construct a 3D voxel occupancy grid; estimating the motion vectors of each voxel in the time series and performing short-term extrapolation to obtain the occupancy flow; assessing collision risk based on predicted occupancy and relative velocity and establishing a safety potential field, causing the radius of the safety potential field around the aircraft to dynamically shrink according to risk probability and braking capacity, thereby constraining the guidance path and landing point updates, and outputting near-ground landing phase guidance adapted to the preset near-ground distance; during the dynamic prediction process of the spatial RNN model, an adaptive weight scheduling method assigns a medium weight to the hydrodynamic prior data within a preset interval. The weight value of the medium weight in the preset interval is less than the weight value of the high weight in the preset interval and greater than the weight value of the low weight in the preset interval, thus affecting the occupancy... The Flow dynamic obstacle avoidance module assigns a large weight to a preset interval, where the weight value of the large weight in the preset interval is greater than the weight value of the small weight in the preset interval.
[0049] In this embodiment, during the preset near-ground distance phase, dynamic obstacle avoidance using Occupancy Flow is executed: 3D BEV and depth estimation are fused into a 3D voxel occupancy grid, which solves the problem of the separation between 2D semantics and 3D geometry under complex near-ground geometry and occlusion, achieving a unified expression of voxel-level visibility and load-bearing capacity. Estimating the motion vectors of each voxel in a time series and extrapolating them in a short time to obtain the occupancy flow solves the problem that static grids cannot reflect the dynamic risks of moving objects such as people, vehicles, and floating objects, enabling joint temporal and spatial prediction of potential collisions. Based on the predicted occupancy and relative velocity, collision risk is assessed and a safety potential field is established. The radius of the safety potential field around the vehicle dynamically shrinks according to the risk probability and braking capacity to constrain the update of the guidance path and landing point. This solves the problem of a fixed safety radius being too conservative or too aggressive, achieving adaptive allocation of obstacle avoidance margin and real-time replanning of the landing point during the near-ground phase, thereby improving the safety and success rate of the final landing.
[0050] In addition, by assigning a large weight to the Occupancy Flow and a medium weight to the hydrodynamic priors during the near-ground phase, we can prioritize meeting the hard safety constraints for collision avoidance while retaining moderate sensitivity to ground effects and eddy current disturbances, thus maintaining controllable convergence of attitude / velocity.
[0051] Preferably, in addition to performing Transformer feature encoding, semantic segmentation, and depth estimation on the purely visual landing environment image data, the end-to-end multi-task perception model also performs attitude estimation on the purely visual landing environment image data. The output landing environment analysis data includes not only a three-dimensional BEV space map, a semantic segmentation map of the landing area and obstacles, and image resolution depth estimation, but also aircraft attitude estimation. The control commands issued to the aircraft's actuators to guide the aircraft to automatic landing include aircraft attitude control commands and speed control commands.
[0052] In this embodiment, the end-to-end multi-task perception model also performs attitude estimation on the pure visual landing environment image data and outputs the attitude as part of the landing environment analysis data. This solves the problems of complex interface coupling, time asynchrony, and information time difference caused by attitude provided by independent VIO / IMU links. It achieves consistent expression of attitude, BEV, semantics, and depth in the same feature space and unified time series, shortening the closed-loop delay from perception to control. The common-source joint output of attitude estimation, 3D BEV, landing area / obstacle segmentation, and depth estimation can solve the problem of lack of consistent coordinates and explicit coupling between body attitude, terrain normal / slope, and obstacle boundary in landing point assessment and path planning. This enables landing area quality judgment with attitude constraints and adaptive guidance of inclined surface / slope, reducing attitude correction overshoot before touchdown. The issued control commands include aircraft attitude control commands and velocity control commands. This addresses the issues of oscillation and delay superposition caused by secondary calculations by the lower-level controller when only position / track references are provided. It enables coordinated design of the outer-loop velocity profile and inner-loop attitude tracking, improving guidance responsiveness and suppressing overshoot and sway during the near-ground phase. This embodiment performs attitude estimation purely visually within the end-to-end model, resolving the single-source vulnerability to attitude / velocity reliability degradation under conditions such as IMU saturation, magnetic interference, and GNSS obstruction. It achieves cross-source verification and redundancy fault tolerance between vision, inertial navigation, and satellite navigation, improving continuous availability in abnormal environments. Directly feeding the attitude estimation into the temporal state of the spatial RNN for staged guidance generation solves the problem of unclosed-loop modeling of the coupling between attitude, velocity, and landing point selection under near-ground wind fields / ground effects. This allows for dynamic adjustment of descent rate and lateral maneuvers based on attitude margin and handling capability, improving convergence stability under disturbance conditions. Overall, this embodiment solves key problems such as timing asynchrony, coordinate inconsistency, rigid velocity management, and single-source vulnerability by moving attitude estimation forward to the end-to-end perception layer and issuing attitude and velocity commands simultaneously on the control side, thus establishing a low-latency closed loop for perception, attitude, guidance, and control. This enables stable convergence and real-time engineering deployment in highly dynamic and strongly disturbed near-ground segments.
[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An autonomous takeoff and landing guidance method based on an end-to-end large model, characterized in that, include: Hardware synchronization and spatial calibration are performed on the simulated landing vehicle to achieve spatiotemporal alignment of the simulated landing vehicle, generate a unified time series of the vehicle hardware, and collect image data reflecting the landing environment of the vehicle in real time through multiple cameras of the vehicle. The image data is then deblurred to obtain pure visual landing environment image data with enhanced motion fuzzy control. The pure visual landing environment image data is input into the end-to-end multi-task perception model. The end-to-end multi-task perception model performs Transformer feature encoding, semantic segmentation and depth estimation on the pure visual landing environment image data, and outputs landing environment analysis data including a three-dimensional BEV space map, a semantic segmentation map of the landing area and obstacles and image resolution depth estimation. The landing environment analysis data, fluid dynamics prior data, and time series generated during spatiotemporal alignment are input into a spatial RNN model. The spatial RNN model dynamically predicts the landing point of the aircraft based on the landing environment analysis data, the fluid dynamics prior data, and the time series to generate landing phase guidance outputs adapted to different landing altitudes. The landing phase guidance is acquired, and the landing of the aircraft is guided and controlled in stages according to the acquired landing phase guidance. Control commands are generated and sent to the aircraft's actuators for execution to guide the aircraft to land automatically.
2. The method as described in claim 1, characterized in that, The hardware synchronization and spatial calibration process includes: synchronizing the hardware timestamps of the multiple cameras and the inertial navigation system of the proposed landing vehicle; and aligning the coordinate system of the multiple cameras and the inertial navigation system of the proposed landing vehicle to the coordinate system of the vehicle body through the external parameter calibration matrix.
3. The method as described in claim 2, characterized in that, The inertial navigation system includes an IMU and a GNSS. The IMU outputs high-frequency attitude data to construct a unified local time series, which serves as a reference for aligning multi-camera image data. The timestamps of each image data are mapped to the local time series, and interpolation calculations are performed between the high-frequency attitude data output by adjacent IMUs to obtain the camera attitude parameters at the corresponding time, thereby achieving multi-sensor time alignment in the body coordinate system. The GNSS provides absolute position data to align the body coordinate system to the global coordinate system and outputs a UTC time reference, which serves as a globally unified time reference for multiple sensors and the system. Combined with the extrinsic parameter calibration matrix, the absolute position, and the UTC time reference, the spatiotemporal alignment of the body coordinate system and the global coordinate system is completed, and the relative pose increment of the visual odometry is corrected for drift using GNSS.
4. The method as described in claim 3, characterized in that, The process of deblurring the image data includes: acquiring IMU data of the aircraft to be landed, and using the IMU data to deblurr the image data reflecting the landing environment of the aircraft to obtain pure visual landing environment image data with enhanced motion fuzzy control.
5. The method as described in claim 4, characterized in that, When using the IMU data to deblur image data reflecting the aircraft landing environment, the process includes: acquiring the angular velocity and linear acceleration parameters of the IMU under a unified time series, and converting the parameters into camera rotation matrix and motion amplitude estimates; embedding the camera rotation matrix and motion amplitude as prior information for attention weights into the image feature matching and recovery process, so as to assign higher matching confidence to key feature points in motion-blurred regions and reduce the weight in unstable regions, thereby enhancing the image feature matching capability under motion-blurred conditions and generating pure visual landing environment image data with enhanced motion-blurred control.
6. The method according to any one of claims 1-5, characterized in that, When generating landing phase guidance adapted to different landing altitudes, if the aircraft is at a landing altitude of a preset altitude, the spatial RNN model dynamically predicts the landing point of the aircraft based on the landing environment analysis data and the time series and hydrodynamic prior data to generate high-altitude landing phase guidance adapted to the preset altitude. When the spatial RNN model performs dynamic prediction, the adaptive weight scheduling method uses the landing environment analysis data as input to schedule the weights of the hydrodynamic prior data and the OccupancyFlow dynamic obstacle avoidance module, where the weight of the OccupancyFlow dynamic obstacle avoidance module is 0.
7. The method as described in claim 6, characterized in that, When generating the high-altitude landing guidance, the spatial RNN model calls the Transformer-based BEV global planning module to perform long-term path optimization on the 3D BEV spatial map under the unified time series, and performs global path fusion with the satellite map to form the target path and candidate landing points as the global prior of the spatial RNN model to output the high-altitude landing guidance. During the generation of the high-altitude landing guidance, the spatial RNN model uses an adaptive weight scheduling method to schedule the weights of the hydrodynamic prior data to low weights in a preset range based on the landing environment analysis data.
8. The method as described in claim 7, characterized in that, When generating landing phase guidance adapted to different landing altitudes, if the aircraft is at a landing altitude with a preset transition distance, and the preset transition distance is lower than the preset high-altitude distance, the spatial RNN model dynamically predicts the aircraft's landing point based on the landing environment analysis data, the hydrodynamic prior data, and the time series to generate transition landing phase guidance adapted to the landing altitude with the preset transition distance. During the dynamic prediction process of the spatial RNN model, the adaptive weight scheduling method assigns high weights to the hydrodynamic prior data within a preset interval, where the weight values of the high weights are greater than the weight values of the low weights, and assigns low weights to the OccupancyFlow dynamic obstacle avoidance module within a preset interval.
9. The method as described in claim 8, characterized in that, When generating the transition landing phase guidance, the spatial RNN model employs a Convolutional Long Short-Term Memory (ConvLSTM) submodule to perform spatiotemporal prediction of the wind field state under the unified time series. The wind field state includes a three-dimensional vector field of wind speed and direction. Using the prior hydrodynamic data as constraints, continuity and momentum conservation constraints are applied to the predicted field to achieve regularization. Simultaneously, based on IMU time series calculations, turbulence intensity feedback indices including acceleration variance and angular velocity spectral density are used to perform online compensation of the predicted wind field, correcting the guidance output of the spatial RNN model to suppress jitter and overshoot, and generating transition landing phase guidance at a landing altitude adapted to the preset transition distance.
10. The method as described in claim 8, characterized in that, When generating landing phase guidance adapted to different landing altitudes, if the aircraft is at a landing altitude with a preset near-ground distance, and the preset near-ground distance is lower than the preset transition distance, the spatial RNN model performs Occupancy Flow dynamic obstacle avoidance based on the landing environment analysis data, the time series, and the hydrodynamic prior data. Specifically, this includes: fusing a 3D BEV spatial map with depth estimation to construct a 3D voxel occupancy grid; estimating the motion vectors of each voxel in the time series and performing short-term extrapolation to obtain the occupancy flow; assessing collision risk based on predicted occupancy and relative velocity and establishing a safety potential field, dynamically shrinking the radius of the safety potential field around the aircraft body according to risk probability and braking capacity to constrain the guidance path and landing point updates, and outputting near-ground landing phase guidance adapted to the preset near-ground distance; during the dynamic prediction process of the spatial RNN model, an adaptive weight scheduling method assigns a medium weight within a preset interval to the hydrodynamic prior data, where the medium weight is less than the high weight and greater than the low weight, thus affecting the Occupancy Flow. The dynamic obstacle avoidance module assigns a large weight to a preset range, where the weight value of the large weight is greater than the weight value of the small weight.
Citation Information
Patent Citations
Unmanned aerial vehicle automatic landing control method based on visual guidance
CN116893682A
Automatic driving planning method, device and equipment of two-wheeled mobile robot and medium
CN120686816A