Target system and method for following flight and precise landing of mooring unmanned aerial vehicle

By using a multi-level, multi-scale Aruco-marked target structure and intelligent recognition algorithm, combined with traditional and deep learning, the problem of stable flight and high-precision autonomous landing of tethered UAVs on dynamic targets has been solved, achieving high-precision positioning and robust recognition in complex environments.

CN120964057APending Publication Date: 2025-11-18ZHUOYI ZHINENG
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511302829.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing tethered drones face technical bottlenecks in achieving stable tracking and high-precision autonomous landing on dynamic targets, especially in complex environments where positioning errors are large and recognition robustness is poor. Furthermore, existing target designs cannot simultaneously ensure both long-range tracking stability and short-range positioning accuracy.

Method used

A multi-level, multi-scale Aruco marker target structure design is adopted, which combines traditional visual recognition and deep learning algorithms. Through the partitioned layout of multi-size Aruco markers and intelligent recognition mode switching, stable positioning of UAVs at different altitudes and attitudes is achieved. Error state Kalman filter is used to fuse multi-marker pose information.

Benefits of technology

In complex environments, stable and continuous high-precision positioning of UAVs from long-distance tracking to close-range landing has been achieved, which improves the system's environmental robustness and real-time processing capabilities, reduces computing resource requirements, and expands the scope of application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120964057A_ABST
    Figure CN120964057A_ABST
Patent Text Reader

Abstract

The invention introduces a tethered unmanned aerial vehicle follow-up and precise landing target system, which comprises a multi-level visual target, and the visual target sequentially comprises a central positioning area, a target positioning area and a target positioning area from the center to the periphery, and the central positioning area is formed by arranging four aruco marks of a first size in a 2 * 2 array mode; the middle transition area is formed by arranging eight aruco marks of the first size in the upper direction, the lower direction, the left direction and the right direction of the center positioning area; four aruco marks of a second size are respectively arranged at the left upper part, the right upper part, the left lower part and the right lower part of the central positioning area; the peripheral auxiliary area is formed by arranging four aruco marks of a third size at four corners of the whole visual target respectively; the maximum peripheral frame is an aruco mark of a fourth size, and the outer frame of the aruco mark of the fourth size forms the outermost layer area of the whole target. According to the invention, continuous and stable visual features can be provided for the mooring unmanned aerial vehicle at different flight heights and attitudes, and the landing reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of autonomous control and machine vision technology for unmanned aerial vehicles (UAVs), specifically relating to a tethered UAV tracking and precise landing target system and method. This invention is particularly suitable for application scenarios with stringent requirements for landing accuracy. Background Technology

[0002] Tethered drones are unmanned aerial platforms connected to ground facilities via cables. Relying on continuous external power, they can achieve ultra-long-duration operations and have significant advantages in scenarios such as communication relay, wide-area monitoring, and emergency rescue. However, significant technical bottlenecks remain in achieving stable tracking and high-precision autonomous landing on dynamic targets. Currently, common landing guidance methods mainly rely on GPS, inertial navigation systems, and visual navigation methods, but these methods all have certain limitations in dynamic and complex environments.

[0003] Specifically, GPS positioning is susceptible to environmental factors such as building obstruction and electromagnetic interference. In complex urban or nearby indoor areas, its positioning error often reaches the meter level, failing to meet the centimeter-level accuracy requirements for precise landing. While inertial navigation can provide short-term, high-precision pose estimation, its errors accumulate over time, leading to significant drift in long-term tracking tasks, necessitating correction by other sensors. Existing visual navigation solutions often use traditional markers as landing targets, but in real-world dynamic scenarios, they are easily affected by factors such as changes in lighting, partial occlusion, motion blur caused by high-speed movement, and large-angle deflections in the relative pose of the target and the drone, resulting in poor robustness. Furthermore, such visual algorithms typically have high computational complexity, making low-latency real-time processing difficult to achieve onboard embedded platforms for tethered drones.

[0004] Furthermore, existing target designs often lack comprehensive optimization for the entire landing phase. To ensure stable identification by UAVs at altitudes above 30 meters, excessively large targets are often required. However, this causes the target to exceed the camera's field of view during the near-ground phase (below 2 meters), resulting in the loss of positioning information and severely impacting final landing accuracy. Currently, integrated target design solutions that balance long-range tracking stability with short-range positioning accuracy are lacking, and a complete technical system supporting continuous tracking of high-speed moving targets and high-precision landing is yet to be established. Summary of the Invention

[0005] This invention aims to provide a target system and its identification method specifically for dynamic landing scenarios, addressing key technical challenges in achieving stable tracking and high-precision autonomous landing of tethered UAVs on mobile platforms. By optimizing the target structure and identification algorithm, the system effectively improves positioning accuracy, environmental robustness, and real-time performance at different altitudes throughout the entire process from long-distance tracking to near-ground landing, meeting reliable operational requirements in complex real-world application scenarios such as changing lighting, target movement, and partial obstruction.

[0006] This invention introduces a tethered drone tracking and precise landing target system, comprising a multi-level visual target, wherein the visual target includes multiple Aruco markers of varying sizes arranged at different spatial locations; the visual target, from the center to the periphery, includes:

[0007] The central positioning area consists of four first-size Aruco markers arranged in a 2×2 array;

[0008] The intermediate transition area is composed of eight Aruco markers of the first size arranged in the four directions of the central positioning area: above, below, left, and right, with two markers arranged in each direction.

[0009] The corner enhancement area is composed of four second-sized Aruco markers arranged at the upper left, upper right, lower left, and lower right positions of the central positioning area, respectively, where the second size is larger than the first size;

[0010] The peripheral auxiliary area consists of four Aruco markers of a third size, which are respectively placed at the four corners of the entire visual target. The third size is larger than the second size.

[0011] The maximum outer bounding box is an Aruco mark of the fourth size. The outer border of the Aruco mark of the fourth size constitutes the outermost region of the entire target. The fourth size is larger than the third size.

[0012] This invention, through an Aruco marker partition layout with increasing size from the inside out, can provide continuous and stable visual features for tethered UAVs at different flight altitudes and attitudes. It effectively overcomes the technical contradiction that stability and accuracy of single-size targets cannot be achieved at different distances, and significantly improves the environmental adaptability and reliability of UAVs landing on dynamic targets.

[0013] According to the optimized scheme of the above system, the first size is 5cm×5cm, the second size is 15cm×15cm, the third size is 30cm×30cm, and the fourth size is 1.5m×1.5m. By specifically defining the physical size of the Aruco markers in each area, the visibility and recognition performance of the target at different operating distances are optimized. This ensures both long-range acquisition capability and near-field positioning accuracy, providing a standardized target scheme with reasonable size design and efficient space utilization for UAV landing guidance throughout all stages.

[0014] According to the optimized scheme of the above system, the four Aruco markers in the central positioning area are used for the tethered UAV to achieve initial positioning and initialize the error state Kalman filter in a stationary state. This invention clarifies the specific functions of the four markers in the central positioning area, enabling them to provide a reliable initial pose reference for the UAV during the system startup phase and complete the filter initialization, thereby improving the accuracy of the system's initial state estimation and laying the foundation for subsequent dynamic tracking and high-precision control.

[0015] According to the optimized scheme of the above system, the eight Aruco markers in the intermediate transition zone are used to provide continuously identifiable positioning features when the UAV's altitude is below 50cm and it is in a pitch or roll attitude. This emphasizes the functional advantages of the eight markers in the intermediate transition zone during low-altitude attitude disturbances of the UAV. The redundant layout of multiple markers significantly enhances the fault tolerance and robustness of the vision system, ensuring that effective positioning information can still be continuously obtained during the critical stages of takeoff and landing.

[0016] The present invention also discloses a method for identifying the target system described above, comprising the following steps:

[0017] 1) Image acquisition: Images of the visual target are acquired in real time using imaging equipment mounted on the UAV;

[0018] 2) Recognition mode selection: Select a recognition algorithm suitable for Aruco tags of different sizes based on the relative height or relative motion state between the UAV and the visual target;

[0019] 3) Target Recognition and Localization: Based on the recognition algorithm selected in step 2), the Aruco markers in the image are detected and decoded, outputting the image region and pose information of the visual target. This invention intelligently selects the recognition algorithm based on height and motion state, giving full play to the advantages of different computer vision methods, balancing real-time performance and robustness, thereby improving the overall performance of the entire system in complex environments.

[0020] A further optimization of the above recognition method involves using traditional visual recognition methods for the first-sized Aruco markers within the central positioning area and intermediate transition area, including:

[0021] The acquired images undergo preprocessing including grayscale conversion, filtering, and edge enhancement.

[0022] Coarse localization of the target region based on template matching or Hough transform;

[0023] The aruco region is extracted and decoded, and its pose is calculated. By limiting the use of traditional image processing methods for small and medium-sized markers, efficient and low-latency recognition and localization can be achieved on embedded platforms with limited computing resources. This is beneficial for the system to achieve rapid response and real-time control in low-altitude and near-field environments.

[0024] Further optimization of the above recognition method involves employing a deep learning-based detection algorithm for the third and fourth dimensions of the Aruco markers within the maximum bounding box and the peripheral auxiliary region, including:

[0025] The target detection model is trained using target images acquired under various environmental conditions;

[0026] The trained model is deployed on the UAV's onboard computing platform;

[0027] The system performs model inference on real-time images and directly outputs candidate target regions and recognition results. This invention clarifies that using a deep learning-based detection algorithm for large-sized markers can effectively address the problem of decreased recognition rate of traditional methods under long-distance, occluded, and complex lighting conditions, significantly improving the system's recognition success rate and pose estimation accuracy in challenging environments.

[0028] Further optimizations to the aforementioned identification method include a deep learning detection algorithm designed to improve robustness and accuracy when the target is partially occluded, lighting conditions change drastically, or the drone is at a high altitude. This further highlights the advantages of deep learning algorithms in complex scenarios such as occlusion, changing lighting, and high altitudes, demonstrating that this method significantly enhances the system's anti-interference capabilities and operational reliability under extreme conditions, thus expanding the applicable mission range of tethered drones.

[0029] A further optimization of the aforementioned recognition method includes a pose fusion step: fusing the pose information of multiple identified Aruco markers and inputting it into an error state Kalman filter to calculate the precise pose of the UAV relative to the dynamic target. This invention, by introducing multi-marker pose information fusion and error state Kalman filtering, effectively aggregates redundant observation data, suppresses noise and error accumulation, and ultimately outputs a stable and high-precision relative pose, providing a reliable state input for precise UAV tracking and landing control.

[0030] Compared with the prior art, the beneficial effects of the present invention are:

[0031] 1. Firstly, in terms of positioning accuracy, by introducing a multi-layered, multi-scale target structure design and a high-precision visual positioning algorithm, the system can provide stable and continuous position and attitude estimation throughout the entire process of UAV tracking from a long distance to a close-range landing. This design effectively overcomes the shortcomings of traditional targets in terms of insufficient accuracy at different altitudes, especially in the final landing stage, where it can achieve centimeter-level positioning accuracy, significantly improving the accuracy and reliability of landing on dynamic targets.

[0032] 2. Secondly, the system exhibits strong environmental robustness and real-time processing capabilities. The target design has been optimized in terms of shape, color, and spatial structure, and can maintain a high recognition rate under challenging conditions such as complex lighting, partial occlusion, and target movement or posture changes. At the same time, the recognition algorithm has been lightweighted and optimized, with high computational efficiency and low resource consumption. It does not need to rely on a high-performance computing platform and can be directly deployed in the embedded processing system carried by the tethered drone, which helps to reduce the overall system cost and power consumption.

[0033] 3. Finally, the solution has broad applicability and scalability. It can be applied to various mobile platforms, such as ground vehicles, ships at sea, and temporary landing platforms, and can also adapt to various mission requirements, including logistics delivery, emergency rescue, and communication relay scenarios. Its integrated flight and landing guidance capabilities, as well as its stable performance in different environments, make it a key technology for improving the autonomy and mission effectiveness of tethered UAVs, and it has important engineering application and promotion value. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the tethered drone tracking and precise target landing system of the present invention;

[0035] Figure 2 This is a flowchart of the identification method for the tethered drone tracking and precise landing target system in this invention. Detailed Implementation

[0036] This invention introduces a multi-level visual target system for dynamic landing of tethered unmanned aerial vehicles (UAVs). The system adopts a composite structure design with gradually increasing size from the inside out to provide continuous, stable and high-precision visual positioning features for UAVs at different altitudes and flight stages.

[0037] The visual target system specifically includes the following hierarchical structure:

[0038] The central positioning area consists of four Aruco markers, each 5cm x 5cm (first dimension), arranged in a 2x2 square matrix. The core function of this area is to provide a stable and clear positioning reference for the tethered UAV when it is stationary (e.g., before takeoff), thereby completing initial positioning and initializing the Error State Kalman Filter (ESKF), establishing an accurate initial state estimate for subsequent dynamic tracking and landing.

[0039] Intermediate Transition Zone: This zone consists of eight Aruco markers, each 5cm x 5cm (first dimension), arranged precisely in the four directions above, below, left, and right of the central positioning area, with two markers in each direction. This area is specifically designed to handle the flight phase of the UAV from the moment of takeoff to approximately 50cm above the ground. During this phase, the UAV often experiences pitch or roll attitude changes due to takeoff disturbances or control adjustments. This arrangement ensures that even when such attitude changes result in limited or angular deflected vision, the UAV can still capture the complete set of markers within its field of view, thereby continuously calculating the relative attitude and achieving a smooth initial climb and early attitude stabilization.

[0040] Corner Enhancement Area: Four Aruco markers, each 15cm x 15cm (the second size, larger than the first size), are positioned at the upper left, upper right, lower left, and lower right corners of the central positioning area. This area primarily assists in the UAV's positioning in low to medium altitudes above 50cm. Its larger size offers two major advantages: firstly, it can compensate for the UAV's horizontal position shift in the air to some extent, providing a wider field of view; secondly, the larger physical size means occupying more pixels in the image, thus helping to improve the accuracy of relative pose calculations at long distances.

[0041] The outer auxiliary area consists of four 30cm x 30cm Aruco markers (the third size, larger than the second size) placed at the four corners of the target layout. This area is designed to ensure that the UAV can still achieve high-precision positioning and attitude determination results within the critical landing altitude range of 1 to 3 meters. It works in conjunction with the maximum outer frame to ensure that the UAV can smoothly transition the identification and positioning reference from the inner markers to the outer large markers within this altitude range.

[0042] Maximum outer bounding box: Consists of a large Aruco marker measuring 1.5m × 1.5m (the fourth dimension, larger than the third dimension), whose outer border forms the outermost visual boundary of the entire target system. Its main function is to provide robust pose calculation data that is clearly visible from a long distance for UAVs flying at altitudes of 2 meters and above.

[0043] This invention, through multi-scale and multi-level partitioned target design and optimized configuration of space and size, constitutes a complete visual guidance system that can comprehensively cover the entire process of tethered UAVs from takeoff, climb, high-altitude follow-up flight to precise near-ground landing.

[0044] The present invention also introduces a method for identifying the above-mentioned target system, the method comprising the following steps: 1) Image acquisition: image data containing the multi-layer visual target is acquired in real time by an imaging device (such as a visible light camera or an infrared camera) on the airborne drone; the imaging device should have sufficient resolution to ensure that Aruco markers of different sizes in the target structure can be distinguished at different altitudes; the image data is input to the airborne processing unit in the form of a video stream or continuous frames;

[0045] 2) Recognition Mode Selection: Based on the relative height information or dynamic motion state of the UAV and the visual target, an adaptive recognition algorithm suitable for Aruco markers of different sizes is selected; specifically:

[0046] When the drone is at low altitude (below 1 meter) or the attitude angle changes slightly, priority is given to identifying the 5cm×5cm Aruco marker in the central positioning area and the intermediate transition area using traditional image processing methods.

[0047] When the drone is at a high altitude (above 1 meter), or when the target is obviously moving, obstructed, or under complex lighting conditions, a deep learning-based detection model is used to identify the Aruco markers with a maximum outer bounding box (1.5m×1.5m) and 30cm×30cm in the four corners.

[0048] The switching of recognition modes can be automatically triggered based on altitude sensor data or the confidence level in the recognition results of the previous moment;

[0049] 3) Target recognition and localization: Based on the selected recognition algorithm, the Aruco markers in the current image are detected, decoded, and their poses calculated.

[0050] If traditional image processing methods are used, the process includes: preprocessing the input image by converting it to grayscale, filtering and denoising, and enhancing the edges; using Hough transform or template matching to achieve coarse localization of the Aruco region; extracting the region and decoding it, and calculating the relative pose using the PnP algorithm;

[0051] If a deep learning recognition method is chosen, a pre-trained target detection model deployed on an embedded platform is used for forward inference to directly output the bounding box position and recognition number of the Aruco code, and then the pose is calculated based on this. Finally, the region information of the visual target in the image and the pose estimation result of the UAV relative to the target are output. This result will be used for subsequent UAV pose fusion and flight control.

[0052] When using traditional visual recognition methods to identify the first-sized (5cm × 5cm) Aruco mark within the central positioning area and intermediate transition area, the method specifically includes the following steps:

[0053] a) Image preprocessing: First, the acquired original RGB image is converted to grayscale to reduce the amount of computation; then, Gaussian filtering or median filtering is used to smooth the grayscale image to suppress noise interference; finally, edge enhancement algorithms (such as the mature Canny operator or Sobel operator) are used to enhance the contour features in the image and improve the distinguishability of the marked area in subsequent steps.

[0054] b) Coarse localization of target regions: Based on the preprocessed image, one or more of the following methods are used to initially locate possible Aruco-marked regions:

[0055] Template matching: Using pre-stored Aruco-labeled templates, potential target regions are identified by sliding the matching process in the image by calculating similarity metrics (such as normalized cross-correlation coefficients).

[0056] Hough Transform: For obvious straight or rectangular features in an image, Hough line detection is used to fit a quadrilateral contour, and the position and boundary of candidate Aruco regions are initially determined.

[0057] c) Aruco region decoding and pose calculation: For each candidate region obtained by coarse localization, firstly, a perspective transformation is performed to correct image distortion and obtain a front view; then, the binary encoding pattern in the region is extracted and decoded according to the predefined Aruco dictionary to obtain the unique ID of the marker; finally, based on the known physical size of the marker and the relationship between the corresponding points in the image, the precise 3D pose of the camera relative to the Aruco marker is calculated, where the pose includes position and orientation.

[0058] In the above recognition method, a deep learning-based detection algorithm is used to detect the Aruco markers of the third and fourth dimensions in the maximum outer bounding box and the outer auxiliary area. The specific implementation of the deep learning-based detection algorithm includes the following steps:

[0059] a) Training the object detection model:

[0060] A large-scale target image dataset covering various environmental conditions will be collected and constructed to train a deep learning target detection model. The dataset should include target images under different lighting conditions (e.g., strong light, weak light, backlight), different weather conditions (e.g., rain, fog), different viewpoints (e.g., tilted, overhead), different distances, and partial occlusion (e.g., tethered cable occlusion, shadow occlusion). Each image's aruco marker requires accurate bounding boxes and class annotations. A convolutional neural network-based detection architecture (e.g., existing YOLO, SSD, or Faster R-CNN) will be used for training. The loss function will be optimized (e.g., a weighted sum of cross-entropy loss and localization loss) to enable the model to robustly recognize large-sized aruco markers in various complex environments.

[0061] b) Model deployment and optimization:

[0062] The trained deep learning model is deployed on an airborne computing platform for unmanned aerial vehicles (UAVs) (such as existing Jetson series embedded modules). Before deployment, the model needs to be compressed and accelerated, including but not limited to weight quantization, model pruning, and layer fusion, to ensure that it can achieve real-time inference in an embedded environment with limited computing resources. The optimized model is loaded through mature TensorRT, OpenVINO, or corresponding hardware inference engines and integrated into the airborne vision processing program.

[0063] c) Real-time inference and result output:

[0064] During flight, the drone's onboard camera captures images in real time and inputs them into a deployed deep learning model for forward inference. The model directly outputs the bounding box of each candidate region for each large-size Aruco marker in the image, along with the category confidence score and the corresponding marker ID. The system filters out low-quality detection results based on a confidence score threshold, ultimately outputting stable and reliable target region locations and recognition information for subsequent pose calculation modules.

[0065] In this invention, the core feature and application advantage of the deep learning-based detection algorithm lies in its ability to significantly improve the system's recognition robustness and positioning accuracy when the visual target is partially occluded, ambient lighting conditions change drastically, or the drone is flying at a high altitude. The specific technical implementation is described below:

[0066] The deep learning model, trained on massive and diverse datasets, learns the essential feature representations of Aruco tags, rather than relying on traditional, environmentally susceptible handcrafted image features (such as edges and corners). This model possesses powerful feature abstraction and contextual reasoning capabilities, effectively handling the following complex scenarios:

[0067] Partial occlusion: When the target is occluded by the tether of a drone, a temporarily passing obstacle, or its own shadow, traditional algorithms are prone to failure due to missing feature points. Deep learning models, however, can infer the target region and output accurate bounding boxes based on the marked visible parts and the overall structural context information they have learned, thus ensuring the continuity of pose calculation.

[0068] Dramatic changes in lighting conditions: Under complex lighting conditions such as strong light overexposure, weak light low illumination, backlighting, or alternating light and dark, the contrast and color of the markers will change significantly, causing traditional thresholding and edge extraction methods to fail. Deep learning models trained on large amounts of such data are highly invariant to changes in lighting and can directly extract robust features from the original pixels, maintaining a high recognition rate.

[0069] At higher altitudes, targets occupy smaller pixel areas in images, have blurred details, and are easily affected by atmospheric disturbances. Deep learning models can effectively detect small targets, and their large receptive fields help separate weak target signals from the background, achieving high-precision long-distance recognition. This solves the problems of limited recognition distance and decreased accuracy of traditional methods in this scenario.

[0070] This invention utilizes deep learning detection algorithms, through their data-driven learning approach and powerful generalization capabilities, to directly address the core challenges faced by tethered drones in practical applications, becoming a key technological guarantee for achieving reliable identification and accurate landing under harsh conditions.

[0071] The invention also includes a pose fusion step: this step aims to comprehensively utilize the pose information of multiple Aruco markers identified from the image, and perform optimal fusion and filtering through an Error State Kalman Filter (ESKF) to finally output a high-precision and robust pose estimate of the UAV relative to a dynamic target.

[0072] The specific implementation method for this step is as follows:

[0073] a) Extraction of multi-source observation information:

[0074] After the visual recognition module successfully decodes one or more Aruco markers in the image, it outputs independent pose observations for each marker, including relative position (x, y, z) and attitude (such as quaternions or Euler angles). These observations together constitute a multi-source set of observations of the UAV's current state, which may have different confidence levels. b) Error state modeling and ESKF prediction update:

[0075] The system maintains a state vector centered on the current pose of the UAV. ESKF is designed to recursively estimate the error of this state (i.e., the difference between the true state and the predicted state), rather than directly estimating the state itself. This approach effectively handles nonlinear motion and improves numerical stability.

[0076] Prediction Phase (Time Update): Based on accelerometer and gyroscope data provided by the UAV's own inertial measurement unit (IMU), the system's state (such as position, velocity, and attitude) and its errors are predicted. Update Phase (Measurement Update): Multiple visual pose observations obtained in step a) are used as measurement inputs. ESKF establishes an observation model for each valid marker observation and assigns an appropriate observation noise covariance based on the physical size of each marker, recognition confidence, and historical observation quality. Subsequently, the filter performs standard Kalman gain calculation and state update, optimally fusing all visual observation information with the IMU's predicted state to correct accumulated errors and combat transient interference that may exist from a single marker observation.

[0077] c) Output precise relative pose:

[0078] After ESKF fusion and filtering, the output is an optimized six-DOF pose estimate of the UAV relative to a dynamic target. This result significantly outperforms the solution provided by any single marker, and its accuracy and stability benefit from:

[0079] Spatial redundancy: By utilizing information from multiple spatially distributed markers, the limitations of a single viewpoint and potential occlusion issues are overcome.

[0080] Temporal continuity: ESKF integrates historical state information with current observations, effectively smoothing observation noise and suppressing pose jumps caused by image recognition jitter.

[0081] Sensor complementarity: It deeply integrates the high-frequency short-time accuracy of the IMU with the absolute positioning information of visual observation, effectively calibrating the drift of the IMU, and providing continuous, reliable and accurate pose input for the UAV to follow and land on dynamic targets.

[0082] This pose fusion step is the core of this recognition method to achieve high-precision autonomous landing, ensuring that the system can maintain excellent positioning performance even in complex dynamic environments.

[0083] Example 1:

[0084] Application scenario: Tethered drones need to achieve autonomous flight and precise landing on the top platform of a military communication relay vehicle moving at a constant speed, in order to recover the equipment or recharge it.

[0085] Implementation Process: The multi-level visual target described in this invention is securely deployed at the center of the mobile vehicle platform. The target is arranged strictly according to the structure described in claims 1-4: four 5cm×5cm Aruco markers are placed at the center; two markers of the same size are placed in each of the four directions (top, bottom, left, and right); 15cm×15cm and 30cm×30cm markers are placed at the four corners respectively; the outermost layer is a large 1.5m×1.5m Aruco frame. The tethered UAV is equipped with a high-definition global shutter camera and an embedded processing unit (such as the existing NVIDIA Jetson AGX Xavier). The onboard fusion recognition algorithm includes a traditional visual processing flow for small-sized markers and an optimized deployment of the YOLOv5 deep learning detection model for large-sized markers. When the UAV takes off from the vehicle platform, the onboard camera clearly captures multiple 5cm markers in the central area and intermediate transition area, completes the ESKF filter initialization, and stabilizes its departure accordingly. The vehicle continues to move. The drone, tethered by a mooring tether, flew above the vehicle. When its altitude exceeded 2 meters, the onboard system prioritized detecting the outermost 1.5-meter marker and the 30-centimeter marker. Deep learning algorithms effectively overcame the effects of vehicle dust, changing lighting, and tether obstruction, continuously providing stable attitude information. As the drone began its descent and its altitude gradually decreased to below 3 meters, the system successively and stably identified the 30-centimeter and 15-centimeter corner markers. Below 1 meter, even with minor attitude disturbances due to airflow, multiple 5-centimeter markers were still reliably identified. All identification results were input into the ESKF for fusion, ultimately outputting a relative attitude with centimeter-level accuracy. Based on this information, the drone precisely adjusted its attitude and position, landing smoothly at the center of the target on the mobile platform.

[0086] In this embodiment, the UAV achieved continuous and stable tracking of the dynamic target throughout the flight and landing process, and finally successfully landed on the moving vehicle, verifying the high reliability and accuracy of the target system and method in dynamic scenarios.

[0087] Example 2:

[0088] Application scenario: In emergency communication support, tethered drones need to fly alongside and autonomously land on the deck of a rescue boat that rises and falls with the waves for recovery.

[0089] Implementation Process: The multi-layered targets are fixedly positioned at the center of the rescue boat's deck. Due to the severe shaking in its working environment, the targets are made of waterproof and wear-resistant materials, ensuring all markings are firmly attached. To cope with strong sunlight, water mist reflection, and drastic changes in the boat's attitude, the deep learning model in the airborne recognition algorithm has been enhanced with training data under interference conditions such as sea waves and sunlight glare, improving the model's robustness under extreme conditions. The drone approaches the swaying rescue boat from a distance. At high altitude, despite the target's shaking and water mist interference, the large outer markings are still reliably detected by the deep learning model, providing initial guidance for the drone. During descent, the boat's pitch and roll cause drastic changes in the target's viewing angle, and occasional splashes of water create partial obstruction. The system intelligently switches algorithms based on altitude and recognition confidence: relying on the deep learning model at high altitudes or with partial obstruction; and utilizing traditional visual calculations of multiple markings of different sizes at low altitudes with good visibility. The ESKF filter continuously fuses pose observation data from different markers with IMU data, effectively filtering out measurement noise caused by the ship's high-frequency swaying, and outputting a smooth and accurate relative pose estimate. Ultimately, based on the ship's dynamic swaying, the UAV successfully achieved a precise soft landing at the center of the deck.

[0090] This embodiment verifies the effectiveness and excellent robustness of the target system and identification method of the present invention in emergency rescue scenarios with extreme dynamics and high interference, and successfully achieves safe and accurate recovery on a shaking platform.

[0091] In summary, the multi-level visual target system and its recognition method for dynamic tracking and landing of tethered UAVs provided by this invention cleverly cover the entire process from takeoff and high-altitude tracking to precise near-ground landing through an Aruco marker array layout that gradually increases in size from the inside out. This provides continuous and stable visual positioning features for UAVs at different altitudes and attitudes. Combining traditional image processing and deep learning recognition algorithms, the system effectively overcomes challenges such as complex lighting, partial occlusion, target movement, and platform sway, significantly improving the robustness and accuracy of recognition. By fusing multi-marker pose information through an error-state Kalman filter, a high-precision relative pose estimate is finally output, providing a reliable technical guarantee for the safe, autonomous, and precise landing of tethered UAVs on dynamic platforms such as moving vehicles and ships.

[0092] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the technical solutions. Although the applicant has described the present invention in detail with reference to preferred embodiments, those skilled in the art should understand that any modifications or equivalent substitutions made to the technical solutions of the present invention cannot depart from the spirit and scope of the present invention and should be covered within the scope of the claims of the present invention.

Claims

1. A tethered unmanned aerial vehicle (UAV) tracking and precise landing target system, characterized in that, It includes a multi-level visual target, which comprises multiple Aruco markers of varying sizes arranged at different spatial locations; the visual target, from its center to its periphery, includes: The central positioning area consists of four first-size Aruco markers arranged in a 2×2 array; The intermediate transition area is composed of eight Aruco markers of the first size arranged in the four directions of the central positioning area: above, below, left, and right, with two markers arranged in each direction. The corner enhancement area is composed of four second-sized Aruco markers arranged at the upper left, upper right, lower left, and lower right positions of the central positioning area, respectively, where the second size is larger than the first size; The peripheral auxiliary area consists of four Aruco markers of a third size, which are respectively placed at the four corners of the entire visual target. The third size is larger than the second size. The maximum outer bounding box is an Aruco mark of the fourth size. The outer border of the Aruco mark of the fourth size constitutes the outermost region of the entire target. The fourth size is larger than the third size.

2. The tethered UAV tracking and precise landing target system according to claim 1, characterized in that, The first dimension is 5cm×5cm, the second dimension is 15cm×15cm, the third dimension is 30cm×30cm, and the fourth dimension is 1.5m×1.5m.

3. The tethered UAV tracking and precise landing target system according to claim 1, characterized in that, The four Aruco markers in the central positioning area are used for the tethered UAV to achieve initial positioning and initialize the error state Kalman filter when stationary.

4. The tethered UAV tracking and precise landing target system according to claim 1, characterized in that, The eight Aruco markers in the intermediate transition zone are used to provide continuously identifiable positioning features when the UAV is below 50cm above the ground and is in a pitch or roll attitude.

5. A method for identifying a target system as described in any one of claims 1-4, characterized in that, Includes the following steps: 1) Image acquisition: Images of the visual target are acquired in real time using imaging equipment mounted on the UAV; 2) Recognition mode selection: Select a recognition algorithm suitable for Aruco tags of different sizes based on the relative height or relative motion state between the UAV and the visual target; 3) Target recognition and localization: Based on the recognition algorithm selected in step 2), the aruco marker in the image is detected and decoded, and the image region and pose information of the visual target are output.

6. The identification method according to claim 5, characterized in that, Traditional visual recognition methods are used for the first-sized Aruco markings within the central positioning area and intermediate transition area, including: The acquired images undergo preprocessing including grayscale conversion, filtering, and edge enhancement. Coarse localization of the target region based on template matching or Hough transform; Extract the aruco region and perform decoding and pose calculation.

7. The identification method according to claim 5, characterized in that, A deep learning-based detection algorithm is used for the third and fourth dimensions of Aruco markers in the maximum outer bounding box and the outer auxiliary region, including: training the target detection model using target images acquired under multiple environmental conditions; The trained model is deployed on the UAV's onboard computing platform; The model performs inference on real-time images and directly outputs candidate target regions and recognition results.

8. The identification method according to claim 7, characterized in that, The deep learning detection algorithm is used to improve the robustness and accuracy of recognition when the target is partially occluded, lighting conditions change drastically, or the drone is at a high altitude.

9. The identification method according to claim 5, characterized in that, It also includes a pose fusion step: fusing the pose information of multiple identified Aruco markers and inputting it into an error state Kalman filter to calculate the precise pose of the UAV relative to the dynamic target.

Citation Information

Cited By

  • Visual positioning method and system for building robot

    CN121810807A