Quad-rotor unmanned aerial vehicle autonomous landing system based on AprilTag improved algorithm
Through the binocular camera and ROI acceleration strategy combined with binocular vision optimization, the problem of autonomous drone landing affected by GPS signal instability and camera shaking is solved, achieving higher positioning accuracy and stability.
Patent Information
- Application Number
- CN202510577758.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-12
AI Technical Summary
Under certain operating conditions, the unstable GPS signal leads to a decrease in positioning accuracy of the quadrotor UAV, affecting the accuracy of autonomous landing, and camera swaying affects the identification and positioning effect of AprilTag tags.
Binocular camera optimization strategy and region of interest (ROI) are used to accelerate visual detection, combined with binocular vision for redundant checks and optimize positioning, and precise position locking is used to use AprilTag tags.
It improves the recognition reliability and positioning accuracy of AprilTag tags, and enhances the stability and real-time performance of autonomous landing of drones.
Smart Images

Figure CN120469446A_ABST
Abstract
Description
Technical Field
[0001] The invention discloses a quadrotor unmanned aerial vehicle autonomous landing system based on an improved AprilTag algorithm, and belongs to the fields of computer vision, artificial intelligence and automation. Background Art
[0002] Global Positioning System (GPS) technology is widely used in the positioning systems of quadrotor drones. However, under certain operating conditions, factors such as building obstructions, electromagnetic interference, and changing weather conditions can severely affect the stability and accuracy of GPS signals. In such cases, the quadrotor's positioning navigation may fail or its accuracy may drop sharply.
[0003] To address these issues, improved algorithms based on visual recognition systems like AprilTag have become a highly sought-after option. AprilTag accurately identifies a specific QR code tag and doesn't rely on full GPS signal availability. Compared to traditional GPS-dependent systems, AprilTag-based solutions offer higher real-time response speeds and improved anti-interference capabilities, particularly in urban environments or locations with complex lighting conditions. Summary of the Invention
[0004] Since the camera shakes during drone flight, it affects tag recognition and positioning. Therefore, the present invention adopts a binocular camera optimization strategy to improve the reliability of tag recognition and the accuracy of its positioning solution. In addition, the Region of Interest (ROI) is used to accelerate the speed of visual tag detection and improve the efficiency of target positioning in visual navigation. In summary, the proposed method has certain application prospects in autonomous landing of drones.
[0005] In order to achieve the above functions, the present invention adopts the following technical solutions:
[0006] Step 1: Place a TAG36H11 AprilTag on the ground and equip the drone with a binocular camera and an onboard computer. The drone receives a mission instruction, and the flight control system responds and takes off to a preset altitude.
[0007] Step 2: Use a binocular camera to detect the AprilTag using an ROI-based acceleration strategy and mark it as a region of interest.
[0008] Step 3: Process and decode the AprilTag in the region of interest, and calculate the relative 3D coordinates P0 and P1 of the binocular camera and the tag based on the camera's intrinsic parameters.
[0009] Step 4: Calculate the measurement deviation based on the distance between the two cameras for redundancy check and determine whether to output; use the known binocular camera baseline distance to construct an optimization objective function to obtain the optimized three-dimensional coordinates P′0 and P′1, and take the average of the results to obtain the three-dimensional coordinate P of the camera center. c ;
[0010] Step 5: The drone obtains its relative position to the tag in real time, controls the flight control system to place the drone directly above the landing platform, and then begins to land, completing the landing mission. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0012] Figure 1 This is a schematic diagram of the overall solution provided by an embodiment of the present invention;
[0013] Figure 2 is a structural diagram of a label positioning system provided by an embodiment of the present invention;
[0014] Figure 3 This is a schematic diagram of the ROI accelerated detection tag strategy provided by an embodiment of the present invention;
[0015] Figure 4 Schematic diagram of graphic homography transformation provided by an embodiment of the present invention;
[0016] Figure 5 Schematic diagram of AprilTag encoding provided by an embodiment of the present invention;
[0017] Figure 6 This is a working logic diagram of the binocular vision tag positioning system provided by an embodiment of the present invention.
[0018] Figure 7 This is an optimization flowchart based on the gradient descent method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] ROI accelerated detection labeling strategy:
[0020] The AprilTag algorithm uses a full-image search approach, which inevitably limits computer processing speed and makes it difficult to meet the real-time requirements of drone positioning. Therefore, this paper introduces an acceleration strategy based on a region of interest (ROI). By specifying the algorithm's detection area using the ROI, the detection range is narrowed, eliminating the need to perform target recognition on the entire image. This significantly reduces processing time and improves the algorithm's recognition efficiency.
[0021] For any image containing an AprilTag, its four corner points can be detected. These four corner points are represented as (u1, v1), (u2, v2), (u3, v3), and (u4, v4). Once the coordinates of the four corner points of the tag are obtained, the pixel size of the current AprilTag tag can be solved based on the detected coordinates. Specifically, first find the maximum value u of the horizontal and vertical coordinates based on the coordinates of the four corner points. max ,v max and minimum value u min , v min , and then calculate the pixel length and width of the detected AprilTag label based on this. Finally, through a magnification factor f roi Specify the current detection area and determine the size of the ROI area using formula (1):
[0022]
[0023] W in formula (1) roi and H roi Represent the width and height of the selected ROI area respectively. Then, according to formula (2), the coordinates of the upper left corner of the ROI area are solved:
[0024]
[0025] In formula (2), X roi and Y roi They represent the coordinates of the upper left corner of the selected ROI area. It should be noted that the coordinates of the upper left corner need to meet the X roi ≥0 and Y roi ≥0. After the camera acquires a frame containing the AprilTag image, the next frame detection only needs to be completed after identification and positioning. Figure 3 Quickly locate the label position in the red box.
[0026] Label detection and pose calculation:
[0027] In order to achieve visual tag-assisted positioning, it is necessary to decode the tags detected by the camera. When the AprilTag tag is detected, the relative position of the AprilTag is calculated using the homography matrix estimated in the tag recognition step and the camera intrinsic parameter matrix. The specific steps are as follows:
[0028] 1. Image preprocessing: Grayscale the image. Since the AprilTag is designed in black and white, it is not necessary to retain color information when processing the image. Therefore, grayscale images can be used to reduce the algorithm calculation overhead and improve the algorithm's calculation efficiency. The grayscale of the image is calculated using formula (3):
[0029] Gray=0.299R+0.578G+0.114B (3)
[0030] Among them, R, G, and B represent the red, green, and blue channels of the image respectively. By calculating the formula (3), a single-channel grayscale image can be obtained.
[0031] 2. Threshold segmentation: First, the image is divided into 4×4 pixel blocks; then the maximum and minimum grayscale values are calculated in each block; finally, the extreme values in all blocks are filtered with a 3×3 neighborhood maximum and minimum filter, and the filter is applied as the adaptive threshold for image segmentation.
[0032] 3. Contour Search: Contour search is used to find contours that may contain AprilTags. To avoid errors caused by shared edges, the Union-Find algorithm is used to handle dynamic connectivity issues and ensure that each connected domain is assigned a unique ID.
[0033] 4. Fitting quadrilaterals: By fitting continuous edge segmentation points into quadrilaterals and decoding the effective information contained therein, the function of identifying the target is achieved.
[0034] 5. Homography transformation and external parameter estimation: such as Figure 4 The quadrilateral pattern shown here has a homography that may adversely affect the subsequent acquisition of pose information. Therefore, it is necessary to perform a homography on the image to correct the image in shape and angle for subsequent decoding and position and pose determination.
[0035] 6. Decoding: Identify the pixel area inside the quadrilateral to determine whether it conforms to the valid encoding specified by the system. For example, for an AprilTag tag of type TAG36H11, its pixel encoding is as follows: Figure 5 As shown in the figure, the outer circle of pixels has a value of 0, which represents a black edge. Therefore, when a quadrilateral is identified, the quadrilateral is divided into 8×8 pixel patterns and the pixel value of the outermost circle is further checked to see if it is 0. If this condition is met, the inner 6×6 pixel area is then identified to verify whether it meets the valid encoding set by the system, thereby achieving target detection.
[0036] 7. Pose Calculation: When the AprilTag is detected, the relative pose of the AprilTag is calculated using the homography matrix estimated in the tag recognition step and the camera intrinsic parameter matrix. The homography matrix is a transformation matrix that converts from the image coordinate system to the tag coordinate system. It can be decomposed into the product of two matrices, as shown in Equation (4). The two matrices represent the camera's intrinsic parameter matrix P and extrinsic parameter matrix E, which are used to describe the camera's internal parameters and external transformation and rotation relationships, respectively.
[0037]
[0038] Where s is a constant scale factor used to adjust the scale of the mapping. Expanding equation (4) yields the following equation (5):
[0039]
[0040] In formula (5), the elements of the homography matrix H are known to be h ij and the element f of the camera intrinsic parameter matrix P x and f y With this information, the rotation matrix R can be further calculated ij At the same time, the column vectors of the rotation matrix are orthogonal, so R ij Calculate the third column of the external parameter matrix E, that is, t k The value of the translation matrix [t x , t y , t z ] describes the relative position information between the visual tag and the camera. In this way, the relative poses P0 and P1 of the visual tag and the two cameras of the binocular camera can be calculated respectively.
[0041] Positioning optimization based on binocular vision
[0042] Compared with the method of using only a single visual AprilTag algorithm for tag positioning, the proposed system uses binocular vision to optimize positioning, which can achieve higher positioning accuracy. At the same time, the robustness of positioning is improved through multi-source information redundancy verification. Figure 2 The details are as follows:
[0043] Figure 6 The working logic of the entire binocular vision tag positioning system is demonstrated. The system workflow mainly includes binocular vision tag recognition and positioning, redundancy verification and binocular vision optimization. First, the images of the binocular cameras are acquired and the three-dimensional coordinates P0 and P1 of the tag relative to the camera are calculated. Then, the measurement deviation is calculated based on the known distance between the two cameras for redundancy verification, that is, it is compared with the set verification threshold to determine whether to output. Finally, an optimization objective function is constructed using the known binocular camera baseline distance to obtain the optimized three-dimensional coordinates P′0 and P′1, and the average of the results is taken to obtain the three-dimensional coordinate P of the camera center. c .
[0044] In the redundancy check part, since the binocular cameras used in the present invention are fixedly connected, their distance remains unchanged. In addition, the posture states of the camera positioning tags are independent of each other, which enables the present invention to evaluate the accuracy of positioning by comparing the difference between the posture estimation value and the actual value between the binocular cameras. Specifically, after the binocular camera obtains the environmental image containing the AprilTag tag, it obtains the posture information P0 (x0, y0) and P1 (x1, y1) of the two cameras respectively. Then, the error estimation is performed. It is known that the distance between the binocular cameras is l. According to formula (6), the measurement error is:
[0045]
[0046] This value can be used to avoid the situation where positioning is unstable and the positioning result has a large error due to external factors, thereby completing the function of the redundancy check part. In the binocular vision optimization part, the optimization flow chart based on the gradient descent method is as follows Figure 7 As shown. First, define the cost function as follows:
[0047] J=A(||P0-P1|| 2 -l 2 ) 2 (7)
[0048] In formula (7), A is a constant coefficient, which is set to 500 here; P0 is the three-dimensional coordinate of the tag relative to camera 0; P1 is the three-dimensional coordinate of the tag relative to camera 1; and l is the distance between the two cameras. Then, according to formula (8), the partial derivatives of the cost function with respect to the two positioning coordinates are calculated respectively:
[0049]
[0050] In formula (8), P m are the two positioning results of the binocular camera, m∈{0,1}. The number of iterations is then limited to ε≤100 to ensure the real-time positioning. When ε≤100, the convergence condition is determined by formula (9):
[0051] J t -J t-1 |<ζ (9)
[0052] In formula (9), J t and J t-1 are the cost function values at time t and time t-1 respectively; ζ is the convergence threshold, and the present invention sets ζ = 0.001. If it does not converge, iterate according to formula (10):
[0053]
[0054] In formula (10), △ is the iteration step size, and the present invention sets A=0.001. m Continue to calculate the cost function and judge the convergence condition until the conditions are met to obtain the positioning results P′0 and P′1 of the binocular camera. Take the average of the results to obtain the three-dimensional coordinates P of the camera center. c , and sends the position information to the UAV flight controller to complete autonomous landing.
Claims
1. A quadrotor drone autonomous landing system based on the improved AprilTag algorithm is characterized by: Step 1: Place a TAG36H11 AprilTag on the ground and equip the drone with a binocular camera and an onboard computer. The drone receives a mission instruction, and the flight control system responds and takes off to a preset altitude. Step 2: Use a binocular camera to detect the AprilTag using an ROI-based acceleration strategy and mark it as a region of interest. Step 3: Process and decode the AprilTag in the region of interest, and calculate the relative 3D coordinates P0 and P1 of the binocular camera and the tag based on the camera's intrinsic parameters. Step 4: Calculate the measurement deviation based on the distance between the two cameras for redundancy check and determine whether to output; use the known binocular camera baseline distance to construct an optimization objective function to obtain the optimized three-dimensional coordinates P′0 and P′1, and take the average of the results to obtain the three-dimensional coordinate P of the camera center. c ; Step 5: The drone obtains its relative position to the tag in real time, controls the flight control system to place the drone directly above the landing platform, and then begins to land, completing the landing mission.
2. The method according to step 2 of claim 1, characterized in that: Use a binocular camera to detect AprilTag tags using an ROI-based acceleration strategy and mark them as regions of interest, including: For any image containing an AprilTag, its four corner points can be detected. These four corner points are represented as (u1, v1), (u2, v2), (u3, v3), and (u4, v4). After obtaining the coordinates of the four corner points of the tag, the pixel size of the current AprilTag tag is solved according to the detected coordinates. Specifically, first find the maximum value u of the horizontal and vertical coordinates based on the coordinates of the four corner points. max , v max and minimum value u min , v min , and then calculate the pixel length and width of the detected AprilTag label based on this; finally, through a magnification factor f roi Specify the current detection area W roi and H roi , W roi and H roi Represent the width and height of the selected ROI area respectively; then according to u max , v max ,W roi and H roi Calculate the X coordinate of the upper left corner of the ROI area roi and Y roi ; It should be noted that the coordinates of the upper left corner need to meet the X roi ≥0 and Y roi Conditions ≥ 0: After the camera acquires a frame containing the AprilTag image, after identification and positioning, the next frame detection only needs to quickly locate the tag position in the ROI area selected in the previous frame.
3. The method according to step 3 of claim 1, characterized in that: The AprilTag in the region of interest is processed and decoded to calculate the relative 3D coordinates P0 and P1 of the binocular camera and the tag, including: (1) Image preprocessing: Grayscale the image. Since the AprilTag is designed in black and white, it is not necessary to retain color information when processing the image. Grayscale the image reduces the algorithm computational overhead and improves the algorithm's computational efficiency. (2) Threshold segmentation: First, the image is divided into 4×4 pixel blocks; then, the maximum and minimum grayscale values are calculated within each block; finally, the extreme values in all blocks are filtered using the maximum and minimum filtering of a 3×3 neighborhood, and the filter is used as the adaptive threshold for image segmentation; (3) Contour search: Contour search is to find contours that may contain AprilTag labels. To avoid errors caused by shared edges, the Union-Find algorithm is used to handle dynamic connectivity issues and ensure that each connected domain is assigned a unique identification ID. (4) Quadrilateral Fitting: By fitting continuous edge segmentation points into quadrilaterals and decoding the effective information contained in them, the function of identifying the target is realized; (5) Homography transformation and extrinsic parameter estimation: The existence of homography transformation in quadrilateral patterns may have an adverse effect on the subsequent acquisition of pose information; therefore, the present invention performs homography transformation on the image to correct the image in shape and angle for subsequent decoding and position and pose solution; (6) Decoding: Identify the pixel area inside the quadrilateral to determine whether it meets the valid code specified by the system. For example, for an AprilTag tag of type TAG36H11, the pixel value of the outer circle of its pixel code is 0, which represents a black edge. Therefore, when a quadrilateral is identified, the quadrilateral is divided into 8×8 pixel patterns and the pixel value of the outermost circle is further checked to see whether it is 0. If this condition is met, the inner 6×6 pixel area is then identified to verify whether it meets the valid code set by the system, thereby achieving target detection. (7) Pose calculation: When the AprilTag is detected, the relative pose of the AprilTag is calculated with the help of the homography matrix estimated in the tag recognition step and the camera intrinsic parameter matrix. The homography matrix H is a transformation matrix that realizes the conversion from the image coordinate system to the tag coordinate system. It is decomposed into the product of the camera's intrinsic parameter matrix P and the extrinsic parameter matrix E and s. The intrinsic parameter matrix P and the extrinsic parameter matrix E are used to describe the internal parameters of the camera and the external transformation and rotation relationship respectively. s is a constant scale factor used to adjust the scale of the mapping. It is known that the elements of the homography matrix H are h ij and the element f of the camera intrinsic parameter matrix P x and f y ; Through this information, calculate the elements R in the external parameter matrix E i0 and R i1 The value of, at the same time, the column vectors of the matrix E are orthogonal, by R i0 and R i1 Calculate the third column of the external parameter matrix E, namely R 02 , R 12 , R 22 The value of R 02 , R 12 , R 22 Describe the relative three-dimensional position information between the visual tag and the camera; then calculate the relative poses P0 and P1 of the visual tag and the two cameras of the binocular camera respectively.
4. The method according to step 4 of claim 1, characterized in that: The measurement deviation of the distance between the two cameras is redundancy checked, and the optimization objective function is constructed to optimize the three-dimensional coordinates P0 and P1 of the camera and the tag. The average of the optimization results is the three-dimensional coordinate P of the camera center. c ,include: First, according to step 3 of claim 1, the images of the binocular cameras are obtained respectively and the three-dimensional coordinates P0 and P1 of the tag relative to the camera are calculated respectively; then, the measurement deviation is calculated based on the known distance between the two cameras for redundancy verification, that is, it is compared with the set verification threshold to determine whether to output; finally, an optimization objective function is constructed using the known binocular camera baseline distance to obtain the optimized three-dimensional coordinates P′0 and P′1, and the results are averaged to obtain the three-dimensional coordinate P of the camera center c ; The details are as follows: In the redundancy check part, since the binocular cameras used in the present invention are fixedly connected, their spacing remains unchanged. In addition, the posture states of the camera positioning tags are independent of each other, which enables the present invention to evaluate the accuracy of positioning by comparing the difference between the estimated and actual posture values of the binocular cameras. Specifically, after the binocular cameras acquire the environmental image containing the AprilTag tag, the posture information P0 (x0, y0) and P1 (x1, y1) of the two cameras are obtained respectively. Then, the error estimation is performed. Knowing that the spacing between the binocular cameras is l, the measurement error is calculated as w. This value can be used to avoid unstable positioning and large errors in positioning results due to external factors, thereby completing the function of the redundancy check part. In the binocular vision optimization part, the specific steps of optimization based on the gradient descent method are as follows: First define the cost function J, and find P for the cost function J. m Partial derivative, P m are the two positioning results of the binocular camera, m∈{0,1}; then the number of iterations is limited to ε≤100 to ensure the real-time positioning of the system; when ε≤100, the difference between J at time t and time t-1 is calculated to determine whether it has converged; if it does not converge, the positioning result P at this time is used m Subtract J from P m The product of the partial derivative and Δ, where Δ is the iterative step size, is the result of the new P updated after the descent. m point; for the new P after the drop m Continue to calculate the cost function and judge the convergence condition until the conditions are met to obtain the positioning results P′0 and P′1 of the binocular camera. Take the average of the optimization results to obtain the three-dimensional coordinates P of the camera center. c .