Unmanned aerial vehicle landing guidance method based on monocular depth estimation network
Through the monocular depth estimation network based on the DepthMaster model and YOLOv8 model recognition combined with the PnP algorithm, the three-stage precise landing of the drone is achieved, solving the problem of precise landing of the drone in complex environments, reducing costs and improving accuracy and reliability.
Patent Information
- Application Number
- CN202510458890.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing drone landing technology is difficult to achieve accurate and reliable landing in complex environments. The traditional methods are costly or have insufficient accuracy, especially the universality and accuracy of monocular vision methods.
A monocular depth estimation network based on the DepthMaster model is used, and a YOLOv8 model is combined with the landing mark recognition and PnP algorithm for pose information solving. It is divided into three stages: global positioning, identification positioning and precise positioning. Infrared LED markings and bandpass filters are used to reduce environmental interference and achieve accurate landing of the drone.
It reduces the cost and time cost of drone landing, improves accuracy and reliability in complex environments, is efficient and universal, and is suitable for accurate landing in various environments.
Smart Images

Figure CN120333445A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV landing, and particularly relates to a UAV landing guidance method based on a monocular depth estimation network. Background Art
[0002] In the field of image processing, deep learning networks have achieved great success in object classification and detection. With the rapid development of deep learning and its combination with various fields, it has shown great potential and commercial value. Deep learning also shows strong analysis and expression capabilities in the field of computer vision, making it possible to estimate depth from a single image. The problem of monocular image depth estimation based on deep learning has also become one of the hot topics concerned by researchers in recent years. In the problem of monocular depth estimation, deep learning networks also have great advantages over traditional image algorithms. For example, Chinese Patent with publication number CN111402310B and Chinese Patent Applications with publication numbers CN111539922A and CN116188555A.
[0003] In recent years, with the continuous progress of UAV technology, its actual application scenarios have become increasingly complex and diverse, which undoubtedly puts more stringent requirements on the accuracy and reliability of UAV landing technology. How to ensure that the UAV can achieve precise landing has become a key technical problem that needs to be solved urgently in the UAV field and has attracted much attention from all walks of life.
[0004] In the field of UAV landing guidance, traditional methods include instrument landing system (ILS), global navigation satellite system (GNSS), and inertial navigation system (INS), etc. Among them, ILS uses the precision landing radar of the landing field to analyze the relative position data between the landing field and the UAV, and the positioning accuracy depends on the landing radar, which cannot meet the landing requirements of UAVs in complex and changeable environments. GNSS is widely used in UAV landing guidance due to its stable signal and mature technology, but it has disadvantages such as low signal update rate and being easily interfered. INS realizes navigation through inertial sensors, but cumulative errors will occur during the working process and is usually combined with GNSS to improve the reliability and accuracy of UAV landing guidance.
[0005] Accurate position estimation is the key for the UAV to complete precise landing. The method based on binocular vision can accurately estimate the spatial distance, but the high-precision binocular depth camera is costly, and the measurement range is proportional to the baseline (the distance between two cameras), and the baseline cannot be increased infinitely, which limits its ability to measure long distances. Summary of the Invention
[0006] In view of the above problems, the present invention provides a method for guiding the landing of an unmanned aerial vehicle (UAV) based on a monocular depth estimation network. The present invention divides the UAV landing into three stages. The first stage is the global positioning stage, which roughly locates the UAV above the landing mark using GPS / IMU. The second stage is the recognition and positioning stage, which performs target detection on the landing mark and uses a monocular depth estimation network based on the DepthMaster model for depth estimation to make the UAV approach the landing mark. The third stage is the fine positioning stage. By using the PnP algorithm, the three-dimensional information of the known landing mark is matched with the two-dimensional information of the landing mark obtained from image processing to calculate the pose information. Combining with the depth estimation information, the precise landing of the UAV is achieved, improving the safety and reliability of the system.
[0007] The present invention provides a method for guiding the landing of an unmanned aerial vehicle based on a monocular depth estimation network, characterized by comprising:
[0008] Step S1: Let i = 1. When i = 1, it represents the first UAV.
[0009] Step S2: Perform global positioning on the i-th UAV to obtain the i-th UAV after global positioning.
[0010] Step S3: Based on the i-th UAV after global positioning, identify the corresponding original landing area image S i ;
[0011] Step S4: Let n = 1. When n = 1, it represents the first landing mark.
[0012] Step S5: Input the original landing area image S i into the YOLOv8 model A i , identify the n-th landing mark Y i,n , and adjust the position of the i-th UAV after global positioning so that the landing mark Y i,n is at the center position of the landing area image to obtain the i-th UAV with updated position Z i,n ;
[0013] Based on the i-th UAV with updated position Z i,n obtain the landing area image S′ i,n ;
[0014] Step S6: Input the landing area image S′ i,n into the monocular depth estimation network B i,n , and obtain the depth estimation feature of the i-th UAV with updated position Z i,n ;
[0015] Adjust the i-th UAV with updated position Z i,nDistance from the landing area, and complete the position update of the i-th drone Z i,n Identify and locate the landing area, and obtain the drone Z' with the second position update of the i-th one i,n ;
[0016] Step S7: Based on the drone Z' with the second position update of the i-th one i,n Obtain the updated landing area image S″ i,n ; Based on the updated landing area image S″ i,n Obtain the pose information T of the drone with the second position update of the i-th one i,n ;
[0017] Step S8: Utilize the pose information T of the drone with the second position update of the i-th one i,n And the depth estimation feature of the drone Z with the position update of the i-th one i,n To adjust the position of the drone Z with the position update of the i-th one i,n And complete the positioning of the i-th drone Z i,n For the n-th landing marker;
[0018] Step S9: Judge whether n is greater than or equal to N, where N represents the total number of landing markers. If so, complete the positioning of each landing marker by the i-th drone. If not, let n = n + 1 and return to step S5;
[0019] Step S10: Judge whether i is greater than or equal to I, where I represents the total number of drones in the training set. If so, complete the positioning of each landing marker by each drone and obtain the final monocular depth estimation network; If not, let i = i + 1 and return to step S2;
[0020] Step S11: Based on the final monocular depth estimation network, conduct landing guidance for the drone.
[0021] Optionally, the landing marker is an infrared LED.
[0022] Optionally, the specific steps for obtaining the depth estimation feature of the drone Z with the position update of the i-th one in step S6 include: i,n :
[0023] Input the landing area image S′ i,n Into the monocular depth estimation network B i,n , perform latent space encoding on the landing area image S′ i,n , and then input it into the U-Net model D i,n , and extract the corresponding local image feature F inet,i,n ;
[0024] Extract the external image feature F of the landing area image S′ i,n ; ext,i,n ;
[0025] Based on the feature alignment loss, the local features F of the image unet,i,n are feature-aligned with the external features F of the image ext,i,n to obtain the aligned features of the landing area image S′ i,n ;
[0026] The aligned features of the landing area image S′ i,n are input into the Fourier transform module F i,n to obtain the frequency domain features of the landing area image S′ i,n ;
[0027] The frequency domain features of the landing area image S′ i,n are input into the modulator G i,n to obtain the balanced frequency domain features of the landing area image S′ i,n ;
[0028] The balanced frequency domain features of the landing area image S′ i,n are inverse-transformed back to the spatial domain using the inverse fast Fourier transform to obtain the optimized landing area image F f,i,n ;
[0029] The spatial features F of the landing area image are obtained s,i,n ;
[0030] The optimized landing area image F f,i,n is cascaded with the spatial features F of the landing area image s,i,n for feature complementarity to obtain the complementary features of the landing area image S′ i,n ;
[0031] The complementary features of the landing area image S′ i,n are decoded to obtain the depth estimation features of the i-th position-updated drone Z i,n ;
[0032] Optionally, the expression of the feature alignment loss is:
[0033]
[0034] where L fa is the feature alignment loss, dist(.) represents the distance between the two feature distributions, is the feature after projection of the local features of the n-th landing mark image of the i-th position-updated drone, and F ext,i,n is the external feature of the n-th landing mark image of the i-th position-updated drone.
[0035] Optionally, the specific steps of obtaining the pose information of the drone for the second position update in step S7 include:
[0036] Step S71: Obtain the updated landing area image S″ i,n the pixel coordinates and three-dimensional coordinates of each landing marker in;
[0037] Obtain the updated landing area image S″ i,n the conversion relationship between the pixel coordinates and three-dimensional coordinates of each landing marker in;
[0038] Step S72: Construct the least-squares problem related to the conversion relationship between the pixel coordinates and three-dimensional coordinates of each landing marker in the updated landing area image S″ i ;
[0039] Step S73: Minimize the least-squares problem related to the conversion relationship between the pixel coordinates and three-dimensional coordinates of each landing marker in the updated landing area image S″, and obtain the pose information T of the i-th drone with a second position update i,n ; i,n .
[0040] Optionally, for the least-squares problem related to the conversion relationship between the pixel coordinates and three-dimensional coordinates of each landing marker in the updated landing area image S″ i ;
[0041] The expression is:
[0042]
[0043] where E is the least-squares error value, represents obtaining the corresponding T value according to the least-squares error value of the conversion relationship between the pixel coordinates and three-dimensional coordinates of the landing marker, s i,n is the scale factor of the n-th landing marker of the i-th drone, i = 1, 2, 3,..., I, P i,n is the projected pixel coordinate of the n-th landing marker of the i-th drone, K is the internal parameter matrix of the camera, and T is the Lie group representation of the camera pose R and the translation vector t.
[0044] Optionally, the pose information T of the i-th drone with a second position update in step S7 i,n includes the rotation matrix of the i-th drone with a second position update and the translation vector of the i-th drone with a second position update.
[0045] Compared with the prior art, the present invention has at least the following beneficial effects:
[0046] (1) The present invention uses the DepthMaster model to establish a monocular depth estimation network for depth estimation. The DepthMaster model is based on the pre-trained Stable Diffusion v2 model and is trained on the large-scale LAION-5B dataset, containing rich image prior knowledge. Therefore, DepthMaster does not need to collect and train data on a large scale, which can effectively reduce the time cost and labor cost;
[0047] (2) The monocular depth estimation network used in the present invention adopts a two-stage training strategy. In the first stage, the feature alignment module is used to reduce texture overfitting and learn the global structure of the scene. In the second stage, the Fourier enhancement module is used to balance the frequency domain features and improve the ability to retain details, combining high efficiency and accuracy, and showing excellent generalization performance and depth detail capture ability in monocular depth estimation;
[0048] (3) The present invention uses the YOLOv8 model to identify and track the landing mark, uses DepthMaster to estimate the depth information, and uses the PnP algorithm to measure the position information of the landing mark, without the assistance of a laser rangefinder, lidar, etc., which can effectively reduce the cost;
[0049] (4) The present invention solves the problems of low accuracy and poor universality of the autonomous landing of drones when only relying on monocular vision, can meet the actual needs of the precise landing of drones, and has high practical value;
[0050] (5) The present invention uses an infrared LED as the landing mark, adopts a checkerboard arrangement, and installs a band-pass filter in front of the camera, which can effectively reduce the influence of ambient light and increase the robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings are only for the purpose of showing specific embodiments and are not considered to be a limitation of the present invention.
[0052] Figure 1 It is a schematic diagram of the overall process of the drone for landing guidance in the embodiment of the present invention;
[0053] Figure 2 It is a schematic diagram of the landing trajectory of the drone for landing guidance in the embodiment of the present invention;
[0054] Figure 3 It is a schematic diagram of the overall framework of the DepthMaster model in the embodiment of the present invention;
[0055] Figure 4 It is a schematic diagram of obtaining the pose information of the drone using the PnP algorithm in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0056] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other. In addition, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0057] A specific embodiment of the present invention, as Figures 1-4 , discloses a method for guiding a drone to land based on a monocular depth estimation network, and the specific implementation steps are as follows:
[0058] Step S1: Let i = 1. When i = 1, it represents the first drone;
[0059] Step S2: Perform global positioning on the i-th drone, and control the i-th drone to be above the landing mark in the landing area; obtain the i-th drone after global positioning;
[0060] Optionally, the specific steps of the control in step S2 for the i-th drone to be above the landing mark in the landing area include:
[0061] Obtain the global position information of the i-th drone based on satellite map GPS; the global position information includes longitude, latitude and altitude;
[0062] Obtain the angular velocity and acceleration of the i-th drone based on the inertial measurement unit IMU;
[0063] Based on the global position information of the i-th drone, the angular velocity of the i-th drone and the acceleration of the i-th drone, perform long-distance navigation and positioning of the i-th drone, control the i-th drone to be above the landing mark in the landing area, and achieve global positioning of the i-th drone;
[0064] Optionally, the global positioning is the position where the drone is 60 m above the landing mark;
[0065] Optionally, choose the layout of a checkerboard to design the landing mark;
[0066] The landing mark is an infrared LED with a specific wavelength of 940 nm, which can effectively penetrate obstacles such as smoke and dust and is suitable for monitoring in complex environments;
[0067] Optionally, the number of landing marks in the landing area is not less than 3.
[0068] Step S3: Based on the monocular camera carried by the i-th drone after global positioning, identify the corresponding original landing area image S i;
[0069] Step S4: Let n = 1. When n = 1, it represents the first landing marker;
[0070] Step S5: Input the original landing area image S i into the YOLOv8 model A i , identify the nth landing marker Y i,n , and adjust the position of the ith drone after global positioning so that the landing marker Y i,n is at the center position of the landing area image, obtaining the ith position-updated drone Z i,n ;
[0071] Based on the ith position-updated drone Z i,n obtain the landing area image S' i,n ;
[0072] Optionally, the installation method of the monocular camera is the installation method of an optoelectronic pod, which increases the field of view of the monocular camera and improves the stability of the monocular camera system;
[0073] Install a band-pass filter with a wavelength similar to that of the landing marker in front of the lens of the monocular camera to reduce the influence of ambient light and make it easier for the monocular camera to obtain information about the landing marker.
[0074] Step S6: Input the landing area image S' i,n into the monocular depth estimation network B i,n , and obtain the depth estimation features of the ith position-updated drone Z i,n ;
[0075] Adjust the distance between the ith position-updated drone Z i,n and the landing area, and complete the identification and positioning of the ith position-updated drone Z i,n with the landing area, obtaining the ith position-secondarily updated drone Z' i,n ;
[0076] Optionally, adjust the distance between the ith position-updated drone Z i,n and the landing area, and complete the identification and positioning of the ith position-updated drone Z i,n with the landing area within the range of 5 - 10 m between the drone and the landing marker;
[0077] Optionally, the specific steps of step S6 for inputting the landing area image S' i,n into the monocular depth estimation network and obtaining the depth estimation features of the ith position-updated drone Z i,n include:
[0078] Input the landing area image S' i,nInput monocular depth estimation network B i,n , and utilize the I2L encoder C i,n , to encode the landing area image S′ i,n into the latent space, obtaining the latent state of the landing area image S′ i,n ;
[0079] Input the latent state of the landing area image S′ i,n into the U-Net model D i,n , and extract the corresponding local image features F unet,i,n based on the intermediate layer;
[0080] Input the landing area image S′ i,n into the external encoder DINOv2 E i,n , to obtain the corresponding external image features F ext,i,n ;
[0081] Use a multi-layer perceptron to project the local image features F unet,i,n into the feature space of the external image features F ext,i,n , and based on the feature alignment loss, align the local image features F unet,i,n with the external image features F ext,i,n to obtain the aligned features of the landing area image S′ i,n ;
[0082] Input the aligned features of the landing area image S′ i,n into the Fourier transform module F i,n , and perform a fast Fourier transform to the frequency domain to obtain the frequency domain features of the landing area image S′ i,n ;
[0083] Input the frequency domain features of the landing area image S′ i,n into the modulator G i,n , balance the information of different frequency bands, and obtain the balanced frequency domain features of the landing area image S′ i,n ;
[0084] Use the inverse fast Fourier transform to transform the balanced frequency domain features of the landing area image S′ i,n back to the spatial domain to obtain the optimized landing area image F f,i,n ;
[0085] Input the latent state F i,n of the landing area image S′ mid,i,n into the spatial channel H i,n , and obtain the spatial features F s,i,n of the landing area image through two convolutional transformations;
[0086] Input the optimized landing area image F f,i,nPerform a cascading operation with the spatial feature F of the landing area image s,i,n to perform feature complementarity and obtain the landing area image S′ i,n The complemented features;
[0087] Use the I2L decoder J i,n to decode the landing area image S′ i,n and its complemented features to obtain the depth estimation features of the landing area image S′ i,n and the depth estimation features of the i-th position-updating drone Z i,n and the depth estimation features of the i-th position-updating drone Z
[0088] Optionally, the monocular depth estimation network described in step S6 is the DepthMaster model;
[0089] Furthermore, the monocular depth estimation network includes an I2L encoder, a U-Net model, a feature alignment module, a Fourier enhancement module, and an I2L decoder;
[0090] The Fourier enhancement module includes a spatial channel and a frequency channel;
[0091] Optionally, the expression for the transformation relationship of projecting the local image feature F unet,i,n onto the feature space of the external image feature F ext,i,n is:
[0092]
[0093] where, is the feature after projection of the n-th local feature of the landing mark image of the i-th position-updating drone, and h φ represents the projection process.
[0094] Optionally, the expression for the feature alignment loss is:
[0095]
[0096] where, L fa is the feature alignment loss, and dist(·) represents the distance between the two feature distributions.
[0097] In the present invention, the optimized landing area image F f,i,n performs a cascading operation with the spatial feature F of the landing area image s,i,n so that the DepthMaster model adaptively balances the low-frequency structural features and high-frequency detail features within a single forward channel, effectively improving the visual quality of depth prediction.
[0098] In the present invention, the monocular depth estimation network DepthMaster model is a single-step diffusion model, which reduces texture overfitting through a feature alignment module and balances frequency domain features through a Fourier enhancement module. Combining a two-stage training strategy, it realizes the depth estimation of each pixel point in the image, as Figure 3 shown.
[0099] Step S7: Based on the drone Z' with the second update at the i-th position i,n obtain the corresponding updated landing area image S″ i,n ; Use the corresponding updated landing area image S″ i,n to obtain the pose information T of the drone with the second update at the i-th position by using the PnP algorithm i,n ;
[0100] Optionally, the specific steps of obtaining the pose information of the drone with the second update at the i-th position in step S7 include:
[0101] Step S71: Obtain the pixel coordinates and three-dimensional coordinates of each landing marker in the updated landing area image S″ i,n ;
[0102] Obtain the conversion relationship between the pixel coordinates and three-dimensional coordinates of each landing marker in the updated landing area image S″ i,n , and the expression is:
[0103] s i,n p i,n = KT i,n P i,n
[0104] where s i,n is the scale factor of the n-th landing marker of the i-th drone, i = 1, 2, 3,..., I, P i,n is the projected pixel coordinate of the n-th landing marker of the i-th drone, P i,n = [u i,n , v i,n T , u i,n is the pixel abscissa of the n-th landing marker of the i-th drone, v i,n is the pixel ordinate of the n-th landing marker of the i-th drone, K is the internal parameter matrix of the camera, T is the Lie group representation of the camera pose R and the translation vector t, that is, the conversion relationship between the pixel coordinates and three-dimensional coordinates of the landing marker, P i,n is the three-dimensional coordinate of the n-th landing marker of the i-th drone, P i,n = [X i,n , Y i,n , Z i,n T , X i,nis the X-axis coordinate of the nth landing marker of the ith drone, Y i,n is the Y-axis coordinate of the nth landing marker of the ith drone, Z i,n is the Z-axis coordinate of the nth landing marker of the ith drone.
[0105] Step S72: Construct an updated landing area image S″ i The least squares problem related to the conversion relationship between the pixel coordinates and three-dimensional coordinates of each landing marker in
[0106]
[0107] where E is the least squares error value, represents obtaining the corresponding T value according to the least squares error value of the conversion relationship between the pixel coordinates and three-dimensional coordinates of the landing marker.
[0108] Step S73: Minimize the least squares problem related to the conversion relationship between the pixel coordinates and three-dimensional coordinates of each landing marker in the updated landing area image S″ i,n to obtain the pose information T of the ith drone with a second position update i,n ;
[0109] Optionally, the pose information T of the ith drone with a second position update described in step S7 i,n includes the rotation matrix R of the ith drone with a second position update i,n and the translation vector t of the ith drone with a second position update i,n ;
[0110] It can be understood that the rotation matrix R i,n and the translation vector t i,n represent the attitude and position of the ith drone with a second position update relative to the landing marker; among them, the rotation matrix R i,n provides the attitude information of the ith drone with a second position update, including pitch angle, yaw angle, and roll angle; the translation vector t i,n provides the position information (Tx, Ty, Tz) of the ith drone with a second position update relative to the landing marker. As Figure 4 shown;
[0111] Step S8: Based on the pose information T of the ith drone with a second position update i,n and the depth estimation feature of the ith drone with a position update Z i,n perform position adjustment on the ith drone with a position update Z i,n to complete the precise positioning of the nth landing marker by the ith drone with a position update Z i,n i.e., complete the landing guidance of the ith drone for the nth landing marker;
[0112] Step S9: Determine whether n is greater than or equal to N, where N represents the total number of landing markers. If so, complete the landing guidance of the i-th drone for each landing marker. If not, set n = n + 1 and return to Step S4;
[0113] Step S10: Determine whether i is greater than or equal to I, where I represents the total number of drones in the training set. If so, complete the precise positioning of the i-th drone for each landing marker to obtain the final monocular depth estimation network. If not, set i = i + 1 and return to Step S2;
[0114] Step S11: Conduct landing guidance for the drone based on the final monocular depth estimation network.
[0115] As described above, only the preferred specific embodiments of the present invention are provided, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.
Claims
1. A drone landing guidance method based on a monocular depth estimation network, characterized in that Including: Step S1: Let i = 1. When i = 1, it represents the first drone; Step S2: Perform global positioning on the i-th drone to obtain the i-th drone after global positioning; Step S3: Identify the corresponding original landing area image S of the i-th drone after global positioning i ; Step S4: Let n = 1. When n = 1, it represents the first landing marker; Step S5: Input the original landing area image S i into the YOLOv8 model A i to identify the nth landing marker Y i,n , and adjust the position of the ith drone after global positioning to obtain the ith position-updated drone Z i,n ; Update the drone Z based on the position of the i-th aircraft i,n Obtain the landing area image S′ i,n ; Step S6: Input the landing area image S′ i,n into the monocular depth estimation network B i,n to obtain the depth estimation feature of the i-th position-updated drone Z i,n . Adjust the position of the i-th drone to update drone Z i,n Obtain the distance to the landing area and get the second updated drone Z' of the i-th position i,n ; Step S7: Drone Z' with the second update based on the position of the i-th drone i,n Obtain the pose information T of the drone with the second update based on the position of the i-th drone i,n ; Step S8: Use the pose information T of the drone with the second updated position of the i-th frame i,n and the depth estimation feature of the drone Z with the updated position of the i-th frame i,n to adjust the position of the drone Z with the updated position of the i-th frame i,n and complete the positioning of the n-th landing marker by the drone Z with the updated position of the i-th frame i,n ; Step S9: Determine whether n is greater than or equal to N, where N represents the total number of landing markers. If so, complete the positioning of each landing marker by the i-th drone. If not, let n = n + 1 and return to Step S5; Step S10: Determine whether i is greater than or equal to I, where I represents the total number of drones in the training set. If so, complete the positioning of each landing marker by each drone to obtain the final monocular depth estimation network. If not, let i = i + 1 and return to Step S2; Step S11: Perform landing guidance on the drone based on the final monocular depth estimation network.
2. The method for guiding a drone to land based on a monocular depth estimation network according to claim 1, wherein: The landing marker is an infrared LED.
3. The method for guiding a drone to land based on a monocular depth estimation network according to claim 1, wherein: The specific steps of obtaining the depth estimation feature of the i-th position-updated drone Z described in step S6 i,n include: Input the landing area image S′ i,n into the monocular depth estimation network B i,n , perform latent space encoding on the landing area image S′ i,n , and then input it into the U-Net model D i,n to extract the corresponding local image feature F unet,i,n ; Extract the image S′ of the landing area i,n of the external features F of the image ext,i,n ; Based on the feature alignment loss, align the local feature F of the image unet,i,n with the external feature F of the image ext,i,n to perform feature alignment and obtain the aligned feature of the landing area image S′ i,n ; Input the alignment features of the landing area image S′ i,n into the Fourier transform module F i,n to obtain the frequency domain features of the landing area image S′ i,n ; Input the frequency domain features of the landing area image S′ i,n into the modulator G i,n to obtain the frequency domain features of the landing area image S′ i,n after balancing; Use the inverse fast Fourier transform to transform the landing area image S′ i,n The balanced frequency-domain features are transformed back to the spatial domain to obtain the optimized landing area image F f,i,n ; Obtain the spatial feature F of the landing area image s,i,n ; Optimize the landing area image F f,i,n With the spatial features of the landing area image F s,i,n Perform a cascading operation to complement features and obtain the landing area image S′ i,n Complementary features; Decode the complementary features of the landing area image S′ i,n to obtain the depth estimation feature of the i-th position-updated drone Z i,n 4. The method for guiding a drone to land based on a monocular depth estimation network according to claim 1, wherein: The expression of the feature alignment loss is: Among them, L fa is the feature alignment loss, and dist(.) represents the distance between two feature distributions. is the feature after projection of the local feature of the nth landing mark image of the ith position-updated UAV, and F ext,i,n is the external feature of the image of the nth landing mark of the ith position-updated UAV.
5. The method for guiding a drone to land based on a monocular depth estimation network according to claim 1, wherein: The specific steps of obtaining the pose information of the i-th drone with the second position update in Step S7 include: Step S71: Drone Z' with the second updated position based on the position of the i-th aircraft i,n Obtain the updated landing area image S″ i,n ; Obtain the updated landing area image S″ i,n and the pixel coordinates and three-dimensional coordinates of each landing marker in it; Obtain an updated image S″ of the landing area i,n The conversion relationship between the pixel coordinates and three-dimensional coordinates of each landing mark in Step S72, construct the updated landing area image S″ i a least squares problem related to the conversion relationship between the pixel coordinates and the three-dimensional coordinates of each landing marker in Step S73, minimize and update the landing area image S″ i,n For the least squares problem related to the conversion relationship between the pixel coordinates and the three-dimensional coordinates of each landing marker in i,n , obtain the pose information T of the i-th drone with the second position update i,n .
6. The method for guiding the landing of a drone based on a monocular depth estimation network according to claim 1, wherein The updated landing area image S″ i Regarding the least squares problem related to the conversion relationship between the pixel coordinates and the three-dimensional coordinates of each landing marker, the expression is: where E is the least squares error value, denotes obtaining the corresponding T value according to the conversion relationship between the pixel coordinates and the three-dimensional coordinates of the landing marker, and s i,n is the scale factor of the nth landing marker of the ith UAV, i = 1, 2, 3,..., I, and P i,n is the projected image pixel coordinate of the nth landing marker of the ith UAV, K is the internal parameter matrix of the camera, and T is the Lie group representation of the camera pose R and the translation vector t.
7. The method for guiding a drone to land based on a monocular depth estimation network according to claim 1, wherein: The pose information T of the drone with the second updated position in step S7 i,n includes the rotation matrix of the drone with the second updated position and the translation vector of the drone with the second updated position.
Citation Information
Patent Citations
A Monocular Image Depth Estimation Method and System Based on Depth Estimation Network
CN111402310B
Monocular depth estimation and surface normal vector estimation method based on multi-task network
CN111539922A
Monocular indoor depth estimation algorithm based on deep network and motion information
CN116188555A
Unmanned aerial vehicle autonomous landing guiding method based on vision
CN115793677A
Vision and radar information fused unmanned aerial vehicle cluster target positioning method
CN118584470A