Homography-assisted drone landing method based on computer vision and deep learning
By employing a homography-assisted method based on computer vision and deep learning, the sensor dependence and complexity issues in UAV autonomous landing technology were addressed, enabling high-precision autonomous landing of UAVs in complex environments, reducing hardware costs and improving safety and real-time performance.
Patent Information
- Application Number
- CN202511277068.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing autonomous landing technologies for drones rely on multiple sensors, leading to high costs, system complexity, and data fusion challenges. Furthermore, their accuracy is low in complex environments. Traditional methods are susceptible to changes in lighting conditions and are highly dependent on hardware, making real-time operation difficult.
A homography-assisted method based on computer vision and deep learning is adopted. Visible light images of special patterns are acquired by UAV cameras, and the homography matrix is calculated using a deep learning model. Combined with physical constraints and projection verification, the UAV can achieve precise autonomous landing on the docking platform.
It enables high-precision autonomous landing of UAVs in complex environments, reduces hardware costs, simplifies system complexity, improves safety, real-time performance, adaptability and robustness, and reduces accident risks.
Smart Images

Figure CN120780007B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of unmanned aerial vehicle control, and particularly relates to a homography assisted unmanned aerial vehicle landing method based on computer vision and deep learning. BACKGROUND
[0002] When the unmanned aerial vehicle flies outdoors, it is affected by light changes, motion blurring, partial occlusion and other disturbances, which affects the accuracy of homography matrix estimation; traditional feature-based homography estimation methods cannot extract enough feature points in weak texture or repetitive texture areas, resulting in matching failure; and for deep learning-based homography estimation methods, most of them estimate homography matrix for homologous (such as visible light) images, and have large errors in complex environments such as heavy fog, night, and low light.
[0003] Traditional SIFT feature extraction and matching methods are time-consuming (>100ms), which is difficult to meet the real-time control requirements; and deep learning-based homography estimation methods also rely on a large amount of learning parameters, and also have a high inference time, which is difficult to run in real time on an embedded unmanned aerial vehicle platform.
[0004] Existing unmanned aerial vehicle autonomous landing technology mainly relies on multi-sensor fusion and environment perception algorithms, and the core goal is to achieve accurate positioning, obstacle avoidance and stable landing. The mainstream technical solutions mainly include the following:
[0005] Multi-sensor fusion scheme: combine visual, IMU, GPS, barometer and other multi-sensor data, and improve robustness through Kalman filtering or AI algorithms.
[0006] Supersonic / infrared ranging assisted method: measure the relative height of the unmanned aerial vehicle and the ground through supersonic or infrared sensors to assist in slow landing.
[0007] GPS / RTK-based positioning guidance method: use GPS or RTK (real-time dynamic positioning) to provide position positioning, and combine with preset coordinates for landing.
[0008] Vision recognition-based landing method: recognize the landing platform markers (such as QR codes, ArUco markers, special patterns, etc.) or natural features (texture, edge) through the camera, and combine with SLAM (simultaneous localization and mapping) algorithm to locate the landing point.
[0009] For the above points, the respective deficiencies are as follows: first, for the multi-sensor fusion scheme, the scheme relies on high hardware cost; and the system complexity is high, and the calibration and maintenance are difficult; for the ultrasonic / infrared ranging auxiliary method, it is easy to be affected by rain, dust and other environment, and the ranging range is usually limited; for the positioning guidance method based on GPS / RTK, the signal is easy to be blocked by buildings or electromagnetic interference, and cannot be used indoors; at the same time, GPS cannot achieve complete accurate positioning, and has a certain deviation; and RTK needs to preset a base station; finally, for the landing method based on visual recognition, it is sensitive to light changes; at the same time, the method still needs the data provided by the hardware device (such as IMU) to assist the calculation, in addition, the IMU has zero bias error, and the cumulative error is large after long time flight, thereby causing inaccurate pose estimation. SUMMARY
[0010] The present application aims to provide a homography assisted unmanned aerial vehicle landing method based on computer vision and deep learning, which solves the problems of high cost, system complexity and data fusion caused by relying on multiple sensors (such as IMU, GPS, laser radar, etc.) in the existing unmanned aerial vehicle automatic landing technology.
[0011] To achieve the above purpose, the present application provides a homography assisted unmanned aerial vehicle landing method based on computer vision and deep learning, comprising the following steps:
[0012] At the center position of the horizontal landing platform containing a special pattern, a visible light image containing a special pattern is collected by the vertical downward camera of the unmanned aerial vehicle, and stored as a fixed preset picture image;
[0013] When the unmanned aerial vehicle enters the range of the landing platform, the landing platform related image is photographed by the camera, and the homography matrix is solved for the preset picture image and the real-time picture image by using the visual model based on deep learning;
[0014] According to the calculated homography matrix, the three-dimensional coordinate offset matrix of the current position of the unmanned aerial vehicle relative to the position of the unmanned aerial vehicle under the preset picture image is derived in the case of taking the platform center as the origin, so as to calculate the three-dimensional coordinates of the current position of the unmanned aerial vehicle relative to the origin;
[0015] According to the obtained world coordinates of the unmanned aerial vehicle relative to the platform, the unmanned aerial vehicle pose and flight path are adjusted in real time by a control algorithm, so that it autonomously lands towards the landing point in the preset picture.
[0016] In the process of "when the unmanned aerial vehicle enters the range of the landing platform, the landing platform related image is photographed by the camera, and the homography matrix is solved for the preset picture image and the real-time picture image by using the visual model based on deep learning", the visual model based on deep learning comprises:
[0017] The features of the preset picture image and the real-time picture image are extracted by two extractors that do not share parameters, wherein each of the extractors comprises five ordinary convolution modules and one SimAM parameter-free attention convolution module;
[0018] The feature maps of the two types of images are aggregated, and are converted into graph structure features by an Embedding layer;
[0019] The graph structure features are input into an improved ViG model, and a coordinate offset vector is output, and a homography matrix is obtained by a DLT algorithm.
[0020] In the process of inputting the graph structure features into the improved ViG model, outputting the coordinate offset vector, and obtaining the homography matrix by the DLT algorithm, the basic block processing of the ViG model comprises:
[0021] The input features are processed by two 1x1 convolutions and one 3x3 convolution;
[0022] The output of the MLP block is added to the input features by residual addition;
[0023] The features after the residual addition are processed by a max-relative graph convolution block and an FFN block;
[0024] The output features of the basic block are obtained by residual addition and feedforward processing again.
[0025] In the process of deriving, from the calculated homography matrix, a three-dimensional coordinate offset matrix of the current position of the unmanned aerial vehicle relative to the position of the unmanned aerial vehicle in the preset picture image with the center of the platform as the origin, and calculating the three-dimensional coordinates of the current position of the unmanned aerial vehicle relative to the origin, the specific steps of calculating the three-dimensional coordinates of the current position of the unmanned aerial vehicle relative to the origin comprise:
[0026] The influence of the camera intrinsic parameters is eliminated to obtain a normalized homography matrix;
[0027] The normalized matrix is singular value decomposed to obtain a rotation matrix and a coordinate offset vector;
[0028] The current three-dimensional coordinates of the unmanned aerial vehicle are derived in combination with the camera optical center coordinates corresponding to the preset picture;
[0029] The unique correct solution is screened by physical constraints and projection verification.
[0030] In the process of screening the unique correct solution by the physical constraints and the projection verification, the method specifically comprises:
[0031] The pitch angle and the roll angle of the unmanned aerial vehicle are limited to be within the range of ±30°, and the vertical component of the coordinate offset vector is greater than 0;
[0032] Four marked points on the platform are selected, and the matching homographic transformation result is verified through coordinate conversion and projection relationship.
[0033] In the "adjusting the attitude and flight path of the UAV in real time through a control algorithm according to the obtained world coordinates of the UAV relative to the platform, so that the UAV autonomously lands towards the landing point in the preset picture", the autonomous landing specifically comprises:
[0034] The attitude of the UAV is corrected according to the rotation matrix, so that the UAV body is kept horizontally parallel to the platform;
[0035] The moving direction and position are calculated based on the coordinate offset vector, and the flight path of the UAV is adjusted;
[0036] The image is repeatedly collected every 0.5s, the homographic matrix is calculated, and the attitude is adjusted until landing.
[0037] In the "acquiring a visible light image containing a special pattern by a vertical downward camera of the UAV at the center position of the horizontal landing platform containing the special pattern, and storing the image as a fixed preset picture image", the method specifically comprises:
[0038] The UAV is accurately landed at the center of the horizontal landing platform containing the special pattern, so that the UAV body is kept horizontally parallel to the platform;
[0039] The camera parameters of the vertical downward camera installed at the bottom of the UAV are adjusted to ensure that the captured picture clearly contains the complete special pattern and the platform edge;
[0040] The captured visible light image is marked as a positioning reference image and stored in the UAV storage unit as a reference standard for subsequent homographic matrix solving.
[0041] The homography assisted unmanned aerial vehicle landing method based on computer vision and deep learning of the application is that, on a horizontal platform for unmanned aerial vehicle landing containing a special pattern (such as an H letter on a helicopter landing platform), the unmanned aerial vehicle is pre-landed at the center of the platform, a camera arranged below the unmanned aerial vehicle and perpendicular to the plane of the unmanned aerial vehicle captures the landing point image (containing the special pattern) of the platform, at this time, the image captured by the camera is used as a reference for the positioning of the unmanned aerial vehicle landing, and is stored as a fixed preset image; when the unmanned aerial vehicle needs to land, the camera captures the current image (which can be an infrared image) captured by the unmanned aerial vehicle in real time; when the unmanned aerial vehicle enters a certain range of the landing platform, the camera can capture the relevant image of the landing platform, and a visual model based on deep learning is used to solve the homography matrix of the preset image and the real-time image; according to the calculated homography matrix, the three-dimensional coordinate offset matrix of the current position of the unmanned aerial vehicle relative to the position of the unmanned aerial vehicle in the preset image is derived (multi-solution elimination is required) with the center of the platform as the origin, so that the three-dimensional coordinates of the current position of the unmanned aerial vehicle relative to the origin are calculated; according to the obtained world coordinates of the unmanned aerial vehicle relative to the platform, the posture and flight path of the unmanned aerial vehicle are adjusted in real time through a corresponding control algorithm, so that the unmanned aerial vehicle accurately lands towards the landing point in the preset image; the above process completely depends on the image information collected by the camera and the calculation results of the deep learning model, and does not need to rely on other sensors such as IMU and laser radar. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced.
[0043] Fig. 1 is a flowchart of the homography assisted unmanned aerial vehicle landing method based on computer vision and deep learning of the application.
[0044] Fig. 2 is a homography transformation relationship diagram, and a perspective transformation relationship between two planes.
[0045] Fig. 3 is a visual model based on deep learning of the application. DETAILED DESCRIPTION
[0046] The embodiments of the application will be described in detail below, examples of which are shown in the drawings, and the embodiments described below by referring to the drawings are exemplary and are intended to explain the application, and cannot be understood as a limitation of the application.
[0047] Please refer to Fig. 1 to Fig. 3 , wherein Fig. 1is a flowchart of a homography-assisted UAV landing method based on computer vision and deep learning, Fig. 2 is a schematic diagram of a homographic transformation relationship, and a perspective transformation relationship of two planes, Fig. 3 is a deep learning-based visual model.
[0048] The present application provides a homography-assisted UAV landing method based on computer vision and deep learning, comprising the following steps:
[0049] S100: At the center of the horizontal landing platform containing a special pattern, a visible light image containing a special pattern is captured by a vertical downward camera of the UAV, and stored as a fixed preset picture image.
[0050] S101: The UAV is precisely landed at the center of the horizontal landing platform containing a special pattern, ensuring that the UAV body is horizontally parallel to the platform.
[0051] S102: Adjust the camera parameters installed vertically downward at the bottom of the UAV to ensure that the captured image clearly contains the complete special pattern and the platform edge.
[0052] S103: The captured visible light image is labeled as a positioning reference image and stored in the UAV storage unit as a reference standard for subsequent homography matrix solving.
[0053] In this embodiment, on a horizontal platform for UAV landing containing a special pattern (such as an H letter on a helicopter landing platform), the UAV is pre-landed at the center of the platform. The camera installed vertically downward below the UAV and parallel to the UAV plane captures the landing point image of the platform (which needs to contain the special pattern). At this time, the camera-captured image is used as a reference for UAV landing positioning, and is stored as a fixed preset picture image.
[0054] S200: When the UAV enters the landing platform range, the camera captures the landing platform-related image, and a deep learning-based visual model is used to solve the homography matrix of the preset picture image and the real-time picture image.
[0055] In this embodiment, homography estimation is a basic technique in computer vision, which is usually defined as the projection mapping relationship between images of the same planar object taken by two cameras without lens distortion from different positions, aiming to determine a projection transformation between two images and express it with a homography matrix Through this transformation, the geometric differences between images can be effectively understood and compensated, achieving better image alignment. As Fig. 2 shown. Wherein, and Two different position coordinates of the same point in two different planes mapped by two different perspectives. and Homogeneous coordinates of two coordinates. R represents an orthogonal matrix used to describe the rotation relationship of two camera perspectives . t represents the coordinate offset vector of two camera positions in the world coordinate system . H represents the homography matrix indicating the projection transformation relationship between two planes.
[0056] In the autonomous landing of a UAV, the homography matrix can be used to calculate the pose (position and attitude) of the UAV relative to the landing platform. The existing technology mainly includes the following two methods:
[0057] Traditional feature matching method: based on manual feature (such as SIFT, SURF, ORB) to extract image key points, and then filter matching points and calculate homography matrix through RANSAC algorithm.
[0058] End-to-end method based on deep learning: the method of deep learning converts the traditional feature matching-based method into a unified, trainable deep learning framework model, so as to estimate the homography matrix. This method greatly improves the efficiency and accuracy of feature extraction. This method can also be divided into supervised and unsupervised algorithms.
[0059] Most of the existing homography estimation techniques based on infrared and visible light images use relatively complex networks to compensate for the influence of modal differences, such as Res-Net, GAN, Swin Transformer, etc. These network models have a very large number of parameters and require a high response time. Therefore, the present application uses an improved GNN-based infrared and visible light image model architecture (a deep learning-based visual model) for the intelligent landing of a UAV, as shown in Fig. 3 .
[0060] S201: Extract the features of the preset picture image and the real-time picture image through two extractors that do not share parameters, wherein each of the extractors includes 5 ordinary convolution modules and 1 SimAM parameter-free attention convolution module.
[0061] S2011: The preset picture image and the real-time picture image pass through the first 3 ordinary convolution modules respectively, to preliminarily extract the basic features such as image texture and edge.
[0062] S2012: The feature maps processed by the first 3 ordinary convolution modules are respectively input into the 1 SimAM parameter-free attention convolution module of the corresponding extractor, to strengthen the key area features through the feature map and SimAM attention operation.
[0063] S2013: The number of channels of the optimized feature map is restored to 1 through the last two ordinary convolution modules, and the preset picture image feature and the real-time picture image feature are obtained.
[0064] S202: The features of the two types of images are aggregated, and the graph structure features are converted through an Embedding layer.
[0065] S203: The graph structure features are input into the improved ViG model, and the coordinate offset vector is output, and then the homography matrix is obtained through the DLT algorithm.
[0066] S2031: The input features are processed through two 1x1 convolution and one 3x3 convolution.
[0067] S2032: The MLP block output and the input features are added by residual.
[0068] S2033: The features after residual addition are processed through the maximum relative graph convolution block and the FFN block.
[0069] S2034: The basic block output features are obtained through residual addition and feedforward processing again.
[0070] In the embodiment, the deep learning-based visual model mainly includes the following steps:
[0071] 1. First, for two images of size , the input images and , the application uses two lightweight shallow feature extractors that do not share weights to extract shallow features. Among them, H and W represent the height and width of the images and . For the previous infrared and visible image homography estimation method, attention mechanisms such as SENet, CBAM, TripletAttention are mostly used for joint modeling in channel and spatial dimensions to improve performance. However, these attention mechanisms mostly have high parameter quantities, or have more branch structures. In view of this, the application selects to replace the above attention module with SimAM, a simple, parameter-free convolutional neural network attention module, which can effectively reduce the parameter quantity of the model and speed up the model inference speed. The feature extractor of the application contains 5 ordinary convolution modules and one parameter-free attention convolution module. The parameter-free attention convolution module can be written as the following formula:
[0072]
[0073] Among them, represents the feature map in the feature extractor, The first layer output. represents a convolutional layer containing a SimAM module. represents a parameter-free attention module. represents element-wise multiplication. In the feature extractor, three convolutional layers are used to extract shallow features, then a convolutional layer with a SimA attention module is used, and finally two convolutional layers are used to restore the channel number to 1. The feature extractor can be written as follows:
[0074]
[0075] 2、After obtaining two feature maps and with a size of after passing through the feature extractor, the present application aggregates them to obtain , which is used as the input of the homography estimator to calculate the homography matrix. The present application uses an improved ViG model as the basic framework of the homography estimator. This architecture is inspired by the ViG model and the MobileViG model, and some ViG blocks are deleted to replace simpler and faster convolutional blocks. Specifically, the homography estimator of the present application contains basic blocks, each of which is composed of two MLP blocks and a maximum relative graph convolution block, as shown in the main block in Fig. 1 . Among them, the MLP block contains two convolutions and one convolution; the maximum relative graph convolution block contains a maximum relative graph convolution layer and a feedforward neural network FFN block. Such a design balances the size of the parameter quantity of the model and the performance of the model, and can be written as follows:
[0076]
[0077]
[0078] wherein represents the output of the xth main block, represents the input of the xth main block. represents an MLP block, represents a maximum relative graph convolution block, is a feedforward neural network composed of two fully connected layers. The addition in the formula is regarded as residual processing.
[0079] Specifically, the present application first converts the feature map to a graph structure feature Subsequently, it is input into N backbone basic blocks, and finally passed through two fully connected layers to obtain an 8-dimensional coordinate offset vector. Finally, this invention uses the DLT algorithm to solve for the image. and homography matrix image .
[0080] S300: Based on the calculated homography matrix, the three-dimensional coordinate offset matrix of the UAV's current position relative to the UAV's position in the preset image is derived with the platform center as the origin, thereby calculating the three-dimensional coordinates of the UAV's current position relative to the origin.
[0081] S301: Eliminate the influence of camera intrinsic parameters to obtain the normalized homography matrix.
[0082] S302: Perform singular value decomposition on the normalized matrix to solve for the rotation matrix and coordinate offset vector.
[0083] S303: Based on the camera optical center coordinates corresponding to the preset image, deduce the current three-dimensional coordinates of the drone.
[0084] S304: Select the unique correct solution through physical constraints and projection verification.
[0085] S3041: Limit the pitch and roll angles of the UAV to within ±30°, and ensure that the vertical component of the coordinate offset vector is greater than 0;
[0086] S3041: Select 4 marker points on the platform and verify the homography transformation results through coordinate transformation and projection relationship verification.
[0087] In this embodiment, during the intelligent landing of the UAV, after obtaining the homography matrix between two images captured by the camera through a homography estimation model, a specific algorithm is used for matrix decomposition and outlier elimination. The specific calculation method is as follows:
[0088] First, we need to clarify the existing conditions. We have now assumed that the center point of the drone's docking plane is in three-dimensional world coordinates. The origin The normal vector of the plane Based on a fixed preset image, the coordinates of the optical center of the camera mounted on the drone at position A under the docking platform can be obtained. ,in This refers to the height of the optical center of the drone's camera relative to the plane of the docking platform, and the camera is configured to be mounted vertically downwards. The invention establishes that there exists a first camera coordinate at position A. Location coordinates for The origin. Based on the drone's parameters (such as length, width, and height), the range occupied by the drone in the three-dimensional coordinate system can be obtained. However, for now, the drone's position coordinates are simply considered as the coordinate point of the optical center of the camera. The purpose of this invention is to allow the real-time position B to be determined by the coordinate point of the camera's optical center. Move to In terms of location. Simultaneously, this invention sets the optical center coordinates of the drone camera at location B. For the second camera coordinate system The origin. The homography matrix has now been obtained. It needs to be solved. , and First, the homography matrix is normalized to eliminate the influence of camera intrinsic parameters:
[0089]
[0090] Where K is the intrinsic parameter matrix of the camera, and then, for the normalized homography matrix, the present invention has the following linear transformation formula:
[0091]
[0092] in, It is a Translation vector , represents the coordinate offset vector between two position coordinates. This indicates the height of the camera's optical center relative to the plane of the platform when the drone is docked. Represent a An orthogonal matrix is used to describe the rotation relationship. Then, according to... , and To establish a linear transformation relationship between them, we need to first perform SVD singular value decomposition on R:
[0093]
[0094] in, and It is an orthogonal matrix. It is a diagonal singular value matrix containing three singular values. Based on the decomposition results, a rotation matrix can be constructed. and coordinate offset The solution expression is:
[0095]
[0096] in, Represents the determinant of a matrix; From the singular value matrix represents the third singular value; is the third column of the matrix . Further, since the rotation matrix is the product of three primitive rotation matrices, it can be written as follows:
[0097]
[0098] where, , and represent the yaw angle, the pitch angle, and the roll angle, respectively. It should be noted that the specific and obtained by decomposing H are and have a relationship similar to the following formula:
[0099]
[0100] Therefore, essentially, represents the rotation of the coordinate system with the camera optical center as the origin, which can be regarded as the deflection of the unmanned aerial vehicle during flight. And represents the position offset vector of the origin of the second camera coordinate system relative to the origin of the first camera coordinate system. Therefore, the position of the second camera coordinate point can be obtained as follows:
[0101]
[0102] At this point, given , , , the real-time position coordinates of the unmanned aerial vehicle are obtained; however, the decomposition of the homography matrix may have multiple solutions; the previous technical solutions have no clear algorithmic solution to the multiple solution problem, and the screening relies on hardware such as IMU or laser radar to determine the solution by height determination. Or use the motion model to predict the trajectory to screen the solution. This to some extent brings higher computational cost and hardware burden. To this end, the present application uses certain physical constraints to regulate the solution, such as is always true. Or according to the provisions of some industry standard organizations and the provisions of some manufacturers' unmanned aerial vehicle products (such as ISO 21384-3, PX4 user manual, DJI user manual, etc.), the pitch angle and the roll angle are generally limited to , which can be changed according to the specific use. In the present application, the pitch angle and the roll angle the range of the .
[0103] In addition, a homography assisted projection verification method can be used to screen out the unique correct solution. Specifically, four marker points can be found on the UAV landing plane, denoted as . Note that the coordinate system of the four coordinate points is the world coordinate system with the center of the plane as the origin. Then, in the fixed preset image, the pixel coordinate system of the four points can be obtained . With the help of the pixel internal parameter matrix , the coordinate point in the first camera coordinate system can be obtained:
[0104]
[0105] Subsequently, according to the coordinate conversion relationship, the point in the first camera coordinate system is converted to the point in the second camera coordinate system, and the internal parameter matrix of the camera is used again to convert it to the pixel coordinate point in the real-time image:
[0106]
[0107] Finally, according to the homographic transformation relationship between , the multiple solution problem of R and t can be verified and screened.
[0108] S400: According to the obtained world coordinates of the UAV relative to the platform, the UAV attitude and flight path are adjusted in real time through a control algorithm, so that the UAV autonomously lands towards the landing point in the preset image.
[0109] S401: The UAV attitude is corrected according to the rotation matrix, so that the UAV body is horizontally parallel to the platform.
[0110] S402: The moving direction and position are calculated based on the coordinate offset vector, and the UAV flight path is adjusted.
[0111] S403: Repeat the image acquisition, homographic matrix calculation and attitude adjustment every 0.5s until landing.
[0112] In this embodiment, after the rotation matrix and the offset vector are calculated, the rotation offset of the UAV is first corrected to make it as horizontally parallel to the landing platform as possible. The correction process does not calculate the time, and the corrected angle deviation is set to Afterwards, the UAV calculates the direction and position to be moved according to the offset vector, and the moving process is counted in a time interval, generally set as 0.5s, when the calculation time and the moving time reach 0.5s, the UAV stops moving, and the homography matrix is calculated again according to the real-time position picture image obtained by re-shooting and the fixed preset picture image, and the above operation is repeated.
[0113] The training step of the deep learning-based visual model specifically comprises:
[0114] S105: Collecting preset picture images and real-time images when the UAV cruises around the platform, and recording real poses;
[0115] S106: Labeling and preprocessing the collected images, and constructing a training data set;
[0116] S107: Optimizing the model by using quantization-aware training and channel pruning to reduce the parameter quantity and improve the inference speed.
[0117] In the embodiment, data collection is first performed, and the data collection is mainly divided into static data collection and dynamic data collection; the static data collection mainly collects and preprocesses images of the predetermined landing position of the UAV, so as to obtain fixed preset picture images of the UAV at the landing position ; then the UAV platform above and around a certain area are cruised, and are photographed in a short interval time, so as to obtain picture images photographed at real-time positions ; the real three-dimensional coordinate positions and the yaw angle, the pitch angle and the roll angle of the UAV during the photographing process are recorded; then the collected data are labeled and preprocessed, and are used as a training data set to train and optimize the model; the model training process adopts the quantization-aware training mode to further reduce the parameter quantity of the model and speed up the inference speed of the model. At the same time, the channel pruning and other operations are adopted to adapt the UAV platform.
[0118] Finally, it should be noted that the technical solution of the present application is only for the process of intelligent autonomous landing of the UAV; in the process of cruising or flying to the specified location according to the predetermined flight route, other technologies such as GPS are still needed; only the GPS cannot accurately locate the landing platform position of the UAV, and it is more likely to locate the area near the landing platform of the UAV; and the present application aims to solve the problem of intelligent autonomous landing of the UAV in the area near the landing platform.
[0119] The application adopts pure vision and deep learning technology, realizes the accurate autonomous landing of the unmanned aerial vehicle on the preset landing platform through real-time calculation and matrix decomposition of homography matrix. This technical scheme not only meets the requirements of high precision and real-time of the physical and chemical indicators in the technical aspect, but also brings significant beneficial effects to the society, mainly in the following aspects:
[0120] 1. Safety and reliability improvement: The landing accuracy of the unmanned aerial vehicle is improved (the landing error can be reduced to centimeter level), which greatly reduces the risk of accidents caused by landing errors, ensures the safety of personnel, facilities and the public, and greatly improves the application safety of unmanned aerial vehicles in key fields such as rescue, security and city management.
[0121] 2. Cost reduction and resource conservation: By relying only on computer vision and deep learning models, no additional sensors (such as lidar, IMU, etc.) are needed, which significantly reduces the overall system cost. Simplified hardware architecture and maintenance process help to reduce production and operation costs, indirectly promoting the technological upgrading and economic benefit improvement of related industries.
[0122] 3. Intelligent and high-performance application: This technical scheme realizes the autonomous intelligent landing of the unmanned aerial vehicle, improves the real-time response speed and accuracy of the system, and helps to build an efficient and intelligent unmanned aerial vehicle operation system, promoting the development of smart cities, smart logistics and unmanned delivery systems.
[0123] 4. Environmental and social benefits: While reducing costs and improving efficiency, the application helps to widely promote the application of unmanned aerial vehicle intelligent landing technology in remote disaster relief, emergency material transportation and other fields, bringing long-term benefits to the safety, environmental protection and economic sustainable development of the society. The performance indicators of the application are summarized in the following table:
[0124] Index Ideal performance Positioning accuracy Horizontal positioning accuracy deviation is about 0.2~0.5 m, and height deviation is within ±0.05 m Calculation real-time Through hardware acceleration and algorithm optimization, the processing time delay of a single image can be controlled within 50 milliseconds, and the response delay is less than 20 milliseconds. Adaptability and robustness Homography estimation success rate > 95% Resource occupancy Only rely on camera and computing unit. Model parameter amount is about 5M, and memory peak occupancy is <25M
[0125] The above disclosure is only one or more preferred embodiments of the application, and cannot limit the scope of the rights of the application. Those skilled in the art can understand that the implementation of all or part of the above embodiments, and the equivalent changes made according to the claims of the application, still belong to the scope covered by the application.
Claims
1. A homography-assisted unmanned aerial vehicle landing method based on computer vision and deep learning, characterized in that, The method comprises the following steps: acquiring a visible light image containing a special pattern by a vertical downward camera of the UAV at a center position of a horizontal landing platform containing the special pattern, and storing the image as a fixed preset picture image; when the UAV enters the range of the landing platform, capturing a landing platform related image by the camera, and solving a homography matrix of the preset picture image and a real-time picture image by using a visual model based on deep learning; deriving a three-dimensional coordinate offset matrix of the current position of the UAV relative to the position of the UAV in the preset picture image with the center of the platform as the origin, so as to calculate the three-dimensional coordinates of the current position of the UAV relative to the origin; adjusting the attitude and flight path of the UAV in real time by a control algorithm according to the obtained world coordinates of the UAV relative to the platform, so that the UAV autonomously lands towards the landing point in the preset picture; the visual model based on deep learning comprises: extracting the features of the preset picture image and the real-time picture image by two extractors that do not share parameters, wherein each extractor comprises five ordinary convolution modules and one SimAM parameter-free attention convolution module; aggregating the feature maps of the two types of images, and converting them into graph structure features by an Embedding layer; inputting the graph structure features into an improved ViG model to output a coordinate offset vector, and then obtaining the homography matrix by a DLT algorithm; the basic block processing of the ViG model comprises: processing the input features by two 1x1 convolutions and one 3x3 convolution; performing residual addition on the output of the MLP block and the input features; processing the residual added features by a maximum relative graph convolution block and an FFN block; performing residual addition and feedforward processing again to obtain the output features of the basic block; the specific steps of calculating the three-dimensional coordinates of the current position of the UAV relative to the origin comprise: eliminating the influence of the camera internal parameters to obtain a normalized homography matrix; performing singular value decomposition on the normalized matrix to solve a rotation matrix and a coordinate offset vector; deriving the current three-dimensional coordinates of the UAV in combination with the camera optical center coordinates corresponding to the preset picture; selecting a unique correct solution by physical constraints and projection verification. 2.The computer vision and deep learning based homography-assisted UAV landing method of claim 1, wherein, In the step of "selecting a unique correct solution by physical constraints and projection verification", the method specifically comprises: limiting the pitch angle and roll angle of the UAV to be within ±30°, and limiting the vertical component of the coordinate offset vector to be greater than 0; selecting four marker points on the platform, and verifying the homographic transformation result by coordinate conversion and projection relationship. 3.The computer vision and deep learning based homography-assisted UAV landing method of claim 1, wherein, In the step of "adjusting the attitude and flight path of the UAV in real time by a control algorithm according to the obtained world coordinates of the UAV relative to the platform, so that the UAV autonomously lands towards the landing point in the preset picture", the autonomous landing specifically comprises: correcting the attitude of the UAV based on the rotation matrix, so that the UAV body is horizontally parallel to the platform; calculating the moving direction and position based on the coordinate offset vector, and adjusting the flight path of the UAV; repeating the steps of image acquisition, homography matrix calculation and attitude adjustment every 0.5 seconds until landing. 4.The computer vision and deep learning based homography-assisted UAV landing method of claim 1, wherein, In "acquiring visible light image containing special pattern by vertical downward camera of unmanned aerial vehicle at the center position of horizontal parking platform containing special pattern, and storing it as fixed preset picture image", the method specifically comprises: Accurately parking the unmanned aerial vehicle at the center of the horizontal parking platform containing special pattern, ensuring that the unmanned aerial vehicle body is horizontally parallel to the platform; Adjusting the parameters of the camera vertically downwardly installed at the bottom of the unmanned aerial vehicle, ensuring that the captured picture clearly contains the complete special pattern and the platform edge; Marking the captured visible light image as a positioning reference image, and storing it to the storage unit of the unmanned aerial vehicle as a reference standard for subsequent homography matrix solving.
Citation Information
Patent Citations
Unmanned aerial vehicle pose adaptive estimation method based on active vision
CN110865650A
Autonomous flight method and system based on GAAS, and storage medium
CN111338383A