Runway relative pose estimation method in smooth approach process of machine vision guided unmanned aerial vehicle
By improving the YOLOv5 and MobileSAM models, combined with the P4P position estimation problem, the problem of stable approach and autonomous and safe landing of fixed-wing drones in complex environments is solved, and the autonomous landing and high-precision positioning of the drone without auxiliary equipment is achieved.
Patent Information
- Application Number
- CN202510296792.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively guide fixed-wing drones for smooth approach and autonomous and safe landing in complex environments, especially when satellites and navigation positioning systems are disturbed.
By creating the drone's smooth approach process object detection and image segmentation data set in real scenes, the YOLOv5 target detection network and MobileSAM image segmentation model are trained and improved, the runway area images are obtained for detection and segmentation, the corner location of the landing runway is further extracted, and the position and attitude information of the drone relative to the runway is obtained through the P4P pose estimation problem solution.
It realizes that without ground auxiliary equipment, the drone can only rely on its own on-board camera and processing system to complete autonomous landing, improving the autonomous positioning and landing accuracy and safety of the drone in complex environments.
Smart Images

Figure CN120182379A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle applications, and particularly relates to a method for estimating the relative pose of a runway during the smooth approach process of a machine vision-guided unmanned aerial vehicle. Background Art
[0002] Fixed-wing unmanned aerial vehicles are highly favored in the unmanned aviation equipment industry due to their excellent endurance and load-carrying capabilities. Currently, the mainstream unmanned aerial vehicle equipment mainly uses manual or ground equipment assistance for landing, which requires high experience from the pilot and lacks a certain degree of autonomy. Moreover, a number of fact-finding investigations and studies have shown that the landing process of fixed-wing unmanned aerial vehicles is the stage with the highest accident rate. Therefore, developing a more robust and accurate unmanned aerial vehicle landing guidance platform to reduce the risk of crashing of unmanned aerial vehicles during the frequently occurring accident landing stage is one of the urgent problems to be solved.
[0003] Traditionally, unmanned aerial vehicles usually use a satellite navigation and positioning system combined with an inertial navigation system to provide coordinate and attitude information. Their high mobility itself poses higher requirements for the control system, and more rapid and accurate feedback results need to be given for pose estimation and motion state. On the one hand, unmanned aerial vehicles are limited by the sensitivity of sensors, and the satellite and navigation positioning system are extremely vulnerable to interference, and the positioning information given may have a large deviation. Especially in a denied environment, it is even prone to failure. On the other hand, the cost of deploying an unmanned aerial vehicle landing assistance system on the ground is relatively high, with poor mobility and certain communication delays, and the ability of unmanned aerial vehicles to complete the entire process of tasks autonomously is relatively low. The method of measuring and guiding the landing of unmanned aerial vehicles based on machine vision has the advantages of low cost, low communication delay, and high accuracy, which can greatly enhance the autonomous positioning ability of unmanned aerial vehicles, thereby enhancing the ability of unmanned aerial vehicles to perform tasks autonomously outdoors.
[0004] Currently, most of the methods for guiding the smooth approach of unmanned aerial vehicles based on machine vision are for small rotor unmanned aerial vehicles, and there are fewer guiding methods for the smooth approach of fixed-wing unmanned aerial vehicles. Among the machine vision schemes for guiding fixed-wing unmanned aerial vehicles to perform a smooth approach, traditional image processing operations are mainly used. For example, the airport runway extraction methods mentioned in patents CN202210766485 and CN202310052860 are based on traditional visual image processing methods to identify the runway by using edge detection and calculating the gradient direction. However, in a more complex environment, such as a dark environment, the processing results of such methods are not satisfactory. In addition, there are also some methods for visual image processing based on deep learning. For example, the method of using the YOLO object detection network for runway target recognition in patent CN202311632901 improves the recognition accuracy, but it cannot directly obtain the segmentation result of the runway contour and cannot handle the runway situation under multiple modalities. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a method for estimating the relative pose of a UAV during a smooth approach process and a runway, which is used to guide a fixed-wing UAV to autonomously and safely land without ground auxiliary equipment, or to guide the UAV to smoothly approach and land when the satellite and navigation positioning system are interfered, and can also be used to assist in improving the accuracy and safety of the UAV landing under normal circumstances.
[0006] Technical solution of the present invention: A method for estimating the relative pose of a UAV during a smooth approach process and a runway, the key lies in making a dataset for target detection and image segmentation of the UAV's smooth approach process in a real scene, training an improved YOLOv5 target detection network to detect and locate the runway area image obtained by the on-board vision, and inputting the detection result of the target detection network into the MobileSAM image segmentation model to further accurately segment the landing runway, further extracting the corner positions of the landing runway, matching them with relevant points in the real environment to form control point pairs, constructing a P4P pose estimation problem, and solving it to obtain the position and attitude information of the UAV relative to the runway.
[0007] To achieve the above object, the technical solution of the present invention application is as follows: A method for estimating the relative pose of a runway during the autonomous landing process of a machine vision-guided UAV, comprising the following steps: Step 1: Collect runway pictures from the perspective of a fixed-wing UAV during a smooth approach through publicly available network data and on-board cameras. Images in multiple modalities, multiple time periods, and multiple scenarios need to be collected; Step 2: Annotate the collected runway pictures and make them into a dataset, and a total of a target detection dataset and an image segmentation dataset need to be made. Use the Labelimg annotation tool to annotate the images. The annotation format of the target detection dataset is [class_id, x1, y1, x2, y2], where class_id is the category of the target, and (x1, y1)(x2, y2) are the pixel coordinates of the pixel points at the upper left and lower right corners of the runway area. The core of the annotation format of the image segmentation dataset is the two keys "points" and "label", and its segmentation area is obtained by surrounding with a polygon, and the pixel coordinate data of the corresponding points are stored in "points"; Step 3: Improve on the basis of the YOLOv5 object detection network, and use the improved YOLOv5 object detection network to train on the object detection dataset; the improved YOLOv5 object detection network is optimized based on the original YOLOv5 architecture, mainly focusing on the improvement of the backbone network. The optimized backbone network contains an initial convolutional layer and four downsampling layers, named Layer1, Layer2, Layer3, and Layer4 respectively. Among them, the first and second downsampling layers use the C2F module to extract features to improve the feature extraction efficiency; the third and fourth downsampling layers are composed of a general inverted bottleneck search module, effectively reducing the computational complexity of the network. In terms of feature fusion, the output features of the second, third, and fourth downsampling layers enter the next layer of the network after being fused. Specifically, the output data of Layer4 first undergoes a convolutional operation, and then through the CARAFE upsampling module, it performs feature fusion with the output data of Layer3 to obtain an intermediate tensor. This intermediate tensor is further processed by the C2F module and undergoes another convolutional operation, and then through the CARAFE upsampling module, it performs feature fusion with the output data of Layer2 to finally obtain an optimized feature tensor. This process realizes the effective extraction and fusion of multi-scale features, providing richer semantic information and spatial details for subsequent object detection tasks.
[0008] Step 4: Use the trained object detection network to detect the runway image from the drone's perspective, and obtain the pixel coordinates of the detection box of the runway in the image; Step 5: Input the pixel coordinates of the runway detection box and the center point coordinates of the detection box in the drone's perspective image into the MobileSAM image segmentation model to obtain the mask image of the runway in the picture. To reduce the computing overhead, only the runway within the detection box is segmented. The runway detection box is a rectangular box, and the pixel coordinates of its upper left corner point and lower right corner point can be used to calculate the pixel coordinates of the center point of the detection box, expressed as:
[0009] Step 6: Post-process the mask image in combination with the geometric information of the runway, refine the mask image and obtain the corner point coordinates of the runway. Use the runway processing mechanism PRP (Precision Runway Processing) for processing, and its specific steps are as follows: S01: Perform an image erosion operation on the mask image, using a convolution kernel of size for processing; S02: Perform an image dilation operation on the mask image, using a convolution kernel of size for processing; S03: The center point of the mask area is used as the calculation origin , in the x - direction and y - direction consistent with the definition of the runway image pixel coordinates, map the contour into the defined coordinate system, and use the Canny operator for edge detection to obtain runway edge data ; S04: Perform Harris corner detection on the mask image to obtain the rough positions of the four corner points ; S05: Divide the edge data into four groups corresponding to the four sides of the runway respectively, perform linear function fitting on each of them, and the intersection of the fitting results obtains the accurate coordinates of the runway corner points .
[0010] Step 7: Establish control point pairs by means of the runway corner pixel coordinates to determine the relative pose of the UAV with respect to the runway, which is used to guide the UAV to approach and land smoothly; through the previous step, the coordinates of the four corner points of the runway on the runway mask image can be obtained, and the four corner point coordinates of the runway in the inertial coordinate system in the real world can be constructed into control point pairs to establish a P4P problem. Further use the EPnP method to solve the P4P problem to obtain the pose of the UAV relative to the runway.
[0011] The present invention has the following beneficial effects compared with the prior art: A method for estimating the relative pose of the runway during the machine - vision - guided autonomous landing of a UAV proposed by the present invention obtains a runway mask image under airborne vision by improving the YOLOv5 object detection network and the MobileSAM image segmentation network, and uses the PRP post - processing mechanism to obtain runway corner information. By establishing control point pairs to solve the P4P problem, the pose of the UAV relative to the runway is obtained to guide the UAV to land smoothly. In an environment without ground - assisted landing equipment, through lightweight network design, the UAV can complete autonomous landing only relying on its own on - board camera and its own processing system, providing a solution for autonomous landing. Description of the Drawings
[0012] Figure 1 It is a schematic diagram for object detection and image segmentation annotation.
[0013] Figure 2 It is a structural diagram of the improved YOLOv5 object detection neural network.
[0014] Figure 3 It is a schematic diagram of the input of the MobileSAM image segmentation model.
[0015] Figure 4 It is an example diagram of the result after performing opening operation on the mask.
[0016] Figure 5 It is a schematic diagram of the edge data partitioning method.
[0017] Figure 6 It is a result graph of the performance indicators of the image segmentation model.
[0018] Figure 7 It is a graph of the algorithm output results.
[0019] Figure 8 It is a schematic diagram of the establishment of the scene coordinate system and the relative pose estimation.
[0020] Figure 9 It is a flowchart of the implementation mode of the present invention. Specific implementation mode
[0021] In order to further clarify the purpose, technical solution and advantages of the present invention application, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0022] A method for estimating the relative pose of a runway in the process of a machine vision-guided UAV autonomous landing proposed by the present invention includes the following steps: Step 1: Collect runway pictures from the stable approach perspective of a fixed-wing UAV and label them to make a data set.
[0023] It is necessary to collect airport runway pictures from the landing approach perspective of fixed-wing UAVs in different weather, different scenarios and different time periods to make a data set. In one embodiment, the collected pictures include runway pictures of different airports in different weather such as sunny, cloudy, rainy, foggy, etc. and different time periods such as early morning, noon, evening, night, etc. for the production of a multi-modal data set. As Figure 1 shown, it is necessary to label the collected pictures in the format of an object detection data set and label them in the format of an image segmentation data set respectively. Since the annotation format of the object detection data set uses an external rectangle to enclose the runway area, the object detection data set can be obtained by conversion on the basis of the image segmentation data set. The image segmentation data set uses polygon annotation. For a completely visible runway, four points are required to determine the runway contour, while for a partially occluded runway, more points are required to enclose the runway contour. The annotation points of the pictures using the image segmentation data set annotation method can form an array, which stores all the pixel coordinates and is expressed as:
[0024] Respectively extract the x coordinates and y coordinates among them, and we can get:
[0025] Extract the maximum and minimum values, denoted as , then the pixel coordinates of the upper-left and lower-right pixel points of the rectangular box in the target detection annotation format can be determined:
[0026] Step 2: Use the improved YOLOv5 object detection network to train on the dataset, and use the trained improved YOLOv5 object detection network to perform object detection to obtain the pixel coordinates of the detection box of the runway in the picture. The evaluation metrics of the object detector are Precision, Recall, and Average Precision, and their calculation methods are as follows:
[0027] As Figure 2 shown, the improved YOLOv5 object detection network consists of three structures: backbone, neck, and head. The backbone includes an initial convolutional layer and four downsampling layers, namely Layer1, Layer2, Layer3, and Layer4. The first and second downsampling layers use the C2F module for downsampling; the third and fourth downsampling layers are composed of a general inverted bottleneck search module to reduce the computational complexity of the backbone network. The downsampling modules of Layer2, Layer3, and Layer4 are respectively output to the neck, and the data between different layers are upsampled by the CARAFE module and then concatenated, and finally output to the detection head. The detection head consists of three modules, which respectively detect objects at different resolutions to adapt to the change in the resolution of the runway area in the images captured by the on-board camera during the process of the drone landing from far to near.
[0028] Step 3: Input the coordinates of the detection box into the MobileSAM image segmentation model to obtain the mask image of the runway in the picture, and post-process the mask image to obtain the processed mask image; the MobileSAM image segmentation model needs to input the area to be segmented to specify the object to be segmented, or use the point-specification method to specify the object to be segmented. In this embodiment, the region + point method is adopted to improve the segmentation accuracy. The object detection model can give the detection box of the runway as the input of the MobileSAM model region. Considering the position of the runway at the center of the detection box, a point input can be added by determining the center of the detection box. The pixel coordinates of the upper left corner point and the lower right corner point are , and the pixel coordinates of the center point of the detection box are calculated as follows:
[0029] The above coordinate system definition and input points are as shown in Figure 3 .
[0030] For the output result of the MobileSAM model, the PRP post-processing mechanism (Precision Runway Processing) also needs to be used for processing to obtain a more accurate runway segmentation mask. PRP includes the following steps: S1: First perform five image erosion operations on the mask image, and then perform five image dilation operations, using a convolution kernel of size for processing. This operation can be expressed as:
[0031] where represents the erosion operation, represents the dilation operation, is the graphic operation convolution kernel, represents the image to be processed. After this step of processing, the noise of the image can be filtered and the contour of the picture can be smoothed. The processed picture is as shown in Figure 4 .
[0032] S2: Use the Canny edge operator for edge detection, take the center point of the mask area as the calculation origin, and the x-direction and y-direction consistent with the pixel coordinate definition of the runway image, and map the contour to the defined coordinate system to obtain the runway edge data .
[0033] S3: Use the Harris corner detection method for the mask image to obtain the rough positions of the four corner points of the runway , and divide the edge data into four groups according to the corner detection results , and perform function fitting on them.
[0034] The fitting basic model is as follows: , where n and m are undetermined parameters.
[0035] The fitting results of the four groups of data by the linear function intersect to obtain four accurate runway corner pixel coordinates , and the division and determination method can be seen in Figure 5 .
[0036] The effect of image segmentation is evaluated by the intersection over union (IoU):
[0037] The results of its various indicators can be seen in Figure 6 .
[0038] The effect of the algorithm on image segmentation can be seen in Figure 7 . The runway at night is usually a very challenging task in image segmentation. The present invention can handle the airport runway situation under multi-modal, and has a better effect on segmenting the runway at night.
[0039] Step 4: Further determine the pixel coordinates of the runway corners through the mask image, and establish control point pairs with the help of the runway corner pixel coordinates to determine the relative pose of the UAV with respect to the runway, so as to guide the UAV to approach and land smoothly.
[0040] The pixel coordinates and the three-dimensional world coordinates are described in two coordinate systems. The definition of the three-dimensional world coordinate system is given in Figure 8 , so only the conversion problem between the two coordinate systems needs to be considered. The transformation from a point in the world coordinate system to the camera coordinate system can be represented by an orthogonal rotation matrix R and a translation rotation matrix T, and its conversion formula is:
[0041] In the formula, is the coordinate of the origin of the world coordinate system in the camera coordinate system.
[0042]
[0043] Among them, is the pitch angle (positive when the UAV looks up), the roll angle (positive when the UAV tilts to the right), and the yaw angle (positive when the UAV yaws to the right).
[0044] After constructing the coordinate system conversion relationship, the next step is to solve using the corner pixel coordinates in the mask image and the control points in the three-dimensional world. The three-dimensional coordinate points in the world coordinate system are represented by four control points, and then the PnP problem is simplified to estimate the coordinates of the four control points in the camera coordinate system. The solution method is to represent these unknown coordinates as a The weighted sum of the eigenvectors of the matrix is calculated, and a constant quadratic equation is solved to select the correct weights. Since the optimization is only performed on four control points, the solution speed is very fast and the accuracy is high.
[0045] In the world coordinate system, four control points can be calculated through all three-dimensional space points, and the coordinates of the control points in the sight coordinate system are obtained. The three-dimensional space points are represented as , and the corresponding four control points are represented as . The calculation formula of the control points is expressed as:
[0046] Each reference point is represented as a weighted sum of the control point coordinates using four control points:
[0047] In the formula are homogeneous barycentric coordinates. Similarly, there is the same weighted sum relationship in the camera coordinate system:
[0048] Next, it is necessary to solve the coordinates of the control points in the camera coordinate system. The coordinates of the control points in the camera coordinate system are set as , the pixel coordinates , the focal length parameter of the camera is , the principal point coordinates of the image are , then there is:
[0049] After eliminating the projection parameters , the following two linear equations can be obtained:
[0050] After concatenating the linear equations obtained from all control points, the coordinates of the control points in the camera coordinate system can be solved. By obtaining the coordinates of the four control points in the world coordinate system and the camera coordinate system and using the 3D-3D matching method, R and T can be calculated to determine the pose of the UAV relative to the runway.
[0051] It should be noted that there are already many mature pose estimation schemes in the technical field. The present invention can provide accurate runway corner point positioning data, and the present invention is not limited to any pose estimation algorithm. The above EPnP method is an example of the combination of the present invention with a pose estimation method.
Claims
1. A method for estimating the relative pose of a runway during autonomous landing of a UAV guided by machine vision, including four steps: Step 1: Collect runway images from the perspective of a fixed-wing drone’s steady approach and annotate them to create a dataset; Step 2: Use an improved YOLOv5 target detection network to train on the data set, and use the trained improved YOLOv5 target detection network to perform target detection to obtain the pixel coordinates of the detection box of the runway in the picture; Step 3: Input the detection frame coordinates into the MobileSAM image segmentation model to obtain a mask image of the runway in the picture, and post-process the mask image to obtain a processed mask image; Step 4: The pixel coordinates of the runway corner points are further determined by the mask image, and control point pairs are established with the help of the pixel coordinates of the runway corner points to determine the relative position of the UAV to the runway, so as to guide the UAV to smoothly approach and land.
2. The method for estimating the relative position of a runway during autonomous landing of a UAV guided by machine vision as claimed in claim 1, characterized in that: Step 1 includes: using airborne vision to collect images of fixed-wing UAV take-off and landing runways in real scenes at multiple time periods and in multiple scenes, annotating the runway contours in the images using the Labelimg tool to obtain the corresponding mask image, and at the same time selecting the runway area to obtain the pixel coordinates of the minimum circumscribed rectangular frame of the runway. The original image and its corresponding mask image, and the original image and its corresponding pixel coordinates are combined to produce a target detection dataset and a target segmentation dataset.
3. The method for estimating runway relative posture during autonomous landing of a UAV guided by machine vision as claimed in claim 1, characterized in that: Step 2 includes: the improved YOLOv5 target detection network optimizes the backbone network based on the original YOLOv5 target detection network, and the optimized backbone includes an initial convolution layer and four downsampling layers, namely Layer 1, Layer 2, Layer 3 and Layer 4. The first and second downsampling layers are sampled using the C2F module; the third and fourth downsampling layers are composed of a general inverted bottleneck search module to reduce the computational complexity of the backbone network. The second, third and fourth downsampling layers are processed by feature fusion to obtain tensors Enter the next layer of network structure, where Layer4 data After convolution, the CARAFE upsampling module and Layer3 data Perform feature fusion to obtain tensor , the tensor After being processed by the C2F module and convolved, it is then passed through the CARAFE upsampling module and combined with the Layer2 data. Perform feature fusion to obtain tensor , complete the extraction and fusion of data.
4. The method for estimating runway relative posture during autonomous landing of a UAV guided by machine vision as claimed in claim 1, characterized in that: Obtaining the mask image in step 3 includes: inputting the frame selection result obtained according to the target detection neural network into the MobileSAM image segmentation model, outputting the contour of the drone runway detected in the target detection frame selection area, and segmenting the mask image of the drone runway contour area.
5. The method for estimating runway relative posture during autonomous landing of a UAV guided by machine vision as claimed in claim 1, characterized in that: The post-processing of the mask image in step 3 includes: performing image processing on the mask image output by the MobileSAM image segmentation model through edge detection and opening operation, determining the corner points of the mask area according to the target detection frame selection result, and fitting the edge of the runway contour according to the edge of the mask area to determine the precise segmentation position of the runway.
6. The method for estimating runway relative posture during autonomous landing of a UAV guided by machine vision as claimed in claim 1, characterized in that: Step 4 includes: S601: Process the image using the PRP mechanism. Specifically, the mask image is first eroded and then expanded. The Canny operator is further used for edge detection. The center point of the mask area is taken as the calculation origin. The contour is mapped to the defined coordinate system in the x and y directions consistent with the pixel coordinates of the runway image to obtain the runway edge data. ; Then use the Harris corner detection method on the mask image to obtain the rough positions of the four corner points of the runway , according to the corner point detection results, the edge data is divided into four groups , a function fitting is performed on it, and the four sets of data are intersected to obtain the four precise pixel coordinates of the runway corner points ; S602: The four precise pixel coordinates of the runway corner points are matched with the three-dimensional coordinates of the runway corner points in space to form a P4P solution problem. The P4P problem is solved using the EPnP method to obtain the position and attitude information of the UAV relative to the runway.
7. The method of claim 4, wherein the target detection neural network obtains the frame selection result and inputs it into the MobileSAM image segmentation model, comprising: According to the runway location area output by the improved YOLOv5 target detection neural network, four corner points of a detection box are used, and the center point of the runway location area is calculated based on the four corner points. A total of five points are used as the input of the MobileSAM image segmentation model.
Citation Information
Patent Citations
Visual positioning method for fixed-wing unmanned aerial vehicle in landing stage
CN115272458A
Airport runway extraction method based on image processing
CN115797374A
Unmanned aerial vehicle autonomous landing method and device based on visual image guidance
CN117636284A