A high-precision relative positioning method for unmanned aerial vehicles based on binocular depth point cloud

By using binocular depth point cloud and model predictive control algorithms, high-precision relative positioning of UAVs during aerial refueling is achieved, solving the problem of unstable sensor data and improving the safety and efficiency of refueling operations.

CN118210329BActive Publication Date: 2026-03-31BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-precision relative positioning of UAVs during aerial refueling, especially under high-speed flight and complex weather conditions. Unstable sensor data processing limits the safety and efficiency of refueling operations.

Method used

A high-precision relative positioning method for UAVs based on binocular depth point clouds is adopted. Images are captured by binocular cameras to generate depth point clouds, the YOLOv7 detection model is optimized to generate 3D point cloud detection areas, and the UAV attitude is adjusted using model predictive control algorithms to achieve accurate target recognition and control.

Benefits of technology

It improves the positioning accuracy and stability of UAVs during aerial refueling, ensuring the safety and efficiency of refueling operations, and is suitable for real-time applications in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118210329B_ABST
    Figure CN118210329B_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned aerial vehicle high-precision relative positioning method based on binocular depth point cloud, belong to sensor information fusion technical field.The method includes the following steps: using binocular camera to capture target image in the air, according to the image of binocular camera the depth point cloud in entire field of view is calculated;YOLOv7 model is detected to binocular image using modified, and the 2D detection frame of conical sleeve target is obtained;According to 2D detection frame, the 3D point cloud detection area of target is generated, and the mass point of dense point cloud is calculated as target point;Based on model predictive control algorithm, the attitude of unmanned aerial vehicle is accurately adjusted, and unmanned aerial vehicle is controlled.The application uses the above-mentioned one kind of unmanned aerial vehicle high-precision relative positioning method based on binocular depth point cloud, so that unmanned aerial vehicle is in air relative positioning, quickly and accurately identify target object, and realize accurate attitude control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sensor information fusion technology, and in particular to a high-precision relative positioning method for unmanned aerial vehicles based on binocular depth point clouds. Background Technology

[0002] Relative positioning technology between unmanned aerial vehicles (UAVs) is a complex field developed on a multidisciplinary basis, primarily driven by military needs. Initially, this technology was designed to perform complex formation flying and tactical missions to improve the efficiency and safety of military operations. Over time, the increasing civilian applications of UAVs, such as disaster monitoring, traffic management, and agricultural monitoring, have further spurred the development of relative positioning technology. Relative positioning technology enables UAVs to accurately determine their position and motion status without external reference points, which is crucial for performing precise team missions.

[0003] In modern warfare, unmanned aerial vehicles (UAVs) are playing an increasingly important role in conducting ultra-long-duration reconnaissance and long-range strike missions. However, the execution of these missions is limited by the payload and fuel capacity of UAVs. To overcome these limitations, aerial refueling technology is used, aiming to significantly improve the flight range, loiter time, and payload of UAVs without making major modifications to their structure. This technology has significantly enhanced the performance of UAVs in long-range flights, target detection, and weapon carrying missions. During aerial refueling, precise relative positioning technology is particularly important because the tanker and the UAV must maintain a strict relative position in flight to ensure the safety and efficiency of the refueling operation. The challenge of this technology lies in the need for both UAVs to achieve extremely high positioning accuracy and stability at high speeds, while also considering aerodynamic effects and potential environmental interference. In summary, the development of aerial refueling technology based on high-precision relative positioning of UAVs has not only improved the combat capabilities of UAVs in modern warfare but also promoted advancements in UAV technology in the fields of automatic navigation and precision control. Through continuous technological innovation, the application scope and efficiency of UAVs will continue to expand, providing greater flexibility and stronger strategic advantages for modern military operations.

[0004] The docking phase of aerial refueling requires the receiving UAV to automatically analyze sensor data and generate corresponding control commands within limited autonomous operating permissions to achieve precise, efficient, and safe refueling operations. This is not only a challenge in flight technology but also a test of the UAV's autonomous navigation and automatic control systems.

[0005] First, from the perspective of the receiver aircraft's attitude control, ensuring the safety and stability of the docking is crucial. This requires the UAV to possess highly precise flight path control and speed matching capabilities. The receiver UAV must be able to coordinate closely with the tanker aircraft and provide real-time feedback and adjustments to cope with any changes in the dynamic environment. This not only requires the UAV to have an advanced flight control system, but also to be able to process large amounts of sensor data in real time to achieve precise dynamic flight adjustments.

[0006] Meanwhile, precise detection and positioning of the tanker's drogue is crucial for successful in-flight refueling. This requires the UAV's sensors and navigation systems to accurately identify the tanker's position and status, and make precise flight adjustments accordingly. The UAV needs to be able to accurately dock with the tanker under varying weather conditions and potential electronic interference. Furthermore, any minute errors that may occur during refueling must be detected and corrected in real time to avoid collisions or other dangerous situations.

[0007] The detection and positioning of the refueling boom is a crucial aspect of aerial refueling technology, but existing technologies have certain limitations. While various methods have been used to address this challenge, each has its own limitations. For example, while inertial measurement units (IMUs) and global positioning systems (GPS) perform well in most situations, they may fail to provide sufficiently accurate navigation information in certain complex flight environments. This is primarily because GPS signals can be interfered with during high-speed flight and under complex weather conditions, and IMUs can accumulate significant errors. Sensor technologies relying on infrared and lidar (LIDAR), while providing accurate distance and velocity data, are highly sensitive to changes in ambient light, which can lead to performance instability under certain lighting conditions. For instance, the effectiveness of infrared and lidar sensors may be affected by strong sunlight or cloud cover variations. Furthermore, meeting the specific relative positioning requirements of these sensors may necessitate significant design improvements to the refueling probe and boom. This not only increases the difficulty of system implementation but also raises the overall complexity and cost of the refueling system.

[0008] Therefore, to address these challenges, this patent employs a model predictive control (MPC) approach as an innovative and efficient solution. Summary of the Invention

[0009] The purpose of this invention is to provide a high-precision relative positioning method for UAVs based on binocular depth point clouds, enabling UAVs to quickly and accurately identify targets and achieve precise attitude control during aerial relative positioning.

[0010] To achieve the above objectives, this invention provides a high-precision relative positioning method for unmanned aerial vehicles (UAVs) based on binocular depth point clouds, comprising the following steps:

[0011] S1. Use a binocular camera to capture images of the target object;

[0012] S2. Based on the image data captured by the binocular camera in step S1, use the binocular stereo vision algorithm to generate a depth point cloud in the entire field of view in real time.

[0013] S3. Optimize the backbone network of the YOLOv7 detection model;

[0014] S4. Use the optimized YOLOv7 model to accurately detect images acquired from the stereo camera and generate corresponding 2D detection boxes;

[0015] S5. Based on the 2D detection box, locate the three-dimensional spatial position of the cone-shaped target in the depth point cloud, generate the corresponding 3D point cloud detection area, optimize the calculation of dense areas of the point cloud, and apply mean filtering to improve the quality and accuracy of the point cloud data.

[0016] S6. Continuously track the determined target points and store the position information of the target points in continuous images;

[0017] S7. Calculate the average relative velocity of the target point by calculating the position information;

[0018] S8. Based on model predictive control algorithm, accurately adjust the attitude of the UAV and control the UAV.

[0019] Preferably, step S2 generates a depth point cloud covering the entire field of view in real time, including image correction, matching, and depth calculation, to obtain the three-dimensional coordinates of each point in space, as shown below:

[0020] S21. Image correction, the specific steps are as follows:

[0021] S211. Use Zhang's calibration method to obtain the distortion coefficient of the camera. Combine the intrinsic parameters provided by the camera manufacturer to calibrate the relative position and orientation between the two cameras, i.e., the extrinsic parameters include rotation and translation.

[0022] S212. Based on the parameters obtained in step S211 above, calculate the transformation matrix required for stereo correction; this includes two rotation matrices, used for the correction of the left and right camera images respectively, and two new projection matrices.

[0023] S213. Using the rotation matrix and projection matrix described above, the original image is remapped as follows:

[0024] p1'=K1'R1R -1 K1 -1 p1

[0025] p2'=K2'R2(R-1 K2 -1 p2+T)

[0026] Where K1 and K2 are the intrinsic parameter matrices of the left and right cameras respectively, R and T are the rotation matrix and translation vector between the two cameras respectively, K1', K2' and R1 and R2 are the intrinsic parameter matrices and rotation matrices that need to be recalculated, and p1' and p2' are the corrected pixel positions. Using the above transformation, the images of the two cameras can be converted into a common view plane, and the projection of the same object point in the two images has the same y coordinate.

[0027] S22. Based on the improved cost function, image matching is performed by introducing an adaptive Gaussian weight matrix and normalization, as shown in the following formula:

[0028]

[0029] Where W is the window region, used to define the size of the region considered around each pixel; I L ,I R Here are two images to be compared; p and q are images I and II respectively. L ,I R The corresponding pixel in the window is i,j is the relative position offset within the window area; W(i,j,p) is the adaptive Gaussian weight, and (p+i,p+j) is the offset pixel;

[0030] S23. The depth calculation formula is as follows:

[0031]

[0032] Where Z is the pixel depth, f is the camera focal length; B is the baseline distance between the two cameras, i.e., the horizontal distance between the two cameras; and D is the parallax value, i.e., the difference in horizontal position of the same scene point in the images of the two cameras.

[0033] S24. Convert the 2D image coordinates and depth values ​​to 3D spatial coordinates using the following formula:

[0034]

[0035] Where (x,y) are the coordinates of a pixel in the image plane, (X,Y,Z) are the coordinates of a pixel in the world coordinate system, and (c x ,c y ) represents the optical center coordinates of the image.

[0036] Preferably, in step S3, the optimization of the backbone network of the YOLOv7 detection model involves the following specific steps:

[0037] S31. Set a threshold for the convolutional layer weights and remove weights that are less than the threshold, using the following formula:

[0038]

[0039] Where, ω i,j ' represents the pruned weight matrix, θ is the preset threshold, and ω i,j Let be the element in the i-th row and j-th column of the weight matrix;

[0040] S32. Use the CBAM module to perform spatial attention weighting, and adjust the weight allocation method of CBAM as follows:

[0041] S cone (x,y)=σ(f centre (I(x,y))+f dark (I(x,y)))

[0042] Where σ is the activation function, f centre and f dark These correspond to the attention weights of the central area of ​​the refueling cone sleeve and the dark-colored block, respectively.

[0043] S33, Spatial attention map S cone Element-wise multiplication is applied to the original feature map F, as shown below:

[0044]

[0045] S34. Depthwise separable convolution is used to extract features of the refueling cone sleeve more efficiently, and a medium-scale feature map is used for detection. The formula is as follows:

[0046] P cone =F mid

[0047] Among them, P cone It is a feature map used for detecting refueling cone sleeves, F mid It is a feature map of moderate scale in the network.

[0048] Preferably, in step S5, the specific steps are as follows:

[0049] S51. Locate the 3D spatial position of the cone-shaped target in the depth point cloud based on the 2D detection bounding box, and generate the corresponding 3D point cloud detection area. Remove objects that are far from the camera or too close to it, as shown below:

[0050]

[0051] Among them, D roi (x,y) represents the point cloud depth value of the detection area;

[0052] S52. Optimize the calculation of dense regions in the point cloud, using the following formula:

[0053]

[0054] Where l is an indicator function that generates a binary output for a given condition. It returns 1 if the condition is true and 0 otherwise.

[0055] S53. Apply mean filtering to the effective point cloud, using the following formula:

[0056]

[0057]

[0058] Preferably, in step S6, when storing the location information of the target point, the position of the centroid of the dense point cloud is calculated as the target point, using the following formula:

[0059]

[0060]

[0061]

[0062] Among them, (X) center ,Y center Z center () is a relative target point in 3D space.

[0063] Preferably, the average velocity of the target point is calculated in step S7 as follows:

[0064]

[0065] Where, p i =(x i ,y i ,z i ) represents the position coordinates of the target point in the i-th frame of the image.

[0066] Preferably, in step S8, when constructing the state and control input model of the UAV, based on the model predictive control algorithm, the state scalars of the UAV include position (x,y,z), yaw angle θ, and pitch angle ψ;

[0067] The control scalars include: velocity v, yaw rate, and roll rate pitch; therefore, the dynamic model of the UAV can be expressed as:

[0068]

[0069] Preferably, based on the dynamics model of the UAV, the objective function for the aerial refueling scenario is as follows:

[0070] S811. First, calculate the minimum distance between the UAV and the target point, using the following formula:

[0071]

[0072] Among them, (X) center,k ,Y center,k Z center,k ω represents the current position of the target point relative to the drone. pos It is the weight of the position error;

[0073] S812. Secondly, minimize the rate of change of the UAV's flight speed, as shown in the following formula:

[0074]

[0075] Where ω ctrl These are the weights that control changes in input;

[0076] S813. Finally, the objective function is as follows:

[0077]

[0078] Preferably, based on the dynamics model of the UAV, the objective function for the relative positioning scenario of multiple UAVs is as follows:

[0079] S821. Tracking error is calculated by determining the Euclidean distance between the current position of the UAV and the target position, as shown below:

[0080]

[0081] Among them, X k ω1 is the position of the UAV at the kth time step, and ω1 is the weighting factor.

[0082] S822, Field of view correction, as shown below:

[0083]

[0084] Where ω2 is the weighting factor, and ellipse is a function of the field of view and the UAV's position; this function can be calculated based on the UAV's position, orientation, and field of view.

[0085] S823. In 3D space, the field of view is considered as an ellipsoid, and the target is located inside this ellipsoid. The formula for the elliptic function is as follows:

[0086]

[0087] Where (x,y,z) is the position of the drone, (x...u ,y u ,z u ) represents the location of the target point.

[0088] Therefore, the present invention employs the above-mentioned high-precision relative positioning method for UAVs based on binocular depth point clouds, and the beneficial effects are as follows:

[0089] 1. Use the centroid of the dense point cloud region within the bounding box instead of the center point of the bounding box as the target point. In actual 3D space, the shape and pose of an object may cause part of its surface to be closer to the camera, and the target may not necessarily fill the bounding box. Using the centroid can more accurately represent the actual position of the object, while avoiding detection jitter caused by bounding box jitter.

[0090] 2. When the target moves or rotates, the position change of the centroid is smoother than the position change of the detection box center, thus providing more stable target tracking.

[0091] 3. By considering the changes in the UAV's state and environment over several future time steps (discrete time intervals), advanced path planning and UAV docking are achieved. By continuously updating the target position, the proposed MPC scheme enables the UAV to dynamically track moving targets, making it suitable for aerial refueling scenarios. It is also particularly important for applications such as search and rescue, and surveillance.

[0092] 4. Visibility constraints due to limited field of view have been specifically considered. This ensures that the target remains within the camera's field of view when the UAV is performing tracking tasks, and the MPC controller can generate control commands at a high frequency, which is feasible on most embedded computing platforms, indicating that this solution is suitable for real-time applications.

[0093] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0094] Figure 1 This is a flowchart of an aerial refueling docking process for UAVs, which is based on a high-precision relative positioning method for UAVs using binocular depth point clouds.

[0095] Figure 2 It is an improved SAD algorithm based on binocular depth point cloud for high-precision relative positioning of UAVs, and a disparity map of binocular image weights.

[0096] Figure 3 This is a schematic diagram illustrating the weight allocation of dark color blocks in an improved SAD algorithm for high-precision relative positioning of UAVs based on binocular depth point clouds.

[0097] Figure 4 This is a block diagram of an improved YOLOv7 recognition algorithm for high-precision relative positioning of UAVs based on binocular depth point clouds;

[0098] Figure 5 This is a high-precision relative positioning method for UAVs based on binocular depth point clouds, used for the identification of UAV target objects.

[0099] Figure 6 This is a simulation trajectory of MPC-controlled UAV tracking dynamic targets, based on a high-precision relative positioning method for UAVs using binocular depth point clouds. Detailed Implementation

[0100] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0101] like Figure 1 As shown, a high-precision relative positioning method for UAVs based on binocular depth point clouds includes the following steps:

[0102] S1. Use a binocular camera to capture images of the refueling machine and the cone-shaped sleeve;

[0103] S2. Based on the image data captured by the binocular camera in step S1, use the binocular stereo vision algorithm to generate a depth point cloud in the entire field of view in real time.

[0104] S3. Optimize the backbone network of the YOLOv7 detection model;

[0105] S4. Use the optimized YOLOv7 model to accurately detect images acquired from the stereo camera and generate corresponding 2D detection boxes;

[0106] S5. Based on the 2D detection box, locate the three-dimensional spatial position of the cone-shaped target in the depth point cloud, generate the corresponding 3D point cloud detection area, optimize the calculation of dense areas of the point cloud, and apply mean filtering to improve the quality and accuracy of the point cloud data.

[0107] S6. Continuously track the determined target point and store the target point's location information in 10 consecutive frames of images;

[0108] S7. By calculating the position information, the average relative velocity of the target point over 10 frames is obtained;

[0109] S8. Based on model predictive control algorithm, accurately adjust the attitude of the UAV and control the UAV.

[0110] Example

[0111] S1. Use a binocular camera to capture images of the refueling aircraft and the cone in the aerial refueling docking scenario.

[0112] Within the ROS framework, the stereo camera creates nodes usb_cam_node1 and usb_cam_node2, where 1 and 2 are the stereo camera node numbers. Each node is an independent process within the ROS framework, describing different functionalities. The stereo camera, with its image capture capability, is a camera sensor that publishes RGB image messages to cam_image1 and cam_image2, respectively.

[0113] S2. Based on the image data captured by the binocular camera, a depth point cloud covering the entire field of view is generated in real time using a binocular stereo vision algorithm. This step includes image correction, matching, and depth calculation to obtain the three-dimensional coordinates of each point in space.

[0114] S21. The specific steps for image correction are as follows:

[0115] S211. First, perform image correction. Use the "Zhang's calibration method" to obtain the distortion coefficient of the camera. Combine the intrinsic parameters provided by the camera manufacturer to calibrate the relative position and orientation between the two cameras, i.e., the extrinsic parameters include rotation and translation.

[0116] S212. Based on the parameters obtained in step S211, calculate the transformation matrix required for stereo correction. The transformation matrix includes two rotation matrices, used for correcting the left and right camera images respectively, and two new projection matrices.

[0117] S213. Finally, using the rotation matrix and projection matrix described above, the original image is remapped as follows:

[0118] p1'=K1'R1R -1 K1 -1 p1

[0119] p2'=K2'R2(R -1 K2 -1 p2+T)

[0120] Where K1 and K2 are the intrinsic parameter matrices of the left and right cameras, respectively; R and T are the rotation matrix and translation vector between the two cameras, respectively; K1', K2', R1, and R2 are the intrinsic parameter matrices and rotation matrices that need to be recalculated; and p1' and p2' are the corrected pixel positions. Using the above transformation, the images of the two cameras can be converted into a common view plane, and the projection of the same object point in the two images will have the same y-coordinate.

[0121] The above steps can be completed automatically using the stereo correction function in the OpenCV library.

[0122] S22. Based on the improved cost function, image matching is achieved by introducing an adaptive Gaussian weight matrix and normalization.

[0123] like Figure 2 As shown, the disparity map represents the pixel offset between two images. For each pixel in the left image, a search is conducted within a certain range in the right image for its corresponding pixel. This search process involves a special aerial refueling vision environment, using a modified SAD (Sum of Absolute Differences) cost function to find the optimal matching position to obtain the offset or disparity value of that pixel in the disparity map.

[0124] The following are the main features of the improved cost function:

[0125] 1. Adaptive Gaussian Weight Matrix: The cost function introduces an adaptive Gaussian weight matrix. During pixel matching, the surrounding environment of each pixel is considered. In particular, within the matching window, dark pixel regions will receive higher weights than blue or white background pixels. Figure 3 As shown, this helps to accurately match pixels even under complex lighting conditions.

[0126] 2. Normalization: Normalizing pixel values ​​helps reduce the impact of lighting and contrast on image matching. The cost function can better handle images under different lighting conditions, ensuring the stability and accuracy of matching.

[0127] The improved cost function is shown below:

[0128]

[0129] Where W is the window region, used to define the size of the region around each pixel for contrast; I L ,I R Here are two images to be compared; p and q are images I and II respectively. L ,I R The corresponding pixels in the diagram, i,j are the relative position offsets within the window region; W(i,j,p) represents the adaptive Gaussian weights, as shown below:

[0130] W(i,j,p)=G(i,j)×f(p+i,p+j)

[0131] Where (p+i,p+j) are the offset pixel points, and G(i,j) is the standard Gaussian weight function, as shown below:

[0132]

[0133] Where σ is the standard deviation of the Gaussian function, controlling the range of the weight distribution; the function f(x) is the pixel value adaptive adjustment of the Gaussian weights, as shown below:

[0134]

[0135] Where α and β are adaptive parameters set according to the specific lighting environment, and L(p) is a function that calculates the image grayscale value based on the human eye's sensitivity to different colors, as shown below:

[0136] L(p)=0.299Rp+0.587Gp+0.114Bp

[0137] Here, minL and maxL are the set grayscale value ranges, which can be simply set to 0 and 255; when the pixel block color is darker, the adaptive parameter is more biased towards β; when the pixel block is sky blue or white, the adaptive parameter is more biased towards α.

[0138] S23. After obtaining the disparity map, the depth of each pixel is calculated using the following formula:

[0139]

[0140] Where Z is the pixel depth, f is the camera focal length, B is the baseline distance between the two cameras, i.e., the horizontal distance between the two cameras; and D is the parallax value, i.e., the difference in horizontal position of the same scene point in the images of the two cameras.

[0141] S24. Convert the 2D image coordinates and depth values ​​to 3D spatial coordinates using the following formula:

[0142]

[0143] Where (x,y) are the coordinates of a pixel in the image plane, (X,Y,Z) are the coordinates of a pixel in the world coordinate system, and (c x ,c y The coordinates of the optical center (principal point) of the image are the intrinsic parameters of the camera. This gives each image pixel a 3D coordinate, thus forming a depth point cloud.

[0144] S3. Optimize the backbone network of the YOLOv7 detection model.

[0145] like Figure 4 As shown, an attention mechanism is introduced to improve the model's ability to recognize cone sleeve features; the pooling layer is improved to enhance the model's robustness when handling cone sleeves of different sizes; the loss function is optimized to improve detection accuracy; the improved YOLOv7 detection model adapts to the real-time and lightweight requirements of UAV computing platforms.

[0146] S31. To address the specific characteristics of aerial refueling scenarios, the YOLOv7 backbone network was modified. A novel convolutional layer was added to CSPDarkNet, and redundant convolutional and fully connected layers were removed to simplify the model structure. The model was pruned by setting a threshold for the weights of each layer. If the absolute value of a weight is less than the threshold, it is either deleted or set to zero. The formula is as follows:

[0147]

[0148] Where, ω i,j ' represents the pruned weight matrix, θ is the preset threshold, and ω i,j is the element in the i-th row and j-th column of the weight matrix.

[0149] Through the above operations, the model is made lightweight and more suitable for real-time aerial flight missions. This not only improves the model's computational efficiency but also helps reduce hardware resource requirements to meet the requirements of UAVs performing vision tasks, such as cone detection and relative pose tracking.

[0150] S32. To ensure the model focuses more on the central region and dark areas of the gas cone, the Convolutional Block Attention Module (CBAM) is used for spatial attention weighting. Furthermore, considering that the background is mostly blue sky, the weight distribution of CBAM is adjusted accordingly.

[0151] S cone (x,y)=σ(f centre (I(x,y))+f dark (I(x,y)))

[0152] Where σ is the activation function, f centre and f dark The attention weights correspond to the central area of ​​the refueling cone and the dark-colored block, respectively.

[0153] S33, Spatial attention map S cone Element-wise multiplication is applied to the original feature map F, as shown below:

[0154]

[0155] S34. Due to the singular detection target and usage scenario, depthwise separable convolution is used to extract features from the refueling cones more efficiently. Furthermore, considering the relatively fixed size of the refueling cones, a medium-scale feature map is used directly for detection, without requiring a complete feature pyramid structure. The formula is as follows:

[0156] P cone =F mid

[0157] Among them, P cone It is a feature map used for detecting refueling cone sleeves, F mid It is a feature map of moderate scale in the network.

[0158] S4. Use the optimized YOLOv7 model to accurately detect targets in images acquired from the stereo camera. This model can accurately identify targets in stereo images and generate corresponding 2D detection boxes, such as... Figure 5 As shown;

[0159] S5. Locate the three-dimensional spatial position of the cone-shaped target in the depth point cloud based on the 2D detection box, generate the corresponding 3D point cloud detection area, optimize the calculation of dense areas of the point cloud, and apply mean filtering to improve the quality and accuracy of the point cloud data.

[0160] S51. Extract the regions corresponding to the 2D detection boxes from the complete depth map. Values ​​outside the sub-regions are set to 0 to reduce the amount of data for subsequent processing and ensure focus on regions that may contain targets. Define a depth range to remove objects that are far from the camera or too close, as shown below:

[0161]

[0162] Among them, D roi (x,y) represents the point cloud depth value of the detection area.

[0163] S52. Dense regions refer to areas in a point cloud with a high number of points. These regions typically represent the surface of an object, rather than its space or noise. The local density of a point is estimated by counting non-zero points within a small window surrounding that point. The size of the window is determined by r, which represents the window radius. The calculation of dense regions in a point cloud is optimized using the following formula:

[0164]

[0165] Here, l is an indicator function that generates a binary output for a given condition; specifically, it returns 1 if the condition is true, and 0 otherwise.

[0166] S53. To reduce the impact of noise and outliers, mean filtering is applied to the effective point cloud, using the following formula:

[0167]

[0168]

[0169] S6. Continuously track the determined target point and store its location information in continuous images, using the centroid of the dense point cloud as the target point, as shown below:

[0170]

[0171]

[0172]

[0173] Among them, (X) center ,Y center Z center () is a relative target point in 3D space.

[0174] S7. Calculate the average velocity of the target point by calculating the position information;

[0175] For single-target recognition scenarios, Kalman filtering is used to smooth noise in the detection algorithm, reducing its interference with target recognition. Furthermore, Kalman filtering can predict the inter-frame position of the target point even when it is temporarily lost, providing a continuous and consistent target position estimate even when the target cannot be directly observed.

[0176] To further reduce the impact of target point jitter, this step also includes storing estimated position data for 10 consecutive frames. The average velocity of the target point across the 10 frames is calculated as follows:

[0177]

[0178] Where, p i =(x i ,y i ,z i ) represents the position coordinates of the target point in the i-th frame of the image.

[0179] This calculation not only smooths out fluctuations in the target's position but also provides crucial dynamic information about the target's movement.

[0180] S8. Based on model predictive control algorithm, accurately adjust the attitude of the UAV and control the UAV.

[0181] Based on the Model Predictive Control (MPC) algorithm, the UAV automatically detects its attitude angles, obtaining pitch and yaw angles. This data is used to construct a model of the UAV's state and control inputs. The MPC then controls the UAV to track a simulated trajectory of a dynamic target. Figure 6 As shown.

[0182] The state scalars of the UAV include: position (x, y, z), yaw angle θ, and pitch angle ψ; the control scalars include: velocity v, yaw rate, and pitch rate; therefore, the dynamic model of the UAV can be expressed as:

[0183]

[0184] S81, designed for aerial refueling scenarios;

[0185] S811. First, minimize the distance between the drone and the target point, as shown below:

[0186]

[0187] Among them, (X) center,k ,Y center,k Z center,k ω represents the current position of the target point relative to the drone. pos It is the weight of the position error.

[0188] S812. Secondly, to ensure stable flight, it is necessary to minimize the rate of change of the UAV's flight speed, as shown in the following formula:

[0189]

[0190] Where, ω vel This represents the weighting of speed changes; to ensure the smoothness of control inputs (such as yaw rate and pitch rate), the formula is as follows:

[0191]

[0192] Where, ω ctrl It is the weight that controls the changes in input.

[0193] S813. Finally, for the aerial refueling scenario, the objective function is expressed as:

[0194]

[0195] S82, designed for relative positioning scenarios involving multiple drones;

[0196] The objective function consists of the following parts:

[0197] S821. Tracking error is calculated by determining the Euclidean distance between the UAV's current position and the target position, ensuring the UAV follows the target path and aiming to minimize the acceptable distance difference between the UAV and the target. The tracking error calculation formula is as follows:

[0198]

[0199] Among them, X k ω1 is the position of the UAV at the k-th time step, and ω1 is the weight factor.

[0200] S822, Field of View Correction, ensures the target remains within the UAV camera's field of view at all times. The formula for calculating the field of view correction is as follows:

[0201]

[0202] Where ω2 is the weighting factor, and ellipse is a function of the field of view and the UAV's position; this function can be calculated based on the UAV's position, orientation, and field of view.

[0203] S823. In 3D space, the field of view is considered as an ellipsoid, and the target should be located inside the ellipsoid. The formula for the elliptic function is as follows:

[0204]

[0205] Where (x,y,z) is the position of the drone, (x... u ,y u ,z u ) represents the location of the target point.

[0206] Therefore, the present invention employs the above-mentioned high-precision relative positioning method for UAVs based on binocular depth point clouds, enabling UAVs to quickly and accurately identify targets and achieve precise attitude control during aerial relative positioning.

[0207] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A high-precision relative positioning method for a UAV based on binocular depth point cloud, characterized in that, The method comprises the following steps: S1, capturing target image using binocular camera; S2, generating depth point cloud in the entire field of view in real time based on the image data captured by the binocular camera in step S1, using a binocular stereo vision algorithm; S3, optimizing the backbone network of the YOLOv7 detection model; S4, using the optimized YOLOv7 model to accurately detect the image obtained from the binocular camera, and generating the corresponding 2D detection frame; S5, locating the three-dimensional space position of the cone target in the depth point cloud according to the 2D detection frame, generating the corresponding 3D point cloud detection area, optimizing the calculation of the point cloud dense area, and applying mean filtering to improve the quality and accuracy of the point cloud data, the specific steps are as follows: S51, locating the three-dimensional space position of the cone target in the depth point cloud according to the 2D detection frame, and generating the corresponding 3D point cloud detection area, eliminating objects far away from the camera or too close, as follows: ; wherein, a point cloud depth value representing a detection region; S52, optimizing the calculation of the point cloud dense area, the formula is as follows: ; wherein is an indicator function that generates a binary output for a given condition, returning 1 if the condition is true and 0 otherwise; S53, mean filtering of the effective point cloud, the formula is as follows: ; S6, continuously tracking the determined target point, and storing the position information of the target point in the continuous image; Calculating the centroid position of the dense point cloud as the target point, the formula is as follows: ; ; ; wherein, is a relative target point in 3D space; S7, calculating the average relative speed of the target point by calculating the position information, as follows: ; wherein, is the position coordinate of the target point in the i-th image; S8, accurately adjusting the attitude of the unmanned aerial vehicle based on the model predictive control algorithm, and controlling the unmanned aerial vehicle; In constructing the state and control input model of the UAV, based on the model predictive control algorithm, the state scalar of the UAV includes position , yaw angle , pitch angle ; The control scalars include: velocity , yaw rate , and roll rate ; the dynamics model of the UAV can be represented as: ; S81, based on the dynamics model of the unmanned aerial vehicle, for the aerial refueling scene, the objective function is as follows: S811, first, calculate the minimum distance between the unmanned aerial vehicle and the target point, the formula is as follows: ; wherein, is the position of the current target point relative to the drone, is a weight of the position error; S812, second, minimize the flight speed change rate of the unmanned aerial vehicle, the formula is as follows: ; wherein is a weight controlling the input variation; S813, finally, the objective function is as follows: ; S82, based on the dynamics model of the unmanned aerial vehicle, for the multi-unmanned aerial vehicle relative positioning scene, the objective function is as follows: S821, the tracking error is realized by calculating the Euclidean distance between the current position of the unmanned aerial vehicle and the target position, as follows: ; wherein, is the position of the drone at the kth time step, is a weight factor; S822, field of view angle correction, as follows: ; wherein, is a weight factor, is a function of the field of view angle and the drone position; this function can be calculated depending on the position, orientation and field of view angle of the drone; S823, in 3D space, the field of view range is regarded as an ellipsoid, the target is located inside the ellipsoid, and the elliptic function formula is as follows: ; wherein, is a position of the UAV, is a position of the target point.

2. The method of claim 1, wherein the method is a high-precision relative positioning method for a UAV based on binocular depth point clouds. Step S2 generates depth point cloud in the entire field of view in real time, including image correction, matching and depth calculation to obtain the three-dimensional coordinates of each point in space, as follows: S21, image correction, the specific steps are as follows: S211, use Zhang's calibration method to obtain the distortion coefficient of the camera, combine the internal parameters given by the camera manufacturer, and calibrate the relative position and attitude between the two cameras, that is, the external parameters including rotation and translation; S212, calculate the transformation matrix required for stereo correction according to the parameters obtained in step S211; this includes two rotation matrices, respectively for left and right camera image correction, and two new projection matrices; S213, remap the original image using the above rotation matrix and projection matrix, as follows: ; ; wherein, and are intrinsic matrices of left and right cameras respectively, and are rotation matrix and translation vector between two cameras respectively, , and and are intrinsic matrix and rotation matrix which need to be recalculated, and are corrected pixel positions, using the above transformation, images of two cameras can be converted into a common view plane, and the projections of the same object point in two images have the same y coordinate. S22, based on the improved cost function, by introducing an adaptive Gaussian weight matrix and normalization, the image is matched, the formula is as follows: ; wherein, is a window region for defining the size of the region considered around each pixel; are two images to be compared; are the corresponding pixel points in the images respectively, is a relative position offset in the window region; is an adaptive Gaussian weight, is the offset pixel point; S23, the depth calculation formula is as follows: ; wherein, is the depth of the pixel, is the focal length of the camera; is the baseline distance between the two cameras, i.e. the horizontal distance between the two cameras; is the disparity value, i.e. the horizontal position difference of the same scene point on the two camera images; S24, the 2D image coordinates and depth values are converted into 3D space coordinates, the formula is as follows: ; wherein, is the coordinate of the pixel in the image plane, is the coordinate of the pixel in the world coordinate system, is the principal point coordinate of the image.

3. The method of claim 1, wherein, In step S3, the main network of YOLOv7 detection model is optimized, and the specific steps are as follows: S31, set the convolution layer weight threshold, and remove the weights less than the threshold, the formula is as follows: ; wherein, is the weight matrix after pruning, is a preset threshold, is an element in the weight matrix in the i-th row and the j-th column, is an element in the weight matrix in the i-th row and the j-th column, is an element in the weight matrix in the i-th row and the j-th column. S32, use CBAM module for spatial attention weighting, and adjust the weight distribution mode of CBAM, as follows: ; wherein, is an activation function, and correspond to the attention weights of the center region of the fuel cone sleeve and the dark color block, respectively; S33, the spatial attention map apply using element-wise multiplication to the original feature map As follows: ; S34, depth separable convolution is used to more efficiently extract the features of the oiling cone sleeve, and medium scale feature map is used for detection, the formula is as follows: ; wherein is a feature map for detecting a fueling nozzle sleeve, is a mid-sized feature map in the network.

Citation Information

Patent Citations

  • SYSTEMS AND METHODS FOR TESTING A COMPONENT OR ASSEMBLIES

    DE102021128624A1

  • Three-dimensional modeling

    US20230267682A1