Unmanned aerial vehicle target identification and positioning method considering adaptive learnable parameters

Through the method of depth camera and adaptive anchor box generation, combined with the depth separation convolution and learnable parameter attention mechanism, the YOLOv5s model is improved, and the accuracy and accuracy of drone target recognition and positioning are solved, which is suitable for drone search and rescue and patrol tasks.

CN120339874APending Publication Date: 2025-07-18HARBIN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510386146.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The drone has low detection accuracy of ground target recognition and poor target positioning accuracy, especially in complex dynamic environments, it is difficult to meet the high-precision and high-efficiency operation needs.

Method used

The depth camera is used to obtain RGB-D image information, combine the depth separable convolution and learnable parameter attention mechanism (L-ECA), and generate an adaptive anchor box, object detection is performed through the improved YOLOv5s model, and the depth information is processed using the median filtering algorithm to improve positioning accuracy.

Benefits of technology

It significantly improves the accuracy of drone target detection and positioning accuracy, meets practical application needs, and is suitable for drone search and rescue and patrol scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339874A_ABST
    Figure CN120339874A_ABST
Patent Text Reader

Abstract

The invention aims to solve the problems of low ground target detection precision mAP and poor target positioning accuracy of an unmanned aerial vehicle, and discloses an unmanned aerial vehicle target identification and positioning method considering adaptive learnable parameters, and the method specifically comprises the steps: in a detection stage, collecting RGB-D image information through employing a depth camera, employing an adaptive anchor frame generation strategy, and carrying out the recognition and positioning of the RGB-D image information; a self-adaptive anchor frame suitable for the view angle of the unmanned aerial vehicle is generated, model parameters are optimized by adopting depth separable convolution, an L-ECA module containing learnable parameters is proposed to enhance the extraction capability of the model on image features, and image detection is performed through an image detection model; in the target position calculation stage, the depth value of the center point of the bounding box is obtained by using a median filtering algorithm, and the position information of the target is obtained through coordinate transformation in combination with the internal reference of the camera and the depth value, so that target positioning is realized. The problems of low detection precision mAP and poor target positioning accuracy in the prior art are solved, and the method is suitable for scenes such as search and rescue and inspection of the unmanned aerial vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent perception and control of unmanned systems and robots, especially in the direction of image detection in computer vision, and particularly relates to a method for UAV target recognition and positioning considering adaptive learnable parameters. Background Technique

[0002] With the wide application of UAVs in fields such as disaster relief, inspection, agricultural plant protection, and logistics transportation, the rapid recognition and high-precision positioning of ground targets have become the core technical requirements for realizing autonomous operation of UAVs. In a complex dynamic environment, UAVs need to continuously perceive the position information of ground targets in real time. However, due to problems such as environmental interference, variable target scales, and limited computing resources, traditional methods are difficult to meet the requirements of high-precision and high-efficiency operations. In recent years, with the breakthroughs in computer vision, deep learning algorithms, and sensor technology, the UAV ground target recognition and positioning technology has also been further developed.

[0003] The mainstream method in current target recognition technology is the image detection model based on deep learning. In recent years, the single-stage detection model based on deep learning has shown excellent performance in terms of real-time performance. However, there are still significant problems in UAV ground tasks. The default anchor box mechanism is not optimized for the aspect ratio distribution of targets in the UAV's top-down view, resulting in low target detection accuracy. In terms of target positioning, existing technologies mostly adopt monocular vision positioning schemes, relying on feature matching and structure from motion techniques. However, the feature points in dynamic scenes are unstable and prone to cumulative errors. At the same time, there is a lack of an efficient mapping mechanism from the image coordinate system to the global geographic coordinate system. Existing improvement schemes mostly focus on single performance improvement and do not systematically solve the problem of coordinated optimization of recognition and positioning, resulting in difficulty in achieving both detection accuracy and positioning accuracy in practical applications. Therefore, there is an urgent need to design an integrated method for target recognition and positioning with few model parameters, high detection accuracy, and good system robustness to break through the technical bottleneck in complex UAV ground tasks.

[0004] The journal paper "Improved Algorithm for UAV Small Target Detection Based on YOLOv5" aims at the problems of large computational complexity and insufficient accuracy in UAV small target detection by the traditional YOLOv5 algorithm. It optimizes on the basis of the YOLOv5s algorithm. The proposed algorithm improves the detection speed while maintaining the original accuracy. However, it still uses the fixed anchor boxes of the original YOLOv5, which suppresses the detection accuracy for images with uneven aspect ratio distributions of targets in the UAV's top-down view different from conventional datasets.

[0005] The journal paper "Research on the Defect Detection Method of Transmission Line Insulators Based on Improved YOLOv5s" aims at the problem of low accuracy in identifying defects of transmission line insulators, and proposes a defect detection method for transmission line insulators based on improved YOLOv5s. The proposed method has improved the detection accuracy of the model. However, when enhancing image features, the traditional ECA module is used, and the calculation of the one-dimensional convolution kernel size k is a fixed parameter, resulting in insufficient extraction of image features by the traditional ECA module.

[0006] The invention patent CN 117935085 A "A Small Target Detection Method and System for UAV-Captured Images Based on YOLO" improves the target detection ability of the model to a certain extent by adding a multi-level feature aggregation center module and a refined feature detection head to the YOLOv5m model. However, its model is complex, with a large amount of calculation, and the default anchor box of YOLOv5 is used without optimizing the aspect ratio of the target under the top-down view of the UAV.

[0007] The invention patent CN 115079229 B "A Method for UAV Ground Target Positioning Based on a Monocular Camera" uses the YOLOv5 neural network detection model for target detection, and calculates the target position by the collinearity condition equation and the least squares method; this method uses the YOLOv5 model for image detection in the top-down view of the UAV without improving the model. Due to the uneven distribution of the aspect ratio of the target under the top-down view of the UAV, the original YOLOv5 model has low detection accuracy for images in the top-down view of the UAV, resulting in a large target positioning error; this method uses the collinearity condition equation to calculate the position information of the target point. The collinearity condition equation assumes that the image point, the photography center, and the target center point are on the same straight line. However, in practical applications, due to factors such as camera distortion and atmospheric refraction, this assumption may not be completely valid, resulting in calculation errors and affecting the target positioning accuracy. Summary of the Invention

[0008] Aiming at the problems existing in the prior art, the present invention provides a method for UAV target recognition and positioning considering adaptive learnable parameters, aiming to solve the problems of low recognition and detection accuracy of UAV ground targets and poor target positioning accuracy. First, RGB-D image information is collected through a depth camera, and then the number of model parameters is reduced by introducing depthwise separable convolution (DSConv). A learnable parameter attention mechanism (Learnable Efficient Channel Attention, L-ECA) and an adaptive anchor box generation strategy are proposed to improve the target detection accuracy. The improved model is used for image detection and the bounding box information is returned. Finally, the depth value of the center point of the bounding box is obtained through the median filtering algorithm, and the position information of the target point is realized by combining the camera internal parameters, the depth value, and the coordinate transformation.

[0009] To achieve the above object, the present invention adopts the following specific technical solutions to solve:

[0010] A method for UAV target recognition and positioning considering adaptive learnable parameters, comprising the following steps:

[0011] S1: Obtain RGB-D image information through a depth camera carried by the UAV, and align the RGB image information and the depth image information;

[0012] S2: Use the YOLO-SDK algorithm to detect the RGB image to obtain the target point detection box information, specifically including the following sub-steps:

[0013] S2.1: Use the VisDrone2019 dataset as the dataset for model training. It contains 6,471 training images, 548 validation images, and 1,610 test images, with a total of 10 object categories. Convert the image annotations into the input format of the YOLOv5 network;

[0014] S2.2: Select the YOLOv5s model as the basic model, and improve the model on this basis. Adjust the training set image size to 640×640 as the input of the network model, and adopt data augmentation strategies such as random inversion and random cropping to expand the dataset;

[0015] S2.3: Since the anchor boxes of the traditional YOLOv5 are designed based on the conventional COCO dataset, however, the aspect ratio distribution of the targets in the top-down view of the drone is different from that of the conventional dataset, which will lead to low detection accuracy when the traditional YOLOv5 model detects images from the drone's perspective. Therefore, the present invention adopts an adaptive anchor box generation strategy, which can better adapt to images of different sizes and ratios and improve the detection accuracy. The adaptive anchor box generation strategy uses the K-means++ algorithm. First, a sample is randomly selected from the dataset as the first clustering center, then the shortest distance from each sample in the dataset to the known clustering centers is calculated, and then the probability of each sample being selected as the next clustering center is calculated, and the next clustering center is selected according to the probability distribution. The calculation method of the probability is as follows:

[0016]

[0017] In the formula, D(x i ) represents the shortest distance between the i-th sample and the current known clustering center, and n is the total number of samples; repeat the above method until K clustering centers are selected, and substitute them into the traditional K-means algorithm to obtain the final clustering centers, generate 9 anchor boxes, corresponding to the 3 detection layers of YOLOv5, with 3 anchor boxes in each layer, so as to realize the generation of adaptive anchor boxes and improve the detection accuracy of the model;

[0018] S2.4: In order to reduce the number of parameters of the model, improve the detection rate of the model and ensure that the model can be deployed on the low-computing-power drone embedded platform, DSConv is used to replace the standard convolution of the C3 module in the Backbone of the YOLOv5s network structure to reduce the number of parameters of the model, so as to facilitate the deployment on the drone embedded platform and improve the detection rate. The depthwise separable convolution algorithm decomposes the traditional convolution algorithm into pointwise convolution (merging channels with a 1×1 convolution kernel) and depthwise convolution (processing each channel independently). When the traditional convolution operation uses M convolution kernels of a specific size, and the input feature map has N channels and a size of D F When, the computational complexity of the traditional convolution is:

[0019] F conv = D F × D F × M × N × D K × D K (2)

[0020] In the formula, D F is the size of the input feature map, M is the number of convolution kernels, N is the number of channels of the input feature map, and D K is the size of the convolution kernel; the calculation steps of the depthwise separable convolution mainly include two parts: depthwise convolution and pointwise convolution. The computational complexity of the depthwise separable convolution is:

[0021] F DSConv = D K × D K × N × D F + N × M × D F × D F (3)

[0022] Dividing Equation (2) by Equation (3), the final result is:

[0023]

[0024] By using DSConv to replace the traditional convolution in the C3 module of the YOLOv5s network structure Backbone, the model parameters and computational complexity can be significantly reduced, ensuring that the UAV embedded platform can stably run the image detection model;

[0025] S2.5: Introduce L-ECA in some C3 modules of the YOLOv5s network structure Backbone to enhance important features and avoid the loss of important information. As a lightweight attention mechanism, L-ECA can learn feature channels at a relatively small computational cost. W, H, and C represent the width, height, and number of channels of the feature map respectively. When the feature map is input into L-ECA, it first undergoes a global average pooling (GAP) operation, and then the one-dimensional convolution with a convolution kernel size of k captures the correlation between channels. The output of the one-dimensional convolution is transformed into the attention weight corresponding to each feature channel through the Sigmoid activation function. Finally, the elements of each channel in the input feature map are multiplied element-wise with the corresponding attention weight to obtain the feature map output by L-ECA. To enhance the coverage of cross-channel information, the model can adaptively adjust the convolution kernel size k according to the number of channels of the input feature map. In the deep feature map, the model can generate a larger k to expand the receptive field to capture long-range dependencies. In the shallow feature map, the model generates a smaller k to avoid computational redundancy. There is a non-linear mapping relationship between the convolution kernel size k of the one-dimensional convolution in L-ECA and the number of channels C of the feature map, and its calculation method is:

[0026]

[0027] In the formula, C represents the number of channels of the feature map, |t| odd represents the odd number closest to t, γ learnable and b learnable are learnable parameters. At the beginning of training, initial values are set for γ learnable and b learnable according to experience, then the output k is calculated through backpropagation, and then the loss function L is calculated. Finally, γ learnable and b learnableThe gradient of γ is updated to obtain the updated γ learnable and b learnable The calculation method is as follows:

[0028]

[0029]

[0030] In the formula, η is the learning rate of the model, and L is the loss function of the model;

[0031] S2.6: The YOLOv5s model is improved through the above steps S2.3 to S2.5 to obtain the YOLO-SDK, and then the improved model is trained using the VisDrone2019 dataset. The mAP and the number of parameters are selected as the core evaluation indicators. Among them, the mAP indicator can comprehensively reflect the precision rate and recall rate of the model under different categories, providing a more comprehensive and in-depth evaluation perspective. While considering the detection effect of the model, the size of the model also needs to be considered, and the number of parameters is used to evaluate the size of the model; The calculation method of mAP is as follows:

[0032]

[0033] In the formula, N is the number of categories of detection targets, and AP i is the average accuracy of category i, and the calculation method of AP i is as follows:

[0034]

[0035] In the formula, P i and R i are used to represent the accuracy rate and recall rate of detecting category i respectively. The calculation methods of P and R are as follows:

[0036]

[0037]

[0038] In the formula, TP represents the number of correctly identified positive examples, FP represents the number of negative examples misidentified as positive examples, and FN represents the number of positive examples misidentified as negative examples;

[0039] S2.7: Use the trained YOLO-SDK model to detect the input RGB image and output the detection box information containing the target points;

[0040] S3: Extract the depth information within the detection box and process it to obtain the depth information of the target points within the detection box, which specifically includes the following sub-steps:

[0041] S3.1: The coordinates of the upper left corner (x1, y1) and the lower right corner (x2, y2) of the bounding box are obtained from the detection box information. The coordinates of the center point of the bounding box (u, v) are obtained from the upper left corner position information and the lower right corner position information of the bounding box. The calculation method of the center point of the bounding box is:

[0042]

[0043]

[0044] Where x1 and y1 are the coordinates of the upper left corner of the bounding box, and x2 and y2 are the coordinates of the lower right corner of the bounding box;

[0045] S3.2: After obtaining the coordinates of the center point of the bounding box, in order to reduce the depth value error of the center point and improve robustness, the median filtering algorithm is used to process the depth information of the target point, select the depth values within 5×5 pixels around the center point, remove the invalid depth values of 0, and then take the median of the remaining depth values as the final depth value d of the target center point;

[0046] S4: Combine the camera intrinsic parameters, the pixel coordinates (u, v) of the target center point, and the depth information d of the target center point to convert the two-dimensional coordinates (u, v) of the target point in the pixel coordinate system to the three-dimensional coordinates P in the camera coordinate system c =(X c ,Y c ,Z c ), the three-dimensional coordinates of the target point in the camera coordinate system are calculated as follows:

[0047]

[0048]

[0049] Z c =d (16)

[0050] Where u and v are the pixel coordinates of the center point, d is the depth value of the center point, (c x ,c y ) is the optical center of the camera, (f x ,f y ) is the focal length of the depth camera;

[0051] S5: Obtain the three-dimensional coordinates P of the target point in the camera coordinate system from step S4 c =(X c ,Y c ,Z c ), the coordinates of the target point are transformed from the camera coordinate system to the drone body coordinate system P b =(X b ,Yb , Z b ), the calculation method is:

[0052] P b = R c→b · P c + T c→b (17)

[0053] In the formula, R c→b is the rotation matrix from the camera coordinate system to the body coordinate system, and T c→b is the translation vector from the camera coordinate system to the body coordinate system;

[0054] Perform coordinate transformation again to transform the coordinates P b = (X b , Y b , Z b ) of the target point in the UAV body coordinate system to P w = (X w , Y w , Z w ) in the global coordinate system with the UAV take-off point as the coordinate origin, thereby realizing the positioning function of the UAV for the detected target point, and the calculation method of its coordinate transformation is:

[0055] P w = R b→w · P b + T b→w (18)

[0056] In the formula, R b→w is the rotation matrix from the body coordinate system to the global coordinate system, and T b→w is the translation vector from the body coordinate system to the global coordinate system.

[0057] The present invention has the following beneficial effects:

[0058] 1. Aiming at the problem that the aspect ratio distribution of targets in the UAV's perspective varies greatly, resulting in poor anchor box matching effect, the present invention uses the K-means++ algorithm to cluster the width and height of the targets, generate new anchor box sizes to adapt to the target distribution in the UAV's perspective, filter out unmatched candidate boxes through the aspect ratio threshold, reduce false detections, and improve detection accuracy;

[0059] 2. When the traditional ECA module faces complex scenes, it cannot dynamically and adaptively adjust the size k of the convolution kernel, and the model has a weak ability to extract image features, resulting in low detection accuracy of the model, and image misdetection and missed detection. The present invention proposes an L-ECA module. The L-ECA module introduces learnable parameters into the calculation of the one-dimensional convolution kernel size k on the basis of the traditional ECA module, and dynamically optimizes the learnable parameters through neural network back propagation. The improved model can dynamically adjust the size k of the convolution kernel through the learnable parameters, has stronger adaptability to complex scenes, significantly improves the detection performance of the model, and avoids the errors caused by artificial parameter adjustment to the model.

[0060] 3. The present invention uses a strategy of fusing RGB image information with depth image information to identify and locate targets. A median filter algorithm is used to extract depth information of target points. First, the depth information within 5×5 pixels around the target center point is extracted. Then, invalid depth values are removed and the median is taken to obtain the final depth information of the target point, thereby reducing depth information errors and improving robustness.

[0061] 4. By improving the detection model, the present invention has improved various detection indicators compared with the original YOLOv5s model, and the model parameters are further optimized, among which mAP@0.5 and mAP@0.5:0.95 are increased by 7.9% and 6.5% respectively, the parameter amount is reduced by 5.5%, and the efficiency is significantly improved; at the same time, in terms of target positioning accuracy, the target point positioning deviation is less than 0.15m, which meets the actual application needs and is suitable for drone search and rescue, inspection and other scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0063] Figure 1 It is a flowchart of a method for UAV target recognition and localization considering adaptive learnable parameters;

[0064] Figure 2 Depthwise Separable Convolution (DSConv) and the introduction of positional structure graph;

[0065] Figure 3 For the learnable parameter attention mechanism (L-ECA) and the introduction of position structure graph;

[0066] Figure 4This is the backbone network structure diagram of the YOLO-SDK network;

[0067] Figure 5 Flow chart of coordinate transformation for target point positioning;

[0068] Figure 6 This is a scene diagram of the UAV target recognition and positioning experiment;

[0069] Figure 7 This is the target detection result diagram of the YOLOv5 model;

[0070] Figure 8 This is a target detection result diagram of the present invention. DETAILED DESCRIPTION

[0071] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and understandable, a method for identifying and locating unmanned aerial vehicle targets considering adaptive learnable parameters is provided. Figure 1 As shown, the following steps are included:

[0072] S1: Obtain RGB-D image information through the Intel RealSense D435i depth camera carried by the drone, and align the RGB image information and depth image information through the program provided by the Intel RealSense SDK;

[0073] S2: Use the YOLO-SDK algorithm to detect the RGB image and obtain the target point detection box information, which specifically includes the following sub-steps:

[0074] S2.1: Use the VisDrone2019 dataset as the dataset for model training. It contains 6471 training images, 548 validation images, and 1610 test images. There are 10 object categories in total. Convert the image annotations to the input format of the YOLOv5 network.

[0075] S2.2: The YOLOv5s model is selected as the basic model, and the model is improved on this basis. The training set image size is adjusted to 640×640 as the input of the network model, and data enhancement strategies such as random inversion and random cropping are used to expand the data set;

[0076] S2.3: Since the anchor boxes of the traditional YOLOv5 are designed based on the conventional COCO dataset, however, the aspect ratio distribution of the targets in the top-down view of the drone is different from that of the conventional dataset, which will lead to low detection accuracy when the traditional YOLOv5 model detects images from the drone's perspective. Therefore, the present invention adopts an adaptive anchor box generation strategy, which can better adapt to images of different sizes and ratios and improve the detection accuracy. The adaptive anchor box generation strategy uses the K-means++ algorithm. First, a sample is randomly selected from the dataset as the first clustering center, and then the shortest distance from each sample in the dataset to the known clustering centers is calculated. Next, the probability of each sample being selected as the next clustering center is calculated, and the next clustering center is selected according to the probability distribution. The calculation method of the probability is as follows:

[0077]

[0078] In the formula, D(x i ) represents the shortest distance between the i-th sample and the current known clustering center, and n is the total number of samples. Repeat the above method until K clustering centers are selected and substitute them into the traditional K-means algorithm to obtain the final clustering centers, generate 9 anchor boxes, corresponding to the 3 detection layers of YOLOv5, with 3 anchor boxes in each layer, so as to achieve adaptive anchor box generation and improve the detection accuracy of the model.

[0079] S2.4: In order to reduce the number of parameters of the model, improve the detection rate of the model and ensure that the model can be deployed on a low-computing-power drone embedded platform, DSConv is used to replace the standard convolution in the C3 module of the YOLOv5s network structure Backbone to reduce the number of parameters of the model. Its network structure and the introduction position in the C3 module are as Figure 2 shown; the depthwise separable convolution algorithm decomposes the traditional convolution algorithm into pointwise convolution (merging channels with a 1×1 convolution kernel) and depthwise convolution (processing each channel independently). When the traditional convolution operation uses M convolution kernels of a specific size, and the input feature map has N channels and a size of D F , then the computational complexity of the traditional convolution is:

[0080] F conv =D F ×D F ×M×N×D K ×D K (2)

[0081] In the formula, D F is the size of the input feature map, M is the number of convolution kernels, N is the number of channels of the input feature map, and D K is the size of the convolution kernel; the calculation steps of the depthwise separable convolution mainly include two parts: depthwise convolution and pointwise convolution. The computational complexity of the depthwise separable convolution is:

[0082] F DSConv = D K × D K × N × D F + N × M × D F × D F (3)

[0083] Dividing equation (2) by equation (3), the final result is:

[0084]

[0085] By using DSConv instead of traditional convolution in the C3 module of the YOLOv5s network structure Backbone, the model parameters and computational volume can be significantly reduced, ensuring that the UAV embedded platform can stably run the image detection model;

[0086] S2.5: To avoid the loss of important features due to the reduction of model parameters after using DSConv, L-ECA is introduced in some C3 modules of the YOLOv5s network structure Backbone to enhance important features and avoid the loss of important information. As a lightweight attention mechanism, L-ECA can learn feature channels at a relatively small computational cost. It can be seen from Figure 3 that W, H, and C represent the width, height, and number of channels of the feature map respectively. When the feature map is input into L-ECA, it first undergoes a global average pooling (Global Average Pooling, GAP) operation, and then the one-dimensional convolution with a kernel size of k captures the correlation between channels. The output of the one-dimensional convolution is transformed into the attention weight corresponding to each feature channel through the Sigmoid activation function. Finally, the elements of each channel in the input feature map are multiplied element-wise with the corresponding attention weight to obtain the feature map output by L-ECA. To enhance the coverage rate of cross-channel information, the model can adaptively adjust the kernel size k according to the number of channels of the input feature map. In the deep feature map, the model can generate a larger k to expand the receptive field to capture long-range dependencies. In the shallow feature map, the model generates a smaller k to avoid computational redundancy. There is a non-linear mapping relationship between the kernel size k of the one-dimensional convolution in L-ECA and the number of channels C of the feature map, and its calculation method is:

[0087]

[0088] where C represents the number of channels of the feature map, |t| odd represents the odd number closest to t, γ learnable and b learnable are learnable parameters. At the beginning of training, for γ learnable and b learnableSet the initial value according to experience, then calculate the output k through backpropagation, calculate the loss function L, and finally update γ learnable and b learnable by calculating the gradients of the loss function L with respect to γ learnable and b learnable , and the calculation method is as follows:

[0089]

[0090]

[0091] In the formula, η is the learning rate of the model, and L is the loss function of the model;

[0092] S2.6: Improve the YOLOv5s model through the above steps S2.3 to S2.5 to obtain YOLO-SDK. The backbone part of its model network structure is as Figure 4 shown. Then use the VisDrone2019 dataset to train the improved model. When training the model, the CPU model of the hardware device used is Intel(R) Core(TM) i7-14700HX, and the GPU model is NVIDIA GeForce RTX 4070. In terms of software, the programming language used is Python 3.8, and the deep learning framework is PyTorch 1.7.1 to train the model. Some parameter settings during model training are shown in Table 1:

[0093] Table 1 Training parameter settings

[0094]

[0095] Select the mean Average Precision (mAP) and the number of parameters (Parameters, Params) as the core evaluation indicators. Among them, the mAP indicator can comprehensively reflect the precision rate and recall rate of the model under different categories, providing a more comprehensive and in-depth evaluation perspective. When considering the detection effect of the model, the size of the model also needs to be considered, and the number of its parameters is used to evaluate the size of the model; the calculation method of mAP is as follows:

[0096]

[0097] In the formula, N is the number of categories of detection targets, and AP i is the average accuracy of category i, and the calculation method of AP i is as follows:

[0098]

[0099] In the formula, P iand R i They are used to represent the precision and recall of the detection category i, respectively. The calculation methods of P and R are:

[0100]

[0101]

[0102] In the formula, TP represents the number of positive examples correctly identified, FP represents the number of negative examples mistakenly identified as positive examples, and FN represents the number of positive examples mistakenly identified as negative examples.

[0103] S2.7: Use the trained YOLO-SDK model to detect the input RGB image and output the detection box information containing the target point;

[0104] S3: extract the depth information in the detection frame and process it to obtain the depth information of the target point in the detection frame, which specifically includes the following sub-steps:

[0105] S3.1: The coordinates of the upper left corner (x1, y1) and the lower right corner (x2, y2) of the bounding box are obtained from the detection box information. The coordinates of the center point of the bounding box (u, v) are obtained from the upper left corner position information and the lower right corner position information of the bounding box. The calculation method of the center point of the bounding box is:

[0106]

[0107]

[0108] Where x1 and y1 are the coordinates of the upper left corner of the bounding box, and x2 and y2 are the coordinates of the lower right corner of the bounding box;

[0109] S3.2: After obtaining the coordinates of the center point of the bounding box, in order to reduce the depth value error of the center point and improve robustness, the median filtering algorithm is used to process the depth information of the target point, select the depth values within 5×5 pixels around the center point, remove the invalid depth values of 0, and then take the median of the remaining depth values as the final depth value d of the target center point;

[0110] S4: Combine the camera intrinsic parameters, the pixel coordinates (u, v) of the target center point, and the depth information d of the target center point to convert the two-dimensional coordinates (u, v) of the target point in the pixel coordinate system to the three-dimensional coordinates P in the camera coordinate system c =(X c ,Y c ,Z c ), the three-dimensional coordinates of the target point in the camera coordinate system are calculated as follows:

[0111]

[0112]

[0113] Z c = d(16)

[0114] where u and v are the pixel coordinates of the center point, d is the depth value of the center point, and (c x , c y ) is the optical center of the camera, and (f x , f y ) is the focal length of the depth camera;

[0115] S5: Obtain the three-dimensional coordinates P of the target point in the camera coordinate system c = (X c , Y c , Z c ), and transform the coordinates of the target point from the camera coordinate system to the UAV body coordinate system P b = (X b , Y b , Z b ), and the calculation method is:

[0116] P b = R c→b ·P c + T c→b (17)

[0117] where R c→b is the rotation matrix from the camera coordinate system to the body coordinate system, and T c→b is the translation vector from the camera coordinate system to the body coordinate system;

[0118] Perform coordinate transformation again to transform the coordinates of the target point in the UAV body coordinate system P b = (X b , Y b , Z b ) to the global coordinate system with the UAV takeoff point as the coordinate origin P w = (X w , Y w , Z w ), and the coordinate transformation process of steps S4 and S5 is as shown in Figure 5 , thus realizing the positioning function of the UAV for the detected target point, and the calculation method of its coordinate transformation is:

[0119] P w = R b→w ·P b + T b→w (18)

[0120] where R b→w is the rotation matrix from the body coordinate system to the global coordinate system, and Tb→w is the translation vector from the body coordinate system to the global coordinate system;

[0121] S6: To evaluate the performance of the improved model, the improved model and the control model were trained using the VisDrone2019 dataset, and the experimental results are shown in Table 2:

[0122] Table 2 Experimental results of each detection model on the VisDrone2019 dataset

[0123]

[0124] From the above experimental results, it can be seen that the values of mAP@0.5 and mAP@0.5:0.95 of the present invention in terms of detection accuracy are slightly lower than those of the YOLOv5m model, but the number of model parameters is much lower than that of the YOLOv5m model, which is more conducive to being deployed on the UAV embedded platform. Compared with the YOLOv5s model, all detection indexes of the present invention have been improved, and the number of model parameters has been further optimized. The mAP@0.5 and mAP@0.5:0.95 have increased by 7.9% and 6.5% respectively, and the number of parameters has decreased by 5.5%, and the detection accuracy has been significantly improved;

[0125] The method of the present invention was verified through physical experiments, and the experimental scenario is as Figure 6 shown, mainly including a quadrotor UAV, a target object with a QR code, and a motion capture system. The motion capture system consists of 8 motion capture cameras, which can calculate the position information of the moving target in the three-dimensional space in real time; an object with a QR code is used as the target to be recognized, and the coordinate (X, Y, Z) of the target point under the motion capture system is used as the benchmark to evaluate the positioning accuracy of the target point of the present invention. The UAV takes off at the coordinate origin under the motion capture system to ensure that the coordinate origin of the motion capture system is consistent with the coordinate origin of the UAV's global coordinate system. Four points at different positions under the motion capture system coordinate system are selected, and the target object with a QR code is placed at these 4 positions in turn, and the UAV is used to perform target detection and positioning on it. The target detection results of the original YOLOv5 model are as Figure 7 shown, the target detection results of the present invention are as Figure 8 shown, the target positioning results of the present invention are shown in Table 3. By calculating the distance deviation of the target point in the two coordinate systems the positioning accuracy of the target point of the present invention can be obtained:

[0126] Table 3 Experimental analysis of the positioning accuracy of the target point of the present invention

[0127]

[0128] At Figure 7Figures (a), (b), (c), and (d) in the middle are the detection results of the original YOLOv5 model at 4 different position points in the coordinate system of the motion capture system. In Figure (a) and Figure (b), the targets are correctly detected. In Figure (c), the target is misdetected. In Figure (d), the target is missed. It can be seen that the detection effect of the original YOLOv5 model in the UAV ground target recognition is not very good; Figure 8 Figures (a), (b), (c), and (d) in the middle are the detection results of the present invention at 4 different position points in the coordinate system of the motion capture system. From Figure 8 all the detection results in it, it can be known that the QR codes pasted on the target object are accurately detected, and there are no misdetected targets or missed detected targets. From Table 3, it can be known that in the 4 groups of UAV ground target recognition and positioning results, the present invention can accurately calculate the position information of the target point in the global coordinate system of the UAV. The distance deviation Δ of the coordinates of the target point in the two coordinate systems is less than 0.15m, that is, the positioning error of the UAV ground target recognition and positioning is less than 0.15m. Therefore, the method of the present invention improves the detection accuracy of the model and improves the target positioning accuracy, meeting the actual application requirements, especially in scenarios such as UAV search and rescue, inspection, etc.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting. Although the technical solutions of the present invention have been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present invention, and they should all be covered by the protection scope of the present invention.

Claims

1. A method for UAV target recognition and positioning considering adaptive learnable parameters, characterized in that It includes the following steps: S1: Obtain RGB-D image information through a depth camera carried by a drone, and align the RGB image information and the depth image information; S2: Use the YOLO-SDK algorithm to detect the RGB image to obtain target point detection box information, which specifically includes the following sub-steps: S2.1: Use the VisDrone2019 dataset as the dataset for model training. It contains 6,471 training images, 548 validation images, and 1,610 test images, with a total of 10 object categories. Convert the image annotations into the input format of the YOLOv5 network; S2.2: Select the YOLOv5s model as the base model, and improve the model on this basis. Adjust the training set image size to 640×640 as the input of the network model, and adopt data augmentation strategies such as random inversion and random cropping to expand the dataset; S2.3: Improve the YOLOv5s model using the adaptive anchor box generation strategy. The adaptive anchor box generation strategy uses the K-means++ algorithm. First, randomly select a sample from the dataset as the first clustering center, then calculate the shortest distance from each sample in the dataset to the known clustering centers, and then calculate the probability that each sample is selected as the next clustering center, and select the next clustering center according to the probability distribution. The calculation method of the probability is: where D(x i ) represents the shortest distance between the i-th sample and the currently known cluster center, and n is the total number of samples; repeat the above method until K cluster centers are selected, substitute them into the traditional K-means algorithm to obtain the final cluster centers, generate 9 anchor boxes, corresponding to 3 detection layers of YOLOv5, with 3 anchor boxes in each layer; S2.4: Replace the standard convolution in the C3 module of the YOLOv5s network structure Backbone with DSConv to reduce the number of model parameters. The DSConv algorithm decomposes the traditional convolution algorithm into pointwise convolution (merging channels with a 1×1 convolution kernel) and depthwise convolution (independently processing each channel). When the traditional convolution operation uses M convolution kernels of a specific size, and the input feature map has N channels and a size of D F , the computational complexity of the traditional convolution is: F conv = D F × D F × M × N × D K × D K (2) where D F is the size of the input feature map, M is the number of convolutional kernels, N is the number of channels of the input feature map, and D K is the size of the convolutional kernel; The calculation steps of DSConv mainly include two parts: depth convolution and pointwise convolution. The computational cost of depthwise separable convolution is: F DSConv = D K × D K × N × D F + N × M × D F × D F (3) Divide Equation (2) by Equation (3) to get the final result: S2.5: Introduce L-ECA in the C3 module of the Backbone part of the YOLOv5s network structure to enhance important features. W, H, and C represent the width, height, and number of channels of the feature map respectively. The feature map is input into L-ECA and first undergoes a global average pooling (GlobalAverage Pooling, GAP) operation, then captures the correlation between channels through a one-dimensional convolution with a convolution kernel size of k. The output of the one-dimensional convolution is transformed into the attention weight corresponding to each feature channel through the Sigmoid activation function. Finally, perform an element-wise product operation on the elements of each channel in the input feature map and the corresponding attention weight to obtain the feature map output by L-ECA; To enhance the coverage of cross-channel information, the model can adaptively adjust the convolution kernel size k according to the number of channels of the input feature map. In the deep feature map, the model can generate a larger k to expand the receptive field to capture long-range dependencies. In the shallow feature map, the model generates a smaller k to avoid computational redundancy. There is a non-linear mapping relationship between the convolution kernel size k of the one-dimensional convolution in L-ECA and the number of channels C of the feature map, and its calculation method is: Where C represents the number of channels of the feature map, |t| odd represents the odd number closest to the distance t, γ learnable and b learnable are learnable parameters. At the beginning of training, the initial values of γ learnable and b learnable are set according to experience. Then, the output k is calculated through backpropagation, and then the loss function L is calculated. Finally, the updated γ learnable and b learnable are obtained by calculating the gradients of the loss function L with respect to γ learnable and b learnable , and its calculation method is as follows: In the formula, η is the learning rate of the model, and L is the loss function of the model; S2.6: Improve the YOLOv5s model through the above steps S2.3 to S2.5 to obtain YOLO-SDK, and then use the VisDrone2019 dataset to train the improved model; select mAP and parameter quantity as the core evaluation indicators, among which the mAP indicator can comprehensively reflect the precision and recall rate of the model under different categories, providing a more comprehensive and in-depth evaluation perspective. While considering the model detection effect, it is also necessary to consider the size of the model, and use its parameter quantity to evaluate the size of the model; S2.7: Use the trained YOLO-SDK model to detect the input RGB image and output the detection box information containing the target point; S3: extract the depth information in the detection frame and process it to obtain the depth information of the target point in the detection frame, which specifically includes the following sub-steps: S3.1: The coordinates of the upper left corner (x1, y1) and the lower right corner (x2, y2) of the bounding box are obtained from the detection box information. The coordinates of the center point of the bounding box (u, v) are obtained from the upper left corner position information and the lower right corner position information of the bounding box. The calculation method of the center point of the bounding box is: Where x1 and y1 are the coordinates of the upper left corner of the bounding box, and x2 and y2 are the coordinates of the lower right corner of the bounding box; S3.2: After obtaining the coordinates of the center point of the bounding box, in order to reduce the depth value error of the center point and improve robustness, the median filtering algorithm is used to process the depth information of the target point, select the depth values within 5×5 pixels around the center point, remove the invalid depth values of 0, and then take the median of the remaining depth values as the final depth value d of the target center point; S4: Combine the camera internal parameters, the pixel coordinates (u, v) of the target center point, and the depth information d of the target center point to convert the two-dimensional coordinates (u, v) of the target point in the pixel coordinate system to the three-dimensional coordinates P in the camera coordinate system. c =(X c , Y c , Z c ), and the calculation method of the three-dimensional coordinates of the target point in the camera coordinate system is as follows: Z c = d (12) where u and v are the pixel coordinates of the center point, d is the depth value of the center point, and (c x , c y ) is the optical center of the camera, and (f x , f y ) is the focal length of the depth camera; S5: Obtain the three-dimensional coordinates P of the target point in the camera coordinate system from step S4 c =(X c , Y c , Z c ). Through coordinate transformation, transform the coordinates of the target point from the camera coordinate system to the UAV body coordinate system P b =(X b , Y b , Z b ). The calculation method is as follows: P b = R c→b · P c + T c→b (13) where R c→b is the rotation matrix from the camera coordinate system to the body coordinate system, and T c→b is the translation vector from the camera coordinate system to the body coordinate system; Perform coordinate transformation again to obtain the coordinates P of the target point in the UAV body coordinate system b =(X b , Y b , Z b ), and transform it to the global coordinate system with the UAV take-off point as the coordinate origin, P w =(X w , Y w , Z w ). Thus, the positioning function of the UAV for the detected target point is realized, and the calculation method of its coordinate transformation is as follows: P w = R b→w ·P b + T b→w (14) where, R b→w is the rotation matrix from the body coordinate system to the global coordinate system, and T b→w is the translation vector from the body coordinate system to the global coordinate system.

Citation Information

Patent Citations

  • A method for ground target positioning of UAV based on monocular camera

    CN115079229B

  • Small target detection method and system for images captured by unmanned aerial vehicle based on YOLO

    CN117935085A

Cited By

  • Low-illumination tunnel camera identification and positioning method based on depth camera

    CN120976527A

  • Unmanned aerial vehicle target detection and tracking method

    CN121170655A