Unmanned Aerial Vehicle Aerial Image Target Detection Method Based on Learnable Non-Uniform Sampling
By introducing learning non-uniform sampling technology in the object detection of aerial images of drones, using lightweight convolutional networks and target magnification loss function, the target detector network is optimized, and the problems of insufficient micro-object detection performance and high computing resource consumption in the existing technology are solved, and the efficient small-object detection effect is achieved.
Patent Information
- Application Number
- CN202411853706.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-12-16
AI Technical Summary
The existing drone aerial image object detection algorithm is relatively large in terms of calculation cost and inference speed, and it is difficult to effectively detect small-sized targets to be detected, which may lead to the problem of the target being ignored.
A drone aerial image object detection method based on learning non-uniform sampling is proposed. Through a lightweight convolutional network, a non-uniform sampling grid is constructed, and an image non-uniform downsampling and target box annotation conversion is performed, the object detection loss function and the loss function based on the target magnification are calculated, and the object detector network is optimized.
It significantly improves the detection accuracy of small targets, reduces the consumption of computing resources, solves the problem of insufficient detection performance of small targets in aerial images, and maintains efficient computing performance.
Smart Images

Figure CN119762995B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) aerial image target detection, and particularly relates to a UAV aerial image target detection method based on learnable non-uniform sampling. Background Technique
[0002] With the development of deep convolutional neural networks (CNNs) and the popularization of UAV technology, UAV aerial image target detection has played an important role in many remote sensing applications, such as environmental monitoring, urban management, precision agriculture, and disaster response systems. Due to the existence of a large number of small-sized targets to be detected in aerial images, as well as the problems of uneven and sparse target distribution, existing aerial image target detection algorithms mainly improve the performance of small target detection through the method of image block detection. However, such methods face problems such as high computational cost and slow inference speed, and may cause targets to be ignored during the process of image segmentation. Therefore, how to provide a target detection method that is both computationally efficient and can significantly improve the performance of small target detection is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0003] In order to solve the above problems, the present invention proposes a UAV aerial image target detection method based on learnable non-uniform sampling.
[0004] The technical solution of the present invention is: A UAV aerial image target detection method based on learnable non-uniform sampling includes the following steps:
[0005] S1. Input the image to be detected into a lightweight convolutional network, output the offset positions of the sampling points, and obtain a non-uniform sampling grid according to the offset positions of the sampling points;
[0006] S2. Use the non-uniform sampling grid to obtain the non-uniformly sampled image, and convert the target box annotation;
[0007] S3. Use the non-uniformly sampled image to obtain the target box prediction result, and calculate the target detection loss function using the converted target box annotation;
[0008] S4. Construct a loss function based on the target magnification ratio;
[0009] S5. Optimize the target detector network according to the target detection loss function and the loss function based on the target magnification ratio, and complete the target detection using the optimized target detector network.
[0010] In S5, non-uniform downsampling of the image to be detected and object detection on the distorted image space can be directly performed. Since the foreground objects are explicitly magnified in the distorted image space, better detection results for small objects can be obtained. At the same time, after obtaining the prediction results on the distorted image space, the prediction results can be mapped back to the original image space through this step to complete the object detection of the image to be detected.
[0011] The object detector network can adopt any convolutional neural network structure, such as the YOLO series detectors or Faster-RCNN, etc.
[0012] Further, S1 includes the following sub-steps:
[0013] S11: Input the image to be detected into the lightweight convolutional network to obtain the horizontal offset and vertical offset of each pixel coordinate;
[0014] S12: Obtain the final sampling points of each pixel coordinate according to the horizontal offset and vertical offset of each pixel coordinate;
[0015] S13: Use the set of all final sampling points as the non-uniform sampling grid.
[0016] The beneficial effect of the above further solution is that in the present invention, a lightweight convolutional neural network is used to construct the sampling grid matrix G for non-uniform sampling, and the process is fully differentiable. Therefore, this module can be optimized end-to-end with the object detection network through backpropagation.
[0017] To reduce the computational cost increased by this operation, the original image x θ fed into the offset estimation network f ori is replaced with its uniformly downsampled low-resolution image In this way, through S12 - S13, the same low-resolution sampling grid matrix G will be obtained d , and it is restored to the original resolution through bilinear interpolation to obtain the final sampling grid matrix G. By scaling the size of the input image of the convolutional network, the computational time and video memory occupancy required in its calculation process are reduced. Since the original image size of aerial images is often very large (such as 1500 * 2000), using a low-resolution image input can significantly reduce the additional computational cost of the proposed module to maintain the efficiency of the present invention.
[0018] Further, in S12, the calculation formula for the final sampling point (x ′ , y ′ ) of the pixel coordinate is:
[0019] x ′ = x + Δx;
[0020] y ′ = y + Δy;
[0021] In the formula, x represents the abscissa of the pixel coordinates in the image to be detected, y represents the ordinate of the pixel coordinates in the image to be detected, Δx represents the horizontal offset of the pixel coordinates, and Δy represents the vertical offset of the pixel coordinates.
[0022] Furthermore, S2 includes the following sub-steps:
[0023] S21. Obtain the image after non-uniform sampling according to the sampling grid matrix of the non-uniform sampling grid;
[0024] S22. Use the image to be detected and the image after non-uniform sampling to convert the target box annotation.
[0025] The beneficial effect of the above further solution is that in the present invention, for the target box annotation b required in the neural network training process, the upper left corner point c1 = (x1, y1) and the lower right corner point c2 = (x2, y2) are mapped from the original image space to the distorted image space after non-uniform sampling to obtain c ′ 1 and c ′ 2, thus completing the conversion of the target box annotation b ′ in the distorted space. The label mapping in the non-uniform sampling process is realized, so that the target box annotation in the distorted image space can be obtained, and the training process of the target detection network in the distorted image space can be realized.
[0026] Grid sampling method F grid_sample is the lookup process of the lookup table T defined by the sampling grid matrix G. The label mapping method is the reverse lookup process of the lookup table T By looking up the nearest neighbors of c1 and c2 in the space, their coordinates c in the distorted image space are obtained through the lookup table mapping ′ 1 and c ′ 2, and then the target box annotation b ′ in the distorted image space is finally obtained. Through one storage process of the lookup table, the non-uniform mapping of the image and the conversion process of the label are realized under the condition of reducing the additional storage cost.
[0027] Furthermore, in S22, using the original image space of the image to be detected, the upper left corner point and the lower right corner point of the target box annotation are mapped to the distorted image space of the image after non-uniform sampling, and the mapping expression is: In the formula, x represents the abscissa of the pixel coordinates in the image to be detected, y represents the ordinate of the pixel coordinates in the image to be detected to be monitored, Denote the warped image space, where \(u\) represents the abscissa of the pixel coordinates in the warped image space, and \(v\) represents the ordinate of the pixel coordinates in the warped image space. Denote the original image space.
[0028] Furthermore, S3 includes the following sub-steps:
[0029] S31: Input the non-uniformly sampled image into the target detector network to obtain the target box prediction results of the non-uniformly sampled image in the warped image space;
[0030] S32: Calculate the target detection loss function according to the target box prediction results and the transformed target box annotations.
[0031] The beneficial effect of the above further solution is: In the present invention, the network could have been directly supervised by the final target detection loss function for parameter update, enabling the offset estimation network to directly learn the non-uniform downsampling pattern beneficial to improving the target detection performance.
[0032] Furthermore, in S32, the target detection loss function \(L\) Detection has the following expression:
[0033]
[0034] In the formula, \(\alpha\) represents the first hyperparameter, \(x''\) represents the abscissa of the center point of the target box, \(y''\) represents the ordinate of the center point of the target box, \(w\) represents the length of the target box, \(h\) represents the width of the target box, \(i\) represents the parameters of the target box, represents the predicted target box, \(b'\) i represents the true target box, \(\beta\) represents the second hyperparameter, \(p'\) c represents the predicted probability of the true class of the target box, \(c\) represents the parameters of the predicted target box, represents the predicted probability of the true class of the predicted target box, and \(\log(\cdot)\) represents the logarithmic function.
[0035] The first term is the localization loss of the target, which is used to measure the spatial difference between the predicted target box \(b\) pred and the true target box \(b\) ′ , and \(\|\cdot\|\) represents the L1 norm. The second term is the classification loss of the target, which is used to measure whether the classes of the predicted target box \(b\) pred and the true target box \(b'\) are the same. In this term, \(p'\) c and respectively represent the true class of the target box and the predicted probability of the predicted target box for this class. \(\alpha\) and \(\beta\) are used as hyperparameters to control the weights of the localization loss and the classification loss in \(L\) Detection .
[0036] Furthermore, S4 includes the following sub-steps:
[0037] S41. Calculate the target magnification ratio according to the area sizes of the target bounding boxes in the image to be detected and the area sizes of the target bounding boxes in the non-uniformly sampled image.
[0038] S42. Construct a loss function based on the target magnification ratio.
[0039] The beneficial effect of the above further solution is that in the present invention, based on the loss function L of the supervised magnification ratio, Zoom the offset estimation network f θ can be guided to output an offset position that can focus on the foreground target area, realizing an explicit magnification of the foreground target area.
[0040] Further, in S41, the calculation formula for the target magnification ratio m is:
[0041]
[0042] where a represents the area size of the target bounding box in the image to be detected, and a ′ represents the area size of the target bounding box in the non-uniformly sampled image, and sum(·) represents the summation function.
[0043] Further, in S42, the expression of the loss function L Zoom based on the target magnification ratio is:
[0044]
[0045] where max(·) represents the maximum value function, log(·) represents the logarithmic function, α represents the first hyperparameter, β represents the second hyperparameter, γ represents the third hyperparameter, and m represents the target magnification ratio.
[0046] The beneficial effects of the present invention are:
[0047] (1) The present invention introduces a lightweight convolutional network module before the target detector to predict the offset positions of the sampling points during non-uniform sampling; according to the learned non-uniform sampling points, a target bounding box conversion method based on a lookup table is proposed to realize the conversion of the target bounding boxes in the original image space to the distorted space during training, and the restoration of the target bounding boxes in the distorted image space to the original image space during prediction, thereby supporting end-to-end network optimization;
[0048] (2) The present invention proposes a loss function based on the target magnification ratio, explicitly guiding the offset network to focus on the foreground target area, realizing a visual magnification of the target to be detected under the same canvas size;
[0049] (3) The present invention significantly improves the detection accuracy of existing object detection algorithms for small objects in UAV aerial images by introducing extremely low additional computational costs, and solves the problem of detecting small objects in aerial images with limited computational resources. Description of the Drawings
[0050] Figure 1 It is a flowchart of an object detection method for UAV aerial images based on learnable non-uniform sampling. Detailed Embodiments
[0051] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0052] As Figure 1 shown, the present invention provides an object detection method for UAV aerial images based on learnable non-uniform sampling, including the following steps:
[0053] S1. Input the image to be detected into a lightweight convolutional network, output the offset positions of the sampling points, and obtain a non-uniform sampling grid according to the offset positions of the sampling points;
[0054] S2. Use the non-uniform sampling grid to obtain the non-uniformly sampled image and convert the target box annotation;
[0055] S3. Use the non-uniformly sampled image to obtain the target box prediction result, and calculate the object detection loss function using the converted target box annotation;
[0056] S4. Construct a loss function based on the target magnification ratio;
[0057] S5. Optimize the object detector network according to the object detection loss function and the loss function based on the target magnification ratio, and complete object detection using the optimized object detector network.
[0058] In S5, non-uniform downsampling of the image to be detected and object detection in the distorted image space can be directly performed. Since the foreground objects will be explicitly magnified in the distorted image space, better detection results for tiny objects can be obtained. At the same time, after obtaining the prediction result in the distorted image space, the prediction result can be mapped back to the original image space through this step to complete object detection for the image to be detected.
[0059] The object detector network can adopt any structure based on convolutional neural networks, such as YOLO series detectors or Faster-RCNN, etc.
[0060] In the embodiments of the present invention, S1 includes the following sub-steps:
[0061] S11. Input the image to be detected into the lightweight convolutional network to obtain the horizontal offset and vertical offset of each pixel coordinate.
[0062] S12. Obtain the final sampling points of each pixel coordinate according to the horizontal offset and vertical offset of each pixel coordinate.
[0063] S13. Use the set of all final sampling points as the non-uniform sampling grid.
[0064] In the present invention, a lightweight convolutional neural network is used to construct a sampling grid matrix G for non-uniform sampling. The process is fully differentiable, so this module can be optimized end-to-end with the target detection network through backpropagation.
[0065] To reduce the computational cost increased by this operation, the original image x θ fed into the offset estimation network f ori is replaced with its uniformly downsampled low-resolution image In this way, through S12 - S13, a sampling grid matrix G of the same low resolution will be obtained d , which is restored to the original resolution through bilinear interpolation to obtain the final sampling grid matrix G. By scaling the input image size of the convolutional network, the computational time and video memory occupancy required in its calculation process are reduced. Since the original image size of aerial images is often very large (such as 1500 * 2000), using a low-resolution image input can significantly reduce the additional computational cost of the proposed module to maintain the efficiency of the present invention.
[0066] In the embodiment of the present invention, in S12, the formula for the final sampling point (x ′ , y ′ ) of the pixel coordinate is:
[0067] x ′ = x + Δx;
[0068] y ′ = y + Δy;
[0069] In the formula, x represents the abscissa of the pixel coordinate in the image to be detected, y represents the ordinate of the pixel coordinate in the image to be detected, Δx represents the horizontal offset of the pixel coordinate, and Δx represents the vertical offset of the pixel coordinate.
[0070] In the embodiment of the present invention, S2 includes the following sub-steps:
[0071] S21. Obtain the non-uniformly sampled image according to the sampling grid matrix of the non-uniform sampling grid.
[0072] S22. Use the image to be detected and the non-uniformly sampled image to convert the target box annotation.
[0073] In the present invention, for the target box annotation b required in the neural network training process, the upper left corner point c1=(x1, y1) and the lower right corner point c2=(x2, y2) are mapped from the original image space to the warped image space after non-uniform sampling to obtain c ′ 1 and c ′ 2, thus completing the conversion of the target box annotation b in the warped space. The label mapping during the non-uniform sampling process is realized, so that the target box annotation in the warped image space can be obtained, and the training process of the target detection network in the warped image space can be realized. ′ The conversion is achieved. The label mapping during the non-uniform sampling process is realized, so that the target box annotation in the warped image space can be obtained, and the training process of the target detection network in the warped image space can be realized.
[0074] Grid sampling method F grid_sample is the lookup process of the lookup table T defined by the sampling grid matrix G. The label mapping method is the reverse lookup process of the lookup table T By looking up the nearest neighbors of c1 and c2 in the space, their coordinates c in the warped image space are obtained through the lookup table mapping 1 and c ′ 1 and c ′ 2, and finally the target box annotation b in the warped image space is obtained. ′ . Through one storage process of the lookup table, the non-uniform mapping of the image and the conversion process of the label are realized under the condition of reducing the additional storage cost.
[0075] In the embodiment of the present invention, in S22, the upper left corner point and the lower right corner point of the target box annotation are mapped from the original image space of the image to be detected to the warped image space of the image after non-uniform sampling, and the mapping expression is: In the formula, x represents the abscissa of the pixel coordinate in the image to be detected, y represents the ordinate of the pixel coordinate in the image to be detected in the image to be monitored, represents the warped image space, u represents the abscissa of the pixel coordinate in the warped image space, v represents the ordinate of the pixel coordinate in the warped image space, represents the original image space.
[0076] In the embodiment of the present invention, S3 includes the following sub-steps:
[0077] S31. Input the image after non-uniform sampling into the target detector network to obtain the target box prediction result of the image after non-uniform sampling in the warped image space;
[0078] S32. Calculate the target detection loss function according to the target box prediction result and the converted target box annotation.
[0079] In the present invention, the network parameters could have been directly updated by the final object detection loss function, enabling the offset estimation network to directly learn the non-uniform downsampling pattern beneficial for improving object detection performance.
[0080] In an embodiment of the present invention, in S32, the object detection loss function L Detection has the following expression:
[0081]
[0082] In the formula, α represents the first hyperparameter, x″ represents the abscissa of the center point of the target box, y″ represents the ordinate of the center point of the target box, w represents the length of the target box, h represents the width of the target box, i represents the parameters of the target box, represents the predicted target box, b ′ i represents the ground truth target box, β represents the second hyperparameter, p ′ c represents the predicted probability of the true class of the target box, c represents the parameters of the predicted target box, represents the predicted probability of the true class of the predicted target box, and log(·) represents the logarithmic function.
[0083] The first term is the localization loss of the target, which is used to measure the spatial difference between the predicted target box b pred and the ground truth target box b′, and || represents the L1 norm. The second term is the classification loss of the target, which is used to measure whether the class of the predicted target box b pred is the same as that of the ground truth target box b′. In this term, p′ c and respectively represent the true class of the target box and the predicted probability of the predicted target box for this class. α and β are used as hyperparameters to control the weights of the localization loss and the classification loss in L Detection .
[0084] In an embodiment of the present invention, S4 includes the following sub-steps:
[0085] S41. Calculate the target magnification factor according to the area sizes of the target boxes in the image to be detected and the area sizes of the target boxes in the image after non-uniform sampling;
[0086] S42. Construct a loss function based on the target magnification factor.
[0087] The beneficial effect of the above further solution is that in the present invention, based on the loss function L Zoom that supervises the magnification factor, the offset estimation network f θ can be guided to output offset positions that focus on the foreground target area, achieving explicit magnification of the foreground target area.
[0088] In the embodiment of the present invention, in S41, the calculation formula of the target magnification m is as follows:
[0089]
[0090] wherein, a represents the area size of the target box of the image to be detected, and a ′ represents the area size of the target box of the image after non-uniform sampling, and sum(·) represents the summation function.
[0091] In the embodiment of the present invention, in S42, the loss function L Zoom based on the target magnification has the following expression:
[0092]
[0093] wherein, max(·) represents the maximum value function, log(·) represents the logarithmic function, α represents the first hyperparameter, β represents the second hyperparameter, γ represents the third hyperparameter, and m represents the target magnification.
[0094] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention according to the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A method for detecting targets in drone aerial images based on learnable non-uniform sampling, characterized in that: The following steps are involved: S1. Input the image to be detected into the lightweight convolutional network, output the offset position of the sampling point, and obtain a non-uniform sampling grid according to the offset position of the sampling point; S2, using a non-uniform sampling grid to obtain a non-uniformly sampled image, and converting the target box annotation; S3, using the non-uniformly sampled image to obtain the target box prediction result, and using the converted target box annotation to calculate the target detection loss function; S4, constructing a loss function based on the target magnification; S5. Optimizing the target detector network according to the target detection loss function and the target magnification-based loss function, and completing target detection using the optimized target detector network; The S2 comprises the following sub-steps: S21, obtaining an image after non-uniform sampling according to a sampling grid matrix of the non-uniform sampling grid; S22, converting the target frame annotation using the image to be detected and the image after non-uniform sampling; In S22, the upper left corner point and the lower right corner point of the target box annotation are mapped to the distorted image space of the image after non-uniform sampling using the original image space of the image to be detected, and the mapping expression is: ; In the formula, Represents the horizontal coordinate of the pixel coordinate in the image to be detected, Represents the ordinate of the pixel coordinates in the image to be detected, represents the distorted image space, represents the horizontal coordinate of the pixel coordinate in the distorted image space, represents the ordinate of the pixel coordinate in the distorted image space, represents the original image space; The S4 comprises the following sub-steps: S41, calculating the target magnification according to the area size of each target frame of the image to be detected and the area size of each target frame of the image after non-uniform sampling; S42, constructing a loss function based on the target magnification; In S41, the target magnification The calculation formula is: ; In the formula, Indicates the area size of the target box of the image to be detected, Represents the area size of the target box of the image after non-uniform sampling, represents the sum function; In S42, the loss function based on the target magnification is The expression is: ; In the formula, represents the maximum value function, represents the logarithmic function, represents the first hyperparameter, represents the second hyperparameter, represents the third hyperparameter, Indicates the target magnification.
2. The method for detecting targets in unmanned aerial images based on learnable non-uniform sampling according to claim 1, characterized in that: The S1 comprises the following sub-steps: S11, inputting the image to be detected into a lightweight convolutional network to obtain the horizontal offset and vertical offset of each pixel coordinate; S12, obtaining a final sampling point of each pixel coordinate according to the horizontal offset and the vertical offset of each pixel coordinate; S13. The set of all final sampling points is used as a non-uniform sampling grid.
3. The method for detecting targets in unmanned aerial images based on learnable non-uniform sampling according to claim 2, characterized in that: In S12, the final sampling point of the pixel coordinates The calculation formula is: ; ; In the formula, Represents the horizontal coordinate of the pixel coordinate in the image to be detected, Represents the ordinate of the pixel coordinates in the image to be detected, Indicates the horizontal offset of the pixel coordinates. Indicates the vertical offset in pixel coordinates.
4. The method for detecting targets in unmanned aerial images based on learnable non-uniform sampling according to claim 1, characterized in that: The S3 comprises the following sub-steps: S31, inputting the non-uniformly sampled image into the target detector network to obtain the target box prediction result of the non-uniformly sampled image in the distorted image space; S32. Calculate the target detection loss function based on the target box prediction result and the converted target box annotation.
5. The method for detecting targets in unmanned aerial images based on learnable non-uniform sampling according to claim 4, characterized in that: In S32, the target detection loss function The expression is: ; In the formula, represents the first hyperparameter, Indicates the horizontal coordinate of the center point of the target frame, Indicates the ordinate of the center point of the target box. Indicates the length of the target box, Indicates the width of the target box. The parameters representing the target box, represents the predicted target box, represents the true target box, represents the second hyperparameter, Represents the predicted probability of the true category of the target box, Represents the parameters of the predicted target box, Represents the predicted probability of the true category of the predicted target box, Represents a logarithmic function.
Citation Information
Patent Citations
Remote sensing image super-resolution reconstruction method based on multi-angle linear array CCD sensors
CN104574338A
Unmanned aerial vehicle video single-target long-term tracking method based on improved twin network
CN110443827A