Unmanned aerial vehicle end scale adaptive target detection method and system
By deploying an adaptive target detection method on the drone end, and adjusting the network structure of the target detection algorithm using the drone sensor information, the problem of reducing the accuracy of the drone target detection is solved, and efficient target detection is achieved when the focal length changes.
Patent Information
- Application Number
- CN202510298824.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-04
AI Technical Summary
When the existing drone target detection algorithm changes in flight altitude, the target detection accuracy is reduced, and multi-scale strategies fail when there are large changes or small targets, and computing resources are consumed too much.
Deploy an adaptive target detection method on the drone end, calculate the world coordinates and camera coordinates of the target object through the drone sensor information, adjust the network structure of the target detection algorithm, cut the deep network layer and add shallow detection heads, and adjust the level of the detection heads according to the target scale.
In the case of changes in the focal length of the drone-mounted camera, accurately calculate the target scale, reduce the calculation amount, improve detection accuracy and save computing resources.
Smart Images

Figure CN120259627A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of aerial detection and surveillance of unmanned aerial vehicles and target detection technology in computer vision; specifically, it relates to a method and system for scale-adaptive target detection on an unmanned aerial vehicle side. Background Art
[0002] With the support of artificial intelligence technology, the field of computer vision has made great progress, and the average detection precision (mAP) of general target detection algorithms on public datasets has reached 70%+ or even higher. After the appearance of the RCNN algorithm, the accuracy of general target detection algorithms on the Pascal-VOC2007 dataset (http: / / host.robots.ox.ac.uk / pascal / VOC / ) was directly increased to 58.5%. After RCNN, general target detection algorithms have emerged continuously. In 2018, the highest precision rate on the Pascal-VOC2007 public dataset was RefineDet, reaching 83.8%. On the Pascal-VOC2012 (http: / / host.robots.ox.ac.uk / pascal / VOC / ) dataset, the highest precision rate reached 83.5%. The highest accuracy on the COCO public dataset also reached 62.9%, and this accuracy was further increased to 69.7% by TridentNet in 2019. The best target detection algorithm on the Pascal-VOC2007 public dataset is Cascade Eff-B7NAS-FPN, with an accuracy of 89.3%. As of February 2024, for the precision rates of current various target detection algorithms on the publicly available COCO test-dev dataset for target detection, Co-DETR proposed in 2023 achieved the highest correct rate of 66% box mAP on this dataset. If mAP@0.5 is used as the standard, the EVA algorithm proposed in 2023 can even reach 81.9%.
[0003] The improvement of the performance of general target detection algorithms has reached a bottleneck period. It is very difficult to increase the average precision rate by 0.1%, while the algorithm complexity increases exponentially. In the process of algorithm application, this approach of improving the average precision rate has extremely low cost performance and no practical significance, resulting in difficulties in promoting and applying the algorithm. The target detection algorithm based on a deep neural network often faces the problem that the flight altitude of the unmanned aerial vehicle changes, resulting in a change in the distance from the optoelectronic sensor to the target, and thus a scale change of the target, which further leads to a reduction in the target detection accuracy. Although most current target detection algorithms adopt a multi-scale strategy to cope with the problem of reduced target detection accuracy caused by changes in the flight altitude of the unmanned aerial vehicle within a certain range, when the altitude change is relatively large or the target size itself is relatively small, the multi-scale strategy will fail and consume computing resources. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a method for scale-adaptive object detection on the drone side, which is deployed on the drone side and introduces some sensor information of the drone itself into the object detection algorithm of the deep neural network and its implementation, including:
[0005] Determine the world coordinates of the object in the image data collected by the camera on the drone side; calculate the camera system coordinates of the object based on the world coordinates and attitude information of the camera on the drone side;
[0006] Convert the camera system coordinates of the object into 2D pixel coordinates;
[0007] Determine the network layer corresponding to the object in the object detection algorithm deployed on the drone based on the pixel coordinates of the object;
[0008] Process the network structure of the object detection algorithm based on the network layer;
[0009] Use the processed network structure to perform object detection on the drone.
[0010] Preferably, the determination of the world coordinates of the object in the image data collected by the camera on the drone side includes:
[0011] Collect image data containing the object;
[0012] Record the world coordinates of the center point of the object, the attitude angles in the world coordinates, and the external dimensions; the external dimensions include the length, width, and height of the object;
[0013] Determine the world coordinates of the eight vertices of the minimum circumscribed cuboid of the object according to the world coordinates of the center point of the object, the attitude angles in the world coordinates, and the external dimensions; the attitude angles include: yaw angle, pitch angle, and roll angle.
[0014] Preferably, the calculation of the camera system coordinates of the object based on the world coordinates and attitude information of the camera on the drone side includes:
[0015] Taking the world coordinates of the camera as the origin, rotate and translate the world coordinates of the object according to the attitude information of the camera to obtain the camera coordinates of the object.
[0016] Preferably, the conversion of the camera system coordinates of the object into a 2D pixel coordinate system includes:
[0017] Based on the perspective projection relationship, convert the camera system coordinates of the center point and each vertex of the object into 2D camera coordinates;
[0018] The two-dimensional camera coordinates of the center point and each vertex of the target object are expressed in pixels to obtain the two-dimensional pixel coordinates of the center point and each vertex of the target object.
[0019] Preferably, determining the network layer corresponding to the target object in the target detection algorithm deployed on the drone based on the pixel coordinates of the target object includes:
[0020] Based on the two-dimensional pixel coordinates of each vertex of the target object, obtain four points corresponding to the minimum and maximum values in the two coordinate axes of the two-dimensional pixel coordinates;
[0021] Based on the four points, draw lines parallel to the two coordinate axes of the two-dimensional pixel coordinates respectively to obtain a rectangle;
[0022] Taking the short side of the rectangle as the standard, determine the network layer corresponding to the target object in the target detection algorithm deployed on the drone.
[0023] Preferably, taking the short side of the rectangle as the standard to determine the network layer corresponding to the target object in the target detection algorithm deployed on the drone includes:
[0024] Define the parameter n, and let l = min{A l A r , A t A d}, where l is the short side of the rectangle; A l is the coordinate of the vertex with the minimum value in the U-axis direction among the pixel coordinates of the eight vertices, A r is the coordinate of the vertex with the maximum value in the U-axis direction among the pixel coordinates of the eight vertices, A t is the coordinate of the vertex with the minimum value in the V-axis direction among the pixel coordinates of the eight vertices, A d is the coordinate of the vertex with the maximum value in the V-axis direction among the pixel coordinates of the eight vertices;
[0025] According to 2 n <l ≤ 2 n+1 , determine the value of n;
[0026] Take the parameter n as the network layer.
[0027] Preferably, processing the network structure of the target detection algorithm based on the network layer includes:
[0028] When the network layer n is less than the number of sampling layers set by the target detection algorithm of the deep neural network, delete the detection head and the corresponding network layer corresponding to the nth and deeper sampling layers in the target detection algorithm, otherwise do not perform network structure adjustment.
[0029] Preferably, using the processed network structure for drone target detection includes:
[0030] When network structure adjustment is performed:
[0031] Upsample the feature map corresponding to the shallowest detection head in the current detection network;
[0032] Fuse the upsampled feature map with the feature map of a network layer that is shallower and has not been fused with the upsampled layer feature map, and add a new detection head to detect the fused feature map.
[0033] Based on the same inventive concept, the present invention also provides a scale adaptive target detection system for an unmanned aerial vehicle (UAV), including:
[0034] A camera coordinate determination module for determining the world coordinates of a target object in the image data collected by the camera on the UAV; calculating the camera system coordinates of the target object based on the world coordinates and attitude information of the camera on the UAV;
[0035] A pixel coordinate determination module for converting the camera system coordinates of the target object into 2D pixel coordinates;
[0036] A network layer number calculation module for determining the network layer number corresponding to the target object in the target detection algorithm deployed on the UAV based on the pixel coordinates of the target object;
[0037] A network structure processing module for processing the network structure of the target detection algorithm based on the network layer number;
[0038] A detection module for performing UAV target detection using the processed network structure.
[0039] Preferably, the network structure processing module is specifically configured to: when the network layer number n is less than the number of sampling layers set by the target detection algorithm of the deep neural network, delete the detection head and the corresponding network layer corresponding to the nth and deeper sampling layers in the target detection algorithm, otherwise do not perform network structure adjustment.
[0040] The detection module is specifically configured to: when network structure adjustment is performed: upsample the feature map corresponding to the shallowest detection head in the current detection network; fuse the upsampled feature map with the feature map of a network layer that is shallower and has not been fused with the upsampled layer feature map, and add a new detection head to detect the fused feature map.
[0041] Based on the same inventive concept, the present invention also provides an electronic device, including: at least one processor and a memory; the memory and the processor are connected by a bus;
[0042] The memory for storing one or more programs;
[0043] When the one or more programs are executed by the at least one processor, a method for scale adaptive target detection on a drone provided by the present invention is implemented.
[0044] Based on the same inventive concept, the present invention also provides a readable storage medium with an executable program stored thereon. When the executable program is executed, a method for scale adaptive target detection on a drone provided by the present invention is implemented.
[0045] Compared with the closest prior art, the present invention has the following beneficial effects:
[0046] 1. The present invention provides a method and system for scale adaptive target detection on a drone, including: converting the world coordinates of a target object in the image data collected by a drone-mounted camera into the camera system coordinates of the target object; converting the camera system coordinates of the target object into 2D pixel coordinates; and performing pruning processing on the network structure in the target detection algorithm deployed on the drone based on the pixel coordinates of the target object. The present invention can accurately calculate the scale information of the target in the image by using the imaging parameters, pose information of the drone camera, pose information of the drone, and the true size information of the target in the case of the change of the focal length of the drone-mounted camera.
[0047] 2. The present invention uses the scale information of the target in the image, deletes the over-sampled network layers and detection heads according to the sampling times of the deep neural network, and only retains the detection heads on the target scale, so that the limited computing power on the drone is concentrated on the detection heads consistent with the target scale, and the calculation amount is reduced to a certain extent while maintaining the detection accuracy.
[0048] 3. While the present invention deletes the over-sampled network layers and detection heads by using the scale information of the target, it adds detection heads for the target to the shallow network layers, enhancing the detection ability for the target. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is the flowchart of a method for scale adaptive target detection on a drone according to the present invention;
[0050] Figure 2 The overall flowchart of Embodiment 1 of the present invention;
[0051] Figure 3 is the conversion diagram of world coordinates to camera coordinates according to the present invention;
[0052] Figure 4 is the conversion diagram of camera coordinates to image coordinates according to the present invention;
[0053] Figure 5 is the conversion diagram of image coordinates to pixel coordinates according to the present invention;
[0054] Figure 6It is a schematic diagram of the target scale in the image of the present invention;
[0055] Figure 7 It is the network structure diagram before modification in the embodiment of the present invention;
[0056] Figure 8 It is the network structure diagram after modification in Embodiment 1 of the present invention;
[0057] Figure 9 It is the network structure diagram after modification in Embodiment 2 of the present invention;
[0058] Figure 10 It is the structure diagram of a scale adaptive target detection system at the UAV end of the present invention;
[0059] Figure 11 It is the structure diagram of an electronic device of the present invention.
[0060] In the figure, Input is the input layer, stride is the step size, Block is the block, Maxpool is the max pooling, Conv is the convolutional layer, batchnorm is the normalization layer, leaky is the leaky activation layer, linear is the linear activation layer, yolo - head is the detection head, route is the routing layer, and Upsamle is the upsampling. Detailed implementation manners
[0061] Under the condition of knowing the position and attitude information of the target object, the present invention uses the pose sensor on the UAV to obtain the position information of the UAV itself and the attitude information of the camera. Combining the actual size of the target and the internal parameters of the camera, the number of pixels occupied by the target in the image collected by the camera can be calculated, that is, the scale of the target in the image. According to the number of pixels of the target, the scale on each layer of the feature map can be calculated after the image is input into the target detection deep neural network. When the target scale is large, there is no need to crop the deep network and the detection head, but the detection head in the shallow layer can be appropriately increased according to the computing power and real - time requirements to improve the detection accuracy. When the target scale is small, the deep network can be cropped to reduce invalid calculations, and at the same time, the detection head in the shallow layer is increased to improve the detection accuracy. This method can ignore the change of the focal length of the UAV - borne camera, accurately calculate the target scale in the image, and use a single - scale target detection network to adaptively change the layer where the detection head is located according to the target scale in the image, so that it has the ability to detect multi - scale targets, not only greatly improving the target detection accuracy but also suppressing the excessive consumption of computing resources on the UAV. The technical solution of the present invention will be further described below with reference to the accompanying drawings, but the present invention is not limited to these embodiments.
[0062] Embodiment 1:
[0063] As Figure 1 shown, the present invention provides a scale adaptive target detection method at the UAV end, including:
[0064] S1. Determine the world coordinates of the target object in the image data collected by the camera on the UAV; calculate the camera system coordinates of the target object based on the world coordinates and attitude information of the camera on the UAV;
[0065] S2. Convert the camera system coordinates of the target object into 2D pixel coordinates;
[0066] S3. Determine the network layer corresponding to the target object in the target detection algorithm deployed on the UAV based on the target object pixel coordinates;
[0067] S4. Process the network structure of the target detection algorithm based on the network layer;
[0068] S5. Perform UAV target detection using the processed network structure.
[0069] The following explains a UAV-side scale adaptive target detection method provided by the present invention according to a specific calculation example, specifically as follows Figure 2 shown:
[0070] 1. Point the camera towards the target object so that the target object appears in the camera's field of view, collect the image data containing the target object, and at the same time record the world coordinates P w value as [x w y w z w T , the attitude angles (yaw angle θ t , pitch angle and roll angle ω t ) of the target object in the world coordinate system, the length l t , width w t and height h t of the target object, the world coordinates O c value of the camera as [x o y o z o T and the attitude angles (yaw angle θ, pitch angle and roll angle ω) of the camera in the world coordinate system.
[0071] 2. According to the world coordinates P w =[x w y w z w T of the center point of the target object, the length l t , width w t and height h t of the target object, and the attitude angles (yaw angle θ t , pitch angle and the roll angle ω t ) Calculate the world coordinates P of the eight vertices of the minimum circumscribed cuboid of the target object w1 = [x w1 y w1 z w1 T , P w2 = [x w2 y w2 z w2 T , P w3 = [x w3 y w3 z w3 T , P w4 = [x w4 y w4 z w4 T , p w5 = [x w5 y w5 z w5 T , P w6 = [x w6 y w6 z w6 T , P w7 = [x w7 y w7 Z w7 1 , P w8 = [x w8 y w8 z w8 T . The specific calculation method is as follows:
[0072] In the target object coordinate system, the center of the target object is the origin, and the target object coordinate system coordinates of the eight vertices are respectively
[0073]
[0074] Each point t can be transformed in the following way
[0075] First, the coordinate transformation from the world coordinate system to the target object coordinate system is as follows:
[0076]
[0077] where R t is the rotation matrix, and T t is the translation vector.
[0078] Rotate by θ around the Z axis t , rotation matrix
[0079]
[0080] Rotation about the Y-axis Rotation matrix
[0081]
[0082] Rotation about the X-axis by ω t , rotation matrix
[0083]
[0084] R t = R zt R yt R xt , ………………………………………(5)
[0085]
[0086] From equations (1), (2), (3), (4), (5), and (6), the coordinate transformation from the object coordinate system to the world coordinate system is as follows:
[0087]
[0088] Thus, the world coordinates of the eight vertices of the minimum bounding box of the object are:
[0089]
[0090] 3. According to the world coordinate P w = [x w y w Z w T of the center point of the object, the world coordinate O c = [x o y o z o T of the camera, and the attitude information (yaw angle θ, pitch angle and roll angle ω), calculate the coordinate P c = [x c y c z c T .
[0091] The transformation from the world coordinate system to the camera coordinate system is a rigid body transformation, which requires rotation and translation, as Figure 3 shown.
[0092]
[0093] where R is the rotation matrix and T is the translation vector.
[0094] Rotating by θ about the Z-axis, the rotation matrix
[0095]
[0096] Rotating about the Y-axis the rotation matrix
[0097]
[0098] Rotating by ω about the X-axis, the rotation matrix
[0099]
[0100] R = R z R y R x , ……………………………………… (20)
[0101] The translation matrix
[0102]
[0103] Combining (1)-(6) gives:
[0104]
[0105] 4. From the camera coordinate system to the image coordinate system, it belongs to a perspective projection relationship, converting from 3D to 2D. At this time, point P has been converted from the above through the world coordinate system to be expressed as p(x c , y c , Z c ) in the camera coordinate system. As Figure 4 shown.
[0106]
[0107] In the formula: f is the camera focal length.
[0108] 5. Both the pixel coordinate system and the image coordinate system are on the imaging plane, but their respective origins and measurement units are different. The origin of the image coordinate system is the intersection of the camera optical axis and the imaging plane. Usually, the midpoint of the imaging plane is called the principal point. The unit of the image coordinate system is mm, which is a physical unit, while the unit of the pixel coordinate system is pixel. We usually describe a pixel point in terms of rows and columns. So the conversion between the two is as follows: where dx and dy represent how many mm each column and each row represent respectively, that is, 1 pixel = dx mm. As Figure 5 shown.
[0109]
[0110] In the formula, u is the abscissa of the pixel coordinate system; u0 is the abscissa of the origin of the pixel coordinate system; v is the ordinate of the pixel coordinate system; v0 is the ordinate of the origin of the pixel coordinate system; y is the ordinate of the image coordinate system; x is the abscissa of the image coordinate system;
[0111] 6. Combining 2 to 5, we can get:
[0112]
[0113] Thus, the central pixel coordinates A of the target object can be obtained as A = [u v] T . In the formula, R is the rotation matrix; T is the translation matrix; refer to the definition in Equation (16).
[0114] 7. According to the eight vertex world coordinates P w1 = [x w1 y w1 z w1 T 、P w2 = [x w2 y w2 z w2 T 、P w3 = [x w3 y w3 z w3 T 、P w4 = [x w4 y w4 z w4 T 、P w5 = [x w5 y w5 z w5 T 、P w6 = [x w6 y w6 z w6 T 、P w7 = [x w7 y w7 z w7 T 、P w8 = [x w8 y w8 z w8 T Following the process from Step 3 to Step 6 and through the world coordinates P w = [x w y w zw T The same conversion process can obtain the pixel coordinates of the eight vertices of the minimum bounding cuboid of the target object as A1 = [u1 v1] T , A2 = [u2 v2] T , A3 = [u3 v3] T , A4 = [u4 v4] T , A5 = [u5 v5] T , A6 = [u6 v6] T , A7 = [u7 v7] T , A8 = [u8 v8] T .
[0115] 8. Respectively take the minimum and maximum values along the two coordinate axes from the eight pixel coordinates obtained in step 7. A l is the coordinate of the vertex with the minimum value in the U-axis direction among the eight vertex pixel coordinates, and A r is the coordinate of the vertex with the maximum value in the U-axis direction among the eight vertex pixel coordinates, and A t is the coordinate of the vertex with the minimum value in the V-axis direction among the eight vertex pixel coordinates, and A d is the coordinate of the vertex with the maximum value in the V-axis direction among the eight vertex pixel coordinates. As Figure 6 shown.
[0116] 9. Define the parameter n, and let l = min{A l A r A t A d}. According to 2 n < l ≤ 2 n+1 , the value of n can be obtained. n is just a parameter, which can determine the deepest layer number of the network, and the network layers greater than this number can be deleted.
[0117] 10. When l = 30, n = 4 can be obtained.
[0118] 11. The target detection algorithm deployed on the drone is yolov4-tiny-3l. The network structure of this target detection algorithm is as Figure 7 shown. This network has 5 sampling layers, that is, N = 5. Since n ≤ N, directly execute step 12 of the technical solution.
[0119] 12. Prune the network of the target detection algorithm, and prune the feature layers and detection heads with more than n (that is, 4 times) sampling. The network structure of the target detection network pruned according to this solution is as Figure 8 shown. The network structure to be pruned is in the blue box. This operation can reduce the amount of invalid calculations because after n times of sampling, the target has been completely submerged and no valid pixels will be retained.
[0120] 13. According to step 12 of the technical solution, perform upsampling on the network layer corresponding to the current deepest detection head, fuse it with the feature map output by block1, and add a new detection head. The specific operation is to add a routing layer after the last convolutional layer, normalization layer, and leaky activation layer module, add an upsampling layer after the module composed of the convolutional layer, normalization layer, and leaky activation layer, then input it into the routing layer together with the output of block1 for an addition operation, then pass through the module composed of the convolutional layer, normalization layer, and leaky activation layer, and finally enter the detection head (yolo-head) after passing through the module composed of the convolutional layer, normalization layer, and linear activation layer. As Figure 8 shown, the network structure newly added is in the red box.
[0121] Use the processed network structure to continuously track and detect the target object and the target objects with sizes similar to that of this target object. When the size deviation of the newly tracked target object is relatively large, adjust the network structure according to the above method.
[0122] Among them, the objects with sizes similar to that of the target object can be target objects with the same or similar models or styles;
[0123] The newly tracked target object with a relatively large size deviation from the original target object can be target objects of different types, such as cars, drones, etc.; or they can both be cars but be a sedan or a bus respectively;
[0124] Of course, in order to improve the detection accuracy, target objects of the same model can be considered as similar target objects, for example, SUV cars of the same category. And consider a mini car and an SUV car as target objects with a relatively large deviation, etc. This application does not limit whether the target objects are similar, and can be set according to the actual needs of tracking and detection.
[0125] Example 2:
[0126] Steps 1 to 9 are the same as in Example 1.
[0127] 10. When l = 12, n = 3 can be obtained.
[0128] 11. The target detection algorithm deployed on the drone is yolov4-tiny-3l, and the network structure of this target detection algorithm is as Figure 7 shown. This network has 5 sampling layers, that is, N = 5. Since n ≤ N, execute step 12 of the technical solution.
[0129] 12. Prune the network of the target detection algorithm, and prune the special adjustment layer and detection head with more than n (i.e., 3 times) samplings. The network structure of the target detection network pruned according to this solution is as Figure 9 shown. The network structure that needs to be pruned is in the blue box.
[0130] 13. According to step 12 of the technical solution, perform upsampling on the network layer corresponding to the deepest detection head, fuse it with the feature map output by block1, and add a new detection head. As Figure 9 shown, the newly added network structure is in the red box.
[0131] Embodiment 3:
[0132] In order to implement an end - scale adaptive target detection method for drones provided by the present invention, the present invention also provides an end - scale adaptive target detection system for drones. As Figure 10 shown, it includes:
[0133] A camera coordinate determination module, which is used to determine the world coordinates of the target object in the image data collected by the camera on the drone; calculate the camera - system coordinates of the target object based on the world coordinates and attitude information of the camera on the drone;
[0134] A pixel coordinate determination module, which is used to convert the camera - system coordinates of the target object into 2 - D pixel coordinates;
[0135] A network layer number calculation module, which is used to determine the network layer number corresponding to the target object in the target detection algorithm deployed on the drone based on the pixel coordinates of the target object;
[0136] A network structure processing module, which is used to process the network structure of the target detection algorithm based on the network layer number;
[0137] A detection module, which is used to perform drone target detection using the processed network structure.
[0138] Specifically, the network structure processing module is used to: when the network layer number n is less than the number of sampling layers set by the target detection algorithm of the deep neural network, delete the detection head and the corresponding network layer corresponding to the n - th and deeper sampling layers in the target detection algorithm, otherwise do not perform network structure adjustment.
[0139] Specifically, the detection module is used to: when network structure adjustment is performed: upsample the feature map corresponding to the shallowest detection head in the current detection network; fuse the upsampled feature map with the feature map of the network layer that is shallower and has not been fused with the upsampled layer feature map, and add a new detection head to detect the fused feature map.
[0140] The specific content implemented by each functional module is as in the above - mentioned embodiments and will not be elaborated here.
[0141] Embodiment 4:
[0142] As Figure 11As shown in the figure, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected through a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and this data can be called and / or modified when the instructions are executed.
[0143] The processor may be a Central Processing Unit (CPU), or it may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a method for scale adaptive target detection on a drone side in the above embodiment.
[0144] Embodiment 5
[0145] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in the electronic device, and is used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device, and of course, it can also include the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, in this storage space, one or more instructions suitable for being loaded and executed by the processor are also stored. These instructions can be one or more execution programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory, or it can also be a non-volatile memory, such as at least one disk memory. By the processor loading and executing one or more instructions stored in the storage medium, the steps of a method for scale adaptive target detection on a drone side in the above embodiment can be implemented.
Claims
1. A method for scale-adaptive target detection on an unmanned aerial vehicle, characterized in that Including: Determine the world coordinates of the target object in the image data collected by the camera on the UAV; calculate the camera system coordinates of the target object based on the world coordinates and attitude information of the camera on the UAV; Convert the camera system coordinates of the target object into 2D pixel coordinates; Determine the network layer corresponding to the target object in the target detection algorithm deployed on the UAV based on the target object pixel coordinates; Process the network structure of the target detection algorithm based on the network layer; Perform UAV target detection using the processed network structure.
2. The method according to claim 1, characterized in that, The determination of the world coordinates of the target object in the image data collected by the camera on the UAV includes: Collect image data containing the target object; Record the world coordinates of the center point of the target object, the attitude angles in the world coordinates, and the external dimensions; the external dimensions include the length, width, and height of the target object; Determine the world coordinates of the eight vertices of the minimum circumscribed cuboid of the target object according to the world coordinates of the center point of the target object, the attitude angles in the world coordinates, and the external dimensions; the attitude angles include: yaw angle, pitch angle, and roll angle.
3. The method according to claim 1, characterized in that, The calculation of the camera system coordinates of the target object based on the world coordinates and attitude information of the camera on the UAV includes: Taking the world coordinates of the camera as the origin, rotate and translate the world coordinates of the target object according to the attitude information of the camera to obtain the camera coordinates of the target object.
4. The method according to claim 2, wherein The conversion of the camera system coordinates of the target object into a 2D pixel coordinate system includes: Based on the perspective projection relationship, convert the camera system coordinates of the center point and each vertex of the target object into 2D camera coordinates; Express the 2D camera coordinates of the center point and each vertex of the target object in pixels to obtain the 2D pixel coordinates of the center point and each vertex of the target object.
5. The method according to claim 1, characterized in that The determination of the network layer corresponding to the target object in the target detection algorithm deployed on the UAV based on the target object pixel coordinates includes: Based on the 2D pixel coordinates of each vertex of the target object, obtain the 4 points corresponding to the minimum and maximum values in the two coordinate axes of the 2D pixel coordinates; Based on the 4 points, draw lines parallel to the two coordinate axes of the 2D pixel coordinates to obtain a rectangle; Taking the short side of the rectangle as the standard, determine the network layer corresponding to the target object in the target detection algorithm deployed on the UAV.
6. The method according to claim 5, wherein The determination of the network layer corresponding to the target object in the target detection algorithm deployed on the UAV taking the short side of the rectangle as the standard includes: Define the parameter n, and let l = min{A l A r , A t A d}, where l is the shorter side of the rectangle; A l is the coordinate of the vertex with the minimum value in the U-axis direction among the eight vertex pixel coordinates, A r is the coordinate of the vertex with the maximum value in the U-axis direction among the eight vertex pixel coordinates, A t is the coordinate of the vertex with the minimum value in the V-axis direction among the eight vertex pixel coordinates, A d is the coordinate of the vertex with the maximum value in the V-axis direction among the eight vertex pixel coordinates; According to 2 n <l≤2 n+1 , determine the value of n; Taking the parameter n as the network layer; Preferably, the processing of the network structure of the target detection algorithm based on the network layer includes: When the network layer n is less than the number of sampling layers set by the target detection algorithm of the deep neural network, delete the detection head and the corresponding network layer corresponding to the nth and deeper sampling layers in the target detection algorithm, otherwise do not perform network structure adjustment; Preferably, the UAV target detection using the processed network structure includes: When network structure adjustment is performed: Upsample the feature map corresponding to the shallowest detection head in the current detection network; Fuse the upsampling with the feature maps of network layers in shallower layers that have not been fused with the feature maps of the upsampling layer, and add a new detection head to detect the fused feature maps.
7. A scale-adaptive target detection system for the UAV side, characterized in that, It includes: A camera coordinate determination module for determining the world coordinates of the target object in the image data collected by the camera on the drone; Calculate the camera system coordinates of the target object based on the world coordinates and attitude information of the camera on the drone; A pixel coordinate determination module for converting the camera system coordinates of the target object into 2D pixel coordinates; A network layer number calculation module for determining the network layer number corresponding to the target object in the target detection algorithm deployed on the drone based on the target object pixel coordinates; A network structure processing module for processing the network structure of the target detection algorithm based on the network layer number; A detection module for performing drone target detection using the processed network structure.
8. The system according to claim 7, wherein The network structure processing module is specifically configured to: when the network layer number n is less than the number of sampling layers set by the target detection algorithm of the deep neural network, delete the detection head and the corresponding network layer corresponding to the nth and deeper sampling layers in the target detection algorithm, otherwise do not perform network structure adjustment; The detection module is specifically configured to: when network structure adjustment is performed: upsample the feature map corresponding to the shallowest detection head in the current detection network; fuse the upsampling with the feature maps of network layers in shallower layers that have not been fused with the feature maps of the upsampling layer, and add a new detection head to detect the fused feature maps.
9. An electronic device, characterized in that, It includes: At least one processor and a memory; The memory and the processor are connected by a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the method for scale-adaptive target detection at the drone end according to any one of claims 1 to 8 is implemented.
10. A readable storage medium, characterized in that, There is an execution program stored thereon, and when the execution program is executed, the method for scale-adaptive target detection at the drone end according to any one of claims 1 to 6 is implemented.