Dynamic optimal rotating target detection method based on multi-view images of unmanned aerial vehicle
By constructing a detection strategy for adaptive rotating target frames, the rotation angle of the target frame in the drone's perspective image is dynamically adjusted to maximize effective information coverage, which solves the problems of insufficient information coverage and weak robustness in traditional methods and achieves high-precision target detection.
Patent Information
- Application Number
- CN202510812291.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-26
AI Technical Summary
Traditional axis-aligned frames and existing rotated target frame methods have problems such as insufficient effective information coverage, weak angle prediction robustness, and low computational efficiency in drone-viewing image scenes, making it difficult to meet high-precision detection requirements, especially in the detection of hidden targets.
A dynamic optimal rotating target detection method based on multi-view UAV images is constructed. Through an adaptive rotating target box optimization algorithm, the rotation angle of the target box is dynamically adjusted to maximize the effective information coverage. The multi-view image dataset is combined for model training and detection.
It significantly improves the accuracy and robustness of target detection in drone-perspective images, adapts to target recognition and threat assessment tasks in complex environments, and improves the accuracy and applicability of detection.
Smart Images

Figure BDA0005454263430000061 
Figure BDA0005454263430000062 
Figure BDA0005454263430000063
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and drone remote sensing, and more specifically, to a method for dynamically optimally rotating target detection based on multi-view drone imagery. This method analyzes the effective information distribution of targets in drone-captured images and dynamically rotates the target frame to maximize target feature coverage. This method addresses the technical challenges of traditional axis-aligned frames, which suffer from insufficient target feature capture and low detection accuracy in multi-view drone imagery. The method is particularly suitable for the precise identification of ground and aerial targets required in drone inspection, mapping, security, and other tasks. Background Art
[0002] As efficient aerial perception platforms, drones are widely used in scenarios such as power inspections, geographic mapping, and emergency rescue. Their onboard cameras can capture images of ground or aerial targets from multiple perspectives (e.g., overhead, side, and oblique). In these scenarios, target detection requires accurately locating the target and describing its posture in images from the drone's perspective. The accuracy of the target box directly impacts the reliability of subsequent target recognition, trajectory tracking, and mission decision-making. In traditional target detection tasks, the target box is typically represented as an axis-aligned bounding box, where the edges of the box are parallel to the image coordinate axes. However, images from the drone's perspective exhibit significant "perspective distortion" and "posture diversity."
[0003] On the one hand, when a drone shoots from a non-orthogonal angle, the projection of the target in the image will undergo geometric deformation. The axis-aligned frame can only cover the projection area of the target and cannot accurately represent its real spatial orientation. A large number of pixels in the frame are background or obstructions, and the proportion of effective target information is low, resulting in a significant decrease in the accuracy of subsequent feature extraction and recognition.
[0004] On the other hand, while existing methods for rotating target frames describe the true orientation of a target by introducing a rotation angle parameter, they still have several drawbacks in drone-viewed imagery, making them difficult to meet high-precision detection requirements. First, existing methods typically predefine a fixed number of angles. However, targets can appear tilted at any angle from a drone's perspective. The fixed spacing of the predefined angles can lead to significant deviations between the frame and the target's true pose, resulting in insufficient coverage of valid information within the frame, which in turn affects the determination of target attributes. Second, some methods predict the rotation angle through models, but their input is typically target edge features or keypoint coordinates. In drone-viewed images, targets are prone to blurred edges and missing keypoints due to long-distance capture or tilted viewing angles. This makes angle prediction methods based on shallow features prone to failure, seriously affecting target recognition reliability. Finally, in existing methods, the choice of rotation angle is not directly correlated with the coverage of valid information within the frame. When drones perform covert reconnaissance missions, targets may evade detection through camouflage or posture adjustments. Even if a rotation angle exists that maximizes the proportion of valid information within the frame, existing methods fail to actively search for this angle through optimization algorithms. This results in a disconnect between the target frame's orientation description and actual physical properties, making it difficult to meet the requirements of covert target detection.
[0005] In summary, traditional axis-aligned frames and existing rotating target frame technologies have problems such as insufficient effective information coverage, weak angle prediction robustness, and low computational efficiency in drone-perspective image scenarios. There is an urgent need for an adaptive rotating target frame method driven by effective information to improve the detection accuracy and task applicability of targets in drone-perspective images, and provide reliable technical support for drone remote sensing missions. Summary of the Invention
[0006] The present invention provides a dynamic optimal rotation target detection method based on multi-view images of unmanned aerial vehicle (UAV), aiming to solve the problem of low detection accuracy of targets in UAV perspective images due to posture distortion and insufficient effective information coverage.
[0007] Specifically, the steps include:
[0008] S1, Image acquisition and preprocessing from the perspective of UAV;
[0009] S2. Construct a dynamic optimal rotating target detection method based on multi-view images of UAV;
[0010] S3. Detect drone images in low-altitude scenes based on the dynamic optimal rotation target detection method.
[0011] Furthermore, step S1 constructs a UAV perspective image dataset for model training and verification. The dataset is constructed in the following way:
[0012] S11. Construct a UAV perspective image dataset, with the flight altitude gradient set to three levels: 50m (ground resolution 2cm), 200m (resolution 8cm), and 500m (resolution 20cm). The camera tilt angle achieves full coverage of -45° to +45° in pitch direction (looking down at ground targets) and -30° to +30° in yaw direction (looking sideways at aerial targets). The acquisition period covers three typical light environments: summer noon, cloudy dusk, and night during the lunar phase cycle. The target samples include three categories: static targets (12 types of civilian vehicle wheel textures, 24 high-rise building facades), dynamic targets (5 groups of walking posture sequences), and mixed scenes (parking lot roof reflective matrix, intersection traffic trajectory). The target samples in the dataset include:
[0013] Static targets: 12 civilian vehicle wheel textures, 24 high-rise building facades.
[0014] Dynamic target: 5 groups of people walking posture sequences.
[0015] Mixed scenarios: reflective matrix of car roofs in parking lots, traffic trajectories at intersections, etc.
[0016] S12. Based on the characteristics of drone perspective images, a distortion-preserving enhancement strategy is designed, including simulating the perspective distortion of drone high-angle shooting (random scaling 0.8 to 1.2 times, affine transformation to simulate tilted perspective), adding Gaussian noise (σ = 15 to 25), and adjusting brightness / contrast (± 20%) to simulate different lighting conditions. Finally, a drone perspective target detection dataset containing more than 8,000 images is formed.
[0017] Furthermore, the dynamic optimal rotation target detection method in S2 is specifically as follows:
[0018] S21. Input the drone's perspective image set, extract the initial axis-aligned bounding box through the target detection network, divide the image within the bounding box into M×N grid cells, determine whether each cell contains the target's key features (such as contours and textures), generate a binary mask (1 indicates valid area, 0 indicates background / occlusion), and calculate the effective information value for each cell.
[0019] S22. Generate a set of candidate rotation angles within a preset angle range ([-90°, 90°]) using the center of the initial bounding box as the rotation center. An optimization algorithm is used to iteratively search for the optimal angle that maximizes effective coverage, ensuring maximum coverage of effective information within the box.
[0020] S23. Generate the final target frame using the optimal rotation angle as the rotation angle, and output its coordinates (center position, size), rotation angle θ, and effective information coverage. This frame is directly used in subsequent detection tasks, solving the problem of insufficient feature coverage of traditional axis-aligned frames in multi-view UAV scenarios, and significantly improving target detection accuracy in complex environments.
[0021] Furthermore, in step S3, the drone detection method based on the adaptive rotating target frame is added to the original test model to realize the detection of the test set.
[0022] The beneficial effects of the present invention are:
[0023] Improve the ability to capture target features: By block-wise weighted quantization of effective information within the frame and dynamically optimizing the target frame rotation angle, the coverage of key features within the frame (such as contours, textures, and logos) is significantly improved, solving the problem of feature loss caused by perspective distortion in traditional axis-aligned frames. This is especially suitable for multi-angle shooting scenarios using drones.
[0024] Multi-view geometric attitude adaptation: During the data collection phase, the multi-sensor camera onboard the drone is used to collect target images from multiple angles, including overhead, side, and oblique views, covering complex shooting scenes such as large tilt angles (>60° overhead), medium tilt angles (30°~60° side views), and small tilt angles (0°~30° oblique views). This provides real perspective distortion data for model training and enhances the model's understanding of the multi-view imaging characteristics of drones. During the target detection phase, the direction of the target frame is adjusted through dynamic rotation angles based on the center position and size of the initial detection frame. The algorithm aims to maximize the coverage of effective information within the frame, automatically compensating for geometric distortion caused by the drone's shooting perspective (such as building tilt), so that the rotated target frame always fits the target's real spatial orientation, significantly improving the detection robustness in large-angle and oblique shooting scenarios.
[0025] Wide applicability in multiple scenarios: compatible with multiple fields such as drone inspection, security monitoring, logistics distribution, etc., supports multimodal data fusion, and maintains stable detection performance in severe weather (rain, fog, snow) and low light conditions.
[0026] This invention significantly improves the accuracy and robustness of target detection in drone-viewed imagery, effectively addressing the insufficient feature coverage of traditional detection methods in scenarios with distorted viewpoints, diverse poses, and concealed targets. This technology provides high-precision, real-time core detection capabilities for applications such as drone inspections, security monitoring, logistics distribution, and low-altitude defense. It is particularly well-suited for target identification and threat assessment in complex environments, and has broad engineering application value and technological advancement prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 Flowchart
[0028] Figure 2 Schematic diagram of the dynamic optimal rotating target detection method
[0029] Figure 3 Test results comparison chart DETAILED DESCRIPTION
[0030] This paper provides a dynamic optimal rotating target detection method based on multi-view drone imagery. It aims to address the low detection accuracy of targets in drone imagery due to pose distortion and insufficient effective information coverage. By constructing an adaptive rotating target box detection strategy, this method achieves better detection results in multi-view scenarios, thereby improving target detection accuracy. Figure 1 The following is a schematic diagram of the overall process:
[0031] S1. Image acquisition and preprocessing from the perspective of drones
[0032] The purpose of step S1 is to construct a multi-scene drone-view image dataset for training and validating the adaptive rotation target box detection algorithm.
[0033] S11. The dataset is constructed by mainly acquiring aerial image data with rich scenes and targets from public datasets, including the settings of various flight altitudes and viewing angles: the flight altitude gradient is set to three levels: 50m (ground resolution 2cm), 200m (resolution 8cm), and 500m (resolution 20cm). The camera tilt angle covers a full range of viewing angles from -45° to +45° in pitch direction (looking down at ground targets) and -30° to +30° in yaw direction (looking sideways at aerial targets). The collection period covers typical lighting conditions such as summer noon, cloudy dusk, and night during the lunar phase cycle. The target samples in the dataset include:
[0034] Static targets: 12 civilian vehicle wheel textures, 24 high-rise building facades.
[0035] Dynamic target: 5 groups of people walking posture sequences.
[0036] Mixed scenarios: reflective matrix of car roofs in parking lots, traffic trajectories at intersections, etc.
[0037] S12. Aiming at the distortion characteristics of drone-view images, a set of enhancement strategies was designed, including:
[0038] Simulate the perspective distortion caused by high-angle shooting of drones (random scaling of 0.8 to 1.2 times, affine transformation to simulate tilted perspective).
[0039] Gaussian noise (σ=15-25) is added to simulate noise interference.
[0040] Image brightness and contrast were adjusted (±20%) to simulate different lighting conditions.
[0041] Finally, a training dataset containing more than 8,000 images is formed for training the target detection model.
[0042] S2. Construct a dynamic optimal rotating target detection method based on multi-view images of UAVs
[0043] The core of step S2 is the construction of an adaptively rotated target frame. This method optimizes the rotation angle of the target frame through the following sub-steps to maximize the effective information coverage within the frame. Figure 2 Schematic diagram of the framework of the adaptive rotation target box detection method.
[0044] S21. Extract initial axis-aligned bounding box
[0045] Input the drone-view image set and use the object detection network to extract the corresponding image area in the initial axis-aligned bounding box. The specific formula is:
[0046] I bbox =I(x:x+w,y:y+h)
[0047] Where I is the original image, (x, y) is the coordinate of the upper left corner of the target box, w and h are the width and height of the target box respectively. Through this formula, we cut out a rectangular area of size w×h from the original image I, and this area is the image I within the target box bbox .
[0048] Each bounding box is divided into M×N grid cells.
[0049] Each grid cell is analyzed to determine whether it contains the key features of the target (such as contour, texture), and a binary mask is generated based on the result (1 represents the valid area, 0 represents the background or occluded area).
[0050] Calculate the effective information value g of each grid cell i,j , the specific formula is:
[0051]
[0052] where g i,j Indicates the valid information flag of the grid in row i and column j. 1 indicates that the grid contains valid information, and 0 indicates no valid information. By calculating the average pixel value in each grid, if the value is greater than the set threshold, the grid is considered to contain valid information (i.e., part of the target). The pixel average calculation formula is:
[0053]
[0054] Among them I k Is the value of each pixel in the grid. If the average pixel value of the grid is greater than the threshold, the grid is considered to contain valid information of the target and marked as 1; otherwise, it is marked as 0.
[0055] S22, rotation angle optimization
[0056] The center of the initial bounding box is used as the rotation center, and the preset rotation angle range is set ([-90°, 90°]). Through the iterative search of the optimization algorithm, the optimal rotation angle that maximizes the effective information coverage is found. The optimal rotation angle selection formula is:
[0057] θ best =arg maxCoverage(θ)
[0058] θ is the search range of the rotation angle, and all possible angles are traversed within the range of [-90°, 90°]. By calculating the coverage (θ) at each rotation angle, the angle that maximizes the coverage is selected as the optimal angle θ best Calculate the effective information coverage within the frame after each angle rotation. The specific formula is:
[0059]
[0060] Among them, 1 condition is an indicator function, which takes the value 1 if the condition is met (i.e. the rotated grid mask is the same as the original grid mask) and 0 otherwise. The number of grids that have the same original grid mask and the rotated grid mask is calculated, that is, the number of overlapping valid information grids. M×N is the total number of grids. The formula for calculating the rotated grid mask is:
[0061] Grid Mask rotated ={g′ i.j ∣g′ i.j ∈{0,1},1≤i,j≤max(M,N)}
[0062] Among them, g' i.j It is the valid information flag of the grid at row i and column j in the rotated image.
[0063] S23, generate the final target frame
[0064] The optimal rotation angle is used as the rotation angle to generate the final target frame. The rotation matrix formula is:
[0065]
[0066] Where θ is the rotation angle, positive values indicate clockwise rotation and negative values indicate counterclockwise rotation.
[0067] Rotation matrix R θ Is a 3x3 matrix that acts on each point in the image to perform a rotation operation. Assuming there is a point (x, y), this rotation matrix can be used to rotate the point by an angle θ to obtain the new coordinates (x', y') after rotation:
[0068]
[0069] Rotation matrix R θ By performing a linear transformation on x and y, we get the rotated point (x', y').
[0070] The final target box generation formula is:
[0071] Final BBox=Rotate(BBox,θ best )
[0072] The optimal rotation angle θ is determined by S22 best Then, use it to rotate the original target frame, Rotate(BBox,θ best ) represents the rotation angle θ around the center of the target frame best Get the final target frame
[0073] S3. Detection of drone images in low-altitude scenes based on dynamic optimal rotation target detection method
[0074] In step S3, the dynamic optimal rotation target detection method is integrated into the existing target detection model, and the method is used to detect UAV images in low-altitude scenes. Figure 3 A schematic diagram comparing detection results shows that this method can significantly improve detection accuracy and robustness. The rotation angle optimization mechanism in S2 enables the target frame to more accurately adapt to the target's posture changes, thus addressing the shortcomings of traditional axis-aligned frames in low-altitude scenes.
[0075] In summary, this paper proposes a dynamic optimal rotation target detection method based on multi-view drone imagery. This method aims to address the low detection accuracy of traditional target detection techniques in drone imagery, which is often caused by target pose distortion and insufficient effective information coverage. By constructing a detection strategy based on an adaptive rotating target frame, this method not only optimizes the rotation angle of the target frame to maximize coverage of effective information within the frame, but also adapts to the dynamic changes of targets in multi-view scenarios, thereby improving the accuracy and robustness of target detection. Specifically, by precisely optimizing the rotation of the target frame, this method effectively overcomes the shortcomings of target detection in low-altitude scenes, significantly improving detection accuracy when dealing with complex backgrounds, changing target poses, and densely distributed targets. Through careful design and targeted preprocessing of the image dataset, this method provides a richer and more diverse training foundation for target detection models, further enhancing the algorithm's adaptability in diverse environments. This detection method based on an adaptive rotating target frame has broad application prospects and can provide more accurate and reliable technical support for fields such as drone visual perception, autonomous driving, and intelligent monitoring.
Claims
1. A dynamic optimal rotating target detection method based on multi-view images of unmanned aerial vehicles, the method comprising: S1, Image acquisition and preprocessing from the perspective of UAV; S2. Construct a dynamic optimal rotating target detection method based on multi-view images of UAV; S3. Detect drone images in low-altitude scenes based on the dynamic optimal rotation target detection method.
2. The method for detecting a dynamically optimal rotating target based on drone-view images according to claim 1, wherein: In step S1, a multi-view image dataset is constructed, and different flight altitudes and tilt angles are set. The flight altitude gradient is set to three levels: 50m (ground resolution 2cm), 200m (resolution 8cm), and 500m (resolution 20cm). The camera tilt angle achieves full coverage of the pitch direction of -45° to +45° (looking down at ground targets) and the yaw direction of -30° to +30° (looking sideways at aerial targets). UAV-perspective images are collected under multiple typical lighting environments (summer noon, cloudy dusk, and night during the lunar phase cycle, etc.), including static targets, dynamic targets, and mixed scenes. In addition, a distortion-preserving enhancement strategy is designed based on the characteristics of drone-perspective images, including simulating the perspective distortion of drone-perspective shooting at large angles (random scaling of 0.8 to 1.2 times, affine transformation to simulate tilted perspective), adding Gaussian noise (σ = 15 to 25), and adjusting the brightness / contrast (±20%) to simulate different lighting conditions.
3. The dynamic optimal rotating target detection method according to claim 1, wherein: The step S2 constructs a dynamic optimal rotating target detection method based on multi-view images of drones, specifically: S21. Input the drone-view image set and extract the image region corresponding to the initial axis-aligned bounding box through the object detection network. The specific formula is: I bbox =I(x:x+w,y:y+h) Where I is the original image, (x, y) is the coordinate of the upper left corner of the target box, w and h are the width and height of the target box respectively. Through this formula, we cut out a rectangular area of size w×h from the original image I, and this area is the image I within the target box bbox . Each bounding box is divided into M×N grid cells. Analyze each grid cell to determine whether it contains the key features of the target (such as contour, texture), and generate a binary mask based on the result (1 represents the valid area, 0 represents the background or occluded area). Calculate the effective information value g of each grid cell i,j , the specific formula is: where g i,j The valid information flag for the grid at row i and column j. 1 indicates that the grid contains valid information, and 0 indicates no valid information. By calculating the average pixel value within each grid, if this value is greater than a set threshold, the grid is considered to contain valid information (i.e., part of the target). S22, rotation angle optimization The center of the initial bounding box is used as the rotation center, and the preset rotation angle range is set ([-90°, 90°]). Through the iterative search of the optimization algorithm, the optimal rotation angle that maximizes the effective information coverage is found. The optimal rotation angle selection formula is: θ best =arg maxCoverage(θ) θ is the search range of the rotation angle, and all possible angles are traversed within the range of [-90°, 90°]. By calculating the coverage (θ) at each rotation angle, the angle that maximizes the coverage is selected as the optimal angle θ best Calculate the effective information coverage within the frame after each angle rotation. The specific formula is: Among them, 1 condition is an indicator function, which takes the value 1 if the condition is met (i.e. the rotated grid mask is the same as the original grid mask) and 0 otherwise. The number of grids that have the same original grid mask and the rotated grid mask is calculated, that is, the number of overlapping valid information grids. M×N is the total number of grids. The formula for calculating the rotated grid mask is: GridMask rotated ={g′ i.j ∣g′ i.j ∈{0,1},1≤i,j≤max(M,N)} Among them, g' i.j It is the valid information flag of the grid at row i and column j in the rotated image. S23. Generate the final target frame using the optimal rotation angle as the rotation angle. The formula for generating the final target frame is: Final BBox=Rotate(BBox,θ best ) The optimal rotation angle θ is determined by S22 best Then, use it to rotate the original target frame, Rotate(BBox,θ best ) represents the rotation angle θ around the center of the target frame best Get the final target box.
4. The method for detecting drone images in low-altitude scenes based on the dynamic optimal rotating target detection method according to claim 1 is characterized in that: After constructing a drone-view image dataset based on step S1, the optimal target frame rotation angle is obtained using step S2, and the target frame is rotated to maximize the effective information coverage within the frame. Integrating this method into a target detection model and performing detection on drone images in low-altitude scenes can significantly improve detection accuracy and robustness, especially when dealing with complex backgrounds, changing target poses, and densely distributed multiple targets. This method can provide a richer and more diverse training foundation for target detection models, further improving the algorithm's adaptability in different environments.