Track generation method for unmanned aerial vehicle to track and aerially photograph target vehicle

By using the MS-TSOF-YOLO moving target detection neural network and the multi-scale adaptive dense optical flow algorithm, combined with wavelet transform and Kalman filtering, the inaccuracy and instability problems of target vehicle trajectory generation in complex motion scenes of drones are solved, and high-precision trajectory generation and compensation are achieved.

CN120635496APending Publication Date: 2025-09-12NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511059005.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing drone tracking systems cannot accurately capture the instantaneous position of the target vehicle in fast-moving scenarios, and the decoupling compensation algorithm cannot effectively eliminate the errors caused by motion, resulting in unstable trajectory generation and large errors.

Method used

The MS-TSOF-YOLO moving target detection neural network model is used to process video frames, combined with a bidirectional temporal optical flow enhancement module and a multi-scale adaptive dense optical flow algorithm, and accurate trajectories are generated through wavelet transform and extended Kalman filtering.

Benefits of technology

The accuracy and robustness of target vehicle detection by UAVs in complex motion scenarios are improved, high-precision trajectory generation and stable motion decoupling compensation are achieved, and the accuracy and integrity of the generated trajectory in the geodetic coordinate system are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635496A_ABST
    Figure CN120635496A_ABST
Patent Text Reader

Abstract

The invention relates to the field of digital image detection and signal processing, and particularly discloses an unmanned aerial vehicle tracking aerial target vehicle trajectory generation method, which comprises the following steps of: constructing a moving target detection neural network model, and performing stage processing on micro, medium and fast moving optical flow features on an input image by the model through a hierarchical cascade optical flow attention mechanism to obtain a moving target detection neural network model; motion processing of video frames is improved using bidirectional timing optical flow enhancement. Inputting a target vehicle video into a model to obtain a center coordinate of a target vehicle detection frame in each frame as a position coordinate of a vehicle, and connecting the position coordinates according to a time sequence to form a preliminary track; performing decoupling compensation of the motion of the unmanned aerial vehicle on the initial track through a multi-scale adaptive dense optical flow algorithm; performing coordinate transformation to obtain a roughly estimated trajectory of the trajectory after decoupling compensation in a geodetic coordinate system; and identifying an abnormal frequency through wavelet transform, removing noise points by using Lagrange interpolation, and de-noising by applying extended Kalman filtering to generate an accurate trajectory of the target vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of digital image detection technology and signal processing technology, and in particular to a method for generating a trajectory of a target vehicle for aerial photography by an unmanned aerial vehicle. Background Art

[0002] With the implementation of low-altitude economic policies, the use of drones in fields such as traffic monitoring, express delivery, and military operations is increasing. This is particularly true for tracking ground vehicles. During drone tracking, the trajectory of the target vehicle directly impacts the accuracy of tracking control. Therefore, accurately and efficiently capturing the target vehicle's trajectory from a drone's aerial perspective has become a pressing technical challenge.

[0003] Existing drone tracking systems mostly rely on static image detection models and lack adaptability to dynamic scenarios. In fast-moving scenarios, existing technologies often fail to accurately capture the instantaneous position of the target vehicle, affecting the accuracy of trajectory generation. Furthermore, existing decoupling compensation algorithms often fail to effectively eliminate motion-induced errors during drone movement, resulting in unstable data displays and, in turn, affecting the true reflection of the trajectory.

[0004] While some signal processing techniques have been proposed to filter out noise and abnormal data during trajectory post-processing, these techniques often fail to fully consider the temporal characteristics of trajectory data, resulting in significant errors in the resulting trajectory. Therefore, to accurately generate target vehicle trajectories, a novel approach integrating deep learning models with efficient signal processing techniques is urgently needed to overcome the limitations of traditional techniques and achieve high-precision, real-time tracking and trajectory generation. Summary of the Invention

[0005] 1. Technical problems to be solved:

[0006] In response to the above technical problems, the present invention provides a method for generating trajectories of target vehicles tracked by an unmanned aerial vehicle (UAV), which solves the technical problem in the prior art of the lack of systematic target vehicle identification and trajectory generation in tracking perspective aerial video streams.

[0007] 2. Technical solution:

[0008] A method for generating a trajectory of a target vehicle for aerial photography by an unmanned aerial vehicle, characterized by comprising:

[0009] Step 1: Use the UAV's visual measurement device to shoot a video of the target vehicle, process the video frame by frame, and annotate the target vehicle to generate training samples;

[0010] Step 2: Build an MS-TSOF-YOLO moving target detection neural network model. This model uses a hierarchical cascade optical flow attention mechanism to hierarchically process the optical flow features of micro-motion, medium-motion, and fast-motion in the input image, and uses a bidirectional temporal optical flow enhancement module to improve motion processing in video frames. The model is trained using training samples, using a combination of multiple losses to obtain the final weighted results, resulting in a target vehicle detection model.

[0011] Step 3: Obtain the target vehicle video captured by the UAV visual measurement device for which the trajectory needs to be generated, and process the video frame by frame to obtain frame-by-frame images of the target vehicle; use the target vehicle detection model to detect the frame-by-frame images, obtain the center coordinates of the target vehicle detection frame in each frame as the position coordinates of the vehicle in that frame, and connect the vehicle position coordinates of each frame in chronological order to form a preliminary trajectory;

[0012] Step 4: Decouple and compensate the UAV motion of the preliminary trajectory using a multi-scale adaptive dense optical flow algorithm;

[0013] Step 5: Convert the UAV coordinate system to the earth coordinate system to obtain the rough estimated trajectory of the target vehicle after decoupling compensation in the earth coordinate system;

[0014] Step 6: Use wavelet transform to identify abnormal frequencies of the rough estimated trajectory, use Lagrange interpolation to remove noise points, and apply extended Kalman filter to denoise and generate the precise trajectory of the target vehicle.

[0015] Furthermore, step one specifically includes:

[0016] S11: The visual measurement device of the drone tracks and photographs the target vehicle on the road at a preset aerial perspective to obtain an aerial video; the aerial perspective is the tilt angle of the camera on the drone;

[0017] S12: Process the aerial video frame by frame, convert each frame into an image format, and annotate the vehicles in the image to generate training samples.

[0018] Furthermore, the MS-TSOF-YOLO moving target detection neural network model constructed in step 2 specifically includes an input module, a bidirectional temporal optical flow enhancement module, a dynamic optical flow feature decoupling network, a backbone network, a hierarchical optical flow mechanism, a feature fusion module, a loss function, and a target position prediction module; the process of achieving target vehicle recognition between each module includes:

[0019] S21: The input module inputs the bidirectional optical flow data of the continuous multi-frame input video sequence (t-1, t, t+1) and its corresponding two adjacent frames of video into the bidirectional temporal optical flow enhancement module and the backbone network respectively; where t-1, t, and t+1 represent video frames at three consecutive moments respectively;

[0020] S22: The bidirectional temporal optical flow enhancement module adopts a forward and backward optical flow cross-validation mechanism to calculate the forward (t-1→t) and backward (t→t+1) bidirectional optical flows, and adaptively adjusts the reliability of the optical flow information through the confidence weight, wherein the confidence is calculated as follows:

[0021] Confidence=exp(-||Flow forward -Flow backward ||2 / σ) (1)

[0022] In the above formula, Confidence represents confidence; ||Flow forward -Flow backward ||2 represents the Euclidean distance between the forward and backward optical flows; σ represents the standard deviation parameter that controls the confidence decay rate;

[0023] S23: The dynamic optical flow feature decoupling network first decomposes the input optical flow feature Flow(x,y)=[u(x,y),v(x,y)] into the motion direction feature component Flow through polar coordinate transformation. dir and motion amplitude feature component Flow mag ;

[0024]

[0025] In the above formula, u and v are the u component and v component of the optical flow feature respectively;

[0026] The motion direction characteristic component Flow dir and motion amplitude feature component Flow mag The following formula is processed by the direction convolution layer and the amplitude convolution layer respectively;

[0027] F dir =Conv dir (Flow dir )

[0028] F mag =Conv mag (Flow mag ) (3)

[0029] In the above formula, Conv dir 、Conv mag They represent the direction feature extraction convolution layer and the amplitude feature extraction convolution layer respectively; F dir 、F mag They are the corresponding features extracted by the convolutional layer;

[0030] Then, F dir With F mag Apply bidirectional attention:

[0031] A D2M = Softmax(Q(F dir ) T K(F mag ))

[0032] A M2D = Softmax(Q(F mag ) T K(F dir ))

[0033] F′ dir = F dir + γA D2M V(F mag )

[0034] F′ mag = F mag + γA M2D V(F dir ) (4)

[0035] In the above formula, Q, K, and V are the mappings of query, key, and value in the attention respectively; γ is a learnable scaling coefficient; A D2M , A M2D respectively represent the unidirectional attention results between the direction feature and the amplitude feature; F′ dir , F′[[ID=�3]] mag respectively represent the direction feature and the amplitude feature updated by bidirectional attention; [[ID=�6]]

[0036] Concatenate the two features updated by bidirectional attention along the channel dimension to obtain the fused dynamic optical flow feature F dof , and directly output it to the hierarchical optical flow mechanism;

[0037] S24: The backbone network adopts the core architecture of YOLOv11, and extracts features with different receptive fields and spatial scales from consecutive multi-frame video sequences from the input module to the detection head;

[0038] S25: The hierarchical optical flow mechanism divides the optical flow information into three levels of micro motion, medium motion, and fast motion according to the motion amplitude feature ||Flow|| of the fused dynamic optical flow feature F dof , and feeds the hierarchical optical flow information back to the P3, P4, and P5 detection heads respectively; The three levels of micro motion, medium motion, and fast motion are stratified as follows:

[0039] Micro motion optical flow: ||Flow|| < threshold1, for P3 small target detection; [[ID=7Ⲑ]]

[0040] [[ID=7Ⲑ]]Medium motion optical flow: threshold1 ≤ ||Flow|| < threshold2, for P4 medium target detection;

[0041] Fast motion optical flow: ||Flow||≥threshold2, used for P5 large object detection;

[0042] In the above formula, threshold1 and threshold2 are the preset first and second motion amplitude thresholds respectively;

[0043] S26: The feature fusion module deeply fuses the multi-scale high-level features extracted by the backbone network with the optical flow features of the corresponding level. The fused features are input into the corresponding P3, P4, and P5 detection heads;

[0044] S27: The P3, P4, and P5 detection heads input the corresponding micro-motion, medium-motion, and fast-motion level features into the target position detection module;

[0045] S28: The target position prediction module detects the position and motion trend of the target in the next frame based on the decoupling results of the dynamic optical flow features. Position :

[0046] Next Position =Current Position +α·F dir +β·F mag (5)

[0047] In the above formula, Current Position is the current frame position and motion trend; α represents the weight coefficient of the preset direction feature, and β represents the weight coefficient of the preset amplitude feature;

[0048] S29: The detection result of the target position prediction module is input into the vehicle anchor frame, and the target tracking method is used to determine the detected vehicle ID and generate the trajectory of the subsequent target vehicle.

[0049] Furthermore, step four specifically includes:

[0050] S41: Setting the adaptive local weight distribution mechanism of the input image I(x,y), and selecting the region R using the elliptical adaptive window selection strategy as follows:

[0051]

[0052] In the above formula, a and b are the major and minor semi-axes of the ellipse respectively; θ1 is the main direction angle of the ellipse; i and j represent the row and column offsets of the ellipse relative to the current center pixel (x, y);

[0053] S42: Establish a hierarchical polynomial approximation model and adaptively select the approximation order according to the local motion complexity index MC(x,y):

[0054]

[0055] In the above formula, σ 2 is the gradient variance, is the gradient vector field, H is the historical entropy of optical flow;

[0056] After the calculation is completed, fit is performed:

[0057]

[0058] In the above formula, a n , b n is the polynomial fitting coefficient, i.e. the parameter to be determined, n is an integer from 1 to 5; T1 is the preset complexity threshold;

[0059] S43: Construct a multi-constraint fusion cost function E(u,v), integrating the triple constraints of brightness preservation, gradient preservation, and temporal consistency:

[0060]

[0061] In the above formula, (x, y) is the pixel coordinate, (u, v) is the component of the optical flow displacement field, R anis is the anisotropic regularization term; μ,τ, ρ is the weight coefficient of brightness preservation, gradient preservation, temporal consistency constraint and regularization, and the components in the above formula are specifically:

[0062] L photo =(I(x,y)-I′(x+u,y+v)) 2

[0063]

[0064] L temporal =|Flow {t-1} (x,y)-(u,v)| 2 (9)

[0065] S44: Design a 5-layer non-uniform scale pyramid structure with multi-level and multi-granularity resolution expression capabilities, and use an adaptive scale factor s n : Its pyramid is specifically as follows: L0 layer, original resolution, s0 = 1.0; L1 layer, fine resolution, s1 = 0.8; L2 layer, standard resolution, s2 = 0.5; L3 layer, coarse resolution, s3 = 0.25; L4 layer, global resolution, s4 = 0.125; each layer constructs three-channel features of original brightness, gradient amplitude and local direction consistency, and adopts a bidirectional iterative refinement mechanism, that is, global motion estimation is transmitted from top to bottom, and local motion correction is fed back from bottom to top;

[0066] S45: Establish a confidence-guided adaptive fusion mechanism and define the multi-dimensional confidence evaluation Con(x,y):

[0067] Con(x,y)=C con (x,y)·C sta (x,y)·C bou (x,y) (10)

[0068] In the above formula, C con is the neighborhood optical flow consistency confidence; C sta is the confidence level of cross-scale stability; C bou Maintain confidence for motion boundaries;

[0069] Final fusion:

[0070]

[0071] In the above formula, the subscript s represents the S-th layer of the pyramid; Flow s (x, y) refers to the optical flow obtained by minimizing energy within its own resolution at the sth layer;

[0072] S46: Output enhanced optical flow information [Flow MSADOF (x,y),Con(x,y)], specifically including: motion vector field, confidence distribution map, providing high-precision motion prior information for subsequent UAV motion decoupling compensation and coordinate system transformation;

[0073] S47: The optical flow estimation results at all scales are fused and the final optical flow field [u(x,y),v(x,y)] is output to represent the motion information of each pixel in the image, which is the trajectory after structural compensation of the UAV motion.

[0074] Furthermore, step five specifically includes:

[0075] S51: Assume that the target is in the UAV coordinate system (X u ,Y u ,Z u ), the geodetic coordinate system is (X b ,Y b ), the drone’s aerial photography angle θ, the drone’s height is h t ; Considering the influence of the lens tilt angle, first calculate the actual depth distance Z of the target real :

[0076]

[0077] Perform coordinate transformation on the target:

[0078]

[0079]

[0080] In the above formula, f x ,f y is the focal length of the camera; (X′ b ,Y′ b ) is the coordinate after rotation transformation;

[0081] S52: The coordinates after rotation transformation (X′ b ,Y′ b ) is integrated with the optical flow data [u(x,y),v(x,y)] of the target vehicle in the image obtained in step 4 as follows, while considering the effect of the lens tilt angle on the optical flow:

[0082]

[0083] In the above formula (X I , Y I ) is the integrated coordinate; p x ,p y is the pixel size; the cos(θ) term in the formula is used to correct the effect of lens tilt on the projection relationship;

[0084] S53: Generate the trajectory of the target vehicle in the earth coordinate system: obtain the trajectory point sequence {(t i ,X I (t i ),Y I (t i ))|i=1,2,...,N} is the rough estimated trajectory in the geodetic coordinate system.

[0085] Furthermore, step six includes:

[0086] S61: Add the trajectory point (t i ,X I (t i ),Y I (t i The corresponding trajectory data is decomposed using the Daubechies wavelet basis to obtain the wavelet transform results under the X and Y coordinate axes;

[0087] S62: Calculate the wavelet coefficient energy E j Distribution; when E j When it is greater than the preset threshold, the point k is judged to be an abnormal frequency and the corresponding trajectory point is a noise point;

[0088] S63: For each noise point corresponding to the abnormal frequency (t k ,X I (tk ),Y I (t k ), select n normal points as references and use the Lagrange interpolation formula to calculate the new position; the new position is used to replace the position of the noise point;

[0089] S64: Use the extended Kalman filter to further optimize the denoised trajectory data to obtain a complete and accurate trajectory.

[0090] 3.Beneficial effects:

[0091] (1) The present invention provides a method for generating trajectories for drone-tracked aerial photography target vehicles. The MS-TSOF-YOLO motion target detection neural network model is designed to identify target vehicles in videos. The model uses a hierarchical cascade optical flow attention mechanism to hierarchically process the optical flow features of micro-motion, medium motion, and fast motion. A bidirectional temporal optical flow enhancement module is used to improve the motion processing of video frames. A multi-loss training method is used to obtain the final weighted results. Adding optical flow prediction to the model successfully improves the accuracy and robustness of moving vehicle detection in video streams. Locally enhanced optical flow analysis is used to enhance motion decoupling compensation for fast-moving drones.

[0092] (2) The present invention provides a method for generating a trajectory of a target vehicle for tracking aerial photography by a drone, which decouples and compensates the drone's motion for the preliminary trajectory through a multi-scale adaptive dense optical flow algorithm (MSADOF), thereby enhancing the motion adaptive decoupling compensation of fast-moving drones.

[0093] (3) The present invention provides a method for generating a trajectory of a target vehicle for tracking aerial photography by a UAV, which generates a vehicle trajectory for the target vehicle under the tracking aerial photography perspective, and can significantly optimize the trajectory acquisition and control of the UAV target tracking task.

[0094] In summary, the present invention can detect ground-moving target vehicles from the perspective of drone-tracked aerial photography and reconstruct their trajectory in the geodetic coordinate system from the drone's overhead perspective. This solution optimizes the network structure and loss function to improve video detection of small target vehicles in complex motion scenes. Furthermore, it uses the multi-scale adaptive dense optical flow (MSADOF) algorithm to decouple and compensate for drone motion. Finally, by employing digital signal processing techniques such as wavelet transform, Lagrange interpolation, and extended Kalman filtering, it accurately reconstructs the trajectory of target vehicles in drone-tracked aerial videos in the geodetic coordinate system. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 It is the overall flow chart of the present invention;

[0096] Figure 2This is a schematic diagram of the UAV involved in the present invention collecting data from a target vehicle;

[0097] Figure 3 This is a structural diagram of the MS-TSOF-YOLO moving target detection neural network model in the present invention. DETAILED DESCRIPTION

[0098] The present invention will be described in detail below with reference to the accompanying drawings.

[0099] As attached Figure 1 A method for generating a trajectory of a target vehicle for aerial photography by an unmanned aerial vehicle, characterized by comprising:

[0100] Step 1: Use the UAV's visual measurement device to shoot a video of the target vehicle, process the video frame by frame, and annotate the target vehicle to generate training samples;

[0101] Step 2: Build an MS-TSOF-YOLO moving target detection neural network model. This model uses a hierarchical cascade optical flow attention mechanism to hierarchically process the optical flow features of micro-motion, medium-motion, and fast-motion in the input image, and uses a bidirectional temporal optical flow enhancement module to improve motion processing in video frames. The model is trained using training samples, using a combination of multiple losses to obtain the final weighted results, resulting in a target vehicle detection model.

[0102] Step 3: Obtain the target vehicle video captured by the UAV visual measurement device for which the trajectory needs to be generated, and process the video frame by frame to obtain frame-by-frame images of the target vehicle; use the target vehicle detection model to detect the frame-by-frame images, obtain the center coordinates of the target vehicle detection frame in each frame as the position coordinates of the vehicle in that frame, and connect the vehicle position coordinates of each frame in chronological order to form a preliminary trajectory;

[0103] Step 4: Decouple and compensate the UAV motion of the preliminary trajectory using the multi-scale adaptive dense optical flow algorithm MSADOF;

[0104] Step 5: Convert the UAV coordinate system to the earth coordinate system to obtain the rough estimated trajectory of the target vehicle after decoupling compensation in the earth coordinate system;

[0105] Step 6: Use wavelet transform to identify abnormal frequencies of the rough estimated trajectory, use Lagrange interpolation to remove noise points, and apply extended Kalman filter to denoise and generate the precise trajectory of the target vehicle.

[0106] As attached Figure 2 As shown in FIG. 1 , a schematic diagram of a UAV collecting video of a moving target vehicle is shown. In the figure, 201 represents a visual measurement device, 202 represents a target vehicle, θ represents the aerial perspective, and V represents the target vehicle. v Indicates the speed of the target vehicle; V u Indicates the speed of the drone.

[0107] As attached Figure 3 Figure 2 shows the architecture of the MS-TSOF-YOLO moving target detection neural network model constructed using this method. This model utilizes a dual-path feature acquisition architecture. The upper path extracts three levels of motion features using a bidirectional optical flow enhancement module (BTOSTE) and a hierarchical optical flow mechanism (HCF). The lower path uses the YoloV11 backbone network to extract multi-scale, high-level features from continuous multi-frame video sequences. These dual-path features are enhanced and decomposed into micro-motion, medium-motion, and fast-motion optical flow features, which are then fused with deep features to ultimately output detection results. The model incorporates multiple loss functions, including detection loss, optical flow consistency loss, and temporal continuity loss, to ensure comprehensive performance in both detection and trajectory generation.

[0108] The method for generating the trajectory of a target vehicle for drone tracking aerial photography in this embodiment has certain requirements for drone videos. The resolution of the drone video should at least meet the requirements of a resolution of not less than 1920×1080, a frame rate of 25-30fps, and a flight altitude of 50-120m. The shooting altitude of the drone video should ensure that the proportion of the target vehicle in the entire image in the video image sequence and the proportion of the target vehicle in the entire image in the training set are within 3% to ensure the adaptability of the MS-TSOF-YOLO model weights to the detection video.

[0109] The specific information of the target video used in this embodiment is as follows:

[0110] Video information Test video parameters Traffic conditions Medium flow Frame rate 30fps Resolution 1920px×1080px Section length 500m Duration 300s Shooting height 80m

[0111] Taking into account the detection requirements of different lighting conditions and vehicle types, we built the MS-TSOF-YOLO moving object detection model. The model was trained using manually annotated vehicle samples, including different types of vehicles such as cars, bicycles, and trucks, totaling 13,240 annotated images.

[0112] The MS-TSOF-YOLO model training parameters are as follows:

[0113] Parameter name Parameter value Input size 640px×640px Batch size 16 Learning rate 0.001 Number of training rounds 300 Optimizer Adam Weight decay 0.0005

[0114] The constructed MS-TSOF-YOLO model is used to detect targets in aerial videos. During detection, the optical flow information of three consecutive frames is first calculated, using the Lucas-Kanade optical flow algorithm:

[0115] I x (x,y,t)v x +I y (x,y,t)v y +I t (x,y,t)=0

[0116] Among them, I x , I y is the image gradient, I t is the time gradient, v x 、v y is the optical flow velocity component.

[0117] After each frame is detected, the output detection results include the pixel coordinates of the detection box (x, y), the length and width of the detection box (w, h), and the detection confidence. The output order is sorted according to the detection confidence.

[0118] In this example, target detection is performed according to the above steps. The detection results are shown in the following table:

[0119] Detection method Detection accuracy (mAP@0.5) Detection speed (fps) False detection rate (%) YOLOv11 88.7 35 8.3 MS-TSOF-YOLO 92.3 45 5.2

[0120] It can be seen that the MS-TSOF-YOLO model has significantly improved detection accuracy and speed compared to the traditional YOLO model, and the false positive rate has also been greatly reduced.

[0121] Step 2: Generate geodetic coordinate system trajectory under UAV motion

[0122] Based on the target detection results obtained in step 1, this step uses a multi-scale adaptive optical flow algorithm to perform motion decoupling compensation, achieve coordinate system transformation, and generate an optimized vehicle trajectory. This step requires good continuity of the input image sequence, with the time interval between adjacent frames not exceeding 0.1 seconds, and consistent image resolution to ensure the accuracy of the optical flow calculation and the stability of the trajectory generation.

[0123] Optical flow calculations are performed using an elliptical adaptive window selection strategy and a hierarchical polynomial approximation model. A five-layer non-uniform scale pyramid structure is constructed, integrating the triple constraints of brightness preservation, gradient preservation, and temporal consistency. Enhanced optical flow information is output through a confidence-guided adaptive fusion mechanism.

[0124] The optical flow calculation parameters are set as follows:

[0125]

[0126]

[0127] The UAV coordinate system is converted into the geodetic coordinate system, the tilt angle of the UAV is rotated, and the real-time position of the target vehicle in the geodetic coordinate system is obtained by combining the optical flow data to generate a trajectory point sequence.

[0128] The coordinate transformation parameters are set as follows:

[0129] Parameter name Parameter value Camera focal length fx 1200.5px Camera focal length fy 1201.2px Pixel size px 4.8μm Pixel size py 4.8μm Flight altitude ht 80m Lens tilt angle θ 15°

[0130] Subsequently, the Daubechies wavelet basis is used to decompose the trajectory data, the abnormal frequency is identified by the wavelet coefficient energy, the Lagrange interpolation is used to eliminate the noise points, and finally the extended Kalman filter is used to further optimize the trajectory data.

[0131] In this example, motion decoupling compensation, coordinate transformation, and trajectory optimization are performed according to the above steps. The processing results are shown in the following table:

[0132] Treatment method Trajectory accuracy (m) Noise suppression rate (%) Processing speed (fps) Traditional optical flow method ±2.3 45.2 25 Method of the present invention ±0.5 85.3 30

[0133] Comparison of the effects of each treatment stage:

[0134] Processing stage Trajectory completeness rate (%) Smoothness Coordinate accuracy (m) After optical flow calculation 82.1 0.65 ±1.2 After coordinate conversion 89.7 0.72 ±0.8 After trajectory optimization 92.7 0.89 ±0.5

[0135] As can be seen, the multi-scale adaptive optical flow algorithm significantly improves trajectory accuracy and noise suppression compared to traditional optical flow methods. Coordinate transformation achieves accurate mapping from image coordinates to geodetic coordinates. Trajectory optimization further improves trajectory quality through wavelet denoising and Kalman filtering. The final trajectory accuracy reaches ±0.5 meters and the completeness rate reaches 92.7%, meeting practical application requirements.

[0136] Although the present invention has been disclosed above in terms of preferred embodiments, they are not intended to limit the present invention. Anyone skilled in the art can make various changes or modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection defined by the claims of this application.

Claims

1. A method for generating a trajectory of a target vehicle for tracking aerial photography by an unmanned aerial vehicle, characterized by: include: Step 1: Use the UAV's visual measurement device to shoot a video of the target vehicle, process the video frame by frame, and annotate the target vehicle to generate training samples; Step 2: Build the MS-TSOF-YOLO moving target detection neural network model. This model uses a hierarchical cascade optical flow attention mechanism to hierarchically process the optical flow features of micro-motion, medium-motion, and fast-motion in the input image, and uses a bidirectional temporal optical flow enhancement module to improve motion processing in video frames. The model is trained using training samples. The training method is to use a combination of multiple losses to obtain the final weight result and obtain the target vehicle detection model. Step 3: Obtain a target vehicle video captured by the UAV visual measurement device for which a trajectory needs to be generated, and process the video frame by frame to obtain a frame-by-frame image of the target vehicle; Use the target vehicle detection model to detect each frame of the image, obtain the center coordinates of the target vehicle detection frame in each frame as the position coordinates of the vehicle in that frame, and connect the vehicle position coordinates of each frame in chronological order to form a preliminary trajectory; Step 4: Decouple and compensate the UAV motion of the preliminary trajectory using a multi-scale adaptive dense optical flow algorithm; Step 5: Convert the UAV coordinate system to the earth coordinate system to obtain the rough estimated trajectory of the target vehicle after decoupling compensation in the earth coordinate system; Step 6: Use wavelet transform to identify abnormal frequencies of the rough estimated trajectory, use Lagrange interpolation to remove noise points, and apply extended Kalman filter to denoise and generate the precise trajectory of the target vehicle.

2. The method for generating a trajectory of a target vehicle for tracking aerial photography by an unmanned aerial vehicle according to claim 1, characterized in that: Step 1 specifically includes: S11: The visual measurement device of the drone tracks and photographs the target vehicle on the road at a preset aerial perspective to obtain an aerial video; the aerial perspective is the tilt angle of the camera on the drone; S12: Process the aerial video frame by frame, convert each frame into an image format, and annotate the vehicles in the image to generate training samples.

3. The method for generating a trajectory of a target vehicle for tracking aerial photography by an unmanned aerial vehicle according to claim 1, characterized in that: The MS-TSOF-YOLO moving target detection neural network model constructed in step 2 specifically includes an input module, a bidirectional temporal optical flow enhancement module, a dynamic optical flow feature decoupling network, a backbone network, a hierarchical optical flow mechanism, a feature fusion module, a loss function, and a target position prediction module. The process of achieving target vehicle recognition between each module includes: S21: The input module inputs the bidirectional optical flow data of the continuous multi-frame input video sequence (t-1, t, t+1) and its corresponding two adjacent frames of video into the bidirectional temporal optical flow enhancement module and the backbone network respectively; where t-1, t, and t+1 represent video frames at three consecutive moments respectively; S22: The bidirectional temporal optical flow enhancement module adopts a forward and backward optical flow cross-validation mechanism to calculate the forward (t-1→t) and backward (t→t+1) bidirectional optical flows, and adaptively adjusts the reliability of the optical flow information through the confidence weight, wherein the confidence is calculated as follows: Confidence=exp(-||Flow forward -Flow backward ||2 / σ) (1) In the above formula, Confidence represents confidence; ||Flow forward -Flow backward ||2 represents the Euclidean distance between the forward and backward optical flows; σ represents the standard deviation parameter that controls the confidence decay rate; S23: The dynamic optical flow feature decoupling network first decomposes the input optical flow feature Flow(x,y)=[u(x,y),v(x,y)] into the motion direction feature component Flow through polar coordinate transformation. dir and motion amplitude feature component Flow mag ; In the above formula, u and v are the u component and v component of the optical flow feature respectively; The motion direction characteristic component Flow dir and motion amplitude feature component Flow mag The following formula is processed by the direction convolution layer and the amplitude convolution layer respectively; F dir =Conv dir (Flow dir ) F mag =Conv mag (Flow mag ) (3) In the above formula, Conv dir 、Conv mag They represent the direction feature extraction convolution layer and the amplitude feature extraction convolution layer respectively; F dir 、F mag They are the corresponding features extracted by the convolution layer; Then, F dir With F mag Apply bidirectional attention: A D2M =Softmax(Q(F dir ) T K(F mag )) A M2D =Softmax(Q(F mag ) T K(F dir )) F′ dir =F dir +γA D2M V(F mag ) F′ mag =F mag +γA M2D V(F dir ) (4) In the above formula, Q, K, and V are the mappings of query, key, and value in attention respectively; γ is a learnable scaling factor; A D2M 、A M2D Respectively represent the unidirectional attention results between direction features and amplitude features; F′ dir , F′ mag Respectively represent the direction features and amplitude features updated by bidirectional attention; The two features updated by bidirectional attention are spliced ​​according to the channel dimension to obtain the fused dynamic optical flow feature F dof , directly output to the layered optical flow mechanism; S24: The backbone network adopts the core architecture of YOLOv11, extracting features of different receptive fields and spatial scales from the continuous multi-frame video sequence of the input module to the detection head; S25: Hierarchical optical flow mechanism based on fusion dynamic optical flow features F dof The motion amplitude feature ||Flow|| divides the optical flow information into three levels: micro-motion, medium-motion, and fast-motion. The classified optical flow information is fed back to the P3, P4, and P5 detection heads respectively. The micro-motion, medium-motion, and fast-motion levels are divided into the following categories: Micro-motion optical flow: ||Flow|| < threshold1, for P3 small target detection; Medium-motion optical flow: threshold1 ≤ ||Flow|| < threshold2, for P4 medium target detection; Fast-motion optical flow: ||Flow|| ≥ threshold2, for P5 large target detection; In the above formula, threshold1 and threshold2 are respectively the preset first and second motion amplitude thresholds; S26: The feature fusion module deeply fuses the multi-scale high-level features extracted by the backbone network with the optical flow features of the corresponding levels, and the fused features are input into the corresponding P3, P4, P5 detection heads; S27: The P3, P4, P5 detection heads input the features of the corresponding micro-motion, medium-motion, and fast-motion levels into the target position detection module; S28: The target position prediction module detects the position and motion trend of the target in the next frame based on the decoupling results of the dynamic optical flow features. Position : Next Position =Current Position +α·F dir +β·F mag (5) In the above formula, Current Position is the current frame position and motion trend; α represents the weight coefficient of the preset direction feature, and β represents the weight coefficient of the preset amplitude feature; S29: The detection results of the target position prediction module are input into the vehicle anchor box, and the target tracking method is used to determine the detection vehicle ID and generate the trajectory of the subsequent target vehicle.

4. The method for generating a trajectory of a target vehicle for tracking aerial photography by an unmanned aerial vehicle according to claim 3, characterized in that: Step four specifically includes: S41: Set the adaptive local weight allocation mechanism of the input image I(x, y), and select the region R using the elliptical adaptive window selection strategy as follows: In the above formula, a and b are respectively the long and short semi-axes of the ellipse; θ1 is the main direction angle of the ellipse; i and j represent the row and column offsets of the ellipse relative to the current center pixel (x, y); S42: Establish a hierarchical polynomial approximation model and adaptively select the approximation order according to the local motion complexity index MC(x, y): In the above formula, σ 2 is the gradient variance, is the gradient vector field, H is the historical entropy of optical flow; After the calculation is completed, perform fitting: In the above formula, a n , b n is the polynomial fitting coefficient, i.e. the parameter to be determined, n is an integer from 1 to 5; T1 is the preset complexity threshold; S43: Construct a multi-constraint fusion cost function E(u, v), integrating three constraints of brightness preservation, gradient preservation, and temporal consistency: In the above formula, (x, y) is the pixel coordinate, (u, v) is the component of the optical flow displacement field, R anis is the anisotropic regularization term; They are brightness preservation, gradient preservation, temporal consistency constraint terms and regularization weight coefficients respectively. The components in the above formula are specifically: L photo =(I(x,y)-I′(x+u,y+v)) 2 L temporal =|Flow {t-1} (x,y)-(u,v)| 2 (9) S44: Design a 5-layer non-uniform scale pyramid structure with multi-level and multi-granularity resolution expression capabilities, and use an adaptive scale factor s n : Its pyramid is specifically as follows: L0 layer, original resolution, s0 = 1.0; L1 layer, fine resolution, s1 = 0.8; L2 layer, standard resolution, s2 = 0.5; L3 layer, coarse resolution, s3 = 0.25; L4 layer, global resolution, s4 = 0.125; each layer constructs three-channel features of original brightness, gradient amplitude and local direction consistency, and adopts a bidirectional iterative refinement mechanism, that is, global motion estimation is transmitted from top to bottom, and local motion correction is fed back from bottom to top; S45: Establish a confidence-guided adaptive fusion mechanism and define a multi-dimensional confidence evaluation Con(x, y): With(x,y)= C con (x,y)·C sta (x,y)·C bou (x,y) (10) In the above formula, C con is the neighborhood optical flow consistency confidence; C sta is the confidence level of cross-scale stability; C bou Maintain confidence for motion boundaries; Final fusion: In the above formula, the subscript s represents the S-th layer of the pyramid; Flow s (x, y) refers to the optical flow obtained by minimizing energy within its own resolution at the sth layer; S46: Output enhanced optical flow information [Flow MSADOF (x,y),Con(x,y)], specifically including: motion vector field, confidence distribution map, providing high-precision motion prior information for subsequent UAV motion decoupling compensation and coordinate system transformation; S47: Fuse the optical flow estimation results at all scales, and output the final optical flow field [u(x, y), v(x, y)] to represent the motion information of each pixel in the image, which is the trajectory after the structural compensation of the UAV motion.

5. The method for generating a trajectory of a target vehicle for tracking aerial photography by an unmanned aerial vehicle according to claim 4, characterized in that: Step five specifically includes: S51: Assume that the target is in the UAV coordinate system (X u ,Y u ,Z u ), the geodetic coordinate system is (X b ,Y b ), the drone’s aerial photography angle θ, the drone’s height is h t ; Considering the influence of the lens tilt angle, first calculate the actual depth distance Z of the target real : Perform coordinate transformation on the target: In the above formula, f x ,f y is the focal length of the camera; (X′ b ,Y′ b ) is the coordinate after rotation transformation; S52: The coordinates after rotation transformation (X′ b ,Y′ b ) is integrated with the optical flow data [u(x,y),v(x,y)] of the target vehicle in the image obtained in step 4 as follows, while considering the effect of the lens tilt angle on the optical flow: In the above formula (X I , Y I ) is the integrated coordinate; p x ,p y is the pixel size; the cos(θ) term in the formula is used to correct the effect of lens tilt on the projection relationship; S53: Generate the trajectory of the target vehicle in the earth coordinate system: obtain the trajectory point sequence {(t i ,X I (t i ),Y I (t i ))|i=1,2,...,N} is the rough estimated trajectory in the geodetic coordinate system.

6. The method for generating a trajectory of a target vehicle for tracking aerial photography by an unmanned aerial vehicle according to claim 5, characterized in that: Step six includes: S61: Add the trajectory point (t i ,X I (t i ),Y I (t i The corresponding trajectory data is decomposed using the Daubechies wavelet basis to obtain the wavelet transform results under the X and Y coordinate axes; S62: Calculate the wavelet coefficient energy E j Distribution; when E j When it is greater than the preset threshold, the point k is judged to be an abnormal frequency and the corresponding trajectory point is a noise point; S63: For each noise point corresponding to the abnormal frequency (t k ,X I (t k ),Y I (t k ), select n normal points as references and use the Lagrange interpolation formula to calculate the new position; the new position is used to replace the position of the noise point; S64: Use the extended Kalman filter to further optimize the denoised trajectory data to obtain a complete and accurate trajectory.

Citation Information

Cited By

  • Bridge cable multi-target stable detection, matching tracking and vibration identification method based on large-view-field unmanned aerial vehicle video

    CN121616990A

  • Unmanned aerial vehicle dynamic target speed measurement method based on optical flow self-motion compensation

    CN122066736A

  • An unmanned aerial vehicle dynamic target speed measurement method based on optical flow self-motion compensation

    CN122066736B

  • Ground-air seamless connection cargo transportation control method based on improved neural network

    CN122198820A