A method and system for tracking targets using unmanned aerial vehicles (UAVs)
The UAV target tracking method, which combines Zernike moments and phase flow fields with the Lie group motion principle, solves the problems of high computational resource consumption and tracking errors caused by occlusion. It achieves efficient and stable target tracking, extends the UAV's endurance, and improves tracking accuracy.
Patent Information
- Application Number
- CN202510876633.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-06-27
AI Technical Summary
Existing drone target tracking methods rely on artificial intelligence for target recognition and tracking, which consumes high computing resources, resulting in reduced endurance, and are prone to target loss or tracking errors when there is background interference or occlusion.
A Zernike moment is used to extract phase information structure descriptors. Combined with the phase flow field and Lie group motion principle, target tracking is performed by minimizing the phase gradient cross product, which optimizes the computational load and predicts the target position. The occlusion factor is used to determine the occlusion situation, thereby improving the target acquisition accuracy and response speed.
It reduces computing resource consumption, extends the drone's endurance, improves target identification accuracy and tracking stability, ensures continuous and accurate target tracking in complex environments, and reduces tracking errors caused by occlusion.
Smart Images

Figure CN120807571B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a UAV target tracking method and system. BACKGROUND
[0002] A UAV is an autonomous flying vehicle, usually powered by batteries, and controlled through remote control or autonomous navigation systems for its flight and task execution. The application of UAVs covers military, agriculture, environmental monitoring, photography, express delivery and many other fields. They are widely used for their flexibility, cost-effectiveness and efficiency, especially when entering dangerous or hard-to-reach areas.
[0003] UAV target tracking is crucial in many applications, by tracking the target in real time, the UAV can accurately locate and track the movement of the target, ensuring continuous monitoring and data collection of the target. For automated tasks such as search and rescue, it can improve the efficiency and accuracy of operations and reduce the waste of human resources. At the same time, accurate target tracking can help the UAV avoid obstacles, optimize the flight path, ensure flight safety, and improve the success rate of tasks.
[0004] However, with the rapid development of artificial intelligence, current UAV target tracking largely applies artificial intelligence for target recognition and tracking, which consumes high computing resources, reduces the endurance of the UAV, and most of the image processing algorithms based on artificial intelligence are based on real-time pixel-by-pixel scanning to identify target color differences or local contour shape differences to track the target, which has insufficient response capability, further increasing the consumption of computing resources, and in the presence of background interference or occlusion, it is prone to target loss or tracking errors. SUMMARY
[0005] In view of the above deficiencies of the prior art, the purpose of the embodiments of the present application is to provide a UAV target tracking method, which can solve the technical problems of the prior art that with the rapid development of artificial intelligence, current UAV target tracking largely applies artificial intelligence for target recognition and tracking, which consumes high computing resources, reduces the endurance of the UAV, and most of the image processing algorithms based on artificial intelligence are based on real-time pixel-by-pixel scanning to identify target color differences or local contour shape differences to track the target, which has insufficient response capability, further increasing the consumption of computing resources, and in the presence of background interference or occlusion, it is prone to target loss or tracking errors.
[0006] The first aspect of the embodiments of the present application proposes a UAV target tracking method, comprising:
[0007] S1: acquiring a plurality of video frames with continuous acquisition time of target tracking label information;
[0008] S2: according to the to-be-tracked target annotation information, extracting a phase information structure descriptor of the to-be-tracked target from each video frame by using a Zernike matrix;
[0009] S3: constructing a phase flow field describing a motion feature of the to-be-tracked target according to the phase information structure descriptor;
[0010] S4: judging whether the to-be-tracked target is occluded or not by combining an occlusion factor based on the phase information structure descriptor, if yes, entering step S5, otherwise, entering step S8;
[0011] S5: determining a plurality of predicted positions of the to-be-tracked target by combining the phase flow field and a Lie group motion principle;
[0012] S6: collecting a plurality of predicted video frames by taking each predicted position as a focus center of the unmanned aerial vehicle camera;
[0013] S7: capturing the to-be-tracked target from each predicted video frame by taking minimizing a phase gradient cross product as a target;
[0014] S8: capturing the to-be-tracked target according to the phase information structure descriptor;
[0015] S9: tracking the to-be-tracked target according to the capturing result.
[0016] The second aspect of the embodiment of the present application provides an unmanned aerial vehicle target tracking system, comprising a processor and a memory.
[0017] The memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the unmanned aerial vehicle target tracking method according to the first aspect.
[0018] The third aspect of the embodiment of the present application provides a readable storage medium, and the readable storage medium stores programs or instructions, and the programs or instructions are executed by a processor to implement the steps of the unmanned aerial vehicle target tracking method according to the first aspect.
[0019] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:
[0020] In the embodiment of the present application, the phase information structure descriptor is extracted from the video frame by the Zernike moment, which can effectively capture the phase characteristics of the target, thereby reducing the influence of background interference and avoiding the loss of target caused by complex background or target blur. In addition, by constructing the phase flow field, the motion characteristics of the target can be described, and the prediction process of the target position is optimized by combining the Lie group motion principle, thereby effectively reducing the calculation amount based on image pixel scanning, reducing the consumption of computing resources, and prolonging the endurance time of the unmanned aerial vehicle. Especially when the occlusion occurs, the method can predict multiple positions of the target by combining the phase flow field and the occlusion factor, and accurately track the target, solving the problem that the traditional method cannot stably track the target when the target is occluded. In addition, the minimization method based on the phase gradient cross product further improves the accuracy and response speed of target capture, enhances the real-time performance, ensures the continuous tracking of the target, and further improves the accurate long-time tracking ability of the unmanned aerial vehicle in complex environments. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments of the application and together with the description serve to explain the principles of the application. In the drawings:
[0022] Figure 1 is a flow diagram of an unmanned aerial vehicle target tracking method provided by an embodiment of the present application;
[0023] Figure 2 is a structural diagram of an unmanned aerial vehicle target tracking system provided by an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order for those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions of the present application will be described clearly and completely below in conjunction with the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.
[0025] The unmanned aerial vehicle target tracking method provided by the embodiments of the present application will be described in detail below in conjunction with the drawings and specific embodiments and application scenarios.
[0026] Reference is made to the drawingsFigure 1 The diagram shows a flowchart of a UAV target tracking method provided by an embodiment of the present invention.
[0027] This invention provides a method for tracking unmanned aerial vehicle (UAV) targets, which may include the following steps:
[0028] S1: Acquire multiple consecutive video frames with annotation information of the target to be tracked at the acquisition time.
[0029] The target annotation information refers to the specific location, shape, and other characteristic information of the target that has been determined through preprocessing or manual annotation when the drone first captures the target. Continuous acquisition time means that the drone continuously captures multiple video frames, and these frames are consecutive in time.
[0030] It should be noted that at least two video frames are required to capture motion features. Before tracking, the drone continuously acquires multiple video frames of the target to be tracked, then transmits them back for annotation to identify the target. These frames are used to capture the target's initial position and motion features. By continuously acquiring video frames, it ensures a more accurate description of the target's dynamic behavior, providing accurate initial data for subsequent target identification and tracking.
[0031] S2: Based on the target annotation information, use Zernike moments to extract the phase information structure descriptor of the target from each video frame.
[0032] Zernike moments are mathematical tools based on Zernike phase fields used to extract phase information from images. By processing the phase of an image, Zernike moments can effectively describe the microstructure of a target and capture detailed information within the image. The phase information structure descriptor is a mathematical description extracted from the image, describing the phase structure of the target, particularly its edge features and local shape. Higher-order components of Zernike moments can capture the fine structure of the target, and phase information is robust to changes in illumination, local occlusion, and background noise, making it less susceptible to environmental variations and thus enhancing the stability of target recognition.
[0033] It should be noted that by utilizing Zernike moments to extract phase information structural descriptors, a target's "structural fingerprint" can be effectively constructed. This fingerprint is insensitive to local occlusion, illumination changes, and background noise. Compared to traditional image feature extraction methods, Zernike moments better preserve the stability of the target in complex environments, thereby improving the accuracy and reliability of target tracking, especially under conditions of significant occlusion or background interference, enabling continuous and accurate target tracking.
[0034] In one possible implementation, S2 specifically includes:
[0035] S201: Locate the information region containing the target in the video frame based on the target annotation information.
[0036] The information area is a roughly circular labeled area, which is the smallest circle that can completely encompass the target to be tracked.
[0037] S202: Calculate the weighted integral of the grayscale image and the Zernike polynomial under different order combinations within the information area to obtain the Zernike moments describing the target to be tracked, where the order combination is a combination of the radial and angular orders of the Zernike polynomial.
[0038] The radial order describes the degree of variation of the Zernikal polynomial in the radial direction, while the angular order describes the degree of variation in the angular direction. Their orders correspond and can be set according to actual needs.
[0039] The specific formula for calculating the Zernike moment is as follows:
[0040]
[0041] Where n represents the radial order and m represents the angular order. This represents the Zernike moment of the target to be tracked corresponding to the video frame at time t, which is related to n and m. Represents pi (π). This represents the normalization coefficient used to eliminate the influence of the information region size on the Zernike moment. This represents the information region corresponding to the video frame at time t. This represents the pixel grayscale value at position x in the video frame at time t. Represents the polar coordinates of each pixel in the information area. Indicates and Zernikic polynomials related to n and m.
[0042] Zernike polynomials are a class of orthogonal polynomials defined in polar coordinates and whose shape is controlled by radial and angular orders. Zernike polynomials can effectively represent the geometry of objects in images, especially within circular or near-circular regions. The advantage of Zernike polynomials lies in their orthogonality, enabling the extraction of independent features and avoiding interference between different features.
[0043] S203: Extract the phase components of the Zernike moments under different order combinations.
[0044] The specific formula for extracting the phase component is as follows:
[0045]
[0046] in, This represents the phase component corresponding to the video frame at time t, which is related to n and m. This represents the argument function.
[0047] S204: Combine the phase components corresponding to each video frame to obtain the phase information structure descriptor of each video frame.
[0048] The specific formula for combining phase information structure descriptors is as follows:
[0049]
[0050] in, This represents the phase information structure descriptor corresponding to the video frame at time t. This indicates the preset maximum radial order.
[0051] It should be noted that those skilled in the art can set the size of the preset maximum radial step according to actual needs, and this invention does not limit this.
[0052] Specifically, the core of the entire process is to describe and extract the phase information of the target using Zernike moments to improve the stability and accuracy of target tracking. First, the target's information region in the video frame is located and labeled with a minimum circle to ensure the target is completely contained within it. Next, the Zernike moments of the target are extracted by calculating the weighted integral of the grayscale image and the Zernike polynomial, accurately modeling the target's shape and structure. This involves combining radial and angular orders to capture details at different levels of the target, and a normalization coefficient is used to eliminate the influence of the information region size. Step S203 extracts the phase components and obtains accurate phase information through argument functions. Finally, step S204 combines the phase components into a phase information structure descriptor, providing a stable feature description for subsequent target tracking. This method, through phase information, exhibits strong robustness to local occlusion, illumination changes, and background noise. Compared to traditional feature extraction methods based on color, shape, or edges, it maintains greater stability and accuracy in target tracking, especially in complex environments and under rapid motion, significantly improving the continuous tracking capability of the target.
[0053] S3: Construct a phase flow field describing the motion characteristics of the target to be tracked based on the phase information structure descriptor.
[0054] The phase flow field is constructed based on the extracted phase information structure descriptor and is used to represent the target's motion features in video frames. The phase flow field describes the target's trajectory by tracking changes in phase information across consecutive video frames, thus transforming static phase features into a dynamic motion field. This process solves the problem of separation between static features and dynamic motion estimation in traditional methods, enabling a close integration of motion estimation and target tracking. Through the phase flow field, the target's trajectory can still be predicted even when briefly occluded, ensuring continuous target tracking and occlusion detection.
[0055] Understandably, by establishing motion continuity constraints, this method can maintain tracking by predicting the trajectory when the target is briefly occluded, thus avoiding the trajectory interruption problem in traditional methods. This motion continuity constraint significantly improves the stability and accuracy of tracking, especially in complex environments, effectively handling occlusion and motion changes.
[0056] In one possible implementation, S3 specifically includes:
[0057] S301: Determine the correlation between each phase information structure descriptor and the phase flow field based on the principle of phase energy conservation.
[0058] The specific relationships are as follows:
[0059]
[0060] in, The phase information structure descriptor of the video frame at time t Spatial gradient, Represents the phase flow field, This indicates the partial derivative.
[0061] It should be noted that, based on the principle of phase energy conservation, by ensuring the conservation of phase information in space and time, the motion characteristics of the target can be accurately described. The combination of the spatial gradient of phase information and the phase flow field ensures the consistency of phase during motion, which makes the target motion estimation more stable and accurate. Especially when dealing with complex motion and occlusion, it can effectively reduce errors and improve the robustness and accuracy of target tracking.
[0062] S302: The phase flow field is optimized with the goal of minimizing the energy function with data terms and regularization terms. The data terms are used to constrain the phase of the video frame pixels to satisfy phase conservation during motion, and the regularization terms are used to constrain the adjacent pixels of the video frame to be in a smooth state in space.
[0063] The energy function is as follows:
[0064]
[0065] in, This represents the phase flow field to be optimized. express Energy function value, express Spatial gradient, Represents a data item. Indicates adjustment of regular expression terms Weighting coefficients for importance.
[0066] It should be noted that by minimizing the energy function, which includes data and regularization terms, the phase flow field can be effectively optimized, ensuring the continuity and smoothness of the phase in both time and space. The data term guarantees the conservation of phase during motion, making motion prediction more accurate during target tracking. The regularization term, by smoothing the phase flow field, avoids errors caused by local noise or drastic changes. The overall optimization process, by balancing the data and regularization terms, can stably track targets in complex environments while maintaining computational efficiency and accuracy.
[0067] Specifically, the energy function can be solved using variational methods or iterative algorithms such as the Gauss-Newton method or the conjugate gradient method to obtain a unique phase flow field.
[0068] S303: Output the optimized phase flow field.
[0069] Specifically, the entire process accurately estimates and optimizes the phase flow field using the phase energy conservation principle and energy optimization methods, thereby ensuring continuous target tracking and stability. More specifically, firstly, the correlation between the phase information structure descriptor and the phase flow field is established using the phase energy conservation principle, ensuring the continuous stability of the phase information through spatial gradient and temporal variation terms. Then, the phase flow field is optimized by minimizing an energy function containing data and regularization terms. The data terms guarantee the conservation of phase during motion, while the regularization term constrains the smoothness of the phase flow field, avoiding inaccurate tracking results due to drastic changes. The optimal phase flow field can be accurately solved using variational methods or iterative algorithms (such as the Gauss-Newton method or the conjugate gradient method). In step S303, the optimized phase flow field is output as the final result for subsequent target tracking. This method can handle complex motion patterns and occlusion situations, accurately predict targets, reduce errors and computational burdens in traditional methods, and improve tracking stability.
[0070] S4: Determine whether the target to be tracked is occluded by combining the occlusion factor based on the phase information structure descriptor. If yes, proceed to step S5; otherwise, proceed to step S8.
[0071] The occlusion factor is a metric based on the phase information structure descriptor used to determine whether the target being tracked is occluded. Specifically, the occlusion factor analyzes the changes in the target's phase information in the current frame to assess whether the target is partially or completely occluded. If the target's phase information changes abnormally, or the target's features are lost, it is determined that the target may be occluded, thus triggering subsequent processing steps, such as proceeding to step S5 for occlusion processing. The introduction of the occlusion factor can effectively detect whether the target is occluded, avoiding tracking errors caused by occlusion. The occlusion factor can accurately determine whether the target is occluded, which provides crucial information for subsequent target tracking. Compared with traditional occlusion detection methods based on appearance or edge features, this method, based on phase information occlusion judgment, has stronger robustness to illumination changes, background interference, and target deformation, and can more accurately identify whether the target is occluded, thereby reducing tracking loss or errors caused by occlusion and enhancing the stability and reliability of target tracking.
[0072] In one possible implementation, S4 specifically includes:
[0073] S401: Under each order combination, calculate the phase difference between adjacent video frames at different times.
[0074] The specific formula for calculating the phase difference between adjacent video frames is as follows:
[0075]
[0076] in, This represents the phase component of the video frame at time t-1 under the combination of orders n and m. This represents the phase component of the video frame at time t under the order combination of n and m. This indicates the phase difference between adjacent video frames.
[0077] S402: Convert the phase difference between adjacent video frames into a unit vector.
[0078] The unit vector obtained after the transformation is:
[0079]
[0080] in, represents the natural constant, and i represents the imaginary unit.
[0081] S403: Summing the obtained unit vectors under different order combinations yields a total vector describing the consistency of structural changes in the target to be tracked.
[0082] S404: Normalize the total vector by dividing the magnitude of the total vector by the number of order combinations.
[0083] The total vector magnitude reflects whether the changes in the target being tracked are consistent. If they are consistent, the total length is large; if the directions are chaotic, the total length is small.
[0084] S405: Calculate the difference between 1 and the normalized total vector to obtain the occlusion factor.
[0085] The occlusion factor is calculated as follows:
[0086]
[0087] in, This represents the occlusion factor corresponding to the video frame at time t.
[0088] It should be noted that the occlusion factor is obtained by calculating the difference between the normalized total vector magnitude and 1. By comparing the normalized total vector magnitude, the phase consistency of the target across consecutive frames is determined. If the target is not occluded, the phase is consistent, the total vector magnitude is close to 1, and the occlusion factor is close to 0; if the target is occluded or the phase changes are inconsistent, the total vector magnitude decreases, and the occlusion factor increases. This method has strong anti-interference capabilities and can accurately distinguish between target occlusion and motion changes based on phase, especially in complex scenes and with a lot of background noise, effectively reducing misjudgments and improving stability.
[0089] S406: If the occlusion factor is greater than the preset occlusion factor, determine that the target to be tracked is occluded; otherwise, determine that the target to be tracked is not occluded.
[0090] It should be noted that those skilled in the art can set the size of the preset occlusion factor according to actual needs, and this invention does not limit it.
[0091] Optionally, the occlusion factor can be set to 0.8 or 0.9.
[0092] Specifically, the entire process calculates the phase difference between adjacent video frames and uses unit vectors to describe the consistency of structural changes in the target, thereby determining whether the target is occluded. In step S401, the phase difference is calculated to obtain the changes in adjacent video frames. In step S402, the phase difference is converted into a unit vector to ensure numerical stability of the calculation. In step S403, the unit vectors are summed to obtain a total vector describing the consistency of structural changes in the target. Further, in step S404, the total vector is normalized to measure the consistency of target changes. In step S405, the difference between the magnitude of the total vector and 1 is calculated to obtain the occlusion factor, and step S406 determines whether the target is occluded. The advantage of this method is that by accurately calculating the occlusion factor, it can dynamically assess whether the target is occluded, avoiding misjudgments in traditional methods.
[0093] This method calculates the phase difference between video frames and converts it into a unit vector to accurately capture the consistency of target changes under different order combinations, thus effectively assessing whether the target is occluded. Through normalization processing and occlusion factor calculation, it can accurately distinguish between target changes and occlusion, avoiding tracking errors caused by background interference or partial target occlusion. Compared with traditional methods, this phase difference-based occlusion detection is more robust, can continuously track targets in complex scenes, and improves tracking accuracy and stability.
[0094] S5: Combining the phase flow field, multiple predicted positions of the target to be tracked are determined by the Lie group motion principle.
[0095] The Lie group motion principle is a mathematical method based on Lie group and Lie algebra theory used to describe and handle rigid body motion. Lie groups are a class of groups that describe rigid body motion (such as translation and rotation transformations). Combining geometric transformations and group theory, it can effectively model and predict motion. In target tracking, the Lie group motion principle can be used to infer the target's possible future position based on its current motion state (position and velocity, etc.), thus predicting the target's trajectory. By incorporating phase flow fields, the Lie group motion principle can more accurately predict multiple target positions, enhancing the accuracy of motion estimation.
[0096] It's worth noting that by utilizing the motion principles of Lie groups to predict multiple target positions, a more accurate motion estimate is provided for continuous target tracking. This method can precisely capture the target's motion patterns, maintaining high efficiency during both translation and rotation. In cases of occlusion or rapid movement, predicting multiple possible target positions avoids target loss or mistracking, enhancing the stability and accuracy of target tracking. This Lie group-based motion prediction significantly improves the robustness and reliability of tracking in complex environments.
[0097] In one possible implementation, S5 specifically includes:
[0098] S501: Obtain the real-time position of the target to be tracked from the video frame at time t-1.
[0099] S502: Based on the video frame at time t-1, use the Lie algebra prior template to determine the Lie algebra generators that describe the motion law of the target to be tracked. The Lie algebra generators include translation generators, rotation generators and scaling generators.
[0100] Among them, the Lie algebra prior template is the basic template or model used to represent the motion of a target, which includes the mathematical forms of basic motions such as translation, rotation, and scaling. The Lie algebra generator is a set of basic operators that describe the elements of the Lie group (such as transformations such as translation, rotation, and scaling).
[0101] Optionally, the prior templates for extracting translation generators specifically include an x-axis translation prior template and a y-axis translation prior template, wherein the formula for the x-axis translation prior template is as follows:
[0102] .
[0103] The specific formula for the y-axis translation prior template is as follows:
[0104] .
[0105] The formula for the prior template of the rotation generator is as follows:
[0106] .
[0107] The prior templates for scaling generators specifically include isotropic scaling prior templates and anisotropic scaling prior templates. The formula for the isotropic scaling prior template is as follows:
[0108] .
[0109] The specific formula for the anisotropic prior template is as follows:
[0110] .
[0111] Among them, in the formula form of each prior template, to These represent the generators of the corresponding basic motions.
[0112] S503: The intensity coefficients describing the motion of the target to be tracked are obtained by fitting the phase flow field. The intensity coefficients include translational velocity, rotational angular velocity, and scaling velocity.
[0113] Specifically, the fitting process involves decomposing the phase flow field into various basic motions to obtain the contribution of each basic motion (i.e., how many meters it moves per second in the x direction, how many meters it moves per second in the y direction, how many degrees it rotates per second, the shrinkage rate per second, and the magnification rate per second). Then, the decomposition results are matched or mapped to the corresponding prior templates to obtain the corresponding intensity coefficients.
[0114] S504: Combines intensity coefficients and Lie algebra generators into predicted locations through exponential mapping.
[0115] The specific formula for calculating the predicted location is as follows:
[0116]
[0117] in, The exp function represents the real-time position of the target at time t-1. This represents the predicted position of the target at time t. Denotes the generator of the k-th basic motion. This represents the intensity coefficient of the k-th basic motion at time t. This indicates the number of basic motion types, which include translation, rotation, and scaling.
[0118] It should be noted that by combining exponential mapping with Lie algebra generators and intensity coefficients, the translation, rotation, and scaling motions of the target can be accurately unified into a single prediction model, providing precise target position prediction. This method can simultaneously consider multiple motion types, adapt to the diverse changes of targets in complex dynamic environments, and handle nonlinear motion more naturally, improving the accuracy and robustness of target tracking. In particular, it ensures more stable tracking performance when the target has multiple motion modes (such as rotation or scaling).
[0119] Specifically, the entire process accurately predicts the target's trajectory based on Lie algebra generators and a phase flow field. In S501, the position of the target in the previous video frame is first obtained. Next, in S502, the target's motion is determined using a Lie algebra prior template, generating generators for translation, rotation, and scaling to describe different types of motion. In S503, intensity coefficients describing the target's motion are calculated by fitting the phase flow field; these coefficients specifically include translational velocity, rotational angular velocity, and scaling velocity. In S504, these intensity coefficients are combined with the corresponding Lie algebra generators through exponential mapping to obtain the predicted position of the target. The advantage of this method is that it effectively captures the target's translational, rotational, and scaling changes using Lie algebra generators and combines this with phase flow field fitting for accurate motion prediction. Compared to traditional methods, this approach can more accurately handle complex changes in target motion and improves the accuracy and robustness of target tracking in complex environments.
[0120] S6: Collect multiple predicted video frames using each predicted location as the focus center of the drone camera.
[0121] It should be noted that by acquiring multiple predicted video frames with each predicted position as the focus center of the UAV camera, it is possible to ensure that the UAV always accurately observes and captures the target. By focusing on the predicted position of the target, the UAV can adjust the camera focus more quickly, reducing the risk of the target deviating from the lens and increasing the probability of target capture. In addition, this method can effectively deal with situations of rapid movement or brief occlusion, ensuring stable target tracking while reducing image loss caused by excessive target displacement, thus improving the real-time performance and reliability of tracking.
[0122] S7: Capture the target to be tracked from each predicted video frame with the goal of minimizing the phase gradient cross product.
[0123] Minimizing the phase gradient cross product is an optimization method used to accurately capture targets. The phase gradient cross product is a metric for measuring changes in the target's phase information, representing the direction and magnitude of these changes. When the target's phase information changes in consecutive predicted video frames, the phase gradient is calculated and a cross product operation is performed to obtain a quantified value of the change. The goal of minimizing the phase gradient cross product is to adjust the UAV camera parameters to minimize the change in the target's phase gradient across the predicted video frames, thereby accurately capturing the target's position. This process ensures minimal error capture of target features, thus improving tracking accuracy.
[0124] In one possible implementation, S7 specifically includes:
[0125] S701: Calculate the predicted phase information structure descriptor within the preset neighborhood of each predicted location.
[0126] It should be noted that the calculation method of the phase information structure descriptor in step S701 is the same as that in step S2, and the entire preset neighborhood in S701 is similar to the annotation information of the target to be tracked.
[0127] S702: Calculate the phase gradient cross product between each predicted phase information structure descriptor and the phase information structure descriptor.
[0128] The specific method for calculating the phase gradient cross product is as follows:
[0129]
[0130] in, Describing the L1 norm, and These represent the predicted phase information structure descriptors. The phase information structure descriptor of the target to be tracked at time t-1, i.e., in the unoccluded state. Spatial gradient.
[0131] The spatial gradient is a quantity that describes the rate of change of an image or function in space. It represents the direction and rate of change of grayscale value at a point in an image, and is calculated by differentiating the image. In an image, the magnitude of the spatial gradient reflects the drasticness of the grayscale value change, while the direction indicates the direction of the greatest change.
[0132] S703: Select the predicted video frame corresponding to the minimum phase gradient cross product.
[0133] S704: The predicted position corresponding to the selected predicted video frame is used as the capture position of the target to be tracked for capture.
[0134] Specifically, firstly, the phase information structure descriptor within the neighborhood of the predicted location is calculated. Next, the difference between the predicted phase information descriptor and the target phase descriptor is evaluated by calculating the phase gradient cross product, and then the predicted video frame corresponding to the minimum phase gradient cross product is selected. Then, the optimal predicted location is used as the target capture location. By calculating the phase gradient cross product, the matching degree of the predicted location can be accurately evaluated, and the predicted frame that best matches the target's current state can be selected, thereby improving the accuracy and stability of target tracking. Especially when the target moves rapidly or is occluded, it can effectively reduce tracking errors and ensure stable target capture.
[0135] S8: Capture the target to be tracked based on the phase information structure descriptor.
[0136] In one possible implementation, S8 specifically includes:
[0137] S801: Acquire the video frame at time t+1, and divide the video frame at time t+1 into multiple sub-regions according to the preset number of regions.
[0138] It should be noted that those skilled in the art can set the size of the preset area according to actual needs, and this invention does not limit this.
[0139] S802: Calculate the sub-region phase information structure descriptor for each sub-region.
[0140] S803: Calculate the sub-region phase gradient cross product between the sub-region phase information structure descriptor and the phase information structure descriptor corresponding to the video frame at time t.
[0141] S804: Select the sub-region corresponding to the cross product of the phase gradient of the smallest sub-region as the capture region.
[0142] S805: Use the midpoint of the capture area as the capture position of the target to be tracked.
[0143] S806: Capture the target to be tracked based on the capture location.
[0144] Specifically, the entire process improves the accuracy and stability of target tracking by refining the target acquisition region. First, the video frame at time t+1 is acquired and segmented into multiple sub-regions to more accurately analyze target details. Next, the phase information structure descriptor for each sub-region is calculated to capture features within the region. Then, the phase gradient cross product between each sub-region and the frame at time t is calculated, the matching degree is evaluated, and the sub-region corresponding to the smallest phase gradient cross product is selected as the acquisition region. The target acquisition position is then determined by selecting the midpoint of the acquisition region. The advantage of this method is that by subdividing the video frame and accurately calculating the phase information of each sub-region, the target position can be captured more accurately, especially when the target pose or background is complex, reducing errors and ensuring stable and efficient tracking.
[0145] Following S8, it also includes:
[0146] If no capture results are obtained, perform a hover traversal scan to acquire multiple environmental video frames.
[0147] The target to be tracked is captured from each environmental video frame using phase information structure descriptors.
[0148] Specifically, when a target capture fails, the system acquires multiple environmental video frames by performing a hover traversal scan. These video frames provide more data sources for further target capture. Then, using a phase information structure descriptor, target features are extracted from each environmental video frame to improve the accuracy and robustness of target capture. The advantage of this method is that by hovering and scanning the environment, the system can provide more perspectives and data even when the target is not captured, enhancing the stability of target tracking in dynamic or complex environments, avoiding errors caused by local occlusion or environmental changes, and ensuring the accuracy of continuous tracking.
[0149] S9: Track the target to be tracked based on the capture results.
[0150] In practical applications, the system first constructs a "structural fingerprint" of the target by continuously acquiring video frames and extracting the target's phase information. This allows the target to maintain stable tracking even in complex environments (such as occlusion and changes in lighting). By combining phase flow field and Lie group motion principles, the system can accurately predict the target's trajectory, further reducing tracking errors caused by occlusion or rapid target movement. The matching degree between the target and the predicted frame is accurately evaluated by minimizing the phase gradient cross product, ensuring accurate capture. Further improvements in target capture accuracy and robustness are achieved by subdividing video frames and calculating phase descriptors for sub-regions. When the target is not successfully captured, the system performs hover scanning to collect more environmental video frames, ensuring efficient and accurate target capture even in dynamic environments. This method maintains high-precision tracking in complex backgrounds, occlusion, and irregular motion environments, improving the stability and robustness of target tracking and ensuring continuous and accurate operation of the UAV in changing environments.
[0151] The beneficial effects of the technical solutions provided by the embodiments of the present invention include at least the following:
[0152] In this embodiment of the invention, phase information structure descriptors are extracted from video frames using Zernike moments. This method effectively captures the phase features of the target, thus reducing the impact of background interference and eliminating reliance on traditional color or shape differences. In complex backgrounds or situations where the target is similar to the background, it improves target identification accuracy and avoids target loss due to complex backgrounds or blurred targets. Secondly, by constructing a phase flow field, the motion characteristics of the target can be described, and prediction is performed using the Lie group motion principle, optimizing the target position prediction process. This effectively reduces the computational load based on image pixel scanning, decreases computational resource consumption, and extends the UAV's endurance. Especially when occlusion occurs, the method, by combining the phase flow field and occlusion factor judgment, can predict multiple target positions and perform accurate tracking, solving the problem of unstable tracking in traditional methods when the target is occluded. Furthermore, the method based on minimizing the phase gradient cross product further improves the accuracy and response speed of target acquisition, enhances real-time performance, ensures continuous target tracking, and thus improves the UAV's accurate long-term tracking capability in complex environments.
[0153] The drone target tracking method provided in this application can be executed by a drone target tracking device. This application uses a drone target tracking device executing the drone target tracking method as an example to illustrate the drone target tracking device provided in this application.
[0154] Reference manual attached Figure 2 The diagram shows a schematic representation of a drone target tracking system provided in an embodiment of the present invention.
[0155] This invention provides an unmanned aerial vehicle (UAV) target tracking system 20, comprising: a processor 201 and a memory 202;
[0156] The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the steps of the above-described UAV target tracking method and achieve the same technical effect. To avoid repetition, the present invention will not elaborate further.
[0157] It should be understood that the processor 201 in this embodiment of the invention may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0158] It should also be understood that the memory 202 in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM).
[0159] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0160] It should be understood that, in various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0161] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0162] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0163] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0164] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0165] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0166] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0167] This invention provides a readable storage medium comprising: storing a program or instructions on the readable storage medium, wherein when the program or instructions are executed by a processor, the program or instructions implement the steps of the above-described UAV target tracking method and achieve the same technical effect. To avoid repetition, this invention will not elaborate further.
[0168] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the protection scope of the present invention.
Claims
1. A method for tracking targets using an unmanned aerial vehicle (UAV), characterized in that the method... The method comprises the following steps: S1: acquiring a plurality of video frames with continuous acquisition time and target annotation information to be tracked; S2: extracting phase information structural descriptor of the target to be tracked from each video frame by using Zernike moments according to the target annotation information to be tracked; S3: constructing a phase flow field describing the motion characteristics of the target to be tracked according to the phase information structural descriptor; S4: combining the occlusion factor based on the phase information structural descriptor to determine whether the target to be tracked is occluded, if yes, entering step S5, otherwise, entering step S8; S5: combining the phase flow field to determine a plurality of predicted positions of the target to be tracked by using Lie group motion principle; S6: acquiring a plurality of predicted video frames with each predicted position as the focus center of the unmanned aerial vehicle camera; S7: capturing the target to be tracked from each predicted video frame by minimizing the phase gradient cross product; S8: capturing the target to be tracked according to the phase information structural descriptor; S9: tracking the target to be tracked according to the capture result. 2.The UAV target tracking method of claim 1, wherein, The S2 specifically comprises: S201: locating an information region including the target to be tracked in the video frame according to the target annotation information to be tracked; S202: calculating the weighted integral of the gray scale graph and Zernike polynomials under different order combinations in the information region to obtain Zernike moments describing the target to be tracked, wherein the order combination is the combination of the radial order and the angular order of the Zernike polynomials; S203: extracting the phase component of the Zernike moments under different order combinations respectively; S204: combining the phase components corresponding to each video frame to obtain the phase information structural descriptor of each video frame. 3.The UAV target tracking method of claim 1, wherein, The S3 specifically comprises: S301: determining the correlation between each phase information structural descriptor and the phase flow field based on the phase energy conservation principle; S302: optimizing the phase flow field by minimizing the energy function with data items and regularization terms, wherein the data items are used to constrain the phase of the video frame pixels to satisfy the phase conservation in the motion, and the regularization terms are used to constrain the adjacent pixels of the video frame to be in a smooth state in space; S303: outputting the phase flow field obtained by optimization. 4.The UAV target tracking method of claim 2, wherein, The S4 specifically comprises: S401: calculating the adjacent video frame phase difference between video frames at different time under each order combination; S402: converting each adjacent video frame phase difference into a unit vector; S403: summing up the unit vectors obtained under different order combinations to obtain a total vector describing the consistency of the structural changes of the target to be tracked; S404: normalizing the total vector by dividing the total vector module length by the number of order combination groups; S405: calculating the difference between 1 and the normalized total vector to obtain the occlusion factor; S406: determining that the target to be tracked is occluded if the occlusion factor is greater than a preset occlusion factor, otherwise, determining that the target to be tracked is not occluded. 5.The UAV target tracking method of claim 1, wherein, The S5 specifically comprises: S501: acquiring the real-time position of the target to be tracked from the video frame at time t-1; S502: determining a Lie algebra generator describing a motion rule of the target to be tracked according to a video frame at time t-1, by using a Lie algebra prior template, wherein the Lie algebra generator comprises a translation generator, a rotation generator and a scaling generator; S503: fitting an intensity coefficient describing the motion rule of the target to be tracked according to the phase flow field, wherein the intensity coefficient comprises a translation velocity, a rotation angular velocity and a scaling velocity; S504: synthesizing the intensity coefficient and the Lie algebra generator into the predicted position by exponential mapping. 6.The UAV target tracking method of claim 1, wherein, The S7 specifically comprises: S701: calculating a predicted phase information structure descriptor in a preset neighborhood of each predicted position; S702: calculating a phase gradient cross product between each predicted phase information structure descriptor and the phase information structure descriptor; S703: selecting a predicted video frame corresponding to a minimum phase gradient cross product; S704: capturing a predicted position corresponding to the selected predicted video frame as a capture position of the target to be tracked. 7.The UAV target tracking method of claim 1, wherein, The S8 specifically comprises: S801: obtaining a video frame at time t+1, and dividing the video frame at time t+1 into a plurality of sub-regions according to a preset number of regions; S802: calculating a sub-region phase information structure descriptor of each sub-region; S803: calculating a sub-region phase gradient cross product between the sub-region phase information structure descriptor and a phase information structure descriptor corresponding to the video frame at time t; S804: selecting a sub-region corresponding to a minimum sub-region phase gradient cross product as a capture region; S805: taking a midpoint of the capture region as a capture position of the target to be tracked; S806: capturing the target to be tracked according to the capture position. 8.The UAV target tracking method of claim 1, wherein, After the S8, further comprising: In a case where a capture result is not obtained, performing hovering traversal scanning to obtain a plurality of environmental video frames; Capturing the target to be tracked from each of the environmental video frames by using the phase information structure descriptor.
9. An unmanned aerial vehicle target tracking system, comprising: Comprise: A processor and a memory; The memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the steps of the UAV target tracking method according to any one of claims 1 to 8.
10. A readable storage medium, characterized by, The readable storage medium stores programs or instructions, and the programs or instructions are executed by the processor to implement the steps of the UAV target tracking method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Target tracking method and device based on unmanned aerial vehicle video and computer equipment
CN113936036A
Cross-regional phase unwrapping method and device and computer readable storage medium
CN118565329A