Unmanned aerial vehicle single-target tracking method, heterogeneous embedded system and device
By deploying the core computing tasks of target tracking on a programmable logic unit through a heterogeneous embedded system, the problem of limited computing power of UAV onboard platforms is solved, and real-time target tracking of high frame rate video streams is realized, meeting the requirements of low latency and high real-time performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN QIYANG SPECIAL EQUIP TECH ENG CO LTD
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-29
AI Technical Summary
The computing power of UAV onboard computing platforms is limited, making it difficult to meet the real-time target tracking requirements of high frame rate video streams, especially when the movement speed is extremely fast. Traditional solutions are difficult to meet the stringent requirements of low latency and high real-time tracking due to high power consumption and large computing power requirements.
A heterogeneous embedded system is adopted, which uses a general-purpose processor unit to perform initial target detection and deploys the core computing tasks on a programmable logic unit. Through a preset deep neural network and a preset tracking model, the real-time state information of the object to be tracked and the motion state of the UAV are determined, thus meeting the real-time tracking requirements.
Without relying on a high-performance graphics processor, frame-by-frame tracking in high frame rate video streams was achieved, meeting the real-time and robustness requirements of UAV airborne platforms, reducing power consumption and size, and improving computational load offloading efficiency.
Smart Images

Figure CN122115500A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of unmanned aerial vehicle (UAV) control, and more specifically, to a single-target tracking method, heterogeneous embedded system, and device for UAVs. Background Technology
[0002] The image information processed by UAV onboard image acquisition equipment is mostly 30 frames per second (FPS), while most target tracking algorithms need to process a large amount of image information in real time, which puts high demands on the computing platform's computing performance.
[0003] In related technologies, target tracking algorithms primarily rely on high-performance computers or devices equipped with dedicated graphics processing cards. These computers are bulky and consume a lot of power, limiting their application in mobile tracking scenarios.
[0004] Unmanned aerial vehicle (UAV) onboard computing platforms are often limited in computing power due to payload and power consumption constraints. Because UAVs move at extremely high speeds in actual operation, the real-time requirements for image processing are even more stringent. Therefore, the real-time performance issue of UAV ground target tracking platforms urgently needs to be addressed. Summary of the Invention
[0005] In view of the above problems, this application proposes a single-target tracking method, heterogeneous embedded system and device for UAVs, which can solve the above problems.
[0006] In a first aspect, embodiments of this application provide a single-target tracking method for unmanned aerial vehicles (UAVs), applied to a heterogeneous embedded system. The heterogeneous embedded system includes a general-purpose processor unit and a programmable logic unit (PLU). The method includes: acquiring a video stream containing an object to be tracked; running a preset deep neural network through the general-purpose processor unit to detect the object to be tracked contained in the target image frames of the video stream and determine the initial state information of the object to be tracked; inputting the initial state information and the video stream into a preset tracking model deployed in the PLU to determine the real-time state information of the object to be tracked in each image frame of the video stream; and adjusting the spatial motion state of the UAV according to the offset of the real-time state information so that the object to be tracked remains continuously located in the center region of the field of view.
[0007] Secondly, embodiments of this application also provide a heterogeneous embedded system, which includes a general-purpose processor unit and a programmable logic unit. The system includes: an acquisition module for acquiring a video stream containing an object to be tracked; a general-purpose processor unit for running a preset deep neural network to detect the object to be tracked contained in the target image frames of the video stream and determine the initial state information of the object to be tracked; a programmable logic unit for determining the real-time state information of the object to be tracked in each image frame of the video stream based on the initial state information and the video stream; and an execution module for adjusting the spatial motion state of the UAV based on the offset of the real-time state information so that the object to be tracked remains in the center of the field of view.
[0008] Thirdly, embodiments of this application also provide a drone, including a processor, a memory, and one or more applications; the one or more applications are stored in the memory and configured to be executed by the processor to implement the above-described drone single-target tracking method.
[0009] The technical solution provided in this application is applied to a heterogeneous embedded system, which includes a general-purpose processor unit and a programmable logic unit. The method includes: acquiring a video stream containing an object to be tracked; running a preset deep neural network through the general-purpose processor unit to detect the object to be tracked contained in the target image frames in the video stream and determine the initial state information of the object to be tracked; inputting the initial state information and the video stream into a preset tracking model deployed in the programmable logic unit to determine the real-time state information of the object to be tracked in each image frame of the video stream; and adjusting the spatial motion state of the UAV according to the offset of the real-time state information so that the object to be tracked remains in the center of the field of view. Therefore, by deploying the core computational task of target tracking on the programmable logic unit (PLU), while only the general-purpose processor unit performs the initial target detection once, the computational load is efficiently offloaded. Since the PLU has hardware-level parallel processing capabilities and low power consumption and small size, frame-by-frame tracking in high frame rate video streams can be completed in real time on the resource-constrained UAV onboard platform, and real-time status information for flight control adjustments can be directly output. Thus, without relying on a high-performance graphics processor, the traditional solution, which is difficult to meet the stringent requirements of UAVs for low latency and high real-time tracking due to high power consumption and large computing power requirements, is solved. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments and drawings obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0011] Figure 1 A schematic diagram of the structure of a heterogeneous embedded system according to an embodiment of this application is shown.
[0012] Figure 2 A flowchart illustrating a single-target tracking method for unmanned aerial vehicles (UAVs) provided in an embodiment of this application is shown.
[0013] Figure 3 A schematic diagram of another heterogeneous embedded system provided in an embodiment of this application is shown.
[0014] Figure 4 A schematic diagram of the structure of a drone provided in an embodiment of this application is shown. Detailed Implementation
[0015] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.
[0016] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0017] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0018] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0019] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0020] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0021] This application provides a method for tracking a single target on a UAV, a heterogeneous embedded system, and an apparatus. The method, applied to a heterogeneous embedded system including a general-purpose processor unit and a programmable logic unit, includes: acquiring a video stream containing the object to be tracked; running a preset deep neural network through the general-purpose processor unit to detect the object to be tracked in the target image frames of the video stream and determine the initial state information of the object; inputting the initial state information and the video stream into a preset tracking model deployed in the programmable logic unit to determine the real-time state information of the object to be tracked in each image frame of the video stream; and adjusting the spatial motion state of the UAV according to the offset of the real-time state information to keep the object to be tracked continuously located in the center of the field of view.
[0022] Therefore, by deploying the core computational task of target tracking on the programmable logic unit (PLU), while only the general-purpose processor unit performs the initial target detection once, the computational load is efficiently offloaded. Since the PLU has hardware-level parallel processing capabilities and low power consumption and small size, frame-by-frame tracking in high frame rate video streams can be completed in real time on the resource-constrained UAV onboard platform, and real-time status information for flight control adjustments can be directly output. Thus, without relying on a high-performance graphics processor, the traditional solution, which is difficult to meet the stringent requirements of UAVs for low latency and high real-time tracking due to high power consumption and large computing power requirements, is solved.
[0023] Please see Figure 1 , Figure 1 This paper illustrates a schematic diagram of a heterogeneous embedded system according to an embodiment of this application, as shown below. Figure 1 As shown, the heterogeneous embedded system 100 includes a processing unit 110 (Processing System, PS), a programmable logic unit 120 (Programmable Logic, PL), an interconnection bus between the processing unit 110 and the programmable logic unit 120, and multiple peripherals 130.
[0024] The heterogeneous embedded system 100 is deployed on an airborne computing platform or a drone, and is used to achieve high real-time performance and low power consumption for single-target detection and tracking. In some implementations, the target to be detected and tracked can be a drone, a bird, a small flying object, a ground moving object, a specific marker, or a region of interest, etc.
[0025] The processing unit 110 integrates a multi-core ARM Cortex-A53 processor (based on the ARMv8-A instruction set architecture, supporting 64-bit and 32-bit execution states), a double data rate fourth generation synchronous dynamic random-access memory controller (DDR4 SDRAM controller), a peripheral controller, and an internal interconnect structure.
[0026] In one specific implementation, the DDR4 SDRAM controller is coupled to an external high-capacity DDR4 SDRAM memory to manage image frame data, intermediate calculation results, and memory read / write operations required for operating system operation.
[0027] Furthermore, the ARM Cortex-A53 multi-core processor in the processing unit 110 constitutes the general-purpose processing unit 111 of the heterogeneous embedded system 100, which has the ability to run a complete embedded operating system (e.g., Linux or FreeRTOS) and is responsible for executing irregular control tasks, task scheduling, human-computer interaction logic, and high-level decision-making algorithms.
[0028] In contrast to the general-purpose processor unit 111, the programmable logic unit 120 serves as a dedicated accelerator, focusing on performing computationally intensive and structurally regular signal processing operations. By deeply integrating the general-purpose processor unit 111 with the dedicated hardware accelerator, this application achieves an optimal balance between flexibility and efficiency.
[0029] In some implementations, the plurality of peripherals 130 may include at least one of Double Data Rate Fourth Generation Synchronous Dynamic Random-Access Memory (DDR4 SDRAM), a serial port, a Secure Digital Memory Card (SD card), a Universal Serial Bus Hub (USB Hub), an Ethernet interface, and a display.
[0030] In some implementations, the processing unit 110 is also connected to a Secure Digital Memory Card (SD card), a Universal Serial Bus Hub (USB Hub), an Ethernet interface, and a display port.
[0031] In one specific implementation, the SD card is used to store the boot configuration file of the heterogeneous embedded system 100 and the root file system of the embedded operating system. A Universal Serial Bus (USB) hub is used to expand the human-machine interface peripherals or mobile storage devices of the heterogeneous embedded system 100, such as a mouse, keyboard, and removable storage media. An Ethernet interface is used to access a network to support remote monitoring and command issuance of the heterogeneous embedded system 100. A display port is used to output tracking results and the status of the heterogeneous embedded system 100 to an external display.
[0032] In some implementations, the programmable logic unit 120 comprises a large number of reconfigurable logic resources and embeds multiple hardware intellectual property (IP) cores. Specifically, the programmable logic unit 120 integrates computationally intensive modules from a kernelized correlation filter (KCF) algorithm in the form of IP cores. These include histogram of oriented gradients (FHOG) feature extraction, fast Fourier transform (FFT / IFFT), kernel correlation calculation, and online filter updates—mapped into dedicated hardware IP cores after high-level synthesis (HLS) and deployed within the programmable logic unit 120 for parallel acceleration.
[0033] The processing unit 110 and the programmable logic unit 120 communicate efficiently via the AXI (Advanced eXtensible Interface) bus. The control channel uses the AXI-Lite protocol. The processing unit 110 sends configuration instructions to the programmable logic unit 120 through the AXI-Master port and reads the tracking status and results. The high-performance AXI-HP (High Performance) interface is used to enable the processing unit 110 to transmit the image sub-window data to be processed to the programmable logic unit 120 at high speed, and the programmable logic unit 120 to transmit the target position coordinates back to the processing unit 110, ensuring that the data throughput meets the real-time tracking requirements of more than 30 FPS.
[0034] Through the aforementioned hardware and software collaborative architecture, general task scheduling and operating system management are handled by the processing unit 110, while the highly parallel and rule-based KCF core computing is offloaded to the programmable logic unit 120 for execution. This significantly improves the real-time performance and robustness of target tracking under limited onboard power consumption and size conditions, effectively solving the technical bottleneck of the traditional solution where it is difficult to balance computing power, power consumption and real-time performance.
[0035] Furthermore, its onboard online processing capabilities eliminate errors caused by manual operation, continuously providing accurate and stable tracking results, and freeing it from the limitations of ground-based remote control, enabling a larger operating radius. Therefore, it can be widely used in fields such as aerial video recording, outdoor rescue, traffic management, and military reconnaissance.
[0036] In other words, this application selects an algorithm with low computational complexity that meets the accuracy requirements for airborne ground target tracking (i.e., the KCF algorithm), and implements a suitable computing platform (i.e., a heterogeneous embedded system 100) for parallel acceleration of the selected algorithm to meet real-time requirements. Specifically: Please see Figure 2 , Figure 2 This diagram illustrates a flowchart of a single-target tracking method for unmanned aerial vehicles (UAVs) according to an embodiment of this application, which can be applied to the aforementioned heterogeneous embedded system. For example... Figure 2 As shown, the method may include steps 210 to 240.
[0037] In step 210, a video stream containing the object to be tracked is acquired.
[0038] The object to be tracked can be a target entity initially specified by the user or automatically identified by the front-end detection module in a video sequence, which has distinguishable visual features (e.g., shape, texture, color, or motion pattern) and is suitable for target tracking algorithms based on correlation filtering.
[0039] In one specific implementation, the objects to be tracked include, but are not limited to: ground vehicles, pedestrians, drones, animals, or specific markers.
[0040] In some implementations, a video stream containing the object to be tracked is acquired using an image acquisition device. This image acquisition device can be a visible light camera, an infrared imager, a multispectral sensor, or a combination thereof. The image acquisition device is mounted on the drone's fuselage and is used to capture continuous image frames containing the object to be tracked in real time, forming a video stream.
[0041] Since the raw video frames acquired by the image acquisition device in a dynamic environment often contain sensor noise, directly inputting them into subsequent models will lead to gradient feature distortion, thereby causing positioning drift. Therefore, denoising processing is required first to ensure the reliability of subsequent steps (e.g., feature extraction). Specifically, in some embodiments, after the step of "acquiring a video stream containing the object to be tracked", the UAV single-target tracking method further includes: preprocessing the acquired video stream by a general-purpose processor unit to obtain the final video stream.
[0042] In one specific implementation, the preprocessing of the acquired video stream by the general-purpose processor unit may include: denoising the acquired video stream (e.g., using Gaussian filtering or nonlocal mean denoising) to suppress sensor noise; performing brightness / contrast correction on the acquired video stream (e.g., adaptive histogram equalization) to adapt to changes in illumination; extracting the region of interest (ROI) from the acquired video stream to crop a sub-image window containing the object to be tracked based on the initial state information, thereby reducing the subsequent computational load, etc.
[0043] The preprocessed video stream is transmitted from the general-purpose processor unit to the programmable logic unit (PLU) via a high-performance AXI-HP bus interface, serving as input for KCF tracking. Simultaneously, initial state information, scale, and template information are configured by the general-purpose processor unit to the tracking accelerator within the PLU to initiate the tracking process. Further: In step 220, a preset deep neural network is run by a general-purpose processor unit to detect the target object contained in the target image frame in the video stream and determine the initial state information of the target object.
[0044] The preset deep neural network can be a convolutional neural network model that has been pre-trained on a large-scale object detection dataset. It is deployed in a general-purpose processor unit to locate and identify one or more candidate targets (i.e., objects to be tracked) from the input video stream.
[0045] In one specific implementation, the pre-defined deep neural network structure includes, but is not limited to, YOLO (YouOnly Look Once), SSD (Single Shot MultiBox Detector), or Faster R-CNN. It is understood that this application is not limited to these, and any deep detection model capable of outputting the target bounding box and class confidence score corresponding to the object to be tracked is applicable.
[0046] The target image frame can be the first frame of the video stream or any keyframe triggered by the user's tracking command.
[0047] In embodiments of this application, the initial state information includes at least the spatial location (represented by bounding box coordinates) and scale (represented by width and height) of the object to be tracked within the target image frame. This initial state information serves as the startup parameter for the KCF tracking algorithm in subsequent steps, used to initialize filter templates and search windows, etc. In some embodiments, the initial state information may also include the target category label, confidence score, or center point coordinates of the object to be tracked.
[0048] After receiving the first frame of the video stream (or any key frame triggered by the user's tracking command), the heterogeneous embedded system loads a preset deep neural network model by a general-purpose processor unit, performs forward inference calculations on the frame image, and outputs all detected target candidate boxes.
[0049] Subsequently, based on user interaction selection (e.g., clicking the screen to specify a target) or an automatic selection strategy (e.g., selecting the pedestrian / vehicle with the highest confidence), the object to be tracked is determined, and its corresponding bounding box parameters are extracted as initial state information. This initial state information is then configured to the KCF tracking accelerator in the programmable logic unit via the AXI-Lite bus to initiate the real-time tracking process for subsequent frames. Specifically: In step 230, the initial state information and video stream are input to the preset tracking model deployed in the programmable logic unit to determine the real-time state information of the object to be tracked in each image frame of the video stream.
[0050] The preset tracking model can be a target tracking algorithm hardware acceleration module that is pre-trained and deployed in a programmable logic unit. It is configured to receive the initial target state and continuous video frames, and output the updated state of the object to be tracked in each frame.
[0051] In one specific implementation, the preset tracking model can be based on correlation filtering, Siamese Network, attention mechanism or Transformer architecture, but is not limited to a specific algorithm type.
[0052] In one specific implementation, the preset tracing model can be a KCF accelerator optimized by High-Level Synthesis (HLS). In another specific implementation, the preset tracing model can also be a lightweight Siamese network inference engine embedded in a programmable logic unit or a Transformer-based tracing IP core.
[0053] Real-time status information includes at least the spatial location of the object to be tracked in the current image frame (e.g., bounding box center coordinates or vertex coordinates) and scale parameters (represented by width and height). In some implementations, real-time status information may also include target confidence score, response map peak, motion direction, or category consistency index.
[0054] After obtaining initial state information, the heterogeneous embedded system uses a general-purpose processor unit to transmit the region of interest (ROI) data of subsequent video frames to the programmable logic unit (PLU) via a high-performance AXI-HP bus. A preset tracking model performs feature extraction, template matching, or attention association calculations in parallel within the PLU. Combining the initial template or historical states, it dynamically estimates the optimal position and scale of the object to be tracked in the current frame. The calculation results are fed back to the processing unit via the AXI-Lite interface, forming a closed-loop feedback loop, and are used to update the search window for the next frame or trigger a re-detection mechanism.
[0055] In related technologies, the development of programmable logic units relies on hardware description languages such as Verilog / VHDL, which has a long development cycle and high barriers to entry. To improve development efficiency, this application uses High-Level Synthesis (HLS) tools to automatically convert the C++ reference code of the KCF algorithm into register-transfer-level (RTL) hardware circuitry. This process improves the performance and resource efficiency of heterogeneous embedded systems through baseline optimization, loop optimization, memory optimization, and dataflow optimization. Specifically: In baseline optimization, arbitrary precision data types (e.g., ap_int) are used. <12> This replaces standard C++ types, reducing the waste of register resources. Variable loop boundaries are marked using the LOOP_TRIPCOUNT instruction to aid in timing analysis.
[0056] In loop optimization, the Pipeline instruction is applied to critical computation loops to enable overlapping execution of operations between iterations, which significantly improves throughput.
[0057] In storage optimization, identify the bandwidth bottleneck array and use the ARRAY_PARTITION instruction to divide it into multiple parallel memory blocks, increasing the number of access ports.
[0058] In data flow optimization, the DATAFLOW instruction is used to implement function-level pipelining, enabling modules such as feature extraction, FFT, and correlation calculation to run in parallel, thereby reducing end-to-end latency.
[0059] After baseline optimization, loop optimization, storage optimization, and data flow optimization, the KCF accelerator can stably achieve real-time tracking performance of over 30 FPS in the programmable logic unit of Zynq UltraScale+MPSoC, meeting the stringent latency requirements of UAV onboard platforms.
[0060] It should be noted that the core idea of correlation filtering-based tracking algorithms (e.g., KCF) is to learn a discriminative filter template in the initial frame and then locate the target in subsequent frames based on the similarity response between this template and the candidate region. To improve computational efficiency, KCF utilizes the Fourier diagonalization property of the circulant matrix to transform dense correlation operations in the time domain into pointwise complex multiplications in the frequency domain, thereby significantly reducing computational complexity. Specifically, in some implementations, the step "inputting the initial state information and video stream into a preset tracking model deployed in a programmable logic unit to determine the real-time state information of the object to be tracked in each image frame of the video stream" may include the following steps: (1) Based on the initial state information, extract the first sub-window image corresponding to the initial state information from the image frame to be tracked; the image frame to be tracked is any image frame in the video stream other than the target image frame; (2) Based on the first sub-window image and the preset window function, extract the first multi-channel gradient direction histogram features, and perform fast Fourier transform on the first multi-channel gradient direction histogram features to obtain the first frequency domain features; (3) Multiply the first frequency domain feature by the preset Fourier domain filter coefficients point by point to generate the frequency domain response; the preset Fourier domain filter coefficients are stored in the programmable logic unit; (4) Perform a fast inverse Fourier transform on the first frequency domain response to obtain the spatial domain response map, and determine the real-time status information of the image frame to be tracked based on the location information of the maximum response value in the spatial domain response map.
[0061] The first sub-window image can be a local image region cropped from the current image frame, centered on the estimated position of the object to be tracked in the previous frame. Its size is determined according to the scale parameters (width and height) in the initial state information, and can be expanded according to a preset ratio to cover the movement range of the object to be tracked.
[0062] The first sub-window image is captured by the general-purpose processor unit: the general-purpose processor unit calculates the search center coordinates and window size of the current frame based on the real-time status information (or initial status information) output from the previous frame, extracts the corresponding pixel region from the complete video frame, and transmits the sub-window image data to the programmable logic unit as tracking input via the AXI-HP bus.
[0063] A preset window function can be used to smoothly weight the edges of the first sub-window image to reduce spectral leakage and improve the accuracy of the Fourier transform. The preset window function can include a Hanning window, a Hamming window, or a Gaussian window, and its coefficients are pre-configured and stored in the read-only memory of the programmable logic unit (PLU) during system initialization. Before feature extraction, the windowing module in the PLU multiplies the preset window function point-by-point with the first sub-window image to generate the windowed image data.
[0064] The first multi-channel gradient orientation histogram feature is the FHOG (Felzenszwalb-style Histogram of Oriented Gradients) feature, which is a local shape descriptor containing three-channel gradient information. The first multi-channel gradient orientation histogram feature can include: horizontal gradient component, vertical gradient component, and gradient magnitude.
[0065] Within the programmable logic unit (PLU), FHOG feature extraction is implemented by a dedicated hardware module: first, the x- and y-axis gradients of the first sub-window image are calculated using the Sobel or Scharr operator; then, the gradients are decomposed into multiple directional intervals (e.g., 9 directional bins), and the distribution of gradient magnitudes is statistically analyzed within each spatial unit, ultimately forming a multi-channel feature map. This process is executed entirely in parallel pipeline within the PLU, without intervention from any processing unit.
[0066] The first frequency domain feature can be a complex frequency domain representation obtained by performing a Fast Fourier Transform (FFT) on the FHOG feature map. To improve efficiency, a dedicated FFT IP core (e.g., a parallel pipeline structure based on the Cooley-Tukey algorithm) is integrated into the programmable logic unit (PLU) to support simultaneous two-dimensional FFT operations on multi-channel feature maps. The transformation result is stored in complex form (real and imaginary parts) in the on-chip buffer of the PLU for subsequent related calculations.
[0067] The preset Fourier domain filter coefficients can be frequency domain coefficients generated by FFT transformation of the KCF filter template trained by the general-purpose processor unit based on the target image frame during the system initialization phase. These preset Fourier domain filter coefficients are written by the processing unit to the dual-port general-purpose processor unit or register array inside the programmable logic unit via the AXI-Lite interface and remain unchanged during each subsequent frame tracking process (or are dynamically adjusted according to an online update strategy). During the relevant calculation phase, the programmable logic unit directly reads the preset Fourier domain filter coefficients from local memory and performs pointwise complex multiplication with the frequency domain features of the current frame.
[0068] The frequency domain response is the result of pointwise complex multiplication of the first frequency domain feature with the coefficients of a preset Fourier domain filter. Mathematically, it is equivalent to the cyclic correlation operation between the FHOG feature and the filter template in the time domain. The frequency domain response reflects the similarity between each position in the current sub-window and the target template, and its energy concentration region corresponds to the potential target position. Since it is calculated in the frequency domain, this operation only requires O(N) complex multiplications, which is much lower than the O(N²) complexity of time-domain convolution.
[0069] After performing an Inverse Fast Fourier Transform (IFFT) on the first frequency domain response, a spatial domain response map is obtained, where each pixel value represents the target matching score at the corresponding location. The programmable logic unit integrates an IFFT IP core and a peak detection module: the IFFT module transforms the first frequency domain response back into the spatial domain, and the peak detection module traverses the spatial domain response map to find the maximum response value and its coordinate offset (relative to the center of the sub-window). This offset, combined with the initial scale, allows the calculation of the absolute position and scale of the object to be tracked in the current frame, which is then output as real-time status information.
[0070] The spatial domain response map is a two-dimensional real-valued matrix generated by the inverse fast Fourier transform (IFFT) of the first frequency domain response, and its dimensions are consistent with those of the first sub-window image. The magnitude of each pixel in the first sub-window image reflects the degree of local similarity between that spatial location and the initial target template (i.e., the target image frame), and the peak region is the location where the target is most likely to exist.
[0071] The programmable logic unit integrates a dedicated peak detection module, which traverses all pixels in the spatial domain response map to find the coordinates of the pixel with the maximum response value. To improve positioning accuracy, the peak detection module can also employ sub-pixel interpolation strategies (e.g., quadratic function fitting or centroid method) to interpolate the neighborhood around the maximum response value, obtaining sub-pixel-level offsets, thereby improving tracking accuracy.
[0072] The above process fully utilizes the efficiency of correlation filtering algorithms in frequency domain calculations, transforming the originally complex time-domain convolution operation into point-by-point multiplication in the frequency domain, which greatly reduces computational overhead and allows the single-frame tracking processing time to be controlled within tens of milliseconds, meeting the real-time tracking requirements of high frame rates above 30 FPS.
[0073] Meanwhile, since FHOG features are highly robust to changes in illumination, background clutter, and texture degradation, and combined with a preset window function to effectively suppress spectral leakage, the accuracy of peak localization in the response map is further improved, thereby enhancing the stability and anti-interference capability of tracking.
[0074] More importantly, all the aforementioned computationally intensive operations are executed in parallel by dedicated hardware acceleration modules in the programmable logic unit (PLU). The general-purpose processing unit is only responsible for initialization, data scheduling, and result parsing. The two work together efficiently through the AXI bus, which not only relieves the computational pressure on the processing unit but also fully leverages the parallel processing advantages of the PLU. Under the constraints of limited onboard power consumption and size, it achieves an organic unity of high real-time performance, low power consumption, and strong robustness, effectively solving the technical bottleneck of the traditional solution where it is difficult to balance computing power, power consumption, and flexibility.
[0075] However, in practical applications, especially for high frame rate video streams or applications with extremely high real-time requirements, generating the frequency domain response efficiently and accurately is particularly important. Therefore, in some implementations, the step "multiplying the first frequency domain feature by the preset Fourier domain filter coefficients point-by-point to generate the frequency domain response" may include the following steps: (1) The conjugate filter coefficients are obtained by performing complex conjugate operation on the preset Fourier domain filter coefficients through the first hardware calculation module; (2) The first frequency domain features and the conjugate filter coefficients are multiplied point by point by the second hardware calculation module to obtain the intermediate frequency domain result; (3) By calling the DFT function in the advanced synthesis library through the third hardware computing module, the intermediate frequency domain result is subjected to inverse fast Fourier transform to obtain the spatial domain correlation value; (4) The Gaussian kernel response value is calculated by calling the exponential function in the advanced synthesis library based on the spatial domain correlation value through the fourth hardware computing module, and the Gaussian kernel response value is used as the first frequency domain response.
[0076] In order to improve throughput, at least two of the first, second, third and fourth hardware computing modules are configured to execute in parallel.
[0077] Kernel correlation calculations are implemented based on the Gaussian kernel function, whose mathematical expression is: in, Features of the current frame For the characteristics of the reference template, The bandwidth parameter of the Gaussian kernel function. This is the inverse discrete Fourier transform. To multiply point by point, for Frequency domain representation after Fast Fourier Transform (FFT) for Frequency domain representation after Fast Fourier Transform (FFT).
[0078] The above mathematical expression shows that kernel correlation calculation can be divided into the following steps: For the features of the current frame... and reference template Perform an FFT transform to obtain the frequency domain representation. and ;right Take conjugate and with Perform pointwise complex multiplication to obtain intermediate frequency domain results; perform IFFT on intermediate results to obtain spatial domain correlation values; substitute these correlation values into the exponential function to calculate the final Gaussian kernel response value.
[0079] Based on the above Gaussian kernel calculation process, in the heterogeneous architecture of this application, kernel-related calculations are completely offloaded to programmable logic units and implemented in parallel by a first hardware calculation module, a second hardware calculation module, a third hardware calculation module and a fourth hardware calculation module.
[0080] The first hardware computing module can be a dedicated complex number arithmetic unit deployed within a programmable logic unit (PLU), configured to perform complex conjugation operations on preset Fourier domain filter coefficients. Specifically, for each complex filter coefficient, the first hardware computing module converts it to its conjugation form, i.e., keeping the real part unchanged and inverting the sign of the imaginary part. This operation is implemented through a signed subtractor and registers within the PLU, with a delay of only one clock cycle.
[0081] The conjugate filter coefficients can be the result of performing complex conjugate operations on the original Fourier domain filter coefficients (that is, the frequency domain representation obtained by converting the time domain filter template learned by the KCF algorithm in the target image frame through fast Fourier transform, and stored in the programmable logic unit). They are used to implement cyclic correlation rather than convolution in the frequency domain to ensure that the peak value of the response map correctly corresponds to the target displacement.
[0082] The second hardware computing module can be a complex multiplication array integrated in a programmable logic unit, used to perform point-by-point complex multiplication of the first frequency domain features with the conjugate filter coefficients. The second hardware computing module is implemented in parallel by multiple DSP slices, supporting simultaneous processing of multi-channel features, and the output is the "intermediate frequency domain result".
[0083] The intermediate frequency domain result can be a complex matrix obtained by multiplying the first frequency domain features by the conjugate filter coefficients. Mathematically, it corresponds to the cyclic correlation spectrum of the feature map and the filter template in the time domain and is the intermediate data for generating the final response map.
[0084] High-level synthesis libraries can be standard function libraries provided by High-Level Synthesis (HLS) tools (e.g., the built-in libraries of Xilinx Vitis HLS or Intel HLS Compiler). These contain optimized mathematical operation IPs, such as Fast Fourier Transform (FFT), exponential functions, trigonometric functions, etc. These functions can be called directly from C++ code and are automatically mapped to efficient RTL circuits by HLS tools.
[0085] The DFT function can be the Inverse Fast Fourier Transform (IFFT) function, which is a standard signal processing component in the high-level synthesis library.
[0086] The third hardware computing module calls the DFT function to perform a two-dimensional IFFT operation on the intermediate frequency domain result, converting it back from the frequency domain to the spatial domain. This IFFT module is based on the Cooley-Tukey algorithm, adopts a pipelined architecture, supports parallel input and output, and outputs real or complex "spatial domain correlation values".
[0087] The spatial domain correlation value can be a two-dimensional matrix generated by IFFT, where each element represents the linear correlation strength between the corresponding position in the sub-window and the target template, without undergoing nonlinear mapping by the kernel function.
[0088] The fourth hardware computing module can be a nonlinear function calculation unit within the programmable logic unit. It calls exponential functions (e.g., exp()) from the high-level synthesis library to calculate the Gaussian kernel response value for each element of the spatial domain correlation value. This operation implements the nonlinear mapping of the Gaussian kernel, and the output is the "Gaussian kernel response value," which has a sharper peak and higher positioning accuracy.
[0089] The Gaussian kernel response value can be the final frequency domain response (the spatial domain response after IFFT), which can be used for subsequent peak detection.
[0090] Through the computation process of the first hardware computing module, the second hardware computing module, the third hardware computing module and the fourth hardware computing module, this application efficiently implements the Gaussian kernel related calculation in the KCF algorithm in a programmable logic unit, achieving significant technical results.
[0091] In other words, by mapping complex conjugate operations, complex multiplication, IFFT transformation, and exponential function calculation to the first to fourth hardware computing modules respectively, the kernel-related operations that originally required hundreds of milliseconds to complete on a general-purpose processor are completely offloaded to the programmable logic unit for parallel execution, reducing the single-frame related calculation time to the microsecond level.
[0092] Secondly, each module calls optimization functions from the advanced synthesis library and combines pipeline and data flow optimization strategies, which not only ensures the calculation accuracy of the Gaussian kernel response value, but also significantly improves the throughput.
[0093] More importantly, by solidifying intensive operations such as "pointwise complex multiplication" and "exponential function calculation" into dedicated hardware circuits, frequent intervention at the processing unit level is avoided, effectively reducing system power consumption and communication overhead.
[0094] Ultimately, on a heterogeneous embedded system, this design enabled the entire KCF tracking process (including FHOG feature extraction, FFT / IFFT, kernel correlation, and peak detection) to run stably at over 30 FPS, meeting the stringent requirements of UAV airborne platforms for high real-time performance and low latency target tracking.
[0095] It should be noted that although the Gaussian kernel correlation calculation described above has high accuracy, in some application scenarios with extremely high real-time requirements or small changes in the target's appearance, a simplified linear correlation model can be used to further reduce computational complexity. Based on this, in some implementations, the step "multiplying the first frequency domain feature by the preset Fourier domain filter coefficients point-by-point to generate the frequency domain response" may include the following steps: (1) The conjugate filter coefficients are obtained by performing complex conjugate operation on the preset Fourier domain filter coefficients through the fifth hardware calculation module; (2) The first frequency domain feature and the conjugate filter coefficients are matched one by one according to their frequency domain positions through the sixth hardware computing module, and the multiple independent hardware multiplier units generated by the high-level synthesis library are used to perform point-by-point complex multiplication on each pair of complex elements in parallel to generate the frequency domain response.
[0096] Among them, the fifth and sixth hardware computing modules are configured to execute in parallel.
[0097] The core formula for "fast localization" (i.e., target prediction) in the KCF algorithm can be: in, For reference template, For the features of the current frame, Features of the current frame With reference template The Gaussian kernel correlation result in the frequency domain is obtained by multiplying the result in the frequency domain using the inverse discrete Fourier transform (IFFT). These are the regression coefficients output by the training module.
[0098] The above formula shows that the "fast localization" in the KCF algorithm can be broken down into: using the reference template... Frequency domain representation As preset Fourier domain filter coefficients, they are subjected to complex conjugate operations by the fifth hardware calculation module to obtain conjugate filter coefficients; the current frame features are then used to... Frequency domain representation and Perform pointwise complex multiplication to generate intermediate frequency domain results. .
[0099] Based on the "fast location" calculation process in the KCF algorithm described above, in the heterogeneous architecture of this application, kernel-related calculations are completely offloaded to the programmable logic unit and implemented in parallel by the fifth hardware computing module and the sixth hardware computing module.
[0100] The fifth hardware computing module can be a dedicated complex number arithmetic circuit deployed in a programmable logic unit (PLU). Its function is to perform complex conjugate operations on preset Fourier domain filter coefficients. Specifically, the fifth hardware computing module stores the filter coefficients in complex number form (the real and imaginary parts are represented as signed fixed-point numbers, respectively). Through registers and a signed subtractor in the PLU, the imaginary part of each complex number is inverted while the real part remains unchanged, thereby generating conjugate filter coefficients. This operation is completed within a single clock cycle with extremely low latency, meeting the real-time requirements of the KCF tracing process.
[0101] The sixth hardware computing module can be a dedicated parallel complex multiplication array deployed in a programmable logic unit. Its function is to map the first frequency domain features and the conjugate filter coefficients to each other in the frequency domain according to their positions, and to perform point-by-point complex multiplication synchronously on each pair of complex elements to generate a frequency domain response.
[0102] Specifically, the sixth hardware computing module is automatically generated by the High-Level Synthesis (HLS) tool based on the C++ algorithm description. It calls the complex number operation component (e.g., hls::complex) in the high-level synthesis library and utilizes the DSPSlice hardware resources in the programmable logic unit to construct multiple independent hardware multiplier units. These multiplier units are fully parallelized according to the frequency domain grid (e.g., 32×32 or 64×64), so that all points in the entire frequency domain feature map can complete the multiplication operation simultaneously without the need for loop iteration.
[0103] Since all calculations are performed within the programmable logic unit, no general-purpose processing unit is required, significantly reducing data transfer overhead and processing latency, thus meeting the real-time requirements of UAV airborne platforms for high frame rate tracking.
[0104] Therefore, by mapping complex conjugate operations and pointwise complex multiplication to dedicated hardware circuits, frequency domain operations that originally needed to be executed serially on a general-purpose processor (PS) are completely offloaded to the programmable logic unit for parallel processing, reducing the latency of a single correlation calculation to the microsecond level.
[0105] Secondly, the sixth hardware computing module automatically generates multiple independent hardware multiplier units using an advanced synthesis library, performing fully parallel complex multiplication on all positions of the frequency domain feature map, fully exploiting the spatial parallelism capability of the programmable logic unit, and improving throughput by tens of times compared to the software implementation.
[0106] Ultimately, through the cooperation between the fifth and sixth hardware computing modules, the KCF tracking process meets the stringent requirements of the UAV airborne platform for high real-time performance and low latency target tracking.
[0107] Although the fifth and sixth hardware computing modules have achieved efficient target localization, the appearance of the tracked object may drift significantly due to scale changes, rotation, or occlusion during long-term tracking. To improve tracking robustness, this application further introduces an online filter update mechanism. Specifically, in some embodiments, the UAV single-target tracking method further includes the following steps: (1) Based on the real-time status information of the image frame to be tracked, extract the second sub-window image corresponding to the real-time status information in the image frame to be tracked; (2) Based on the second sub-window image and the preset window function, extract the second multi-channel gradient direction histogram features, and perform fast Fourier transform on the second multi-channel gradient direction histogram features to obtain the second frequency domain features; (3) Based on the second frequency domain features and the preset Gaussian shape regression label, generate the frequency domain filter coefficients corresponding to the image frame to be tracked; (4) Linearly interpolate the frequency domain filter coefficients corresponding to the image frame to be tracked with the historical filter coefficients to obtain the updated frequency domain filter coefficients; (5) Use the updated frequency domain filter coefficients as the preset Fourier domain filter coefficients corresponding to the next frame of the image to be tracked.
[0108] Based on the real-time status information output in the current frame (including the target center coordinates and scale), a local region is cropped from the current image frame to be tracked, centered on that location, as the second sub-window image. Its size is determined by the current scale parameters (width and height) and can be expanded according to a preset ratio to cover the possible range of motion of the target.
[0109] Edge weighting is applied to the second sub-window image using a preset window function, followed by extraction of the second multi-channel gradient orientation histogram (FHOG) feature. The second multi-channel gradient orientation histogram feature includes three channels: horizontal gradient, vertical gradient, and gradient magnitude. Next, a Fast Fourier Transform (FFT) is performed on the FHOG feature of each channel to obtain the corresponding second frequency domain feature.
[0110] Based on the second frequency domain features and a preset Gaussian shape regression label (which has a Gaussian peak at the center of the target and decays at other positions), the frequency domain filter coefficients corresponding to the current frame are calculated using the frequency domain ridge regression formula. This calculation is performed by a dedicated hardware module in PL, avoiding matrix inversion operations.
[0111] The updated frequency domain filter coefficients are stored in the on-chip dual-port RAM or register array of the programmable logic unit and used as the preset Fourier domain filter coefficients for tracking in the next frame. The fifth and sixth hardware calculation modules in subsequent frames will directly read the updated coefficients for relevant calculations.
[0112] Through the aforementioned online filter update mechanism, this invention implements a complete closed-loop filter update mechanism within the programmable logic unit. This mechanism not only inherits the low-latency advantage of the main tracking process but also significantly improves the system's adaptability to changes in target appearance, effectively solving the model drift problem in long-term tracking, while simultaneously meeting the comprehensive requirements of UAV airborne platforms for high real-time performance, low latency, and low power consumption.
[0113] Furthermore, in some implementations, the step "generating frequency domain filter coefficients corresponding to the image frame to be tracked based on the second frequency domain features and the preset Gaussian shape regression label" may also include the following steps: (1) The autocorrelation of the second frequency domain features is calculated by the seventh hardware computing module to obtain the first frequency domain representation of the cross-correlation matrix; (2) The preset Gaussian shape regression label is converted into a second frequency domain representation through the eighth hardware computing module; (3) Through the ninth hardware computing module, based on the first frequency domain representation, the second frequency domain representation and the regularization parameter, multiple independent hardware divider units perform point-by-point complex division in parallel to obtain the frequency domain filter coefficients.
[0114] Among them, at least two of the seventh, eighth and ninth hardware computing modules are configured to execute in parallel to fully utilize the parallel processing capabilities of the FPGA and achieve the technical goal of "using its parallelism to achieve high-speed computing" in the document.
[0115] This application transfers the online training process of the KCF algorithm from a general-purpose processor (CPU) to a programmable logic unit (PLU), leveraging the parallel computing capabilities of the PLU to achieve high-speed training, thereby meeting the stringent real-time tracking requirements of UAV onboard systems. The core of the training process is the calculation of regression coefficients. Its mathematical expression can be: in, The frequency domain representation of the pre-defined Gaussian shape regression label has its peak located at the center of the target, reflecting the ideal response distribution. Features of the current frame The frequency domain representation of its complex conjugate cross-correlation matrix. This is a second frequency domain feature. This is the regularization parameter.
[0116] Based on the above formula, in the heterogeneous architecture of this application, The computation is mapped to be performed by a dedicated hardware module. Specifically: the second frequency domain features are processed through the seventh hardware computation module. Perform complex conjugate operation to obtain Then it was combined with By performing point-by-point complex multiplication according to the one-to-one correspondence of frequency domain positions, the following is generated: This is the first frequency domain representation of the cross-correlation matrix. The frequency domain representation is obtained by performing a Fast Fourier Transform (FFT) on the preset Gaussian shape regression labels using the eighth hardware computing module. Through the ninth hardware computing module, according to , and The complex division function in the advanced synthesis library is called, and multiple independent hardware divider units perform point-by-point complex division in parallel to obtain the regression coefficients. , which are the frequency domain filter coefficients.
[0117] In some implementations, the step "linearly interpolating the frequency domain filter coefficients corresponding to the image frame to be tracked with the historical filter coefficients to obtain the updated frequency domain filter coefficients" may include the following steps: (1) Through the tenth hardware computing module, the frequency domain filter coefficients and historical filter coefficients corresponding to the image frame to be tracked are represented as real part matrices and imaginary part matrices, respectively; (2) Through the eleventh hardware calculation module, according to the preset exponential smoothing formula, and by calling multiple independent hardware adder units, the real part matrix and the imaginary part matrix are respectively subjected to weighted summation operation to obtain the calculation result; (3) The calculation results are merged through the twelfth hardware calculation module to generate updated frequency domain filter coefficients.
[0118] Among them, at least two of the tenth, eleventh and twelfth hardware computing modules are configured to execute in parallel.
[0119] This application represents the update process of the KCF algorithm's model parameters as follows: in, These are the updated regression coefficients. These are the preset Fourier domain filter coefficients currently in use (i.e., historical filter coefficients). For the new regression coefficients generated based on the current frame, The learning rate is preset to control the fusion ratio of the old and new models.
[0120] The above formula (i.e., the preset exponential smoothing formula) can be viewed as a simple weighted addition operation, that is, a linear combination of two complex matrices element by element. Since the matrices involved are all complex matrices, the real and imaginary parts must be calculated separately according to the rules of complex number operations. This operation has high parallelism and can be completed simultaneously on an FPGA using multiple independent hardware adder units, thereby achieving high-efficiency acceleration.
[0121] Based on the above formula, in the heterogeneous architecture of this invention, the model parameter update process is mapped to be executed by a dedicated hardware module. Specifically: the historical filter coefficients are processed by the tenth hardware computing module. The newly generated regression coefficients are processed through the eleventh hardware computing module. Multiply by ( The twelfth hardware computing module performs a point-by-point complex summation of the two sets of results to obtain the updated regression coefficients. .
[0122] The aforementioned model parameter update mechanism relies on the accurate generation of regression coefficients and, in turn, on high-quality second multi-channel gradient direction histogram features. Similarly, the first frequency domain features in the main tracking process also need to be extracted from the first sub-window image using FHOG features. To ensure the accuracy and efficiency of feature extraction, this application implements a unified image preprocessing and gradient calculation pipeline in a programmable logic unit. Specifically, in some embodiments, this UAV single-target tracking method further includes the following steps: (1) In the programmable logic unit, if the size of the first sub-window image or the second sub-window image is inconsistent with the size of the preset Hanning window, the size of the first sub-window image or the second sub-window image is adjusted by the image processing function in the advanced synthesis library to obtain the adjusted first sub-window image or the second sub-window image. (2) Perform convolution operation on the adjusted first sub-window image or second sub-window image using the two-dimensional convolution filtering function in the advanced synthesis library to determine the horizontal and vertical gradient components of each pixel in the adjusted first sub-window image or second sub-window image; (3) Calculate the horizontal and vertical gradient components of each pixel using the square root calculation function in the advanced synthesis library to obtain the gradient magnitude, and determine the gradient direction based on the horizontal and vertical gradient components of each pixel. (4) Generate the first multi-channel gradient direction histogram feature or the second multi-channel gradient direction histogram feature based on the gradient magnitude and gradient direction.
[0123] In the programmable logic unit, if the image of the first or second sub-window is inconsistent with the size of the preset Hanning window, its size is adjusted using the image processing function hls::Resize in the high-level synthesis library to obtain the adjusted image. This function has been pre-optimized by HLS, supports parallel execution, and can perform image scaling with high quality and high performance.
[0124] The adjusted image is convolved using the 2D convolution filtering function hls::Filter2D from the advanced synthesis library, with predefined gradient operators [-1,0,1] and... Calculate the horizontal gradient component for each pixel. and vertical gradient components .
[0125] The horizontal and vertical gradient components of each pixel are calculated using the square root calculation function hls::sqrtf from the advanced synthesis library, yielding the gradient magnitude. and according to and Determine the gradient direction .
[0126] Based on the gradient magnitude and gradient direction, generate either a first multi-channel gradient direction histogram feature or a second multi-channel gradient direction histogram feature.
[0127] Each of the above steps calls the optimized library functions provided by HLS to ensure high-throughput, low-latency parallel computing on the FPGA.
[0128] In summary, this application constructs a fully hardware-based high-speed target tracking architecture by offloading the core processes of the KCF tracking algorithm—including sub-window truncation, FHOG feature extraction, FFT / IFFT transformation, frequency domain correlation calculation, Gaussian kernel response generation, and online filter training and updating—to a programmable logic unit.
[0129] All computation modules are implemented based on High-Level Synthesis (HLS) library functions, making full use of the FPGA's parallel processing capabilities: image resizing uses hls::Resize; gradient calculation uses hls::Filter2D and hls::sqrtf; frequency domain operations are performed in parallel by dedicated hardware units such as complex conjugation, pointwise multiplication, and complex division; filter updates achieve model adaptation through linear interpolation.
[0130] The entire system requires no PS intervention, avoiding frequent data transfer and control overhead, and significantly reducing power consumption while meeting high real-time and low latency requirements. Ultimately, by "transferring the online training and detection process of KCF from the CPU to the FPGA and utilizing its parallelism to achieve high-speed computing," the technical bottlenecks of "insufficient computing power and excessive power consumption" in traditional software implementations are effectively solved, providing an efficient, robust, and low-power single-target tracking solution for UAV airborne platforms.
[0131] In step 240, the spatial motion state of the UAV is adjusted according to the offset of the real-time status information so that the object to be tracked remains in the center of the field of view.
[0132] In order to achieve dynamic control of the drone's flight status, the system calculates the positional offset of the object to be tracked in the image based on the real-time status information in the tracking results, and adjusts the spatial motion state of the drone accordingly to ensure that the object to be tracked remains in the center of the camera's field of view.
[0133] Specifically, the target center coordinates, determined by inverse Fourier transform and peak detection, are first obtained and compared with the geometric center of the image frame to obtain the pixel offsets in the horizontal and vertical directions. Then, based on the pre-calibrated proportional control relationship, the horizontal offset is converted into an angular velocity command to control the UAV's rotation around the vertical axis to adjust the heading, and the vertical offset is converted into an angular velocity command to control the UAV's pitch around the horizontal axis to adjust the camera's pitch angle. These commands are sent to the flight controller in real time through the onboard high-speed serial communication interface. After receiving the commands, the flight controller adjusts the speed of each rotor motor to drive the UAV to produce corresponding yaw or pitch movements, thereby changing the pointing of the onboard camera.
[0134] Based on this, the system repeatedly performs target localization, offset calculation, command generation and attitude adjustment operations within each video frame cycle, forming a closed-loop control process of "visual tracking - deviation detection - flight correction - re-tracking", which effectively ensures that the target remains stably located in the center of the field of view during long-term tracking, significantly improving the robustness and practicality of UAV autonomous tracking.
[0135] Please see Figure 3 , Figure 3 This illustration shows a schematic diagram of another heterogeneous embedded system provided in an embodiment of this application. The heterogeneous embedded system includes a general-purpose processor unit and a programmable logic unit. The heterogeneous embedded system 300 includes: a data acquisition module 310 and an execution module 320. Specifically: Acquisition module 310 is used to acquire video streams containing objects to be tracked; The general-purpose processor unit 111 is used to run a preset deep neural network to detect the target object contained in the target image frame in the video stream and determine the initial state information of the target object. The programmable logic unit 120 is used to determine the real-time status information of the object to be tracked in each image frame of the video stream based on the initial state information and the video stream. The execution module 320 is used to adjust the spatial motion state of the UAV based on the offset of the real-time status information so that the object to be tracked remains in the center of the field of view.
[0136] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0137] In the several embodiments provided in this application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interface, device or module, and may be electrical, mechanical or other forms.
[0138] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0139] Please see Figure 4 , Figure 4 The diagram shows a structural schematic of a cooking device provided in an embodiment of this application. The drone 400 in this application may include one or more of the following components: processor 410, memory 420, and one or more applications, wherein the one or more applications may be stored in memory 420 and configured to be executed by one or more processors 410, and the one or more applications are configured to perform the drone single target tracking method as described in the foregoing method embodiments.
[0140] Processor 410 may include one or more processing cores. Processor 410 connects to various parts within the UAV 400 via various interfaces and lines, executing instructions, programs, code sets, or instruction sets stored in memory 420, and calling data stored in memory 420 to perform various functions and process data of the UAV 400. Optionally, processor 410 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 410 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 410 and may be implemented separately through a communication chip.
[0141] The memory 420 may include random access memory (RAM) or read-only memory (ROM). The memory 420 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 420 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the various method embodiments described below, etc. The data storage area may also store data created by the drone 400 during use.
[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A single-target tracking method for unmanned aerial vehicles (UAVs), characterized in that, Applied to heterogeneous embedded systems, the heterogeneous embedded systems including general-purpose processor units and programmable logic units, the method includes: Acquire a video stream containing the object to be tracked; The general-purpose processor unit runs a preset deep neural network to detect the target object contained in the target image frame in the video stream and determine the initial state information of the target object. The initial state information and the video stream are input to a preset tracking model deployed in the programmable logic unit to determine the real-time state information of the object to be tracked in each image frame of the video stream; Based on the offset of the real-time status information, the spatial motion state of the UAV is adjusted so that the object to be tracked remains in the center of the field of view.
2. The UAV single-target tracking method according to claim 1, characterized in that, The step of inputting the initial state information and the video stream into a preset tracking model deployed in the programmable logic unit to determine the real-time state information of the object to be tracked in each image frame of the video stream includes: Based on the initial state information, a first sub-window image corresponding to the initial state information is extracted from the image frame to be tracked; the image frame to be tracked is any image frame in the video stream other than the target image frame. Based on the first sub-window image and the preset window function, the first multi-channel gradient direction histogram features are extracted, and the first multi-channel gradient direction histogram features are subjected to fast Fourier transform to obtain the first frequency domain features. The first frequency domain feature is multiplied point-by-point by the preset Fourier domain filter coefficients to generate the frequency domain response; the preset Fourier domain filter coefficients are stored in the programmable logic unit. Perform an inverse fast Fourier transform on the first frequency domain response to obtain a spatial domain response map, and determine the real-time status information of the image frame to be tracked based on the location information of the maximum response value in the spatial domain response map.
3. The UAV single-target tracking method according to claim 2, characterized in that, The method further includes: Based on the real-time status information of the image frame to be tracked, a second sub-window image corresponding to the real-time status information is extracted from the image frame to be tracked; Based on the second sub-window image and the preset window function, the second multi-channel gradient direction histogram features are extracted, and the second multi-channel gradient direction histogram features are subjected to fast Fourier transform to obtain the second frequency domain features. Based on the second frequency domain features and the preset Gaussian shape regression label, the frequency domain filter coefficients corresponding to the image frame to be tracked are generated; The frequency domain filter coefficients corresponding to the image frame to be tracked are linearly interpolated with the historical filter coefficients to obtain the updated frequency domain filter coefficients. The updated frequency domain filter coefficients are used as the preset Fourier domain filter coefficients corresponding to the next frame of the image to be tracked.
4. The UAV single-target tracking method according to claim 2 or 3, characterized in that, The method further includes: In the programmable logic unit, if the size of the first sub-window image or the second sub-window image is inconsistent with the size of the preset Hanning window, the size of the first sub-window image or the second sub-window image is adjusted by the image processing function in the advanced synthesis library to obtain the adjusted first sub-window image or the second sub-window image. The horizontal and vertical gradient components of each pixel in the adjusted first sub-window image or second sub-window image are determined by performing a convolution operation on the two-dimensional convolution filtering function in the advanced synthesis library. The horizontal and vertical gradient components of each pixel are calculated using the square root calculation function in the advanced synthesis library to obtain the gradient magnitude, and the gradient direction is determined based on the horizontal and vertical gradient components of each pixel. Based on the gradient magnitude and the gradient direction, generate a first multi-channel gradient direction histogram feature or a second multi-channel gradient direction histogram feature.
5. The UAV single-target tracking method according to claim 2, characterized in that, The step of multiplying the first frequency domain feature by the preset Fourier domain filter coefficients point-by-point to generate the frequency domain response includes: The first hardware computing module performs a complex conjugate operation on the preset Fourier domain filter coefficients to obtain the conjugate filter coefficients. The first frequency domain feature is multiplied point-by-point by the conjugate filter coefficients using the second hardware computing module to obtain the intermediate frequency domain result. The third hardware computing module calls the DFT function in the advanced synthesis library to perform an inverse fast Fourier transform on the intermediate frequency domain result to obtain the spatial domain correlation value. The fourth hardware computing module calculates the Gaussian kernel response value based on the spatial domain correlation value by calling the exponential function in the advanced synthesis library, and uses the Gaussian kernel response value as the first frequency domain response. Among them, at least two of the first hardware computing module, the second hardware computing module, the third hardware computing module and the fourth hardware computing module are configured to execute in parallel.
6. The UAV single-target tracking method according to claim 2, characterized in that, The step of multiplying the first frequency domain feature by the preset Fourier domain filter coefficients point-by-point to generate the frequency domain response includes: The fifth hardware computing module performs a complex conjugate operation on the preset Fourier domain filter coefficients to obtain the conjugate filter coefficients. The first frequency domain feature and the conjugate filter coefficients are mapped one-to-one according to their frequency domain positions through the sixth hardware computing module, and the frequency domain response is generated by performing point-by-point complex multiplication on each pair of complex elements in parallel through multiple independent hardware multiplier units generated by the high-level synthesis library. The fifth hardware computing module and the sixth hardware computing module are configured to execute in parallel.
7. The UAV single-target tracking method according to claim 3, characterized in that, The step of generating the frequency domain filter coefficients corresponding to the image frame to be tracked based on the second frequency domain features and the preset Gaussian shape regression label includes: The autocorrelation of the second frequency domain feature is calculated by the seventh hardware computing module to obtain the first frequency domain representation of the cross-correlation matrix; The preset Gaussian shape regression label is converted into a second frequency domain representation through the eighth hardware computing module; Through the ninth hardware computing module, based on the first frequency domain representation, the second frequency domain representation and the regularization parameter, multiple independent hardware divider units perform point-by-point complex division in parallel to obtain the frequency domain filter coefficients. Among them, at least two of the seventh hardware computing module, the eighth hardware computing module and the ninth hardware computing module are configured to execute in parallel.
8. The UAV single-target tracking method according to claim 3, characterized in that, The step of linearly interpolating the frequency domain filter coefficients corresponding to the image frame to be tracked with the historical filter coefficients to obtain the updated frequency domain filter coefficients includes: The tenth hardware computing module represents the frequency domain filter coefficients and the historical filter coefficients corresponding to the image frame to be tracked as real matrix and imaginary matrix, respectively. The eleventh hardware computing module performs a weighted summation operation on the real part matrix and the imaginary part matrix respectively, according to a preset exponential smoothing formula and by calling multiple independent hardware adder units, to obtain the calculation result. The twelfth hardware computing module merges the calculation results to generate the updated frequency domain filter coefficients. Among them, at least two of the tenth hardware computing module, the eleventh hardware computing module and the twelfth hardware computing module are configured to execute in parallel.
9. A heterogeneous embedded system, characterized in that, include: The acquisition module is used to acquire video streams containing the objects to be tracked. A general-purpose processor unit is used to run a preset deep neural network to detect the target object contained in the target image frame in the video stream and determine the initial state information of the target object. A programmable logic unit is used to determine the real-time status information of the object to be tracked in each image frame of the video stream based on the initial state information and the video stream. The execution module is used to adjust the spatial motion state of the UAV according to the offset of the real-time status information, so that the object to be tracked remains in the center of the field of view.
10. A drone, characterized in that, include: One or more processors; Memory; One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the one or more processors, the one or more applications being configured to perform the UAV single target tracking method as described in any one of claims 1-8.