Target detection and real-time tracking control method based on FPGA+GPU heterogeneous cooperation
Patent Information
- Application Number
- CN202611009826.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]目前实现目标检测与实时跟踪的主流技术路径有基于通用GPU的计算方案、基于FPGA的专用计算方案及基于FPGA+GPU异构协同的计算方案,其中,基于FPGA+GPU异构协同的计算方案利用FPGA完成数据采集、预处理和接口转换等前端任务,利用GPU执行深度学习推理等后端计算任务,实现任务的合理拆分与协同加速;现有的FPGA+GPU异构方案大多采用简单的任务级并行划分,由FPGA完成数据预处理,GPU完成深度学习推理;这种划分方式未能充分发挥两种计算资源的互补优势,FPGA的并行计算能力在完成预处理后即处于闲置或低负载状态,GPU则需独自承担所有与目标状态估计相关的计算任务,导致GPU计算负载不均衡,整体处理延迟仍然较高;同时,现有方案没有考虑FPGA+GPU异构方案中固定的流水线延迟对实时控制的影响,将估计出的目标状态直接用于控制指令生成,导致控制指令所对应的目标状态与实际物理世界中的目标位置之间存在时间错配,在高速运动场景下会产生显著的控制误差;因此,如何根据粒子滤波各子任务的计算特性进行精细化硬件映射并对固有流水线延迟进行补偿是本发明要解决的根本问题之一
本发明将数据采集、图像预处理、点云采样等低延迟、高确定性需求的任务部署于FPGA,利用FPGA的硬件并行性和确定性时延特性,实现了微秒级的多传感器时间同步和毫秒级的数据预处理,满足实时系统对前端数据处理的确定性要求;同时,将深度学习目标检测、点云特征提取、多模态似然计算等计算密集、高度并行的任务部署于GPU,利用GPU强大的浮点计算能力和对深度学习框架的原生支持,实现了高精度的目标检测与跟踪;此外,本发明通过对异构系统进行延迟补偿,实现了目标状态估计与物理世界真实状态的时空同步,使后续模型预测控制指令所基于的目标状态与实际物理世界中的目标位置在时间上严格对齐。
Smart Images

Figure CN122821094A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of target detection and tracking technology, and in particular to a target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration. Background Technology
[0002] Target detection and real-time tracking control are widely used in autonomous driving, industrial automation, drone navigation, intelligent security and medical imaging. In these application scenarios, the system needs to complete the rapid detection, continuous tracking and precise control of multiple moving targets under strict time constraints, which places extremely high demands on the real-time performance, accuracy and stability of the computing platform.
[0003] Currently, the mainstream technical paths for achieving target detection and real-time tracking include general-purpose GPU-based computing solutions, dedicated FPGA-based computing solutions, and FPGA+GPU heterogeneous collaborative computing solutions. Among these, the FPGA+GPU heterogeneous collaborative computing solution utilizes the FPGA to complete front-end tasks such as data acquisition, preprocessing, and interface conversion, while the GPU performs back-end computing tasks such as deep learning inference, achieving reasonable task decomposition and collaborative acceleration. Most existing FPGA+GPU heterogeneous solutions employ simple task-level parallel partitioning, with the FPGA handling data preprocessing and the GPU performing deep learning inference. This partitioning approach fails to fully leverage the complementary advantages of the two computing resources, limiting the parallel computing capabilities of the FPGA. After preprocessing, the particle filter is idle or under low load, while the GPU must handle all computational tasks related to target state estimation, resulting in an unbalanced GPU computational load and a still high overall processing latency. Furthermore, existing solutions do not consider the impact of fixed pipeline delays in FPGA+GPU heterogeneous solutions on real-time control, directly using the estimated target state for control command generation. This leads to a time mismatch between the target state corresponding to the control command and the target position in the actual physical world, causing significant control errors in high-speed motion scenarios. Therefore, one of the fundamental problems this invention aims to solve is how to perform refined hardware mapping based on the computational characteristics of each subtask of particle filtering and compensate for inherent pipeline delays. Summary of the Invention
[0004] To perform refined hardware mapping based on the computational characteristics of each subtask of particle filtering and to compensate for inherent pipeline delays, this application provides a target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration, employing the following technical solution: A target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration includes: The sensory data of the target scene is collected, and the sensory data is preprocessed by the FPGA to obtain preprocessed sensory data. The sensory data includes at least visual images and point cloud data. The GPU performs object detection on the preprocessed visual image and extracts features from the preprocessed point cloud data; The CPU performs data association between the target detection results and the prediction results of existing target trajectories to obtain matching relationships; it initializes new target trajectories for unmatched detection results and terminates unmatched trajectories for multiple consecutive frames. The FPGA, GPU, and CPU perform state estimation and tracking on the associated target trajectory, and output the optimal state estimate of the target. The CPU performs latency compensation on the heterogeneous system and obtains the compensated target state. The CPU constructs a finite-time optimal control problem based on the compensated target state and solves the optimal control sequence, with the first control variable of the optimal control sequence as the output result.
[0005] Optionally, the data association process includes: For an existing target trajectory, a predicted bounding box is obtained by extrapolation using a uniform motion model; a cost matrix is constructed, the elements of which are calculated based on the intersection-union ratio of the detected bounding box and the predicted bounding box, as well as the similarity between visual depth features and the historical average appearance features of the trajectory. The cost matrix is solved using the optimal binary matching algorithm to obtain the optimal matching relationship between the detection result and the existing target trajectory.
[0006] Optionally, the process of state estimation and tracking of the associated target trajectory includes: The state of each particle is predicted in parallel using multiple models via FPGA. The GPU uses the target detection results and feature extraction results as a combined observation vector to calculate the multimodal likelihood probability of each particle relative to the combined observation vector, and updates the weight of each particle based on the multimodal likelihood probability. The effective particle count is calculated by the FPGA, and resampling is triggered when the effective particle count is lower than the dynamic resampling threshold; the dynamic resampling threshold is adaptively adjusted according to the particle distribution entropy. The GPU performs iterative optimization on the subset of particles with the highest weights after resampling, moving each particle towards the region with the maximum a posteriori probability. The CPU performs weighted fusion based on all updated particles and their weights, and outputs the optimal state estimate of the target.
[0007] Optionally, the calculation process of the particle distribution entropy includes: The weights of each particle are normalized to obtain normalized weights; Take the natural logarithm of the normalized weight for each particle, and calculate the product of the normalized weight and its natural logarithm for each particle; then sum the products corresponding to all particles after taking the opposite of the original values to obtain the particle distribution entropy.
[0008] Optionally, the motion models used in the multi-model parallel prediction process include uniform velocity model, uniform acceleration model and cooperative turning model; The predicted state of each particle under each motion model is calculated using the state transition matrix and the adaptive process noise covariance; the adaptive process noise covariance is dynamically adjusted according to the target acceleration and the observation residual.
[0009] Optionally, the adaptive process noise covariance dynamic range adjustment process includes: Calculate the acceleration amplification factor based on the current target acceleration; Calculate the observation bias amplification factor based on the current observation residual; The adaptive process noise covariance is obtained by multiplying the basic process noise covariance matrix by the acceleration amplification factor and the observation bias amplification factor in sequence.
[0010] Optionally, the combined observation vector includes the three-dimensional coordinates of the target detection bounding box center coordinates, width, height, and point cloud centroid. The multimodal likelihood probability is obtained by fusing visual modal likelihood and point cloud modal likelihood. Both visual modal likelihood and point cloud modal likelihood are calculated by a weighted mixture of Gaussian distribution and robust kernel function. The weights corresponding to Gaussian distribution and robust kernel function are adjusted in real time according to the signal quality feedback of the corresponding sensor. The observation noise covariance in the Gaussian distribution calculation is adaptively adjusted based on the detection confidence level.
[0011] Optionally, the process of performing iterative optimization includes: The maximum a posteriori particle iteration enhancement strategy is used to perform at least two iterations on the top 50% of particles with the highest weights; Particles that have not undergone iteration are not included in the maximum a posteriori particle iteration enhancement strategy and are directly retained to the next time step.
[0012] Optionally, the process of delay compensation includes: Based on the pipeline processing latency of FPGA and GPU, determine the fixed system latency from the sensor data acquisition time to the state estimation output time; The optimal state estimate is extrapolated to the current control time based on the motion model to obtain the target state after delay compensation.
[0013] Optionally, the finite-time optimal control problem aims to minimize target tracking error and control energy consumption, and uses the physical limit of the actuator as a constraint to solve for the optimal control sequence. The target tracking error is the deviation between the expected position and the actual predicted position of the target at each time point in the prediction time domain.
[0014] In summary, this application includes at least one of the following beneficial technical effects: This invention deploys tasks requiring low latency and high determinism, such as data acquisition, image preprocessing, and point cloud sampling, on an FPGA. Leveraging the hardware parallelism and deterministic latency characteristics of the FPGA, it achieves microsecond-level multi-sensor time synchronization and millisecond-level data preprocessing, meeting the deterministic requirements of real-time systems for front-end data processing. Simultaneously, it deploys computationally intensive and highly parallel tasks, such as deep learning target detection, point cloud feature extraction, and multimodal likelihood calculation, on a GPU. Utilizing the GPU's powerful floating-point computing capabilities and native support for deep learning frameworks, it achieves high-precision target detection and tracking. Furthermore, by performing latency compensation on heterogeneous systems, this invention achieves spatiotemporal synchronization between target state estimation and the actual physical world state, ensuring that the target state upon which subsequent model prediction control commands are based is strictly aligned in time with the target position in the actual physical world. Attached Figure Description
[0015] Figure 1 This is a flowchart of the steps of the target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration in this invention.
[0016] Figure 2 This is a flowchart illustrating the process of state estimation and tracking of associated target trajectories in an embodiment of the present invention. Detailed Implementation
[0017] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.
[0018] In the description of this specification, the references to "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples" refer to specific features, structures, materials, or characteristics described in connection with the described embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0019] Please see Figure 1 The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration disclosed in this application includes: S1. Collect perception data of the target scene, and preprocess the perception data by the FPGA to obtain preprocessed perception data. The perception data includes at least visual images and point cloud data. S2. The GPU performs target detection on the preprocessed visual image and feature extraction on the preprocessed point cloud data; S3. The CPU performs data association between the target detection results and the prediction results of the existing target trajectories to obtain the matching relationship; initializes new target trajectories for unmatched detection results, and terminates unmatched trajectories for multiple consecutive frames. S4. The FPGA, GPU and CPU perform state estimation and tracking on the associated target trajectory, and output the optimal state estimate of the target. S5. The CPU performs delay compensation on the heterogeneous system and obtains the compensated target state. S6. The CPU constructs a finite-time optimal control problem based on the compensated target state and solves the optimal control sequence, taking the first control variable of the optimal control sequence as the output result.
[0020] This invention deploys tasks requiring low latency and high determinism, such as data acquisition, image preprocessing, and point cloud sampling, on an FPGA. Leveraging the hardware parallelism and deterministic latency characteristics of the FPGA, it achieves microsecond-level multi-sensor time synchronization and millisecond-level data preprocessing, meeting the deterministic requirements of real-time systems for front-end data processing. Simultaneously, it deploys computationally intensive and highly parallel tasks, such as deep learning target detection, point cloud feature extraction, and multimodal likelihood calculation, on a GPU. Utilizing the GPU's powerful floating-point computing capabilities and native support for deep learning frameworks, it achieves high-precision target detection and tracking. Furthermore, by performing latency compensation on heterogeneous systems, this invention achieves spatiotemporal synchronization between target state estimation and the actual physical world state, ensuring that the target state upon which subsequent model prediction control commands are based is strictly aligned in time with the target position in the actual physical world.
[0021] In one embodiment, the perceived data in step S1 includes visual images, point cloud data, and sensor data. Preprocessing the perceived data includes: denoising and enhancing the visual images through guided filtering; performing voxel mesh filtering on the point cloud data, which can significantly reduce the amount of data while preserving the geometric structure of the point cloud; using the high-precision clock inside the FPGA to unify all sensor data to the same time base; and finally, the FPGA transmits the packaged data frames to the GPU and CPU through the PCIe Gen4 x16 channel.
[0022] The feature extraction process in step S2 includes: Object detection is performed on the preprocessed visual image based on YOLOv8n+, and the output is... Target bounding box and feature vector , , Here, x and y are the center coordinates of bounding box i, respectively. Here, represents the width and height of bounding box i, respectively. The depth visual features of bounding box i; Describes the probability distribution of bounding box i across K categories, with the maximum confidence level being... .
[0023] Run PointNet++ on the point cloud data to extract geometric features from the in-view frustum point cloud corresponding to each bounding box i. .
[0024] It should be noted that the process of object detection using YOLOv8n+ and the process of extracting geometric features from point cloud data using PointNet++ are both implemented using existing solutions.
[0025] The data association process in step S2 includes: Given an existing target trajectory, the predicted bounding box is obtained by extrapolation using a uniform motion model. ,in, This is the final state estimate for the previous time step (including bounding box position and size-related components and other components). To extract the observation matrix for position and size, it selectively extracts only the components related to the bounding box position and size from the state vector, ignoring other components; The uniform state transition matrix describes the evolution of the target state over time under the assumption of uniform linear motion.
[0026] Next, a cost matrix is constructed, with its elements based on the intersection-union ratio (IU) of the detected bounding boxes and the predicted bounding boxes. And the similarity between visual depth features and trajectory history average appearance features. The calculation yields the following result, which is expressed as: , express and The intersection-union ratio, i.e., the bounding box and The ratio of the intersection area to the union area; The historical average appearance characteristics of the trajectory. Indicates taking and The Euclidean distance; The balance coefficient is set to 0.3 in this embodiment. The cost matrix is solved using the optimal bipartite matching algorithm to obtain the optimal matching relationship between the detection result and the existing target trajectory. The optimal bipartite matching algorithm can be implemented by various mathematical models for solving allocation problems, such as the existing Hungarian algorithm and JV algorithm. In this embodiment, the matching relationship is obtained by solving the problem using the Hungarian algorithm.
[0027] Please see Figure 2 Step S4, which involves state estimation and tracking of the associated target trajectory, includes: S41. The FPGA performs parallel prediction of the state of each particle using multiple models. In this embodiment, the motion models used are the uniform velocity model, the uniform acceleration model, and the cooperative turning model. The predicted state of each particle under each motion model is calculated by the state transition matrix and the adaptive process noise covariance, and is expressed as follows: ,in, Let be the process noise vector of the p-th particle under the m-th motion model. Follows a vector with zero mean and noise covariance of The multivariate normal distribution is given, where p is the index of the particle and m is the index of the motion model, therefore m = {1, 2, 3}. , For the three-dimensional position in the coordinate system, For velocities in all directions, For acceleration in all directions, For roll, pitch, and yaw angles.
[0028] Among them, noise covariance Dynamic range adjustment is performed based on the target acceleration and the observed residual. First, the acceleration amplification factor is calculated based on the current target acceleration. ; Calculate the observation bias amplification factor based on the current observation residual. The basic process noise covariance matrix Multiplying the acceleration amplification factor and the observation bias amplification factor in sequence yields the adaptive process noise covariance, which is expressed as: , Let m be the fundamental noise matrix of model m, which is constant. The acceleration response coefficient is set to 0.5 in this embodiment. To observe the residual suppression coefficient, it is set to 0.8 in this embodiment; H is the observation matrix, used to extract position, velocity, and attitude angles; Let represent the acceleration vector of the p-th particle at time k-1. For the optimal state estimate at time k-2, This is the combined observation vector at time k-1.
[0029] This invention establishes a dual adaptive adjustment mechanism based on target acceleration and observation residuals, enabling the process noise covariance to dynamically adjust in real time according to the target's maneuvering degree and model prediction deviation, thereby improving the tracking robustness of particle filtering in scenarios with severe maneuvering. Specifically, the process noise covariance is dynamically adjusted based on target acceleration: when the target acceleration increases, it is determined that the target is in a rapid maneuvering state, and the process noise covariance is automatically increased to expand the particle prediction dispersion range, covering various possible maneuvering directions of the target; when the acceleration decreases, the process noise covariance is automatically decreased to concentrate particles near the prediction position, improving prediction accuracy. The process noise covariance is also dynamically adjusted based on observation residuals: when the observation residuals increase, it is determined that there is a large deviation between the current motion model's prediction of the target state and the actual observation, and the process noise covariance is automatically increased to expand the search range, compensating for the impact of model prediction inaccuracies; when the observation residuals decrease, the process noise covariance is automatically decreased to improve the focus of the prediction. Through the independent yet synergistic interaction of these two adjustment mechanisms, the process noise covariance can be adaptively adjusted in real time according to changes in the target's motion state and observation quality.
[0030] Next, the mixture probability of each motion model is calculated. , where l is the motion model index of the previous time step. The posterior probability that the p-th particle belongs to the l-th motion model at time k-1. This represents the prior probability of the target transitioning from the l-th motion model to the m-th motion model, satisfying... It should be noted that, In this embodiment, a fixed value is preset for the system. Initial model probability As a fixed value, in this embodiment .
[0031] S42. The GPU uses the object detection results and feature extraction results as a combined observation vector. This combined observation vector includes the three-dimensional coordinates of the bounding box center coordinates, width, height, and the centroid of the point cloud, as shown below. , The coordinates of the centroid of the point cloud are in three-dimensional space.
[0032] Next, the multimodal likelihood probability of each particle relative to the combined observation vector is calculated. This multimodal likelihood probability is obtained by fusing the visual modal likelihood and the point cloud modal likelihood, and is expressed as follows: s is the modality index, where s=1 represents the visual modality and s=2 represents the point cloud modality. Both the visual modality likelihood and the point cloud modality likelihood are calculated using a weighted mixture of a Gaussian distribution and a robust kernel function. The weights corresponding to the Gaussian distribution and the robust kernel function are adjusted in real-time based on the signal quality feedback from the corresponding sensors, thus providing dynamic confidence. , The parameters of mode s are estimated (based on image contrast when s=1, and on point cloud density when s=2). ; This represents the standard Gaussian probability density function. To observe the noise covariance, it is adaptively adjusted based on the detection confidence level, and is expressed as follows: ,in, The initial noise covariances for the visual modality and the point cloud modality are respectively, which are determined through offline calibration or empirical setting, and the confidence level is detected. Adaptive adjustment of the initial noise covariance of the visual modality; For robust kernel functions (in this embodiment, the student t-distribution kernel function is used), ), where d is the dimension of the observation vector (d=3 in point cloud modes), and R is the observation noise covariance matrix. Let R be the determinant of matrix R. Here, V represents the Gamma function (used for normalization to ensure the integral of the probability density function is 1), and V represents the degrees of freedom, which is 3 in this embodiment. To observe the residual, it is represented as , Let be the actual observation vector of the s-th mode in the current frame k. Let be the observation matrix for the s-th mode.
[0033] The weights of each particle are updated based on the multimodal likelihood probability, expressed as: ,in, This represents the total number of particles at time k-1. The unnormalized weights are expressed as: , Let be the weight of particle p at time k. Let be the weight of particle p at time k-1.
[0034] It should be noted that the standard Gaussian probability density function mentioned above is existing technology, therefore its expression is not listed.
[0035] This invention employs differentiated likelihood functions for different modalities: the visual modality is calculated using a Gaussian distribution, while the point cloud modality is calculated using a robust kernel function to suppress outlier interference in the LiDAR point cloud. The confidence level of each modality is adjusted in real time based on the sensor signal quality feedback. When the detection confidence is high, the observation noise covariance of the visual modality automatically decreases, and the system places greater trust in visual observations. When the detection confidence is low, the observation noise covariance automatically increases, and the system automatically reduces its dependence on visual observations. This dynamic confidence level mechanism enables the system to automatically adjust the fusion weights according to the real-time operating status of each sensor, achieving adaptive optimal fusion of multi-source information.
[0036] S43. Calculate the effective number of particles using FPGA. Simultaneously calculate the tracking confidence level. , The number of bounding boxes in the current frame (number of detected objects); the total number of particles at time k is calculated using the tracking confidence score. "round" means rounding to the nearest integer. , .
[0037] Next, the particle distribution entropy is calculated. The weights of each particle are normalized to obtain normalized weights. The natural logarithm of the normalized weights of each particle is taken, and the product of the normalized weight and its natural logarithm is calculated. The negatives of the products for all particles are then summed to obtain the particle distribution entropy, expressed as: The dynamic resampling threshold is adaptively adjusted based on the particle distribution entropy. The dynamic resampling threshold is expressed as... Resampling is triggered when the number of effective particles is lower than the dynamic resampling threshold.
[0038] This invention adaptively adjusts the resampling trigger threshold based on the particle weight distribution dispersion (particle distribution entropy). When the dispersion is large, it is determined that the current particle set has sufficient diversity, and the resampling threshold is automatically lowered to reduce unnecessary resampling computation overhead. When the dispersion is small, it is determined that the current particle set is severely degraded, and the resampling threshold is automatically raised to trigger resampling earlier and suppress particle degradation in time. This makes the resampling timing match the actual degradation state of the particle set, achieving a dynamic balance between computational resource utilization and tracking robustness.
[0039] S44. The GPU performs iterative optimization on the subset of particles with the highest weights after resampling. This process includes: performing at least two iterations on the top 50% of particles with the highest weights using a maximum a posteriori (MAP) particle iteration enhancement strategy; particles that have not undergone iteration do not participate in the MAP particle iteration enhancement strategy and are directly retained until the next time step, and each particle is moved towards the MAP probability region; in this embodiment, the number of iterations is 2, and the MAP particle iteration enhancement strategy uses an extended Kalman filter (EKF); therefore, the iterative process is expressed as: , where l is the iteration index, and in this embodiment l=1, 2.
[0040] S45. The CPU performs weighted fusion based on all updated particles and their weights, and outputs the optimal state estimate of the target. .
[0041] The delay compensation process in step S5 includes: Based on the pipeline processing latency of the FPGA and GPU, determine the system fixed latency from the sensor data acquisition time to the state estimation output time. Based on the motion model, the optimal state estimate is extrapolated to the current control time to obtain the target state after delay compensation. ,in, Let be the mixture probability of the m-th motion model. Let be the continuous-time state transition matrix of the m-th motion model under a delay duration τ. , This is the continuous form of the state transition matrix of the m-th motion model, which belongs to the fixed parameters of the motion model.
[0042] This invention compensates for the fixed pipeline delay of an FPGA+GPU heterogeneous system by looking ahead, enabling the target state corresponding to the control command to be synchronized with the target position in the actual physical world. In high-speed target motion scenarios, it can effectively eliminate control lag errors caused by pipeline delay, significantly improving the accuracy and stability of tracking control.
[0043] In step S6, the finite-time optimal control problem aims to minimize the target tracking error and control energy consumption, using the physical limit of the actuator as a constraint, and solves for the optimal control sequence; where the target tracking error... The deviation between the expected position and the actual predicted position of the target at each time point in the prediction time domain is expressed as: , To control the output matrix, The target state after delay compensation. To determine the desired tracking reference value, it should be noted that the control output matrix... The preset constant matrix is used to extract the state components related to the tracking error from the state vector; the expected tracking reference value is a preset time-varying vector sequence, representing the expected output of the system at each time point in the prediction time domain.
[0044] The finite-time optimal control problem is expressed as: ,in, To predict the time domain, To control the time domain, For the tracking error weight matrix, To control the incremental weight matrix, To control variables, To control the increment, , , , and Based on the physical limits of the actuator The target tracking error vector is represented by The weighted square norm, Indicates control variables according to The weighted square norm.
[0045] It should be noted that the tracking error weight matrix and the control increment weight matrix determine the degree of importance the control system attaches to tracking accuracy and control energy consumption, respectively. The diagonal elements of the tracking error weight matrix are set according to the maximum permissible deviation of each error component, and the diagonal elements of the control increment weight matrix are set according to the maximum permissible response speed of the actuator; the relative ratio of the two determines the response characteristics of the controller.
[0046] This invention utilizes an FPGA to handle tasks with high time determinism requirements, such as particle prediction, resampling, and sensor preprocessing. By leveraging the hardware parallelism and pipelined architecture of the FPGA, it achieves microsecond-level processing latency and nanosecond-level timing determinism. Conversely, it uses a GPU to handle computationally intensive tasks such as deep learning inference and likelihood calculation. By utilizing the high throughput parallel computing capabilities of the GPU, it achieves high-precision perception and tracking. The complementary advantages of these two architectures enable the system to possess both low latency and high precision characteristics.
[0047] This invention forms a complete end-to-end processing link, from multi-sensor data acquisition, FPGA preprocessing, GPU target detection, particle filter tracking, delay compensation to MPC control command output. Based on the target state after delay compensation, the optimal control sequence is solved in the finite time domain to generate smooth control commands to drive the actuator, realizing an integrated closed loop of perception, decision-making, and control, which can be directly deployed in autonomous systems such as autonomous driving, drones, and intelligent robots.
[0048] This invention achieves significantly lower computational resource consumption than a fixed particle number scheme while maintaining tracking accuracy through an adaptive particle number control mechanism (reducing the number of particles to save computing power when tracking confidence is high, and increasing the number of particles to enhance coverage when tracking confidence is low) and fine-grained task partitioning. It is particularly suitable for deployment on embedded heterogeneous computing platforms with limited power consumption and computing power.
[0049] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration, characterized in that, include: The sensory data of the target scene is collected, and the sensory data is preprocessed by the FPGA to obtain preprocessed sensory data. The sensory data includes at least visual images and point cloud data. The GPU performs object detection on the preprocessed visual image and extracts features from the preprocessed point cloud data; The CPU performs data association between the target detection results and the prediction results of existing target trajectories to obtain matching relationships; it initializes new target trajectories for unmatched detection results and terminates unmatched trajectories for multiple consecutive frames. The FPGA, GPU, and CPU perform state estimation and tracking on the associated target trajectory, and output the optimal state estimate of the target. The CPU performs latency compensation on the heterogeneous system and obtains the compensated target state. The CPU constructs a finite-time optimal control problem based on the compensated target state and solves the optimal control sequence, with the first control variable of the optimal control sequence as the output result.
2. The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration according to claim 1, characterized in that, The data association process includes: For an existing target trajectory, a predicted bounding box is obtained by extrapolation using a uniform motion model; a cost matrix is constructed, the elements of which are calculated based on the intersection-union ratio of the detected bounding box and the predicted bounding box, as well as the similarity between visual depth features and the historical average appearance features of the trajectory. The cost matrix is solved using the optimal binary matching algorithm to obtain the optimal matching relationship between the detection result and the existing target trajectory.
3. The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration according to claim 1, characterized in that, The process of state estimation and tracking of associated target trajectories includes: The state of each particle is predicted in parallel using multiple models via FPGA. The GPU uses the target detection results and feature extraction results as a combined observation vector to calculate the multimodal likelihood probability of each particle relative to the combined observation vector, and updates the weight of each particle based on the multimodal likelihood probability. The effective particle count is calculated by the FPGA, and resampling is triggered when the effective particle count is lower than the dynamic resampling threshold; the dynamic resampling threshold is adaptively adjusted according to the particle distribution entropy. The GPU performs iterative optimization on the subset of particles with the highest weights after resampling, moving each particle towards the region with the maximum a posteriori probability. The CPU performs weighted fusion based on all updated particles and their weights, and outputs the optimal state estimate of the target.
4. The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration according to claim 3, characterized in that, The calculation process of the particle distribution entropy includes: The weights of each particle are normalized to obtain normalized weights; Take the natural logarithm of the normalized weight for each particle, and calculate the product of the normalized weight and its natural logarithm for each particle; then sum the products corresponding to all particles after taking the opposite of the original values to obtain the particle distribution entropy.
5. The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration according to claim 3, characterized in that, The motion models used in the multi-model parallel prediction process include uniform velocity model, uniform acceleration model, and cooperative turning model; The predicted state of each particle under each motion model is calculated using the state transition matrix and the adaptive process noise covariance; the adaptive process noise covariance is dynamically adjusted according to the target acceleration and the observation residual.
6. The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration according to claim 5, characterized in that, The adaptive process for adjusting the dynamic range of noise covariance includes: Calculate the acceleration amplification factor based on the current target acceleration; Calculate the observation bias amplification factor based on the current observation residual; The adaptive process noise covariance is obtained by multiplying the basic process noise covariance matrix by the acceleration amplification factor and the observation bias amplification factor in sequence.
7. The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration according to claim 3, characterized in that, The combined observation vector includes the three-dimensional coordinates of the target detection bounding box center coordinates, width, height, and point cloud centroid. The multimodal likelihood probability is obtained by fusing visual modal likelihood and point cloud modal likelihood. Both visual modal likelihood and point cloud modal likelihood are calculated by a weighted mixture of Gaussian distribution and robust kernel function. The weights corresponding to Gaussian distribution and robust kernel function are adjusted in real time according to the signal quality feedback of the corresponding sensor. The observation noise covariance in the Gaussian distribution calculation is adaptively adjusted based on the detection confidence level.
8. The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration according to claim 3, characterized in that, The process of performing iterative optimization includes: The maximum a posteriori particle iteration enhancement strategy is used to perform at least two iterations on the top 50% of particles with the highest weights; Particles that have not undergone iteration are not included in the maximum a posteriori particle iteration enhancement strategy and are directly retained to the next time step.
9. The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration according to claim 4, characterized in that, The process of delay compensation includes: Based on the pipeline processing latency of FPGA and GPU, determine the fixed system latency from the sensor data acquisition time to the state estimation output time; The optimal state estimate is extrapolated to the current control time based on the motion model to obtain the target state after delay compensation.
10. The target detection and real-time tracking control method based on FPGA+GPU heterogeneous collaboration according to claim 1, characterized in that, The finite-time optimal control problem aims to minimize target tracking error and control energy consumption, with the physical limit of the actuator as the constraint, and solves for the optimal control sequence. The target tracking error is the deviation between the expected position and the actual predicted position of the target at each time point in the prediction time domain.