Multi-target tracking method and device based on diffusion model
By introducing a diffusion model-based detection and trajectory prediction method in the multi-objective tracking technology, combined with a matching association strategy of mixed IoU and ReID distances, the problem that the existing technology is difficult to deal with linear and nonlinear motion targets and dense scene ID switching at the same time, achieving a more stable and robust multi-objective tracking effect.
Patent Information
- Application Number
- CN202411811835.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-10
AI Technical Summary
The existing multi-objective tracking technology is difficult to cope with linear and nonlinear moving targets at the same time, and ID switching problems are prone to occur in dense scenarios, and the stability and robustness of tracking are poor.
Using a multi-objective tracking method based on the diffusion model, the ROI features of the image are acquired through the preset feature extraction module and input them to the diffusion detection head module for object detection. At the same time, the target displacement is calculated and input to the decoupled diffusion trajectory prediction module for trajectory prediction. Finally, the target matching association is performed based on the hybrid IoU and ReID distance calculation strategy to generate the final target tracking result.
It realizes accurate tracking of multiple mobile targets in complex sports states, and is suitable for security, rescue, sports events and autonomous driving, improving the stability and robustness of tracking and reducing ID switching problems.
Smart Images

Figure CN119941785A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a multi-target tracking method and device based on a diffusion model. Background Art
[0002] Multi-target tracking technology has a wide range of applications in video surveillance, autonomous driving, smart security, sports events and other fields. Its core goal is to accurately and continuously track multiple moving targets from a video stream, thereby providing reliable target trajectory information for downstream tasks (such as behavior analysis, decision support, etc.). Traditional multi-target tracking methods mainly include two stages: target detection and data association. First, the target in the video frame is identified through the target detection model, and then the same target in consecutive frames is associated through the data association algorithm to form a motion trajectory. However, the complexity of actual application scenarios poses many challenges to existing multi-target tracking technology.
[0003] In real-world scenarios, targets usually have complex motion patterns, such as linear and nonlinear motion, frequent occlusions, and appearance similarities between targets. In addition, different camera devices and environmental conditions (such as lighting, weather, etc.) can also significantly affect the performance of target detection and tracking. Therefore, an efficient multi-target tracking system needs to be highly robust and able to effectively handle occlusion and target overlap problems in dense target scenes while maintaining high-precision tracking of linear and nonlinear moving targets.
[0004] Although deep learning-based multi-target tracking technology has made significant progress, existing methods have difficulty in handling both linear and nonlinear moving targets, and the tracking effect in scenes with dense targets is still unsatisfactory. How to achieve an efficient and robust multi-target tracking system in complex scenes is still a technical problem that needs to be solved urgently.
[0005] Target tracking is an important research direction in computer vision. In recent years, existing technologies have introduced diffusion models into the field of multi-target tracking, and regarded target detection and association as a consistent denoising diffusion process from paired noise boxes to target associations, which successfully achieved target tracking. However, it only considers the position transformation of the target without considering the appearance of the target, and does not consider the scene of nonlinear motion. Therefore, it is less effective in dense scenes and when processing nonlinear motion scenes. In addition, DiffMOT is another multi-target tracking method based on the diffusion model. It introduces the diffusion model in trajectory prediction, which improves the tracking effect of nonlinear motion targets. However, DiffMOT is not ideal when tracking linear motion targets, and is prone to ID switching problems in dense scenes.
[0006] In summary, the existing multi-target tracking technology is difficult to deal with linear and nonlinear moving targets at the same time, and is prone to ID switching problems in dense scenes. The tracking stability and robustness are poor and need to be solved urgently. Summary of the invention
[0007] The present application provides a multi-target tracking method and device based on a diffusion model to solve the problems that the existing multi-target tracking technology is difficult to deal with linear and nonlinear moving targets at the same time, is prone to ID switching problems in dense scenes, and has poor tracking stability and robustness.
[0008] The first aspect of the present application provides a multi-target tracking method based on a diffusion model, comprising the following steps: obtaining ROI features of two frames of images to be detected through a preset feature extraction module, and inputting the ROI features into a preset diffusion detection head module, and performing target detection operations on the two frames of images to be detected to generate target detection results for each target in the two frames of images to be detected, wherein the two frames of images to be detected include a previous frame of image to be detected and a current frame of image to be detected; calculating the target displacement of each target in the two frames of images to be detected according to the target detection results, and inputting the target displacement into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of image to be detected to obtain a trajectory prediction result; based on a preset hybrid IoU and ReID distance calculation strategy, performing matching and association operations on the target detection results and the trajectory prediction results to generate a final target tracking result.
[0009] Optionally, in one embodiment of the present application, the ROI features of the two frames of images to be detected are obtained through a preset feature extraction module, and the ROI features are input into a preset diffusion detection head module, and a target detection operation is performed on the two frames of images to be detected to generate a target detection result for each target in the two frames of images to be detected, including: inputting the two frames of images to be detected into the feature extraction module to output the ROI features of the target size corresponding to the two frames of images to be detected; generating a plurality of random frames in the two frames of images to be detected, and determining the plurality of random frames as the initial detection frames corresponding to the two frames of images to be detected; based on a preset multi-scale attention module, a dynamic convolution module and a data The association head module constructs the diffusion detection head module, and inputs the ROI feature and the initial detection frame into the diffusion detection head module to generate the prediction frames corresponding to the two frames of images to be detected before and after through the multi-scale attention module and the dynamic convolution module, and calculates the noise between the multiple random frames and the prediction frame; the prediction frame and the noise are input into the data association head module to obtain the detection frame of the current iteration, and the initial detection frame is updated using the detection frame of the current iteration; the ROI feature and the updated initial detection frame are re-input into the diffusion detection head module to iteratively perform the target detection operation until the preset time step iteration end condition is met to generate the target detection result.
[0010] Optionally, in one embodiment of the present application, the target displacement is input into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of the image to be detected to obtain a trajectory prediction result, including: constructing the decoupled diffusion trajectory prediction module through a preset plurality of multi-scale attention modules and a motion fusion module; adding Gaussian noise to the target displacement, and inputting the target displacement after adding Gaussian noise into the decoupled diffusion trajectory prediction module, so as to perform post-diffusion on the target displacement after adding Gaussian noise through the multi-scale attention module and the motion fusion module to generate the target predicted displacement of the current frame of the image to be detected; and calculating the trajectory prediction result of the current frame of the image to be detected based on the target predicted displacement and the target detection result of the previous frame of the image to be detected.
[0011] Optionally, in one embodiment of the present application, the target detection result and the trajectory prediction result are matched and associated based on the preset hybrid IoU and ReID distance calculation strategy to generate a final target tracking result, including: calculating the IoU distance and ReID distance between the target detection result and the trajectory prediction result, and calculating the distance weight corresponding to each coordinate distance between the IoU distance and the ReID distance; calculating the final hybrid distance corresponding to the IoU distance and the ReID distance according to the distance weight and the preset weight factor; based on the final hybrid distance and the preset linear allocation algorithm, calculating the target matching pair corresponding to the target detection result and the trajectory prediction result, and updating the target trajectory through the target matching pair to obtain the final target tracking result.
[0012] The second aspect of the present application provides a multi-target tracking device based on a diffusion model, including: a target detection module, which is used to obtain the ROI features of the previous and next two frames of images to be detected through a preset feature extraction module, and input the ROI features into a preset diffusion detection head module, and perform target detection operations on the previous and next two frames of images to be detected to generate target detection results for each target in the previous and next two frames of images to be detected, wherein the previous and next two frames of images to be detected include a previous frame of image to be detected and a current frame of image to be detected; a trajectory prediction module, which is used to calculate the target displacement of each target in the previous and next two frames of images to be detected according to the target detection results, and input the target displacement into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of image to be detected to obtain a trajectory prediction result; a matching association module, which is used to perform matching association operations on the target detection results and the trajectory prediction results based on a preset hybrid IoU and ReID distance calculation strategy to generate a final target tracking result.
[0013] Optionally, in one embodiment of the present application, the target detection module includes: an input unit, used to input the two frames of images to be detected before and after into the feature extraction module to output the ROI features of the target size corresponding to the two frames of images to be detected before and after; a generation unit, used to generate a plurality of random frames in the two frames of images to be detected before and after, and determine the plurality of random frames as initial detection frames corresponding to the two frames of images to be detected before and after; a first calculation unit, used to construct the diffusion detection head module based on a preset multi-scale attention module, a dynamic convolution module and a data association head module, and input the ROI features and the initial detection frame into the diffusion detection head module; In the detection head module, the prediction frames corresponding to the two frames of images to be detected are generated by the multi-scale attention module and the dynamic convolution module, and the noise between the multiple random frames and the prediction frames is calculated; the updating unit is used to input the prediction frame and the noise into the data association head module to obtain the detection frame of the current iteration, and use the detection frame of the current iteration to update the initial detection frame; the iteration unit is used to re-input the ROI feature and the updated initial detection frame into the diffusion detection head module to iteratively perform the target detection operation until the preset time step iteration end condition is met to generate the target detection result.
[0014] Optionally, in one embodiment of the present application, the trajectory prediction module includes: a construction unit, used to construct the decoupled diffusion trajectory prediction module through a preset plurality of multi-scale attention modules and a motion fusion module; a noise adding unit, used to add Gaussian noise to the target displacement, and input the target displacement after adding Gaussian noise into the decoupled diffusion trajectory prediction module, so as to perform post-diffusion on the target displacement after adding Gaussian noise through the multi-scale attention module and the motion fusion module to generate the target predicted displacement of the current frame to be detected; a second calculation unit, used to calculate the trajectory prediction result of the current frame to be detected based on the target predicted displacement and the target detection result of the previous frame to be detected.
[0015] Optionally, in one embodiment of the present application, the matching association module includes: a third calculation unit, used to calculate the IoU distance and ReID distance between the target detection result and the trajectory prediction result, and calculate the distance weight corresponding to each coordinate distance between the IoU distance and the ReID distance; a fourth calculation unit, used to calculate the final mixed distance corresponding to the IoU distance and the ReID distance according to the distance weight and a preset weight factor; a linear allocation unit, used to calculate the target matching pairs corresponding to the target detection result and the trajectory prediction result based on the final mixed distance and a preset linear allocation algorithm, and update the target trajectory through the target matching pair to obtain the final target tracking result.
[0016] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-target tracking method based on the diffusion model as described in the above embodiment.
[0017] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-mentioned multi-target tracking method based on a diffusion model.
[0018] The fifth aspect of the present application provides a computer program product, including a computer program, which is executed to implement the above-mentioned diffusion model-based multi-target tracking method.
[0019] Therefore, the embodiments of the present application have the following beneficial effects:
[0020] The embodiments of the present application can obtain the ROI features of the two frames of images to be detected by a preset feature extraction module, and input the ROI features to the preset diffusion detection head module, and perform target detection operations on the two frames of images to be detected to generate target detection results for each target in the two frames of images to be detected, wherein the two frames of images to be detected include the previous frame of images to be detected and the current frame of images to be detected; calculate the target displacement of each target in the two frames of images to be detected according to the target detection results, and input the target displacement into the preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of images to be detected to obtain the trajectory prediction results; based on the preset hybrid IoU and ReID distance calculation strategy, match and associate the target detection results and the trajectory prediction results to generate the final target tracking results. The present application realizes accurate tracking of multiple moving targets in complex motion states by combining the diffusion detection module and the diffusion trajectory prediction module, and can be applied to multiple fields such as security, rescue, sports events and autonomous driving. This solves the problems that the existing multi-target tracking technology is difficult to deal with linear and nonlinear moving targets at the same time, is prone to ID switching problems in dense scenes, and has poor tracking stability and robustness.
[0021] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0023] Figure 1A flowchart of a multi-target tracking method based on a diffusion model provided according to an embodiment of the present application;
[0024] Figure 2 A schematic diagram of execution logic of a multi-target tracking method based on a diffusion model provided for one embodiment of the present application;
[0025] Figure 3 A schematic diagram of a multi-scale attention module network structure provided for an embodiment of the present application;
[0026] Figure 4 A schematic diagram of a motion fusion module network structure provided for an embodiment of the present application;
[0027] Figure 5 A schematic diagram of a multi-target tracking effect provided by an embodiment of the present application;
[0028] Figure 6 is an example diagram of a multi-target tracking device based on a diffusion model according to an embodiment of the present application;
[0029] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0030] Among them, 10-a multi-target tracking device based on a diffusion model; 100-a target detection module, 200-a trajectory prediction module, 300-a matching association module; 701-a memory, 702-a processor, 703-a communication interface. DETAILED DESCRIPTION
[0031] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0032] The following describes the multi-target tracking method and device based on the diffusion model of the embodiment of the present application with reference to the accompanying drawings. In response to the problems mentioned in the above background technology, the present application provides a multi-target tracking method based on the diffusion model, in which the ROI features of the two frames of images to be detected are obtained by a preset feature extraction module, and the ROI features are input into a preset diffusion detection head module, and the target detection operation is performed on the two frames of images to be detected to generate the target detection results of each target in the two frames of images to be detected, wherein the two frames of images to be detected include the previous frame of the image to be detected and the current frame of the image to be detected; the target displacement of each target in the two frames of images to be detected is calculated according to the target detection result, and the target displacement is input into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of the image to be detected to obtain the trajectory prediction result; based on the preset hybrid IoU and ReID distance calculation strategy, the target detection result and the trajectory prediction result are matched and associated to generate the final target tracking result. This application achieves accurate tracking of multiple moving targets in complex motion states by combining a diffusion detection module and a diffusion trajectory prediction module, and can be applied to multiple fields such as security, rescue, sports events, and autonomous driving. This solves the problems that existing multi-target tracking technologies are difficult to handle both linear and nonlinear moving targets at the same time, are prone to ID switching problems in dense scenes, and have poor tracking stability and robustness.
[0033] Specifically, Figure 1 A flowchart of a multi-target tracking method based on a diffusion model provided in an embodiment of the present application.
[0034] like Figure 1 As shown, the multi-target tracking method based on the diffusion model includes the following steps:
[0035] In step S101, the ROI features of the two frames of images to be detected are obtained through a preset feature extraction module, and the ROI features are input into a preset diffusion detection head module, and a target detection operation is performed on the two frames of images to be detected to generate target detection results for each target in the two frames of images to be detected, wherein the two frames of images to be detected include a previous frame of image to be detected and a current frame of image to be detected.
[0036] The embodiment of the present application can firstly obtain the ROI features of the input two frames of images to be detected through the feature extraction module, wherein the feature extraction module is the backbone network in YOLOX; secondly, the embodiment of the present application can input the ROI features into the diffusion detection head module to detect each target in the two frames of images to be detected.
[0037] Optionally, in one embodiment of the present application, the ROI features of the two frames of images to be detected are obtained by a preset feature extraction module, and the ROI features are input into a preset diffusion detection head module, and the target detection operation is performed on the two frames of images to be detected to generate target detection results for each target in the two frames of images to be detected, including: inputting the two frames of images to be detected into the feature extraction module to output the ROI features of the target size corresponding to the two frames of images to be detected; generating multiple random frames in the two frames of images to be detected, and determining the multiple random frames as the initial detection frames corresponding to the two frames of images to be detected; based on the preset multi-scale attention module, dynamic Convolution module and data association head module, construct diffusion detection head module, and input ROI features and initial detection frame into the diffusion detection head module, so as to generate prediction frames corresponding to the two frames of images to be detected before and after through multi-scale attention module and dynamic convolution module, and calculate the noise between multiple random frames and prediction frames; input the prediction frame and noise into the data association head module to obtain the detection frame of the current iteration, and use the detection frame of the current iteration to update the initial detection frame; re-input the ROI features and the updated initial detection frame into the diffusion detection head module to iteratively perform target detection operation until the preset time step iteration end condition is met to generate the target detection result.
[0038] It should be noted that the embodiment of the present application can be the previous frame of the image to be detected F T-1 and the current frame to be detected image F T Input feature extraction module to obtain ROI feature R with scale of 4000×320×7×7 T-1 and R T , and recorded as P T =(R T ,R T-1 ).
[0039] Secondly, the embodiment of the present application can input the ROI feature into the diffusion detection head module to detect each target in the two frames of images to be detected, wherein the diffusion detection head module is composed of a multi-scale attention module, a dynamic convolution module and a data association head module, such as Figure 2 shown.
[0040] Specifically, the steps of performing target detection using ROI features by the diffusion detection head module in the embodiment of the present application are as follows:
[0041] Step 1: Assume N = 500, in the previous frame of the image to be detected F T-1 and the current frame to be detected image F T Generate N random boxes B on T-1 and B T , and use it as the initial detection box, denoted as Z T =(BT ,B T-1 );
[0042] Step 2: The ROI feature P obtained above T and the initial detection box Z T At the same time, the diffusion detection head module is input, and its detection process can be regarded as a denoising diffusion process from the noise frame to the target frame;
[0043] The embodiment of the present application can obtain the prediction box B through the multi-scale attention module and the dynamic convolution module pred ; Given the data distribution X0~q(X0), the forward noise perturbation process at time t is defined as q(X T |X T-1 ), and according to the variance table β1, ...β T Gradually add Gaussian noise as shown below:
[0044]
[0045] Given X0, the embodiments of the present application can obtain X by sampling Gaussian vector ε~N(0,I) and applying the transformation T , specifically expressed as:
[0046]
[0047] in,
[0048] Afterwards, the embodiment of the present application can be B pred and Z T Substitute X T and X0 to calculate the noise introduced by the process from random box to predicted box Neural Network θ (X T ,T) is used to train T Predict X0, and its loss is as follows:
[0049]
[0050] Step 3: The noise obtained in step 2 T and prediction box B pred Input data association head module, and obtain the detection box B of the current iteration through denoising diffusion temp , the process is expressed as:
[0051]
[0052] in, and B pred and Noise T, σ T It can be expressed as:
[0053]
[0054] Therefore, the X calculated above is T-1 That is the detection box B generated by the current iteration temp ;
[0055] Step 4: Use the generated detection box B obtained in step 3 temp Update Z in step 2 T , and the updated Z T and ROI feature P T Input the diffusion detection head module again for a new round of iteration until time step T = 0. At this time, the obtained B temp This is the final detection box B det .
[0056] In step S102, the target displacement of each target in the two frames of images to be detected is calculated according to the target detection result, and the target displacement is input into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of the image to be detected, so as to obtain a trajectory prediction result.
[0057] Furthermore, the embodiment of the present application also needs to calculate the displacement of the target in the two previous and next frames of the image to be detected, and input it into the decoupled diffusion trajectory prediction module to predict the trajectory of the target in the previous frame of the image to be detected.
[0058] Therefore, the embodiments of the present application achieve accurate tracking of multiple targets in complex motion states by innovatively combining a diffusion detection module and a diffusion trajectory prediction module.
[0059] Optionally, in one embodiment of the present application, the target displacement is input into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of the image to be detected to obtain a trajectory prediction result, including: constructing a decoupled diffusion trajectory prediction module through a preset plurality of multi-scale attention modules and a motion fusion module; adding Gaussian noise to the target displacement, and inputting the target displacement after adding Gaussian noise into the decoupled diffusion trajectory prediction module, so as to perform post-diffusion on the target displacement after adding Gaussian noise through a multi-scale attention module and a motion fusion module to generate a target predicted displacement of the current frame of the image to be detected; and calculating the trajectory prediction result of the current frame of the image to be detected based on the target predicted displacement and the target detection result of the previous frame of the image to be detected.
[0060] Specifically, the specific steps of the embodiment of the present application for performing target trajectory prediction by decoupling the diffusion trajectory prediction module are as follows:
[0061] Step 1: Based on the target detection results D of the two frames of images to be detected T and D T-1 , calculate its target displacement M T , the calculation process is shown as follows:
[0062] M T =D T -D T-1 =(ΔX T ,ΔY T ,ΔW T ,ΔH T )
[0063] Among them, X and Y are the horizontal and vertical coordinates of the center point of the target, and W and H are the width and height of the detection box;
[0064] Step 2: Displace the target obtained in step 1 by M T After adding Gaussian noise, the input consists of a multi-scale attention module, a dynamic convolution module, and a data association head module to form a decoupled diffusion trajectory prediction module, and is implemented as follows: Figure 3 The multi-scale attention module shown in Figure 4 The motion fusion module shown in the figure generates the target prediction displacement M of the current frame to be detected. pred ;
[0065] Step 3: Use the target prediction displacement M from step 2 pred and the detection result D of the previous frame of the image to be detected T-1 Get the predicted trajectory D of the image to be detected in the current frame pred , the process is expressed as:
[0066] D pred =D T-1 +M pred
[0067] Therefore, the embodiment of the present application predicts the trajectory of each target in the previous frame of the image to be detected by decoupling the diffusion trajectory prediction module, thereby providing guidance and basis for the execution of subsequent matching association operations.
[0068] In step S103, based on the preset hybrid IoU and ReID distance calculation strategy, a matching association operation is performed on the target detection result and the trajectory prediction result to generate a final target tracking result.
[0069] Furthermore, the embodiments of the present application also need to perform matching association based on the target detection results and the predicted trajectory through a hybrid IoU and ReID distance calculation algorithm to obtain the final target tracking result, and perform iterative updates to continuously track subsequent images.
[0070] Therefore, the embodiments of the present application establish a distance calculation strategy that mixes IoU and ReID, thereby effectively reducing ID switching during target occlusion and interaction and improving the stability of target tracking.
[0071] Optionally, in one embodiment of the present application, based on a preset hybrid IoU and ReID distance calculation strategy, a matching association operation is performed on the target detection result and the trajectory prediction result to generate a final target tracking result, including: calculating the IoU distance and ReID distance between the target detection result and the trajectory prediction result, and calculating the distance weight corresponding to each coordinate distance between the IoU distance and the ReID distance; calculating the final hybrid distance corresponding to the IoU distance and the ReID distance according to the distance weight and the preset weight factor; based on the final hybrid distance and the preset linear allocation algorithm, calculating the target matching pairs corresponding to the target detection result and the trajectory prediction result, and updating the target trajectory through the target matching pairs to obtain the final target tracking result.
[0072] Specifically, the embodiment of the present application performs data association on the detection result and predicted trajectory of the current frame to be detected image according to the hybrid IoU and ReID distance calculation strategy to obtain the steps of the final target tracking result as follows:
[0073] Step 1: Calculate the detection box B obtained above det and predicted trajectory D pred The IoU distance d between I and ReID distance d E , wherein the embodiments of the present application may use the SBS-S50 network to extract target feature embedding;
[0074] Step 2: For step 1, I and d E The distance of each coordinate of is used to calculate the corresponding distance weight, and the value of each coordinate can be expressed as:
[0075] D I =d I [i][j]D E =d E [i][j]
[0076] Set the distance thresholds of IoU and ReID to τ respectively I and τ E , the scaled distance of each coordinate is calculated as follows:
[0077]
[0078] Among them, the weight calculation strategy is F~1 / e X , the embodiments of the present application can be based on I Δ and E ΔCalculate the distance weight W of IoU and ReID I and W E , in D I <τ I &D E <τ E When , the hybrid distance is calculated as follows:
[0079] d mix [i][j]=W I ·D I +W E ·D E
[0080] In other cases, a weight factor η needs to be added to the IoU distance calculation, which is expressed as:
[0081] d mix [i][j]=W I ·D I ·η+W E ·D E
[0082] After one round of iteration, the final mixed distance d can be obtained mix ;
[0083] Step 3: Use the linear assignment algorithm to calculate the mixed distance d from step 2 mix Calculate and get the predicted trajectory D pred And the test result B det The target matching pair m τ , according to the matching pair m τ Update the target trajectory to obtain the target tracking result of the current frame.
[0084] Subsequently, the embodiment of the present application can perform target tracking in subsequent frames. At this time, the current frame F T Change to F T-1 , by iterating the above-mentioned target detection, trajectory prediction and other operations, to obtain the tracking results of subsequent frames until all frames are tracked, and the effect of the embodiment of the present application on multi-target tracking is as follows Figure 5 shown.
[0085] Therefore, the detection and trajectory prediction results in the embodiments of the present application are generated using the diffusion model, which takes into account the linear and nonlinear motion of the target, while considering the occlusion of dense scenes, and has better tracking effects and stronger robustness in real scenes; in addition, the embodiments of the present application can also improve the sampling speed of the diffusion process through parallel sampling technology, thereby reducing its inference time, and is suitable for application scenarios with high real-time requirements.
[0086] To summarize, the embodiment of the present application can firstly obtain the ROI features of the input two frames of images before and after through the feature extraction module; secondly, use the diffusion detection head structure based on the DDIM process to detect the targets in the two frames before and after, and use the trajectory prediction module based on the decoupled diffusion model to predict the trajectory of the previous frame of the image; then, use a hybrid IoU and ReID distance calculation algorithm to effectively associate the predicted trajectory with the obtained detection result, so as to obtain the final target tracking result. Therefore, the embodiment of the present application can be well applied to multiple fields such as security, rescue, sports events and autonomous driving, and can effectively and accurately track moving targets in complex motion states.
[0087] The present application can also construct a corresponding multi-target tracking system based on a diffusion model according to the multi-target tracking method based on a diffusion model. The following is a brief description and introduction of the multi-target tracking system based on a diffusion model of the present application.
[0088] The multi-target tracking system based on the diffusion model of the present application mainly includes a feature extraction unit, a target detection unit, a motion trajectory prediction unit and a target tracking unit.
[0089] The feature extraction unit is used to extract the ROI features of the input images to be detected in the previous and next frames;
[0090] The target detection unit is used to diffusely generate target detection results of input images to be detected in previous and next frames;
[0091] The motion trajectory prediction unit is used to diffusely generate the motion trajectory of the target to be detected in the current frame;
[0092] The target tracking unit is used to perform data association on the detection result and predicted trajectory of the current frame to be detected to obtain the final target tracking result.
[0093] According to the multi-target tracking method based on the diffusion model proposed in the embodiment of the present application, the ROI features of the two frames of images to be detected are obtained through a preset feature extraction module, and the ROI features are input into a preset diffusion detection head module, and the target detection operation is performed on the two frames of images to be detected to generate the target detection results of each target in the two frames of images to be detected, wherein the two frames of images to be detected include the previous frame of images to be detected and the current frame of images to be detected; the target displacement of each target in the two frames of images to be detected is calculated according to the target detection results, and the target displacement is input into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of images to be detected to obtain the trajectory prediction result; based on the preset hybrid IoU and ReID distance calculation strategy, the target detection results and the trajectory prediction results are matched and associated to generate the final target tracking result. The present application realizes the accurate tracking of multiple moving targets in complex motion states by combining the diffusion detection module and the diffusion trajectory prediction module, and can be applied to multiple fields such as security, rescue, sports events and autonomous driving.
[0094] Secondly, a multi-target tracking device based on a diffusion model proposed in an embodiment of the present application is described with reference to the accompanying drawings.
[0095] Figure 6 It is a block diagram of a multi-target tracking device based on a diffusion model according to an embodiment of the present application.
[0096] like Figure 6 As shown, the multi-target tracking device 10 based on the diffusion model includes: a target detection module 100 , a trajectory prediction module 200 and a matching association module 300 .
[0097] Among them, the target detection module 100 is used to obtain the ROI features of the two frames of images to be detected through a preset feature extraction module, and input the ROI features to a preset diffusion detection head module, and perform target detection operations on the two frames of images to be detected to generate target detection results for each target in the two frames of images to be detected, wherein the two frames of images to be detected include the previous frame of the image to be detected and the current frame of the image to be detected.
[0098] The trajectory prediction module 200 is used to calculate the target displacement of each target in the two frames of images to be detected before and after according to the target detection result, and input the target displacement into the preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of the image to be detected, so as to obtain the trajectory prediction result.
[0099] The matching and association module 300 is used to perform matching and association operations on the target detection results and the trajectory prediction results based on a preset hybrid IoU and ReID distance calculation strategy to generate a final target tracking result.
[0100] Optionally, in one embodiment of the present application, the target detection module 100 includes: an input unit, a generation unit, a first calculation unit, an update unit and an iteration unit.
[0101] Among them, the input unit is used to input the previous and next two frames of images to be detected into the feature extraction module to output the ROI features of the target size corresponding to the previous and next two frames of images to be detected.
[0102] The generating unit is used to generate a plurality of random frames in the two frames of images to be detected, and determine the plurality of random frames as initial detection frames corresponding to the two frames of images to be detected.
[0103] The first computing unit is used to construct a diffusion detection head module based on a preset multi-scale attention module, a dynamic convolution module and a data association head module, and input the ROI feature and the initial detection frame into the diffusion detection head module to generate prediction frames corresponding to the two frames of images to be detected before and after through the multi-scale attention module and the dynamic convolution module, and calculate the noise between multiple random frames and the prediction frames.
[0104] The updating unit is used to input the prediction box and the noise into the data association head module to obtain the detection box of the current iteration, and use the detection box of the current iteration to update the initial detection box.
[0105] The iteration unit is used to re-input the ROI features and the updated initial detection frame into the diffusion detection head module to iteratively perform the target detection operation until a preset time step iteration end condition is met to generate a target detection result.
[0106] Optionally, in one embodiment of the present application, the trajectory prediction module 200 includes: a construction unit, a noise adding unit and a second calculation unit.
[0107] Among them, the construction unit is used to construct a decoupled diffusion trajectory prediction module through a preset multiple multi-scale attention modules and motion fusion modules.
[0108] The noise adding unit is used to add Gaussian noise to the target displacement, and input the target displacement after adding Gaussian noise into the decoupled diffusion trajectory prediction module, so as to perform post-diffusion on the target displacement after adding Gaussian noise through the multi-scale attention module and the motion fusion module to generate the target predicted displacement of the current frame to be detected.
[0109] The second calculation unit is used to calculate the trajectory prediction result of the current frame of the image to be detected based on the target prediction displacement and the target detection result of the previous frame of the image to be detected.
[0110] Optionally, in one embodiment of the present application, the matching association module 300 includes: a third calculation unit, a fourth calculation unit and a linear allocation unit.
[0111] Among them, the third calculation unit is used to calculate the IoU distance and ReID distance between the target detection result and the trajectory prediction result, and calculate the distance weight corresponding to each coordinate distance between the IoU distance and the ReID distance.
[0112] The fourth calculation unit is used to calculate the final mixed distance corresponding to the IoU distance and the ReID distance according to the distance weight and the preset weight factor.
[0113] The linear allocation unit is used to calculate the target matching pairs corresponding to the target detection results and the trajectory prediction results based on the final mixed distance and the preset linear allocation algorithm, and update the target trajectory through the target matching pairs to obtain the final target tracking result.
[0114] It should be noted that the above explanation of the embodiment of the multi-target tracking method based on the diffusion model is also applicable to the multi-target tracking device based on the diffusion model of this embodiment, which will not be repeated here.
[0115] According to the multi-target tracking device based on the diffusion model proposed in the embodiment of the present application, it includes a target detection module 100, which is used to obtain the ROI features of the two frames of images to be detected before and after through a preset feature extraction module, and input the ROI features to the preset diffusion detection head module, and perform target detection operations on the two frames of images to be detected before and after to generate target detection results for each target in the two frames of images to be detected before and after, wherein the two frames of images to be detected before and after include the previous frame of images to be detected and the current frame of images to be detected; a trajectory prediction module 200, which is used to calculate the target displacement of each target in the two frames of images to be detected before and after according to the target detection results, and input the target displacement into the preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of images to be detected to obtain the trajectory prediction result; a matching association module 300, which is used to match and associate the target detection results and the trajectory prediction results based on the preset hybrid IoU and ReID distance calculation strategy to generate the final target tracking result. The present application realizes accurate tracking of multiple moving targets in complex motion states by combining the diffusion detection module and the diffusion trajectory prediction module, and can be applied to security, rescue, sports events, automatic driving and other fields.
[0116] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0117] A memory 701 , a processor 702 , and a computer program stored in the memory 701 and executable on the processor 702 .
[0118] When the processor 702 executes the program, the multi-target tracking method based on the diffusion model provided in the above embodiment is implemented.
[0119] Furthermore, the electronic device further comprises:
[0120] The communication interface 703 is used for communication between the memory 701 and the processor 702 .
[0121] The memory 701 is used to store computer programs that can be executed on the processor 702 .
[0122] The memory 701 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0123] If the memory 701, the processor 702 and the communication interface 703 are implemented independently, the communication interface 703, the memory 701 and the processor 702 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0124] Optionally, in a specific implementation, if the memory 701, the processor 702 and the communication interface 703 are integrated on a chip, the memory 701, the processor 702 and the communication interface 703 can communicate with each other through an internal interface.
[0125] The processor 702 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0126] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned multi-target tracking method based on a diffusion model.
[0127] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned multi-target tracking method based on the diffusion model.
[0128] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0129] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0130] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0131] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.
[0132] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0133] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0134] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0135] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A multi-target tracking method based on a diffusion model, characterized in that: The following steps are involved: The ROI features of the two frames of images to be detected are obtained by a preset feature extraction module, and the ROI features are input into a preset diffusion detection head module, and a target detection operation is performed on the two frames of images to be detected to generate a target detection result for each target in the two frames of images to be detected, wherein the two frames of images to be detected include a previous frame of image to be detected and a current frame of image to be detected; Calculating the target displacement of each target in the two frames of images to be detected based on the target detection result, and inputting the target displacement into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of the image to be detected, so as to obtain a trajectory prediction result; Based on the preset hybrid IoU and ReID distance calculation strategy, the target detection result and the trajectory prediction result are matched and associated to generate the final target tracking result.
2. The method according to claim 1, characterized in that The ROI features of the two frames of images to be detected are obtained by a preset feature extraction module, and the ROI features are input into a preset diffusion detection head module, and a target detection operation is performed on the two frames of images to be detected to generate a target detection result for each target in the two frames of images to be detected, including: Input the two frames of images to be detected to the feature extraction module to output ROI features of the target size corresponding to the two frames of images to be detected; Generate a plurality of random frames in the two frames of images to be detected, and determine the plurality of random frames as initial detection frames corresponding to the two frames of images to be detected; Based on the preset multi-scale attention module, dynamic convolution module and data association head module, the diffusion detection head module is constructed, and the ROI feature and the initial detection frame are input into the diffusion detection head module, so as to generate the prediction frames corresponding to the two frames of images to be detected before and after through the multi-scale attention module and the dynamic convolution module, and calculate the noise between the multiple random frames and the prediction frame; Inputting the predicted frame and the noise into the data association header module to obtain a detection frame of a current iteration, and using the detection frame of the current iteration to update the initial detection frame; The ROI feature and the updated initial detection frame are re-input into the diffusion detection head module to iteratively perform the target detection operation until a preset time step iteration end condition is met to generate the target detection result.
3. The method according to claim 1, characterized in that The target displacement is input into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of the image to be detected to obtain a trajectory prediction result, including: The decoupled diffusion trajectory prediction module is constructed by using a plurality of preset multi-scale attention modules and motion fusion modules; Adding Gaussian noise to the target displacement, and inputting the target displacement after adding Gaussian noise into the decoupled diffusion trajectory prediction module, so as to perform post-diffusion on the target displacement after adding Gaussian noise through the multi-scale attention module and the motion fusion module to generate the target prediction displacement of the current frame to be detected image; Based on the target predicted displacement and the target detection result of the previous frame of the image to be detected, the trajectory prediction result of the current frame of the image to be detected is calculated.
4. The method according to claim 3, characterized in that The matching and associating operation is performed on the target detection result and the trajectory prediction result based on the preset hybrid IoU and ReID distance calculation strategy to generate the final target tracking result, including: Calculating the IoU distance and the ReID distance between the target detection result and the trajectory prediction result, and calculating the distance weight corresponding to each coordinate distance between the IoU distance and the ReID distance; Calculating a final mixed distance corresponding to the IoU distance and the ReID distance according to the distance weight and a preset weight factor; Based on the final hybrid distance and a preset linear assignment algorithm, a target matching pair corresponding to the target detection result and the trajectory prediction result is calculated, and the target trajectory is updated through the target matching pair to obtain the final target tracking result.
5. A multi-target tracking device based on a diffusion model, characterized in that: include: A target detection module, used for obtaining ROI features of the two frames of images to be detected through a preset feature extraction module, inputting the ROI features into a preset diffusion detection head module, and performing target detection operations on the two frames of images to be detected to generate target detection results for each target in the two frames of images to be detected, wherein the two frames of images to be detected include a previous frame of image to be detected and a current frame of image to be detected; A trajectory prediction module, used for calculating the target displacement of each target in the two frames of images to be detected before and after according to the target detection result, and inputting the target displacement into a preset decoupled diffusion trajectory prediction module to predict the target trajectory of each target in the previous frame of the image to be detected, so as to obtain a trajectory prediction result; The matching and association module is used to perform matching and association operations on the target detection result and the trajectory prediction result based on a preset hybrid IoU and ReID distance calculation strategy to generate a final target tracking result.
6. The device according to claim 5, characterized in that The target detection module comprises: An input unit, used for inputting the two frames of images to be detected to the feature extraction module to output ROI features of the target size corresponding to the two frames of images to be detected; A generating unit, configured to generate a plurality of random frames in the two frames of images to be detected, and determine the plurality of random frames as initial detection frames corresponding to the two frames of images to be detected; A first calculation unit is used to construct the diffusion detection head module based on a preset multi-scale attention module, a dynamic convolution module and a data association head module, and input the ROI feature and the initial detection frame into the diffusion detection head module, so as to generate the prediction frames corresponding to the two frames to be detected before and after through the multi-scale attention module and the dynamic convolution module, and calculate the noise between the multiple random frames and the prediction frame; An updating unit, configured to input the prediction frame and the noise into the data association header module to obtain a detection frame of a current iteration, and update the initial detection frame using the detection frame of the current iteration; The iteration unit is used to re-input the ROI feature and the updated initial detection frame into the diffusion detection head module to iteratively perform the target detection operation until a preset time step iteration end condition is met to generate the target detection result.
7. The device according to claim 5, characterized in that The trajectory prediction module comprises: A construction unit, used to construct the decoupled diffusion trajectory prediction module through a plurality of preset multi-scale attention modules and motion fusion modules; a noise adding unit, configured to add Gaussian noise to the target displacement, and input the target displacement after adding Gaussian noise into the decoupled diffusion trajectory prediction module, so as to perform post-diffusion on the target displacement after adding Gaussian noise through the multi-scale attention module and the motion fusion module, and generate the target prediction displacement of the current frame to be detected image; The second calculation unit is used to calculate the trajectory prediction result of the current frame of the image to be detected based on the target predicted displacement and the target detection result of the previous frame of the image to be detected.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the multi-target tracking method based on the diffusion model as described in any one of claims 1 to 4.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the multi-target tracking method based on a diffusion model as described in any one of claims 1 to 4.
10. A computer program product, comprising a computer program, characterized in that The computer program is executed to implement the multi-target tracking method based on the diffusion model as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Multi-target tracking method for synchronous moving target
CN113723190A
Space-time fusion multi-target tracking method, device, equipment and medium
CN117314965A
Infrared target tracking method based on diffusion model joint association mechanism
CN117671522A
Vehicle multi-target intelligent dynamic fusion tracking method
CN118608563A
Multi-target tracking method for sea surface scene
CN118941595A