Multi-camera multi-target tracking-oriented artificial intelligence model training data set generation method

By generating diffusion models and target motion simulation technology, the problem that the existing technology cannot generate the appearance and motion trajectory of the three-dimensional model is solved, and efficient and customized cross-camera multi-objective tracking data set generation is achieved, supporting automated large-scale data set generation.

CN120163996APending Publication Date: 2025-06-17SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311714187.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art cannot generate model appearance data and model motion trajectory in three-dimensional space, and cannot be directly used for data set generation in target tracking model training scenarios, and data sets and labeling require high costs.

Method used

By constructing and training, a diffusion model is generated, a diversified three-dimensional map of tracking targets is generated, and a physical model of the target motion trajectory is established. Combined with the Unity 3D engine to simulate the target tracking scene, and a test data set picture frame and target labeling data are generated.

Benefits of technology

It realizes end-to-end generation of multi-target tracking data sets and labeling data across cameras without relying on external data injection, supports customization of target number, appearance and camera position parameters, and outputs target box labeling data that complies with CoCoVid or MOT standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163996A_ABST
    Figure CN120163996A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic generation method for a multi-camera multi-target tracking data set. A diffusion model is generated through construction and training and is used for generating diversified tracking target three-dimensional chartlets; the end-to-end cross-camera multi-target tracking model training data set generation method comprises the steps of establishing a target motion trail simulation physical model, modeling a target tracking scene through Unity three-dimensional engine simulation, and generating a corresponding test data set picture frame and target annotation data, and generating an end-to-end cross-camera multi-target tracking model training data set by means of an artificial intelligence generation technology. And a large number of various data sets required by target tracking model training can be quickly generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of image processing, specifically a method for generating an artificial intelligence model training dataset (AIMCG, Artificial Intelligence MTA Content Generation) for multi-camera multi-object tracking. Background Art

[0002] The cross-camera multi-object tracking dataset is the basis for training the object tracking model, but the cost of shooting and annotating the dataset is relatively high. Existing image generation technologies can only be used to generate image samples in the field of defective samples, and cannot generate the appearance data of the model in three-dimensional space or the movement trajectory of the model, and cannot be directly used for generating the dataset in the object tracking model training scenario. Summary of the Invention

[0003] In view of the above deficiencies of the prior art, the present invention proposes an automatic generation method for a multi-camera multi-object tracking dataset, which uses artificial intelligence generation technology to generate an end-to-end cross-camera multi-object tracking model training dataset, and can quickly generate a large number of diverse datasets required for training the object tracking model.

[0004] The present invention is realized through the following technical solutions:

[0005] The present invention relates to a method for generating a cross-camera multi-object tracking dataset. A generative diffusion model is constructed and trained to generate diverse three-dimensional maps of tracking objects; a physical model for simulating the target movement trajectory is established, and the target tracking scenario is modeled through the Unity 3D engine, and the corresponding test dataset frame and target annotation data are generated. Specifically, the generative diffusion model (SDM) is used to assign clothing maps with different materials, colors, and features to randomly generated tracking objects, so that the generated data shows a certain degree of diversity; then, according to the characteristics of the target movement, a physical model of the target movement including displacement, speed, and acceleration is established, and a random perturbation factor is combined to simulate the random movement trajectory of the target; then, based on the Unity graphics 3D engine, a character model is constructed and the texture material is constructed using the map generated by the artificial intelligence generated content (AIGC) model, and an animation is generated in combination with the movement simulation physical model; finally, computer graphics related algorithms are used to simulate the camera shooting content from a fixed position perspective, and the image is stored as a dataset; at the same time, the position of the projected target box of the corresponding target in each camera is calculated, and the position information is saved as annotation data.

[0006] The present invention relates to a cross-camera multi-object tracking dataset generation system for implementing the above method, including: a clothes texture map generation unit (CTMGU, Clothes Texture Material Generating Unit), a pedestrian trajectory simulation unit (RMTS, Randomized Motion Tracklet Simulator), and an algorithm main module, where: the clothes texture map generation unit is based on a generative diffusion model implemented based on PyTorch, and generates texture maps according to random latent parameters without relying on external input image samples or text prompts; the pedestrian trajectory simulation unit supports approximate solutions with any number of iterations for the motion state of pedestrian targets in any state according to a differential equation model describing the walking trajectories of pedestrians; the algorithm main module communicates with the CTMGU and RMTS modules and requests texture map data and motion trajectory data as needed, and uses a renderer implemented by the Unity 3D graphics engine to complete 3D modeling visualization, intercepts camera frame data at corresponding positions, calculates the target box coordinates at corresponding positions, and generates a dataset output.

[0007] The described clothes texture map generation unit is based on a diffusion generative network. During its training process, after preprocessing, expanding, and scaling the incoming training dataset, a Gaussian diffuser is used to gradually increase the noise interference factor, and a U-Net reverse diffusion neural network model with extremely strong feature extraction ability is gradually trained to make it learn the ability to remove noise interference and restore effective features; after the training is completed, a full-noise random tensor is generated by the inference process manager, and the U-Net model is repeatedly used for denoising inference with this as the input, and the obtained results are output as a new data generation set.

[0008] The differential equation model describing the walking trajectories of pedestrians uses, but is not limited to, the motion differential equation recorded in the Dirk Helbing social force model [Helbing D, Molnar P. Social force model for pedestrian dynamics[J]. Physical review E, 1995, 51(5):4282]. In view of the situation where pedestrians gather in groups and act together in real life, a pedestrian trajectory simulation algorithm with random perturbation is proposed by considering three factors: obstacle avoidance, maintaining distance between individuals, and pedestrians forming groups, and an approximate numerical iteration method is used to approximately solve the differential equation to obtain a simulation result containing multiple trajectories. Technical effects

[0009] The present invention can generate an end-to-end cross-camera multi-object tracking dataset and annotated data without relying on external data injection, and supports customization of the number of tracking objects, the appearance of tracking objects, and the position parameters of the shooting cameras; it supports outputting object box annotation data that conforms to the CoCoVid standard or the MOT standard. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a schematic diagram of the overall architecture of the present invention;

[0011] Figure 2 It is a schematic diagram of the internal process of CTMGU in the present invention;

[0012] Figure 3 It is a schematic diagram of the internal diffusion model network generated by CTMGU in the present invention;

[0013] Figure 4 It is a schematic diagram of the texture result generated by CTMGU in the present invention;

[0014] Figure 5 It is a schematic diagram of the result of simulating and generating pedestrian trajectories by RMTS in the present invention;

[0015] Figure 6 It is a schematic diagram of the dataset generated by the main body module in the present invention;

[0016] Figure 7 It is a comparison chart of the change of the loss rate during the training process between the dataset generated in the present invention and the existing MTA dataset;

[0017] Figure 8 It is a UML sequence diagram of the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] As Figure 1 shown, this embodiment relates to a multi-camera multi-object tracking dataset generation system for road monitoring, which uses the Python language to implement the communication protocol layer and adopts RESTful API defined based on HTTP to communicate with other modules. The main body module of the dataset generation method is implemented based on the C# language. The system includes: CTMGU, RMTS, and a three-dimensional scene modeling and output unit, where: CTMGU contains a diffusion model inference module implemented based on PyTorch, which generates texture map pedestrian trajectory simulations according to internal random hidden parameters; RMTS recursively solves the motion state of pedestrian targets in any state through an SFM differential equation solving module that supports any number of iteration times; the three-dimensional scene modeling and output unit uses a renderer implemented based on the Unity3D graphics engine to achieve three-dimensional modeling visualization, intercepts the frame data of the images captured by the corresponding cameras, calculates the corresponding object box coordinates, and outputs them as the generated dataset.

[0019] Preferably, since the generation process takes a long time, a texture map buffer pool of a fixed size is provided in the CTMGU to reduce possible latency problems when there are a large number of texture map generation requests.

[0020] The method for generating a data set according to the above system in this embodiment includes the following steps:

[0021] Step 1: Construct a generative diffusion network model as shown in Figure 2 and train it using a publicly available texture material data set and labeled sample data so that the model has the ability to generate new texture maps according to the prompt words, specifically including:

[0022] Step 1.1: For each image in the original labeled data set, scale it to the standard texture map size of 1024 by 1024, and perform normalization operations on its contrast and brightness to ensure the overall unity of the data set samples;

[0023] Step 1.2: Perform the following iterative process on each image in the data set: Add a certain amount of Gaussian noise to the image, and use the obtained result to train the SDM model to restore the influence of the noise, and measure its performance by the difference between the restored noise image and the image before processing. Repeat such iterative operations until the image features are unrecognizable due to excessive noise. Let the noise-free sample vector in the initial state be x0, then the sampling result after any t steps can be calculated by iteratively adding noise according to the following recurrence formula: where: x t+1 is the sample distribution data after t + 1 steps of iteration, R |x| is the noise vector generated according to the distribution before diffusion, and λ is the diffusion factor that adjusts the mean value of the distribution before and after the diffusion process. The noise added in each step should conform to the Gaussian distribution given by the following formula: where β1, β2,...β n are always monotonically increasing diffusion factors, which describe that as the diffusion process of adding noise progresses, the obtained distribution mean value will gradually decrease, the variance will gradually increase, and the effective information carried by the samples will gradually be lost.

[0024] Step 1.3: After completing the forward diffusion training, perform reverse diffusion denoising operations on the noise image lacking features by the trained SDM model, and repeat this operation until a stable image sample is obtained. Specifically: where: x t is the sample distribution data after t steps of iteration, x0 is the original noise-free sample data; a t and 1 - a t are respectively the true values of the approximate mean and approximate variance of the noise distribution during the diffusion process.

[0025] This equation describes the distribution relationship between the sample \(x_t\) after \(t\) iterations and the original sample \(x_0\). When the value of \(\beta\) i is small, the reverse process of the diffusion process still approximately conforms to the Gaussian distribution. Therefore, the mean \(\mu\) of the diffused distribution can be sampled and estimated t and the variance \(\Sigma\) t , and the original sample is attempted to be restored through the reverse Markov process. Specifically: where: \(x\) t+1 is the sample distribution data after \(t + 1\) steps of iteration, and \(x\) t is the sample distribution data after \(t\) steps of iteration; \(\mu\) t and \(\Sigma\) t respectively represent the estimated values of the mean and variance of the diffused distribution, and \(N\) represents that the original distribution approximates a normal distribution.

[0026] The described reverse Markov process only approximately holds when the iteration step is small, and there is no analytical solution to completely restore the original state.

[0027] Such as Figure 3 The U-Net neural network model shown can learn and train the reverse Markov process described above, and can restore the original distribution with certain characteristics from the estimated values of the mean \(\mu\) and variance \(\Sigma\), so that the model has the ability to generate new samples. During the learning process, the maximum likelihood method is used to train the model, and the cross-entropy value between the true distribution of the sample and the predicted distribution of the model is used as the objective function. Specifically: where: \(p_ heta\) (x0) is the difference loss obtained by comparing the sample with the true value pixel by pixel; \(E\) q(x0) is the entropy value describing the randomness degree of the true distribution \(x_0\). \(L\) is the calculated loss function value.

[0028] Step 2: Simulate the random movement of the tracking target. Based on the modeling of the target's facing direction, position, movement speed, and movement acceleration, combined with the social force model, simulate the random walk of the target to generate a random target motion simulation trajectory. Specifically include:

[0029] Step 2.1: By applying an improved algorithm based on the social force model, use the target forward direction, interest direction, and distance factor from the wall obstacle to simulate the random movement of the tracking target. The differential equation describing the acceleration during the pedestrian movement (i.e., the change direction of the velocity vector at a certain instant) is: Where: F-target is the invariant target position information of each target, representing the destination coordinates during the target's wandering movement; F-obstacle describes the speed change of the target's action trajectory affected by obstacles, and its magnitude is negatively correlated with the distance between the target and the obstacle, and can be ignored when the distance exceeds a certain value; F-repulsive describes the characteristic that the target tends to keep a distance from other targets during the movement, that is, when the target may be too close to other targets during the movement, they avoid and move away from each other; F-group describes the characteristics of targets in groups getting closer to each other but maintaining a certain distance and moving in the same direction. Each of the above force components is vector-adjusted weights and combined to obtain the overall force situation of the target.

[0030] Step 2.2: Calculate the F-group synthesis factor based on the common forward, approaching, and repulsive effects among groups of people. This component can be further divided into the following three parts: One is the force F-visual for correcting the movement direction and the direction of companions. This vector is used to describe that groups of targets tend to keep their companions within the field of vision during the movement, and at the same time, the target will try to make its movement direction parallel to the line-of-sight direction; the second is the force F-attr that makes companion targets approach each other; the third is the force F-grepl that makes companion targets keep a distance from each other and avoid getting too close. They respectively reflect the situation where grouped pedestrian targets approach each other but do not get too close. Generally speaking, the repulsive effect F-grepl between individuals within a group is much weaker than the repulsive force F-repl between unspecified individuals. Finally, use weight factors to combine them as the instantaneous acceleration estimate value of each moving target in the current state, specifically: Where: α i is the angle between the movement direction of target i and the line-of-sight direction; v i is the instantaneous velocity vector of target i's movement; U i is the direction vector pointing to the companion target that target i is currently following; q A is the distance between target i and the companion target currently. W ik is the direction vector pointing to target k that does not belong to the companion; qR is the distance between target i and that target k.

[0031] Step 2.3: According to the differential equation in Step 2.1, use the iterative approximation method to iteratively solve the position, velocity, and acceleration triples in the current state, and obtain the forward trajectories of each target. The shape of the generated example trajectories is shown in the figure, where pedestrians 0, 1 and 3, 4, 5 move forward in groups respectively.

[0032] Step 3: Based on the 3D engine, simulate and model the target tracking scenario, and generate the frame of the camera captured image and annotation data at the corresponding position, specifically including:

[0033] Step 3.1: Use the 3D graphics engine Unity 3D to build the simulation target tracking scenario environment as the background scenario for generating the dataset;

[0034] Step 3.2: Construct an HTTP Restful data communication message interface for communicating with the appearance texture mapping generation module and the motion trajectory generation module to obtain the corresponding resources.

[0035] Table 1 HTTP Restful data communication message interface

[0036] Step 3.3: By generating multiple Camera 3D objects in Unity 3D, realize the function of capturing images from cameras at different angles and saving them as pictures; by recording the spatial x, y, z coordinates of each camera and its viewing coordinates tx, ty, tz, obtain the two-dimensional plane projection position of the corresponding target through the following spatial projection matrix transformation.

[0037] The spatial projection matrix transformation mentioned above specifically includes:

[0038] a) For a general target located at the position (x, y, z) in the three-dimensional space, represent its coordinates in the form of a 4×1-dimensional vector: where: p is the coordinate vector of the target; x, y, z are the coordinates of the original target on the X-axis, Y-axis, and Z-axis in the three-dimensional space respectively; use the T operator to transpose it into the form of a column vector for storage.

[0039] b) Consider another camera with its viewing origin located at the coordinates (M x , M y , M z ), and its corresponding perspective transformation matrix can be written as where: M x , M y , M zThey are the coordinates of the camera origin on the X-axis, Y-axis, and Z-axis in the three-dimensional space respectively; l, r, t, and b represent left, right, top, and bottom respectively, corresponding to the positions of the four clipping plane edges within the field of view of the camera captured image; f and n represent the nearest and farthest ranges of the spatial plane that the camera can capture. Performing matrix multiplication on the projection vector in the left part of the formula and the camera position vector can obtain the projection transformation matrix P of the camera. M 。

[0040] c) The target coordinate position in the three-dimensional space can be transformed into the (x, y) coordinate position in the two-dimensional plane by using the matrix multiplication shown below. Specifically, where: P M is the projection transformation matrix of the above camera, x i , y i , z i are the coordinates of the target on the X-axis, Y-axis, and Z-axis in the three-dimensional space before transformation; x i ’, y i ’ are the two-dimensional space coordinates after transformation; d i ’, z i ’ are used to describe the front-back order relationship of the coordinate points after the projection process, which is not needed here.

[0041] d) Performing the above transformation projection matrix calculation on the target coordinates (xi, yi, zi) that appear in each camera can generate the coordinate data for the target box annotation.

[0042] Step 4: As Figure 8 shown, use the generative diffusion network model trained in Step 1 to generate several texture maps, and use the target motion data simulation unit in Step 2 to continuously generate target motion and position information. Input the model texture map file and the trajectory coordinate sequence into the Unity 3D renderer to generate the corresponding frame of the image, and use the method described in Step 3 to output the bounding box annotation data of the corresponding target in the image, which can be used as the dataset required for training the multi-target tracking model across cameras.

[0043] Through specific actual experiments, in a computer environment equipped with Intel Xeon Gold 5222 (4c @ 3.8GHz), Nvidia RTX3090 24G GPU, and 32GB x8 ECC DDR4 2933MHz memory, the multi-target tracking model across cameras is trained using the dataset generated by this method and the standard MTA dataset with 6 Epoch training parameters and an initial learning rate of 0.01 respectively, and the IDP, IDR, and IDF1 measurement criteria of the trained result model on the standard dataset are calculated and compared. The indicators are shown in Table 2

[0044] Table 2

[0045] In the table, IDP (Identification Precision) represents the tracking accuracy corresponding to each global target, focusing on measuring the detection quality of successfully detected trajectories; IDR (Identification Recall) represents the overall recall rate of global targets, focusing on measuring the success rate of overall target detection. The IDF1 metric is defined as the harmonic mean of IDP and IDR, and the larger the value, the higher the overall quality of the cross-camera target tracking trajectory.

[0046] By recording the comparison graph of the respective loss rates of the dataset (AIMCG) generated by this method and the existing MTA dataset during the training process, it is observed that the change processes of the two curves are basically the same, and they finally converge to similar loss levels, indicating that both datasets can enable the model to effectively learn and converge, as Figure 7 shown.

[0047] As Figure 4 shown, the results shown in the sample graph of the target appearance map sample data generated by this method are the stage generation result examples after 10k, 20k, 30k, 40k, 50k, 60k, and 70k training iterations respectively; the sample of the target motion trajectory data generated by this method is as Figure 5 shown, where the red, green, and blue trajectories correspond to groups of moving targets respectively; the sample image of the dataset for training the cross-camera multi-target tracking model generated by this method is as Figure 6 shown.

[0048] Compared with the standard MTA dataset provided in the literature [Kohl P, Specker A, Schumann A, et al. The mta dataset for multi-target multi-camera pedestrian tracking by weighted distance aggregation [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition Workshops. 2020: 1042-1043.], the performance of the cross-camera multi-target tracking WDA model trained by this method in the standard test set is basically similar. IDP and IDF1 are respectively 2.7 points and 1.2 points behind, and the difference in IDR is within the error range; this indicates that the dataset generated by this method has high practicality, and this method has certain practical value in the field of model training and evaluation.

[0049] The above data indicates that the training dataset for the cross-camera multi-object tracking model generated by the present method has a relatively similar effect to the original MTA dataset during the model training process, and the differences shown by various evaluation indicators are all within 4 units; compared with the meta-dataset, the present method can achieve computer-automated large-scale dataset generation without relying on external data simulation, manual modeling, target annotation, etc., and has high practical value.

[0050] In summary, the present invention realizes an end-to-end system for generating a training dataset for a cross-camera multi-object tracking model through artificial intelligence generation technology and physical motion simulation technology based on the social force model. By generating maps through the generative diffusion model, the appearance diversity of the obtained target samples is ensured, and the trajectory data conforming to real motion characteristics is generated by the target motion data simulation unit. Through the three-dimensional scene simulation technology that can access external data sources, a large number of diverse datasets required for target tracking model training are quickly generated.

[0051] Those skilled in the art can make local adjustments to the above specific implementation in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the constraints of the present invention.

Claims

1. A method for generating a cross-camera multi-object tracking dataset, characterized in that, Construct and train an SDM to generate diverse 3D maps of tracking targets; establish a physical model for target motion trajectories, and use the Unity 3D engine to simulate and model the target tracking scenario, and generate corresponding test dataset frame images and target annotation data.

2. The method for generating a cross-camera multi-object tracking dataset according to claim 1, characterized in that, Use a generative diffusion model to assign clothing maps with different materials, colors, and features to randomly generated tracking targets. Based on the characteristics of target motion, establish a physical model of target motion including displacement, velocity, and acceleration, and combine it with a random perturbation factor to simulate the random walking motion trajectory of the target; then construct a human model based on the Unity 3D graphics engine, use the maps generated by the artificial intelligence-generated content model to construct its texture materials, and combine the motion simulation physical model to generate animations; finally, use computer graphics-related algorithms to simulate the content captured by a camera from a fixed position perspective, and store the image as a dataset; at the same time, calculate the projected target box positions of the corresponding targets in each camera, and save the position information as annotation data.

3. The method for generating a cross-camera multi-object tracking dataset according to claim 1 or 2, characterized in that, specifically Including: Step 1: Construct a generative diffusion network model and train it using a publicly available map material dataset and annotation sample data to enable the model to generate new maps according to prompts, specifically including: Step 1.1: For each image in the original annotation dataset, scale it to a standard map size of 1024 by 1024, and perform normalization operations on its contrast and brightness. Step 1.2: Perform the following iterative process on each image in the dataset: Add a certain amount of Gaussian noise to the image, and use the obtained result to train the SDM model to restore the influence of the noise, and measure its performance by the difference between the restored noisy image and the image before processing; Repeat such iterative operations until the image features become unrecognizable due to excessive noise; Let the noise-free sample vector in the initial state be x0, then add noise to it iteratively according to the following recurrence formula to calculate the sampling result after any t steps: where: x t+1 is the sample distribution data after t+1 steps of iteration, R |x| is the noise vector generated according to the distribution before diffusion, and λ is the diffusion factor that adjusts the mean value of the distribution before and after the diffusion process; Step 1.3: After completing the forward diffusion training, perform reverse diffusion denoising operation on the noise image lacking features by the trained SDM model, and repeat this operation until a stable image sample is obtained. Specifically: where: x t is the sample distribution data after t-step iteration, and x0 is the original noise-free sample data; a t and 1 - a t are the true values of the approximate mean and approximate variance of the noise distribution during the diffusion process, respectively. Step 2: Simulate the random motion of the tracking target. Based on the modeling of the target's facing direction, position, motion speed, and motion acceleration, combine the social force model to simulate the random walking of the target, and generate a random target motion simulation trajectory, specifically including: Step 2.1: By applying the improved algorithm based on the social force model, simulate the random motion of the tracking target using the target forward direction, the direction of interest, and the distance factor from the wall obstacle; the differential equation describing the acceleration during the pedestrian movement (i.e., the change direction of the velocity vector at a certain instant) is as follows: where: F-target is the invariant target position information of each target, representing the destination coordinates during the target wandering motion; F-obstacle describes the velocity change of the target's action trajectory affected by the obstacle, and its magnitude is negatively correlated with the distance between the target and the obstacle, which can be ignored when the distance exceeds a certain value; F-repulsive describes the characteristic that the target tends to keep a distance from other targets during the movement, that is, when the target may be too close to other targets during the movement, they avoid and move away from each other; F-group describes the characteristic that the targets in a group approach each other but keep a certain distance and move forward in the same direction; each of the above force components is adjusted by the vector weights and combined to obtain the overall force situation of the target. Step 2.2: Calculate the F-group synthesis factor based on the common forward movement, approaching, and repulsion effects among groups of people; this component is further divided into the following three parts: one is the force F-visual for correcting the movement direction and the direction of companions, which is used to describe that the grouped targets tend to keep their companions within the field of view during movement, and at the same time the targets will try to make their movement direction parallel to the line-of-sight direction; the second is the force F-attr that makes the companion targets approach each other; the third is the force F-grepl that keeps the companion targets at a distance and avoids getting too close; they respectively reflect the situation where the grouped pedestrian targets approach each other but do not get too close; generally speaking, the repulsion effect F-grepl among individuals within the group is much weaker than the repulsion force F-repl among unspecified individuals; finally, use The weight factor to combine them as the instantaneous acceleration of each moving target in the current state Degree estimate value, specifically: Where: α i Is the angle between the movement direction of target i and the line-of-sight direction; v i Is the instantaneous velocity vector of the movement of target i; U i Is the direction vector pointing to the companion target currently followed by target i; q A Is the distance between target i and the companion target currently; W ik Is the direction vector pointing to target k that does not belong to the companion; qR is the distance between target i and this target k; Step 2.3: According to the differential equation in Step 2.1, use the iterative approximation solution method to iteratively solve the position, velocity, and acceleration triple in the current state to obtain the forward trajectory of each target; the shape of the generated example trajectory is shown in the figure, where pedestrians 0, 1 and 3, 4, 5 move forward in groups respectively. Step 3: Based on the 3D engine, simulate and model the target tracking scenario, and generate the frame images and annotation data captured by the camera at the corresponding position, specifically including: Step 3.1: Use the 3D graphics engine Unity 3D to build a simulation target tracking scenario environment as the background scenario for generating the dataset. Step 3.2: Construct an HTTP Restful data communication message interface for communicating with the appearance texture map generation module and the motion trajectory generation module to obtain corresponding resources. Step 3.3: By generating multiple Camera 3D objects in Unity 3D, implement the function of capturing images from cameras at different angles and saving them as pictures; by recording the spatial x, y, z coordinates of each camera and its viewing coordinates tx, ty, tz, obtain the two-dimensional plane projection position of the corresponding target through the following spatial projection matrix transformation. Step 4: Use the generative diffusion network model trained in Step 1 to generate a number of texture maps, and use the target motion data simulation unit in Step 2 to continuously generate target motion and position information. Pass the model texture map file and the trajectory coordinate sequence into the Unity 3D renderer to generate the corresponding frame of the image, and use the method described in Step 3 to output the bounding box annotation data of the corresponding target in the image, which can be used as the dataset required for training the multi-target tracking model across cameras.

4. The method for generating a cross-camera multi-object tracking dataset according to claim 3, characterized in that, In the above-mentioned step 1.2, the added noise in each step should conform to the Gaussian distribution given by the following formula: where β1, β2,...β n are diffusion factors that are always monotonically increasing. They describe that as the diffusion process of adding noise progresses, the mean value of the obtained distribution gradually decreases, the variance gradually increases, and the effective information carried by the samples gradually disappears.

5. The method for generating a cross-camera multi-object tracking dataset according to claim 3, characterized in that, For the reverse diffusion denoising operation, the distribution relationship between the sample \(x_t\) after \(t\) iterations and the original sample \(x_0\) still approximately conforms to the Gaussian distribution. Therefore, the mean \(\mu\) of the distribution after diffusion t and the variance \(\Sigma\) t can be sampled and estimated, and the original sample can be attempted to be restored through the reverse Markov process. Specifically: where: \(x\) t+1 is the sample distribution data after \(t + 1\) steps of iteration, and \(x\) t is the sample distribution data after \(t\) steps of iteration; \(\mu\) t and \(\Sigma\) t represent the estimated values of the mean and variance of the distribution after diffusion respectively, and \(N\) represents that the original distribution is approximately a normal distribution.

6. The method for generating a cross-camera multi-object tracking dataset according to claim 3, characterized in that, In the training in Step 1.3, the negative logarithm of the cross entropy between the true distribution of the samples and the predicted distribution of the model is used as the objective function, specifically: where: pθ (x0) is the difference loss obtained by comparing the samples with the true values pixel by pixel; E q(x0) is the entropy value describing the degree of irregularity of the true distribution x0, and L is the calculated loss function value.

7. The method for generating a cross-camera multi-object tracking dataset according to claim 3, characterized in that, The spatial projection matrix transformation described above specifically includes: a) For a general target located at the position (x, y, z) in three-dimensional space, its coordinates are represented in the form of a 4×1-dimensional vector: where: p is the coordinate vector of the target; x, y, and z are the coordinates of the original target on the X-axis, Y-axis, and Z-axis in three-dimensional space respectively; and it is stored in the form of a column vector after being transposed using the T operator; b) Consider another camera with its visual origin located at the coordinates (M x , M y , M z ), and write out its corresponding perspective transformation matrix where: M x , M y , and M z are the coordinates of the camera origin on the X-axis, Y-axis, and Z-axis in three-dimensional space respectively; l, r, t, and b represent left, right, top, and bottom respectively, corresponding to the positions of the four clipping plane edges within the field of view of the camera's captured image; f and n represent the nearest and farthest ranges of the spatial plane that the camera can capture; multiplying the projection vector in the left part of the formula by the camera position vector results in the projection transformation matrix P M ; c) The target coordinate position in three-dimensional space can be transformed into the (x, y) coordinate position in the two-dimensional plane by using the matrix multiplication shown below. Specifically, where: P M is the projection transformation matrix of the above camera, x i , y i , z i are the coordinates of the target on the X-axis, Y-axis, and Z-axis in three-dimensional space before transformation; x i ’, y i ’ are the two-dimensional space coordinates after transformation; d i ’, z i ’ are used to describe the sequence relationship of the coordinate points before and after the projection process and are not needed here; d) Calculate the above transformation projection matrix for the target coordinates (xi, yi, zi) that appear in each camera respectively, and the generation of the target box annotation coordinate data can be achieved.

8. A cross-camera multi-object tracking dataset generation system for implementing the method according to any one of claims 1-7, characterized in that It includes: A clothing texture map generation unit, a pedestrian trajectory simulation unit, and an algorithm main body module. Among them: The clothing texture map generation unit is based on a generative diffusion model implemented with PyTorch, and generates texture maps according to random latent parameters without relying on external input image samples or text prompt words; The pedestrian trajectory simulation unit supports approximate solutions for any number of iterations of the motion state of pedestrian targets in any state through a differential equation model that describes the walking trajectory of pedestrians; The algorithm main body module communicates with the CTMGU and RMTS modules and requests texture map data and motion trajectory data as needed, and uses a renderer implemented with the Unity 3D graphics engine to complete 3D modeling visualization, intercepts the camera frame data at the corresponding position, calculates the target box coordinates at the corresponding position, and generates a dataset for output.

9. The cross-camera multi-object tracking dataset generation system according to claim 8, characterized in that The clothing texture map generation unit is based on a diffusion generative network. During its training process, the incoming training dataset is preprocessed, dilated, and scaled, and then a Gaussian diffuser is used to gradually increase the noise interference factor, and a U-Net reverse diffusion neural network model with extremely strong feature extraction ability is gradually trained to make it learn the ability to remove noise interference and restore effective features; After the training is completed, the inference process manager generates a random tensor of full noise, and uses the U-Net model repeatedly for the denoising inference process with this as the input, and outputs the obtained result as a new data generation set.