Cross-camera multi-target online tracking method and system combining Dinov2 and ByteTracker
By combining the multi-objective cross-camera multi-objective online tracking method of Dinov2 and ByteTracker, the inter-class similarity and intra-class variability problems of multi-camera tracking in complex scenarios are solved, and the performance and robustness of multi-objective tracking are improved, and suitable for real-time monitoring scenarios.
Patent Information
- Application Number
- CN202510545775.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-22
AI Technical Summary
The existing multi-camera multi-objective tracking technology faces the problems of similarity between target classes that increase the difficulty of identity identification and intra-class feature variability in complex scenarios. The offline tracking method is not applicable enough in real-time monitoring scenarios, and offline tracking is incompatible with online tracking, resulting in improved tracking performance and robustness.
Combining the multi-objective cross-camera cross-camera tracking method of Dinov2 and ByteTracker, feature extraction is performed through Dinov2 and trajectory generation is performed. Hierarchical clustering and triple loss training model are adopted to optimize data association strategy and trajectory management to achieve cross-camera target tracking.
It improves the performance and robustness of multi-objective tracking across cameras, is suitable for real-time monitoring scenarios, and enhances the generalization ability and tracking accuracy of the model.
Smart Images

Figure CN120525916A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target tracking, and specifically provides a cross-camera multi-target online tracking method and system combining Dinov2 and ByteTracker. Background Art
[0002] With the large-scale deployment of video surveillance systems and the rapid development of video analytics technology, multi-target multi-camera tracking (MTMCT) technology, which uses video data captured by different cameras to detect and track targets of interest in real time, has shown great application potential and value in a variety of fields, including public security, intelligent transportation, intelligent video surveillance, and the military. MTMCT, also known as cross-camera multi-target tracking, typically consists of four submodules: vehicle detection, re-identification (ReID), single-camera multi-target tracking (SCMT), and inter-camera association (ICA). It is an important research direction in the field of computer vision, aiming to achieve continuous and accurate tracking of multiple dynamic targets and cross-camera trajectory association by integrating video data from multiple cameras.
[0003] In complex real-world scenarios, traditional single-camera systems are often constrained by limited target field of view, occlusion, and dynamic scene changes. For example, in the case of occluded vehicle detection, in busy traffic scenes, some vehicles may be severely obscured by the vehicles in front, which brings difficulties to the single-camera multi-target tracking module and makes it difficult to achieve comprehensive monitoring and continuous tracking of the target. In contrast, the multi-camera system not only expands the monitoring range through multiple spatially distributed perspectives, but also provides complementary target information, thereby effectively alleviating the above problems. MTMCT technology seamlessly splices target trajectories from different cameras by innovatively integrating key technologies such as cross-camera data association, target re-identification, and trajectory optimization, thereby achieving continuous tracking and positioning of targets in three-dimensional space. However, this technology still faces two core challenges: first, the high similarity between target classes increases the difficulty of identity recognition; second, due to differences in viewing angles, lighting conditions, and video quality among cameras, there is a high degree of variability in the features within the target class. As Figure 1This is an example of multi-camera multi-target tracking. Four cameras capture images from different viewpoints. The blue and red boxes, connected by dashed lines, represent two targets captured by different cameras, respectively. However, these targets appear very similar and are difficult to distinguish from the target framed by the yellow box. Furthermore, the same target appears significantly different when viewed from different cameras. This inter-target similarity, combined with the viewpoint differences within the targets themselves, complicates the MTMCT task. Offline multi-target multi-camera tracking methods typically employ a phased approach: first, frame-by-frame target detection and feature extraction are performed on the complete video frame sequence acquired previously. Second, target tracking trajectories are generated at the individual camera level. Finally, feature matching and clustering algorithms are used to cluster the trajectories of the same target from multiple cameras into the same identity.
[0004] Offline tracking requires pre-acquisition of the complete video from each camera and extraction of target features across the entire video sequence. Compared to offline tracking, online tracking can only utilize the video frame sequence prior to the current moment, requiring frame-by-frame target detection, tracking, and cross-camera association. In online tracking, the detector output (typically a bounding box) is used as the minimum unit for matching, rather than complete trajectory information. Furthermore, offline methods pre-set additional constraints for vehicle tracking based on pre-acquired video information, such as travel time, vehicle direction, entry and exit zones, camera linkage models, road topology, and traffic regulations. For example, based on the camera positions and the spatiotemporal characteristics of vehicle travel, only features from adjacent cameras are matched during cross-shot tracking. The time when a vehicle passes the front and rear cameras is calculated based on the reasonableness of time and distance, thereby eliminating unreasonable data. Furthermore, based on vehicle direction and the defined entry and exit zones of intersections, only vehicle information on main roads is retained to eliminate redundant data and achieve better performance. While these factors give offline methods advantages in post-processing and offer superior accuracy, their incompatibility with online applications limits their applicability in monitoring scenarios such as real-time traffic. This is particularly true at intersections where accidents and collisions are frequent and vehicle behavior is diverse and complex, and in key areas such as prisons and airports, where security is paramount. Online, real-time monitoring by multiple cameras is often required, presenting new challenges. In practical applications, the implementation of offline methods can be problematic, often requiring online operation. However, research on online MTMCT algorithms remains underdeveloped, with only a few methods proposed, and their tracking performance and robustness remain to be improved. Summary of the Invention
[0005] In order to solve the above problems, in a first aspect of the present invention, a method for online tracking of multiple targets across cameras combining Dinov2 and ByteTracker is provided, the method comprising:
[0006] Dinov2 is used to extract features from the image of each camera, and ByteTracker is used to track multiple targets with a single camera to obtain trajectories.
[0007] Hierarchical clustering is used to cluster the trajectories in all cameras. If the maximum number of trajectories in a cluster belonging to the same camera is greater than 1, the cluster is further divided into clusters with the maximum number of trajectories. The cluster is matched with the previously identified trajectories based on appearance features, and the average cohesion of the successfully matched clusters is calculated. If there is a mismatched cluster and the cluster cohesion is less than the average, a new ID is assigned to the mismatched cluster. For the elements in the remaining mismatched clusters, the elements are assigned to the successfully identified trajectory that is closest to the element in physical distance.
[0008] Preferably, the feature extraction of the image from each camera using Dinov2 is specifically as follows:
[0009] The patch embed module is used for processing. The patch embed module contains a two-dimensional convolution layer with a 14×14 convolution kernel and a stride of 14, as well as a normal layer, which uniformly converts the input image into a 14×14 patch representation.
[0010] The patches are fed into the ViT Blocks module for feature extraction, which outputs a feature matrix of dimension T×D. The feature matrix is normalized by the Layer Norm module to generate a 1×n-dimensional feature vector; where T represents the number of channels, D represents the feature vector length, and n is a positive integer.
[0011] Preferably, the training process of Dinov2 is:
[0012] Get Image Hard positive samples of:
[0013]
[0014] Get Image Hard negative samples:
[0015]
[0016] Calculate the hard positive loss:
[0017]
[0018] Calculate the hard negative loss:
[0019]
[0020] Calculate triple loss loss tr=max(0,loss hp -loss hn +γ);
[0021] Calculate the total loss Loss = loss tr +loss ex .
[0022] Among them, loss ex is the cross entropy loss, B i represents the number of images belonging to vehicle i, represents the j-th image of vehicle i, Other images representing vehicle i, represents the images of other vehicles except vehicle i, D represents the distance function, V represents the number of vehicles, and γ is the boundary threshold of the triplet loss.
[0023] Preferably, the ByteTracker is used to track multiple targets with a single camera to obtain trajectories, specifically:
[0024] The trajectories of the previous t-1 frames are obtained, and combined with the state attributes of each trajectory during the tracking process, the trajectories are divided into continuous tracking trajectories, broken trajectories, and initialization trajectories. The continuous tracking trajectory is the trajectory that has been tracked until the current frame. The broken trajectory is the trajectory that failed to be tracked in a previous frame but did not exceed the set tracking discard threshold and is still retained in the matching trajectory pool. The initialization trajectory is the trajectory that was tracked only in the previous frame.
[0025] The NSA Kalman filter is used to predict the trajectory of the previous t-1 frame, and the results detected at time t in the current frame are divided into high-confidence targets and low-confidence targets.
[0026] During the first matching, the continuously tracked, broken, and initialized trajectories are simultaneously matched with the high-confidence target based on IOU and REID cosine distance, and only when the distances of the two are both lower than the set thresholds α1 and α2, further matching is performed using the Hungarian algorithm.
[0027] The first unmatched continuous tracking, broken and initialized trajectories are matched with low confidence targets according to IOU. After the matching process is completed, for more than T max The untracked trajectories in the frame are deleted, and the appearance features of the remaining trajectories are updated using the exponential moving average. The continuously tracked trajectories that are not successfully matched to the detection frame target are switched to broken trajectories, and the broken trajectories that are successfully tracked again are switched to continuously tracked trajectories.
[0028] For the remaining high-confidence targets that failed to find matching tracks during the matching process, they are set as new initialization tracks.
[0029] Preferably, the cosine distance is used to calculate the similarity in the hierarchical clustering:
[0030]
[0031] Among them, f(T i ) and f(T j ) are the appearance feature vectors of the two trajectories respectively.
[0032] In a second aspect of the present invention, a cross-camera multi-target online tracking system combining Dinov2 and ByteTracker is provided, the system comprising:
[0033] The single-camera multi-target tracking module is used to extract features from the image of each camera using Dinov2 and use ByteTracker to track multiple targets with a single camera to obtain trajectories.
[0034] The cross-camera association module is used to cluster the trajectories in all cameras using a hierarchical clustering approach. If the maximum number of trajectories in a cluster belonging to the same camera is greater than 1, the cluster is further divided into clusters with the maximum number of trajectories. The cluster is matched with the previously identified trajectories based on appearance features, and the average cohesion of the successfully matched clusters is calculated. If there are mismatched clusters and the cluster cohesion is less than the average, a new ID is assigned to the mismatched cluster. For the elements in the remaining mismatched clusters, the elements are assigned to the track that is closest to the element and has been successfully identified.
[0035] Preferably, the feature extraction of the image from each camera using Dinov2 is specifically as follows:
[0036] The patch embed module is used for processing. The patch embed module contains a two-dimensional convolution layer with a 14×14 convolution kernel and a stride of 14, as well as a normal layer, which uniformly converts the input image into a 14×14 patch representation.
[0037] The patches are fed into the ViT Blocks module for feature extraction, which outputs a feature matrix of dimension T×D. The feature matrix is normalized by the Layer Norm module to generate a 1×n-dimensional feature vector; where T represents the number of channels, D represents the feature vector length, and n is a positive integer.
[0038] Preferably, the training process of Dinov2 is:
[0039] Get Image Hard positive samples of:
[0040]
[0041] Get Image Hard negative samples:
[0042]
[0043] Calculate the hard positive loss:
[0044]
[0045] Calculate the hard negative loss:
[0046]
[0047] Calculate triple loss loss tr =max(0,loss hp -loss hn +γ).
[0048] Calculate the total loss Loss = loss tr +loss ex .
[0049] Among them, loss ex is the cross entropy loss, B i represents the number of images belonging to vehicle i, represents the j-th image of vehicle i, represents other images of vehicle i, represents the images of other vehicles except vehicle i, D represents the distance function, V represents the number of vehicles, and γ is the boundary threshold of the triplet loss.
[0050] Preferably, the ByteTracker is used to track multiple targets with a single camera to obtain trajectories, specifically:
[0051] The trajectories of the previous t-1 frames are obtained, and combined with the state attributes of each trajectory during the tracking process, the trajectories are divided into continuous tracking trajectories, broken trajectories, and initialization trajectories. The continuous tracking trajectory is the trajectory that has been tracked until the current frame. The broken trajectory is the trajectory that failed to be tracked in a previous frame but did not exceed the set tracking discard threshold and is still retained in the matching trajectory pool. The initialization trajectory is the trajectory that was tracked only in the previous frame.
[0052] The NSA Kalman filter is used to predict the trajectory of the previous t-1 frame, and the results detected at time t in the current frame are divided into high-confidence targets and low-confidence targets.
[0053] During the first matching, the continuously tracked, broken, and initialized trajectories are simultaneously matched with the high-confidence target based on IOU and REID cosine distance, and only when the distances of the two are both lower than the set thresholds α1 and α2, further matching is performed using the Hungarian algorithm.
[0054] The first unmatched continuous tracking, broken and initialized trajectories are matched with low confidence targets according to IOU. After the matching process is completed, for more than T max The untracked trajectories in the frame are deleted, and the appearance features of the remaining trajectories are updated using the exponential moving average. The continuously tracked trajectories that are not successfully matched to the detection frame target are switched to broken trajectories, and the broken trajectories that are successfully tracked again are switched to continuously tracked trajectories.
[0055] For the remaining high-confidence targets that failed to find matching tracks during the matching process, they are set as new initialization tracks.
[0056] Preferably, the cosine distance is used to calculate the similarity in the hierarchical clustering:
[0057]
[0058] Among them, f(T i ) and f(T j ) are the appearance feature vectors of the two trajectories respectively.
[0059] This paper proposes a cross-camera multi-target online tracking method that combines DINOv2 and ByteTracker. This method introduces the advanced visual base model DINOv2 in the feature extraction stage, designs and trains a new appearance feature extraction network, and uses the weighted sum of triples and cross-entropy loss as the total loss for training to extract more robust and rich target appearance features; secondly, the impact of additional data sets on training is deeply studied, and the generalization ability of the model is further enhanced by introducing diverse training samples. In addition, the ByteTracker algorithm is constructed based on ByteTrack to generate candidate tracks for each camera, and the tracking performance is further improved by optimizing the data association strategy and track management mechanism; finally, these candidate tracks are matched between different cameras through the cross-camera association module, and the target is associated with the global identity. Ultimately, the performance of cross-camera multi-target tracking is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 Schematic diagram of multi-camera multi-target tracking in the prior art;
[0061] Figure 2 It is the overall flow chart of the present invention;
[0062] Figure 3 This is a flow chart of Example 1;
[0063] Figure 4 This is the structure diagram of the Dinov2 model;
[0064] Figure 5 This is the flow chart of the ByteTracker algorithm;
[0065] Figure 6 Visualize results for multi-object tracking across cameras;
[0066] Figure 7 Tracking visualization results for real-world scenarios. DETAILED DESCRIPTION
[0067] In the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.
[0068] It will be understood that the “embodiment” mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the various embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It will be understood that in the various embodiments of the present application, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.
[0069] In the present invention, unless otherwise specified, the same or similar parts between the various embodiments can refer to each other. In the various embodiments of the present invention, and the various implementation methods / implementation methods / implementation methods in each embodiment, if there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments and the various implementation methods / implementation methods / implementation methods in each embodiment are consistent and can be referenced to each other. The technical features in different embodiments and the various implementation methods / implementation methods / implementation methods in each embodiment can be combined to form new embodiments, implementation methods, implementation methods, or implementation methods according to their inherent logical relationships. The implementation methods of the present application described below do not constitute a limitation on the scope of protection of the present application.
[0070] The overall process of the cross-camera multi-target online tracking algorithm of the present invention is as follows: Figure 2As shown, the video from each camera is processed frame by frame synchronously. First, the vehicle detection module outputs the vehicle coordinates and category in the current frame and extracts vehicle features using a feature extraction model. Then, based on the vehicle position and extracted features, the SCMT module performs tracking in each camera to generate candidate tracks. Finally, the ICA module clusters the current online tracks of each camera and associates the target with a global identity. That is, for each unidentified track, the same identity as an already identified track is assigned. Throughout the MTMCT process, the detection process, the object ReID feature extraction stage, and the single-camera object tracking stage all have a significant impact on the final cross-camera tracking accuracy. Therefore, by introducing the visual foundation model Dinov2, a new appearance feature extraction model is designed and trained to extract robust target features and obtain a highly expressive target image feature vector, thereby improving target tracking performance in complex cross-camera environments. Furthermore, for single-camera tracking, a ByteTracker algorithm is designed and constructed to improve the online single-camera object tracking process and enhance accuracy. After single-camera tracking, the online trajectory of each camera is clustered in a multi-camera environment through a hierarchical clustering method based on the pairwise cosine distance and Dunn index to achieve online cross-camera target tracking.
[0071] Example 1, as Figure 3 As shown, a method for online tracking of multiple targets across cameras combining Dinov2 and ByteTracker is provided, the method comprising:
[0072] S1, use Dinov2 to extract features from the image of each camera, and use ByteTracker to track multiple targets with a single camera to obtain trajectories.
[0073] Dinov2 is a pre-trained visual foundation model. Each frame captured by each camera is input into the Dinov2 model, which outputs a feature vector. The video stream or image sequence of each camera is processed separately to extract the corresponding features. The target is detected in each frame of the video, and the detection results of the same target in different frames are linked to form a complete trajectory of the target in the video. The trajectory shown includes but is not limited to the target's position, size, appearance features, etc. In one embodiment, Dinov2 is used as the backbone network, and the features output by the backbone network are input into the ByteTracker model to obtain the trajectory.
[0074] In one embodiment, the feature extraction of each camera image using Dinov2 is specifically as follows:
[0075] The patch embed module is used for processing. The patch embed module contains a two-dimensional convolution layer with a 14×14 convolution kernel and a stride of 14, as well as a normal layer, which uniformly converts the input image into a 14×14 patch representation.
[0076] The patches are fed into the ViT Blocks module for feature extraction, which outputs a feature matrix of dimension T×D. The feature matrix is normalized by the Layer Norm module to generate a 1×n-dimensional feature vector; where T represents the number of channels, D represents the feature vector length, and n is a positive integer.
[0077] Specifically, if Figure 4 As shown in the figure, the architecture of the Dinov2 model adopts a modular design. At the input stage, the image is first processed by the patch embedding module, which contains a two-dimensional convolutional layer with a 14×14 convolution kernel and a stride of 14, and a normal layer. The input image is uniformly converted into 14×14 patches. Subsequently, these patches are sent to the ViT Blocks module for feature extraction, where the number of ViT blocks can be flexibly configured according to the model scale. After processing by ViT Blocks, the output is a feature matrix with a dimension of T (number of channels) × D (length of feature vector). After normalization by the Layer Norm module, this matrix generates a 1×n-dimensional feature vector. Finally, the model can be adaptively adjusted according to the requirements of specific image tasks through a configurable head module, achieving flexible adaptation to different visual tasks.
[0078] Dinov2 uses knowledge distillation to derive three smaller models from the largest Dinov2g model. These smaller models inherit Dinov2g's superior feature extraction capabilities but are smaller in size. See Table 1 for details.
[0079] Table 1 Parameters of four models of Dinov2
[0080]
[0081] One of the main reasons for Dinov2's outstanding performance is that it was trained on a massive dataset, called LVD-142M. This dataset encompasses ImageNet-22k, ImageNet-1k, Google Landmark, various fine-grained datasets, and images scraped from the web, totaling 142 million images. Just as humans continuously improve their understanding through extensive learning experiences and accumulated knowledge, DINOv2 possesses excellent image understanding capabilities, making it capable of handling a variety of downstream tasks. Therefore, it is believed that this model also excels in ReID feature extraction. Specifically, the pre-trained Dinov2 model is used as a baseline for image feature extraction, and its output image features are aggregated to obtain a reliable global feature representation.
[0082] The present invention pruned the Layer Norm module and the Head module in the Dinov2s model, and only retained ViTBlocks as the backbone network. The feature matrix output by these pruned ViT Blocks modules has the dimension of T (number of channels) × D (length of feature vector). In order to adapt to the subsequent processing flow, the channel dimension is reshaped into a feature map of size H×W, where the number of feature maps corresponds to the length of the feature vector. Considering that the dimension of the Dinov2s output feature is 384, 1×1 convolution is used to expand the feature dimension from 384 to 2048 to enhance the expressive power of the network. Then, an adaptive average pooling layer is used to extract features on a global scale and effectively capture global information, thereby improving the network's understanding of high-level semantics and enhancing its robustness in the face of changes in the input space. In addition, the Dropout and BN (batch normalization) layers are combined to further improve the generalization ability of the model. Finally, the output is passed through the Head layer. During the training phase, a joint loss function is adopted, combined with the cross entropy loss (loss xe ) and triplet loss (loss tr ), where the total loss Loss is defined as the weighted sum of the two.
[0083] In one embodiment, the training process of Dinov2 is:
[0084] Get Image Hard positive samples of:
[0085]
[0086] Get Image Hard negative samples:
[0087]
[0088] Calculate the hard positive loss:
[0089]
[0090] Calculate the hard negative loss:
[0091]
[0092] Calculate triple loss loss tr =max(0,loss hp -loss hn +γ).
[0093] Calculate the total loss Loss = loss tr +loss ex .
[0094] Among them, loss ex is the cross entropy loss, B i represents the number of images belonging to vehicle i, represents the j-th image of vehicle i, Other images representing vehicle i, represents the images of other vehicles except vehicle i, D represents the distance function, V represents the number of vehicles, and γ is the boundary threshold of the triplet loss. feat () represents feature extraction, which takes image I as input and outputs the feature vector of the image.
[0095] Specifically, hard positive samples refer to samples that belong to the same category as the current anchor image but are relatively far away from the anchor in the feature space. For a given anchor image, for example, the jth image of vehicle i In all other images belonging to the same vehicle i (where k is not equal to j), find those samples that are the most dissimilar to the feature representation of the anchor image, that is, the farthest away. The samples of the same category with a farther distance are hard positive samples. Hard negative samples refer to samples that belong to different categories from the current anchor image, that is, samples belonging to other vehicles p except vehicle i, but are relatively close to the anchor in the feature space. For the same anchor image In all other vehicles (where p is not equal to i), find the sample with the closest distance to the feature representation of the anchor image. The samples of different categories with closer distance are hard negative samples. Calculate the hard positive sample loss for all vehicles hp and hard negative loss hn Using loss tr =max(0,loss hp -loss hn +γ) calculate the triple loss and further calculate the total loss Loss = loss tr+loss ex ; The Dinov2 model is trained using total loss.
[0096] In one embodiment, the ByteTracker is used to track multiple targets with a single camera to obtain trajectories, specifically:
[0097] The trajectories of the previous t-1 frames are obtained, and combined with the state attributes of each trajectory during the tracking process, the trajectories are divided into continuous tracking trajectories, broken trajectories, and initialization trajectories. The continuous tracking trajectory is the trajectory that has been tracked until the current frame. The broken trajectory is the trajectory that failed to be tracked in a previous frame but did not exceed the set tracking discard threshold and is still retained in the matching trajectory pool. The initialization trajectory is the trajectory that was tracked only in the previous frame.
[0098] The NSA Kalman filter is used to predict the trajectory of the previous t-1 frame, and the results detected at time t in the current frame are divided into high-confidence targets and low-confidence targets.
[0099] During the first matching, the continuously tracked, broken, and initialized trajectories are simultaneously matched with the high-confidence targets based on IOU and ReID cosine distance, and only when the distances between the two are both lower than the set thresholds α1 and α2, further matching is performed using the Hungarian algorithm.
[0100] The first unmatched continuous tracking, broken and initialized trajectories are matched with low confidence targets according to IOU. After the matching process is completed, for more than T max The untracked trajectories in the frame are deleted, and the appearance features of the remaining trajectories are updated using the exponential moving average. The continuously tracked trajectories that are not successfully matched to the detection frame target are switched to broken trajectories, and the broken trajectories that are successfully tracked again are switched to continuously tracked trajectories.
[0101] For the remaining high-confidence targets that failed to find matching tracks during the matching process, they are set as new initialization tracks.
[0102] Specifically, in the single-camera multi-target tracking task, the present invention is based on the ByteTrack algorithm, called ByteTracker, as shown in Figure 5The specific process is as follows: First, based on the trajectories of the previous t-1 frame and combined with the state attributes of each trajectory during the tracking process, the trajectories are divided into continuous tracking trajectories, broken trajectories, and initialization trajectories. Continuous tracking trajectories are trajectories that have been tracked until the current frame; broken trajectories are trajectories that failed to be tracked in a previous frame but did not exceed the set tracking discard threshold and are still retained in the matching trajectory pool; initialization trajectories are trajectories that were only tracked in the previous frame (i.e., t-1 frame). Next, the NSA Kalman filter is used to predict these trajectories to determine their target box positions at time t.
[0103] like Figure 5 As shown, the detection results at time t in the current frame are first divided into high-confidence targets and low-confidence targets. During the first matching process, continuously tracked, broken, and initialized tracks are simultaneously matched with high-confidence targets based on the IOU and Recognized Integer Cosine distance. Only when the distances are simultaneously below the set thresholds α1 and α2 are they matched using the Hungarian algorithm. The first matching process produces three outputs: matched tracks with high-confidence targets, tracks that were not successfully matched, and tracks that were not successfully matched with high-confidence targets. Next, a second matching process is performed to match the continuously tracked, broken, and initialized tracks that were not matched in the first matching process with low-confidence targets. Unlike the first matching process, this matching process primarily relies on the IOU distance as the judgment criteria, as low-confidence targets typically lack convincing appearance features. After the matching process, the appearance features of the matched tracks are updated using an exponential moving average (EMA) strategy. Furthermore, a track management strategy is used to delete tracks that have not been tracked for more than Tmax frames. Continuously tracked tracks that fail to match the detection box target are switched to broken tracks, and broken tracks that are successfully tracked are switched to continuously tracked tracks. In addition, for the remaining high-confidence targets that failed to find matching trajectories during the matching process, they are set as new initialization trajectories. It is worth noting that, unlike algorithms such as ByteTrack, the continuously tracked, broken, and initialized trajectories are matched because the initialization trajectories are derived from the high-confidence targets that were not successfully matched. It is believed that these initialization trajectories are not caused by short-term tracking error interference caused by detection noise. In previous algorithms, only the continuously tracked and broken trajectories are used for the first and second associations, and the initialization trajectories are matched for the third time with the high-confidence targets that failed to match the first time, such as Figure 5 As shown in the light green box in the middle, this not only increases the number of matches and the amount of cost matrix calculations, but also the initialized trajectory does not participate in the matching with low-confidence targets, resulting in the risk of losing the trajectory target.
[0104] S2: Cluster the trajectories in all cameras using hierarchical clustering. If the maximum number of trajectories in a cluster belonging to the same camera is greater than 1, the cluster is further divided into clusters with the maximum number of trajectories. Match the cluster with the identified trajectories at the previous moment based on the appearance features, and calculate the average cohesion of the successfully matched clusters. If there are mismatched clusters and the cluster cohesion is less than the average, assign a new ID to the mismatched cluster. For the elements in the remaining mismatched clusters, assign the elements to the trajectory that is closest to the element and has been successfully identified.
[0105] Hierarchical clustering is performed on the trajectory data captured by all cameras, initially clustering trajectories that are similar in terms of time, space, and features into distinct clusters. To further refine these clusters, each cluster is examined. If more than one trajectory originates from the same camera within a cluster, this may indicate objects from the same physical location but represented by different trajectory segments, or the clustering granularity is too coarse. In this case, the cluster is further split into several subclusters, with the number of subclusters equal to the maximum number of trajectories belonging to the same camera within the cluster.
[0106] After obtaining the trajectory clusters at the current moment, these clusters are matched with the trajectories that have been successfully identified and tracked at the previous moment. The matching is mainly based on the appearance features of these clusters and existing trajectories, and the correspondence is established by comparing their visual similarities. For successfully matched clusters, their cohesion is further calculated. The cohesion is preferably the average distance of the appearance of the trajectories within the cluster, and the average cohesion of all successfully matched clusters is further calculated. The smaller the cohesion, the more similar the elements within the cluster. For mismatched clusters, cohesion is calculated. If the cohesion of a mismatched cluster is lower than the average cohesion of the successfully matched clusters calculated previously, the trajectories inside the mismatched cluster are more concentrated, indicating a stable or clear new target. A new unique ID is assigned to this mismatched cluster, and it is regarded as a newly appeared target and tracking begins.
[0107] For each individual track element in the remaining mismatched clusters (i.e., those with cohesion no less than the average), the element is compared with all successfully identified and tracked tracks at the current moment, and the physical distance between the element and each identified track is calculated. In one embodiment, the physical distance refers to the spatial distance. The element is then assigned to the successfully identified track closest to it, thereby reassociating these temporarily mismatched track segments that may still belong to a known target. Physical proximity is used to compensate for the deficiencies or uncertainties in appearance feature matching, thereby achieving more robust and coherent multi-target tracking.
[0108] In an alternative embodiment, S2 is specifically as follows: after single-camera tracking, the online trajectory of each current camera is clustered in a multi-camera environment by a hierarchical clustering method to achieve online cross-camera target tracking. First, the current online trajectory generated by each camera is clustered by hierarchical clustering based on the pairwise cosine distance and Dunn index. Based on the clustering results, the unrecognized trajectories in the same cluster are assigned the same identity ID as the recognized trajectories. If some trajectories still need to be identified, the algorithm will match them with the lost trajectories. Among them, the cosine distance is used to measure the similarity between two features, and the formula is:
[0109]
[0110] Among them, f(T i ) and f(T j ) are the appearance feature vectors of the two trajectories respectively.
[0111] Unlike partitioning clustering, k-means and other methods require the number of clusters to be determined in advance, or DBSCAN requires the definition of the minimum number of points to form a cluster. Hierarchical clustering algorithms gradually build cluster structures without the need to pre-specify the number of clusters; therefore, hierarchical clustering has obvious advantages in the absence of prior knowledge about the number of clusters. However, the output of hierarchical clustering is a cluster tree, usually presented in the form of a dendrogram. When correlating cross-camera trajectories, the specific number of vehicle trajectories, that is, the number of clusters, cannot be directly given, but the hierarchical relationship between data points is displayed. In order to determine the optimal number of clusters, clustering verification techniques such as the Dunn Index can be used for evaluation. The formula for the Dunn Index is:
[0112]
[0113] Among them, δ(C i ,C j ) represents cluster C i and C j The minimum distance between l Represents cluster C l The diameter of a cluster (i.e., the maximum distance within the cluster), k is the number of clusters.
[0114] More specifically, by maximizing the Dunn index, an appropriate clustering threshold is determined. This threshold measures both cluster compactness and the degree of separation between clusters, resulting in the most effective clustering results. This ensures that tracks with similar appearance are assigned to the same cluster, while distinct target object instances are adequately distinguished. Then, for each cluster, if any unidentified tracks exist, the track with the smallest feature distance to a recognized track in the same cluster is found, and each unidentified track is assigned the ID of the corresponding recognized track. Finally, any remaining unidentified tracks that failed to cluster with other tracks are matched against the missing tracks. However, if the cosine distance between the matching pairs exceeds a set threshold, these tracks are assigned new IDs. Furthermore, during the clustering process, tracks within the same camera are not clustered, as the same object cannot be detected simultaneously in a single-camera scenario. Similarly, clustering is restricted to tracks outside the possible overlapping region of each camera to avoid false associations.
[0115] Embodiment 2 provides a cross-camera multi-target online tracking system combining Dinov2 and ByteTracker, characterized in that the system includes:
[0116] The single-camera multi-target tracking module is used to extract features from the image of each camera using Dinov2 and use ByteTracker to track multiple targets with a single camera to obtain trajectories.
[0117] The cross-camera association module is used to cluster the trajectories in all cameras using a hierarchical clustering approach. If the maximum number of trajectories in a cluster belonging to the same camera is greater than 1, the cluster is further divided into clusters with the maximum number of trajectories. The cluster is matched with the previously identified trajectories based on appearance features, and the average cohesion of the successfully matched clusters is calculated. If there are mismatched clusters and the cluster cohesion is less than the average, a new ID is assigned to the mismatched cluster. For the elements in the remaining mismatched clusters, the elements are assigned to the track that is closest to the element and has been successfully identified.
[0118] Preferably, the feature extraction of the image from each camera using Dinov2 is specifically as follows:
[0119] The patch embed module is used for processing. The patch embed module contains a two-dimensional convolution layer with a 14×14 convolution kernel and a stride of 14, as well as a normal layer, which uniformly converts the input image into a 14×14 patch representation.
[0120] The patches are fed into the ViT Blocks module for feature extraction, which outputs a feature matrix of dimension T×D. The feature matrix is normalized by the Layer Norm module to generate a 1×n-dimensional feature vector; where T represents the number of channels, D represents the feature vector length, and n is a positive integer.
[0121] Preferably, the training process of Dinov2 is:
[0122] Get Image Hard positive samples of:
[0123]
[0124] Get Image Hard negative samples:
[0125]
[0126] Calculate the hard positive loss:
[0127]
[0128] Calculate the hard negative loss:
[0129]
[0130] Calculate triple loss loss tr =max(0,loss hp -loss hn +γ).
[0131] Calculate the total loss Loss = loss tr +loss ex .
[0132] Among them, loss ex is the cross entropy loss, B i represents the number of images belonging to vehicle i, represents the j-th image of vehicle i, represents other images of vehicle i, represents the images of other vehicles except vehicle i, D represents the distance function, V represents the number of vehicles, and γ is the boundary threshold of the triplet loss.
[0133] Preferably, the ByteTracker is used to track multiple targets with a single camera to obtain trajectories, specifically:
[0134] The trajectories of the previous t-1 frames are obtained, and combined with the state attributes of each trajectory during the tracking process, the trajectories are divided into continuous tracking trajectories, broken trajectories, and initialization trajectories. The continuous tracking trajectory is the trajectory that has been tracked until the current frame. The broken trajectory is the trajectory that failed to be tracked in a previous frame but did not exceed the set tracking discard threshold and is still retained in the matching trajectory pool. The initialization trajectory is the trajectory that was tracked only in the previous frame.
[0135] The NSA Kalman filter is used to predict the trajectory of the previous t-1 frame, and the results detected at time t in the current frame are divided into high-confidence targets and low-confidence targets.
[0136] During the first matching, the continuously tracked, broken, and initialized trajectories are simultaneously matched with the high-confidence target based on IOU and REID cosine distance, and only when the distances of the two are both lower than the set thresholds α1 and α2, further matching is performed using the Hungarian algorithm.
[0137] The first unmatched continuous tracking, broken and initialized trajectories are matched with low confidence targets according to IOU. After the matching process is completed, for more than T max The untracked trajectories in the frame are deleted, and the appearance features of the remaining trajectories are updated using the exponential moving average. The continuously tracked trajectories that are not successfully matched to the detection frame target are switched to broken trajectories, and the broken trajectories that are successfully tracked again are switched to continuously tracked trajectories.
[0138] For the remaining high-confidence targets that failed to find matching tracks during the matching process, they are set as new initialization tracks.
[0139] Preferably, the cosine distance is used to calculate the similarity in the hierarchical clustering:
[0140]
[0141] Among them, f(T i ) and f(T j ) are the appearance feature vectors of the two trajectories respectively.
[0142] The DINOv2 network model was built using the PyTorch framework and trained and tested on NVIDIA GeForce RTX 4090 hardware. The training datasets used included CityFlowV2 and VeRi-776. To expand the dataset and improve the network model's robustness to vehicles in diverse and complex environments, the VehicleX dataset, a synthetic image dataset generated by a 3D engine, was used to augment the training data.
[0143] The CityFlowV2 dataset is collected from 46 cameras at 16 intersections in a medium-sized US city. The dataset covers a variety of location types, including intersections, road stretches, and highways. For city-scale multi-camera vehicle tracking, six scenarios are used: three for training, two for validation, and the remaining for testing.
[0144] The VeRi-776 dataset serves as additional training data to improve feature extraction models. VeRi-776 is one of the largest and most common vehicle re-identification datasets in multi-camera scenarios. It consists of approximately 50,000 bounding boxes of 776 vehicles captured by 20 cameras.
[0145] The VehicleX dataset is a synthetic dataset generated by the publicly available 3D engine VehicleX. This dataset only provides a training set and contains a total of 192,150 images of 1,362 vehicles.
[0146] Training was performed using 309,451 images of 1,362 vehicles across three datasets, and validation and final evaluation were performed using the S02 scene in CityFlowV2. Training process: The network was initialized using Dinov2s pre-trained weights. Since the network requires input images to be multiples of 14×14, all input images were resized to 392×392. Data augmentation, including random color jittering, erasing, horizontal flipping, and affine transformations, was employed during training. Furthermore, the Adam optimizer was used with a weight decay coefficient of 1e-5, and training was performed for 150 epochs using cross-entropy loss and triplet loss. The learning rate started at 0.00005 and was reduced to 1 / 10 of its original value after 40 and 90 epochs of training, respectively.
[0147] Evaluation Metrics: The true label values of MTMC tracking provided by the CityFlowV2 benchmark consist of bounding boxes of vehicles with consistent IDs from multiple cameras. According to the CityFlowV2 benchmark evaluation method, the recognition precision (IDP), recognition recall (IDR), and F1 score (IDF1) are used for evaluation:
[0148]
[0149] IDP (IDR) is the fraction of computed (ground truth) trajectories that are correctly identified. IDF1 is an important metric for evaluating identity consistency in multi-target tracking. The F1 score is the harmonic average of precision and recall. In the context of multi-target tracking, precision refers to the proportion of correctly identified identities to all identified identities, while recall refers to the proportion of correctly identified identities to all actual identities. Among them, IDTP (True Positives) is the number of correctly matched identities, IDFP (False Positives) is the number of incorrectly matched identities, and IDFN (False Negatives) is the number of missed identities.
[0150] To achieve better cross-camera target tracking performance, we tested parameter freezing on different layers of the ViT Block in the Dinov2s Backbone module to explore the optimal number of weight update layers. By systematically analyzing the effects of each layer, we found that the model that only trained the last two Dinov2s Blocks performed better. The results are shown in Table 2:
[0151] Table 2 Experimental results of updating parameters of different layers
[0152]
[0153]
[0154] As shown in Table 2, when updating the parameters of the ViT Block, those with fewer ViT Block parameter updates have more stable accuracy. This indicates that excessively deep parameter updates can destroy the advantages of the Dinov2s model's original training parameters, resulting in a decrease in overall model accuracy. On the one hand, freezing the parameters of early layers can maintain the feature extraction capabilities acquired during training; on the other hand, moderately updating the parameters of later layers can help the model adapt to specific task requirements.
[0155] In order to evaluate the effectiveness of the training data, the models trained using different training sets are compared in Table 3. Based on the previous experiments, only the last two layers of blocks were updated, and the model was trained with images, with the image size adjusted to 392×392 as a benchmark. When trained using only the V2+VeRi dataset, the model's IDF1, IDP, and IDR were 77.63, 78.14, and 77.14, respectively. After adding the VehicleX synthetic dataset, the IDF1, IDP, and IDR increased by 1.44%, 2.25%, and 0.65%, respectively, indicating that the VehicleX synthetic dataset has improved performance and verifying the value of synthetic data in enhancing the generalization ability of the model. In addition, after combining the VehicleX dataset and adopting the improved ByteTracker algorithm, the model's performance in IDF1, IDP, and IDR was further improved, reaching 80.18, 81.67, and 78.74, respectively, verifying the effectiveness of the present invention.
[0156] Table 3 Comparative experiments using synthetic datasets and ByteTracker models
[0157]
[0158] During the cross-camera target track association phase, we evaluated the impact of various methods on the inter-track distance calculation in the MTMCT task through ablation experiments using different feature sampling methods for each track. Best represents sampling from the detection box with the best confidence score, Last represents sampling from the last observed detection box (the most recent detection), and Avg, Weighted_avg, and Ema represent different ways of smoothing the features. Avg refers to a simple average of the features for each track, Weighted_avg is a weighted average based on the average confidence score, and Ema uses an exponentially weighted moving average.
[0159] Table 4. Ablation study of feature sampling method for each track for feature distance comparison
[0160]
[0161] As shown in Table 4, the experimental results show that the features of the detection box with the best confidence in the trajectory history (Best) show the best performance in all evaluation indicators. In contrast, other sampling methods, such as using the last detection box (Last), simple average (Avg), weighted average based on confidence score (Weighted_avg), and exponential moving average (Ema), can provide a certain smoothing effect, but the performance is relatively poor. Therefore, the present invention chooses to use the features extracted from the detection box with the best confidence as the default features to ensure that the best tracking performance is achieved while maintaining a high tracking speed.
[0162] Table 5 Comparison results of the method of the present invention and other methods
[0163]
[0164] Table 5 lists the comparison results of the present invention and other advanced methods. The study shows that the MTMCT algorithm of the present invention surpasses previous online algorithms in multiple key performance indicators, showing significant advantages. Specifically, the method proposed in the present invention achieves the best level of test accuracy. Compared with the optimal algorithm, the IDF1, IDP, and IDR indicators are improved by 1.74%, 2.28%, and 0.74%, respectively, and finally reach an IDF1 of 80.18, an IDP of 81.67, and an IDR of 78.74. This shows that the present invention has achieved a significant improvement in accuracy. In addition, in terms of tracking speed, the present invention not only significantly reduces the running time, but is also only 0.02s slower than the best existing method, showing a great competitive advantage. This means that under the premise of ensuring high accuracy, the present invention also has good performance in execution efficiency and can meet application scenarios with high real-time requirements.
[0165] At the same time, in order to more comprehensively evaluate the algorithm performance and intuitively demonstrate the effect of vehicle tracking, the tracking results were visualized and analyzed, such as Figure 6 As shown, the present invention successfully achieves continuous tracking of targets ID31, ID34, and ID37 in different camera scenarios. Specifically, these targets maintain stable identity consistency under complex conditions such as cross-camera switching, perspective changes, and partial occlusion. These results fully demonstrate that the present invention can effectively identify and track the same target across multiple cameras, validating its effectiveness and applicability.
[0166] In order to further verify the practicality of the invention, not only experimental verification was conducted based on the data set, but also a field shooting experiment was designed. The experiment deployed three small drones to simulate surveillance cameras in road traffic scenes, and built a multi-camera monitoring system to obtain more realistic monitoring video data. The experimental location is near the intersection of Lianhua Street and Chinan Road in Zhengzhou City. The traffic volume in this area is large and has typical urban road characteristics. It can truly reflect the performance of the algorithm in actual scenes. The following figure shows the equipment layout. The vehicle will pass through the three cameras in turn. The experimental results are as follows Figure 7 As shown, targets ID14, ID17, ID19 and ID20 can all be accurately tracked, indicating that the present invention can still maintain high tracking accuracy and stability in complex and changeable urban traffic environments, verifying the practicality and reliability of the present invention in actual traffic scenarios.
[0167] The steps of the methods or algorithms described in the embodiments of the present invention may be directly embedded in hardware, software units executed by a processor, or a combination of the two. The software units may be stored in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. For example, the storage medium may be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and storage medium may be provided in an ASIC.
[0168] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.
[0169] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, the present application is intended to include such modifications and variations if they fall within the scope of the claims of the present application and their equivalents.
Claims
1. A cross-camera multi-target online tracking method combining Dinov2 and ByteTracker, characterized in that: The method comprises: Dinov2 is used to extract features from the image of each camera, and ByteTracker is used to track multiple targets with a single camera to obtain trajectories; Hierarchical clustering is used to cluster the trajectories in all cameras. If the maximum number of trajectories in a cluster belonging to the same camera is greater than 1, the cluster is further divided into clusters with the maximum number of trajectories. The cluster is matched with the previously identified trajectories based on appearance features, and the average cohesion of the successfully matched clusters is calculated. If there is a mismatched cluster and the cluster cohesion is less than the average, a new ID is assigned to the mismatched cluster. For the elements in the remaining mismatched clusters, the elements are assigned to the successfully identified trajectory that is closest to the element in physical distance.
2. The method according to claim 1, wherein The feature extraction of each camera image using Dinov2 is as follows: The patch embed module is used for processing. The patch embed module contains a two-dimensional convolution layer with a 14×14 convolution kernel and a stride of 14, and a normal layer. The input image is uniformly converted into 14×14 patches. The patches are fed into the ViT Blocks module for feature extraction, which outputs a feature matrix of dimension T×D. The feature matrix is normalized by the Layer Norm module to generate a 1×n-dimensional feature vector; where T represents the number of channels, D represents the feature vector length, and n is a positive integer.
3. The method according to claim 1, wherein The training process of Dinov2 is: Get Image Hard positive samples of: Get Image Hard negative samples: Calculate the hard positive loss: Calculate the hard negative loss: Calculate triple loss loss tr =max(0,loss hp -loss hn +γ); Calculate the total loss Loss = loss tr +loss ex ; Among them, loss ex is the cross entropy loss, B i represents the number of images belonging to vehicle i, represents the j-th image of vehicle i, Other images representing vehicle i, represents the images of other vehicles except vehicle i, D represents the distance function, V represents the number of vehicles, and γ is the boundary threshold of the triplet loss.
4. The method according to claim 1, wherein The ByteTracker is used to track multiple targets with a single camera to obtain trajectories, specifically: Obtain the trajectories of the previous t-1 frames of tracking, and combine the state attributes of each trajectory during the tracking process to classify the trajectories into continuous tracking trajectories, broken trajectories, and initialization trajectories. The continuous tracking trajectory is the trajectory that has been tracked up to the current frame. The broken trajectory is the trajectory that failed to be tracked in a previous frame but did not exceed the set tracking discard threshold and is still retained in the matching trajectory pool. The initialization trajectory is the trajectory that was tracked only in the previous frame. Use the NSA Kalman filter to predict the trajectory of the previous t-1 frame, and divide the results detected at the current frame t into high-confidence targets and low-confidence targets; During the first matching, the continuously tracked, broken, and initialized trajectories are simultaneously matched with the high-confidence target based on IOU and ReID cosine distance. Only when the distances between the two are both lower than the set thresholds α1 and α2, further matching is performed using the Hungarian algorithm. The first unmatched continuous tracking, broken and initialized trajectories are matched with low confidence targets according to IOU. After the matching process is completed, for more than T max The untracked trajectories in the frame are deleted, and the appearance features of the remaining trajectories are updated using the exponential moving average. The continuously tracked trajectories that are not successfully matched to the detection frame target are switched to broken trajectories, and the broken trajectories that are successfully tracked again are switched to continuously tracked trajectories. For the remaining high-confidence targets that failed to find matching tracks during the matching process, they are set as new initialization tracks.
5. The method according to claim 1, wherein The cosine distance is used to calculate the similarity in the hierarchical clustering: Among them, f(T i ) and f(T j ) are the appearance feature vectors of the two trajectories respectively.
6. A cross-camera multi-target online tracking system combining Dinov2 and ByteTracker, characterized by: The system comprises: The single-camera multi-target tracking module is used to extract features from each camera's image using Dinov2 and track multiple targets with a single camera using ByteTracker to obtain trajectories. The cross-camera association module is used to cluster the trajectories in all cameras using a hierarchical clustering approach. If the maximum number of trajectories in a cluster belonging to the same camera is greater than 1, the cluster is further divided into clusters with the maximum number of trajectories. The cluster is matched with the previously identified trajectories based on appearance features, and the average cohesion of the successfully matched clusters is calculated. If there are mismatched clusters and the cluster cohesion is less than the average, a new ID is assigned to the mismatched cluster. For the elements in the remaining mismatched clusters, the elements are assigned to the track that is closest to the element and has been successfully identified.
7. The system according to claim 6, wherein: The feature extraction of each camera image using Dinov2 is as follows: The patch embed module is used for processing. The patch embed module contains a two-dimensional convolution layer with a 14×14 convolution kernel and a stride of 14, and a normal layer. The input image is uniformly converted into 14×14 patches. The patches are fed into the ViT Blocks module for feature extraction, which outputs a feature matrix of dimension T×D. The feature matrix is normalized by the Layer Norm module to generate a 1×n-dimensional feature vector; where T represents the number of channels, D represents the feature vector length, and n is a positive integer.
8. The system according to claim 6, wherein: The training process of Dinov2 is: Get Image Hard positive samples of: Get Image Hard negative samples: Calculate the hard positive loss: Calculate the hard negative loss: Calculate triple loss loss tr =max(0,loss hp -loss hn +γ); Calculate the total loss Loss = loss tr +loss ex ; Among them, loss ex is the cross entropy loss, B i represents the number of images belonging to vehicle i, represents the j-th image of vehicle i, Other images representing vehicle i, represents the images of other vehicles except vehicle i, D represents the distance function, V represents the number of vehicles, and γ is the boundary threshold of the triplet loss.
9. The system according to claim 6, wherein: The ByteTracker is used to track multiple targets with a single camera to obtain trajectories, specifically: Obtain the trajectories of the previous t-1 frames of tracking, and combine the state attributes of each trajectory during the tracking process to classify the trajectories into continuous tracking trajectories, broken trajectories, and initialization trajectories. The continuous tracking trajectory is the trajectory that has been tracked up to the current frame. The broken trajectory is the trajectory that failed to be tracked in a previous frame but did not exceed the set tracking discard threshold and is still retained in the matching trajectory pool. The initialization trajectory is the trajectory that was tracked only in the previous frame. Use the NSA Kalman filter to predict the trajectory of the previous t-1 frame, and divide the results detected at the current frame t into high-confidence targets and low-confidence targets; During the first matching, the continuously tracked, broken, and initialized trajectories are simultaneously matched with the high-confidence target based on IOU and ReID cosine distance. Only when the distances between the two are both lower than the set thresholds α1 and α2, further matching is performed using the Hungarian algorithm. The first unmatched continuous tracking, broken and initialized trajectories are matched with low confidence targets according to IOU. After the matching process is completed, for more than T max The untracked trajectories in the frame are deleted, and the appearance features of the remaining trajectories are updated using the exponential moving average. The continuously tracked trajectories that are not successfully matched to the detection frame target are switched to broken trajectories, and the broken trajectories that are successfully tracked again are switched to continuously tracked trajectories. For the remaining high-confidence targets that failed to find matching tracks during the matching process, they are set as new initialization tracks.
10. The system according to claim 6, wherein: The cosine distance is used to calculate the similarity in the hierarchical clustering: Among them, f(T i ) and f(T j ) are the appearance feature vectors of the two trajectories respectively.