Online multi-camera multi-vehicle target tracking method based on deep learning
By using deep learning technology in a multi-camera environment combined with vehicle detection and feature extraction, stable tracking of vehicles and cross-camera matching is achieved, solving the continuity and identification problems of vehicle tracking, and improving the real-time and accuracy of tracking.
Patent Information
- Application Number
- CN202510040439.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In a complex multi-camera environment, there are difficulties in continuous tracking, cross-camera matching and global ID allocation of vehicles, resulting in the problems of target loss, tracking discontinuity and identification difficulties.
The online multi-camera multi-vehicle target tracking method based on deep learning is adopted, combining lightweight vehicle detectors, high-quality appearance feature extraction, single-camera tracking and cross-camera association technology to achieve cross-camera target matching and global ID allocation through hierarchical clustering and Dunn index optimization.
It significantly improves the real-time, accuracy and robustness of vehicle tracking, ensuring stable and reliable vehicle tracking capabilities in a multi-camera environment.
Smart Images

Figure CN119991746A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and intelligent transportation systems, and is particularly applicable to vehicle tracking problems in application scenarios such as urban monitoring, traffic management, and automatic driving assistance systems. Specifically, the present invention relates to a real-time online multi-target, multi-camera vehicle tracking method, which aims to solve the problems of continuous tracking, cross-camera matching, and global ID allocation of vehicles in complex monitoring environments, and provide an efficient and accurate solution. Background Art
[0002] With the acceleration of urbanization and the continuous growth of the number of vehicles, how to effectively manage and monitor road traffic has become an important issue in modern urban management. Most traditional traffic monitoring systems are based on a single camera for vehicle tracking. This method performs well when dealing with simple scenes within a single field of view. However, traditional methods face many challenges when facing complex scenes in a multi-camera environment. For example, when a vehicle moves from the field of view of one camera to another, the target may be lost; the time synchronization problem between different cameras can also lead to discontinuous tracking; in addition, due to factors such as differences in viewing angles and changes in lighting, the appearance characteristics of the vehicle may change, making it difficult to identify the same vehicle under different cameras. These problems not only affect the accuracy and reliability of vehicle tracking, but also limit its scope of application in intelligent traffic management.
[0003] In order to meet the above challenges, researchers have begun to explore vehicle tracking solutions with multi-camera collaboration in recent years. Although some existing studies have tried to combine deep learning and computer vision technology to improve the effect of cross-camera tracking, there are still many difficulties in actual deployment. First, the amount of data in a multi-camera environment is huge, and the real-time processing requirements are high, which puts strict requirements on computing resources; second, the installation positions and angles of different cameras are different, resulting in uneven image quality, which increases the difficulty of feature extraction and matching; finally, the robustness of existing methods for handling occlusion, interference from similar objects, etc. still needs to be improved. Therefore, it is urgent to develop a stable and reliable vehicle tracking method in a multi-camera environment to meet the growing needs of urban traffic management and autonomous driving assistance systems. This method should not only be able to overcome the limitations of existing technologies, but also have efficient real-time processing capabilities and good adaptability to various complex scenarios. Summary of the invention
[0004] Purpose of the invention: In order to solve the above problems, the present invention proposes an online multi-camera multi-vehicle target tracking method based on deep learning, aiming to solve the challenge of vehicle target tracking in complex monitoring environments. The method is optimized for the vehicle target tracking problem in a multi-camera environment. By integrating key technologies such as lightweight vehicle detectors, high-quality appearance feature extraction, single camera tracking (MTSCT) and cross-camera association, the present invention provides a complete solution that significantly improves the real-time, accuracy and robustness of vehicle tracking.
[0005] The technical solution of the present invention is an online multi-camera, multi-vehicle target tracking method based on deep learning, comprising the following steps:
[0006] Step 1: Video image acquisition: collect moving video images of vehicle targets over a period of time from multiple cameras at different positions and angles, use OpenCV library functions to read the video stream frame by frame, and save each frame as a picture to obtain continuous video frame images;
[0007] Step 2: Vehicle detection: Use the YOLOv11 target detection algorithm to process each frame of the image extracted in step 1 to detect and locate all vehicle instances in the image;
[0008] Step 3: Vehicle appearance feature extraction: using the ResNet101_IBN convolutional neural network architecture, extract the appearance features of each vehicle instance determined in step 2 to generate a 2048-dimensional feature vector;
[0009] Step 4: The single camera multi-target (SCMT) tracking algorithm uses the bounding box data of step 2 and the feature vector of step 3 to generate the target motion trajectory under each individual camera, thereby realizing target tracking in a single camera environment;
[0010] Step 5: cluster the target trajectories generated in step 4 by using a hierarchical clustering algorithm combined with feature cosine distance calculation and Dunn index optimization to complete cross-camera target matching;
[0011] Step 6: Iterate the above steps 2 to 5 to achieve multi-camera and multi-target vehicle tracking under continuous video input conditions, and continuously update the feature information and global ID of each target.
[0012] In step 1, the video data is read frame by frame using the OpenCV library function, each frame is saved as a picture, and continuous video frame images are obtained.
[0013] In step 2, an end-to-end target detection method based on deep learning is used, specifically, the YOLOv11s model is used to detect vehicle targets in the surveillance video stream. Each detected vehicle instance is accurately calibrated by a rectangular bounding box, and the output data format is [left, top, right, bottom, confidence], where left and top represent the coordinates of the upper left corner of the bounding box, right and bottom represent the coordinates of the lower right corner, and confidence represents the confidence of the detection result.
[0014] In step 3, a ResNet101_IBN convolutional neural network architecture is used to extract detailed appearance features from each vehicle instance detected in step 2, and fused to generate 2048-dimensional feature vector information for each vehicle target. This feature information captures the visual characteristics of the vehicle, including but not limited to color, shape, and texture.
[0015] In step 4, a single-camera multi-target tracking system based on Kalman filter and feature matching is constructed. The following technologies are used to track multiple targets in the video frame:
[0016] Kalman prediction: predicts the position of the target in the next frame, reduces detection delay and improves tracking continuity; Feature fusion: combines the 2048-dimensional vehicle appearance feature information extracted from the ResNet101_IBN convolutional neural network to enhance the robustness and accuracy of target recognition; Target state management: handles complex situations including target loss, re-detection and deduplication, ensuring that tracking can be resumed after the target temporarily leaves the field of view or is blocked.
[0017] Furthermore, in step 4, the single-camera multi-target tracking algorithm first uses the NSA Kalman filter to predict the bounding box position of each target in the current frame. The state vector is defined as the center position, width, and height of the bounding box, and the trajectories are divided into three categories: tracked trajectories, which have been tracked to the last frame; lost trajectories, which have failed somewhere in the past but are still being tracked; and inactive trajectories, which are new trajectories that have just started tracking from the previous frame.
[0018] Furthermore, the algorithm in step 4 includes three trajectory association processes: In the first association process, the detection results are divided into high-confidence detection frame groups and low-confidence detection frame groups according to the confidence. The high-confidence detection frame group is matched with the tracked trajectory and the lost trajectory. and IoU distance At the same time, it is lower than the threshold value τ set by each cos and τ IoU When , the Hungarian algorithm is used for association by using the cost matrix based on the cosine distance. The cosine distance is used to measure the similarity between two feature vectors and is defined as:
[0019]
[0020] where f i T-1 and f j T They represent the 2048-dimensional feature vectors of trajectory i in the previous frame and trajectory j in the current frame respectively. The dot product operation · and the norm ||·|| are used to calculate the cosine of the angle between the feature vectors.
[0021] The IoU distance is used to measure the degree of overlap between two bounding boxes and is defined as:
[0022]
[0023] in represents the predicted bounding box, represents the detected bounding boxes, and the intersection ∩ and union ∪ are used to calculate the degree of overlap between two bounding boxes.
[0024] The cost matrix of the first association process is:
[0025]
[0026] Cost Matrix Formula C (i,j) Represents the association cost between track i and detection j. If the cosine distance and IoU distance of the two meet the threshold conditions, the cosine distance is used as the association cost, otherwise it is set to 1, indicating mismatch.
[0027] The second association process matches the low-confidence detection frame group with the tracked trajectories and lost trajectories that failed in the first association. The cost matrix of the second association process is:
[0028]
[0029] The third association process matches the unmatched high-confidence detection boxes with the inactive tracks, and the matching method is the same as the first association.
[0030] In step 5, the following sub-steps are included to achieve cross-camera target matching and global ID allocation.
[0031] Hierarchical clustering and Dunn index optimization: A hierarchical clustering algorithm is used in combination with the characteristic cosine distance to generate a link matrix. The quality of different clustering configurations is evaluated by adjusting the distance threshold and calculating the Dunn index. The clustering scheme with the best connectivity is selected to ensure that vehicle targets between different cameras can be accurately matched.
[0032] The connection matrix uses the complete linkage method (Complete Linkage) and is defined as:
[0033]
[0034] where d max (C i ,C j ) is the maximum distance between clusters, C i and C j Represent two different clusters, x and y represent C i and C j Any two trajectory information in d (x,y) Represents the cosine distance between the trajectories x and y.
[0035] The cosine distance is used to measure the similarity between two trajectories and is defined as:
[0036]
[0037] where t x and t y denote trajectory x and trajectory y respectively, and the dot product operation · and the norm ||·|| are used to calculate the cosine of the angle between the trajectories.
[0038] Global ID assignment strategy: For each identified cluster, the system checks whether any track in that cluster already has a global ID. If such a track exists, the rest of the tracks in the same cluster are assigned the same global ID based on feature similarity.
[0039] Hungarian algorithm-assisted matching: For trajectories that have not yet obtained a global ID, the system calculates their pairing distance with the trajectories in the existing clusters and applies the Hungarian algorithm to find the best match. When the confidence of the match exceeds the preset threshold, the global ID of the trajectory is updated.
[0040] Introduction of new global IDs: For those trajectories that cannot be matched to existing clusters through the above steps, the system will assign them new global IDs and establish new clusters. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0042] Figure 1 A flow chart of an online multi-camera multi-vehicle target tracking method based on deep learning provided by an embodiment of the present invention;
[0043] Figure 2A schematic diagram of a three-dimensional association matching process of a single-camera multi-target tracking system provided by an embodiment of the present invention;
[0044] Figure 3 A flowchart of a multi-target trajectory association matching process in a multi-camera environment provided by an embodiment of the present invention; DETAILED DESCRIPTION
[0045] In order to make the purpose and technical solution of the present invention more specific, the present invention will be described in detail below in conjunction with the accompanying drawings and embodiments. The specific embodiments described here are only used to explain the present invention and cannot be understood as limiting the present invention.
[0046] See also Figure 1 , Figure 1 : is a flow chart of an online multi-camera multi-vehicle target tracking method based on deep learning provided by an embodiment of the present invention. Figure 1 As shown, the online multi-target multi-camera vehicle tracking method based on deep learning of this embodiment includes the following steps:
[0047] Step S102, video image acquisition, collects moving video images of the vehicle target within a period of time from multiple cameras at different positions and angles, uses OpenCV library functions to read the video stream frame by frame, and saves each frame as a picture to obtain continuous video frame images.
[0048] Step S104, vehicle detection, uses the YOLOv11 target detection algorithm to process each frame of the image extracted in step 1, and detects and locates all vehicle instances in the image. The algorithm accurately calibrates each detected vehicle instance through a rectangular bounding box, and its output data format is [left, top, right, bottom, confidence], where left and top represent the coordinates of the upper left corner of the bounding box, right and bottom represent the coordinates of the lower right corner, and confidence represents the confidence of the detection result.
[0049] Step S106, vehicle appearance feature extraction, through the ResNet101_IBN convolutional neural network architecture, for each vehicle instance determined in step 2, detailed appearance features are extracted. Specifically, the present invention uses the ResNet-101-IBN model pre-trained on ImageNet as a vehicle appearance feature extractor, pre-trained on 384×384 input images. The step size of the last pooling layer is set to 1 to retain more details of the vehicle, thereby obtaining fine-grained features. The final result is a 2048-dimensional average feature vector for each vehicle instance.
[0050] Furthermore, extensive data augmentation techniques are applied during the training of the vehicle appearance feature extractor, including but not limited to random cropping, horizontal flipping, and color jittering, to improve the generalization and robustness of the model. The feature extraction network is trained using a combination of cross entropy loss and triplet loss, and the loss function can be expressed as:
[0051] L=L CE +λ·L Triplet
[0052] Where L CE represents the cross entropy loss, which is used for classification tasks; L Triplet represents the triplet loss, which is used to learn the ability to distinguish different vehicles; λ is a hyperparameter used to balance the importance of the two losses.
[0053] First, the loss function (Loss) of the feature extraction network of the present invention is divided into two parts; the first part is L CE , the specific formula is as follows:
[0054]
[0055] Where N is the number of training samples per batch, C is the number of vehicle identities, y is the label of the input image, and p ij is the predicted probability.
[0056] The second part is the triplet loss L Triplet , the specific formula is as follows:
[0057]
[0058] where f a is the feature representation of the anchor sample, f p is the feature representation of the positive sample, f n is the feature representation of negative samples, m is the margin, and dist represents the distance between features.
[0059] Step S108, the single camera multi-target (SCMT) tracking algorithm uses the bounding box data of S104 and the feature vector of step S106 to generate the target motion trajectory under each individual camera, thereby realizing target tracking in a single camera environment. The algorithm combines the Kalman filter and feature matching technology to generate the target motion trajectory under each individual camera. The Kalman filter is used to predict the position of the target in the next frame, reduce detection delay and improve tracking continuity. Feature fusion combines the vehicle appearance feature information extracted from the ResNet101_IBN convolutional neural network to enhance the robustness and accuracy of target recognition. Target state management processing includes complex situations such as target loss, re-detection, and deduplication to ensure that tracking can be resumed even after the target temporarily leaves the field of view or is obscured.
[0060] Furthermore, in step S108, the single-camera multi-target tracking algorithm first uses the NSA Kalman filter to predict the bounding box position of each target in the current frame. The state vector is defined as the center position, width and height of the bounding box, and the trajectories are divided into three categories: tracked trajectories, which have been tracked to the last frame; lost trajectories, which have failed somewhere in the past but are still being tracked; and inactive trajectories, which are new trajectories that have just started tracking from the previous frame.
[0061] Figure 2 The three trajectory association processes of the single-camera multi-target tracking system in step S108 are shown: In the first association process, the detection results are divided into high-confidence detection frame groups and low-confidence detection frame groups according to the confidence. The high-confidence detection frame group is matched with the tracked trajectory and the lost trajectory. and IoU distance At the same time, it is lower than the threshold value τ set by each cos and τ IoU When , the Hungarian algorithm is used for association by using the cost matrix based on the cosine distance. The cosine distance is used to measure the similarity between two feature vectors and is defined as:
[0062]
[0063] where f i T-1 and f j T They represent the 2048-dimensional feature vectors of trajectory i in the previous frame and trajectory j in the current frame respectively. The dot product operation · and the norm ||·|| are used to calculate the cosine of the angle between the feature vectors.
[0064] The IoU distance is used to measure the degree of overlap between two bounding boxes and is defined as:
[0065]
[0066] in represents the predicted bounding box, represents the detected bounding boxes, and the intersection ∩ and union ∪ are used to calculate the degree of overlap between two bounding boxes.
[0067] The cost matrix of the first association process is:
[0068]
[0069] Cost Matrix Formula C (i,j) Represents the association cost between track i and detection j. If the cosine distance and IoU distance of the two meet the threshold conditions, the cosine distance is used as the association cost, otherwise it is set to 1, indicating mismatch.
[0070] The second association process matches the low-confidence detection box group with the tracked trajectories and lost trajectories that failed in the first association. The cost matrix of the second association process is:
[0071]
[0072] The third association process matches the unmatched high-confidence detection boxes with the inactive tracks, and the matching method is the same as the first association.
[0073] Step S110, using a hierarchical clustering algorithm to calculate and associate the characteristic distances of the target tracks generated in step S108, and complete the target matching and integration across cameras. Specifically, first, a connection matrix is generated by combining the characteristic cosine distance with the hierarchical clustering algorithm, and the quality of different clustering configurations is evaluated by adjusting the distance threshold and calculating the Dunn index. A clustering scheme with the best connectivity is selected to ensure that vehicle targets between different cameras can be accurately matched. For unidentified tracks, the track showing the minimum characteristic distance between identified tracks is found in the same cluster, and the identifier of the corresponding identified track is assigned to each unidentified track. For the remaining unidentified tracks that cannot be clustered with other tracks, an attempt is made to match them with the missing track group; if the cosine distance between the matching pairs is higher than the set threshold, new IDs are assigned to these tracks. This method does not cluster tracks on the same camera because the same object cannot be detected simultaneously in the same scene.
[0074] Hierarchical clustering and Dunn index optimization: A hierarchical clustering algorithm is used in combination with the characteristic cosine distance to generate a link matrix. The quality of different clustering configurations is evaluated by adjusting the distance threshold and calculating the Dunn index. The clustering scheme with the best connectivity is selected to ensure that vehicle targets between different cameras can be accurately matched.
[0075] The connection matrix uses the complete linkage method (Complete Linkage) and is defined as:
[0076]
[0077] where d max (C i ,C j ) is the maximum distance between clusters, C i and C j Represent two different clusters, x and y represent C i and C j Any two trajectory information in d (x,y) Represents the cosine distance between the trajectories x and y.
[0078] The cosine distance is used to measure the similarity between two trajectories and is defined as:
[0079]
[0080] where t x and t y denote trajectory x and trajectory y respectively, and the dot product operation · and the norm ||·|| are used to calculate the cosine of the angle between the trajectories.
[0081] See also Figure 3 , Figure 3 A flowchart of the multi-target trajectory association matching process in a multi-camera environment provided by an embodiment of the present invention. Figure 3 The multi-camera environment multi-target trajectory association matching process provided by the embodiment of the present invention includes:
[0082] Global ID assignment strategy: For each identified cluster, the system checks whether any track in that cluster already has a global ID. If such a track exists, the rest of the tracks in the same cluster are assigned the same global ID based on feature similarity.
[0083] Hungarian algorithm-assisted matching: For trajectories that have not yet obtained a global ID, the system calculates their pairing distance with the trajectories in the existing clusters and applies the Hungarian algorithm to find the best match. When the confidence of the match exceeds the preset threshold, the global ID of the trajectory is updated.
[0084] Introduction of new global IDs: For those trajectories that cannot be matched to existing clusters through the above steps, the system will assign them new global IDs and establish new clusters.
[0085] Step 6, iteratively execute steps 2 to 5 to achieve consistency and accuracy of multi-camera, multi-target vehicle tracking under continuous video input conditions. By continuously updating and optimizing the tracking results, ensure that the system can respond to new data in real time and continuously update the global ID of each target.
[0086] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention without departing from the principles and intent of the present invention.
Claims
1. An online multi-camera multi-vehicle target tracking method based on deep learning, characterized in that: The following steps are involved: Step 1: Collect moving video images of the vehicle target over a period of time from multiple cameras at different positions and angles, use the OpenCV library function to read the video stream frame by frame, and save each frame as a picture to obtain continuous video frame images; Step 2: Use the YOLO11 target detection algorithm to process each frame of the image extracted in step 1 to detect and locate all vehicle instances in the image; Step 3: Extract the appearance features of each vehicle instance in step 2 through the ResNet101_IBN convolutional neural network architecture to generate a 2048-dimensional feature vector; Step 4: The single camera multi-target (SCMT) tracking algorithm uses the bounding box data of step 2 and the feature vector of step 3 to generate the target motion trajectory under each individual camera, thereby realizing target tracking in a single camera environment; Step 5: Use a hierarchical clustering algorithm combined with feature cosine distance calculation and Dunn index optimization to cluster the target trajectories generated in step 4 to complete cross-camera target matching and integration; Step 6: Iterate the above steps 2 to 5 to achieve multi-camera and multi-target vehicle tracking under continuous video input conditions, and continuously update the feature information and global ID of each target.
2. The method according to claim 1, further characterized in that: In step 1, the video data is read frame by frame using the OpenCV library function, each frame is saved as a picture, and continuous video frame images are obtained.
3. The method for online multi-camera multi-vehicle target tracking based on deep learning according to claim 1, characterized in that: In step 2, an end-to-end target detection method based on deep learning is used, specifically using the YOLOv11s model to detect vehicle targets in the surveillance video stream. Each detected vehicle instance is accurately calibrated by a rectangular bounding box, and the output data format is [left, top, right, bottom, confidence], where left and top represent the coordinates of the upper left corner of the bounding box, right and bottom represent the coordinates of the lower right corner, and confidence represents the confidence of the detection result.
4. The method for online multi-camera multi-vehicle target tracking based on deep learning according to claim 1, characterized in that: In step 3, the ResNet101_IBN convolutional neural network architecture is used to extract detailed appearance features from each vehicle instance detected in step 2 and fused to generate a 2048-dimensional feature vector information for each vehicle target. This feature information captures the visual characteristics of the vehicle, including but not limited to color, shape, and texture.
5. The method for online multi-camera multi-vehicle target tracking based on deep learning according to claim 1, characterized in that: In step 4, a single-camera multi-target tracking system based on Kalman filter and feature matching is constructed. The following techniques are used to track multiple targets in the video frame: Kalman prediction: predicts the target's position in the next frame, reducing detection delays and improving tracking continuity. Feature fusion: Combining the 2048-dimensional vehicle appearance feature information extracted from the ResNet101_IBN convolutional neural network (as described in claim 3) to enhance the robustness and accuracy of target recognition. Target state management: handles complex situations including target loss, re-detection, and de-duplication, ensuring that tracking can be resumed after the target temporarily leaves the field of view or is obscured. Furthermore, the single-camera multi-target tracking system in step 4 includes three trajectory association processes: in the first association process, the detection results are divided into high-confidence detection frame groups and low-confidence detection frame groups according to the confidence. The high-confidence detection frame group is matched with the tracked trajectory and the lost trajectory. and IoU distance At the same time, it is lower than the threshold value τ set by each cos and τ IoU When , the Hungarian algorithm is used for association by using the cost matrix based on the cosine distance. The cosine distance is used to measure the similarity between two feature vectors and is defined as: The IoU distance is used to measure the degree of overlap between two bounding boxes and is defined as: in represents the predicted bounding box, represents the detected bounding box, and the intersection ∩ and union ∪ are used to calculate the overlap between two bounding boxes. The cost matrix of the first association process is. Cost Matrix Formula C (i,j) Represents the association cost between track i and detection j. If the cosine distance and IoU distance of the two meet the threshold conditions, the cosine distance is used as the association cost, otherwise it is set to 1, indicating mismatch. The second association process matches the low-confidence detection box group with the tracked trajectories and lost trajectories that failed in the first association. The cost matrix of the second association process is: The third association process matches the unmatched high-confidence detection boxes with the inactive tracks, and the matching method is the same as the first association.
6. The method for online multi-camera multi-vehicle target tracking based on deep learning according to claim 1, characterized in that: Step 5 further includes the following sub-steps to achieve cross-camera target matching and global ID assignment: Hierarchical clustering and Dunn index optimization: First, a hierarchical clustering algorithm is used to generate a link matrix in combination with the characteristic cosine distance, and the quality of different clustering configurations is evaluated by adjusting the distance threshold and calculating the Dunn index. The clustering scheme with the best connectivity is selected to ensure that vehicle targets between different cameras can be accurately matched. The connection matrix adopts the complete linkage method (Complete Linkage) and is defined as. where d max (C i ,C j ) is the maximum distance between clusters, C i and C j Represent two different clusters, x and y represent C i and C j Any two trajectory information in d (x,y) Represents the cosine distance between the trajectories x and y. The cosine distance is used to measure the similarity between two trajectories and is defined as: where t x and t y denote trajectory x and trajectory y respectively, and the dot product operation · and the norm ||·|| are used to calculate the cosine of the angle between the trajectories. Global ID assignment strategy: For each identified cluster, the system checks whether any track in that cluster already has a global ID. If such a track exists, the rest of the tracks in the same cluster are assigned the same global ID based on feature similarity. Hungarian algorithm-assisted matching: For trajectories that have not yet obtained a global ID, the system calculates their pairing distance with the trajectories in the existing clusters and applies the Hungarian algorithm to find the best match. When the confidence of the match exceeds the preset threshold, the global ID of the trajectory is updated. Introduction of new global IDs: For those trajectories that cannot be matched to existing clusters through the above steps, the system will assign them new global IDs and establish new clusters.
Citation Information
Patent Citations
Cross-camera multi-vehicle tracking method combining road topological structure and overlapped view field
CN117541620A
Delayed online cross-camera multi-target vehicle tracking method based on deep learning
CN118247309A
Cited By
Wharf cross-camera multi-target tracking method and system based on three-dimensional map
CN120876543A
Multi-target continuous motion trail generation method and device for intersection scene
CN121259046A
Real-time cross-camera vehicle tracking method
CN121616627A
Traffic violation behavior detection system based on unmanned aerial vehicle
CN121725635A
Animal target tracking method based on multi-view matching and deep learning
CN121837314A