A multi-camera multi-target vehicle tracking method and system
By integrating the vehicle trajectory and meta-information characteristics in a single camera in a multi-camera system, trajectory matching is optimized, and the problems of vehicle trajectory interruption and identity association errors are solved, and high-precision multi-target vehicle tracking is achieved.
Patent Information
- Application Number
- CN202211274846.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-10-18
AI Technical Summary
In the multi-camera multi-target vehicle tracking, the vehicle trajectory in a single camera is frequently interrupted, the vehicle identity association is incorrect, and the cross-camera re-identification accuracy is not high, especially in the case of occlusion and lighting changes, which are difficult to accurately match.
By obtaining the vehicle trajectory, appearance re-identification features and meta-information features in the single camera video, combining the trajectory fusion in the traffic perception area, using the trajectory smoothness, speed and time differences as auxiliary constraints, vehicle trajectory matching across the camera, and combining vehicle type and color information for joint measurements to optimize trajectory matching.
The vehicle tracking accuracy within a single camera and the vehicle identity correlation accuracy across cameras are improved, and high-precision and robust multi-camera multi-target vehicle tracking is achieved, suitable for complex and congested real-life scenarios.
Smart Images

Figure CN115565157B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle tracking, and particularly relates to a multi-camera multi-target vehicle tracking method and system. Background Art
[0002] In the research of intelligent transportation systems, video analysis using data captured by multiple cameras is of great significance for many applications, such as traffic flow parameter estimation, anomaly detection, multi-camera tracking, etc. Vehicle tracking, as a part of intelligent transportation, has attracted extensive attention from the academic and industrial communities in recent years, especially multi-camera multi-target tracking, which is helpful for traffic flow prediction and analysis.
[0003] Multi-camera multi-target vehicle tracking aims to identify and locate targets in a multi-camera system, and this technology can track multiple detected objects in multiple cameras with overlapping or non-overlapping fields of view. Generally speaking, this technology is divided into 3 sub-tasks: (1) Multi-target tracking within a single camera, usually using a detection-based tracking method. (2) Vehicle re-identification, retrieving the same instance in a large database. (3) Trajectory clustering, aiming to merge the trajectories in the cameras into cross-camera associations. Although good research results have been achieved in tasks such as object detection, tracking, and re-identification, there are still many challenges for a high-performance multi-camera multi-target vehicle tracking framework: (1) Due to unreliable vehicle detection and severe occlusion caused by heavy traffic, it is difficult to track the complete trajectory of a vehicle within a single camera, and trajectory interruptions often occur, resulting in vehicle identity switching. (2) For vehicle re-identification, different shooting angles of the same vehicle, different vehicles of the same model, diverse shooting resolutions, and different lighting conditions in the actual scene are all factors affecting the low accuracy of the re-identification task in the actual scene. Poor performance of single-camera multi-target tracking and re-identification may lead to frequent vehicle identity association errors. In addition, since the re-identification task needs to be based on the results of single-camera vehicle tracking, the vehicle identity association errors introduced within the single-camera field of view will also lead to candidate trajectory association errors in the re-identification task.
[0004] To obtain more accurate multi-camera multi-target vehicle tracking results, it is necessary to enhance the effect of single-camera target tracking and the performance of vehicle re-identification. First, for single-camera target tracking, a trajectory fusion method is needed to fuse interrupted trajectories within a single camera. The approximate range of trajectory fusion can be found based on the characteristics of the interrupted trajectories, and then the range of trajectory fusion can be further narrowed using the characteristics of vehicle movement. Moreover, a robust appearance re-identification feature is required for matching between interrupted trajectories. To associate vehicle trajectories across cameras, appearance-based vehicle re-identification is also one of the most effective methods. For vehicle re-identification, some work focuses on generating discriminative features through deep convolutional neural networks. However, in most methods, the trained re-identification model is used to extract effective embedding features, and the similarity can be estimated based on the Euclidean distance between trajectories during the test phase. On the other hand, vehicle meta-information, such as the type and color of the vehicle, as well as information space and time information, are also key information for assisting multi-camera multi-target tracking, but these information have not been utilized in the prior art, so it is necessary to improve the multi-camera multi-target vehicle tracking in the prior art. Summary of the Invention
[0005] The present invention aims to propose a multi-camera multi-target vehicle tracking method and system based on trajectory fusion and multi-source information assistance, so as to improve the accuracy of vehicle tracking within a single camera under severe vehicle occlusion, as well as the accuracy of vehicle identity association under different shooting angles and different lighting conditions across cameras.
[0006] The multi-camera multi-target vehicle tracking method in the present invention includes the following steps:
[0007] Step 1: Obtain vehicle trajectories within a single camera video, the set of appearance re-identification appearance features of each vehicle trajectory, and the meta-information features of each vehicle trajectory, where the meta-information features include the type feature and color feature of the vehicle;
[0008] Step 2: Obtain the traffic perception area within the shooting range of each camera, and the traffic perception area is the area where vehicle trajectories often get interrupted;
[0009] Step 3: Synthesize the differences between the appearance re-identification features of each vehicle trajectory within a single camera video, as well as the smoothness differences, vehicle speed differences, and time differences between vehicle trajectories, and fuse the interrupted vehicle trajectories within the traffic perception area to obtain complete vehicle trajectories within the single camera video;
[0010] Step 4: Use the joint metric based on the set of appearance re-identification features and meta-information features of vehicle trajectories to perform cross-camera vehicle trajectory matching and merge the complete vehicle trajectories across cameras.
[0011] Further, in step 4, according to the traffic rules and road structure of vehicle driving and the correlation model of the camera, the search space of vehicle trajectory matching is restricted by the time and space constraints during vehicle driving, and the complete vehicle trajectories across cameras are merged by means of hierarchical clustering.
[0012] Further: In step 1, the vehicle trajectories within a single camera video are obtained through a trained object tracking neural network model.
[0013] Further, the object tracking neural network model in step 1 is an object tracking neural network model based on the FairMOT framework and trained by using a dataset annotated with vehicle identity and bounding box position information.
[0014] Further: After obtaining the bounding box position information and vehicle identity information of the vehicle through the object tracking neural network model, the Kalman filter and the Hungarian algorithm are used for matching to obtain the final vehicle trajectories in the single camera video and the vehicle identity information of each vehicle trajectory.
[0015] Further, in step 1, an appearance re-identification feature set of each vehicle trajectory is obtained through a trained video-based re-identification neural network model.
[0016] Further, the re-identification neural network model uses a pre-trained ResNet-50 network as the backbone network, and a BNNeck (Batch Normalization Neck) layer is added between the backbone network and the fully connected layer for classification.
[0017] In the training of obtaining the appearance re-identification feature set of the vehicle trajectory and performing classification, the cross-entropy loss of the classification result is used as the classification loss, and the Hausdorff distance loss formed by the triplet strategy based on the relaxed Hausdorff distance between the appearance re-identification feature sets is used as the metric loss, and the loss function for network training optimization is jointly formed.
[0018] Further, in step 1, the vehicle type meta-information features and vehicle color meta-information features in each frame of the video are respectively extracted through a trained meta-information classification neural network model.
[0019] The meta-information features of each frame in a vehicle trajectory are averaged to obtain the total meta-information features of the vehicle trajectory.
[0020] Further, the meta-information classification neural network model adopts the Light CNN framework, and the network output before the final classification is used as the output of the vehicle's meta-information features.
[0021] Further, the method for obtaining the traffic perception area in step 2 includes: taking the starting points and ending points of the vehicle trajectories in each single camera video as the input of the MeanShift clustering algorithm to cluster out multiple areas;
[0022] Calculate the density of the starting points and ending points of the vehicle trajectories in each area, and find the area with balanced numbers of starting points and ending points as the traffic perception area.
[0023] Further, by calculating the traffic perception area density D ta to measure whether the numbers of starting points and ending points in the area are balanced, the specific formula definition is:
[0024]
[0025] In the formula, N s,k , N e,k respectively represent the numbers of the starting points and ending points of the trajectories in the area;
[0026] If D ta is greater than the threshold ρ ta , then this area is calibrated as the traffic perception area.
[0027] Further, in step 3, calculate the smoothness difference d sm between vehicle trajectories, the speed difference d vc and the time difference d ti , and combine the Euclidean distance d E between the appearance features of vehicle trajectories to obtain the final metric d T for interrupting trajectory fusion in the traffic perception area:
[0028] d T = d E + λ sm d sm + λ vc d vc + λ ti d ti ,
[0029] where λ sm , λ vc , and λ ti are the weights of the smoothness difference d sm , the speed difference d vc and the time difference d ti respectively.
[0030] Further, the calculation process of the smoothness difference d sm is as follows:
[0031]
[0032] where p i,st (t) is the coordinate of the t-th frame among the first n frames of the i-th starting trajectory in the traffic perception area, and p j,nd (t) is the coordinate of the t-th frame among the n frames after the j-th ending trajectory in the traffic perception area, b i,w (t) and b i,h (t) respectively represent the width and length of the vehicle bounding box, X1 to X m are respectively m points evenly distributed on the curve, represents the distance from a point to a line segment.
[0033] Furthermore, the calculation process of the speed difference d vc is as follows:
[0034]
[0035]
[0036] d vc = max(0, |v st - v nd | - γ),
[0037] where γ is the speed boundary value.
[0038] Furthermore, the calculation process of the time difference d ti is as follows:
[0039]
[0040] where t i,st is the timestamp of the starting point of the i-th starting trajectory in the traffic perception area, and t j,nd is the timestamp of the ending point of the j-th ending trajectory in the traffic perception area.
[0041] Furthermore, in step 4, the appearance re-identification feature metric distance and the meta-information feature metric distance between vehicle trajectories are combined in the following way to obtain the final joint metric distance for cross-camera vehicle trajectory fusion:
[0042]
[0043] where T i and T j respectively represent the i-th trajectory and the j-th trajectory;
[0044] M i and M j represent the vehicle type meta-information features of the two vehicle trajectories respectively;
[0045] N i and N jRepresents the vehicle color meta - information features of each of the two vehicle trajectories;
[0046] d E () represents calculating the Euclidean distance between two vectors;
[0047] Represents the set of appearance re - identification features \(S\) of each of the two vehicle trajectories i and \(S\) j the relaxed Hausdorff distance between;
[0048] \(\lambda_1\) and \(\lambda_2\) represent the weights of the distance.
[0049] Furthermore, the relaxed Hausdorff distance is specifically expressed as:
[0050]
[0051]
[0052]
[0053] Wherein, represents selecting the \(k\) - th maximum value in the set .
[0054] Furthermore, the Hausdorff distance loss function is specifically expressed as:
[0055]
[0056] In the formula, \(\tau\) is the distance boundary value, \(P\) represents the number of sampled vehicle identities, \(K\) represents the number of sequences for each identity, \(S\) p represents the positive sample, \(S\) n represents the negative sample.
[0057] Another object of the present invention is to provide a multi - camera multi - target vehicle tracking system, including:
[0058] An acquisition module, configured to acquire vehicle trajectories within a single - camera video, the set of appearance re - identification appearance features of each vehicle trajectory, and the meta - information features of each vehicle trajectory, where the meta - information features include vehicle type features and color features;
[0059] A single - camera trajectory fusion module, which is coupled with a traffic perception area acquisition module, configured to acquire the traffic perception area within the shooting range of each camera;
[0060] And a trajectory fusion module, configured to calculate the differences between the appearance re-identification features of each vehicle trajectory in a single camera video, as well as the smoothness differences, vehicle speed differences, and time differences between vehicle trajectories, and synthesize these differences to fuse the interrupted vehicle trajectories within the traffic perception area to obtain complete vehicle trajectories in the single camera video;
[0061] A cross-camera trajectory fusion module, configured to perform cross-camera vehicle trajectory matching by using the joint metric of the appearance re-identification feature set and meta-information features based on vehicle trajectories, and merge the complete vehicle trajectories across cameras.
[0062] The beneficial effects of the present invention are as follows: For a single camera, the traffic perception area where trajectory interruption is relatively active is found, narrowing the scope for trajectory fusion. By using the characteristics of vehicle trajectories: trajectory smoothness, vehicle speed, and time interval as auxiliary constraint conditions to fuse interrupted trajectories, the problem of frequent ID switching of vehicles, that is, trajectory interruption, caused by poor detection effect, serious vehicle occlusion, and too fast vehicle speed is solved, obtaining a better single camera vehicle tracking effect, so as to better perform vehicle tracking across cameras. And by using meta-information such as vehicle type and color to assist in cross-camera matching of vehicle trajectories, the accuracy of trajectory matching is further improved, realizing a multi-camera multi-target vehicle tracking method with high precision and high robustness that can be applied to different real-world scenarios such as congestion, blur, and complex roads.
[0063] In some embodiments, for multiple cameras, spatio-temporal constraints for vehicle driving are created according to the road structure and camera association model, narrowing the trajectory search space and improving the efficiency. Description of the Drawings
[0064] Figure 1 is a flowchart of the multi-camera multi-target vehicle tracking method in an embodiment of the present invention.
[0065] Figure 2 is a schematic diagram of an exemplary traffic perception area in an embodiment of the present invention.
[0066] Figure 3 is a schematic diagram of the smoothness between adjacent trajectories in an embodiment of the present invention.
[0067] Figure 4 is a schematic diagram of the route between adjacent cameras in an embodiment of the present invention.
[0068] Figure 5 is a schematic logical framework of the multi-camera multi-target vehicle tracking system in an embodiment of the present invention. Detailed Embodiments
[0069] To make the implementation steps and advantages of this invention clearer, I will describe the specific implementation manners in combination with the legends below.
[0070] As Figure 1 shown, the multi-camera multi-target vehicle tracking method based on trajectory fusion and multi-source information assistance of the present invention includes the following steps:
[0071] Step 1: Obtain the vehicle trajectories in the video of a single camera, the appearance re-identification feature sets of each vehicle trajectory, and the meta-information features of each vehicle trajectory, where the meta-information features include the type feature and color feature of the vehicle.
[0072] In this step, the vehicle trajectories in the video of a single camera are obtained through a trained target tracking neural network model. In this embodiment, the target tracking neural network model is a target tracking neural network model based on the FairMOT framework and trained by using a dataset labeled with vehicle identities and bounding box position information.
[0073] FairMOT is a target tracking framework that combines detection and embedding, and the detection method is anchor-free detection based on CenterNet. Therefore, the dataset used to train the FairMOT framework should contain both the position information of the target center and the identity information of the target.
[0074] In this embodiment, vehicle pictures in different scenarios, at different angles, and at different distances are taken, and the vehicle ID of each vehicle, the coordinates of the upper left corner and the lower right corner of the bounding box are labeled, and the center coordinates and length-width values of the vehicle bounding box are obtained through conversion. Then, the color and type of the vehicle are labeled, and the obtained labeled vehicle dataset is used to train the target tracking neural network model in this embodiment.
[0075] The video stream captured by the camera is input into the FairMOT encoding and decoding network, and the detection and re-identification results of the vehicle are obtained from the output feature map. In this embodiment, after obtaining the position and re-identification information of the vehicle, Kalman filtering and the Hungarian algorithm are used for matching to obtain the target vehicle tracking results, that is, the vehicle trajectories, in a single camera, and the local ID of each vehicle trajectory.
[0076] In this step, the appearance re-identification feature sets of each vehicle trajectory are obtained through a trained video-based re-identification neural network model.
[0077] Since vehicle appearance re-identification is mainly used for trajectory association, whether it is the fusion of interrupted trajectories within a single camera or the matching of vehicle trajectories across cameras, the appearance features of the vehicle are required for association. Generally, it is difficult to achieve a good effect by using the features of a single-frame image for trajectory association because the single-frame vehicle image may be a picture with severe occlusion or too much background noise and is difficult to be representative of a trajectory. Therefore, a series of images rather than a single image can be used, that is, video-based vehicle re-identification.
[0078] In the prior art, video-based vehicle re-identification uses temporal attention to perform weighted averaging on the features of each frame or directly averages the features of each frame. These methods all use a series of trajectory frames as the model input, but the effect is not good when the video frame sequence is relatively long or most of the entire trajectory is severely occluded.
[0079] Adopting the measurement method of the Hausdorff distance (PhD) to train the video-based vehicle re-identification model can alleviate the above problems. In this embodiment, the network structure and the corresponding algorithm given in the literature "Zhao, Jianan et al. "PhD Learning: Learning with Pompeiu-hausdorff Distances for Video-based Vehicle Re-Identification." 2021 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021): 2225-2235." are adopted. The ResNet-50 network pre-trained on ImageNet is used as the backbone network of the video-based vehicle appearance re-identification model, and the cross-entropy loss L ID is used as the loss for classification, and the Hausdorff distance loss is used as the measurement loss to jointly form the loss function, and a BNNeck layer is added between the backbone network and the fully connected layer. The Hausdorff distance loss L adopted in this embodiment PhD utilizes a triplet strategy component based on the relaxed Hausdorff distance to evaluate the video-based distance in a set-to-set manner. The formula is as follows:
[0080]
[0081]
[0082]
[0083]
[0084] Where S1 and S2 respectively represent the feature sets of two trajectories, and d E represents the Euclidean distance metric space, represents the relaxed Hausdorff distance between S1 and S2, represents the selection set the k-th maximum value in, and vice versa. τ is the distance boundary value, P represents the number of sampled vehicle identities, K represents the number of sequences for each identity, and S p represents the positive sample, and S n represents the negative sample.
[0085] In this embodiment, here an exemplary total loss L is adopted, which is not limited to the following:
[0086] L = L ID + L PhD ,
[0087] Other ways of using the cross-entropy loss L ID as the loss for classification and the Hausdorff distance loss as the metric loss to jointly form the loss function, such as weighted summation, etc., are well documented in the prior art and will not be elaborated here.
[0088] In this step, a neural network model based on the Light CNN framework is used to classify vehicle meta-information such as vehicle type and vehicle color. The vehicle image is input into this network, and the output obtained is used as the meta-information feature of the vehicle. The meta-information features of a vehicle trajectory are averaged to obtain the final meta-information feature of the vehicle trajectory. The formula is as follows:
[0089]
[0090] ξ i , η i respectively represent the vehicle type and color meta-information features of the i-th frame of the j-th vehicle trajectory. Similarly, for the appearance re-identification features of the vehicle trajectory, they can also be processed in this way. The appearance re-identification features of a vehicle trajectory are averaged to obtain the final appearance re-identification feature of the vehicle trajectory;
[0091] Step 2: Obtain the traffic perception areas within the shooting ranges of each camera. The traffic perception areas are the areas where vehicle trajectories often break;
[0092] The generation of the traffic perception areas is as shown in Figure 2 . First, the starting points and ending points of each trajectory are used as the input for MeanShift clustering to cluster out multiple areas where the starting points or ending points gather. By calculating the densities of the starting points and ending points of each area, traffic perception areas with balanced numbers of starting points and ending points are found. In this embodiment, it is necessary to calculate the density D ta of the traffic perception area. The formula is as follows:
[0093]
[0094] where N s,k and N e,k represent the number of starting points and ending points of the trajectories in the area respectively. If D ta is greater than the threshold ρ ta , then this area is marked as a traffic perception area.
[0095] Step 3: Integrate the interrupted vehicle trajectories in the traffic perception area by synthesizing the differences between the appearance re-identification features of the vehicle trajectories in each single-camera video, as well as the smoothness differences, vehicle speed differences, and time differences between the vehicle trajectories, to obtain the complete vehicle trajectories in the single-camera video;
[0096] The traffic perception area is an area where trajectory interruptions often occur. Therefore, fusing trajectories in the traffic perception area can improve the efficiency and accuracy of trajectory fusion. In this embodiment, the appearance re-identification model of the vehicle is used to obtain the appearance re-identification features of two adjacent trajectories as a measurement benchmark for trajectory fusion. In addition, in this embodiment, the trajectory smoothness, vehicle speed, and time interval between trajectories are used as auxiliary constraint conditions to refine the trajectory fusion.
[0097] As Figure 3 shown, calculate the trajectory smoothness. By Gaussian regression, each point of two adjacent trajectories is regressed to obtain a curve, and then calculate the distance from each point of the two trajectories to the curve, and take the average value to obtain the smoothness difference d sm between the two trajectories. The formula is as follows:
[0098]
[0099] where p i,st (t) is the coordinate of the t-th frame in the first n frames of the i-th starting trajectory, and p j,nd (t) is the coordinate of the t-th frame in the last n frames of the j-th ending trajectory, that is, the coordinates of the adjacent n frames of the two adjacent trajectories respectively. b i,w (t) and b i,h (t) represent the width and length of the vehicle bounding box respectively. X1 to X m are m points evenly distributed on the curve respectively. Then is the distance formula from a point to a line segment.
[0100] Here, the distance of the vehicle position in the pixel coordinates is used to roughly calculate the vehicle speed, and then the change in the vehicle speed of two adjacent trajectories is calculated to measure the speed difference d vc between two adjacent trajectories. The formula is as follows:
[0101]
[0102]
[0103] d vc = max(0, |v st - v nd | - γ),
[0104] where γ is the speed boundary value.
[0105] In this embodiment, the timestamp t i,st of the first frame of the starting trajectory and the timestamp t j,nd of the last frame of the ending trajectory are obtained, and the interval between t i,st and t j,nd is calculated to obtain the time difference d ti of the trajectory. If there is time overlap between two trajectories, that is, t i,st < t j,nd , then they must not be the same trajectory. At this time, d ti is recorded as infinity, and the formula is as follows:
[0106]
[0107] Finally, based on the Euclidean distance d E between the appearance re-identification features of the vehicle trajectories, plus auxiliary constraint conditions such as trajectory smoothness difference, vehicle speed difference, and trajectory time difference, the final metric d T for interrupted trajectory fusion within the traffic perception area is obtained:
[0108] d T = d E + λ sm d sm + λ vc d vc + λ ti d ti ,
[0109] where λ sm , λ vc , and λ ti are the weights of the smoothness difference d sm , speed difference d vc , and time difference d ti , respectively.
[0110] Using this metric for vehicle trajectory matching refines the effect of trajectory fusion.
[0111] Step 4: Use the joint metric based on the appearance re-identification feature set and meta-information features of the vehicle to perform cross-camera vehicle trajectory matching and merge the complete cross-camera vehicle trajectories.
[0112] In vehicle identity association across cameras, it is necessary to combine the appearance feature metric and the meta-information feature metric of vehicle trajectories to obtain the final joint metric distance of vehicle identities:
[0113]
[0114] where T i and T j represent the i-th trajectory and the j-th trajectory respectively, and λ1 and λ2 represent the weights of the distances.
[0115] To perform vehicle trajectory matching across cameras, in this embodiment, first, a joint metric based on the vehicle appearance features and meta-information features in the video is constructed to calculate the distance matrix D between adjacent camera trajectories as follows:.
[0116]
[0117]
[0118] T S,i , i ∈ (1, 2…n) represents one of the n trajectories in the source camera, and T T,j , j ∈ (1, 2…m) represents one of the m trajectories in the destination camera. In this embodiment, k-reciprocal reordering is used to refine the updated distance matrix, thereby generating a stronger distance matrix D.
[0119] Due to road structures and traffic rules, the movement of vehicles follows specific driving patterns. Here, each feasible road is replaced by a route. There may be multiple routes in one camera view, but only specific several routes lead to the adjacent camera view. As Figure 4 shown, only by entering the main road from routes 1-3 of Camera 1 and driving out, can one drive into the main road of Camera 2 from routes 4-6, and then calculate the distance from the vehicle trajectory to the route to assign the vehicle to the specified route. In this way, when calculating the distance matrix between the trajectories of Camera 1 and Camera 2, only the matching situation between the vehicles belonging to routes 1-3 and the vehicles belonging to routes 4-6 needs to be considered, reducing the search space for trajectory matching. In this way, a spatial constraint for the vehicle driving from Camera 1 to Camera 2 can be obtained.
[0120] The camera association model uses the camera topology to generate time constraints. After obtaining the calibrated vehicle trajectories and camera links, the transition time of the vehicle in the camera link can be automatically estimated without manual annotation. The transition time is defined as:
[0121] Δt = t d - t s ,
[0122] t s and t d respectively represent the time when the vehicle leaves the source camera area and the time when it arrives at the destination camera area. Then, a transition time window (Δt min -ε, Δt max +ε) can be obtained for each camera link, where ε is the time boundary value. Only vehicle pairs with a transition time within the time window are considered valid. Therefore, the search space for trajectory matching can be further reduced by an appropriate time window. In this way, a time constraint for the vehicle traveling from Camera 1 to Camera 2 can be obtained.
[0123] By integrating the spatio-temporal constraints of the vehicle, we can obtain the matching of vehicles between two adjacent cameras as follows:
[0124]
[0125] Thus, we can obtain the mask matrix M between adjacent cameras as follows:
[0126]
[0127] Finally, we combine the distance matrix with the mask matrix to obtain the distance matrix after spatio-temporal constraints
[0128]
[0129] ⊙ represents the dot product of matrices, and then replace the 0 values in with infinity. Thus, the final spatio-temporal constraint distance matrix is obtained for subsequent trajectory clustering. Then, the vehicle trajectories in each pair of adjacent cameras are matched by the hierarchical clustering method, and finally the global ID of the vehicle trajectories is obtained.
[0130] This embodiment also discloses a multi-camera multi-target vehicle tracking system, which is basically as shown in Figure 5 and includes:
[0131] An acquisition module, configured to acquire vehicle trajectories in a single camera video, the appearance re-identification appearance feature set of each vehicle trajectory, and the meta-information features of each vehicle trajectory, where the meta-information features include the type feature and color feature of the vehicle;
[0132] As shown in the figure, this module includes a vehicle trajectory acquisition module, and a trained target tracking neural network model, which is not limited to the one described above, is set in this module for acquiring vehicle trajectories in a single camera video;
[0133] Optionally, in this module, after obtaining the bounding box position information and vehicle identity information of the vehicle through the target tracking neural network model, the Kalman filter and the Hungarian algorithm are used for matching to obtain the final vehicle trajectories in the single-camera video and the vehicle identity information of each vehicle trajectory.
[0134] The acquisition module further includes an appearance re-identification feature acquisition module, which is provided with a trained video-based re-identification neural network model, which is not limited to the one described in the previous text, for obtaining the appearance re-identification feature set of each vehicle trajectory. The training method of the re-identification neural network model is also as described in the previous text and is not limited thereto.
[0135] The acquisition module further includes two trained meta-information classification neural network models, which are not limited to the ones described in the previous text, and are respectively used to extract the vehicle type meta-information feature and the vehicle color meta-information feature in each frame of the video.
[0136] The system further includes a single-camera trajectory fusion module, which is coupled with a traffic perception area acquisition module for obtaining the traffic perception area within the shooting range of each camera. This module can obtain the traffic perception area through methods and steps that are not limited to those described in the previous text.
[0137] And a trajectory fusion module for calculating the differences between the appearance re-identification features of each vehicle trajectory in the single-camera video, as well as the smoothness differences, vehicle speed differences, and time differences between vehicle trajectories, and integrating these differences to fuse the interrupted vehicle trajectories within the traffic perception area to obtain the complete vehicle trajectories in the single-camera video. The module can calculate the differences between the appearance re-identification features of vehicle trajectories, as well as the smoothness differences, vehicle speed differences, and time differences between vehicle trajectories, and realize the fusion of vehicle trajectories through methods and steps that are not limited to those described in the previous text.
[0138] The system further includes a cross-camera trajectory fusion module for performing cross-camera vehicle trajectory matching using the joint metric based on the appearance re-identification feature set and meta-information features of vehicle trajectories, and merging the complete cross-camera vehicle trajectories.
[0139] This module can calculate the joint metric through methods and steps that are not limited to those described in the previous text.
[0140] This module can, as in the methods and steps described in the previous text, according to the traffic rules and road structures of vehicle driving and the association model of cameras, and utilize the time and space constraints during vehicle driving to limit the search space for vehicle trajectory matching, and merge the complete cross-camera vehicle trajectories through the hierarchical clustering method.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-camera multi-target vehicle tracking method, characterized in that It includes the following steps: Step 1: Obtain the vehicle trajectories in a single camera video, the appearance re-identification appearance feature set of each vehicle trajectory, and the meta-information features of each vehicle trajectory, where the meta-information features include the type feature and color feature of the vehicle; Step 2: Obtain the traffic perception areas within the shooting ranges of each camera, where the traffic perception areas are the areas where vehicle trajectories often break; The method for obtaining the traffic perception areas includes: taking the starting points and ending points of each vehicle trajectory in a single camera video as the input of the MeanShift clustering algorithm to cluster out multiple areas; Calculate the densities of the starting points and ending points of the vehicle trajectories in each area, and find the area with balanced numbers of starting points and ending points as the traffic perception area; Among them, by calculating the density of the traffic perception area to measure whether the number of starting points and ending points in the area is balanced, and the specific formula definition is as follows: , where , respectively represent the numbers of the starting point and the ending point of the trajectory within the area; If is greater than the threshold , then this area is designated as a traffic perception area; Step 3: Synthesize the differences between the appearance re-identification features of each vehicle trajectory in a single camera video, as well as the smoothness differences, vehicle speed differences, and time differences between vehicle trajectories, and fuse the interrupted vehicle trajectories within the traffic perception area to obtain the complete vehicle trajectories in the single camera video; Among them, calculate the smoothness difference, speed difference, and time difference , and combine the Euclidean distance between the appearance features between vehicle trajectories to obtain the final metric for interrupting trajectory fusion within the traffic perception area : , Among them, are the smoothness difference , speed difference and time difference weights respectively; Step 4: Perform cross-camera vehicle trajectory matching using the joint metric based on the appearance re-identification feature set and meta-information features of the vehicle trajectories, and merge the complete vehicle trajectories across cameras.
2. The method according to claim 1, wherein In Step 4, according to the traffic rules and road structures of vehicle driving and the association model of cameras, use the time and space constraints during vehicle driving to limit the search space for vehicle trajectory matching, and merge the complete vehicle trajectories across cameras through the hierarchical clustering method.
3. The method according to claim 1, wherein In Step 1, obtain the vehicle trajectories in a single camera video through a trained target tracking neural network model.
4. The method according to claim 1, wherein In Step 1, obtain the appearance re-identification feature set of each vehicle trajectory through a trained video-based re-identification neural network model.
5. The method according to claim 1, characterized in that, In Step 1, extract the vehicle type meta-information feature and vehicle color meta-information feature in each frame of the video respectively through a trained meta-information classification neural network model; Calculate the average value of the meta-information features of each frame in a vehicle trajectory to obtain the total meta-information feature of the vehicle trajectory.
6. The method according to claim 1, wherein In Step 4, combine the appearance re-identification feature metric distance and meta-information feature metric distance between vehicle trajectories in the following way to obtain the final joint metric distance for cross-camera vehicle trajectory fusion: , Among them, respectively represent the i-th trajectory and the j-th trajectory; M i and M j represent the vehicle type meta-information features of the respective vehicle trajectories of two vehicles; N i and N j represent the vehicle color meta-information features of the respective vehicle trajectories of two vehicles; () represents calculating the Euclidean distance between two vectors; representing the appearance re-identification feature sets S of the respective vehicle trajectories i and S j the relaxed Hausdorff distance between them; and represent the weights of distances.
7. A multi-camera multi-target vehicle tracking system, characterized in that, It includes: An acquisition module, used to obtain the vehicle trajectories in a single camera video, the appearance re-identification appearance feature set of each vehicle trajectory, and the meta-information features of each vehicle trajectory, where the meta-information features include the type feature and color feature of the vehicle; A single camera trajectory fusion module, which is coupled with a traffic perception area acquisition module, used to obtain the traffic perception areas within the shooting ranges of each camera, including: taking the starting points and ending points of each vehicle trajectory in a single camera video as the input of the MeanShift clustering algorithm to cluster out multiple areas; Calculate the densities of the starting points and ending points of the vehicle trajectories in each area, and find the area with balanced numbers of starting points and ending points as the traffic perception area; Among them, by calculating the traffic perception area density to measure whether the number of starting points and ending points in the area is balanced. The specific formula definition is as follows: , wherein , respectively represent the number of the starting point and the ending point of the trajectory in the area; If is greater than the threshold , then this area is designated as a traffic perception area; And a trajectory fusion module, configured to calculate the differences between the appearance re-identification features of each vehicle trajectory in a single camera video, as well as the smoothness difference, vehicle speed difference, and time difference between vehicle trajectories, and integrate these differences to fuse interrupted vehicle trajectories within the traffic perception area to obtain complete vehicle trajectories in the single camera video; wherein, calculating the smoothness difference between vehicle trajectories , speed difference and time difference , and combining the Euclidean distance between the appearance features between vehicle trajectories to obtain the final metric for interrupted trajectory fusion within the traffic perception area : , Among them, are the smoothness difference , speed difference and time difference weights; A cross-camera trajectory fusion module is used to perform cross-camera vehicle trajectory matching by using the joint metric of the appearance re-identification feature set based on vehicle trajectories and meta-information features, and merge the complete vehicle trajectories across cameras.