Method, device and equipment for tracking multiple vehicles across cameras on expressway

Through the lightweight vehicle object detection model and the domain conversion network ReID model, combined with feature compensation strategy and trajectory clustering algorithm, the vehicle tracking difficulties caused by fast vehicle speed and light changes on the highway are solved, and accurate cross-camera vehicle tracking on the highway is achieved.

CN120339329APending Publication Date: 2025-07-18ZHEJIANG INST OF COMM CO LTD +5

Patent Information

Application Number
CN202510408519.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing cross-camera multi-vehicle tracking technology in urban blocks is difficult to effectively apply in highway scenarios. It is mainly due to the rapid speed of vehicles on the highway, resulting in blurred image, vehicle occlusion and light changes interference, affecting the accuracy of target detection and tracking.

Method used

The lightweight vehicle object detection model and the domain conversion network ReID model are adopted, combined with gated convolution and feature compensation strategies, vehicle detection and appearance feature extraction are carried out, and vehicle tracking across cameras is achieved through trajectory tracking and clustering algorithms.

Benefits of technology

Accurate cross-camera vehicle tracking is implemented on the highway, reducing tracking failures caused by ID identity transformation, and improving vehicle detection accuracy and tracking stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339329A_ABST
    Figure CN120339329A_ABST
Patent Text Reader

Abstract

The invention provides an expressway camera-crossing multi-vehicle tracking method, device and equipment, and relates to the technical field of vehicle trajectory tracking, and the method comprises the steps: obtaining a monitoring video collected by each camera; for any camera, inputting any video frame in the monitoring video collected by the camera into a pre-trained lightweight vehicle target detection model to obtain a vehicle detection result of each vehicle in the video frame; extracting appearance features in the image of each vehicle to obtain the appearance features of each vehicle; performing trajectory tracking on each vehicle; and based on the tracks and the corresponding appearance characteristics of the different vehicles in the monitoring videos acquired by the different cameras, determining the tracking results of the different vehicles across the cameras. According to the method and the device, cross-camera vehicle tracking can be accurately carried out on the expressway in real time, and the situation of tracking failure caused by change of IDs tracked by multiple vehicles among the cross cameras is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of vehicle trajectory tracking. Specifically, it relates to a method, device and equipment for tracking multiple vehicles across cameras on highways. Background Art

[0002] Although there are currently tracking technologies for multiple vehicles across cameras in traditional urban blocks, applying the above tracking technologies to highway scenarios has the following problems:

[0003] First, vehicles on highways are faster than those in urban blocks, resulting in it being difficult for cameras to clearly capture or the vehicle pictures captured by cameras being blurred, reducing the accuracy of target detection;

[0004] Secondly, there is a problem of large vehicles blocking small vehicles on highways, which may cause target tracking to be lost. At the same time, due to the similarity of the brands and colors of some vehicles on highways, it is difficult to extract the appearance features of vehicles.

[0005] Finally, the change in light intensity in highway tunnels is more obvious, which is likely to interfere with the appearance features of vehicles and affect the correct distinction of the features of target vehicles.

[0006] In summary, the current tracking technology for multiple vehicles across cameras in urban streets is difficult to be directly applied to the scenario of tracking multiple vehicles across cameras on highways. Summary of the Invention

[0007] The purpose of the embodiments of the present application is to provide a method, device and equipment for tracking multiple vehicles across cameras on highways, so as to solve the above problems existing in the prior art, and can accurately track vehicles across cameras in real time on highways, reducing the occurrence of tracking failures caused by the change of ID identities of multiple vehicles during tracking across cameras.

[0008] In a first aspect, the present invention provides a method for tracking multiple vehicles across cameras on highways. Cameras are respectively set on each section of the highway, and each camera corresponds to a different monitoring area. The method includes:

[0009] Obtain the monitoring videos collected by each camera;

[0010] For any camera, input any video frame in the monitoring video collected by the camera into a pre-trained lightweight vehicle target detection model to obtain the vehicle detection results of each vehicle in the video frame; wherein, the vehicle detection results include: detection frames of different vehicles and corresponding coordinates;

[0011] Crop images of each vehicle from the video frame according to the coordinates of each detection box;

[0012] Extract appearance features from the images of each vehicle to obtain the appearance features of each vehicle;

[0013] Track the trajectories of each vehicle based on the vehicle detection results and corresponding appearance features of each video frame to obtain the trajectories of each vehicle in the surveillance video captured by the camera;

[0014] Determine the tracking results of different vehicles across cameras based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos captured by different cameras.

[0015] In an optional implementation manner, the lightweight vehicle target detection model includes:

[0016] An input layer for receiving video frames;

[0017] A gated convolution layer for using the video frame as the input feature of the gated convolution layer, performing linear projection on the input feature to obtain a first feature and a second feature; performing depth convolution on the second feature to obtain a third feature; performing first-order interaction on the third feature and the first feature to obtain a first-order spatial feature; using the first-order spatial feature as a new input feature, and returning to perform linear projection on the input feature until a configured termination condition is met, and outputting a high-order spatial feature;

[0018] A classification layer for classifying according to the high-order spatial feature to determine whether there is a vehicle at each position of the video frame;

[0019] A prediction layer for performing regression using the high-order spatial feature to generate vehicle detection results of each vehicle existing in the video frame.

[0020] In an optional implementation manner, extracting appearance features from the images of each vehicle to obtain the appearance features of each vehicle includes:

[0021] Input the images of each vehicle into a pre-trained domain conversion network ReID model to obtain the appearance features of each vehicle.

[0022] In an optional implementation manner, the vehicle detection result further includes: the confidence of different detection boxes and the vehicle IDs of different vehicles;

[0023] The method further includes:

[0024] If the confidence of any detection box is greater than the configured confidence threshold, then use the detection box as the first detection box; otherwise, use the detection box as the second detection box;

[0025] For any vehicle ID, generate the average appearance feature of the vehicle according to the appearance features of the vehicle corresponding to the vehicle ID in each video frame of the monitoring video.

[0026] In an alternative embodiment, perform trajectory tracking on each vehicle according to the vehicle detection results and corresponding appearance features of each video frame, including:

[0027] For any two adjacent video frames in the monitoring video collected by any camera, use the video frame with the earlier frame number in the adjacent video frames as the first target frame, and use the video frame with the later frame number in the adjacent video frames as the second target frame;

[0028] Use the vehicles corresponding to each first detection box in the second target frame as the first vehicles; use the vehicle corresponding to any first detection box in the first target frame as the vehicle to be matched;

[0029] For any first vehicle, if the similarity between the appearance feature of the first vehicle and the appearance feature of any vehicle to be matched is greater than the configured first similarity threshold, then use the vehicle to be matched as the target matching vehicle of the first vehicle;

[0030] Predict the predicted detection box of the target matching vehicle in the second target frame according to the detection box of the target matching vehicle in the first target frame;

[0031] If the intersection over union of the detection box of the first vehicle and the predicted detection box of the target matching vehicle is greater than the configured intersection over union threshold, then modify the vehicle ID of the first vehicle to the vehicle ID of the target matching vehicle;

[0032] For any vehicle ID, generate the trajectory of the corresponding vehicle in the monitoring video collected by the camera according to the coordinates of the vehicle corresponding to the vehicle ID in each video frame of the monitoring video.

[0033] In an alternative embodiment, the method further includes:

[0034] Use the vehicle corresponding to any second detection box in the second target frame as the second vehicle;

[0035] For any second vehicle, if the similarity between the appearance feature of the second vehicle and the appearance feature of any vehicle to be matched is greater than the configured second similarity threshold, then use the vehicle to be matched as the target matching vehicle of the second vehicle;

[0036] Perform feature compensation on the appearance feature of the second vehicle based on the average appearance feature of the target matching vehicle to obtain the updated appearance feature;

[0037] If the similarity between the updated appearance features of the second vehicle and the appearance features of any vehicle to be matched is greater than the configured first similarity threshold, then the vehicle to be matched is taken as the target matching vehicle of the second vehicle;

[0038] According to the detection frame of the target matching vehicle in the first target frame, predict the predicted detection frame of the target matching vehicle in the second target frame;

[0039] If the intersection over union of the detection frame of the second vehicle and the predicted detection frame of the target matching vehicle is greater than the configured intersection over union threshold, then modify the vehicle ID of the second vehicle to the vehicle ID of the target matching vehicle, and adjust the corresponding second detection frame of the second vehicle to the first detection frame.

[0040] In an alternative embodiment, based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos collected by different cameras, determine the tracking results of different vehicles across cameras, including:

[0041] For any camera, use the cameras on the adjacent road sections of the road section where the camera is located as the adjacent cameras of the camera;

[0042] Cluster the trajectories of different vehicles in the surveillance video collected by the camera to obtain a clustering result;

[0043] For any adjacent camera, obtain the updated similarity matrix of the adjacent camera and the camera;

[0044] According to the updated similarity matrix, cluster the trajectories of different vehicles in the surveillance video collected by the adjacent camera and the clustering result to obtain the cross-camera tracking results of each vehicle under the camera and the adjacent camera.

[0045] In a second aspect, the present invention provides a tracking device for multiple vehicles across cameras on a highway. Cameras are respectively arranged on each road section of the highway, and each camera corresponds to a different surveillance area. The device includes:

[0046] An acquisition unit, configured to acquire the surveillance videos collected by each camera;

[0047] A detection unit, configured to, for any camera, input any video frame in the surveillance video collected by the camera into a pre-trained lightweight vehicle target detection model to obtain the vehicle detection results of each vehicle in the video frame; wherein, the vehicle detection results include: detection frames of different vehicles and corresponding coordinates;

[0048] A cropping unit, configured to crop the images of each vehicle from the video frame according to the coordinates of each detection frame;

[0049] An extraction unit, configured to extract appearance features from the images of each vehicle to obtain the appearance features of each vehicle;

[0050] A tracking unit, configured to perform trajectory tracking on each vehicle according to the vehicle detection results of each video frame and the corresponding appearance features, so as to obtain the trajectories of each vehicle in the surveillance video captured by the camera;

[0051] A determination unit, configured to determine the tracking results of different vehicles across cameras based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos captured by different cameras.

[0052] In a third aspect, the present invention provides an electronic device, which includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;

[0053] The memory is used to store a computer program;

[0054] The processor is configured to implement the method according to any one of the foregoing embodiments when executing the program stored on the memory.

[0055] In a fourth aspect, the present invention provides a computer-readable storage medium, in which a computer program is stored, and the computer program realizes the method according to any one of the foregoing embodiments when executed by a processor.

[0056] This application adopts a lightweight vehicle target detection model optimized by gated recurrent convolution, and only completes the information interaction of high-order space using convolutional and fully connected layers, reduces the number of parameters and the amount of computation of the detection model, realizes high-efficiency vehicle detection, and is applicable to vehicle detection in highway scenarios; by combining different foregrounds and backgrounds to increase the training samples to train the domain transformation network ReID model, the domain transformation network ReID model can effectively perform domain transformation on situations such as motion blur and strong light illumination, so as to accurately construct the key structural appearance features of the vehicle.

[0057] This application adopts a tracking strategy of feature compensation, classifies vehicles into a high-confidence group and a low-confidence group, uses different tracking matching methods for vehicles with different confidences, and mines the vehicle features of the un-matched vehicle trajectories in the high-confidence group and uses them as compensation features to compensate for the occluded objects in the low-confidence group, thereby reducing the missing of the single-camera vehicle tracking trajectories of some vehicles due to occlusion.

[0058] This application uses a variety of post-processing techniques to filter the single-camera vehicle tracking trajectories and perform associated matching of the single-camera vehicle tracking trajectories between multiple cameras to form the final cross-camera vehicle tracking trajectories. Description of the Drawings

[0059] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0060] Figure 1 It is a flowchart of a method for tracking multiple vehicles across cameras on a highway provided by an embodiment of the present application;

[0061] Figure 2 It is a schematic diagram of a method for tracking multiple vehicles across cameras on a highway provided by an embodiment of the present application;

[0062] Figure 3 It is an architecture diagram of a method for tracking multiple vehicles across cameras on a highway provided by an embodiment of the present application;

[0063] Figure 4 It is a schematic structural diagram of a device for tracking multiple vehicles across cameras on a highway provided by an embodiment of the present application;

[0064] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments

[0065] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all of them. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.

[0066] The method for tracking multiple vehicles across cameras on a highway provided by the embodiments of the present application first obtains the monitoring videos collected by each camera, and then uses a lightweight vehicle target detection model optimized by gated convolution to process the video frames to obtain detection results including vehicle detection frames, coordinates, confidence levels, and vehicle IDs. Then, the vehicle images are cropped according to the detection frame coordinates, and the appearance features are extracted through the domain conversion network ReID model. On this basis, trajectory tracking is performed according to the vehicle detection results and appearance features, and the vehicles are classified and processed according to the detection frame confidence levels. Specific matching strategies are adopted for vehicles with different confidence levels, effectively reducing the missing of single-camera tracking trajectories.

[0067] When determining the cross-camera tracking results, the present application utilizes trajectory clustering and similarity matrix analysis. First, it clusters the vehicle trajectories within a single camera, then obtains the updated similarity matrix between adjacent cameras, and determines the cross-camera tracking results by synthesizing the above information. Meanwhile, it filters the trajectories with the help of time constraints and region constraints, and optimizes the similarity matrix using the k-complementary nearest neighbor algorithm to ensure accurate results.

[0068] Compared with Comparative Document 1 (CN202311180228.9, a multi-vehicle tracking method and system for cross-cameras on highways combining multiple models), in cross-camera trajectory matching, the present application improves the matching accuracy and anti-interference ability through innovative clustering and matrix optimization methods. In Comparative Document 1, the combination of multiple models is likely to lead to a complex system and a large amount of computation, and its feature extraction ability is weak when dealing with complex environments. The present application, relying on technologies such as gated convolution and domain transformation network, can efficiently handle complex highway scenarios, achieve accurate and stable cross-camera multi-vehicle tracking, and significantly reduce the probability of tracking failure caused by ID identity transformation.

[0069] The multi-vehicle tracking method for cross-cameras on highways provided by the embodiments of the present application can be applied in a server or in a terminal with strong computing power. The server can be a physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), as well as big data and artificial intelligence platforms. The terminal can be a user equipment (UE) such as a mobile phone, a smart phone, a laptop computer, a digital broadcast receiver, a personal digital assistant (PDA), a tablet computer (PAD), a handheld device, a vehicle-mounted device, a wearable device, a computing device or other processing devices connected to a wireless modem, a mobile station (MS), a mobile terminal, etc. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any limitations here.

[0070] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0071] Figure 1 It is a schematic flowchart of a multi-vehicle tracking method for cross-cameras on highways provided by the embodiments of the present application. As Figure 1As shown, the method may include:

[0072] Step S110: Obtain the surveillance videos collected by each camera. For any camera, input any video frame in the surveillance video collected by this camera into a pre-trained lightweight vehicle target detection model to obtain the vehicle detection results of each vehicle in this video frame.

[0073] In the embodiments of the present application, cameras are respectively arranged on each section of the highway, and each camera corresponds to a different surveillance area; the cameras can be numbered according to the order in which each vehicle passes through each section to obtain the camera ID of each camera, and the order in which the vehicle passes through each camera can be determined according to the camera ID. For example, on a 1000-meter-long highway from south to north, a camera is arranged every 200 meters; taking the south end of this highway as the starting point, that is, 0 meters, the camera ID of the camera arranged at the 200-meter mark is 1, and the camera IDs of the cameras at the 400-meter, 600-meter, 800-meter, and 1000-meter marks are 2, 3, 4, and 5 respectively.

[0074] In the embodiments of the present application, the lightweight vehicle target detection model includes:

[0075] An input layer for receiving the video frame;

[0076] A gated convolutional layer for using the video frame as the input feature of the gated convolutional layer, performing linear projection on the input feature to obtain a first feature and a second feature; performing depth convolution on the second feature to obtain a third feature; performing first-order interaction on the third feature and the first feature to obtain a first-order spatial feature; using the first-order spatial feature as the new input feature, and returning to perform linear projection on the input feature until the configured termination condition is met, and outputting a high-order spatial feature;

[0077] A classification layer for classifying according to the high-order spatial feature to determine whether there are vehicles at each position of the video frame;

[0078] A prediction layer for performing regression using the high-order spatial feature to generate the vehicle detection results of each vehicle existing in the video frame.

[0079] In the embodiments of the present application, the gated convolutional layer adopts the gated recursive convolutional structure of Hor-Net. The gated convolutional layer can completely model the high-order spatial interaction based on convolution. By performing convolution on the concatenated features through linear projection mapping, the amount of computation can be greatly reduced; the gated convolution is executed cyclically to achieve interaction with higher-order features.

[0080] In the embodiments of the present application, the interaction formula of the gated convolutional layer is as follows:

[0081]

[0082] P1 = f(q0) ⊙ p0 ∈ R HW×C , y = φ out (p1) ∈ R HW×C

[0083] where φ in represents the channel splitting operation, generating the first feature p0 and the second feature q0; φ out represents the fully connected layer, generating the densified feature; H, W represent the size of the feature, and C represents the channel dimension of the feature; f(·) represents the depthwise separable convolution. The feature extracted from the second feature (or query feature vector) q0 through the depthwise separable convolution is dot - producted with the first feature p0 to obtain the p1 feature, i.e., the first - order spatial feature, thereby realizing the first - order interaction only using convolution and the fully connected layer, and circularly performing high - order linear projection and high - order spatial interaction to obtain the high - order spatial feature;

[0084] The formula for high - order linear projection is as follows:

[0085]

[0086] where C0 represents the channel dimension of the first feature p0, and C k represents the dimension of the second feature q k at the k - th layer; n - 1 represents the number of layers of the second feature, that is, the second feature is split into n - 1 sub - features.

[0087] In an embodiment of the present application, the configurable termination condition may be that the number of interactions reaches 5 times or 6 times.

[0088] In an embodiment of the present application, the classification layer uses the fully connected layer to reduce the channel number of the high - order spatial feature to N (here N is the known number of categories. In highway target tracking, the main target categories are as follows: car, truck, bus, others), and the target category class corresponding to the current video frame image is obtained through the classification layer. The detections corresponding to car, truck, and bus are all vehicle detection frames.

[0089] In an embodiment of the present application, the prediction layer uses NMS and IOU to repeatedly detect and remove the predicted targets, obtaining the final vehicle detection result.

[0090] In an embodiment of the present application, the vehicle detection result includes: the detection frames and corresponding coordinates of different vehicles, as well as the confidence levels of different detection frames and the vehicle IDs of different vehicles.

[0091] In an embodiment of the present application, the lightweight vehicle target detection model is trained with the images marked with target categories, and iteratively trained until the set conditions are met.

[0092] In the embodiments of the present application, a confidence threshold is preset in advance. Among the vehicle detection results of each video frame, the detection boxes with confidence values less than (or not greater than) the confidence threshold are used as the second detection boxes; the detection boxes with confidence values greater than or equal to (or greater than) the confidence threshold are used as the first detection boxes; each first detection box is used as the high-confidence group of the vehicle detection results of the video frame; and each second detection box is used as the low-confidence group of the vehicle detection results of the video frame.

[0093] In the embodiments of the present application, the confidence threshold is determined based on the confidence values of all vehicle detection boxes under each camera, and the formula is as follows:

[0094]

[0095] where N represents the number of cameras, M l represents the number of vehicle images corresponding to the vehicle detection boxes under the l-th camera, and C l,m represents the confidence value of the vehicle detection box of the m-th vehicle under the l-th camera.

[0096] In the embodiments of the present application, the special structure of the gated recurrent convolution causes only the convolution and fully connected layers to be used to complete the information interaction extraction of the features in the high-order space for vehicle detection, thereby obtaining accurate vehicle detection boxes and improving the accuracy of vehicle target detection in blurred images.

[0097] Step S120: According to the coordinates of each detection box, crop the images of each vehicle from the video frame, and extract the appearance features in the images of each vehicle to obtain the appearance features of each vehicle.

[0098] In the embodiments of the present application, extracting the appearance features in the images of each vehicle includes: inputting the images of each vehicle into a pre-trained domain conversion network ReID model to obtain the appearance features of each vehicle.

[0099] In practical applications, due to the motion blur caused by the high-speed movement of vehicles and the change of vehicle lighting conditions, traditional methods cannot effectively extract key structural appearance features. By combining different foregrounds and backgrounds to increase training samples, the trained domain converter can effectively perform domain conversion on situations such as motion blur and strong light irradiation, thereby constructing key structural appearance features. Since the appearance features of occluding objects are unreliable, but these low-confidence groups are often the tracking objects that need to be mined, considering the existence of these vehicle appearance feature information in historical frames, compensation features are introduced, which can reduce the number of missing tracks and ID switches, and more large-camera tracking results can be obtained.

[0100] In the embodiments of the present application, when cropping images of each vehicle from the video frame, the image size is cropped to be the same as the size of each image in the training set of the domain conversion network ReID model; and the cropped images are normalized.

[0101] In the embodiments of the present application, the domain conversion network ReID model includes:

[0102] The domain conversion network, based on VTGAN image-to-image, is used to convert the input vehicle image into the same image style as the training image using GAN or a progressive domain conversion method, while retaining the appearance features of the vehicle; the domain conversion is carried out step by step in stages, first making a rough global adjustment (such as overall brightness adjustment), and then making local refinement (such as detail adjustment of key parts such as headlights and license plates).

[0103] The feature extraction network is used to extract the features of the converted vehicle image using the trained multi-scale feature fusion network to obtain the appearance features of the vehicle image; the feature extraction network uses convolutional kernels of multiple scales to capture feature information at different levels; for example, convolutional kernels of 3x3, 5x5, and 7x7 are used to extract fine-grained, medium-grained, and coarse-grained features respectively.

[0104] In the embodiments of the present application, the purpose of the domain conversion network is to make the images in the source domain (with labels) have the style of the target domain (without labels), while retaining the identity information of the source domain images. The domain conversion network based on the generative adversarial network can reduce the domain bias between different data sets and improve the adaptability of the model on new data sets; the progressive domain conversion method converts the input vehicle image into the same image style as the training image, including: making global and local adjustments to the vehicle image in stages, and combining an adaptive loss function to dynamically adjust the weights of each stage to obtain the images of each vehicle with the same image style as the training image.

[0105] In the embodiments of the present application, the feature extraction network improves the feature expression ability by cascading convolutional layers of different scales and introducing an attention mechanism to weight the importance of features at different scales; at the same time, by introducing a specific local perception area, for key parts of the vehicle (such as the front of the vehicle, the rear of the vehicle, the windows, the headlights, etc.), a local convolution module is used for feature enhancement; by defining a specific local perception area (such as the front of the vehicle, the rear of the vehicle), and combining an adaptive threshold to highlight the features of these areas.

[0106] In the embodiments of the present application, the feature extraction network further enhances the domain adaptation ability through an attention structure, especially in the case of diverse background changes; through the generated images during training, the attention mechanism is used to improve the performance of vehicle ReID. The attention mechanism can help the model focus on the key features in the image, such as the local details and structures of the vehicle, thereby improving the accuracy of cross-domain recognition.

[0107] In the embodiments of the present application, the appearance features of a vehicle include: color features: the overall color of the vehicle, the local color distribution; shape features: the vehicle contour, the body proportion; texture features: the logos, patterns, and detailed textures on the vehicle body; structural features: the license plate position, the window layout, the headlight shape, etc.

[0108] In the embodiments of the present application, when the domain transformation network ReID model is trained, training samples are increased by combining different foregrounds and backgrounds, and the domain transformation network ReID model is trained using the enriched training set, so that the domain transformation network ReID model can effectively perform domain transformation on situations such as motion blur and strong light irradiation, obtain the appearance features with the key structures of the vehicle, and improve the vehicle recognition accuracy.

[0109] In the embodiments of the present application, when the same vehicle is in different video frames collected by the same camera, vehicle detection and feature extraction are performed on each video frame, and the appearance features of the vehicle can be obtained. When it is determined that the vehicles in different video frames are the same vehicle through trajectory tracking, the average appearance feature of the vehicle will be determined based on the appearance features of the vehicle in different video frames, which can also be called the trajectory appearance feature of the vehicle.

[0110] In the embodiments of the present application, the method for determining the average appearance feature of a vehicle includes:

[0111] For any vehicle, obtain the images, appearance features, and the acquisition times of the corresponding video frames of the vehicle in different video frames of the surveillance video collected by the same camera;

[0112] Use the configured image quality assessment model to evaluate the quality of the images of the vehicle in different video frames to obtain a first score; perform occlusion detection on the images of the vehicle in different video frames, and determine a second score based on the occlusion detection and the occlusion degree; determine a third score according to the configured temporal weighting rule and the order of the acquisition times of the respective appearance features; according to the first score, the second score, and the third score, use the normalization method to calculate the weight of each appearance feature; so that the higher the vehicle image quality, the higher the weight of the newer appearance features; use a multi-scale attention mechanism to process each appearance feature, and extract global and local features respectively; perform weighted fusion on the global and local features to generate a feature vector corresponding to each appearance feature; multiply the feature vector corresponding to each appearance feature by the corresponding weight to obtain the feature representation of each appearance feature; input the sequence composed of the feature representation of each appearance feature and time into the pre-trained spatio-temporal graph convolutional network to obtain the average appearance feature of the vehicle.

[0113] Step S130: According to the vehicle detection results and corresponding appearance features of each video frame, perform trajectory tracking on each vehicle to obtain the trajectories of each vehicle in the surveillance video captured by this camera.

[0114] In the embodiment of the present application, performing trajectory tracking on each vehicle according to the vehicle detection results and corresponding appearance features of each video frame includes:

[0115] For any two adjacent video frames in the surveillance video captured by any camera, take the video frame with the earlier frame number in the adjacent video frames as the first target frame, and take the video frame with the later frame number in the adjacent video frames as the second target frame;

[0116] Take the vehicles corresponding to each first detection box in the second target frame as the first vehicles; take the vehicle corresponding to any first detection box in the first target frame as the vehicle to be matched;

[0117] For any first vehicle, if the similarity between the appearance feature of the first vehicle and the appearance feature of any vehicle to be matched is greater than the configured first similarity threshold, then take the vehicle to be matched as the target matching vehicle of the first vehicle;

[0118] According to the detection box of the target matching vehicle in the first target frame, predict the predicted detection box of the target matching vehicle in the second target frame;

[0119] If the intersection over union of the detection box of the first vehicle and the predicted detection box of the target matching vehicle is greater than the configured intersection over union threshold, then modify the vehicle ID of the first vehicle to the vehicle ID of the target matching vehicle;

[0120] For any vehicle ID, generate the trajectory of the corresponding vehicle in the surveillance video captured by the camera according to the coordinates of the vehicle corresponding to the vehicle ID in each video frame of the surveillance video.

[0121] In the embodiment of the present application, randomly select a video frame as the current video frame, and match the vehicle corresponding to any first detection box (i.e., the high-confidence group) in the current video frame with the vehicle corresponding to any first detection box in the previous video frame of the current video frame; among them, the matching is divided into two steps. The first step is the matching of appearance features, by calculating the similarity of the appearance features of the two vehicles and comparing it with the first similarity threshold to determine the similarity degree of the above two vehicles; the second step is the matching of trajectories, by predicting the possible detection box of the vehicle in the current video frame through the detection box of the vehicle in the previous video frame, and comparing the intersection over union of the possible detection box and the actual detection box in the current video frame to determine whether the two detection boxes are the detection boxes of the same vehicle.

[0122] In the embodiments of the present application, since each video frame assigns a vehicle ID to each detected bounding box during vehicle detection, when it is determined that the vehicles in different video frames are the same vehicle, it is necessary to unify their corresponding vehicle IDs to indicate that they are the same vehicle; and based on this vehicle ID, the coordinates of the bounding boxes of the corresponding vehicle in different video frames are extracted to obtain the trajectory of the vehicle in the surveillance video captured by the same camera.

[0123] In the embodiments of the present application, in order to retain more reliable vehicles and avoid the phenomenon of partial vehicle tracking loss caused by the discard of vehicles with short-term low confidence due to undetected at a certain moment caused by partial occlusion, feature compensation will be performed on the vehicles corresponding to each second bounding box in the low-confidence group of the current video frame; specifically including:

[0124] The vehicle corresponding to any second bounding box in the second target frame is used as the second vehicle.

[0125] For any second vehicle, if the similarity between the appearance features of the second vehicle and the appearance features of any vehicle to be matched is greater than the configured second similarity threshold, then the vehicle to be matched is used as the target matching vehicle of the second vehicle.

[0126] Based on the average appearance features of the target matching vehicle, feature compensation is performed on the appearance features of the second vehicle to obtain updated appearance features.

[0127] If the similarity between the updated appearance features of the second vehicle and the appearance features of any vehicle to be matched is greater than the configured first similarity threshold, then the vehicle to be matched is used as the target matching vehicle of the second vehicle.

[0128] According to the bounding box of the target matching vehicle in the first target frame, the predicted bounding box of the target matching vehicle in the second target frame is predicted.

[0129] If the intersection over union (IoU) between the bounding box of the second vehicle and the predicted bounding box of the target matching vehicle is greater than the configured IoU threshold, then the vehicle ID of the second vehicle is modified to the vehicle ID of the target matching vehicle, and the second bounding box corresponding to the second vehicle is adjusted to the first bounding box.

[0130] In the embodiments of the present application, the appearance feature similarity between the vehicle corresponding to any second bounding box in the current video frame and the vehicles corresponding to each first bounding box in the previous video frame of the current video frame is calculated. If the similarity is greater than the configured second similarity threshold, it is determined that the vehicle in the current video frame may be in the low-confidence group due to reasons such as occlusion. At this time, the average appearance features of the vehicle in the previous video frame with a similarity greater than the threshold are compensated to the corresponding vehicle in the current video frame.

[0131] In the embodiment of the present application, after the appearance features of the vehicle corresponding to any second detection box in the current video frame are compensated, the appearance features and corresponding trajectories of the vehicles corresponding to each first detection box in the previous video frame are compared again. If the conditions are met, it is determined to be the same vehicle. At this time, the second detection box corresponding to the vehicle is moved to the high-confidence group, that is, adjusted to the first detection box.

[0132] In the embodiment of the present application, similarity matching is performed on the vehicle appearance features of the vehicles in the high-confidence group. If they match, the single-camera vehicle tracking trajectory of the current vehicle is generated by tracking. If they do not match, the corresponding compensation features are retained and feature compensation is performed on the vehicles in the low-confidence group. Specifically, through the IOU pairs of the vehicle target detection boxes between adjacent video frames, the objects to be compensated are found. After finding the compensation objects, by assigning a certain weight to the compensation features, the feature compensation of the occluded vehicles is realized. Then, secondary matching is performed on the vehicles. If they match, the single-camera vehicle tracking trajectory of the current vehicle is generated. If they still do not match, the tracking threshold of the unmatched vehicles is calculated. If the tracking threshold is greater than the threshold and the vehicle appears in consecutive video frames, the vehicle is re-incorporated into the target tracking set for the next round of tracking.

[0133] In the embodiment of the present application, if the secondary matching of the vehicle still fails, the tracking threshold of the unmatched vehicle is calculated. If the tracking threshold is greater than the threshold and the vehicle appears in consecutive video frames, the vehicle is re-incorporated into the target tracking set for the next round of tracking.

[0134] In one embodiment of the present application, the tracking threshold T e is set to 0.35, the high-confidence detection threshold τ high is 0.6, the low-confidence threshold τ low is 0.15, the matching threshold is 0.5, and the cosine distance threshold is 0.5.

[0135] Step S140: Based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos collected by different cameras, determine the cross-camera tracking results of different vehicles.

[0136] In the embodiment of the present application, based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos collected by different cameras, the cross-camera tracking results of different vehicles are generated through hierarchical clustering and k mutual nearest neighbors. Reordering the single-camera vehicle tracking trajectories refers to calculating the similarity between different single-camera vehicle tracking trajectories based on the vehicle appearance features, reordering the single-camera vehicle tracking trajectories with high similarity of the same target vehicle to the front positions, and establishing a similarity matrix between adjacent cameras according to the similarity between different single-camera vehicle tracking trajectories, and using the k mutual nearest neighbor algorithm to update the similarity matrix.

[0137] In the embodiments of the present application, based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos collected by different cameras, the tracking results of different vehicles across cameras are determined, specifically including:

[0138] For any camera, the cameras on the adjacent sections of the section where the camera is located are used as the adjacent cameras of the camera; the trajectories of different vehicles in the surveillance video collected by the camera are clustered to obtain a clustering result; for any adjacent camera, the updated similarity matrix between the adjacent camera and the camera is obtained; according to the updated similarity matrix, the trajectories and clustering results of different vehicles in the surveillance video collected by the adjacent camera are clustered to obtain the tracking results of each vehicle across the camera and the adjacent camera.

[0139] In the embodiments of the present application, the surveillance area corresponding to each camera is pre-divided into several areas; according to the order in which the vehicle passes through different areas, the divided surveillance areas are determined as entry areas or exit areas; the area that the vehicle passes through first is the entry area, and the area that the vehicle passes through later is the exit area. When the vehicle is in the entry area, it indicates that the current vehicle has just entered the shooting range of the camera; when the vehicle is in the exit area, it indicates that the vehicle is about to drive out of the shooting area.

[0140] In the embodiments of the present application, time constraints and area constraints are set for each camera according to the order in which the vehicle passes through different cameras.

[0141] In the embodiments of the present application, before determining the tracking results of different vehicles across cameras based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos collected by different cameras, the method further includes:

[0142] According to the time when the suspected same vehicle passes through different cameras, combined with the time constraints corresponding to different cameras, the trajectories of different vehicles under different cameras are filtered to obtain the trajectories and corresponding appearance features of different vehicles in the surveillance videos collected by different cameras after filtering.

[0143] In practical applications, since the road segments corresponding to the cameras are different and the order in which vehicles pass through each road segment is fixed, there are differences in the order of the corresponding cameras according to the vehicle's forward direction. If the passing times of a vehicle through two adjacent cameras are contradictory to the order of the cameras, it is considered that there is a false positive in the single-camera vehicle tracking trajectory of the vehicle. For example, if the passing time of a vehicle's single-camera vehicle tracking trajectory at Camera A is 13:00 and the passing time at Camera B is 12:50, and Camera B is set in the forward direction relative to Camera A, then there is a false positive in the single-camera vehicle tracking trajectory at this time; therefore, the corresponding vehicle trajectory is deleted. By using the time relationship of entering and leaving to delete some conflicting and invalid fragmented trajectories, the number of matching vehicles can be further reduced, and time and area constraints can be imposed on the entering area and the leaving area.

[0144] In an embodiment of the present application, a method for obtaining the updated similarity matrix of any two adjacent cameras includes:

[0145] Regarding the trajectories of different vehicles under any one camera as target trajectories; regarding the trajectories of different vehicles under the adjacent camera corresponding to this camera as first trajectories; for any one target trajectory, calculating the similarity between the target trajectory and each first trajectory according to the appearance features of the vehicle corresponding to the target trajectory; sorting each first trajectory in descending order of similarity; constructing a similarity matrix of the camera and the adjacent camera according to the similarity between each target trajectory and each trajectory to be matched; using the k-complementary nearest neighbor algorithm to update the similarity matrix to obtain the updated similarity matrix.

[0146] In an embodiment of the present application, the similarity calculation formula is as follows:

[0147]

[0148] where Sim(T i , T j ) is the similarity between the i-th single-camera vehicle tracking trajectory and the j-th single-camera vehicle tracking trajectory, T i refers to the i-th single-camera vehicle tracking trajectory, T j refers to the j-th single-camera vehicle tracking trajectory, f i refers to the appearance features of the vehicle corresponding to T i , A f refers to the average appearance features of the corresponding vehicle, and f j refers to the appearance features of the vehicle corresponding to T j .

[0149] In an embodiment of the present application, the similarity matrix calculation formula is as follows:

[0150]

[0151] Among them, S m represents the similarity matrix of camera N and camera N + 1; Sim(T1 N , T1 N+1 ) represents the similarity value between the first trajectory under the Nth camera and the first trajectory under the (N + 1)th camera; Sim(T n N , T m N+1 ) represents the similarity value between the nth trajectory under the Nth camera and the mth trajectory under the (N + 1)th camera.

[0152] In the embodiment of the present application, the k-complementary nearest neighbor algorithm is used to update the similarity matrix, including:

[0153] Select the first z items of the sorted first trajectory to obtain the second trajectory; where z is a pre-configured positive integer; calculate the complementary similarity between each target trajectory and each second trajectory; update the similarity matrix according to the calculated complementary similarity.

[0154] In practical applications, the similarity matrix obtained according to the similarity between each target trajectory and each first trajectory can only keep the target vehicle and the similar vehicle at a certain distance from each other, but cannot guarantee the distance between different vehicles. Therefore, the present application further uses the k-complementary nearest neighbor algorithm to refine the update of the similarity matrix, which can ensure that different vehicle trajectories with similar similarities are kept at a certain distance from each other, more accurate than the traditional k-nearest neighbor, and reduce the incorporation of some false positive vehicle noises.

[0155] In the embodiment of the present application, hierarchical clustering is used to cluster different vehicle trajectories after region filtering, that is, first cluster the trajectories under the same camera, and then cluster the trajectories under different cameras, so as to obtain cross-camera vehicle tracking trajectories; using hierarchical clustering can combine the trackers between adjacent cameras, and merge the remaining small trajectories by selecting a pair with the smallest pairwise distance, so as to assign the same global track ID between cameras and generate a completely matching small track result between cameras.

[0156] In the embodiment of the present application, according to the appearance features of each vehicle and the similarity of the trajectories, the trajectories of different vehicles under the same camera are clustered, aiming to exclude the situation where the same vehicle is divided into multiple trajectories, and merge the trajectories that were previously wrongly divided into two or more trajectories but actually belong to the same vehicle into a complete trajectory, so as to ensure that different vehicles in the finally obtained clustering result correspond to different trajectories.

[0157] In the embodiment of the present application, if multiple trajectories in the clustering result are clustered into one class, then spatio-temporal consistency detection is performed on the multiple trajectories according to the time and position corresponding to the trajectories (determined according to the coordinates of each detection box). If the spatio-temporal consistency detection is passed, the trajectories clustered into the same class are used as the trajectories of the same vehicle and assigned the same vehicle ID; otherwise, the trajectories clustered into the same class are used as the trajectories of different vehicles, and their corresponding vehicle IDs are retained.

[0158] In the embodiment of the present application, performing spatio-temporal consistency detection on multiple trajectories includes:

[0159] Regarding the trajectories clustered into one class as the trajectories to be detected; if the coordinates and corresponding time points constituting each trajectory to be detected satisfy the configured conflict rules, it is determined that there is a conflict in the trajectories to be detected, that is, the spatio-temporal consistency detection is not passed.

[0160] In the embodiment of the present application, the conflict rules include time conflict rules, space conflict rules, and speed conflict rules; specifically, the time conflict rules may include: there are multiple trajectories to be detected with consistent coordinates at multiple different consecutive time points; the space conflict rules may include that there are multiple trajectories to be detected with inconsistent coordinates at multiple same time points.

[0161] In the embodiment of the present application, all the trajectories that pass the spatio-temporal consistency detection are used as the trajectories of different vehicles. Regarding the trajectory of any vehicle under this camera as the first node and the trajectory of any vehicle under the adjacent camera as the second node; adding an edge between any first node and any second node according to the updated similarity matrix to obtain a trajectory weighted graph; where the weight of the edge is the similarity between the corresponding first node and the second node; using a clustering algorithm to cluster the trajectory weighted graph to obtain multiple clustering clusters; where each clustering cluster represents the trajectories of a vehicle under different cameras.

[0162] In an embodiment of the present application, as Figures 2 - 3 shown, the method for tracking multiple vehicles across cameras on a highway provided by the embodiment of the present application includes:

[0163] Inputting the video frames of the camera into a vehicle target detection model optimized by a gated recurrent convolutional structure, and using the special structure of the gated recurrent convolution to complete the prediction of the classification network and the bounding box prediction network for the features extracted by the information interaction of the high-order space only using convolutional and fully connected layers, so as to obtain accurate vehicle target detection boxes, improving the accuracy of vehicle target detection in blurred images;

[0164] The ReID feature extraction model based on the domain transformation network and the attention mechanism, and the VTGAN image-to-image translation network. By combining different foregrounds and backgrounds to increase training samples, the trained domain converter can effectively perform domain transformation for situations such as motion blur and strong light irradiation, thereby constructing key structural appearance features.

[0165] Input the vehicle target detection box and the vehicle appearance features into the multi-object tracking module to obtain the single-camera vehicle tracking trajectories of multiple vehicles in each camera; perform region filtering, reordering, and hierarchical clustering on the single-camera vehicle tracking trajectories of multiple cameras to obtain the cross-camera vehicle tracking trajectories under multiple cameras.

[0166] Corresponding to the above method, an embodiment of the present application further provides a tracking device for multiple vehicles across cameras on a highway. Cameras are respectively arranged on each section of the highway, and each camera corresponds to a different monitoring area. For example, Figure 4 As shown, the tracking device for multiple vehicles across cameras on the highway includes:

[0167] An acquisition unit 410, configured to acquire the monitoring videos collected by each camera;

[0168] A detection unit 420, for any camera, input any video frame in the monitoring video collected by the camera into a pre-trained lightweight vehicle target detection model to obtain the vehicle detection results of each vehicle in the video frame; wherein, the vehicle detection results include: detection boxes of different vehicles and corresponding coordinates;

[0169] A cropping unit 430, configured to crop the images of each vehicle from the video frame according to the coordinates of each detection box;

[0170] An extraction unit 440, configured to extract the appearance features in the images of each vehicle to obtain the appearance features of each vehicle;

[0171] A tracking unit 450, configured to perform trajectory tracking on each vehicle according to the vehicle detection results and corresponding appearance features of each video frame to obtain the trajectories of each vehicle in the monitoring video collected by the camera;

[0172] A determination unit 460, configured to determine the cross-camera tracking results of different vehicles based on the trajectories and corresponding appearance features of different vehicles in the monitoring videos collected by different cameras.

[0173] The functions of the functional units of the tracking device for multiple vehicles across cameras on the highway provided in the above embodiments of the present application can be implemented by the above method steps. Therefore, the specific working processes and beneficial effects of each unit in the tracking device for multiple vehicles across cameras on the highway provided in the embodiments of the present application will not be repeated here.

[0174] An embodiment of the present application also provides an electronic device, such as Figure 5 shown, which includes a processor 510, a communication interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communication interface 520, and the memory 530 communicate with each other through the communication bus 540.

[0175] The memory 530 is used to store computer programs;

[0176] When the processor 510 is used to execute the program stored on the memory 530, the following steps are implemented:

[0177] Obtain the surveillance videos collected by each camera;

[0178] For any camera, input any video frame in the surveillance video collected by the camera into a pre-trained lightweight vehicle target detection model to obtain the vehicle detection results of each vehicle in the video frame; wherein, the vehicle detection results include: detection frames of different vehicles and corresponding coordinates;

[0179] Crop the images of each vehicle from the video frame according to the coordinates of each detection frame;

[0180] Extract the appearance features in the images of each vehicle to obtain the appearance features of each vehicle;

[0181] Track the trajectories of each vehicle according to the vehicle detection results and corresponding appearance features of each video frame to obtain the trajectories of each vehicle in the surveillance video collected by the camera;

[0182] Based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos collected by different cameras, determine the cross-camera tracking results of different vehicles.

[0183] The aforementioned communication bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0184] The communication interface is used for communication between the aforementioned electronic device and other devices.

[0185] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0186] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0187] Since the implementation manners and beneficial effects of the various devices of the electronic device in the above embodiments can be realized by referring to the steps in the Figure 1 embodiments shown, therefore, the specific working process and beneficial effects of the electronic device provided by the embodiments of the present application will not be elaborated herein.

[0188] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the multi-vehicle tracking method across cameras on a highway described in any one of the above embodiments.

[0189] In another embodiment provided by the present application, there is also provided a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the multi-vehicle tracking method across cameras on a highway described in any one of the above embodiments.

[0190] Those skilled in the art should understand that the embodiments in the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the embodiments in the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments in the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0191] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0192] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0193] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operating steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.

[0194] Although the preferred embodiments in the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0195] Obviously, those skilled in the art can make various changes and variations to the embodiments in the embodiments of the present application without departing from the spirit and scope of the embodiments in the embodiments of the present application. Thus, if these modifications and variations of the embodiments in the embodiments of the present application fall within the scope of the claims of the embodiments of the present application and their equivalent technologies, the embodiments in the embodiments of the present application are also intended to include these changes and variations.

Claims

1. A tracking method for multiple vehicles across cameras on a highway, characterized in that, Each section of the highway is equipped with cameras, and each camera corresponds to a different monitoring area. The method includes: Obtain the monitoring videos collected by each camera; For any camera, input any video frame in the monitoring video collected by the camera into a pre-trained lightweight vehicle target detection model to obtain the vehicle detection results of each vehicle in the video frame; wherein, the vehicle detection results include: detection frames of different vehicles and corresponding coordinates; Crop the images of each vehicle from the video frame according to the coordinates of each detection frame; Extract the appearance features in the images of each vehicle to obtain the appearance features of each vehicle; Track the trajectories of each vehicle according to the vehicle detection results and corresponding appearance features of each video frame to obtain the trajectories of each vehicle in the monitoring video collected by the camera; Based on the trajectories and corresponding appearance features of different vehicles in the monitoring videos collected by different cameras, determine the cross-camera tracking results of different vehicles.

2. The method according to claim 1, wherein The lightweight vehicle target detection model includes: An input layer for receiving the video frame; A gated convolutional layer for using the video frame as the input feature of the gated convolutional layer, performing linear projection on the input feature to obtain a first feature and a second feature; performing depth convolution on the second feature to obtain a third feature; performing a first-order interaction on the third feature and the first feature to obtain a first-order spatial feature; using the first-order spatial feature as the new input feature, and returning to perform linear projection on the input feature until the configured termination condition is met, and outputting a high-order spatial feature; A classification layer for classifying according to the high-order spatial feature to determine whether there are vehicles at each position of the video frame; A prediction layer for using the high-order spatial feature for regression to generate the vehicle detection results of each vehicle existing in the video frame.

3. The method according to claim 1, wherein Extracting the appearance features in the images of each vehicle to obtain the appearance features of each vehicle includes: Inputting the images of each vehicle into a pre-trained domain transformation network ReID model to obtain the appearance features of each vehicle.

4. The method according to claim 1, wherein The vehicle detection results further include: the confidence levels of different detection frames and the vehicle IDs of different vehicles; The method further includes: If the confidence level of any detection frame is greater than the configured confidence threshold, then use the detection frame as the first detection frame; otherwise, use the detection frame as the second detection frame; For any vehicle ID, generate the average appearance feature of the vehicle according to the appearance features of the vehicle corresponding to the vehicle ID in each video frame of the monitoring video.

5. The method according to claim 4, wherein Tracking the trajectories of each vehicle according to the vehicle detection results and corresponding appearance features of each video frame includes: For any two adjacent video frames in the monitoring video collected by any camera, use the video frame with the earlier frame number in the adjacent video frames as the first target frame, and use the video frame with the later frame number in the adjacent video frames as the second target frame; Use the vehicles corresponding to each first detection frame in the second target frame as the first vehicles; use the vehicle corresponding to any first detection frame in the first target frame as the vehicle to be matched; For any first vehicle, if the similarity between the appearance features of the first vehicle and those of any vehicle to be matched is greater than a configured first similarity threshold, then the vehicle to be matched is taken as the target matching vehicle of the first vehicle; Based on the detection frame of the target matching vehicle in the first target frame, predict the predicted detection frame of the target matching vehicle in the second target frame; If the intersection over union (IoU) between the detection frame of the first vehicle and the predicted detection frame of the target matching vehicle is greater than a configured IoU threshold, then modify the vehicle ID of the first vehicle to the vehicle ID of the target matching vehicle; For any vehicle ID, generate the trajectory of the corresponding vehicle in the surveillance video captured by the camera according to the coordinates of the vehicle corresponding to the vehicle ID in each video frame of the surveillance video; 6. The method according to claim 5, wherein The method further includes: Take the vehicle corresponding to any second detection frame in the second target frame as the second vehicle; For any second vehicle, if the similarity between the appearance features of the second vehicle and those of any vehicle to be matched is greater than a configured second similarity threshold, then the vehicle to be matched is taken as the target matching vehicle of the second vehicle; Based on the average appearance features of the target matching vehicle, perform feature compensation on the appearance features of the second vehicle to obtain updated appearance features; If the similarity between the updated appearance features of the second vehicle and those of any vehicle to be matched is greater than a configured first similarity threshold, then the vehicle to be matched is taken as the target matching vehicle of the second vehicle; Based on the detection frame of the target matching vehicle in the first target frame, predict the predicted detection frame of the target matching vehicle in the second target frame; If the intersection over union (IoU) between the detection frame of the second vehicle and the predicted detection frame of the target matching vehicle is greater than a configured IoU threshold, then modify the vehicle ID of the second vehicle to the vehicle ID of the target matching vehicle, and adjust the second detection frame corresponding to the second vehicle to the first detection frame; 7. The method according to claim 5, wherein Based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos captured by different cameras, determine the cross-camera tracking results of different vehicles, including: For any camera, take the cameras located on the adjacent road sections of the road section where the camera is located as the adjacent cameras of the camera; Cluster the trajectories of different vehicles in the surveillance video captured by the camera to obtain a clustering result; For any adjacent camera, obtain the updated similarity matrix of the adjacent camera and the camera; According to the updated similarity matrix, cluster the trajectories of different vehicles in the surveillance video captured by the adjacent camera and the clustering result to obtain the cross-camera tracking results of each vehicle under the camera and the adjacent camera; 8. A tracking device for multiple vehicles across cameras on a highway, characterized in that, Cameras are respectively arranged on each road section of the highway, and each camera corresponds to a different monitoring area. The device includes: An acquisition unit, configured to acquire the surveillance videos captured by each camera; The detection unit is configured to input any video frame in the surveillance video collected by any camera into a pre-trained lightweight vehicle target detection model for each camera, and obtain vehicle detection results of each vehicle in the video frame; wherein, the vehicle detection results include: detection frames of different vehicles and corresponding coordinates; The cropping unit is configured to crop images of each vehicle from the video frame according to the coordinates of each detection frame; The extraction unit is configured to extract appearance features in the images of each vehicle to obtain appearance features of each vehicle; The tracking unit is configured to perform trajectory tracking on each vehicle according to the vehicle detection results of each video frame and corresponding appearance features, and obtain trajectories of each vehicle in the surveillance video collected by the camera; The determination unit is configured to determine cross-camera tracking results of different vehicles based on the trajectories and corresponding appearance features of different vehicles in the surveillance videos collected by different cameras.

9. An electronic device, characterized in that, The electronic device includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is configured to implement the method according to any one of claims 1-7 when executing the program stored on the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Expressway cross-camera-shooting multi-vehicle tracking method and system combined with multiple models

    CN117218580A

Cited By

  • Cross-camera vehicle tracking and trajectory matching method

    CN120672803A

  • Vehicle tracking data processing method and device, storage medium and program product

    CN120708167A