Multi-Twin Adversarial Network Cross-Camera Vehicle Tracking Method with Coupled Vehicle-Following Enhancement
Through the multi-twin adversarial network method enhanced by coupled fleet follow-up, the stability and accuracy of cross-camera vehicle tracking are solved, vehicle re-identification in a multi-camera environment is realized, and the accuracy and range of traffic information collection are improved.
Patent Information
- Application Number
- CN202210027047.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-01-11
AI Technical Summary
The prior art is difficult to achieve vehicle tracking across cameras, especially in multi-camera environments, and the stability and accuracy of vehicle re-identification are insufficient, due to the influence of environmental conditions such as light and viewing angle.
A multi-twin adversarial network method with coupled fleet following enhancement is adopted to detect vehicle characteristics through neural networks, build a dynamic fleet model, integrate vehicle morphological characteristics and follow-up characteristics, and realize vehicle tracking across cameras.
The stability and accuracy of cross-camera vehicle tracking are improved, and the re-identification effect is enhanced by integrating the following characteristics and dynamic fleet models in the vehicle trajectory, and important means of collecting traffic information is provided.
Smart Images

Figure CN114463390B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-twin adversarial network cross-camera vehicle tracking method for coupling platoon following enhancement, belonging to the technical fields of traffic flow and intelligent transportation. Background Art
[0002] Vehicle trajectories contain rich driving spatio-temporal information and are important channels for obtaining refined macroscopic and microscopic spatio-temporal traffic flow information. Microscopic traffic flow parameters such as transient speed, acceleration, headway, and time headway can be extracted from typical vehicle trajectory spatio-temporal diagrams, and macroscopic traffic flow parameters such as aggregated flow, density, and speed can also be extracted. Typical traffic flow phenomena such as congestion generation, moving wave propagation, and traffic breakdown can be visually identified in vehicle trajectory spatio-temporal diagrams. Traditional methods such as loop collection and radar detection are all single-point collections and cannot obtain continuous trajectories. The monitoring video has advantages such as high clarity, good continuity, and wide coverage, and extracting vehicle trajectories from monitoring videos has gradually become a hot issue.
[0003] A single monitoring camera is limited by its narrow capture range, while the spatio-temporal domain of traffic information and its typical features far exceeds the coverage range of a single-point camera, such as moving waves. Connecting vehicle trajectories of multiple cameras can expand the spatio-temporal range of traffic information, and abnormal event analysis based on traffic flow characteristics also requires the acquisition of full spatio-temporal information of macroscopic and microscopic road traffic flow.
[0004] Emerging machine vision technology provides a convenient and accurate channel for extracting vehicle trajectories from videos. Deep convolutional neural networks comprehensively learn multi-scale features, greatly improving the accuracy and speed compared with recognition based on single prior features. Moreover, machine vision technology provides a technical means for joint recognition of the same target between different videos, but vehicle tracking and re-identification only using vehicle morphological features are limited by many problems such as picture illumination and viewing conditions. Identifying and determining vehicles across cameras and with diverse viewing angles under rich environmental conditions is a novel and challenging task. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a multi-twin adversarial network cross-camera vehicle tracking method for coupling platoon following enhancement, obtain morphological features through vehicle feature enhancement, dynamically predict vehicle time with dynamic platoon construction and trajectory following characteristics, fuse vehicle morphological features and following characteristics to achieve dynamic platoon group matching, and finally implement an individual correspondence method for the vehicle group, so as to form cross-camera vehicle tracking.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] A multi-twin adversarial network cross-camera vehicle tracking method for coupling platoon following enhancement, comprising the following steps:
[0008] S10. Build a basic network for obtaining target feature information, prepare a dataset, and train the basic network. The basic network includes a neural network for target detection and an adversarial network for image normalization. The dataset includes a target detection dataset and an image normalization dataset. The target detection includes obtaining the vehicle position and vehicle size. The image normalization includes pose transformation and background removal.
[0009] S20. Multiply the basic network in a multiple twin manner for multi-threaded synchronous processing of multiple surveillance videos.
[0010] S30. Build a real-time trajectory extraction model in the multiple twin basic network, use the time-position relationship of vehicle operation to track the same target between different frames and assign a target serial number, intercept the trajectory images in the video, take the trajectory as the target unit, and extract and enhance the vehicle morphology features through the adversarial network.
[0011] S40. Analyze the vehicle motion in real time according to the vehicle trajectory, divide the vehicle following state, extract the time-series features of the following behavior, classify the target vehicle into following state vehicles and free state vehicles, build a dynamic platoon model, and record the following features and vehicle morphology distribution features of the platoon to which the target vehicle belongs.
[0012] S50. Establish an iterative formula for predicting the travel time by integrating the dynamic platoon motion features, the overall traffic flow state, and the spatio-temporal positions of multiple surveillance distributions, and predict the spatio-temporal positions where the platoon appears across cameras.
[0013] S60. Integrate the dynamic platoon following characteristics, the predicted spatio-temporal positions where the platoon appears across cameras, and the obtained vehicle morphology distribution features within the platoon to achieve platoon group cross-camera matching under multiple surveillance.
[0014] S70. Based on the results of group cross-camera matching, sequentially track and number the corresponding target morphology feature distributions of individuals within the dynamic platoon to complete vehicle cross-camera tracking.
[0015] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:
[0016] The present invention proposes a multi-twin adversarial network cross-camera vehicle recognition method that couples and strengthens vehicle following characteristics. It uses a neural network to detect vehicle targets in surveillance cameras, and through a twin adversarial network, it eliminates environmental interference for the detected vehicle targets and strengthens the main features. The above network is multiplied and twin for synchronous processing of multi-threaded multi-surveillance videos. The network extracts vehicle trajectories to obtain vehicle queue following motion characteristics and the morphological characteristics of each vehicle. According to the obtained motion characteristics, it divides the vehicle following states and constructs a dynamic fleet feature extraction model. By integrating vehicle motion, traffic state propagation, and camera spatial distribution, it establishes an iterative vehicle spatio-temporal prediction. By integrating the dynamic fleet feature model and the obtained in-team vehicle morphological distribution characteristics, it realizes cross-camera matching of fleet groups under multi-surveillance. Based on the group matching results, it sequentially tracks and numbers the corresponding target morphological feature distributions of individuals within the dynamic fleet to complete cross-camera vehicle tracking. Compared with the previous cross-camera same target determination, the present invention integrates the following characteristics in vehicle trajectories and combines the method of dynamic fleets to add stable spatio-temporal constraints for re-identification, greatly improving the re-identification effect, which is of great significance for cross-camera target tracking of vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic flowchart of the multi-twin adversarial network cross-camera vehicle tracking method that couples and strengthens fleet following of the present invention;
[0018] Figure 2 is a schematic diagram of the dynamic fleet group matching process;
[0019] Figure 3 is a schematic diagram of the neural network vehicle similarity determination result;
[0020] Figure 4 is a schematic diagram of the vehicle feature strengthening result of the GFLA algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] The following details the embodiments of the present invention, and the examples of the embodiments are shown in the drawings. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.
[0022] As Figure 1 shown, it is a schematic flowchart of the multi-twin adversarial network cross-camera vehicle tracking method that couples and strengthens fleet following proposed by the present invention, including the following steps:
[0023] S10: Build a basic network for obtaining target feature information, make a data set, and train the basic network; the basic network includes a neural network for target detection and an adversarial network for image standardization; the data set includes a target detection data set and an image standardization data set; the target detection is to obtain the vehicle position and vehicle size; the image standardization includes pose transformation and background removal;
[0024] S20: Multiply and twin the above basic network for multi-threaded and multi-monitored video synchronization processing;
[0025] S30: Build a real-time trajectory extraction model in the multiply-twinned network, quickly track the same target between different frames using the time-position relationship of vehicle operation, assign a target serial number, intercept the trajectory image in the video, take the trajectory as the target unit, and extract and strengthen the vehicle shape features through the adversarial network model;
[0026] S40: Divide the vehicle following state according to the vehicle movement parsed in real time from the vehicle trajectory, extract the sequential features of the following behavior, classify the target vehicle into a following state vehicle and a free state vehicle, build a dynamic platoon model, and record the following characteristics and vehicle shape distribution characteristics of the platoon to which the target vehicle belongs;
[0027] S50: Establish an iterative formula for predicting the travel time by integrating the dynamic platoon movement characteristics, the overall traffic flow state, and the spatio-temporal positions of multiple monitors, and predict the spatio-temporal positions where the platoon appears across cameras;
[0028] S60: Integrate the dynamic platoon following characteristics, the predicted spatio-temporal positions across cameras, and the obtained in-platoon vehicle shape distribution characteristics to achieve platoon group matching across multiple monitors;
[0029] S70: Based on the group matching results, sequentially track and number the corresponding target shape feature distributions of individuals within the dynamic platoon to complete vehicle tracking across cameras.
[0030] The basic network for obtaining target feature information in S10 specifically includes:
[0031] S11: Include a neural network for target detection; the target detection neural network is trained and detected using the YOLOv5 algorithm under the pytorch framework; during the training process, the images in the training set are scaled to a unified size and then fed into the YOLOv5 algorithm in batches for logistic regression prediction. The training effect of the YOLOv5 model is evaluated by the loss value, and the training effect loss (loss) after one iteration is expressed as follows:
[0032] loss = loss xy + loss lw + loss confidence + loss class
[0033] Among them, loss xy represents the error of the center point of the detection box, loss lw is the error of the length and width of the detection box, loss confidence characterizes the confidence error of the detection box, loss classIndicates the error in the classification of the detection box; when the loss value converges and no longer changes, the YOLOv5 model can be put into use;
[0034] S12: Includes an adversarial network for image normalization; the adversarial network uses the GFLA algorithm, based on the publicly available vehicle dataset VeRi-776, and artificially constructs a training target pair for the feature enhancement process of the adversarial network; the pair of training target pairs includes two images, which are of the same vehicle target. The first image is a picture of the target from any perspective, and the second image is a picture of the target from a specific perspective with the background blurred; multiple sets of training target pairs are made for the same target, combining multiple perspectives with the specific perspective; several vehicle targets in the training set are selected and all are processed as above to form multiple sets of training target pairs with multiple perspectives corresponding to the specific perspective for several vehicle targets. This training target pair will be used as the training dataset for the feature enhancement twin adversarial network;
[0035] S13: The object detection algorithm is applied to the video to obtain the target position x, y, the target pixel size l, w, and the target confidence conf, and the target image is cropped. The image normalization adversarial network is used for the cropped target image to perform image perspective transformation and background deletion.
[0036] S20 Multiple-twins the above basic network, reads in multiple video streams captured by different surveillance cameras, obtains the video frames at the current moment, creates corresponding image libraries for different video streams, allocates separate threads and corresponding networks for each video, and parallelly implements object detection, pose transformation, and background removal of the basic network.
[0037] S30 Builds a real-time trajectory extraction model in the multiple-twin network, specifically including:
[0038] S31: Retains the vehicle target positions in the previous frame of the video, and the set is denoted as P0. Obtains all the vehicle target positions in the current frame, denoted as the set P1. The target positions are obtained from the detection results of the trained YOLOv5 algorithm. The current vehicle target position p(x, y, l, w) to be determined for duplication in the set P1;
[0039] S32: Calculates the intersection over union of the rectangle elements in the set P1 and the set P0. The calculation of the intersection over union IoU is expressed as follows:
[0040] IoU = A inter / A union
[0041] where, A inter represents the area of the intersection of the rectangles in the two sets, and A union represents the total area of the union of the rectangles in the two sets.
[0042] S33: For the targets with an IoU of 0 in the set P1, they are considered as newly added targets in the frame and are individually assigned a newly added vehicle number. The remaining targets execute S34.
[0043] S34: All targets whose IoU is not 0 and the elements in the set P0 are sent to the Hungarian algorithm for multi-target tracking. The Hungarian algorithm uses the Mahalanobis distance as the distance difference between targets. The Mahalanobis distance is calculated as follows:
[0044]
[0045] Among them, D M For vector For the mean value of its distribution The Mahalanobis distance of The covariance matrix of the distribution. During the calculation process, Move pixel vector for vehicle
[0046] S35: For targets whose IoU is not 0, after successful tracking, the targets in set P1 are assigned the corresponding vehicle serial numbers, and unsuccessful matching targets are assigned the newly added vehicle serial numbers.
[0047] S36: Capture the successfully tracked target image and use the GFLA adversarial network to enhance the morphological features; according to the feature enhancement weights trained in S10, call the local attention mechanism in the network to traverse the image pixels to identify the main features.
[0048] S37: Determine whether the feature is background. If yes, blur the sampling area using Gaussian blur. The calculation is as follows:
[0049]
[0050] Among them, G(r) is the Gaussian blur function, which performs blur calculation on the pixel value, and the output value G(r) is the blurred pixel value. r is the blur radius, which is half of the sampling area size, and σ is the standard deviation of the sampling area.
[0051] If the sampling area is judged not to be the background, the original pixel shape is retained and the image is output.
[0052] S38: Add time series to fuse the feature vectors of each frame of the same target image. Use 3×3 and 2×2 small convolution kernels to convert the S37 output image into a corresponding one-dimensional feature vector; multiple feature vectors with the same target label will be merged in sequence according to the following formula:
[0053]
[0054]
[0055] Among them: F is the merged feature vector, F i is each feature vector to be merged, conf i is the confidence corresponding to the i-th feature vector to be merged, which is output simultaneously by the YOLOv5 model for this object detection. The merged confidence corresponding to this merged feature vector is CONF.
[0056] S39: Invoke the GFLA generator according to the feature vector F to generate the standard pose image trained in S10, and complete feature enhancement.
[0057] S40 Parse the vehicle motion in real time according to the vehicle trajectory, specifically including:
[0058] S41: Identify the current road following situation, and judge whether the vehicles on the current road are driving in a following state according to the road flow rate and the average vehicle speed. For vehicles not driving in a following state, directly use the vehicle speed to directly predict the spatio-temporal position of the vehicle. If it is in a following state, continue with S42.
[0059] S42: For vehicles driving in the same lane in a following state, they will be paired one by one according to the following vehicle and the vehicle being followed. A group of following vehicles and the vehicle being followed includes the vehicle walking in front and the vehicle immediately behind it.
[0060] S43: Create a corresponding dynamic vehicle fleet for each monitoring video. The length of the dynamic vehicle fleet includes all vehicles in the total research section, and the vehicle is counted when it drives out from under the current camera; The features included in the dynamic vehicle queue are: the relative following order i of the vehicle, the vehicle shape enhancement feature vector F calculated in S38, the headway sequence SpaceHeadway(t), the time headway sequence TimeHeadway(t), and the following speed difference sequence Δv(t). The headway, time headway, and following speed difference are calculated as follows:
[0061]
[0062] TimeHeadway(t) = (x i-1 (t) - x i (t)) * fps
[0063] Δv(t) = v i-1 (t) - v i (t)
[0064] where t represents the time series, fps represents the frame rate of the current video, x i is the position of vehicle i, and v i is the speed of vehicle i.
[0065] In a dynamic vehicle following queue, the following vehicle is also the followed vehicle. Therefore, the above parameters can be calculated for each vehicle passing through the current section.
[0066] S50 specifically includes:
[0067] S51: Calculate the average speed of vehicles in each monitoring video, refer to the mileage points of the road where the monitoring is located, and fit the average speed distribution curve of the entire research section.
[0068] S52: Take the reciprocal of the vehicle average speed curve to obtain the relationship curve T between the time taken for vehicles to pass through a unit section and the mileage distribution of the entire road. du -X, denoted as
[0069] S53: Due to the driver's reaction time, the traffic wave will transfer upstream over time, and the downstream state will follow the traffic wave upstream. Predict this transfer, that is, calculate the translation speed V of du -X with respect to the time series t. wave (t).
[0070] ΔX = t × V wave (t)
[0071] where ΔX is the dynamic translation mileage of the traffic state and t is the elapsed time. Therefore, we have:
[0072]
[0073] S54: Form an iterative formula for calculating the elapsed time of a typical target vehicle in the research section in a dynamic vehicle fleet:
[0074] X(t) k+1 = X(t) k + v i (t) × Δt
[0075]
[0076] v i (t) is the speed value of vehicle i at time t, which is calculated in this model by calculation.
[0077] Iterate the two formulas until X reaches the spatial position to be predicted or t reaches the time to be predicted, then stop the iteration and calculate the total travel time AT of the typical vehicle i du,i :
[0078]
[0079] where k is the current iteration number; N is the total number of iterations, taking the number of times covering the matching gap area as N, Δt is the iteration step time interval, and ATdu,i Denote the total driving time experienced by the \(i\)-th vehicle on the research section.
[0080] S55: Select typical vehicles in the vehicle platoon at equal intervals according to the vehicle platoon following order \(i\), denote the interval number as \(K\), and calculate the iterative result \(AT\). du,i Estimate the driving time of these \(K\) vehicles passing through the matching gap mileage.
[0081] S60 to achieve cross-camera matching of vehicle platoon groups under multiple monitors specifically includes:
[0082] S61: Select a typical vehicle in the vehicle platoon as the representative of the running spatio-temporal characteristics, and the appearance time of this vehicle is \(DT\). du , select \(K\) targets before and after it to form a vehicle platoon group, with a total of \(2K + 1\) targets;
[0083] S62: Use the spatio-temporal information \(DT\) predicted by the typical vehicle. du + \(AT\) du as the center, in the dynamic queue corresponding to the monitoring camera of the downstream targets to be re-identified, select the \(2K + 1\) vehicles closest to the time center. These \(2K + 1\) vehicles form a group of vehicles to be matched centered on \(DT\). du + \(AT\) du as the center;
[0084] S63: From the two groups of vehicles selected from the spatio-temporal information in S61 and S62, select adjacent \(M\) targets in order, and check the similarity of the vehicle queues in the two groups of targets according to S64. The number of selected targets \(M\) decreases from \(2K + 1\) to \(K + 1\). The two groups of target selections include the first and the last ends, and the overlapping part decreases as \(M\) decreases;
[0085] S64: Specifically, vehicle queue re-identification is realized through a siamese convolutional neural network. The input features of the re-identification siamese neural network are the morphological enhancement feature vector \(F\) corresponding to the vehicle group, the headway sequence, the time headway sequence, and the following speed difference sequence. The output result is the matching result of the two vehicle queues, indicating the degree of complete matching, and the value range is \([0, 1]\).
[0086] S60 to achieve cross-camera matching of vehicle platoon groups under multiple monitors. After realizing vehicle group matching, select the matching group with the highest vehicle group matching degree value as the truly corresponding group for output, and correspond to the remaining vehicle targets in this batch in order, and complete this \(K\) vehicle target group in the dynamic vehicle platoon.
[0087] Example:
[0088] The method for cross-camera vehicle recognition using a multiple twin adversarial network with enhanced coupled vehicle following characteristics in this embodiment has certain requirements for the monitoring camera system. The monitoring camera should meet the resolution of not less than 720×480, have a fixed monitoring angle, a fixed monitoring focal length, and operate in a bright road environment to ensure the clarity of the vehicle shape captured. The computer connected to the monitoring camera should have the condition for running a multi-threaded program to ensure that the algorithm program can extract the target features of the vehicle. In this embodiment, two videos simultaneously captured by two adjacent road monitoring cameras are selected.
[0089] The specific information of the target videos used in the experiment is shown in Table 1:
[0090] Table 1
[0091] Video information Upstream video Downstream video Resolution 1080×720 1080×720 Frame rate 24fps 24fps Duration 3min 3min Section length 30m 60m Shooting angle Rear right Rear right Number of targets passed 152 156
[0092] This embodiment specifically includes the following 3 steps:
[0093] Step 1: Training of the twin adversarial network for vehicle re-identification, which is specifically divided into the following three steps:
[0094] Generation of object detection weights based on deep learning:
[0095] A deep learning training set for vehicle target position recognition under the monitoring perspective is made, and the YOLOv5 algorithm based on CNN is adopted. The training set needs to annotate pictures and target position coordinates. The original model weights of YOLOv5 are detection weights for 20 classes of targets trained based on the coco dataset. In the present invention, in order to better capture vehicle features and make the model adapt to various lighting condition backgrounds, vehicle targets in the coco dataset and the KITTI dataset are mixed to form an enhanced dataset.
[0096] In order to make the enhanced dataset have better robustness, 30% of the dataset content is randomly selected, and the pictures and corresponding labels are changed. The change methods include: image rotation, light enhancement, light dimming, and color adjustment.
[0097] The implementation method of image rotation is as follows: The training image and the rectangular position of the label will be rotated by a random angle around the center of the image. This angle is generated by a random number, and the corner parts are filled with black in the picture. After rotation, a new rectangle formed by passing through the four vertices of the new label rectangle is obtained to ensure that the border of the circumscribed rectangle is parallel to the image edge. To ensure the accurate coincidence of the target circumscribed rectangle and considering that the picture tilt is not too large in practice, the rotation angle here does not exceed 45°.
[0098] The implementation method for light enhancement or dimming is as follows: Read the training images, call the skimage module package in Python for image brightness adjustment, and select the brightness adjustment method based on the gamma value. The call command is: skimage.exposure.adjust_gamma(image, gamma=1), where image is the image to be adjusted and the gamma value is the parameter determining the brightness. When the gamma value is greater than 1, the image brightness increases, and conversely, the image brightness decreases. To retain the target features in the image and prevent highlight overflow or dark part loss, the gamma value is manually limited within the range of [0.7, 1.3].
[0099] The implementation method for color adjustment is as follows: Read the RGB channels of the image color, and superimpose a monochromatic filter on the RGB values of all pixels in the entire image, that is, add or subtract a fixed matrix of the same size to all values of a certain channel.
[0100] Table 2 is the ratio table for dataset mixing and processing:
[0101] Table 2
[0102]
[0103]
[0104] During the training process, two-stage training is adopted. The first stage of training goes through a total of 500 cycles and uses a larger learning rate to make the training model converge faster; the second stage of training goes through a total of 300 cycles and uses a smaller learning rate to make the model reach the optimal state more precisely during the final convergence process.
[0105] Table 3 is the YOLOv5 training parameter table adopted:
[0106] Table 3
[0107] Training parameters Object detection training set Number of training samples 13000 YOLOv5 stride 128 YOLOv5 segmentation 32 YOLOv5 scaled width 672 YOLOv5 scaled height 672 Number of training loops 500 / 300 Learning rate 0.001 / 0.0001
[0108] During the training process, the training effect of the model is represented by the mean intersection over union IoU. After one time, the IoU of the prediction result overlapping with the actual label represents the quality of the prediction. Among them, the overlapping area represents the overlapping part between the prediction box and the ground truth box, and the combined area represents the entire area occupied by the prediction box and the ground truth box. It can be seen that IoU can represent the quality of the model in detecting the target to be determined. The training effect after one iteration is represented by the loss:
[0109] loss = loss pos + loss size + loss confidence + loss class
[0110] Among them, loss pos represents the error of the center point of the detection box. loss size is the error of the length and width of the detection box. loss confidence characterizes the confidence error of the detection box. loss class represents the error of the classification of the detection box. There is only one type of classification in the framework of the present invention, so class is almost 0. loss0 represents the loss value of the previous iteration, and the detection effect of the final image is the superposition of the loss values after all iterations.
[0111] In this example, the loss value of the basic training set converges to below 2, and the loss value of the enhanced training set converges to 0.8. It is regarded as having good effect and is put into use.
[0112] Generation of vehicle feature enhancement weights based on the adversarial network:
[0113] Make the vehicle feature enhancement training weights under the adversarial network and select the GFLA algorithm. The training process requires corresponding training set pairs of multiple vehicle perspectives and specific vehicle perspectives with the background removed. Select the publicly available vehicle dataset VeRi-776, select the right rear perspective of the target as the fixed perspective, and unify all perspective vehicle images to this perspective. Combine all vehicle images in the training set into image pairs of any perspective and the right rear perspective, and artificially blur the background of the vehicle images in the fixed perspective as the basic weights for generating the training set. The parameters in the training process are configured with the default parameters of the GFLA algorithm.
[0114] Generation of similarity discrimination weights based on the siamese neural network:
[0115] Make the similarity discrimination weights of the siamese neural network and build the siamese neural network using the VGG16 framework. The training process integrates the publicly available vehicle re-identification dataset of Peking University and the publicly available dataset VeRi-776 of Beijing University of Posts and Telecommunications. Select the same vehicle images with the same perspective as a group of training groups. The total number of targets in the comprehensively sorted dataset is 2,500 groups, and the number of images is about 130,000. The standard size of the training images is determined by the target size. Usually, the side length of the vehicle detection target size is about 100 pixels. Select 100*100 as the image size of the target input to retain closer features. Use this as the training set of the siamese neural network to generate the basic weights.
[0116] The network training parameters based on VGG16 are shown in Table 4:
[0117] Table 4
[0118] Training parameters Similarity judgment training set Number of training samples 2500 groups Training step size 16 Number of training loops 30 / 80 Learning rate 0.01 / 0.001 Image side length 100
[0119] The training process is represented by a loss function. Since the similarity judgment result is a binary result, that is, there are only two cases of 0 or 1, the binary cross-entropy loss function (BCELoss) applicable to binary results is adopted. The calculation process of BCELoss is as follows:
[0120]
[0121] When BCELoss decreases and converges to less than 0.1, the model can be put into use.
[0122] Step 2: Vehicle target recognition, morphological feature and motion feature extraction. The main idea is to intercept the vehicle target through the object detection algorithm, extract the vehicle trajectory in real time, strengthen the vehicle morphological features by intercepting according to the image position of the trajectory, judge the vehicle following state according to the spatio-temporal information of the trajectory, and extract the following features of the vehicle operation. Vehicle feature enhancement mainly converts it into an image under a specific perspective, extracts the feature vectors of each frame respectively, and merges multiple vectors of the same vehicle. Judging the vehicle following state according to the spatio-temporal information of the vehicle trajectory and extracting the spatio-temporal features mainly judge whether the vehicle is in the following state and extract the parameter vector of its following process. It is specifically divided into the following three steps:
[0123] Detection of vehicle target position in the monitoring video:
[0124] Using the object detection weights of the YOLOv5 algorithm generated in Step 1, detect the vehicle target position in the monitoring screen. The detection process needs to judge whether the target is the original target in this screen. During the detection process, retain the vehicle target position in the previous frame of the detected road screen, and the set is denoted as P0. Calculate the intersection over union (IoU) by traversing the intersection of the target detection position in this frame and the vehicle target in the previous frame. The calculation of IoU is as follows:
[0125] IoU = A inter / A union
[0126] where, A inter represents the area of the intersection of the two rectangles, and A union represents the total area of the union of the two rectangles.
[0127] When the intersection over union is greater than the threshold TH IoU it is considered that the two targets of the current intersecting elements in p and P are the detection results of duplicate vehicles, and the same vehicle serial number label is assigned to these two solved pictures.
[0128] To prevent the vehicle from being misjudged as a newly added vehicle due to discontinuous recognition during the process and being assigned a new label, only retain the new label of the vehicle at the road end. The newly added vehicles appearing in the middle of the road are all judged as invalid newly added targets and discarded.
[0129] In this embodiment, according to the above steps, the trained result weights are used for object detection and vehicle object duplication determination, and the result is used as the final detection result. The detection effect and duplication determination result are shown in Table 5:
[0130] Table 5
[0131]
[0132] It can be seen that the duplication removal step can greatly reduce the number of matching vehicle objects in the final result, and can comprehensively consider the image form of the vehicle over a period of time, preventing the influence of missed detection, blurring, and occlusion problems caused by a single frame on vehicle feature extraction.
[0133] Enhanced feature vector extraction and merging of feature images:
[0134] After successfully detecting the vehicle object, it is cropped according to its bounding rectangle. The bounding rectangle comes from the position output during the detection process. Using the adversarial network feature enhancement weights generated in Step 1, the vehicle object image is input into the adversarial network to generate a corresponding canonical pose image. In this embodiment, the right rear view of the vehicle is selected.
[0135] Since the vehicle object is continuously captured during the video process, the vehicle form does not change significantly in a short period of time but is sent for calculation multiple times, resulting in a slow implementation process. To reduce the computational load, in this embodiment, for the detected and numbered vehicle objects, an interval processing method is adopted. Taking 0.5 seconds as the interval, the vehicle images intercepted by the object detection are sent into the adversarial network for feature enhancement. The generated feature-enhanced images include unified vehicle perspectives, removal of vehicle background information, and retention of the main features of the vehicle form.
[0136] The vehicle background information is mainly removed by blurring the sampling area. Gaussian blurring is used, and the calculation is as follows:
[0137]
[0138] where r is the blurring radius, which is taken as half of the sampling area size here, and σ is the standard deviation of the sampling area.
[0139] The newly generated feature-enhanced image of the vehicle is sent into the siamese neural network to extract the corresponding feature vectors. The image resolution is first standardized to 100*100 pixels and then sent into the neural network. The network structure is as follows: 2 convolutional layers with 64 kernels, a pooling layer, 2 convolutional layers with 128 kernels, a pooling layer, 3 convolutional layers with 256 kernels, a pooling layer, 3 convolutional layers with 512 kernels, a pooling layer, 3 convolutional layers with 512 kernels, and a pooling layer.
[0140] In each convolutional layer, each 3×3 convolutional kernel traverses every pixel of the entire image from the upper right corner. Finally, the output result is obtained by linking together the traversal and calculation results of all the convolutional layers in this layer to complete a set of convolutional operations. The essence of the convolutional operation is to retain necessary features. Using multiple convolutional layers can better retain feature information, so the convolutional calculation will cause the output matrix to continuously increase in dimension.
[0141] Each pooling layer reduces the dimension of the convolutional operation result with a 2×2 kernel. The role of the pooling layer is to reduce the data dimension and accelerate the operation. As the convolutional and pooling layers are stacked, the image data will be compressed into a one-dimensional feature vector containing necessary features, that is, the image feature vector of the final result.
[0142] For the feature vectors collected from the same vehicle in different images, to ensure that the description of the vehicle by the feature vector is not affected by instantaneous unstable factors such as blurring, occlusion, and picture quality, the feature vectors are strengthened through multiple iterations in the following way:
[0143]
[0144]
[0145] Among them: F is the strengthened feature vector, F i are the feature vectors to be merged, conf i are the confidence levels corresponding to the feature vectors to be merged, which are output by the YOLOv5 model for this object detection at the same time. The merged confidence level corresponding to this merged feature vector is conf. In this embodiment, the image generated by strengthening the vehicle features is as Figure 2 shown.
[0146] Following vehicle process determination and parameter extraction:
[0147] All vehicles appearing in the monitoring screen will be divided by lanes and recorded in a vehicle queue in the order of appearance. A minimum macro parameter calculation time unit is defined, and the traffic flow rate and vehicle density within the current time interval unit of the queue are calculated based on the vehicle queue. A threshold Q is selected to determine whether the vehicles appearing within this time unit are in a following vehicle state. When the traffic flow rate is less than the threshold Q, it is determined that this situation is in a free flow state. When the traffic flow threshold is greater than Q, it is determined that the vehicles within the current time unit are in a following vehicle state.
[0148] Generally, since the road range covered by a single monitoring field of view is short, when two vehicles in the same lane appear in the monitoring field of view at the same time, it is determined that the rear vehicle is in a following vehicle state, and the traffic flow threshold Q is selected according to this standard.
[0149] After the following-following state is determined, the vehicles in the following-following state will be selected for the effective following-following parameter period, which indicates that during this period, the following-following process parameters of these vehicles can be accurately calculated. Specifically, during this period, the motion states of both the following-following vehicle and the leading vehicle can be accurately captured by the surveillance cameras. During the effective following-following parameter period, the following-following parameters of the following-following vehicle are calculated. The specific parameters include: the relative following-following order i of the vehicle, the sequence of space headways SpaceHeadway(t), the sequence of time headways TimeHeadway(t), and the sequence of following-following speed differences Δv(t). The space headway, time headway, and following-following speed difference are calculated as follows:
[0150]
[0151] TimeHeadway(t) = (x i-1 (t) - x i (t)) * fps
[0152] Δv(t) = v i-1 (t) - v i (t)
[0153] where t represents the time series, fps represents the frame rate of the current video, x i is the position of vehicle i, and v i is the speed of vehicle i.
[0154] Step 3: Model construction and cross-camera target tracking. The main idea of this step is to construct a dynamic vehicle fleet model. With the dynamic prediction of the travel time of the dynamic vehicle fleet as a reference, based on the distribution characteristics of the dynamic vehicle fleet's morphological features and following-following characteristics, the dynamic vehicle fleet group matching is realized, and then the cross-camera target matching of individual vehicles within the vehicle fleet is completed. It is specifically divided into the following three steps:
[0155] Construct a dynamic vehicle fleet model:
[0156] Construct a dynamic vehicle fleet model for the vehicle queue passing through each camera. The dynamic vehicle fleet includes the vehicles that appear during the current research period arranged in order, and includes vectors of the enhanced morphological features and following-following motion features of these vehicles.
[0157] The research period includes the period not covered by the vehicles that have not been matched within this camera. The upper limit of this period can be set to ensure that the vehicle queue will not be too long and to ensure the search efficiency of the matching process. Once a vehicle is successfully matched, the research period will be changed, and the successfully matched vehicle will be removed from the research period.
[0158] Dynamic prediction of the travel time of the dynamic vehicle fleet:
[0159] Calculate the average speed of vehicles in each monitored video, refer to the mileage points of the road where the monitoring is located, and fit the average speed distribution curve of the entire research section Take the reciprocal of the vehicle average speed curve to obtain the relationship curve T between the time taken for vehicles to pass through a unit section of the road and the mileage distribution of the entire road du -X, denoted as Due to the driver's reaction time, the traffic wave will transfer upstream over time, and the downstream state will follow the traffic wave upstream. Predict this transfer, that is, calculate T du -The translation change speed V of X with respect to the time series t wave (t).
[0160] ΔX = t × V wave (t)
[0161] Where ΔX is the dynamic translation mileage of the traffic state, and t is the elapsed time. So there is:
[0162]
[0163] Select typical vehicles in the platoon at equal intervals according to the platoon following order i in the dynamic platoon. Denote the number of intervals as K, and calculate the iterative result AT according to the iterative formula of the time experienced by typical target vehicles in the research section formed as follows du,i Estimate the travel time of these K vehicles passing through the matching gap mileage:
[0164] X(t) k+1 = X(t) k + v o (t)×Δt
[0165]
[0166] Iterate the two formulas until X reaches the spatial position to be predicted or t reaches the time to be predicted, then stop the iteration and calculate the total travel duration AT of the typical vehicle i du,i :
[0167]
[0168] Where k is the current iteration number; N is the total number of iterations, take the number of times covering the length of the matching gap area as N, Δt is the iteration step time interval, and AT du,i Represents the total travel time of the i-th vehicle in the research section
[0169] Dynamic platoon group matching and vehicle individual matching:
[0170] Select typical vehicles in the platoon as representatives of the operating space-time characteristics. The appearance time of this vehicle is DT du, select the K targets before and after it to form a fleet group, with a total of 2K + 1 targets. Obtain the spatio-temporal information DT of the typical vehicle du +AT du As the center, select 2K + 1 vehicles in the dynamic queue formed by the target cameras, which correspond spatio-temporally to the 2K + 1 vehicles selected before prediction. From the two vehicle groups corresponding to the selected spatio-temporal information, select adjacent M targets in order, and check the similarity of the vehicle queues in the two groups of targets. The number of selected targets M decreases from 2K + 1 to K + 1. The selection of the two groups of targets includes both ends of the first and last positions, and the overlapping part decreases as M decreases.
[0171] The specific inspection method is implemented through a Siamese convolutional neural network. The input features of the re-identification Siamese neural network are the morphological enhancement feature vectors F corresponding to the vehicle groups, the headway sequence, the headway time sequence, and the following speed difference sequence. The output result is the matching result of the two vehicle queues, indicating the degree of complete matching, and the value range is [0, 1].
[0172] After realizing the matching of the vehicle groups, select the matching group with the highest matching degree value of the vehicle groups as the output of the truly corresponding group, and assign the same unique vehicle number to it to complete the cross-camera target tracking.
[0173] Correspond to the remaining vehicle targets in this batch in order to complete the K vehicle target groups in this dynamic fleet.
[0174] In the embodiment, the vehicle dynamic queue group matching process is as Figure 3 shown, the vehicle similarity matching display is as Figure 4 shown, and the accuracy rate of vehicle cross-camera recognition is shown in Table 6. Comparing directly sending the single-vehicle image target recognition result into the similarity determination network for target re-identification, it can prove the rationality of this method for cross-camera recognition.
[0175] Table 6
[0176]
[0177] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the present invention.
Claims
1. A multi-twin adversarial network cross-camera vehicle tracking method with enhanced coupled platoon following, characterized in that It includes the following steps: S10. Build a basic network for obtaining target feature information, make a dataset, and train the basic network. The basic network includes a neural network for target detection and an adversarial network for image normalization. The dataset contains a target detection dataset and an image normalization dataset. The target detection includes obtaining the vehicle position and vehicle size. The image normalization includes pose transformation and background removal. S20. Multiply the above basic network in multiple twins for multi-threaded synchronous processing of multiple surveillance videos. S30. Build a real-time trajectory extraction model in the multiple-twinned basic network, use the time-position relationship of vehicle operation to track the same target between different frames and assign a target serial number, intercept the trajectory images in the video, take the trajectory as the target unit, and extract and strengthen the vehicle morphology features through the adversarial network. The specific process is as follows: S31. Retain all vehicle target positions in the previous frame of the video, denoted as set P0, obtain all vehicle target positions in the current frame, denoted as set P1. The target positions are obtained from the detection results of the trained neural network. Let the current vehicle target position p(x, y, l, w) to be determined for duplication in set P1. S32. Calculate the intersection-over-union ratio for the rectangular elements in set P1 and set P0. The intersection-over-union ratio IoU calculation is expressed as follows: IoU = A inter / A union Among them, A inter represents the area of the intersection of the rectangles in the two sets, and A union represents the total area of the union of the rectangles in the two sets; S33. For the targets in set P1 with IoU equal to 0, consider them as new targets in the current frame and assign them new vehicle serial numbers separately. The remaining targets execute S34. S34. Send all the targets in set P1 with non-zero IoU and the elements in set P0 into the Hungarian algorithm for multi-target tracking. The Mahalanobis distance is used as the distance difference between targets in the Hungarian algorithm. The Mahalanobis distance calculation method is as follows: Among them, D M is a vector For the Mahalanobis distance of its distribution mean C is a vector The covariance matrix of the distribution. During the calculation process is the vehicle moving pixel vector (x0, y0) and (x1, y1) are the vehicle target positions in the previous frame and the current frame respectively; S35. For the targets in set P1 with non-zero IoU, if the tracking is successful, assign the successfully tracked targets in set P1 to the corresponding vehicle serial numbers. If the tracking is unsuccessful, assign new vehicle serial numbers. S36. Intercept the successfully tracked target images and use the GFLA adversarial network to strengthen the morphology features. According to the feature strengthening weights trained in S10, call the local attention mechanism in the network to identify the main features by traversing the image pixels. S37. Determine whether the strengthened morphology features are the background. If so, perform blur processing on the sampling area. Otherwise, retain the original pixel morphology and output the image. The blur processing uses Gaussian blur, and the calculation is as follows: Where G(r) represents the Gaussian blur function, r is the blur radius, which is taken as half of the sampling area size here, and σ is the standard deviation of the sampling area. S38. Add the time series to fuse the feature vectors of each frame of the same target. Change the image output by S37 into the corresponding one-dimensional feature vector by using small convolution kernels of 3×3 and 2×2. Multiple feature vectors with the same target label will be merged in sequence according to the following formula: Among them, F is the merged feature vector, i.e., the vehicle form enhancement feature vector, F i is the i-th feature vector to be merged, conf i is the confidence corresponding to the i-th feature vector to be merged, n represents the total number of frames in which the target vehicle appears in the video, i.e., the total number of features, CONF is the merged confidence corresponding to the merged feature vector, conf i-1 is the confidence corresponding to the (i - 1)-th feature vector to be merged; S39. Call the GFLA generator according to the merged feature vector F to generate the standard pose image trained in S10 to complete the feature strengthening. S40. Analyze the vehicle motion in real time according to the vehicle trajectory, divide the vehicle following state, extract the time-series characteristics of the following behavior, classify the target vehicle into a following state vehicle and a free state vehicle, construct a dynamic platoon model, and record the following characteristics and vehicle shape distribution characteristics of the platoon to which the target vehicle belongs. S50. Establish an iterative formula for predicting the travel time of the spatio-temporal position by integrating the dynamic platoon motion characteristics, the overall traffic flow state, and the multi-monitoring distribution, and predict the spatio-temporal position where the platoon appears across cameras. The specific process is as follows: S51. Calculate the average speed of vehicles in each monitored video, and refer to the mileage points of the road where the monitoring is located to fit the average speed distribution curve of the entire section of the road. S52, take the reciprocal of the vehicle average speed distribution curve to obtain the relationship curve T between the time taken for vehicle passage within a unit road section and the whole road mileage distribution du -X, denoted as S53. Due to the driver's reaction time, the traffic wave propagates upstream over time, and the downstream state follows the traffic wave upstream. Predicting this transfer, i.e., calculating T du - The translational change speed V of X with respect to the time series t wave (t): ΔX = t × V wave (t) Where ΔX is the dynamic translation mileage of the traffic state, then there is: S54. Form an iterative formula for calculating the time experienced by a typical target vehicle in the dynamic platoon on the entire road segment: X(t) k+1 = X(t) k + v i (t) × Δt v i (t) is the speed value of vehicle i at time t, X(t) k+1 is the vehicle position at the (k + 1)-th iteration, X(t) k is the vehicle position at the k-th iteration, is the time taken for the vehicle to pass through the position where the vehicle is located in the (k + 1)-th iteration, obtained from the relationship curve T du -X, is the time taken for the vehicle to pass through the position where the vehicle is located in the k-th iteration; Iterate the two equations until X reaches the spatial position to be predicted or t reaches the time to be predicted, then stop the iteration and calculate the total travel duration AT of the typical vehicle i du,i : where k is the current iteration number; N is the total number of iterations, and the number of times to cover the length of the matching gap region is N, Δt is the time interval of each iteration step, and AT du,i represents the total driving time of vehicle i over the entire road section; S55, Select typical vehicles in the vehicle platoon at equal intervals according to the vehicle platoon following order i, denote the number of intervals as K, and calculate the iterative result AT du,i Estimate the driving time of these K vehicles through the matching gap mileage; S60. Integrate the following characteristics of the dynamic platoon, the predicted spatio-temporal position where the platoon appears across cameras, and the obtained vehicle shape distribution characteristics within the platoon to achieve the cross-camera matching of the platoon group under multi-monitoring. S70. Based on the results of the group cross-camera matching, sequentially track and number the corresponding target shape feature distributions of individuals within the dynamic platoon to complete the cross-camera tracking of vehicles.
2. The multi-twin adversarial network cross-camera vehicle tracking method for enhancing coupled platoon following according to claim 1, wherein As described in S10, build a basic network for obtaining target feature information, make a data set, and train the basic network, which specifically includes: S11. For the neural network used for target detection, use the YOLOv5 model under the pytorch framework for training and detection. During the training process, the images in the target detection data set are scaled to a unified size and then sent into the YOLOv5 model in batches for logistic regression prediction. The training effect of the YOLOv5 model is evaluated by the loss value. The loss loss of the training effect after one iteration is expressed as follows: loss=loss xy +loss lw +loss confidence +loss class Among them, loss xy represents the error of the center point of the detection box, loss lw represents the error of the length and width of the detection box, loss confidence characterizes the confidence error of the detection box, loss class represents the error of the detection box classification; when the loss value converges and no longer changes, the YOLOv5 model is put into use; S12. For the feature-enhanced siamese adversarial network used for image normalization, the adversarial network uses the GFLA algorithm. Based on the vehicle data set VeRi-776, manually construct a training target pair for the feature enhancement process of the adversarial network. One set of training target pairs contains two images, which are of the same vehicle target. The first image is a picture of the target from any perspective, and the second image is a picture of the target from a specific perspective with the background blurred. Make multiple sets of training target pairs for the same target, and combine multiple perspectives with the specific perspective respectively. Select several vehicle targets in the image normalization data set and perform the above combination processing to form multiple sets of perspective-corresponding-to-specific-perspective training target pairs for several vehicle targets, which are used as the training data set of the feature-enhanced siamese adversarial network. S13. The neural network is applied to the surveillance video to obtain the target position x, y, the target pixel size l, w, the target confidence conf, and crop the target image. The adversarial network is applied to the cropped target image for image perspective transformation and background deletion.
3. The method for cross-camera vehicle tracking based on a multi-twin adversarial network with enhanced coupled platoon following according to claim 1, wherein In the above S20, read the video streams captured by different surveillance cameras, obtain the current video frames, create corresponding image libraries for different video streams, allocate separate threads and corresponding basic networks for each video stream, and parallelly implement the target detection, pose transformation, and background removal of the basic network.
4. The method for cross-camera vehicle tracking based on a multi-twin adversarial network with enhanced coupled platoon following according to claim 1, wherein The specific process of the above S40 is as follows: S41. Identify the current road following situation. Determine whether the vehicles on the current road are driving in a following state based on the road flow rate and the average vehicle speed. For vehicles not driving in a following state, directly predict the spatio-temporal position of the vehicle using its speed. If it is in a following state, proceed to S42; S42. For vehicles in the same lane driving in a following state, they will be paired one by one as the following vehicle and the lead vehicle. A group of following vehicles and lead vehicles includes the vehicle in front (the lead vehicle) and the vehicle immediately following it (the following vehicle); S43. Create a corresponding dynamic vehicle fleet for each surveillance video. The length of the dynamic vehicle fleet includes all vehicles in all surveillance videos, and a vehicle is counted when it drives out from under the current camera; The features included in the dynamic vehicle following queue are: the relative following order i of the vehicle, the vehicle shape enhancement feature vector F calculated in S38, the headway sequence SpaceHeadway(t), the time headway sequence TimeHeadway(t), and the following speed difference sequence Δv(t); among them, the headway, time headway, and following speed difference are calculated as follows: TimeHeadway(t)=(x i-1 (t)-x i (t))*fps Δv(t) = v i-1 (t) - v i (t) where t represents the time series, fps represents the frame rate of the current video, x i-1 is the position of vehicle i - 1, x i is the position of vehicle i, v i is the speed of vehicle i, v i-1 is the speed of vehicle i - 1.
5. The method for cross-camera vehicle tracking with a multi-twin adversarial network for enhanced coupled platoon following according to claim 1, wherein, The specific process of S60 is as follows: S61, Select a typical vehicle in the vehicle fleet as the representative of the operating spatio-temporal characteristics. The appearance time of this vehicle is DT du , Select the K targets before and after it to form a vehicle fleet group, with a total of 2K + 1 targets; S62, using the spatio-temporal information DT du +AT du as the center, in the dynamic queue corresponding to the monitoring camera of the target to be re-identified downstream, select 2K + 1 vehicles that are closest to the time center. These 2K + 1 vehicles form a group of vehicles to be matched du +AT du centered on S63. From the two vehicle groups corresponding to the spatio-temporal information selected in S61 and S62, select adjacent M targets in order, and check the similarity of the vehicle queues in the two groups of targets according to S64. The number of selected targets M decreases from 2K + 1 to K + 1. The selection of the two groups of targets includes the head and the tail, and the overlapping part decreases as M decreases; S64. Specifically, vehicle queue re-identification is realized through a siamese convolutional neural network. The input features of the re-identification siamese neural network are the vehicle shape enhancement feature vector F, the headway sequence, the time headway sequence, and the following speed difference sequence corresponding to the vehicle group, and the output result is the matching result of the two vehicle queues, indicating the degree of complete matching, and the value range is [0, 1].
6. The method for cross-camera vehicle tracking based on a multi-twin adversarial network with enhanced coupled platoon following according to claim 1, wherein S60 realizes cross-camera matching of vehicle fleets under multiple surveillances. After realizing vehicle group matching, select the matching group with the highest vehicle group matching degree value as the truly corresponding group for output, and correspond to the remaining vehicle targets in this batch in order, and complete K vehicle target groups in the dynamic vehicle fleet.
Citation Information
Patent Citations
Vehicle recognition and tracking method based on convolutional neural networks
CN108171112A
Vehicle control method based on reinforcement learning control strategy in hybrid fleet
CN112162555A