Vehicle identity information real-time identification and labeling method based on deep learning
By adopting deep learning and multi-camera collaborative tracking technology in vehicle identity information recognition, combining edge computing and adaptive learning, the problems of low recognition accuracy and efficiency in the prior art are solved, and efficient vehicle identity recognition and tracking in complex environments are achieved.
Patent Information
- Application Number
- CN202510189029.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing vehicle identity information identification algorithms have significantly reduced recognition accuracy and efficiency in the case of poor lighting, severe occlusion and complex traffic dynamics, and the centralized architecture is difficult to meet the real-time requirements in large-scale transportation networks.
Real-time identification and labeling of vehicle identity information based on deep learning is adopted, and vehicle feature extraction and preliminary identification are performed through distributed processing of edge nodes and central servers. Combined with multi-camera collaborative tracking and space-time trajectory prediction technology, global optimization labeling is performed, and identification accuracy and efficiency are improved through adaptive learning and model updates.
It significantly improves the accuracy and efficiency of vehicle identification, reduces the false detection rate, realizes vehicle identity tracking and identification in complex environments, and meets the real-time needs of large-scale transportation networks.
Smart Images

Figure CN120126085A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle identity recognition, and in particular to a real-time recognition and annotation method for vehicle identity information based on deep learning. Background Art
[0002] In current vehicle identity information recognition solutions, it mainly relies on technologies such as license plate recognition and vehicle appearance feature extraction. These methods work well in simple traffic scenarios, but in situations with poor lighting, severe occlusion, and complex traffic dynamics, the recognition accuracy and efficiency decrease significantly. In addition, as the scale of the traffic camera network expands, it faces real-time and bandwidth pressures when processing a large amount of data. A single data source and a centralized architecture are difficult to ensure smoothness when dealing with complex environments.
[0003] To address these problems, traditional solutions mainly rely on deep learning models for feature extraction and matching, and at the same time optimize camera layout and rely on a central server to process massive data under specific conditions. However, although deep learning has improved the ability to extract appearance features, its adaptability to perspective changes, lighting conditions, and vehicle occlusion is limited, resulting in unsatisfactory recognition accuracy in complex scenarios. In addition, cross-camera vehicle identity tracking relies on fixed camera perspectives and historical trajectory data, making it difficult to handle dynamic changes in traffic. Although the centralized architecture effectively improves the processing ability, when facing a large-scale camera network, the server is prone to overload, resulting in processing delays and affecting the real-time response of the system.
[0004] Although traditional solutions have improved the efficiency of recognition and annotation to a certain extent, in cases of vehicle occlusion, poor lighting, or similar appearances, the false detection rate of existing recognition algorithms is relatively high. Secondly, cross-camera spatio-temporal trajectory prediction is still difficult to accurately track vehicles in the face of insufficient multi-camera coverage or complex traffic dynamics. Finally, the centralized processing architecture is difficult to meet the real-time requirements in large-scale traffic networks, and the pressure on bandwidth and server load is too large, which will lead to a decline in system performance. Therefore, there is an urgent need for a real-time recognition and annotation method for vehicle identity information based on deep learning to solve such problems. Summary of the Invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] The present invention provides a real-time recognition and annotation method for vehicle identity information based on deep learning to solve the problem of relatively high false detection rate of existing recognition algorithms in cases of vehicle occlusion, poor lighting, or similar appearances.
[0007] To solve the above technical problems, the present invention provides the following technical solutions:
[0008] The present invention provides a method for real-time identification and annotation of vehicle identity information based on deep learning, which includes,
[0009] Step S1, vehicle feature extraction and preliminary identification,
[0010] Distributed processing of vehicle identity information is performed using edge nodes and a central server;
[0011] Step S2, vehicle re-identification,
[0012] Based on the preliminary structured data transmitted by the edge nodes in Step S1, the central server uses a deep learning model to deeply identify the detailed features of the vehicle,
[0013] Step S3, multi-camera collaborative tracking,
[0014] Based on the detailed feature extraction results in Step S2, combined with the spatio-temporal information of the vehicle, spatio-temporal trajectory prediction technology is used to track the moving trajectory of the vehicle, analyze the real-time motion data and historical trajectory of the vehicle, predict its travel path, and notify the next camera node in advance to prepare for identification,
[0015] Step S4, data aggregation and global optimization,
[0016] The central server integrates the multi-camera recognition data transmitted by the edge nodes, and combines the spatio-temporal trajectory prediction results in Step S3. The server performs global optimization annotation on the vehicle information under different cameras,
[0017] Step S5, adaptive learning and model update,
[0018] The vehicle feature data and trajectory information collected in Steps S1 to S4 are incorporated into the existing model for adaptive learning, and online learning technology is used for dynamic update according to new data, so as to improve the recognition accuracy and efficiency under different environments, vehicle details, and multi-camera collaboration, and continuously enhance the recognition ability.
[0019] Further, in Step S1, the edge nodes are deployed near cameras and sensors to process video streams and sensor data in real time, perform preliminary vehicle feature extraction and identification, and the sensor data comes from LIDAR, V2X vehicle networking communication, and GPS.
[0020] Further, in Step S1, the method for performing preliminary vehicle feature extraction and identification is: using a lightweight deep learning model to extract the basic features of the vehicle, generate structured data, which includes the preliminary identification information of the vehicle, license plate, color, and vehicle type, and send it to the central server.
[0021] Further, in Step S1, the methods for performing preliminary vehicle feature extraction and identification include:
[0022] Normalize the original image captured by the camera. Assume the input image is a three-channel color image with a size of H×W, representing the color values corresponding to the pixel coordinates x and y. Perform the following normalization operation on it:
[0023] where, represents the pixel value of the normalized image, represents the pixel value of the original image, μ represents the average pixel value of the training dataset, and σ represents the pixel standard deviation of the training dataset;
[0024] Perform convolution calculation on the input image as follows:
[0025] where, represents the activation value at position (x,y) in the output feature map, K represents the number of convolution kernels, m,n represent the size of the convolution kernel, and the convolution kernel size is (2m + 1)×(2n + 1), represents the weight of the k-th convolution kernel at position (i,j), and b k represents the bias of the k-th convolution kernel, represents the normalized pixel value of the input image, and depthwise separable convolution can be used here as an alternative;
[0026] Pass the feature map after the convolutional layer to the fully connected layer to generate structured data. The fully connected layer is:
[0027] where, z represents the output of the fully connected layer, that is, the basic features of the vehicle extracted, represents the weight matrix of the fully connected layer, and b fc represents the bias of the fully connected layer, represents flattening the feature map into a vector;
[0028] Finally, convert the feature vector z into structured data and format it into the vehicle recognition information:
[0029] Vehicle_Data = {License_Plate, Color, Model}, where Vehicle_Data represents the structured data of the vehicle, including the license plate number, color, and model. License_Plate represents the license plate information recognized by the model, Color represents the color information of the vehicle, and Model represents the model information of the vehicle.
[0030] Further, in step S2, the detailed features identified by the deep learning model include vehicle lights, vehicle logos, and scratches.
[0031] Further, in step S2, the vehicle re-identification method is as follows:
[0032] Assume that the preliminary structured data passed from step S1 is D initial , including license plate, color, and vehicle model:
[0033] D initial = {L, C, M}, where L represents the license plate recognition result, C represents the vehicle color recognition result, and M represents the vehicle model recognition result;
[0034] Use the deep learning model to extract the details of the vehicle image. Assume that the input vehicle image feature is I feat , and the feature map after convolution is F conv :
[0035] Among them, F conv (x, y) represents the activation value of the feature map output by the convolutional layer at the position (x, y), and F input (x, y) represents the input feature map, that is, the local area feature of the vehicle (such as vehicle lights, vehicle logos, etc.), represents the weight of the k-th convolutional kernel at the position (i, j), K represents the number of convolutional kernels, i, j represent the spatial position of the convolutional kernel, and ReLU(·) represents the activation function;
[0036] Use the attention mechanism to weight the importance of different detailed features. Assume that the input feature map is F conv , generate the attention weight matrix A(x, y) and multiply it with the feature map:
[0037] Among them, F att (x, y) represents the feature map after attention weighting, A(x, y) represents the attention weight matrix, which is used to weight the feature map F conv , exp(·) represents the exponential function, which is used to calculate the probability distribution of the attention weights, and x', y' represent the coordinates of all pixel points in the convolutional feature map F conv (x', y');
[0038] Flatten the detailed features extracted by convolution and the attention mechanism into a vector z detail , and then generate high-dimensional features for recognition through the fully connected layer: z detail = flatten, y = σ(W fc ·z detail + b fc ), where z detaildenotes the flattened detailed feature vector, y denotes the output of the fully connected layer, and W fc denotes the weight matrix of the fully connected layer, and b fc denotes the bias of the fully connected layer, σ(·) denotes the activation function for classification, and Softmax and Sigmoid are selected. y contains classification information of vehicle detailed features, such as headlight type, logo style, and scratch position;
[0039] Combine the detailed features with the preliminary recognition information to form complete structured data: Among them, D final denotes the finally generated structured data, and y denotes the detailed feature recognition result obtained from the fully connected layer.
[0040] Furthermore, in step S3, the multi-camera collaborative tracking method is as follows:
[0041] Let the spatial position and speed of the captured vehicle at a certain moment t be p t =(x t , y t ), where p t denotes the two-dimensional spatial position of the vehicle at moment t, x t , y t denote the abscissa and ordinate of the vehicle at moment t, and the speed of the vehicle is where v t denotes the speed vector of the vehicle at moment t, denotes the lateral speed and longitudinal speed of the vehicle at moment t;
[0042] Use a prediction model based on historical trajectories and real-time speeds for prediction. Assume that the trajectory information of the vehicle between time t - k and time t is where, denotes the historical trajectory of the previous k time steps before time t, and p t-k , …, p t denote the sequence of spatial positions of the vehicle from time t - k to t. Use the Kalman filter KalmanFilter to smooth and predict the trajectory. The Kalman filter state update equation is:
[0043] p t+1 = p t + v t ·Δt + w t , where p t+1 denotes the predicted position at the next moment t + 1, p t denotes the position at the current moment t, v t denotes the speed at the current moment t, Δt denotes the time step, that is, the prediction time interval, and w t denotes the process noise, and w t is modeled as zero-mean Gaussian noise: Among them, represents Gaussian noise with zero mean and covariance Q.
[0044] Furthermore, in step S3, the multi-camera collaborative tracking method further includes:
[0045] According to the real-time motion data and predicted trajectory of the vehicle, use the trajectory prediction model to adjust the prediction result. Assume that the future motion trajectory model of the vehicle is a linear acceleration model:
[0046] where p t+n represents the predicted vehicle position at the nth future time step, a t represents the vehicle acceleration at time t, and n represents the number of future prediction time steps;
[0047] Assume that the position of the next camera is c i =(x i , y i ), and the position of the current camera is c 0 =(x 0 , y 0 ). Then the time for the vehicle to travel from the current camera to the next camera is:
[0048] where T i represents the estimated time for the vehicle to travel from the current camera to the next camera, |p t - c i | represents the Euclidean distance between the current vehicle position and the position of the next camera, and |v t | represents the magnitude of the vehicle speed;
[0049] Combining the historical trajectory, current motion state, and environmental information of the vehicle, use the Bayesian prediction model for global optimization. The Bayesian update formula is:
[0050] where represents the probability of predicting the future position of the vehicle based on the historical trajectory, represents the likelihood of the historical trajectory given the future position, P(p t+n ) represents the prior probability of the future position, represents the marginal likelihood of the historical trajectory,
[0051] According to the predicted time T i for the vehicle to reach the next camera, notify the next camera node in advance to prepare for recognition.
[0052] Furthermore, in step S4, the data aggregation and global optimization method is:
[0053] Each camera C i At time t i The recognized vehicle information is: Among them, Indicates the vehicle information recognized by camera C i At time t i The recognized vehicle information, F plate Represents the license plate feature vector of the vehicle, F color Represents the color feature vector of the vehicle, F model Represents the vehicle model feature vector, F detail Represents the detailed feature vector of the vehicle;
[0054] Combined with the spatio-temporal trajectory prediction result in step S3, assuming that camera C i At time t i The captured vehicle position is The movement trajectory of the vehicle is:
[0055] Among them, Indicates the predicted position of the vehicle at the next moment t i+1 The position, Indicates time t i The current position of the vehicle, Indicates time t i The speed vector of the vehicle, Indicates time t i The acceleration vector of the vehicle, Δt represents the time interval between time t i+1 And t i The time interval between;
[0056] Assuming that camera C i And camera C j The recognized vehicle information is respectively And Their feature similarity is:
[0057] Among them, Indicates the recognition similarity of camera C i And C j For the same vehicle, α 1 , α 2 , α 3 , α 4 Represents each feature weight coefficient, cos(·,·) represents calculating the cosine similarity between two vectors.
[0058] Furthermore, in step S4, the data aggregation and global optimization method also includes:
[0059] Based on the previous spatio-temporal trajectory prediction and similarity, the central server performs global annotation on the vehicle recognition results under different cameras, constructs a graph G(V, E) under multiple cameras through weighted graph matching, where each node represents a vehicle recognition result, and the weight of the edge represents the similarity. The optimization goal is to maximize the global similarity, and the optimization objective function is:
[0060] where, represents the global optimization objective function, represents the feature similarity of the vehicle under cameras C i and C j , λ represents the regularization coefficient, represents the spatio-temporal position difference between two cameras, and minimizing the objective function ensures that under the premise of maximizing the feature similarity, the spatio-temporal position difference is minimized to achieve global optimization annotation;
[0061] Based on the global optimization, a unique identifier ID k is assigned to each vehicle, and the final global annotation information is:
[0062] D global ={ID k , F plate , F color , F model , F detail , p t}, where D global represents the vehicle information after global annotation, ID j represents the globally unique identifier assigned to the vehicle, and p t represents the spatio-temporal position of the vehicle in the global coordinate system.
[0063] The beneficial effects of the present invention are as follows:
[0064] In the present invention, edge computing and the central server are used for collaborative processing to reduce data transmission latency. A multi-level feature extraction mechanism is introduced. Basic features such as license plates, colors, and vehicle models are quickly extracted at the edge nodes, and detailed features are extracted on the central server. Additionally, an attention mechanism is introduced to focus on the key details of the vehicle, effectively distinguishing the appearance differences of the same vehicle under different cameras or angles, and greatly improving the re-identification ability.
[0065] In the present invention, spatio-temporal trajectory prediction technology is adopted. Combining the real-time motion state and historical trajectory of the vehicle, the future path of the vehicle is predicted, and the next camera node is notified in advance through the Kalman filter, linear acceleration model, and Bayesian prediction model to prepare for recognition, enabling seamless switching and continuous tracking of the vehicle among multiple cameras.
[0066] The present invention performs global similarity optimization and weighted graph matching, combines the feature similarity of vehicles and the differences in spatio-temporal positions, maintains the consistency of recognition results under different cameras, and through global annotation and unique ID assignment, can effectively integrate recognition data in a large-scale camera network, reduce the situations of misrecognition and duplicate annotation, and improve the consistency. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0068] Figure 1 It is a schematic flowchart of the method for real-time recognition and annotation of vehicle identity information based on deep learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention with reference to the drawings of the specification.
[0070] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0071] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not all refer to the same embodiment, nor is it a separate or selectively exclusive embodiment from other embodiments.
[0072] Embodiment 1, referring to Figure 1 This embodiment provides a method for real-time recognition and annotation of vehicle identity information based on deep learning, including the following steps:
[0073] Step S1, vehicle feature extraction and preliminary recognition,
[0074] Perform distributed processing of vehicle identity information using edge nodes and a central server;
[0075] In step S1, the edge nodes are deployed near cameras and sensors to process video streams and sensor data in real time, perform preliminary vehicle feature extraction and recognition, and the sensor data comes from LIDAR, V2X vehicle networking communication, and GPS.
[0076] In step S1, the method for preliminary vehicle feature extraction and recognition is as follows: a lightweight deep learning model is used to extract the basic features of the vehicle, generate structured data, which includes the preliminary recognition information of the vehicle, license plate, color, and vehicle type, and send it to the central server;
[0077] In step S1, the methods for preliminary vehicle feature extraction and recognition include:
[0078] Normalize the original image collected by the camera. Assume the input image is a three-channel color image with a size of H×W, representing the color values corresponding to the pixel coordinates x and y, and perform the following normalization operation on it:
[0079] where represents the pixel value of the normalized image, represents the pixel value of the original image, μ represents the average pixel value of the training dataset, and σ represents the pixel standard deviation of the training dataset;
[0080] Perform convolution calculation on the input image as follows:
[0081] where represents the activation value at position (x,y) in the output feature map, K represents the number of convolution kernels, m,n represent the size of the convolution kernel, and the convolution kernel size is (2m + 1)×(2n + 1), represents the weight of the kth convolution kernel at position (i,j), and b k represents the bias of the kth convolution kernel, represents the normalized pixel value of the input image, and depthwise separable convolution can be used here as an alternative;
[0082] Transfer the feature map after the convolution layer to the fully connected layer to generate structured data. The fully connected layer is as follows:
[0083] where z represents the output of the fully connected layer, that is, the basic features of the vehicle extracted, represents the weight matrix of the fully connected layer, and b fc represents the bias of the fully connected layer, represents flattening the feature map into a vector;
[0084] Finally, convert the feature vector z into structured data and format it into the recognition information of the vehicle:
[0085] Vehicle_Data = {License_Plate, Color, Model}, where Vehicle_Data represents the structured data of the vehicle, including the license plate number, color, and model. License_Plate represents the license plate information recognized by the model, Color represents the color information of the vehicle, and Model represents the model information of the vehicle;
[0086] Specifically, by leveraging the collaborative work of edge computing and the central server, the processing of real-time video streams and sensor data is decentralized to edge nodes, reducing the latency of data transmission. This enables the rapid identification of basic vehicle information in a congested road environment, significantly improving the system's response speed. At the same time, lightweight deep learning models are deployed on edge nodes, which can operate efficiently in resource-constrained computing environments for the preliminary extraction of vehicle information.
[0087] Step S2, vehicle re-identification
[0088] Based on the preliminary structured data transmitted by the edge nodes in Step S1, the central server uses a deep learning model to deeply identify the detailed features of the vehicle.
[0089] In Step S2, the detailed features identified using the deep learning model include vehicle lights, logos, and scratches.
[0090] In Step S2, the vehicle re-identification method is as follows:
[0091] Assume that the preliminary structured data transmitted from Step S1 is D initial , including the license plate, color, and model:
[0092] D initial = {L, C, M}, where L represents the license plate recognition result, C represents the vehicle color recognition result, and M represents the vehicle model recognition result;
[0093] Use a deep learning model to extract the details of the vehicle image. Assume that the input vehicle image feature is I feat , and the feature map after convolution is F conv :
[0094] Among them, F conv (x, y) represents the activation value of the feature map output by the convolutional layer at the position (x, y), and F input (x, y) represents the input feature map, that is, the local region feature of the vehicle (such as vehicle lights, logos, etc.), represents the weight of the kth convolutional kernel at the position (i, j), K represents the number of convolutional kernels, i, j represent the spatial position of the convolutional kernel, and ReLU(·) represents the activation function;
[0095] The importance of different detailed features is weighted by an attention mechanism. Assume that the input feature map is F conv , and an attention weight matrix A(x, y) is generated and multiplied by the feature map:
[0096] where F att (x, y) represents the feature map after attention weighting, and A(x, y) represents the attention weight matrix, which is used to weight the feature map F conv . exp(·) represents the exponential function, which is used to calculate the probability distribution of the attention weights. x' and y' represent the coordinates of all pixel points in the convolutional feature map F conv (x', y');
[0097] Flatten the detailed features extracted by convolution and the attention mechanism into a vector z detail , and then generate high-dimensional features for recognition through a fully connected layer: z detail = flatten, y = σ(W fc ·z detail + b fc ), where z detail represents the flattened detailed feature vector, y represents the output of the fully connected layer, W fc represents the weight matrix of the fully connected layer, b fc represents the bias of the fully connected layer, and σ(·) represents the activation function for classification. Softmax and Sigmoid are selected. y contains the classification information of vehicle detailed features, such as headlight type, logo style, and scratch position;
[0098] Combine the detailed features with the preliminary recognition information to form complete structured data: where D final represents the finally generated structured data, and y represents the recognition result of the detailed features obtained from the fully connected layer;
[0099] Specifically, in the in-depth recognition of detailed features, the central server further extracts vehicle features by using a convolutional neural network combined with an attention mechanism. Even under the conditions of multiple angles and multiple cameras, it can still effectively extract vehicle detailed features.
[0100] Step S3, multi-camera collaborative tracking,
[0101] Based on the detailed feature extraction results in step S2, combined with the spatio-temporal information of the vehicle, use spatio-temporal trajectory prediction technology to track the moving trajectory of the vehicle, analyze the real-time motion data and historical trajectory of the vehicle, predict its travel path, and notify the next camera node in advance to prepare for recognition.
[0102] In step S3, the multi-camera collaborative tracking method is:
[0103] Let the spatial position and velocity of the captured vehicle at a certain moment \(t\) be \(p\) t =(x t , y t ), where \(p\) t represents the two-dimensional spatial position of the vehicle at moment \(t\), \(x\) t , \(y\) t represent the abscissa and ordinate of the vehicle at moment \(t\), and the velocity of the vehicle is where \(v\) t represents the velocity vector of the vehicle at moment \(t\), represents the lateral velocity and longitudinal velocity of the vehicle at moment \(t\);
[0104] Use a prediction model based on historical trajectories and real-time velocities for prediction. Assume that the trajectory information of the vehicle between moment \(t - k\) and moment \(t\) is where represents the historical trajectory of the previous \(j\) time steps before moment \(t\), \(p\) t-k , …, \(p\) t represent the sequence of spatial positions of the vehicle from moment \(t - k\) to \(t\). Use the Kalman filter KalmanFilter to smooth and predict the trajectory. The Kalman filter state update equation is:
[0105] \(p\) t+1 = \(p\) t + \(v\) t ·\(\Delta t\) + \(w\) t , where \(p\) t+1 represents the predicted position at the next moment \(t + 1\), \(p\) t represents the position at the current moment \(t\), \(v\) t represents the velocity at the current moment \(t\), \(\Delta t\) represents the time step, that is, the prediction time interval, and \(w\) t represents the process noise, and \(w\) t is modeled as zero-mean Gaussian noise: where represents zero-mean Gaussian noise with covariance \(Q\);
[0106] In step S3, the multi-camera collaborative tracking method further includes:
[0107] According to the real-time motion data and predicted trajectory of the vehicle, use the trajectory prediction model to adjust the prediction result. Assume that the future motion trajectory model of the vehicle is a linear acceleration model:
[0108] where \(p\) t+n represents the predicted position of the vehicle at the future \(n\)th time step, \(a\) t represents the acceleration of the vehicle at moment \(t\), and \(n\) represents the number of future prediction time steps;
[0109] Assume that the position of the next camera is \(c\)i =(x i , y i ), the position of the current camera is c 0 =(x 0 , y 0 ), then the time for the vehicle to travel from the current camera to the next camera is:
[0110] where T i represents the estimated time for the vehicle to travel from the current camera to the next camera, |p t - c i | represents the Euclidean distance between the current position of the vehicle and the position of the next camera, |v t | represents the magnitude of the vehicle's speed;
[0111] Combining the historical trajectory, current motion state, and environmental information of the vehicle, a Bayesian prediction model is used for global optimization. The Bayesian update formula is:
[0112] where represents the probability of predicting the future position of the vehicle based on the historical trajectory, represents the likelihood of the historical trajectory given the future position, P(p t+n ) represents the prior probability of the future position, represents the marginal likelihood of the historical trajectory,
[0113] According to the predicted time T i for the vehicle to reach the next camera, notify the next camera node in advance to prepare for recognition;
[0114] Specifically, for vehicle tracking in a multi-camera environment, a spatio-temporal trajectory prediction technology is adopted. By combining the real-time motion data of the vehicle with the historical trajectory for prediction, the tracking effect of multiple cameras is ensured. The Kalman filter and linear acceleration model can effectively predict the travel path of the vehicle and can notify the next camera node in advance to prepare for recognition, thus achieving seamless switching under multiple camera perspectives.
[0115] Step S4, data aggregation and global optimization,
[0116] The central server integrates the multi-camera recognition data transmitted by the edge nodes and, combined with the spatio-temporal trajectory prediction results in step S3, the server globally optimizes and annotates the vehicle information under different cameras.
[0117] In step S4, the data aggregation and global optimization method is:
[0118] The vehicle information recognized by each camera C i at time t i is: Among them, represents the vehicle information recognized by camera C i at time t i F, plate represents the license plate feature vector of the vehicle, F color represents the color feature vector of the vehicle, F model represents the vehicle model feature vector of the vehicle, F detail represents the detailed feature vector of the vehicle;
[0119] Combined with the spatio-temporal trajectory prediction result in step S3, assuming that camera C i at time t i captures the vehicle position as The movement trajectory of the vehicle is:
[0120] Among them, represents the predicted position of the vehicle at the next moment t i+1 , represents the current position of the vehicle at time t i , represents the vehicle speed vector at time t i , represents the vehicle acceleration vector at time t i , Δt represents the time interval between time t i+1 and t i ;
[0121] Assuming that the vehicle information respectively recognized by camera C i and camera C j is and Its feature similarity is:
[0122] Among them, represents the recognition similarity of camera C i and C j for the same vehicle, α 1 , α 2 , α 3 , α 4 represents each feature weight coefficient, and cos(·,·) represents calculating the cosine similarity between two vectors;
[0123] In step S4, the data aggregation and global optimization method also includes:
[0124] Based on the previous spatio-temporal trajectory prediction and similarity, the central server performs global annotation on the vehicle recognition results under different cameras, constructs a weighted graph matching to build a graph G(V, E) under multiple cameras, where each node represents a vehicle recognition result, and the weight of the edge represents the similarity. The optimization goal is to maximize the global similarity, and the optimization objective function is:
[0125] Among them, represents the global optimization objective function, represents the feature similarity of the vehicle under cameras C i and C j , λ represents the regularization coefficient, represents the spatio-temporal position difference between two cameras, and minimizing the objective function ensures that under the premise of maximizing the feature similarity, the spatio-temporal position difference is minimized, achieving global optimization annotation;
[0126] Based on the global optimization, a unique identifier ID is assigned to each vehicle k , and the final global annotation information is:
[0127] D global ={ID k , F plate , F color , F model , F detail , p t}, where D global represents the vehicle information after global annotation, ID k represents the global unique identifier assigned to the vehicle, and p t represents the spatio-temporal position of the vehicle in the global coordinate system.
[0128] Specifically, a weighted graph of multi-camera recognition data is constructed, and global similarity optimization is performed to ensure that the recognition results of the same vehicle under different cameras are consistent. Combining the spatio-temporal trajectory differences of the vehicles, it is ensured that the annotation of each vehicle globally is consistent and accurate, greatly improving the stability and recognition accuracy in a large-scale camera network.
[0129] Step S5, Adaptive learning and model update,
[0130] Incorporate the vehicle feature data and trajectory information collected in steps S1 to S4 into the existing model for adaptive learning, and use online learning technology to perform dynamic updates according to new data, improving the recognition accuracy and efficiency for different environments, vehicle details, and multi-camera collaboration, and continuously enhancing the recognition ability.
[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for real-time identification and labeling of vehicle identity information based on deep learning, characterized in that: include, Step S1, vehicle feature extraction and preliminary identification, using edge nodes and central servers to perform distributed processing of vehicle identity information; Step S2, vehicle re-identification: Based on the preliminary structured data transmitted by the edge node in step S1, the central server uses a deep learning model to deeply identify the detailed features of the vehicle. Step S3, multi-camera collaborative tracking, based on the detailed feature extraction results in step S2, combined with the vehicle's spatiotemporal information, uses spatiotemporal trajectory prediction technology to track the vehicle's movement trajectory, analyzes the vehicle's real-time motion data and historical trajectory, predicts its travel path, and notifies the next camera node in advance to prepare for recognition. Step S4: Data aggregation and global optimization: The central server integrates the multi-camera recognition data transmitted by the edge nodes, and combines the spatiotemporal trajectory prediction results of step S3 to globally optimize and annotate the vehicle information under different cameras. Step S5, adaptive learning and model updating, in which the vehicle feature data and trajectory information collected in steps S1 to S4 are incorporated into the existing model for adaptive learning.
2. The method for real-time identification and labeling of vehicle identity information based on deep learning according to claim 1, characterized in that: In step S1, the edge node is deployed near the camera and sensor to process the video stream and sensor data in real time and perform preliminary vehicle feature extraction and recognition. The sensor data comes from LIDAR, V2X vehicle-to-vehicle communication and GPS.
3. The method for real-time identification and labeling of vehicle identity information based on deep learning according to claim 2, characterized in that: In step S1, the method for performing preliminary vehicle feature extraction and identification is: using a lightweight deep learning model to extract the basic features of the vehicle, generate structured data, which contains the preliminary identification information of the vehicle, license plate, color and model, and send it to the central server.
4. The method for real-time identification and labeling of vehicle identity information based on deep learning according to claim 3 is characterized in that: In step S1, the preliminary vehicle feature extraction and identification method includes: Normalize the original image collected by the camera, assuming that the input image is a three-channel color image The size is H×W, representing the color value corresponding to the pixel coordinates x and y, and normalizing it: in, represents the normalized image pixel value, represents the pixel value of the original image, μ represents the average pixel value of the training data set, and σ represents the pixel standard deviation of the training data set; For the input image Perform convolution calculation: in, represents the activation value of the position (x, y) in the output feature map, K represents the number of convolution kernels, m, n represent the size of the convolution kernel, and the convolution kernel size is (2m+1)×(2n+1). represents the weight of the kth convolution kernel at position (i, j), b k represents the bias of the kth convolution kernel, Represents the normalized pixel value of the input image; The feature map after the convolution layer Passed to the fully connected layer to generate structured data, the fully connected layer is: Among them, z represents the output of the fully connected layer, that is, the extracted basic features of the vehicle. represents the weight matrix of the fully connected layer, b fc represents the bias of the fully connected layer, Indicates that the feature map Flatten to vector; Finally, the feature vector z is converted into structured data and formatted as vehicle identification information: Vehicle_Data = {License_Plate, Color, Model}, where Vehicle_Data represents the structured data of the vehicle, including the license plate number, color and model, License_Plate represents the license plate information recognized by the model, Color represents the color information of the vehicle, and Model represents the model information of the vehicle.
5. The method for real-time identification and labeling of vehicle identity information based on deep learning according to claim 4, characterized in that: In step S2, the detail features identified by the deep learning model include the headlights, vehicle logos and scratches.
6. The method for real-time identification and labeling of vehicle identity information based on deep learning according to claim 5, characterized in that: In step S2, the vehicle re-identification method is: Assume that the preliminary structured data passed from step S1 is D initial , including license plate, color and model: D initial ={L, C, M}, where L represents the license plate recognition result, C represents the vehicle color recognition result, and M represents the vehicle model recognition result; A deep learning model is used to extract vehicle image details. Assuming that the input vehicle image feature is I feat , the feature map after convolution is F conv : Among them, F conv (x, y) represents the activation value of the feature map output by the convolutional layer at position (x, y), F input (x, y) represents the input feature map, that is, the local area features of the vehicle. represents the weight of the kth convolution kernel at position (i, j), K represents the number of convolution kernels, i, j represent the spatial positions of the convolution kernels, and ReLU(·) represents the activation function; The attention mechanism is used to weight the importance of different detail features. Assuming that the input feature map is F conv , generate the attention weight matrix A(x, y) and multiply it with the feature map: Among them, F att (x, y) represents the feature map after attention weighting, and A(x, y) represents the attention weight matrix, which is used to conv Weighted, exp(·) represents the exponential function, which is used to calculate the probability distribution of attention weights, and x′, y′ represent the convolution feature map F conv The coordinates of all pixels in (x′, y′); Flatten the detailed features extracted by convolution and attention mechanism into vector Z detail , and then the high-dimensional features for recognition are generated through the fully connected layer: Z detail = flatten, y = σ(W fc ·z detail +b fc ), where z detail represents the flattened detail feature vector, y represents the output of the fully connected layer, and W fc represents the weight matrix of the fully connected layer, b fc represents the bias of the fully connected layer, σ(·) represents the activation function used for classification, and Softmax and Sigmoid are selected; Combine the detailed features with the preliminary identification information to form complete structured data: final represents the final generated structured data, and y represents the detailed feature recognition result obtained from the fully connected layer.
7. The method for real-time identification and labeling of vehicle identity information based on deep learning according to claim 6, characterized in that: In step S3, the multi-camera collaborative tracking method is: Assume that the spatial position and velocity of the captured vehicle at a certain time t are p t =(x t ,y t ), where p t represents the two-dimensional spatial position of the vehicle at time t, x t ,y t represents the horizontal and vertical coordinates of the vehicle at time t, and the speed of the vehicle is Among them, v t represents the velocity vector of the vehicle at time t, represents the lateral speed and longitudinal speed of the vehicle at time t; Use the prediction model based on historical trajectory and real-time speed for prediction. Assume that the trajectory information of the vehicle between time tk and time t is in, represents the historical trajectory of k time steps before time t, p t-k , ..., p t Represents the spatial position sequence of the vehicle from time tk to t. The Kalman filter is used to smooth and predict the trajectory. The Kalman filter state update equation is: p t+1 =p t +v t ·Δt+w t , where p t+1 represents the predicted position at the next moment t+1, p t represents the position at the current time t, v t represents the speed at the current time t, Δt represents the time step, that is, the predicted time interval, and w t represents the process noise, w t Modeled as zero-mean Gaussian noise: in, represents Gaussian noise with zero mean and covariance Q.
8. The method for real-time identification and labeling of vehicle identity information based on deep learning according to claim 7, characterized in that: In step S3, the multi-camera collaborative tracking method further includes: According to the real-time motion data and predicted trajectory of the vehicle, the trajectory prediction model is used to adjust the prediction results, assuming that the future motion trajectory model of the vehicle is a linear acceleration model: Among them, pt+n represents the predicted vehicle position at the nth time step in the future, a t represents the vehicle acceleration at time t, and n represents the number of time steps predicted in the future; Assume the position of the next camera is c i =(x i ,y i ), the current camera position is c0 = (x0, y0), then the time it takes for the vehicle to travel from the current camera to the next camera is: Among them, T i Indicates the estimated time for the vehicle to travel from the current camera to the next camera, |p t -c i | represents the Euclidean distance between the vehicle’s current position and the next camera position, |v t |Indicates the speed of the vehicle; Combining the historical trajectory, current motion state and environmental information of the vehicle, the Bayesian prediction model is used for global optimization. The Bayesian update formula is: in, represents the probability of predicting the future position of the vehicle based on the historical trajectory, represents the likelihood of the historical trajectory given the future position, P(p t+n ) represents the prior probability of the future position, represents the marginal likelihood of the historical trajectory, According to the predicted time T of the vehicle arriving at the next camera i , notify the next camera node in advance to prepare for recognition.
9. The method for real-time identification and labeling of vehicle identity information based on deep learning according to claim 8, characterized in that: In step S4, the data aggregation and global optimization method is: Each camera C i At time t i The identified vehicle information is: in, Indicates camera C i At time t i Identified vehicle information, F plate represents the vehicle’s license plate feature vector, F color represents the color feature vector of the vehicle, F model Represents the vehicle model feature vector, F detail A detailed feature vector representing the vehicle; Combined with the spatiotemporal trajectory prediction results in step S3, assuming that camera C i At time t i The captured vehicle position is The vehicle's motion trajectory is: in, Indicates the predicted vehicle at the next time t i+1 location, Indicates time t i The vehicle's current location, Indicates time t i The vehicle's velocity vector, Indicates time t i The vehicle's acceleration vector, Δt represents the time t i+1 and t i the time interval between Assume that camera C i and camera C j The identified vehicle information is and The feature similarity is: in, Indicates camera C i and C j For the recognition similarity of the same vehicle, α1, α2, α3, and α4 represent the weight coefficients of each feature, and cos(·, ·) represents the calculation of the cosine similarity between two vectors.
10. The method for real-time identification and labeling of vehicle identity information based on deep learning according to claim 9, characterized in that: In step S4, the data aggregation and global optimization method further includes: Based on the previous spatiotemporal trajectory prediction and similarity, the central server globally annotates the vehicle recognition results under different cameras, and performs weighted graph matching to construct a graph G(V, E) under multiple cameras. Each node represents a vehicle recognition result, and the weight of the edge represents the similarity. The optimization goal is to maximize the global similarity. The optimization objective function is: in, represents the global optimization objective function, Indicates that the vehicle is in front of camera C i and C j The feature similarity under , λ represents the regularization coefficient, Represents the difference in temporal and spatial positions under two cameras; Assign a unique identifier ID to each vehicle based on global optimization k , the final global annotation information is: D global ={ID k , F plate , F color , F model , F detail , p t }, where D global Indicates the vehicle information after global annotation, ID k represents the globally unique identifier assigned to the vehicle, p t Represents the space-time position of the vehicle in the global coordinate system.