A Traffic Police Dynamic Gesture Recognition Method and System Based on In-Vehicle Federated Learning
By deploying key feature extraction models and global models on the on-board terminals, and using on-board federated learning technology, the existing traffic police gesture recognition methods in terms of accuracy, robustness and data utilization are solved, and efficient, safe and highly accurate traffic police gesture recognition is achieved.
Patent Information
- Application Number
- CN202411678360.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-11-22
AI Technical Summary
The existing traffic police gesture recognition methods are insufficient in terms of accuracy and robustness, and have low data utilization, which has data privacy leakage and "data silos".
The traffic police dynamic gesture recognition method based on in-vehicle federated learning is adopted. By deploying key feature extraction models and global models on the in-vehicle terminal, the central server and roadside units are used to initialize, aggregate and distribute model parameters, local model training and global model update are realized.
It improves the accuracy and adaptability of traffic police gesture recognition, protects the data privacy of vehicle users, saves communication costs, and makes full use of the diversity of on-board image data.
Smart Images

Figure CN119600685B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and computer vision, and more particularly, to a traffic police dynamic gesture recognition method and system based on vehicle-mounted federated learning. Background Art
[0002] With the development of autonomous driving and advanced driver assistance technologies, intelligent connected vehicles are required to automatically and real-time recognize the traffic police gestures on the road. The recognition of traffic police dynamic gestures has become a key link to improve driving safety and efficiency. However, the existing data centralized learning method requires collecting a large amount of traffic police gesture images on the road and uniformly transmitting them to the central server for training, which not only causes the burden of data transmission and storage, but also may lead to the risk of privacy leakage of past vehicle users. Traditional methods also fail to fully utilize the traffic police gesture images collected by intelligent connected vehicles through vehicle-mounted cameras at different observation angles and under different environmental conditions, and the accuracy and robustness of the recognition model are low. When using vehicle-mounted image data for model training, it is necessary to consider the data privacy protection requirements of vehicle users. Forcing the collection of vehicle-mounted image data for centralized learning will face the problem of "data islands" in practical applications. Summary of the Invention
[0003] The present invention aims to overcome the defects of low accuracy and low data utilization rate in the recognition of traffic police gestures in the prior art, and provides a traffic police dynamic gesture recognition method and system based on vehicle-mounted federated learning.
[0004] To solve the above technical problems, the technical solution of the present invention is as follows:
[0005] The present invention provides a traffic police dynamic gesture recognition method based on vehicle-mounted federated learning, including:
[0006] Obtaining traffic police gesture images;
[0007] Inputting the traffic police gesture images into a trained key feature extraction model to obtain key nodes and pose vectors in the traffic police gesture images, and constructing a skeleton space feature according to the key nodes and pose vectors;
[0008] Constructing an urban road network scenario, including a central server, several roadside units and several vehicle-mounted terminals; inputting the skeleton space feature into a preset global model, and deploying the key feature extraction model and the global model to all vehicle-mounted terminals;
[0009] Initializing the global model parameters by using the central server, and sending the initialized global model parameters to all roadside units, and the roadside units send them to the selected vehicle-mounted terminals;
[0010] The selected vehicle-mounted terminal performs multiple rounds of local model training based on local traffic police gesture data to update local model parameters. When the preset conditions are met, the updated local model parameters are uploaded to the nearest roadside unit, and then the roadside unit uploads them to the central server;
[0011] The central server aggregates the collected local model parameters to obtain the final global model parameters;
[0012] Based on the final model parameters, the preset global model is trained until the preset number of training times is reached to obtain the final global model;
[0013] The traffic police gestures are recognized according to the final global model.
[0014] Preferably, the key nodes include 14 human key nodes: head, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle; the pose vectors include 3 pose vectors: turning the head, waving the arm, and turning the body.
[0015] Preferably, obtaining a trained key feature extraction model includes:
[0016] Constructing a key feature extraction model, which includes a feature extraction unit and a branch unit connected in sequence;
[0017] Inputting the obtained traffic police gesture image into the feature extraction unit for feature extraction to obtain a key feature map;
[0018] Inputting the key feature map into the branch unit to obtain predicted key nodes and predicted pose vectors;
[0019] Setting a total loss function, calculating the total loss function value according to the predicted key nodes and real key nodes and the predicted pose vectors and real pose vectors, and adjusting the parameters of the constructed key feature extraction model;
[0020] When the total loss function value reaches the preset value or the preset number of training times is reached, a trained key feature extraction model is obtained.
[0021] Preferably, the feature extraction unit includes a first convolutional layer, a second convolutional layer, a first max pooling layer, a third convolutional layer, a fourth convolutional layer, a second max pooling layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a third max pooling layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, and a twelfth convolutional layer connected in sequence; the output end of the twelfth convolutional layer is connected to the input end of the branch unit;
[0022] The branch unit includes a number of stage sub-units and a splicing layer connected in sequence;
[0023] The stage subunit includes an upper branch subunit and a lower branch subunit; the output end of the twelfth convolutional layer is connected to both the input ends of the upper branch subunit and the lower branch subunit; the output ends of the upper branch subunit and the lower branch subunit are both connected to the input end of the splicing layer; the output end of the twelfth convolutional layer is connected to the input end of each splicing layer.
[0024] Preferably, the total loss function includes an upper branch loss function and a lower branch loss function, specifically:
[0025]
[0026] where N is the number of branches; is the upper branch loss function; is the lower branch loss function;
[0027] Preferably, the formula of the upper branch loss function is:
[0028]
[0029] where is the true key node; is the predicted key node; W is a binary mask; j is the serial number of the key node; p is the position of the key node;
[0030] The formula of the lower branch loss function is:
[0031]
[0032] where is the true pose vector; is the predicted pose vector.
[0033] Preferably, obtaining the preset global model includes:
[0034] Construct an initial global model, which includes a first recurrent neural network unit, a first regularization layer, a second recurrent neural network unit, a self-attention layer, a fully connected layer, and a second regularization layer connected in sequence;
[0035] Input the skeleton spatial features into the constructed initial global model to obtain the predicted traffic police gesture category;
[0036] Set the cross-entropy loss function, calculate the value of the cross-entropy loss function according to the predicted traffic police gesture category and the true traffic police gesture category, and adjust the parameters of the constructed initial global model;
[0037] When the value of the cross-entropy loss function reaches the preset value, obtain the preset global model.
[0038] Preferably, the formula of the cross-entropy loss function is as follows:
[0039]
[0040] Wherein, is the predicted probability of the traffic police gesture category; I is the number of skeleton spatial features; y i is the true category of the i-th skeleton spatial feature.
[0041] Preferably, the selected vehicle-mounted terminal performs multiple rounds of local model training on the local traffic police gesture data through an optimization algorithm, and the optimization algorithm is specifically:
[0042]
[0043] Wherein, η is the learning rate; is the momentum term; is the variance term; ε is the gradient parameter;
[0044]
[0045] Wherein, β1 is the decay rate of the momentum term; β2 is the decay rate of the variance term; is the momentum term of the previous time step; is the variance term of the previous time step.
[0046] The present invention also provides a traffic police dynamic gesture recognition system based on vehicle-mounted federated learning for implementing the above method, including:
[0047] An image acquisition module for acquiring traffic police gesture images;
[0048] A skeleton construction module for inputting the traffic police gesture image into a trained key feature extraction model to obtain key nodes and pose vectors in the traffic police gesture image, and constructing skeleton spatial features according to the key nodes and pose vectors;
[0049] A scene construction module for constructing an urban road network scene, including a central server, a plurality of roadside units, and a plurality of vehicle-mounted terminals; inputting the skeleton spatial features into a preset global model, and deploying the key feature extraction model and the global model to all vehicle-mounted terminals;
[0050] A parameter initialization module for initializing the global model parameters by using the central server, sending the initialized global model parameters to all roadside units, and the roadside units distributing them to the selected vehicle-mounted terminals;
[0051] A local training module, where the selected vehicle-mounted terminal performs multiple rounds of local model training based on local traffic police gesture data to update local model parameters. When a preset condition is reached, the updated local model parameters are uploaded to the nearest roadside unit, and then the roadside unit uploads them to the central server;
[0052] An aggregation module, where the central server aggregates the collected local model parameters to obtain the final global model parameters;
[0053] A global training module, which trains a preset global model based on the final model parameters until a preset number of training times is reached to obtain the final global model;
[0054] A gesture recognition module, which recognizes traffic police gestures according to the final global model.
[0055] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0056] The present invention proposes a traffic police dynamic gesture recognition method and system based on vehicle-mounted federated learning, which combines a key feature extraction model and a preset global model as a traffic police gesture recognition model shared by all vehicle-mounted terminals in the network, helping to improve the accuracy of traffic police gesture recognition; through vehicle-mounted federated learning, different vehicle-mounted terminals are organized to train the above models based on the traffic police gesture image data collected by themselves. Subsequently, each vehicle-mounted terminal only shares model update data instead of original training data, protecting the data privacy of vehicle users and saving the communication cost of uploading massive data to the central server. At the same time, vehicle-mounted federated learning makes full use of the real-time traffic police gesture image data scattered in different vehicle-mounted devices. These training data from vehicle-mounted images have different observation angles, different environmental conditions, etc., enriching the diversity of model training data, and thus enhancing the adaptability of the trained traffic police gesture recognition model in a changing traffic environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flowchart of the traffic police dynamic gesture recognition method based on vehicle-mounted federated learning described in Embodiment 1;
[0058] Figure 2 It is a structural schematic diagram of the key node positions described in Embodiment 2;
[0059] Figure 3 It is a structural schematic diagram of the feature extraction model described in Embodiment 2;
[0060] Figure 4 It is a structural schematic diagram of the SA-LSTM network described in Embodiment 2;
[0061] Figure 5 It is a structural schematic diagram of the vehicle-mounted federated learning system described in Embodiment 2;
[0062] Figure 6 Schematic diagram of the loss change and training accuracy during the training of the model described in Embodiment 3;
[0063] Figure 7 Schematic diagram of the structure of the traffic police dynamic gesture recognition system based on vehicle-mounted federated learning described in Embodiment 3. Detailed implementation manners
[0064] The accompanying drawings are only for illustrative purposes and should not be construed as limitations on this patent;
[0065] To better illustrate this embodiment, some components in the accompanying drawings will be omitted, enlarged or reduced, which do not represent the dimensions of the actual product;
[0066] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0067] The technical solutions of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0068] Embodiment 1
[0069] This embodiment provides a traffic police dynamic gesture recognition method based on vehicle-mounted federated learning, as Figure 1 shown, including:
[0070] Obtain traffic police gesture images;
[0071] Input the traffic police gesture images into a trained key feature extraction model to obtain key nodes and pose vectors in the traffic police gesture images, and construct a skeleton space feature according to the key nodes and pose vectors;
[0072] Construct a city road network scenario, including a central server, a number of roadside units and a number of vehicle-mounted terminals; input the skeleton space feature into a preset global model, and deploy the key feature extraction model and the global model to all vehicle-mounted terminals;
[0073] Use the central server to initialize the global model parameters, and send the initialized global model parameters to all roadside units, which are then sent by the roadside units to the selected vehicle-mounted terminals;
[0074] The selected vehicle-mounted terminals perform multiple rounds of local model training according to local traffic police gesture data to update local model parameters. When the preset conditions are met, the updated local model parameters are uploaded to the nearest roadside unit, and then uploaded by the roadside unit to the central server;
[0075] The central server aggregates the collected local model parameters to obtain the final global model parameters;
[0076] Train the preset global model based on the final model parameters until the preset number of training times is reached to obtain the final global model;
[0077] Identify traffic police gestures according to the final global model.
[0078] In this embodiment, the key feature extraction model and the preset global model are combined as the traffic police gesture recognition model shared by all in-vehicle terminals in the network, which helps to improve the accuracy of traffic police gesture recognition; through in-vehicle federated learning, different in-vehicle terminals are organized to train the above models based on the traffic police gesture image data collected by themselves. Subsequently, each in-vehicle terminal only shares the model update data instead of the original training data, protecting the data privacy of vehicle users and saving the communication cost of uploading massive data to the central server. At the same time, in-vehicle federated learning makes full use of the real-time traffic police gesture image data scattered in different in-vehicle devices. These training data from in-vehicle images have different observation angles, different environmental conditions, etc., enriching the diversity of model training data, and then enhancing the adaptability of the trained traffic police gesture recognition model in a changing traffic environment.
[0079] Embodiment 2
[0080] This embodiment provides a traffic police dynamic gesture recognition method based on in-vehicle federated learning, including:
[0081] Obtain traffic police gesture images;
[0082] Input the traffic police gesture image into the trained key feature extraction model to obtain the key nodes and pose vectors in the traffic police gesture image, and construct a skeleton space feature according to the key nodes and pose vectors;
[0083] Construct an urban road network scenario, including a central server, several roadside units and several in-vehicle terminals; input the skeleton space feature into the preset global model, and deploy the key feature extraction model and the global model to all in-vehicle terminals;
[0084] Use the central server to initialize the global model parameters, and send the initialized global model parameters to all roadside units, which are then sent by the roadside units to the selected in-vehicle terminals;
[0085] The selected in-vehicle terminal performs multiple rounds of local model training according to local traffic police gesture data to update the local model parameters. When the preset conditions are met, the updated local model parameters are uploaded to the nearest roadside unit, and then uploaded by the roadside unit to the central server;
[0086] The central server aggregates the collected local model parameters to obtain the final global model parameters;
[0087] Train the preset global model based on the final model parameters until the preset number of training times is reached to obtain the final global model;
[0088] Recognize traffic police gestures according to the final global model.
[0089] As Figure 2 shown, the key nodes include 14 human key nodes: head, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle; the pose vectors include 3 pose vectors: turning the head, waving the arm, and turning the body.
[0090] Obtain a trained key feature extraction model, including:
[0091] Construct a key feature extraction model, which includes a feature extraction unit and a branch unit connected in sequence;
[0092] Input the obtained traffic police gesture image into the feature extraction unit for feature extraction to obtain a key feature map;
[0093] Input the key feature map into the branch unit to obtain predicted key nodes and predicted pose vectors;
[0094] Set the total loss function, calculate the total loss function value according to the predicted key nodes and real key nodes and the predicted pose vectors and real pose vectors, and adjust the parameters of the constructed key feature extraction model;
[0095] When the total loss function value reaches the preset value or reaches the preset number of training times, obtain the trained key feature extraction model.
[0096] As Figure 3 shown, the feature extraction unit includes a first convolutional layer, a second convolutional layer, a first max pooling layer, a third convolutional layer, a fourth convolutional layer, a second max pooling layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, an eighth convolutional layer, a third max pooling layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, and a twelfth convolutional layer connected in sequence; the output end of the twelfth convolutional layer is connected to the input end of the branch unit;
[0097] The branch unit includes a number of stage sub-units and a splicing layer connected in sequence;
[0098] The stage sub-unit includes an upper branch sub-unit and a lower branch sub-unit; the output end of the twelfth convolutional layer is connected to the input ends of the upper branch sub-unit and the lower branch sub-unit; the output ends of the upper branch sub-unit and the lower branch sub-unit are connected to the input end of the splicing layer; the output end of the twelfth convolutional layer is connected to the input end of each splicing layer.
[0099] The total loss function includes an upper-branch loss function and a lower-branch loss function, specifically:
[0100]
[0101] where N is the number of branches; is the upper-branch loss function; is the lower-branch loss function;
[0102] The formula for the upper-branch loss function is:
[0103]
[0104] where is the true key node; is the predicted key node; W is a binary mask; j is the serial number of the key node; p is the position of the key node;
[0105] The formula for the lower-branch loss function is:
[0106]
[0107] where is the true pose vector; is the predicted pose vector.
[0108] Obtaining the preset global model includes:
[0109] Constructing an initial global model, which includes a first recurrent neural network unit, a first regularization layer, a second recurrent neural network unit, a self-attention layer, a fully connected layer, and a second regularization layer connected in sequence;
[0110] Inputting the skeleton spatial features into the constructed initial global model to obtain the predicted traffic police gesture category;
[0111] Setting a cross-entropy loss function, calculating the value of the cross-entropy loss function according to the predicted traffic police gesture category and the true traffic police gesture category, and adjusting the parameters of the constructed initial global model;
[0112] When the value of the cross-entropy loss function reaches the preset value, the preset global model is obtained.
[0113] Preferably, the formula for the cross-entropy loss function is:
[0114]
[0115] where is the predicted probability of the traffic police gesture category; I is the number of skeleton spatial features; y i is the true category of the i-th skeleton spatial feature.
[0116] The selected vehicle-mounted terminal performs multiple rounds of local model training on the basis of local traffic police gesture data through an optimization algorithm. The specific optimization algorithm is as follows:
[0117]
[0118] where η is the learning rate; is the momentum term; is the variance term; ε is the gradient parameter;
[0119]
[0120] where β1 is the decay rate of the momentum term; β2 is the decay rate of the variance term; is the momentum term at the previous time step; is the variance term at the previous time step.
[0121] In a specific embodiment, the traffic police gesture modeling based on key points and skeletons in step 1 includes the following: For vehicles, the traffic police gesture images collected come from the vehicle-mounted video dataset, which is generally from the angle of a single camera (i.e., one direction of the traffic police). Therefore, the present invention models the key nodes and skeleton vectors of the traffic police in 2D form. Considering that the 8 standard gesture signals of Chinese traffic police include various posture actions such as turning the head, waving the arm, waving the hand, and turning around, the key nodes include the head X1, neck X2, left shoulder X3, right shoulder X4, left elbow X5, right elbow X6, left wrist X7, right wrist X8, left hip X9, right hip X 10 , left knee X 11 , right knee X 12 , left ankle X 13 and right ankle X 14 a total of 14 human key nodes. The modeling of the skeleton vector is based on the pairwise links between the key nodes related to the traffic police posture and can be composed of the following parts:
[0122] The posture vector part of turning the head:
[0123] The posture vector part of waving the arm and waving the hand:
[0124] The posture vector part of turning around:
[0125] Furthermore, in step 2, a spatial feature extraction model is used to extract the key nodes of the traffic police gesture, and then the skeleton spatial features are calculated. The specific steps include the following:
[0126] S1. Construct a partial affinity field network model for the key nodes of traffic police gestures. This model takes a color image of size w×h as input and, after passing through a network structure with multiple stages and two branches, produces the two-dimensional positions of the human joint points in the image as output.
[0127] Specifically, the implementation process of the model for spatial feature extraction is as follows:
[0128] (1) First, perform preliminary feature extraction on the image through a convolutional neural network (initialized with the first 10 layers of VGG-19 and fine-tuned as the backbone network) to generate a set of feature maps F. Then, input the feature maps F into the first stage of two branches (the upper branch is used to predict key nodes; the lower branch is used to predict the associations between nodes, i.e., the skeleton).
[0129] (2) In the first stage, after passing through each branch network, output a set of confidence maps E of key nodes 1 =α 1 (F) and a set of partial affinity fields L 1 =β 1 (F). Among them, α 1 and β 1 are the upper branch network and the lower branch network for inference in the first stage, respectively; the confidence map of key nodes is a heat map used to represent the confidence of a specific key point at each pixel position in the image. If Y j ∈R 2 represents the true position of the key node part j of the human body, then the value at position p∈R 2 in the confidence map of key nodes (PCM) can be defined as:
[0130]
[0131] In this formula, σ controls the propagation of the peak. The formula is essentially a Gaussian distribution centered on the key point, and its value decreases rapidly as the distance from the center of the key point increases. In this way, the PCM can generate a clear peak response for each key point, thus helping the model accurately locate the position of the key point. The part affinity field is a vector field composed of multiple two-dimensional vectors with different confidences, which encodes the degree of association between parts. For each skeleton link c, its PAFs vector field can be expressed as:
[0132]
[0133] where v is the unit vector of the skeleton link and p is the pixel point in the image. The calculation formula for the unit vector is:
[0134]
[0135] (3) In the subsequent N stages (N > 1), in each stage n, the predictions E n-1 and L n-1 of the two branches in the previous stage are first connected to the original image feature F, and then input into the two branch networks α n and β n of the current stage, so as to generate more accurate predictions by analyzing the features of the image under different receptive fields. The output results E n and L n in each stage are expressed as follows:
[0136]
[0137] (4) During the training process of the model, two loss functions need to be applied at the end of each stage of the network, one for each branch. The L2 loss is used between the estimated prediction and the ground truth map and field. In this loss function, the intermediate supervision method is used at each stage to solve the problem of vanishing gradients by periodically supplementing gradients, and the loss function is weighted spatially. Specifically, the loss functions of the upper and lower branches in the nth stage are:
[0138]
[0139] where, is the ground truth of the confidence map of the key node, is the ground truth of the part affinity field, W is a binary mask when there is no annotation at the image position p, that is, W(p) = 0, and W is used to avoid penalizing TF (true positive) predictions during training. The overall objective loss function is the sum of the loss values of the two branches in all stages, and its definition is:
[0140]
[0141] (5) When outputting the final prediction result, it is necessary to obtain the candidate node with the highest confidence through non-maximum suppression (NMS) and use it as the finally determined key node That is
[0142] Specifically, in the present invention, the partial affinity fields (PAFs) network model is mainly used to extract key nodes from traffic police dynamic gesture data. However, this model is not trained using traffic police dynamic gesture data, but rather uses the open-source AIChallenger's keypoints dataset for training. This dataset contains 300,000 images and corresponding label files, and the label information includes key node identifiers and coordinates, skeleton link relationships, etc. When training this model, the parameters are set as follows: the batch size is 30, the initial learning rate is 0.001, the Adam optimizer is used, and the number of iterations is 100. The optimal model is saved when the training is completed.
[0143] S2. After completing the training of this model, the trained model is used to extract traffic police gesture key nodes and the association relationships between nodes. However, the association relationships between nodes do not use the final output of the lower branch of this network model, but rather use the coordinates of the output key nodes to calculate more accurate skeleton link relationships through formulas. Specifically, it is necessary to further obtain two spatial feature information items, namely the length of the skeleton link (the distance between key nodes) and the direction of the skeleton link (the unit vector connected by key nodes), from the key node coordinates. The implementation process is as follows:
[0144] (1) For any skeleton link c, the length of the skeleton link can be calculated using the Euclidean distance formula:
[0145]
[0146] where (u1, v1) and (u2, v2) are the position coordinates of the two key nodes X j1 (p) and X j2 (p) of the skeleton link c.
[0147] (2) Select the direction of gravity as the reference direction. Then, the direction of the skeleton link can be defined as the angle between it and the direction of gravity, and this angle can be uniquely determined by calculating the sine value and cosine value of this angle. The calculation formulas for the sine value and cosine value of this angle are as follows:
[0148]
[0149] where g is the unit vector with the same direction as the direction of gravity, and R c represents the direction vector of the skeleton link c.
[0150] Then, the direction of the skeleton link can be determined by vector splicing using the sine value and cosine value of this angle, that is:
[0151]
[0152] (3) For any skeleton link c, the length of the skeleton link is The direction of the skeleton link is Then these two pieces of spatial feature information can be spliced and integrated into the human body spatial feature information S c , which is convenient for further extraction of spatio-temporal features in the follow-up. Its expression is:
[0153]
[0154] Specifically, the trained Part Affinity Fields (PAFs) network model is used to extract the key nodes of each frame of the processed traffic police dynamic gesture video data or real-time data, and then the output of the key node coordinates is combined to calculate the skeleton length, the sine value and cosine value of the angle between the skeleton and the gravity direction of each frame of data. After collecting the spatial feature data of each processed frame in real time, it is input into the subsequent SA-LSTM model under the vehicle-mounted federated framework in the form of a time series for local model training and model update, so as to realize the recognition of traffic police dynamic gestures.
[0155] Step 3: Combine the Self-Attention based Long Short-Term Memory Network (SA-LSTM) for spatio-temporal feature extraction. After the key nodes of the traffic police gestures are extracted by the Part Affinity Fields (PAFs) network model in Step 2 and the skeleton spatial features are calculated, the spatial feature data S of each frame of the traffic police dynamic gesture video data can be obtained c , and these sequence data are used as part of the input of the SA-LSTM model after subsequent data processing, and the final gesture classification result is obtained after network layer inference. The SA-LSTM network consists of a double-layer LSTM, a self-attention layer and an output layer, and its implementation process for spatio-temporal feature extraction is as follows:
[0156] S1. The Long Short-Term Memory Network (LSTM) is a special type of recurrent neural network, which controls the flow of temporal information by introducing three gating mechanisms: the forget gate, the input gate and the output gate. Given an input sequence X = {x t}, where x t is the input at time step t, the LSTM unit calculates the following in the forward propagation at each time step t:
[0157] ① Input gate: It consists of the update gate i t and the candidate cell state These two parts. The update gate i t determines the degree to which new information will be stored in the cell state, and the candidate cell state represents the new information itself. The specific calculation formula is as follows:
[0158] it = σ(W i · [h t-1 , x t + b i )
[0159]
[0160] where i t is the output of the input gate, is the candidate cell state, W i and W C are the weights of the input gate and the candidate cell state respectively, b i and b C are the corresponding bias terms respectively, tanh represents the hyperbolic tangent function, and the symbol · is usually used to denote the dot product, that is, the inner product.
[0161] ② Forget gate: It looks at the current input x t and the hidden state h t-1 at the previous time step, and outputs a value between 0 and 1 through the sigmoid function to represent the degree of forgetting of each unit in the cell state, thus determining which information should be discarded from the cell state. The specific calculation formula is as follows:
[0162] f t = σ(W i · [h t-1 , x t + b f )
[0163] where σ is the sigmoid activation function, W f and b f are the weight and bias term of the forget gate respectively.
[0164] ③ Update cell state: The cell state C t is the core of the LSTM unit, which is responsible for storing information in the sequence. At each time step, the cell state is updated through the output information of the forget gate and the update gate. The implementation expression is as follows:
[0165]
[0166] where C t is the cell state at the current time step, C t-1 is the cell state at the previous time step, and the symbol * represents the element-wise multiplication (Hadamard product). Element-wise multiplication means that the elements at the corresponding positions of two tensors or matrices with the same shape are multiplied to obtain a new tensor or matrix with the same shape as the original tensor or matrix.
[0167] ④ Output gate: It uses the sigmoid function to output a value o between 0 and 1 t , indicating the contribution degree of each unit in the cell state to the hidden state. At the same time, it also updates the hidden state h t , and passes it to the next time step. The relevant calculation formula is as follows:
[0168] h t = σ(W o · [h t-1 , x t + b o )
[0169] h t = o t * tanh(C t )
[0170] Among them, o t is the output of the output gate, W o is the output gate weight, and b o is the output gate bias term.
[0171] As Figure 4 shown, in the LSTM network part, the input is the spatial feature data in each frame after data processing and the initial hidden state h t and the cell state C t . The cascade of two LSTM units is used (and a dropout regularization layer is used between the two LSTM layers to prevent or reduce the overfitting of the model) to first capture the local features of the time series (through LSTM), and finally the hidden state h t+2 , the cell state C t+2 and the LSTM output (including the hidden state output of each time step in the sequence) are output, that is N is the time step length of the sequence.
[0172] S2. The self-attention mechanism can dynamically adjust the attention degree to different time steps according to the content of the input sequence, which helps to capture the long-range dependencies in the gesture sequence. Compared with LSTM, which can only perform sequential learning and local feature extraction on the time series, the self-attention mechanism can provide a summary of global information for each time step by generating context vectors, thus establishing global dependencies. In this SA-LSTM model, after the LSTM extracts the temporal features of the gesture, the constructed self-attention layer is used on the output o t+2Capture the relationships between the LSTM hidden states above, which can enhance the attention to key time steps and help the model consider the information of the entire time series when making the final decision, thus improving the accuracy of gesture recognition. The implementation steps of this self-attention mechanism are as follows:
[0173] ① Apply a linear transformation to the input X of the self-attention layer (actually o t+2 ) to generate the query (Q), key (K), and value (V), which allows the model to learn different representations of the time series. In the self-attention mechanism, the query (Q), key (K), and value (V) are derived from the product of the input matrix X and three biasable weight matrices A q 、A k and A v respectively. It is implemented by three linear layers L q 、L k and L v in code, and its formula is as follows:
[0174]
[0175] Through the weight matrices, the complex relationships between the input features can be learned and mapped into new feature representations, so that the self-attention mechanism can more effectively capture the correlations between different time steps in the sequence.
[0176] ② Calculate the attention weight W A : First, calculate the dot product of the query (Q) and the key (K) to obtain the preliminary attention scores. Then, use a scaling factor (usually the reciprocal of the square root of the key vector dimension) to scale the dot product result to prevent the dot product result from being too large. The scaled attention scores are then converted into a probability distribution through the Softmax function, so that the sum of the attention weights for each time step in the time series is 1. The specific calculation formula is as follows:
[0177]
[0178] where d k is the dimension of the key vector.
[0179] ③ Use the obtained attention weights to perform a weighted sum on the value (V) to obtain the final attention output Att.
[0180] The formula is as follows:
[0181]
[0182] where T is the sequence length of the input time series (in frames).
[0183] S3. After passing through the self-attention layer, the above attention output Att is then passed through a fully connected layer and a dropout regularization layer, and finally a two-dimensional tensor is output. (i is the serial number of the sequence sample, and there are I sequence samples in total, that is, N = the number of time steps * the number of batches; f is the class serial number of the sequence, and there are 9 classes in total), which contains the prediction probabilities of 9 types of gestures (8 standard traffic police gestures and the case of no gesture) at each time step.
[0184] S4. Loss function calculation: First, perform the Softmax function calculation on the output to achieve normalization. For each sequence sample n and each class f, the calculation formula is as follows:
[0185]
[0186] where Z f is the predicted probability of class f, is the probability of class f in the i-th sample after the Softmax transformation. Then, the cross-entropy loss function is used for calculation, and the expression is as follows:
[0187]
[0188] where y i is the actual class index of the i-th sample. Then, the losses of all samples are averaged to obtain the final loss value.
[0189] Specifically, when the SA-LSTM network model is used for local model training in each vehicle-mounted terminal under the vehicle-mounted federated framework, the model parameters are set as follows: the size of the LSTM hidden layer is 64, and the coefficient of the dropout regularization layer is 0.5. The specific federated parameters will be described in the following steps.
[0190] Step 4. Use the federated learning framework based on the FedAvg algorithm to implement distributed learning of each vehicle-mounted terminal in the vehicle-mounted environment. The following specific steps are included in this step:
[0191] S1. Construct the vehicle-mounted federated learning system framework. As Figure 5 shown, in the urban road network scenario, assume that there are H traffic police officers conducting real-time traffic command at each congested urban intersection, and the entities include a central server, N roadside units (RSUs), and K vehicle-mounted terminals. Each vehicle-mounted client uses the traffic police gesture image data collected by itself for local model training and uploads the model gradients. The central server is responsible for aggregating the model gradients of all vehicle-mounted terminals and sending the updated global model parameters back to each vehicle-mounted terminal. The roadside unit serves as a communication relay between the vehicle-mounted terminal and the central server. Usually, the vehicle-mounted terminal performs wireless data transmission with the roadside unit, and the roadside unit performs wireless or wired data transmission with the central server.
[0192] S2. Vehicle terminal and global model initialization: Deploy the part affinity field network model that has been trained with the human key node dataset and the SA-LSTM model to be trained on K vehicle terminals participating in vehicle-mounted federated learning. In the present invention, only the SA-LSTM model is used as the global model and needs to be trained distributively in vehicle-mounted federated learning. The central server initializes the global model parameters ω0 and transmits the initialized global model parameters ω0 to each roadside unit, and then each roadside unit broadcasts the model parameters ω0 to the nearby vehicle terminals that have been confirmed to participate in vehicle-mounted federated learning, from which we can obtain
[0193] S3. Selection of vehicle terminals in each round of global model training: Given that the number of set vehicle terminals is K and the proportion of vehicle terminals selected to participate in local model training is ρ. In each round of federated training, the central server calculates the number m of selected clients according to the following formula:
[0194] m = max(Kρ, 1)
[0195] Then randomly select m vehicle terminals to form the vehicle terminal set S τ , and let the m vehicle terminals carry out local model training in parallel. Among them, τ is the current global model training round.
[0196] S4. Local model training: The vehicle terminals selected in each round of global model training use the traffic police dynamic gesture data collected locally by themselves to train the SA-LSTM model. To reduce communication costs and ensure that each vehicle terminal can fully learn the characteristics of its local data, the vehicle terminals selected in each round of global model training will perform E rounds of local model training and update the local model parameters in each local model training. In each round of local model training, after forward propagation and loss calculation, the model parameters of each vehicle terminal are optimized by gradient descent using the Adam optimization algorithm to minimize the loss function. In this algorithm, first calculate the gradient at the current time step t in the backpropagation, and the calculation formula is as follows:
[0197]
[0198] Among them, are the local model parameters at the current time step t, L k is the loss function, represents the gradient operation. Then, perform gradient descent optimization on the model parameters of each vehicle terminal, and this optimization algorithm can be expressed by the following formula:
[0199]
[0200] Among them, η is the global learning rate, is the first-order moment estimate of the gradient (i.e., the momentum term), is the second-order moment estimate of the gradient (i.e., the variance term), ε is a very small number used to prevent the denominator from being zero, and The calculation formulas of
[0201]
[0202] are as follows: Among them, β1 is the decay rate of the first-order moment estimate, and β2 is the decay rate of the second-order moment estimate. is the first-order moment estimate of the previous time step, is the second-order moment estimate of the previous time step.
[0203] S5. Upload of local model parameters: After the local model training and model update are completed for the selected vehicle-mounted terminals in each round of global model training, the updated model parameters are uploaded to the nearest roadside unit, and it uploads to the central server.
[0204] S6. Global model aggregation and model distribution:
[0205] The central server collects the local model parameters uploaded by all vehicle-mounted terminals and uses the Federated Averaging (FedAvg) algorithm to perform parameter aggregation to obtain the global model parameters ω of the (τ + 1)-th round τ+1 , as follows:
[0206]
[0207] Among them, θ k is the local training data volume of vehicle-mounted terminal k and The central server enters the global model training of the (τ + 1)-th round based on the updated global model data ω τ+1 to obtain the latest global parameters ω τ+1 , and transmits and downloads the updated global parameters ω τ+1 to all vehicle-mounted terminals again through the roadside unit.
[0208] S7. Iterative optimization: Repeat steps S3 to S6 until the model converges or reaches the preset global iteration rounds.
[0209] Regarding the respective traffic police dynamic gesture data collected locally by the vehicle-mounted terminals in step S4, in the actual vehicle-mounted scenario, for the online federated learning of the real-time unlabeled data stream collected by using the vehicle camera, there are the following several solutions:
[0210] (1) Use semi-supervised learning algorithms, combined with data stream processing and window techniques that can adapt to continuous data input. By training on both labeled and unlabeled data simultaneously on each vehicle terminal and fusing the model parameters of each node on the central server, the optimization of the global model is completed, thus realizing online federated learning.
[0211] (2) Use unsupervised learning algorithms (such as clustering algorithms) to process unlabeled data. Such algorithms do not require labels to learn the internal structure and patterns of the data. Through knowledge distillation and contrast learning on both sides, they learn unbiased representations in decentralized and heterogeneous local data, and can realize the federated learning framework FedX combined with unsupervised learning.
[0212] (3) Use self-supervised learning algorithms. By constructing tasks, the vehicle terminal uses the unlabeled video data stream to self-generate pseudo-labels, thereby learning the internal structure and patterns of the data. Then, use the constructed self-supervised tasks to train the model on each vehicle terminal, and complete federated aggregation and global model fine-tuning on the server side.
[0213] For the requirement of the actual vehicle scenario of online federated learning using the real-time unlabeled data stream collected by the vehicle camera, methods including the above methods or real-time automatic data annotation methods on the vehicle terminal or other methods can be used. Since the present invention does not involve the specific method of realizing online federated learning using the real-time unlabeled data stream collected by the vehicle camera, the implementation method is not limited, and any publicly available method in any existing technology can be freely adopted.
[0214] Specifically, during actual training in the present invention, the federated parameters in the federated framework are set as follows: the number of vehicle terminals K is 10, the proportion ρ of the vehicle terminals selected to participate in local model training is 0.7, the global iteration number is 10000, the number of local model training times E in each round of global iteration is 5, the global learning rate is 0.001, and the batch processing number is 1.
[0215] Generally speaking, in the task of traffic police dynamic gesture recognition under the framework of vehicle-mounted federated learning, first, use the trained part affinity field network model deployed on the vehicle-mounted terminal, combined with the video data collected by the vehicle-mounted terminal, to accurately extract the key node information of the traffic police gesture and calculate the skeleton spatial features; then, import the time series data containing skeleton spatial features extracted and calculated from the traffic police dynamic gesture data into the SA-LSTM model deployed on the vehicle-mounted terminal for collaborative training and joint optimization. The optimization process aims to minimize the loss function by adjusting the global model parameters, thereby improving the prediction accuracy; finally, use the trained SA-LSTM global model, combined with the time series data containing skeleton spatial features extracted from the real-time data of traffic police dynamic gestures by the part affinity field network model, to predict the probability of traffic police dynamic gesture categories, and use the visualization method to take the category with the highest probability as the recognition output of traffic police dynamic gestures.
[0216] To verify and compare the model performance of the present invention, tests were carried out on the publicly available Chinese Traffic Police Gesture Dataset (CTPGD). When centralized training was carried out without using the federated framework, the accuracy of the model in recognizing traffic police dynamic gestures in the evaluation test was 96.11% (in the form of edit distance), which was nearly 4% higher than the accuracy of the traditional CPM+LSTM scheme; while when distributed training was carried out using the federated framework, the loss change and training accuracy of the model during training were as Figure 5 shown. The global model training converged well, and the accuracy of the model in recognizing traffic police dynamic gestures in the evaluation test was 95.78% (in the form of edit distance), achieving the effect that the model performance was close to that of centralized training while protecting the data privacy of vehicle users. The present invention can recognize traffic police dynamic gestures from real-time video streams. As Figure 6 shown, it is the effect diagram of one frame in real-time recognition.
[0217] Embodiment 3
[0218] This embodiment also provides a traffic police dynamic gesture recognition system based on vehicle-mounted federated learning, as Figure 7 shown, for implementing the method of Embodiment 1 or Embodiment 2, including:
[0219] An image acquisition module, which acquires traffic police gesture images;
[0220] A skeleton construction module, which inputs the traffic police gesture image into the trained key feature extraction model to obtain the key nodes and pose vectors in the traffic police gesture image, and constructs skeleton spatial features according to the key nodes and pose vectors;
[0221] Scenario construction module, which constructs an urban road network scenario, including a central server, a number of roadside units, and a number of in-vehicle terminals; inputs the skeletal spatial features into a preset global model, and deploys the key feature extraction model and the global model to all in-vehicle terminals;
[0222] Parameter initialization module, which initializes the global model parameters using the central server, and sends the initialized global model parameters to all roadside units, which then distribute them to the selected in-vehicle terminals;
[0223] Local training module, the selected in-vehicle terminals perform multiple rounds of local model training based on local traffic police gesture data to update the local model parameters. When the preset conditions are met, the updated local model parameters are uploaded to the nearest roadside unit, which then uploads them to the central server;
[0224] Aggregation module, the central server aggregates the collected local model parameters to obtain the final global model parameters;
[0225] Global training module, trains the preset global model based on the final model parameters until the preset number of training times is reached to obtain the final global model;
[0226] Gesture recognition module, which recognizes traffic police gestures based on the final global model.
[0227] Identical or similar reference numerals correspond to identical or similar components;
[0228] The terms used to describe the positional relationship in the drawings are for illustrative purposes only and should not be construed as a limitation of this patent;
[0229] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the claims of the present invention.
Claims
1. A traffic police dynamic gesture recognition method based on vehicle-mounted federated learning, characterized in that: include: Get traffic police gesture image; Inputting the traffic police gesture image into a trained key feature extraction model to obtain key nodes and posture vectors in the traffic police gesture image, and constructing skeleton space features according to the key nodes and posture vectors; Constructing an urban road network scenario, including a central server, several roadside units and several vehicle-mounted terminals; inputting skeleton spatial features into a preset global model, and deploying the key feature extraction model and the global model to all vehicle-mounted terminals; Initializing the global model parameters using a central server, and sending the initialized global model parameters to all roadside units, which are then sent by the roadside units to the selected vehicle terminals; The selected vehicle-mounted terminal performs multiple rounds of local model training according to the local traffic police gesture data to update the local model parameters, and when the preset conditions are met, the updated local model parameters are uploaded to the nearest roadside unit, which is then uploaded to the central server by the roadside unit; The central server aggregates the collected local model parameters to obtain the final global model parameters; The preset global model is trained based on the final model parameters until the preset number of training times is reached to obtain the final global model; Traffic police gestures are recognized based on the final global model.
2. The traffic police dynamic gesture recognition method based on vehicle-mounted federated learning according to claim 1 is characterized in that: The key nodes include head, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle and right ankle, a total of 14 key nodes of the human body; the posture vectors include turning the head, waving arms, and turning around, a total of 3 posture vectors.
3. The method for traffic police dynamic gesture recognition based on vehicle-mounted federated learning according to claim 1 is characterized in that: Get the trained key feature extraction model, including: Constructing a key feature extraction model, wherein the key feature extraction model includes a feature extraction unit and a branch unit connected in sequence; The acquired traffic police gesture image is input into the feature extraction unit for feature extraction to obtain a key feature map; Input the key feature graph into the branch unit to obtain the predicted key nodes and predicted posture vector; Set the total loss function, calculate the total loss function value according to the predicted key nodes and the real key nodes and the predicted posture vector and the real posture vector, and adjust the parameters of the constructed key feature extraction model; When the total loss function value reaches a preset value or reaches a preset number of training times, a trained key feature extraction model is obtained.
4. The method for traffic police dynamic gesture recognition based on vehicle-mounted federated learning according to claim 3 is characterized in that: The feature extraction unit includes a first convolution layer, a second convolution layer, a first maximum pooling layer, a third convolution layer, a fourth convolution layer, a second maximum pooling layer, a fifth convolution layer, a sixth convolution layer, a seventh convolution layer, an eighth convolution layer, a third maximum pooling layer, a ninth convolution layer, a tenth convolution layer, an eleventh convolution layer, and a twelfth convolution layer, which are sequentially connected; the output end of the twelfth convolution layer is connected to the input end of the branch unit; The branch unit includes a plurality of sequentially connected stage subunits and splicing layers; The stage subunit includes an upper branch subunit and a lower branch subunit; the output end of the twelfth convolutional layer is connected to the input ends of the upper branch subunit and the lower branch subunit; the output ends of the upper branch subunit and the lower branch subunit are connected to the input end of the splicing layer; the output end of the twelfth convolutional layer is connected to the input end of each splicing layer.
5. The method for traffic police dynamic gesture recognition based on vehicle-mounted federated learning according to claim 3 is characterized in that: The total loss function includes an upper branch loss function and a lower branch loss function, specifically: Where N is the number of branches; is the upper branch loss function; is the loss function of the lower branch.
6. The method for traffic police dynamic gesture recognition based on vehicle-mounted federated learning according to claim 5 is characterized in that: The formula of the upper branch loss function is: in, is the real key node; is the predicted key node; W is the binary mask; j is the sequence number of the key node; p is the position of the key node; The formula of the lower branch loss function is: in, is the true posture vector; is the predicted pose vector.
7. The method for traffic police dynamic gesture recognition based on vehicle-mounted federated learning according to claim 1 is characterized in that: Obtaining the preset global model includes: Constructing an initial global model, wherein the initial global model includes a first recurrent neural network unit, a first regularization layer, a second recurrent neural network unit, a self-attention layer, a fully connected layer, and a second regularization layer connected in sequence; Input the skeleton space features into the constructed initial global model to obtain the predicted traffic police gesture category; Set the cross entropy loss function, calculate the cross entropy loss function value according to the predicted traffic police gesture category and the real traffic police gesture category, and adjust the parameters of the constructed initial global model; When the cross entropy loss function value reaches the preset value, the preset global model is obtained.
8. The method for traffic police dynamic gesture recognition based on vehicle-mounted federated learning according to claim 7 is characterized in that: The formula of the cross entropy loss function is: in, is the predicted probability of the traffic police gesture category; I is the number of skeleton space features; y i is the true category of the i-th skeleton space feature.
9. The method for traffic police dynamic gesture recognition based on vehicle-mounted federated learning according to claim 1 is characterized in that: The selected vehicle-mounted terminal performs multiple rounds of local model training using an optimization algorithm based on the local traffic police gesture data. The optimization algorithm is specifically: Where η is the learning rate; is the momentum term; is the variance term; ε is the gradient parameter; Among them, β1 is the decay rate of the momentum term; β2 is the decay rate of the variance term; is the momentum term of the previous time step; is the variance term of the previous time step.
10. A traffic police dynamic gesture recognition system based on vehicle-mounted federated learning, used to implement the method described in claims 1-9, characterized in that: include: An image acquisition module, which acquires traffic police gesture images; A skeleton construction module, which inputs the traffic police gesture image into a trained key feature extraction model, obtains key nodes and posture vectors in the traffic police gesture image, and constructs skeleton space features according to the key nodes and posture vectors; A scenario construction module constructs a city road network scenario, including a central server, several roadside units and several vehicle-mounted terminals; inputs the skeleton spatial features into a preset global model, and deploys the key feature extraction model and the global model to all vehicle-mounted terminals; A parameter initialization module, which uses a central server to initialize the global model parameters, and sends the initialized global model parameters to all roadside units, which are then sent by the roadside units to the selected vehicle terminals; A local training module, wherein the selected vehicle-mounted terminal performs multiple rounds of local model training according to the local traffic police gesture data, and is used to update the local model parameters, and when the preset conditions are met, the updated local model parameters are uploaded to the nearest roadside unit, and then uploaded to the central server by the roadside unit; Aggregation module: the central server aggregates the collected local model parameters to obtain the final global model parameters; The global training module trains the preset global model based on the final model parameters until the preset number of training times is reached to obtain the final global model; The gesture recognition module recognizes the traffic police’s gestures based on the final global model.
Citation Information
Patent Citations
Traffic police gesture recognition method based on human body key point features
CN111401188A
Personalized behavior recommendation method based on federal learning
CN114003821A