Four-directional traffic police gesture recognition method based on temporal linear human skin model and graph convolutional network

By combining the temporal linear human skin model and the relative height graph convolutional network, a four-directional traffic police gesture recognizer is constructed, which solves the accuracy problem of multi-directional traffic police gesture recognition and achieves efficient recognition in complex environments.

CN115731613BActive Publication Date: 2025-09-23BEIJING UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211424842.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-09-23
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

Existing traffic police gesture recognition technology cannot accurately recognize multi-directional traffic police gestures and is easily affected by environmental changes. Existing methods cannot effectively model camera posture and human posture, resulting in low recognition accuracy.

Method used

A four-directional monocular traffic police gesture recognizer (MTPGR) is constructed by combining the sequential linear human skin model (SMPL) with the relative height-based graph convolutional network (RHGCN). Accurate recognition of traffic police gestures is achieved through the graph convolutional network RHGCN and the spatial mean predictor SMP.

Benefits of technology

In complex multi-pedestrian environments, the recognition accuracy is significantly improved, with the Jaccard coefficient reaching 0.908, an increase of 0.137 compared to existing methods, and it has real-time gesture recognition capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731613B_ABST
    Figure CN115731613B_ABST
Patent Text Reader

Abstract

A four-directional traffic police gesture recognition method based on a temporal linear human skin model and a graph convolutional network belongs to the field of electronic information. This method uses a temporal linear human skin model (SMPL) to reconstruct the dynamic tree of traffic police command gestures, and constructs a temporal dynamic graph model of traffic police gestures based on temporal context information. Secondly, in response to the problem that the existing graph structure partitioning strategy is limited, a relative height-based graph convolution kernel label partitioning strategy (RHPS) is proposed, and a relative height-based graph convolutional network (RHGCN) is designed. Finally, the fusion of RHGCN and the spatial domain average predictor (SMP) is used to design and implement a four-directional traffic police gesture recognizer MTPGR based on monocular vision. The present invention effectively completes the task requirements of four-directional traffic police gesture recognition, and the Jaccard coefficient of the recognition effect reaches 0.908, which is 0.137 higher than the existing single-direction gesture recognition method, and the recognition accuracy is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electronic information and is a traffic police gesture recognition technology based on computer vision and machine learning that can be applied to autonomous driving. Background Art

[0002] Traffic police gesture recognition technology is essential for autonomous vehicles. While purely vision-based autonomous driving technology is maturing and becoming mainstream, the corresponding visual traffic police gesture recognition technology is still insufficient for practical application. Previous research using traffic police gesture datasets failed to account for the four directions of traffic at intersections. Consequently, the corresponding gestures were not defined in the datasets, making the corresponding gesture recognition methods difficult to apply in practice.

[0003] Furthermore, existing methods for directly recognizing traffic police gestures from image sequences are susceptible to environmental fluctuations, affecting gesture recognition accuracy. Keypoint-based 2D skeletal models cannot explicitly represent camera pose parameters (such as the positional relationship between the camera and the traffic officer). Since there is currently no integrated method capable of recognizing traffic police gestures from multiple directions, a method that accurately models both camera and human pose is needed to guide the recognition of traffic police directing gestures. Summary of the Invention

[0004] Based on the consideration of the situation where traffic police direct vehicles in other directions, the present invention uses the dynamic tree in the sequential linear human skin model (SMPL) to estimate the dynamic characteristics of traffic police directing gestures in continuous time, and uses the relative height-based graph convolutional network (RHGCN) to infer gestures. The fusion of RHGCN and the spatial mean predictor (SMP) realizes a four-directional monocular traffic police gesture recognizer (MTPGR). The present invention involves the following three points:

[0005] (1) The visual inference of human posture (VIBE) method is used to infer SMPL parameters from visual information. The dynamic tree of traffic police hand gestures is reconstructed using SMPL, and the graph node parameters in the temporal dynamic graph model (TSKG) are constructed based on the temporal context information.

[0006] (2) A new relative height-based graph convolution kernel label partitioning strategy (RHPS) is proposed to solve the limitations of existing graph structure partitioning strategies.

[0007] (3) A many-to-many graph convolutional network structure RHGCN based on RHPS was proposed, and a four-directional monocular vision traffic police gesture recognizer MTPGR was designed and implemented.

[0008] The core algorithm of the present invention:

[0009] (1) Real-time SMPL parameter prediction algorithm

[0010] The present invention uses a deep learning method and a video-based 3D human pose estimation method (VIBE) to estimate the low-dimensional parameters of a skinned multi-person linear model (SMPL) of traffic police gesture video data, and obtains the three-dimensional coordinates of each human joint by reconstructing the grid.

[0011] Formally, SMPL defines the mesh reconstruction function according to formula (1):

[0012]

[0013] in, Represents body shape parameters, such as fatness, thinness, etc. represents the x-dimensional real number space; Indicates posture parameters, such as standing or sitting; represents the SMPL statistical model parameters; Represents a three-dimensional mesh of the human body, N represents the number of vertices in U, and 3N means that each vertex consists of 3 coordinate values; and Represents the scalar quantity in body shape parameters and posture parameters. In the SMPL model, the parameter and is from a predefined rest pose template mesh Offset value to target grid U, parameter is the distance offset value of the body shape, parameter It is the radian offset value of the human joint rotation.

[0014] In order to reduce the complexity of U and maintain its feature expression ability, this paper uses the regression The positions of the three-dimensional joints of the human body are further calculated by regression from the human body mesh U. As shown in formula (2).

[0015]

[0016] Among them, W joint represents the regressor parameters learned by the deep learning method, represents the joint position matrix, N J Indicates the number of joints. Convert the mesh U to joint positions in 3D space.

[0017] When processing the monocular real-time image sequence of traffic police gestures, the present invention uses the VIBE method to predict and Let X C is the image sequence, that is, X C ={I t |t=1,...,T}, where T represents the time length and I represents the image. C As input prediction and The process includes two stages: Convolutional Network (CNN) and Gated Recurrent Unit (GRU). CNN stage f c From each frame I t The features are extracted according to formula (3).

[0018] C t =f c (I t , Z t ;W VC )#(3)

[0019] in Represents the image color value at time t, W I and H I represents the width and height of the input image, represents the extracted image features, F I Represents the number of features extracted by the predefined CNN, Z t represents the coordinates of the upper left and lower right points of the traffic police’s body bounding box at time t, W VC represents the trained network parameters, Z t From the person tracker to the person bounding box Z t-1 and the current image I t Recursive calculation of .

[0020] The second stage contains GRU and a multi-layer fully connected network to transform the hidden state of GRU into and GRU can be expressed as a recursive function according to formula (4)

[0021]

[0022] in, represents the hidden state at time t, F H represents the number of scalars in the hidden state. For each frame, h t is further input into the multi-layer fully connected network shown in formula (5) middle.

[0023]

[0024] in, and Represents the body shape parameters and posture parameters at time t. and W joint After that, the joint position P t It can be calculated by formula (1) and formula (2).

[0025] (2) Temporal Kinetic Graph Model of Traffic Police Hand Signals (TSKG)

[0026] The joint position matrix P contains the three-dimensional joint coordinates of the gesture, but lacks the key spatial and temporal relationship information of the gesture. Therefore, this paper proposes the TSKG model to express the dynamic relationship between different joints in space, as well as the temporal relationship between the same joints in different frames.

[0027] TSKG is defined as an undirected graph G = (V, E). is the vertex set, E is the edge set, and J represents the number of joints. j,t |j∈1,...,J,t=1,...,T}, each vertex Represents the 3D joint coordinates. The edge set E in the figure consists of two subsets shown in formula (6).

[0028] E=(E S ∪E T )#(6)

[0029] in, Represents the dynamic relationship between joints in each frame, represents the temporal association between spatially identical and temporally adjacent joints, Represents the x-dimensional natural number field space. E S It consists of the union of three sets of edges described by formula (7).

[0030]

[0031] Among them, E K represents the dynamic connection set, E A represents the auxiliary set, E C Represents a circuit set. Set E K Contains edges that are naturally connected according to the human skeleton, such as the left arm and left elbow, as Figure 1 As shown in a. Set E A Contains edges that were manually added to the graph, such as Figure 1 As shown in b, it is used to maintain the connectivity of the graph. Set E C Contains the loop formed by all vertices in the graph and itself, such as Figure 1 As shown in c, it is used to add vertex features to graph convolution calculation.

[0032] Time domain correlation edge set ET Contains edges formed by connecting vertices that are the same in space and adjacent in time, such as formula (8) and Figure 2 shown.

[0033] E T ={(v j,t , v j,t+1 )|j∈1,...,J,t=1,...,T-1}#(8)

[0034] Where, vertex v j,t and v j,t+1 Represents the same joint in 2 consecutive frames. E T The TSKG model adds time domain information on the basis of spatial domain information, which is crucial for the task of traffic police gesture recognition.

[0035] (3) Relative Height-Based Graph Convolution Kernel Label Partitioning Strategy (RHPS)

[0036] This paper designs a graph convolutional network (GCN) that is applicable to both spatial and temporal dimensions to recognize police gestures. Unlike image convolution operations, which have a fixed structure and are matched using square convolution kernels, the graph vertices in this paper have a variable number of adjacent vertices. This requires setting the corresponding relationship between the graph convolution kernel parameters and adjacent vertices, which is called a partitioning strategy (SCPS). The SCPS can be defined according to Equation (9).

[0037]

[0038] Among them, v j,t The vertex representing the center of the graph convolution kernel, v n,t Indicates v j,t The adjacent vertices of r j Indicates v j,t The distance between the center of gravity g and the skeleton, r n Indicates v n,t and the distance between g.

[0039] There are two differences between the implementation and theoretical value of SCPS: (1) The centroid g in the image is actually an arbitrarily selected vertex g′ on the graph. (2) The image distance r is replaced by the graph distance r′, which calculates the shortest distance between two vertices based on the number of edges. Figure 3 The comparison between theory and implementation is shown, where the × symbol represents the center, the dashed range represents the range of graph convolution, and the double circle vertices represent the convolution kernel centers. The dark dashed lines represent the distance r in (a) and the distance r′ in (b).

[0040] These two differences will cause some datasets to be incompatible. If there is no suitable g′ in the dataset annotation that can approximate the center point, SCPS cannot be used. To solve this problem, the present invention proposes a relative height partitioning strategy (RHPS), which assigns a height value to each joint and uses it as a label for comparison to obtain the distance. The center of the convolution kernel is v j,t When joint v n,t The marking function As shown in formula (10).

[0041]

[0042] Among them, s n and s j Represents joint v n,t and joint v j,t The height value of .

[0043] Examples of RHPS are Figure 4 As shown. Figure 4 In (a), the horizontal dashed lines separate the joints into different regions. Within each region, a height value s is assigned as a proxy for the distance r in SCPS. By pre-specifying the height value, we avoid the need for computations associated with a specific central vertex, eliminating the need for a specific central vertex in the graph structure. Figure 4 Example b illustrates the calculation of the relative height d. The double circles indicate the locations of the convolution kernel centers. Standing pose is a common posture across all human keypoint datasets, so RHPS is unaffected regardless of the distribution of keypoint annotations. Therefore, RHPS addresses the data compatibility issue of SCPS.

[0044] (4) Four-directional monocular traffic police gesture recognizer MTPGR integrating RHGCN and SMP

[0045] This patent proposes a new monocular vision traffic police gesture recognizer MTPGR for recognizing continuous four-directional traffic police gestures. MTPGR consists of two networks: (1) Graph Convolutional Network RHGCN. It applies RHPS as the convolution kernel partitioning strategy and keeps the input and output of the time dimension corresponding. (2) Spatial Domain Average Predictor (SMP). SMP performs average pooling operations on the graph in the spatial dimension while retaining the time dimension to support many-to-many prediction mode, and then predicts a gesture category for each frame. The overall architecture of MTPGR is as follows: Figure 5 shown.

[0046] Figure 5 The RHGCN shown in (b) takes the spatiotemporal graph G in the TSKG model as input, where the vertex set V is used as the input feature and the edge set E is used as the adjacency matrix that complies with the RHPS rule. The output is the gesture feature set Y shown in formula (11):G .

[0047]

[0048] Among them, W G represents the model parameters of RHGCN, Y G Represents the output gesture features. The architecture of RHGCN is as follows Figure 5 As shown in b, each block in the figure represents a spatiotemporal graph convolutional layer (STGCL), and the number represents the number of channels in the previous layer.

[0049] Each STGCL consists of 7 parts, such as Figure 6 As shown in the figure, the sequence is the residual layer starting position, spatial graph convolution layer, attention layer, ReLU activation layer, temporal convolution layer, residual layer ending position, and ReLU activation layer. The spatial graph convolution layer uses a 3*3 convolution kernel for graph convolution calculation according to RHPS, as shown in the figure. Figure 7 The dotted area in a is shown. After multiplying each edge by the attention weight, the calculation result is activated using the ReLU activation function, and the result is output to the time convolution layer. The 3*3 convolution kernel is used in the time dimension to perform graph convolution calculation, as shown in Figure 7 (b) The residual layer takes the computation from the spatial graph convolution layer to the temporal convolution layer as the residual and adds it to the direct mapping to improve network training efficiency. Finally, the ReLU activation function is used to activate the summation result of the residual layer and output it to the next STGCL.

[0050] SMP is connected to RHGCN, as shown in 5c, to extract the gesture feature Y G Y G The scalar y can be used G Expand according to formula (12).

[0051]

[0052] Among them, C G Represents the number of channels in the output, N J and T represent the number of joints and duration. Thus, the calculation of the spatial mean predictor (SMP) is shown in formula (13).

[0053]

[0054] in, Represents the spatial eigenvalue y on channel c G The average value at time point t is used, and then the vector representing the score of each type of gesture is obtained using formula (14).

[0055]

[0056] Among them, the vector represents a fully connected network, W E is the network parameter, Is a vector representing the score of each gesture category, K represents the number of gesture categories. Gesture category k at time t t As shown in formula (15).

[0057]

[0058] Effects of the Invention

[0059] A MTPGR consisting of a temporal SMPL kinetic graph model (TSKG), a relative height graph convolutional network (RHGCN), and a spatial domain averaging layer (SMP) was constructed to effectively complete the task requirements of four-directional traffic police gesture recognition. In a complex environment with multiple pedestrians in four command directions, the Jaccard coefficient of the recognition effect reached 0.908, which is 0.137 higher than the existing single-direction gesture recognition method, and the recognition accuracy was significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 Represents the edges corresponding to the spatial dynamics relationship in the graph structure, where (a) is the edge naturally connected according to the human skeleton, (b) is the edge added to the graph due to structural needs, and (c) is the loop edge;

[0061] Figure 2 Represents the edges corresponding to the time domain relationship in the graph structure;

[0062] Figure 3 Represents the difference between theory and implementation in the SCPS partitioning strategy, where (a) is theory and (b) is implementation;

[0063] Figure 4 is an example of the relative height partitioning strategy (RHPS), where (a) represents the calculation of relative height and (b) represents the calculation of height difference;

[0064] Figure 5 This is the overall architecture of MTPGR, where (a) is the TSKG model, (b) is the RHGCN network, and (c) is the SMP network;

[0065] Figure 6 It is the internal structure of the spatiotemporal graph convolutional layer (STGCL);

[0066] Figure 7 Represents the range of different convolution types, (a) is the spatial domain, and (b) is the time domain. DETAILED DESCRIPTION

[0067] The specific implementation method of the present invention, that is, the process of building, training, and deploying MTPGR, includes three steps:

[0068] (1) Establishing a four-directional traffic gesture dataset

[0069] Existing datasets cannot reflect the real traffic police command situation. A dataset containing continuous traffic police gesture videos in four directions is established for model training. The videos are recorded by a fixed camera mounted on the top of the vehicle. The average video length is about 10 minutes. In each video, the traffic police stands at the center of the intersection and makes eight consecutive gestures in different directions: stop, go straight, turn left, wait to turn, turn right, change lanes, slow down, and stop. In addition to the traffic police, the video also records images of all vehicles and pedestrians moving at the intersection. A total of 10 different scenes are recorded and divided equally into training and test sets. The category, start and end time of each gesture in the dataset video are manually annotated, and the time period without gestures is marked with the standby category.

[0070] Before training, the dataset is preprocessed. The steps include: saving the video as a picture set at a uniform frame rate; converting the recorded gesture categories and start and end times into labels corresponding to the picture serial numbers; pre-tracking all the characters in the video to obtain the motion trajectories of all characters and the corresponding character IDs; performing non-maximum suppression on all trajectories and taking the top-ranked trajectory as the tracking trajectory of the traffic police; finding the locations where the traffic police trajectory is disconnected and treating them as the locations where the traffic police are obscured, deleting the pictures corresponding to these locations from the picture set, and deleting the annotations corresponding to these locations from the label.

[0071] (2) Training network model parameters

[0072] This method uses a Jetson Nano hardware environment, running Python, and is built with the PyTorch library. Network parameters are initialized using the PyTorch default initialization behavior. Attention, batch normalization, and dropout are employed to reduce training time and improve training effectiveness. A video from a traffic police hand gesture dataset is used as input to the VIBE network to extract the TSKG model. A graph convolutional network and a SMP network are then used to predict hand gestures. The dataset gesture labels are used as the ground truth, and the distance between them is measured using cross-entropy. Backpropagation is used to optimize network parameters and reduce the cross-entropy distance. The network is trained using the entire training set as an epoch. The cross-entropy value for each epoch is recorded, and this process is repeated until the cross-entropy value remains unchanged after three epochs. This results in the optimal network parameters for the current dataset. These parameters are saved and loaded before the model infers traffic police hand gestures.

[0073] (3) Using the trained model for recognition

[0074] After completing training in step 2, the TSKG model is sequentially connected to the trained RHGCN network and SMP network to produce the four-directional traffic police gesture recognizer MTPGR. MTPGR takes real-time video recorded by an on-board camera as input, processes the data using a 5-second sliding window, extracts TSKG model parameters from the video within the window, then predicts gestures using a graph convolutional network and an SMP network, outputting traffic police gesture predictions synchronized with the real-time video. After the prediction is completed, the image sequence still within the sliding window remains in memory and serves as historical data for predicting the traffic police gesture in the latest frame. This method achieves excellent recognition performance on a four-directional traffic police gesture dataset, significantly improving over existing methods, with a Jaccard index of 0.91 and a confusion rate of 0.1% for mispredicting one gesture as another. With a response latency of 0.2 seconds, it demonstrates the ability to predict gestures in real time.

Claims

1. A four-directional traffic police gesture recognition method based on a temporal linear human skin model and graph convolutional networks, characterized by: It includes the following three parts: Establishing a temporal dynamics graph model for traffic police's hand gestures The temporal dynamic graph model TSKG of traffic police hand gestures is based on the skinned multi-person linear model SMPL and is constructed according to the functional requirements of traffic police hand gesture recognition. TSKG simultaneously represents the dynamic relationship between different joints in space and the relationship between the same joints in different frames in time series; TSKG is defined as an undirected graph G = (V, E); where is the set of vertices, E is the set of all edges, N J represents the number of joints, T represents the total length of the graph record; V = {v j,t |j∈1,…,J,t=1,…,T}, 1,…,J represents N J Different joints, each vertex represents the 3D joint coordinates, represents the x-dimensional real number domain space; the edge set E in the figure consists of the two subsets shown in formula (1); And=(And S ∪E T )#(1) in, Represents the dynamic relationship between joints in each frame, Indicates the temporal correlation of the same joint node between adjacent frames; E S It consists of the union of the three sets of edges described by formula (2); Among them, E K represents the dynamic connection set, E A represents the auxiliary set, E C represents the circuit set; set E K Contains edges that are naturally connected according to the human skeleton, such as the left arm and left elbow; set E A Contains the edges manually added to the graph to maintain the connectivity of the graph; Set E C Contains the loop formed by all vertices in the graph and itself, which is used to add the features of the vertex to the graph convolution calculation; Time domain associated edge set E T It includes the edges formed by connecting vertices that are the same in space and adjacent in time domain, as shown in formula (3); E T ={(v j,t ,v j,t+1 )|j∈1,…,J,t=1,…,T-1}#(3) Where, vertex v j,t and v j,t+1 Represents the same joint in 2 consecutive frames; 2) Graph Convolutional Network RHGCN using Relative Height Partitioning Strategy RHPS The key features of the graph convolutional network of MTPGR include two aspects: the relative height partitioning strategy and the network structure of the graph convolutional network; The relative height partitioning strategy assigns a height value to each joint and uses it as a label for comparison to obtain the distance, thereby determining the correspondence between the graph convolution kernel parameters and adjacent vertices; RHPS assigns a height value s to the joint according to the relative height of the joint in the standard standing posture; the vertex corresponding to the convolution kernel center is v j,t When joint v n,t The marking function As shown in formula (5); Among them, s n and s j Represents joint v n,t and joint v j,t The height value of The graph convolutional network RHGCN using RHPS takes the undirected spatiotemporal graph G in the TSKG model as input, where the vertex set V is used as the input feature and the edge set E is used as the adjacency matrix that complies with the RHPS rule. The output is the gesture feature set Y shown in formula (6): G ; Among them, W G represents the model parameters of RHGCN, Y G Represents the output gesture features; Represents the graph convolutional network RHGCN, which consists of multiple spatiotemporal graph convolutional layers STGCL with different numbers of output channels connected sequentially, set to 4 STGCLs with 64 output channels, 3 STGCLs with 128 output channels, and 3 STGCLs with 256 output channels; Each STGCL is a residual block structure, which consists of 7 parts, namely the starting position of the residual layer, the spatial graph convolution layer, the attention layer, the ReLU activation layer, the temporal convolution layer, the end position of the residual layer, and the ReLU activation layer; the spatial graph convolution layer uses a 3*3 convolution kernel to perform graph convolution calculation in accordance with RHPS, multiplies each edge by the attention weight, and then uses the ReLU activation function to activate the calculation result, and outputs the result to the temporal convolution layer, and uses a 3*3 convolution kernel in the time dimension to perform graph convolution calculation again; the residual layer uses the above calculation part from the spatial graph convolution layer to the temporal convolution layer as the residual part, and adds it to the direct mapping part to improve the network training efficiency; finally, the ReLU activation function is used to activate the result of the residual layer summation and output it to the next STGCL; 3) Four-directional monocular traffic police gesture recognizer MTPGR that integrates RHGCN and spatial domain average predictor SMP The spatial domain average predictor is the last part of the network structure in MTPGR and is connected to RHGCN. It performs average pooling on the graph in the spatial dimension and retains the temporal dimension to cooperate with the many-to-many prediction mode to predict a gesture category for each frame. SMP is used to generate the gesture feature Y output by RHGCN after inputting the graph G. G Classify gestures in Y G The scalar y can be used G Expand into the form of a set according to formula (7); Among them, C G Represents the number of channels in the output, N J and T represent the number of joints and duration; hence, the calculation of the spatial mean predictor (SMP) is shown in formula (8); in, Represents the spatial eigenvalue y on channel c G The average value at time point t is used, and then the vector representing the score of each type of gesture is obtained using formula (9); Among them, the vector represents a fully connected network, W F is the network parameter, Is a vector representing the score of each gesture category, K represents the number of gesture categories; gesture category k at time t t As shown in formula (10);