Traffic police gesture recognition method and system based on 3D attitude reconstruction and LSTM
By combining OpenPose and LSTM's 3D posture reconstruction and sequence modeling network, the problems of traffic police gesture recognition accuracy and real-time in complex traffic environments are solved, and high accuracy and rapid response in lighting changes and complex scenarios are achieved.
Patent Information
- Application Number
- CN202510684406.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art lacks the accuracy and real-time recognition of traffic police gestures in complex traffic environments, and it is difficult to adapt to the challenges of lighting changes, background complexity and traffic police arm posture diversity.
Using 3D pose reconstruction and sequence modeling network based on OpenPose and LSTM, dynamic recognition of traffic police gestures is achieved through a combination of skeleton key point extraction, pose reconstruction and sequence modeling.
This method maintains a high recognition accuracy rate in light changes and complex traffic scenarios, can effectively overcome the interference of irrelevant factors such as traffic police height and clothing, has high recognition accuracy and real-timeness, and is suitable for practical application needs.
Smart Images

Figure CN120198971A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of interactive perception of unmanned vehicles, and in particular to a method and system for dynamic recognition of traffic police gestures based on OpenPose and LSTM, which are applied in complex traffic scenes based on deep learning. Background Art
[0002] At present, driverless technology is developing rapidly in general scenarios, but due to the constraints of complex traffic environments, the realization and full popularization of L4 / L5 systems will still take some time. In the transitional stage of technology and regulations, traffic police gesture recognition has become a key link in the driving of driverless cars and their interaction with the outside world. In actual traffic scenarios, the command gestures of traffic police have higher priority than traffic lights and traffic signs. In special circumstances such as traffic control and traffic accidents, driverless vehicles must obey the command of traffic police. Therefore, constructing a dynamic recognition method for traffic police gestures is the core technology for realizing human-machine collaborative traffic management, which is of great significance to ensuring road safety during the transition period.
[0003] At present, there are two major categories of research on traffic police gesture recognition: methods based on depth images and methods based on gesture feature extraction. For actual traffic scenes, vision-based methods are more practical. Traditional vision-based traffic police gesture recognition methods are usually divided into two steps: first, extracting the spatial features of gestures from videos or depth images, and then classifying and identifying the extracted features. However, due to factors such as lighting changes in traffic scenes, background complexity, and the diversity of traffic police arm postures, traditional methods have large errors in feature extraction and classification accuracy; at the same time, the model generalization ability of traditional methods is weak, making it difficult to adapt to complex and changeable actual scenes, and the real-time performance is poor. Summary of the invention
[0004] In view of the above problems, the present invention innovatively proposes a method based on of Traffic police gesture dynamic recognition method and system based on posture estimation and LSTM. The present invention has excellent robustness and can maintain a high recognition accuracy rate under various interference factors such as changes in light intensity and complex traffic scenes. At the same time, the invention can effectively overcome the interference of irrelevant factors such as the height and clothing of traffic police, and has both high recognition accuracy and real-time performance, which can meet the needs of practical applications.
[0005] In order to achieve the design purpose, the present invention adopts the following technical solutions:
[0006] In a first aspect, the present invention provides a method for traffic police gesture recognition based on 3D posture reconstruction and LSTM, comprising the following steps:
[0007] Extract the skeleton key points from the traffic police command gesture video to obtain multiple key points related to gesture recognition;
[0008] Construct a pose reconstruction network to perform pose estimation on the extracted key points. The pose reconstruction network includes multiple processing modules, each module is used to construct the pose of a specific joint, and fuse the input pose features with the estimated pose features through a feature fusion layer, and finally output a complete pose estimation result;
[0009] Input the pose information output by the pose reconstruction network into the sequence modeling network, model the dynamic sequence of the traffic police command gesture through the sequence modeling network, recognize the semantic meaning of the dynamic gesture, and complete the gesture recognition.
[0010] In a second aspect, the present invention provides a traffic police gesture recognition system based on 3D pose reconstruction and LSTM, including:
[0011] A key point extraction module, which is used to extract the skeleton key points from the traffic police command gesture video to obtain multiple key points related to gesture recognition;
[0012] A pose reconstruction network module, which is used to perform pose estimation on the extracted key points. The pose reconstruction network includes multiple processing modules, each module is used to construct the pose of a specific joint, and fuse the input pose features with the estimated pose features through a feature fusion layer, and finally output a complete pose estimation result;
[0013] A sequence modeling network module, which is used to process the pose information output by the pose reconstruction network module, model the dynamic sequence of the traffic police command gesture through sequence modeling, recognize the semantic meaning of the dynamic gesture, and complete the gesture recognition.
[0014] The present invention has the following beneficial effects:
[0015] 1. Compared with the traditional spatial modeling method based on handcrafted features, the present invention uses the p-LSTMs network to explicitly model the joint spatial dependence and temporal continuity, and effectively overcomes the interference of traffic police height differences and complex lighting through a parallel spatio-temporal feature learning mechanism.
[0016] 2. Aiming at the problem that the traditional method ignores the constraints of the human motion chain, the present invention encodes the skeleton chain anatomical structure into a multi-stage key point detection network, and realizes progressive learning from local features to global constraints through cascaded residual modules.
[0017] 3. Compared with methods such as convolutional neural networks, the LSTM adopted by the present invention has higher recognition accuracy and faster response speed in continuous actions, and can be applied to actual combat. Description of the Drawings
[0018] Figure 1 This is the overall flowchart of the present invention.
[0019] Figure 2 is a schematic diagram showing the key point numbering of the traffic police command gestures in the present invention and the corresponding skeleton structure.
[0020] Figure 3 This is the schematic diagram of the 3D pose reconstruction network in the present invention.
[0021] Figure 4 This is the schematic diagram of the LSTM structure in the present invention. Detailed implementation manners
[0022] To make the technical solution of the present invention clearer and more definite, the following will comprehensively and systematically elaborate on the implementation manners of the present invention in combination with the accompanying drawings and specific embodiments.
[0023] As Figure 1 shown, the embodiment of the present application provides a dynamic traffic police gesture recognition method based on 3D pose estimation of OpenPose and LSTM network, and the specific steps are as follows:
[0024] 1) The embodiment of the present application uses the OpenPose algorithm to perform 2D extraction of skeleton key points on the traffic police command gesture video.
[0025] For the actual application scenario requirements, the embodiment of the present application optimizes the key points of the human pose estimation model of the standard OpenPose:
[0026] Among the 25 whole-body key points output by the original model, according to the business characteristics of traffic command gesture recognition, 14 key points closely related to the upper body posture and hand movements are mainly retained (including key upper limb nodes such as the neck, both shoulders, both elbows, and both wrists, as well as basic posture reference points such as the head and both hips). The 11 redundant key points to be removed can be divided into two categories:
[0027] (1) Lower limb end nodes and over-proximate nodes (right ear, left ear, right big toe, left big toe, right little toe, left little toe, right heel, left heel, hip midpoint);
[0028] (2) Head detail features (right eye, left eye, right ear, left ear).
[0029] This model optimization strategy based on the business scenario effectively reduces the computational load by 42% while ensuring the gesture recognition accuracy, and at the same time improves the algorithm processing speed, laying a technical foundation for subsequent lightweight deployment. Among them, the 14 retained key points are as Figure 2As shown in the figure, they are respectively: 1: nose; 2: neck; 3: right shoulder; 4: right elbow; 5: right wrist; 6: left shoulder; 7: left elbow; 8: left wrist; 9: right hip; 10: right knee; 11: right foot; 12: left hip; 13: left knee; 14: left foot.
[0030] 2) The model training of the current 3D pose estimation method highly depends on the annotated 3D pose data collected in the laboratory environment. At the same time, the mainstream neural network methods generally ignore the biomechanical constraint relationships between human joints, especially the lack of explicit modeling of the structural connectivity, resulting in limited pose estimation accuracy.
[0031] To address the above problems, the embodiments of the present application analyze the kinematic chain characteristics of the human body by constructing a 3D pose reconstruction network and performing pose estimation on the extracted key points to : there is a strong motion correlation between the proximal joints (such as the torso and hips), while the distal joints (such as the ends of the limbs) show relatively independent motion characteristics.
[0032] Furthermore, the embodiments of the present application also perform a unified representation of the 2D traffic police command gesture feature points extracted in 1) in terms of vectors, which helps the model achieve better robustness.
[0033] 2.1) Construct ;
[0034] As Figure 3 shown in (D) of : a consists of 9 units and a connection layer. Each unit includes an LSTM and a deep fusion layer. A
[0035] consists of a memory unit with 100 hidden units and a fully connected layer with 150 hidden units.
[0036] To effectively detect traffic police command gestures, 13 groups of skeleton vectors are constructed in this embodiment. The construction steps are as follows:
[0037] (1) The 14 key points extracted in step 1) form a set ; (2) Represent the skeletons formed by the key point links with vectors and divide them into body seed joint vectors and limb joint vectors. Among them, the body seed joint vectors are:
[0038]
[0039] The limb joint vectors are:
[0040]
[0041] Among them represents the serial number and the serial number of the connection vector. For the sake of convenient representation, hereinafter, will be used to represent the in the vector and vector union in the frame image.
[0042] 2.3) Construct a three-dimensional pose reconstruction network;
[0043] To describe Figure 3 more comprehensively, each module is explained by means of letter annotation: (A) represents the input traffic police command gesture image; (B) represents the extraction of 2D poses from the image using OpenPose; (C) represents the 2D poses extracted from (B); (D) represents the use of to extract depth information from (C); (E) represents a unit of ; (F) represents the three-dimensional pose reconstruction through according to the structural connectivity based on body parts; (G) represents the output of the 3D pose.
[0044] The three-dimensional pose reconstruction network consists of 9 . Starting from the 2D pose, the first s only constructs the 3D joints of the body center of gravity part, and these joints are regarded as seed joints. Each constructs its part of the pose and connects them to each other with fully connected layers.
[0045] Its overall process is as follows: First, the input 2D pose is fused with the estimated 3D joint features through the depth fusion layer, and then the generated fusion result is input into 9 cascaded p-LSTMs for step-by-step processing, and finally the complete overall three-dimensional pose estimation result is output.
[0046] Furthermore, the specific construction parts corresponding to each are as follows:
[0047] is to construct the 3D joints of the nose, neck, left shoulder, right shoulder, left hip, and right hip;
[0048] is to construct the 3D joint of the right elbow;
[0049] is to construct the 3D joint of the left elbow;
[0050] is to construct the 3D joint of the right knee;
[0051] It is a 3D joint for constructing the right hand;
[0052] It is a 3D joint for constructing the left hand;
[0053] It is a 3D joint for constructing the right foot;
[0054] It is a 3D joint for constructing the left foot;
[0055] Finally, unify all the above networks into an end-to-end network structure to achieve efficient and highly accurate 3D pose extraction.
[0056] It should be noted that from the second to the last , it must fully rely on the output of the previous . At the same time, to prevent the initial pose from disappearing, each will use the input pose as auxiliary data and merge it with its own output in the depth fusion layer.
[0057] 2.4) Establish the mapping function of ;
[0058] Its input vector is the 2D pose vector of , and the output vector is the 3D pose vector of .
[0059] The formula of the mapping function is:
[0060]
[0061] Among them, represents the output of the pose, represents the fully connected layer, represents the output of the depth fusion layer, represents the input pose, represents the index of
[0062] The pseudo-code of the mapping function is as follows:
[0063] From to
[0064] If
[0065]
[0066]
[0067] Otherwise:
[0068]
[0069]
[0070] Return
[0071] 3) The pose information generated through step 2) is used as input to achieve dynamic recognition of traffic police command gestures;
[0072] The goal of this step is to model the pose sequence of traffic police and recognize the semantic meaning of dynamic gestures (such as stop, go straight, turn left, etc.). The whole process is divided into four sub-steps, specifically as follows:
[0073] 3.1) Pose temporal feature preservation and gating mechanism;
[0074] Manage the information flow in the time series through the "gating unit". The input at each moment (such as a frame in a video) consists of two parts: the current pose feature, that is, the feature vector composed of the joint coordinates (such as the positions of the right wrist and left elbow) generated by step 2); the hidden state at the previous moment, that is, the vector that stores the historical action information. The memory unit will gradually accumulate the change trend of the pose, forget the irrelevant actions, and finally form the complete gesture action feature. The specific formula of the gating mechanism is:
[0075]
[0076] Among them, represents the temporal feature of the traffic police command gesture output, represents the temporal feature of the traffic police command gesture output in the previous time step, represents the time step of the traffic police command gesture pose, represents feature fusion, represents the weight matrix, represents the bias term, represents the time step of the state, which is used for the memory preservation (hidden state) of traffic police command gestures and serves as the output of the next recurrent neural network.
[0077] Specifically, the forget gate determines which historical information needs to be discarded. For example, if the traffic police gesture has switched from "stop" to "left turn", the forget gate will reduce the memory of the previous "stop" action; the input gate determines which new information needs to be added to the memory. For example, if the left arm suddenly extends in the current frame, the input gate will focus on recording this change; the memory update combines the results of the forget gate and the input gate to update the memory cell of the LSTM. This cell preserves the long-term dependencies of the gesture actions, such as the movement trajectory of continuously waving the left arm.
[0078] 3.2) Calculation of the output vector ;
[0079] The output vector at each moment encodes the action features of the current frame and combines with the historical features to form a complete temporal representation of the dynamic gesture. The updated memory cell needs to be transformed into the output features at the current moment for subsequent classification. The specific formula is as follows:
[0080]
[0081] where, represents function, represents the hyperbolic tangent function. represents vector concatenation, represents matrix multiplication, represents dot product, represents the current time. represents at time the traffic police command gesture features, which is used for the memory preservation (hidden state) of traffic police command gestures and serves as the output of the next recurrent neural network, and represent the weights and biases of the fully connected layer.
[0082] Specifically, the output gate filters the information in the memory cell and only retains the part useful for gesture classification. For example, the output gate may ignore the slight shaking of the traffic police's body and focus on the key displacements of the arm; the non-linear transformation compresses the memory cell to the range through the function to enhance the model's ability to express complex actions.
[0083] 3.3) Calculation of the prediction probability;
[0084] The output vector of the LSTM is input into the fully connected layer and the Softmax function to calculate the probability of each preset gesture category. The specific formula is as follows:
[0085]
[0086] function represents , represents the probability that the current gesture belongs to each gesture category, and represent the weights and biases of the fully connected layer, represents the temporal features of the output traffic police command gesture.
[0087] Specifically, the fully connected layer maps the high-dimensional features to the gesture category space. In the embodiments of the present application, a total of 8 gestures can be recognized (stop, go straight, turn left, turn right, pull over, wait for a left turn, slow down, change lanes), so an 8-dimensional vector is output. The Softmax function normalizes the output values into a probability distribution, indicating the possibility that the current frame belongs to a certain gesture category.
[0088] 3.4) Output the predicted gesture;
[0089] To avoid misjudgment with low confidence, a threshold needs to be set for filtering, and the gesture with the highest predicted probability is used as the predicted gesture. The specific formula is as follows
[0090]
[0091] represents the final gesture classification output, represents the action confidence threshold, and a zero value indicates no gesture.
[0092] 4) Model training and evaluation;
[0093] 4.1) Overall network architecture;
[0094] As Figure 4 shown, and are the inputs of the network. Among them, is the 3D traffic police command gesture posture at time. and are the outputs of the network and are used as the input at in the figure represents the fully connected layer, represents the pointwise operation.
[0095] 4.2) Model training and evaluation;
[0096] The LSTM network is trained with the input traffic police command gesture data so that it can have the ability to dynamically distinguish command gestures. After multiple experiments, when the learning rate, batch number, and number of neurons in the hidden unit of the LSTM are set to 0.0005, 10, and 40 respectively, the model recognition accuracy reaches the highest. During the training process, the mean squared error loss function is used to optimize the network parameters so that it can accurately distinguish different gesture categories. After 20,000 steps of cumulative training, the accuracy of the subsequent trained model no longer improves. Therefore, the embodiment of this application stops training after 20,000 steps.
[0097] Twenty test sets are input into the trained model for testing. The average recognition time is 1.13 s, and the average recognition accuracy is as high as 95.82%, as shown in the following table. Therefore, the embodiment of this application is feasible and effective and can be applied to fields such as driverless cars.
[0098] Comparison table of recognition accuracies of different models
[0099]
[0100] Based on the same inventive concept, the embodiment of this application also provides a traffic police gesture recognition system based on 3D pose reconstruction and LSTM, including:
[0101] A key point extraction module, configured to extract the skeleton key points of the traffic police command gesture video to obtain multiple key points related to gesture recognition;
[0102] A pose reconstruction network module, configured to perform pose estimation on the extracted key points. The pose reconstruction network includes multiple processing modules, each module is configured to construct the pose of a specific joint, and fuse the input pose features and the estimated pose features through a feature fusion layer, and finally output a complete pose estimation result;
[0103] A sequence modeling network module, configured to process the pose information output by the pose reconstruction network module, model the dynamic sequence of the traffic police command gesture through sequence modeling, recognize the semantic meaning of the dynamic gesture, and complete gesture recognition.
[0104] The embodiment of this application is implemented in the Python language and deployed on a high-performance embedded development board equipped with external cameras, a quad-core ARM Cortex-A57 @ 1.43 GHz CPU, a 128-core NVIDIA Maxwell architecture GPU, 4GB LPDDR4 RAM, 16GB eMMC ROM and other hardware resources, which can efficiently process image recognition tasks.
[0105] In a specific implementation, a video stream is collected in real time through a camera, and the traffic police gestures are detected and classified in real time using the embodiments of the present application. The recognition results are sent to the unmanned vehicle control system through serial communication to drive it to perform actions such as moving forward, stopping, and turning. This design is not only applicable to unmanned vehicles but also can be extended to fields such as intelligent traffic lights and robot control, having a broad application prospect. By combining a lightweight model and high-performance hardware, the embodiments of the present application improve the adaptability and safety of unmanned vehicles while providing reliable technical support for the popularization of intelligent transportation and robotics technology.
[0106] Compared with the prior art, this solution has a higher recognition accuracy of 95.82%, the inference delay is controlled within 1.13 s, and at the same time, it can overcome the interference of the height and clothing of traffic police, meeting the actual application requirements of driverless vehicles in real road environments.
[0107] In summary, the present application uses a multi-stage processing architecture to achieve intelligent recognition of traffic police command gestures. First, the traffic police command posture information is collected in real time through an in-vehicle multi-camera vision sensor; second, based on the improved framework, highly robust human key point features are extracted; furthermore, a three-dimensional pose reconstruction network with a spatio-temporal attention mechanism is constructed to map the features to the three-dimensional space to achieve accurate estimation of the coordinates of bone joint points; finally, temporal modeling is performed through the model, and gesture semantic parsing is completed in combination with an adaptive threshold classifier.
[0108] The above embodiments are only used to illustrate the present invention. Each step can be changed. Based on the technical solution of the present invention, any improvement and equivalent transformation made to individual steps according to the principle of the present invention should not be excluded from the protection scope of the present invention.
Claims
1. A traffic police gesture recognition method based on 3D pose reconstruction and LSTM, characterized in that It includes the following steps: Extract the skeleton key points from the traffic police command gesture video to obtain multiple key points related to gesture recognition; Construct a pose reconstruction network to perform pose estimation on the extracted key points. The pose reconstruction network includes multiple processing modules, each module is used to construct the pose of a specific joint, and fuse the input pose features with the estimated pose features through a feature fusion layer, and finally output the complete pose estimation result; Input the pose information output by the pose reconstruction network into the sequence modeling network, model the dynamic sequence of the traffic police command gesture through the sequence modeling network, recognize the semantic meaning of the dynamic gesture, and complete gesture recognition.
2. The traffic police gesture recognition method according to claim 1, wherein In the key point extraction step, use the OpenPose algorithm to extract the skeleton key points from the traffic police command gesture video, and optimize the model to retain 14 key points related to gesture recognition; the key points include: nose, neck, right shoulder, right elbow, right wrist, left shoulder, left elbow, left wrist, right hip, right knee, right foot, left hip, left knee and left foot.
3. The traffic police gesture recognition method according to claim 1, characterized in that The pose reconstruction network is a three-dimensional pose reconstruction network, and the three-dimensional pose reconstruction network includes nine processing modules, and each processing module is respectively used to construct the 3D pose of the following specific joints: The first processing module is used to construct the 3D joints of the nose, neck, left shoulder, right shoulder, left hip and right hip; The second processing module is used to construct the 3D joint of the left elbow; The third processing module is used to construct the 3D joint of the right elbow; The fourth processing module is used to construct the 3D joint of the left knee; The fifth processing module is used to construct the 3D joint of the right knee; The sixth processing module is used to construct the 3D joint of the left hand; The seventh processing module is used to construct the 3D joint of the right elbow; The eighth processing module is used to construct the 3D joint of the left foot; The ninth processing module is used to construct the 3D joint of the right foot.
4. The traffic police gesture recognition method according to claim 3, characterized in that Each processing module of the three-dimensional pose reconstruction network includes: An LSTM unit for processing time series features; A depth fusion layer for fusing the input two-dimensional pose features with the estimated three-dimensional joint features.
5. The traffic police gesture recognition method according to claim 3, characterized in that, The depth fusion layer performs feature fusion through the following steps: Concatenate the input two-dimensional pose features with the three-dimensional joint features output by the previous processing module; Process the concatenated features through a fully connected layer to generate the fused features.
6. According to the traffic police gesture recognition method based on pose estimation and sequence modeling described in claim 3, characterized in that The input of the three-dimensional pose reconstruction network is a two-dimensional pose vector, and the output is a three-dimensional pose vector, and is converted through a mapping function.
7. The traffic police gesture recognition method according to claim 1, wherein The sequence modeling network is a long short-term memory network, and the long short-term memory network includes multiple time steps, and each time step processes the pose information at the current moment and the hidden state at the previous moment to generate the temporal features of the dynamic gesture.
8. The traffic police gesture recognition method according to claim 7, wherein The long short-term memory network manages the information flow in the time series through a gating mechanism, including a forget gate, an input gate and an output gate, to gradually accumulate the change trend of the pose and forget irrelevant actions.
9. The traffic police gesture recognition method according to claim 7, wherein, The output vector of the LSTM network is processed through a fully connected layer and a Softmax function to calculate the probability of each preset gesture category and output the predicted gesture.
10. A traffic police gesture recognition system based on 3D pose reconstruction and LSTM, characterized in that, It includes: A key point extraction module for extracting the skeleton key points from the traffic police command gesture video to obtain multiple key points related to gesture recognition; A pose reconstruction network module for performing pose estimation on the extracted key points. The pose reconstruction network includes multiple processing modules, each module for constructing the pose of a specific joint and fusing the input pose features with the estimated pose features through a feature fusion layer, and finally outputting the complete pose estimation result; A sequence modeling network module for processing the pose information output by the pose reconstruction network module, modeling the dynamic sequence of traffic police command gestures through sequence modeling, identifying the semantic meaning of dynamic gestures, and completing gesture recognition.
Citation Information
Patent Citations
Traffic police gesture recognition method based on convolutional posture machine and long-short-term memory network
CN110287844A
Traffic police command gesture recognition method based on skeleton joint point sequence
CN110837778A
Traffic police gesture recognition method and system based on skeleton recognition and gated loop network
CN115188080A
Behavior recognition artificial intelligence network system and method for efficient recognition of hand signals and gestures
US20230196841A1