A method for identifying park visitors' behaviors characterized by human key-point heat maps
By extracting the heat map of the human body's key points from the park surveillance video and building a behavior recognition network based on deep learning, the accuracy and robustness of behavior recognition in the park environment are solved, and efficient and real-time monitoring of tourists' behavior is achieved.
Patent Information
- Application Number
- CN202310178175.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-02-28
AI Technical Summary
The existing deep learning behavior recognition method based on video frames has problems such as unsatisfactory recognition accuracy, high missed detection rate, poor robustness in park environments, and has high calculation cost and high time consumption.
By extracting frames on the park surveillance video, the human body's key points heat map is extracted, and a behavior recognition network based on the key points heat map is built, combining the attention mechanism and the ResNet18 network to identify and classify tourists' behavior.
It improves the accuracy and robustness of park tourists' behavior recognition, reduces the missed detection rate, and reduces the calculation cost, ensuring real-time and recognition effect.
Smart Images

Figure CN116580447B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of public safety protection for citizens in outdoor public places, and specifically to a method for identifying park visitors' behaviors based on deep learning and featuring human key point heat maps. Background Art
[0002] In the process of urbanization in China, the increasing number of urban parks has met the growing living needs of citizens, promoted the high-quality development of urban construction, and satisfied people's yearning for a better life. However, social hot issues regarding park safety management have also emerged, such as safety incidents like the elderly falling, children tripping, sitting for a long time without getting up, and even fighting occurring in parks from time to time. Currently, the commonly used park safety supervision method is to install video surveillance in key areas and identify visitors' unsafe behaviors by means of manual remote video online viewing or post-event viewing. This method has problems such as low efficiency, high cost, and poor stability, and is prone to causing park safety accidents due to inattentiveness or lack of experience of staff, resulting in the lack of guarantee for visitors' safety and triggering social stability issues. Therefore, how to improve the identification of park visitors' unsafe behaviors is of great significance.
[0003] Currently, the behavior recognition technology based on deep learning is relatively mature and has been successfully applied in aspects such as intelligent monitoring and safety protection. For the behavior recognition task, scholars have used deep learning technology and built various types of behavior recognition models with videos or image frames as inputs. For example, Patent CN202211020833.5 proposes a method and system for identifying abnormal behaviors of pedestrians based on multi-modal deep learning technology. This method integrates three detection processes of pedestrian detection, pedestrian attribute recognition, and single pedestrian abnormal target detection and progressively identifies abnormal behaviors, achieving the identification of abnormal behaviors of pedestrians. Patent CN202211230844.6 proposes a method for identifying dangerous behaviors of pedestrians based on deep learning. A small target detection layer is added to the YOLOv5s algorithm, and through the intelligent analysis of surveillance video images, dangerous behaviors with potential safety hazards are identified, and real-time warnings are given during the event to alert the management staff. CN202110050206.5 proposes a method for identifying pedestrian attributes based on deep learning, builds a pedestrian attribute recognition network with a convolutional neural network as the core, introduces an attention module between the backbone network and the first pooling layer, emphasizes the global features corresponding to high-level semantic information, and improves the accuracy of behavior recognition. The above patents all use a convolutional neural network or a YOLO network as the core and identify behaviors with videos as inputs.
[0004] However, the quality of outdoor surveillance videos represented by park environments is vulnerable to factors such as weather, perspective, and occlusion. Although existing deep learning methods can effectively extract human targets, it is difficult to achieve ideal results in behavior recognition, such as indicators like accuracy, miss detection rate, and robustness. In addition, in video data processing, there are also defects such as high computational costs and high time consumption. To address the above problems, the present invention aims to improve the performance of park visitor behavior recognition. By performing frame extraction on the video, using the heatmap of human key points as features, and introducing an attention mechanism to fuse existing deep learning models, a new behavior recognition model is constructed to overcome the deficiencies of existing deep learning models in behavior recognition in park environments, improve the safety guarantee for visitors, and reduce personal safety accidents caused by ineffective monitoring. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] For the outdoor human behavior recognition task represented by park environments, existing behavior recognition methods based on convolutional neural networks with video frames as input have defects such as unsatisfactory recognition accuracy, high miss detection rate, and poor robustness. The present invention aims to construct new sample data to overcome the deficiencies of video frames, and accordingly design a new deep learning-based park visitor behavior recognition model. Through training and learning, a park visitor behavior recognition method with better performance is constructed. The main technical problems to be solved include the extraction of human key heatmaps, the preprocessing of heatmaps, and the construction, training, and learning of a visitor behavior recognition model with heatmaps as input features.
[0007] Based on the existing convolutional neural network human recognition model, the present invention first crops the visitor targets in the park surveillance video frames, then extracts the corresponding heatmaps of human key points, and finally builds a park visitor behavior recognition model with the heatmaps of human key points as input features to achieve the behavior recognition of visitors, improving the behavior recognition accuracy and robustness and reducing the miss detection rate.
[0008] (2) Technical Solutions
[0009] The present invention at least includes the following steps:
[0010] S1: Obtain real-time video through cameras distributed on site, perform frame extraction on the video at a certain frequency to obtain video frames;
[0011] S2: Input the above video frames into a pre-trained human detection model based on a convolutional neural network to obtain human target boxes, and further process them to obtain a human cropped image a;
[0012] S3: Perform behavior category annotation on the cropped image a to obtain a labeled image a gt, where gt is the behavior category;
[0013] S4: Input picture a gt into a pre-trained convolutional neural network-based human key point detection model to obtain a human key point heat map A gt , and construct a labeled key point heat map data set accordingly;
[0014] S5: Use the key point heat map data set as the training sample to train the built tourist behavior recognition network with the key point heat map as the feature to obtain a heat map behavior recognition network model;
[0015] S6: According to S1 and S2, the cropped picture a passes through the human key point detection model described in S4, and the output heat map A is then input into the heat map behavior recognition network model described in S5 to obtain the behavior recognition result of park tourists.
[0016] Furthermore, the specific process (as shown in the appendix Figure 1 ) of the tourist behavior recognition network with the key point heat map as the feature described in S5 includes:
[0017] SS1: Heat map data preprocessing module: Perform threshold filtering on the obtained original heat map to filter out the key point heat map information with low confidence, and realize the preprocessing of the heat map. The filtering method is as follows:
[0018]
[0019] In the formula, H is the original heat map, is the indicator function, which outputs 1 when the condition in the parentheses is satisfied, otherwise 0, conf i is the confidence of the i-th key point, λ is the key point confidence threshold, represents the element-wise multiplication operator, and H′ is the filtered heat map;
[0020] SS2: Feature extraction and classification module: This module consists of two parallel branches with the same structure and a fusion module. Among them, each parallel branch consists of three sub-modules: an attention module, a ResNet18 network module, and a single-layer fully connected neural network L; the two parallel branches take H′ and H in SS1 as inputs respectively. After obtaining the preliminary recognition results O1 and O2 from the ResNet18 network module, the corresponding weights S1 and S2 of O1 and O2 are obtained after passing through the single-layer fully connected neural network L with the same structure. The specific calculation is as follows:
[0021] S i =σ(L i (O i )) i = 1, 2
[0022] In the formula, Li It represents a single-layer fully-connected neural network, and σ represents the Sigmoid function; finally, the fusion module performs fusion processing on the eigenvalues O1, O2, S1, and S2 extracted by the parallel branches to obtain the fused recognition result O of the behavior, that is, the final recognition result of the tourist behavior, realizing the recognition and classification of tourist behaviors. The specific calculation is as follows:
[0023]
[0024] In the formula, is the element-wise multiplication operator, is the element-wise addition operator.
[0025] (III) Beneficial effects
[0026] In outdoor environments represented by parks, there are generally phenomena such as complex backgrounds, common aggregation and occlusion, and easily variable environments. When common behavior recognition methods are used for park tourist behavior recognition, problems such as low accuracy and poor robustness are likely to occur. Especially for the occlusion phenomenon, the behavior recognition effect is not good. The present invention performs real-time frame extraction on the video, extracts the pedestrian target and the heat map of human body key points through a convolutional neural network model, and sequentially constructs a tourist behavior recognition network based on deep learning. Since the heat map of human body key points not only contains the corresponding features of the video frame, but also contains the internal relationship between key points, and the relationship between key points and behavior categories, by eliminating features with weak or insignificant supporting effects, the influence of factors such as occlusion and unclear features on behavior recognition is overcome, and the park tourist behavior recognition effect is improved. This network model reduces the computational cost by means of frame extraction processing, ensures real-time performance and improves the recognition effect, has strong operability and practicability, and is of great significance for real-time monitoring and recognition of dangerous behaviors of park tourists. Description of the drawings
[0027] Figure 1 is the flow chart of the steps implemented in the present invention;
[0028] Figure 2 is the sample of the heat map of 17 key points of the human body implemented in the present invention;
[0029] Figure 3 is the schematic diagram of the heat map behavior recognition network implemented in the present invention;
[0030] Table 1 shows the effect comparison of various behavior recognition methods. Specific implementation method
[0031] To make the objectives, content, and advantages of the present invention clearer, the following will, in conjunction with the accompanying drawings and embodiments of the present invention, provide a complete and clear description of the implementation solutions of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention fall within the scope of protection of the present invention.
[0032] The following will introduce the specific implementation solutions of the present invention in detail in conjunction with the accompanying drawings and embodiments.
[0033] Refer to Figure 1 , taking the outdoor video data in a certain park environment, the pre-trained YOLOv5 as the human detection model, and the pre-trained AlphaPose as the human key point detection model as embodiments, in combination with a park visitor behavior recognition method based on deep learning featuring human key point heat maps provided by the present invention, which includes:
[0034] (1) Video acquisition and frame extraction processing
[0035] Real-time crowd videos are obtained through 12 ordinary outdoor bullet cameras installed on a 3-meter-high pole, equipped with a 5G network, and the video data is synchronously transmitted to the NVR server in the park management center; the management center is equipped with a GPU server for deploying the intelligent analysis algorithm for visitor behavior safety; the GPU and the NVR server are networked to read the video data from the NVR server and perform frame extraction processing on the video at a certain frequency to obtain video frames;
[0036] (2) Human detection and behavior annotation based on YOLOv5
[0037] The above video frames are input into the pre-trained YOLOv5 model to obtain human target boxes, and after further processing, a human target cropped image a is obtained; the cropped image a is annotated with behavior categories to obtain a labeled picture a gt , where gt is the behavior category, and the behaviors to be recognized include 7 types: lying, squatting, standing, wrestling, sitting on a stool, lying prone, sitting on the ground, etc.;
[0038] (3) Extraction of human key point heat maps based on AlphaPose
[0039] The picture a gt is input into the pre-trained AlphaPose model, and a human key point heat map A gt (as shown in Figure 2 ) is output, and a labeled key point heat map data set is constructed accordingly;
[0040] (4) Construction of a behavior recognition network featuring key point heat maps
[0041] Refer to Figure 3 , the heatmap behavior recognition network consists of two major modules: heatmap data preprocessing, and feature extraction and classification:
[0042] (1) Heatmap data preprocessing
[0043] The obtained original heatmap is processed through a threshold to filter out the heatmap information of key points with low confidence. The filtering method is as follows:
[0044]
[0045] In the formula, H is the original heatmap, is the indicator function, conf i is the confidence of the i-th key point, λ is the key point confidence threshold, set to 0.05, is element-wise multiplication, and H′ is the filtered heatmap;
[0046] (2) Feature extraction and classification
[0047] It consists of two parallel branches and a fusion module. Both branches contain three sub-modules: an attention module, a ResNet18 network, and a single-layer fully connected neural network. Among them:
[0048] The attention module is before the ResNet18 network module. The number of channels in the convolutional layer of the ResNet18 network module is half of that of the original ResNet18 network module. The heatmap passes through the attention module and the ResNet18 network module to obtain preliminary behavior recognition results O1 and O2 respectively;
[0049] The two single-layer fully connected convolutional neural network modules respectively integrate the information contained in O1 and O2, and obtain corresponding weights S1 and S2 respectively. The specific calculation is as follows:
[0050] S i =σ(L i (O i )) i = 1, 2
[0051] In the formula, L i represents a single-layer fully connected neural network, and σ represents the Sigmoid function;
[0052] The fusion module performs fusion processing on the feature values O1, O2 and weights S1, S2 extracted by the parallel branches to obtain the fusion recognition result O of the tourist behavior. The specific calculation is as follows:
[0053]
[0054] In the formula, represents element-wise multiplication, Indicates element-wise addition.
[0055] (5) Training of the key point heatmap behavior recognition network
[0056] Using the key point heatmap dataset as the training sample, train the constructed key point heatmap behavior recognition network to obtain the heatmap behavior recognition network model;
[0057] (6) Park visitor behavior recognition based on the key point heatmap behavior recognition model
[0058] Input the key point heatmap into the trained heatmap behavior recognition network model to obtain the behavior recognition result of park visitors. For comparison, Table 1 shows the effects of various methods in human behavior recognition. From the indicators such as precision, F1 score, and mAP, the performance of the heatmap behavior recognition model proposed in the present invention is better than other methods.
[0059] Table 1 Comparison of human behavior recognition effects
[0060] Network Precision (%) F1(%) mAP AlexNet 45.56 45.96 49.94 VGG16 53.26 50.39 52.47 VGG19 62.19 56.98 53.42 ResNet18 51.12 50.44 54.81 GoogLeNet 44.34 45.05 62.03 DenseNet 52.5 50.47 61.35 MobileNetV2 67.83 57.23 60.02 MobileNetV3 58.42 57.51 61.29 ShuffleNetV2 56.12 56.8 59.07 SqueezeNet 44.22 45.36 41.14 Heatmap Behavior Recognition Network 71.58 62.39 65.74
Claims
1. A method for identifying park visitors' behaviors characterized by human key point heat maps, which is characterized in that, It includes the following steps: S1: Obtain real-time video through cameras distributed on-site, and perform frame extraction on the video at a certain frequency to obtain video frames; S2: Input the above video frames into a pre-trained human detection model based on a convolutional neural network to obtain human target boxes, and further process them to obtain human cropped image a; S3: Perform behavior category annotation on the cropped image a to obtain the labeled image a gt , where gt is the behavior category; S4: Input image a gt into a pre-trained human key point detection model based on a convolutional neural network to obtain a human key point heat map A gt , and construct a labeled key point heat map data set accordingly; S5: Use the key point heat map data set as the training sample to train the tourist behavior recognition network built with the key point heat map as the feature, and obtain the heat map behavior recognition network model; S6: According to S1 and S2, the cropped image a obtained is input into the human key point detection model described in S4, and the output heat map A is then input into the heat map behavior recognition network model described in S5 to obtain the behavior recognition result of park tourists; The specific process of the tourist behavior recognition network with the key point heat map as the feature described in S5 includes two major modules: heat map data preprocessing, feature extraction and classification, where: SS1: Heat map data preprocessing module: Perform threshold filtering on the obtained original heat map to filter out the key point heat map information with low confidence, and achieve the preprocessing of the heat map. The filtering method is as follows: where H is the original heatmap, is the indicator function, which outputs 1 when the condition in the parentheses is satisfied and 0 otherwise, conf i is the confidence of the i-th key point, λ is the key point confidence threshold, represents the element-wise multiplication operator, and H′ is the filtered heatmap; SS2: Feature extraction and classification module: This module consists of two parallel branches with the same structure and a fusion module. Among them, each parallel branch consists of three sub-modules: an attention module, a ResNet18 network module, and a single-layer fully connected neural network L; the two parallel branches use H′ and H in SS1 as inputs respectively. After obtaining the preliminary recognition results O1 and O2 from the ResNet18 network module, the corresponding weights S1 and S2 of O1 and O2 are obtained after passing through the single-layer fully connected neural network L with the same structure. The specific calculation is as follows: Si = σ(Li(Oi)) i = 1, 2 where, L i represents a single-layer fully-connected neural network, and σ represents the Sigmoid function; finally, the fusion module performs a fusion process on the eigenvalue O1, O2 and S1, S2 extracted by the parallel branches to obtain the fused recognition result O of the behavior, that is, the final recognition result of the tourist behavior, so as to realize the recognition and classification of the tourist behavior. The specific calculation is as follows: In the formula, is the element-wise multiplication operator, and ⊕ is the element-wise addition operator.
Citation Information
Patent Citations
Pedestrian attribute identification method based on deep learning
CN114764919A
Pedestrian dangerous behavior recognition method based on deep learning
CN115294661A
Pedestrian abnormal behavior identification method and system based on multi-modal deep learning technology
CN115346175A
Smoking behavior detection method based on human body posture estimation and image classification
CN112528960A
Posture correction pedestrian re-identification method based on convolutional generative adversarial network
CN114639122A