Human body abnormal behavior detection method and system
By using the STGCN-DLSK network model in chemical plants for abnormal behavior detection, the safety accident problems caused by workers' irregular operations and unsafe behaviors are solved, and efficient and accurate abnormal behavior identification and early warning are achieved, reducing accident risks and supervision costs.
Patent Information
- Application Number
- CN202510228782.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
Chemical plant workers may experience irregular operations and unsafe behaviors during operation, resulting in safety accidents. It is difficult for the existing technology to effectively identify and warn of these abnormal behaviors.
The STGCN-DLSK network model based on graph convolutional neural network (GCN) is adopted to obtain human skeleton data through video data acquisition and openpose pose extraction algorithm, and a multi-scale feature extraction module (MS module) and a double-layer fusion block (DLSK module) are constructed to realize real-time detection and recognition of abnormal behaviors in humans.
It improves the efficiency and accuracy of abnormal behavior detection in humans, can promptly identify and warn of unsafe behaviors of workers, reduces the probability of accidents and the risk of casualties, and reduces the supervision costs.
Smart Images

Figure CN120183036A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of monitoring and security, and information technology, and particularly relates to a method and system for detecting abnormal human behaviors. Background Art
[0002] Work safety accidents not only bring serious economic losses to society, but also pose a great threat to life safety. Therefore, we should sound the alarm, attach great importance to and strengthen the safety work in production, and improve the accuracy of prediction and early warning. When seeking effective ways to reduce the accident rate and casualties, the industrial and academic communities have gradually realized that in many cases, accidents may be caused by non-standard operations and unsafe behaviors of workers during the operation process, such as causing the safety device to fail artificially, smoking, falling, running, etc. Therefore, it is particularly important to monitor and identify the behaviors of chemical plant workers. An accurate identification system can more timely detect and handle the unsafe behaviors of workers, enabling plant managers to have more time to handle emergencies, reduce the probability of accidents, and reduce casualties.
[0003] In recent years, deep learning technologies, especially graph convolutional neural networks (GCNs), have made remarkable progress in the field of human behavior recognition. Compared with traditional recognition methods, deep learning technologies have lower costs and higher recognition accuracies, providing new ideas and methods for the detection of abnormal human behaviors. The method and system for detecting abnormal human behaviors based on graph convolutional neural networks can real-time monitor and identify the unsafe behaviors of chemical plant workers, providing timely and accurate early warning information for managers. This not only helps managers quickly take intervention measures to correct the unsafe behaviors of workers, but also effectively prevents potential safety accidents, reducing the probability of accidents and the risk of casualties. Summary of the Invention
[0004] Object of the Invention: In order to effectively improve the efficiency of detecting abnormal human behaviors, reduce the probability of chemical production accidents, reduce casualties, and also reduce the supervision cost, the present invention discloses a method and system for detecting abnormal human behaviors, and realizes the detection of abnormal human behaviors through the established STGCN-DLSK network model.
[0005] Technical Solution: The present invention provides a method for detecting abnormal human behaviors, including the following steps:
[0006] Step 1: For the abnormal behavior characteristics of the human body, collect video data and establish an abnormal behavior data set;
[0007] Step 2: Crop the collected video data, and each frame will pass through the openpose pose extraction algorithm to obtain human bone data;
[0008] Step 3: Construct the STGCN-DLSK network model. An MS module is constructed in the backbone feature extraction network of the STGCN-DLSK network model. Human skeleton data enters the STGCN-DLSK network model and serves as the input to the double-layer fusion block DLSK after passing through the graph convolutional network GCN and the temporal convolutional network TCN, and is output after passing through the double-layer fusion block DLSK. The double-layer fusion block DLSK is a module formed by combining the LSK network and the dense connection network. The MS module processes the data after different layers of STGCN, and multi-scale fusion features are obtained through the MS module.
[0009] Step 4: Train the STGCN-DLSK network model. Use the cross-entropy loss function to evaluate the performance of the model in the abnormal behavior detection task, and adjust the model weights through the optimization algorithm SGD, and save the best model weights obtained during the training process.
[0010] Step 5: Apply the trained STGCN-DLSK network model to the real-time video stream for abnormal behavior detection. The detection results will be displayed on the front-end user interface developed based on PyQt, which is convenient for managers to identify potential risks in a timely manner and take corresponding measures to ensure safety.
[0011] Further, the human skeleton data obtained by the openpose pose extraction algorithm in step 2 is as follows:
[0012] Step 2.1: Crop the video of human behavior with the frame size of W×H as the input, and extract its features through the backbone network VGG19.
[0013] Step 2.2: The feature map passes through two parallel branch networks to obtain the Part Confidence Maps (PCM confidence maps) and the Part Affinity Fields (PAF association fields) respectively.
[0014] Step 2.3: According to the above two sets of information, use the Bipartite Matching algorithm to connect the key points belonging to the same person.
[0015] Step 2.4: Merge the connected key points into the overall skeleton of each person, and output the pose information of each person in the video.
[0016] Furthermore, in the STGCN-DLSK network model, four convolutional layers with 32 channels are added at the very front of the original STGCN network to enhance the low-level feature extraction ability. The optimized STGCN network structure has thirteen layers. The MS module processes the data after the seventh layer (64 channels), the tenth layer (128 channels), and the thirteenth layer (256 channels), and the generated feature maps are concatenated with the features of the fourth layer (32 channels) Figure 1 After upsampling, the concatenated result contains temporal features of different scales. By establishing connections between feature maps of different scales and fusing multi-scale information, the expressiveness of the model is improved.
[0017] Furthermore, each STGCN convolutional layer includes a graph convolutional network (GCN), a temporal convolutional network (TCN), and a double-layer fusion block (DLSK).
[0018] Furthermore, the MS module consists of a dimensionality reduction convolutional layer, a multi-scale feature extraction layer, and a feature fusion layer. The specific steps are as follows:
[0019] Step 3.1: When the human skeleton data enters the MS module, it is first input into the dimensionality reduction convolutional layer for 1×1 convolution to reduce the features of the information and form the original feature map.
[0020] Step 3.2: The multi-scale feature extraction layer performs multi-feature extraction on the original feature map. The multi-feature extraction layer includes dilated convolutions with different dilation rates (2, 4, 8, 16) to capture spatio-temporal features of different scales, and feature maps of different scales are obtained respectively.
[0021] Step 3.3: In the feature fusion layer, the feature maps from different branches are concatenated together by channels. The width and height of the feature maps remain unchanged, and the feature maps of different branches are superimposed in the channel dimension to form a fused feature map. The fused feature map retains the feature information of different scales extracted by each convolutional kernel. The fused feature map is compressed through a 1×1 convolution operation to refine the feature information, and finally, multi-scale fused features are generated.
[0022] Furthermore, the specific structure of the double-layer fusion block DLSK is as follows:
[0023] Step 4.1: The multi-scale fused features output by the MS module enter the STGCN module and are used as the input of the double-layer fusion block DLSK after passing through the GCN and TCN. When the features are input, they are divided into two parallel branches and one serial branch. The parallel branches are dense connection networks, and the serial branch is the LSK network.
[0024] Step 4.2: After extracting features through the dense connection network and the LSK network, the feature dimensions are added residually, and the output features are used as the input of the next module.
[0025] A human abnormal behavior detection system, comprising:
[0026] A data acquisition module, where data acquisition is mainly obtained through a monitoring camera and uploaded locally;
[0027] A pose estimation module, which converts the acquired video data into picture frames and obtains human body bone data through the openpose algorithm;
[0028] An abnormal behavior detection module, which uses the STGCN-DLSK network model to detect abnormal behaviors in the input human body bone data;
[0029] A front-end display and warning module, where the abnormal behaviors in the behavior recognition results will be displayed on the front end so that managers can judge whether it is necessary to remind their behavior status through an alarm bell;
[0030] A data storage and management module, which saves the detected abnormal behavior results and the data that managers remind employees through the alarm bell to the local;
[0031] The STGCN-DLSK network model specifically includes: constructing an MS module in the backbone feature extraction network. The human body bone data enters the STGCN-DLSK network model, and after passing through the graph convolutional network GCN and the temporal convolutional network TCN, it is used as the input of the double-layer fusion block DLSK and then output after passing through the double-layer fusion block DLSK; the double-layer fusion block DLSK is a module formed by combining the LSK network and the dense connection network; the MS module processes the data after different layers of STGCN, and multi-scale fusion features are obtained through the MS module.
[0032] Further, the double-layer fusion block DLSK is specifically as follows: the multi-scale fusion features output by the MS module enter the STGCN module, and after passing through GCN and TCN, they are used as the input of the double-layer fusion block DLSK; when the features are input, they are divided into two parallel branches and one serial branch. The parallel branch is the dense connection network, and the serial branch is the LSK network; after extracting features through the dense connection network and the LSK network, the feature dimensions are added residually, and then the output features are used as the input of the next layer of the module.
[0033] Advantages: (1) The present invention uses the Openpose algorithm to obtain human skeletal joint point information, realizing behavior recognition without the need for human appearance feature information, and being less affected by the external environment. (2) The MS module constructed in the present invention is composed of multi-scale feature extraction modules. It can not only be used to extract features of different resolutions and regions, but also fuse richer and more comprehensive context information, enhancing the model's representation ability and adaptability, and reducing the risks of information loss and overfitting. (3) The double-layer fusion block DLSK adopted in the present invention is a new module formed by combining dense connections and LSK blocks. This module can enhance the expression ability of internal features, reduce the number of parameters and computational complexity, promote information flow and gradient propagation, and enhance the ability to obtain receptive fields and context information. At the same time, it also helps to construct a more efficient and better-performing deep learning model. (4) A simple front-end interface is developed to display and give early warnings of the model recognition results. Supervisors can obtain whether there are abnormalities in the behavior of employees through the front-end interface prompts, and judge whether it is necessary to remind workers of the correct behavior status through an alarm bell. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is the system flowchart of the present invention;
[0035] Figure 2 is the 13-layer network model of the present invention;
[0036] Figure 3 is the network structure diagram of the STGCN-DLSK algorithm;
[0037] Figure 4 is the MS module diagram;
[0038] Figure 5 is the structure diagram of the STGCN convolutional layer;
[0039] Figure 6 is the structure diagram of the DLSK module;
[0040] Figure 7 is the structure diagram of the LSK module;
[0041] Figure 8 is the interface diagram of the human abnormal behavior detection system. DETAILED DESCRIPTION OF THE INVENTION
[0042] The present invention will be further described in detail with reference to the following drawings. The processes, conditions, implementation methods, etc. for implementing the present invention, except for the specifically mentioned content below, are all common knowledge in the art, and the present invention has no particularly restricted content.
[0043] The present invention provides a human abnormal behavior detection system. The system consists of a data acquisition module, a body posture estimation module, an abnormal behavior detection module, a front-end display interface and an early warning module, and a data storage and management module. The data acquisition module includes a network camera for real-time monitoring of employees. The posture estimation module and the abnormal behavior detection module use the openpose algorithm and the STGCN-DLSK model to process the input video data and obtain the detection results. The front-end display and early warning module is used to display the detection results of employees' abnormal behaviors. Supervisors can view the detection results through the display interface and judge whether it is necessary to remind employees of their abnormal behaviors through an alarm bell. The data storage and management module is used for data storage for subsequent management.
[0044] The present invention also proposes a method for detecting human abnormal behaviors using the above system, as Figure 1 shown in the flowchart of this method, which specifically includes the following steps:
[0045] Step 1: Establish an abnormal behavior dataset according to the characteristics of abnormal behaviors of chemical plant workers. First, define abnormal behaviors, such as vomiting, coma caused by leakage of industrial hazardous gases in the factory, or behaviors such as illegally touching chemical raw materials with hands, climbing production equipment, eating, damaging chemical production equipment, and operating with hands instead of tools. Further, obtain a video dataset of workers corresponding to these behaviors.
[0046] Step 2: For the collected video data, convert it into picture frames, use the openpose algorithm to obtain the information of workers' skeletal joints, and then label and organize the data to obtain a skeletal dataset of workers' abnormal behaviors.
[0047] (1) Take the employee behavior video cropped to size W×H as the input, and extract its feature F through the backbone network VGG19;
[0048] (2) The feature map will pass through two parallel branch networks to obtain Part Confidence Maps (PCM confidence maps) and Part Affinity Fields (PAF correlation fields) respectively;
[0049] (3) According to the above two sets of information, use the Bipartite Matching algorithm to connect the key points belonging to the same person;
[0050] (4) Merge the connected key points into the overall skeleton of each person, and output the posture information of each person in the video.
[0051] Step 3: Construct an STGCN-DLSK network model, as Figure 3The STGCN-DLSK network model shown in the figure constructs an MS module in the backbone feature extraction network, and the human skeleton data enters the STGCN-DLSK network model, and is used as the input of the double-layer fusion block DLSK after passing through the graph convolution network GCN and the temporal convolution network TCN, and is output after passing through the double-layer fusion block DLSK; the double-layer fusion block DLSK is a module formed by combining the LSK network and the dense connection network; the MS module processes the data after different layers of STGCN, and obtains multi-scale fusion features through the MS module;
[0052] Step 4: Train the STGCN-DLSK network model, use the cross entropy loss function to evaluate the performance of the model on the abnormal behavior detection task, adjust the model weights through the optimization algorithm SGD, and save the best model weights obtained during the training process;
[0053] Step 5: Apply the trained STGCN-DLSK network model to the real-time video stream for abnormal behavior detection. The detection results will be displayed on the front-end user interface developed based on PyQt, so that managers can identify potential risks in time and take corresponding measures to ensure safety.
[0054] Further, such as Figure 2 The STGCN-DLSK network model adds four 32-channel convolutional layers to the front end of the original STGCN network to enhance the low-level feature extraction capability. The optimized STGCN network structure has thirteen layers. The MS module processes the data after the seventh layer (64 channels), the tenth layer (128 channels), and the thirteenth layer (256 channels). The generated feature map will be compared with the feature map of the fourth layer (32 channels). Figure 1 After upsampling, the images are spliced together, and the spliced results contain temporal features of different scales. By establishing connections between feature maps of different scales, multi-scale information is fused, which improves the expressiveness of the model.
[0055] Further, such as Figure 5 Each STGCN convolutional layer includes a graph convolutional network GCN, a temporal convolutional network TCN, and a double-layer fusion block DLSK.
[0056] like Figure 4 The MS module consists of a dimensionality reduction convolution layer, a multi-scale feature extraction layer and a feature fusion layer.
[0057] When human skeleton data enters the MS module, it is first input into the dimension reduction convolution layer for 1×1 convolution to reduce the dimension of the information and form the original feature map;
[0058] The multi-scale feature extraction layer performs multi-feature extraction on the original feature map. The multi-scale feature extraction layer includes dilated convolutions with different dilation rates (2, 4, 8, 16) to capture spatio-temporal features at different scales, and feature maps at different scales are obtained respectively;
[0059] In the feature fusion layer, the feature maps from different branches are concatenated together by channel. The width and height of the feature maps remain unchanged, and the feature maps of different branches are superimposed in the channel dimension to form a fused feature map. The fused feature map retains the feature information at different scales extracted by each convolutional kernel. The fused feature map is compressed through a 1×1 convolutional operation to refine the feature information, and finally multi-scale fused features are generated.
[0060] The specific structure of the double-layer fusion block DLSK is as follows:
[0061] The multi-scale fused features output by the MS module enter the STGCN module and are used as the input of the double-layer fusion block DLSK after passing through GCN and TCN; when the features are input, they are divided into two parallel branches and one serial branch. The parallel branch is a densely connected network, and the serial branch is an LSK network;
[0062] After extracting features through the densely connected network and the LSK network, the feature dimensions are added residually, and the subsequently output features are used as the input of the next-layer module. Figure 6 is the structural diagram of the DLSK module; Figure 7 is the structural diagram of the LSK module.
[0063] Figure 8 is the interface diagram of the human abnormal behavior detection system. The trained and optimized abnormal behavior recognition model for chemical plant workers is used to recognize the behaviors of workers under the monitoring of the chemical plant. By frameizing the worker videos in the monitoring and sending them into the employee abnormal behavior detection model trained on the local server for detection, the recognition results of worker abnormal behaviors are obtained and saved to the local server. The obtained recognition results of worker abnormal behaviors are displayed and warned in real time on the front-end interface for the convenience of supervisors to handle, and it is judged whether it is necessary to remind the workers of the correctness of their behavior status through an alarm bell.
[0064] The present invention also discloses a chemical industrial park worker abnormal behavior detection system, including:
[0065] A data acquisition module, and data acquisition is mainly obtained through monitoring cameras and uploaded locally;
[0066] A pose estimation module. For the acquired video data, it is converted into picture frames, and human bone data is obtained through the openpose algorithm;
[0067] The abnormal behavior detection module uses the trained STGCN-DLSK model to detect abnormal behaviors in the input human skeleton data;
[0068] The front-end display and warning module will display the abnormal behaviors in the behavior recognition results at the front end so that the management personnel can judge whether it is necessary to use an alarm to remind of the behavior status;
[0069] The data storage and management module saves the detected abnormal behavior results and the data of the management personnel reminding the employees through the alarm to the local for future management.
[0070] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent transformation or modification made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for detecting abnormal human behavior, characterized in that: The following steps are involved: Step 1: Collect video data based on the abnormal behavior characteristics of the human body and establish an abnormal behavior data set; Step 2: Crop the collected video data, and each frame will be subjected to the openpose posture extraction algorithm to obtain human skeleton data; Step 3: Construct an STGCN-DLSK network model. The STGCN-DLSK network model constructs an MS module in the backbone feature extraction network. The human skeleton data enters the STGCN-DLSK network model, and after passing through the graph convolution network GCN and the temporal convolution network TCN, it is used as the input of the double-layer fusion block DLSK, and output after passing through the double-layer fusion block DLSK; the double-layer fusion block DLSK is a module formed by combining the LSK network and the dense connection network; the MS module processes the data after different layers of STGCN, and obtains multi-scale fusion features through the MS module; Step 4: Train the STGCN-DLSK network model, use the cross entropy loss function to evaluate the performance of the model on the abnormal behavior detection task, adjust the model weights through the optimization algorithm SGD, and save the best model weights obtained during the training process; Step 5: Apply the trained STGCN-DLSK network model to the real-time video stream for abnormal behavior detection. The detection results will be displayed on the front-end user interface developed based on PyQt, so that managers can identify potential risks in time and take corresponding measures to ensure safety.
2. A method for detecting abnormal human behavior according to claim 1, characterized in that: In step 2, the human skeleton data obtained by the openpose posture extraction algorithm is as follows: Step 2.1: The human action video with the screen size cropped to W×H is used as input, and its features are extracted through the backbone network VGG19; Step 2.2: The feature map will pass through two parallel branch networks to obtain the PCM confidence map and PAF correlation field respectively; Step 2.3: Based on the above two sets of information, use the even matching algorithm to connect the key points belonging to the same person; Step 2.4: Merge the connected key points into the overall skeleton of each person and output the posture information of each person in the video.
3. A method for detecting abnormal human behavior according to claim 1, characterized in that: The STGCN-DLSK network model adds four 32-channel convolutional layers to the front end of the original STGCN network. The optimized STGCN network has thirteen layers. The MS module processes the data after the seventh, tenth, and thirteenth layers. The generated feature map is spliced with the feature map of the fourth layer after upsampling. The spliced result contains temporal features of different scales.
4. A method for detecting abnormal human behavior according to claim 3, characterized in that: Each STGCN convolutional layer includes a graph convolutional network GCN, a temporal convolutional network TCN, and a double-layer fusion block DLSK.
5. A method for detecting abnormal human behavior according to claim 1, characterized in that: The MS module is composed of a dimensionality reduction convolution layer, a multi-scale feature extraction layer and a feature fusion layer. The specific steps are as follows: Step 3.1: When information enters the MS module, it is first input into the dimension reduction convolution layer for 1×1 convolution to reduce the dimension of the information and form the original feature map; Step 3.2: The multi-scale feature extraction layer performs multi-feature extraction on the original feature map, wherein the multi-feature extraction layer includes dilated convolutions with different dilation rates to capture spatiotemporal features of different scales, and obtain feature maps of different scales respectively; Step 3.3: In the feature fusion layer, feature maps from different branches are connected together by channel. The width and height of the feature map remain unchanged. The feature maps of different branches are superimposed on the channel dimension and combined into a fused feature map. The fused feature map retains the feature information of different scales extracted by each convolution kernel. The fused feature map is compressed through a 1×1 convolution operation to refine the feature information and finally generate multi-scale fused features.
6. A method for detecting abnormal human behavior according to claim 1, characterized in that: The double-layer fusion block DLSK is as follows: The information enters the STGCN module and serves as the input of the double-layer fusion block DLSK after passing through GCN and TCN. When the feature is input, it is divided into two parallel branches and one serial branch. The parallel branch is a densely connected network and the serial branch is an LSK network. After extracting features through the densely connected network and the LSK network, the residuals of the feature dimensions are added, and the output features are then used as the input of the next layer module.
7. A human abnormal behavior detection system, characterized in that: include: Data collection module: data collection is mainly obtained through surveillance cameras and uploaded locally; The posture estimation module converts the collected video data into picture frames and obtains human skeleton data through the openpose algorithm; The abnormal behavior detection module uses the STGCN-DLSK network model to detect abnormal behavior of the input human skeleton data; Front-end display and early warning module: abnormal behaviors in behavior recognition results will be displayed on the front-end, so that managers can judge whether it is necessary to use alarms to remind them of their behavior status; The data storage and management module stores the detected abnormal behavior results and the data that managers use to remind employees through alarms locally; The STGCN-DLSK network model specifically includes: constructing an MS module in the backbone feature extraction network, human skeleton data enters the STGCN-DLSK network model, and after passing through the graph convolution network GCN and the temporal convolution network TCN, it is used as the input of the double-layer fusion block DLSK, and output after passing through the double-layer fusion block DLSK; the double-layer fusion block DLSK is a module formed by combining the LSK network and the densely connected network; the MS module processes the data after different layers of STGCN, and obtains multi-scale fusion features through the MS module.
8. A human abnormal behavior detection system according to claim 7, characterized in that: The double-layer fusion block DLSK is specifically as follows: the multi-scale fusion features output by the MS module enter the STGCN module, and after passing through GCN and TCN, they are used as the input of the double-layer fusion block DLSK; when the features are input, they are divided into two parallel branches and one serial branch, the parallel branch is a densely connected network, and the serial branch is an LSK network; after extracting features through the densely connected network and the LSK network, the residuals of the feature dimensions are added, and the output features are then used as the input of the next layer module.