A method for acquiring and identifying spatiotemporal features of risky behaviors based on three-dimensional depth vision
Through the spatiotemporal and spatial characteristics acquisition and identification method of hazard-induced behavior based on three-dimensional deep vision, and using depth sensors and deep learning models, real-time intelligent monitoring of workshop actors is realized, solving the problem that existing systems cannot achieve intelligent perception and timely feedback.
Patent Information
- Application Number
- CN202011467738.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-12-14
AI Technical Summary
The existing workshop safety monitoring system cannot realize the digitalization and intelligent perception of the behavior of the perpetrator, and cannot promptly provide feedback and solve the risk factors of the perpetrator in production.
The spatiotemporal and spatial feature acquisition and identification method based on three-dimensional depth vision is adopted. The skeleton data of the perpetrator is collected through depth sensors, and an online spatiotemporal and spatial feature acquisition and identification model is designed, including a time feature extraction module and a spatial feature extraction module to realize real-time intelligent monitoring.
Real-time intelligent monitoring of the perpetrator is realized, and predefined behavior categories can be identified without the participation of the perpetrator without the participation of the security management actor, reducing safety hazards, and improving the multi-dimensional and multi-level monitoring.
Smart Images

Figure CN114694240B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of safety monitoring, and in particular relates to a method for acquiring and identifying spatiotemporal features of risky behaviors based on three-dimensional depth vision. Background Art
[0002] Human behavior recognition technology is a basic technology in many fields such as intelligent monitoring and human-computer interaction. Through this technology, we can realize the perception of human behavior in different scenes in life, such as hospital wards and elevators, and timely alarm for abnormal behaviors. The production workshop has the characteristics of large scale, many workstations, complex environment and high-risk key processes. The uncertainty of the actor's state may have a significant impact on the production safety and production efficiency of the workshop. The safety monitoring system of the workshop generally only contains color or infrared video stream data, which cannot realize the digitization and intelligent perception of the actor's behavior. Most of this data is used for post-event accountability. The dangerous factors of the actor in the workshop production cannot be fed back and resolved in time, and other means such as production actors wearing data collection equipment cannot be conveniently applied in complex production processes. Summary of the invention
[0003] In order to solve the above problems, the present invention provides a method for acquiring and identifying spatiotemporal features of risky behaviors based on three-dimensional depth vision. The method collects the behavioral data of the actors and trains an online spatiotemporal feature acquisition and recognition model. It can identify predefined behavior categories without the participation of security management actors to achieve real-time intelligent monitoring.
[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0005] A method for acquiring and identifying spatiotemporal features of risky behaviors based on three-dimensional depth vision, the method comprising the following steps:
[0006] Step 1: For the area where risky behaviors need to be monitored, a visual sensing network composed of depth sensors is constructed to continuously collect information about the people in the area in the form of video streams. The coordinate information of the skeleton joints of each behavior of the person is obtained through a human posture estimation algorithm based on depth sensors, so as to construct a risky behavior dataset;
[0007] Step 2: Preprocess the coordinate information of the skeleton joint points in the risky behavior dataset;
[0008] Step 3: Design an online spatiotemporal feature acquisition and recognition model for risky behaviors based on deep learning. The model includes a temporal feature extraction module and a spatial feature extraction module. The temporal feature extraction module based on the temporal convolutional network is responsible for extracting temporal features, and the spatial feature extraction module based on the graph neural network is responsible for extracting spatial features. The spatial feature extraction module includes a behavior temporal bounding box regression module and a frame-level behavior classification module. The behavior temporal bounding box regression module is used to predict the duration of the current behavior, and the frame-level behavior classification module is used to determine the behavior category at the current time point.
[0009] Step 4: Build a deep learning model, import the preprocessed skeleton joint point coordinate information into the risk-causing behavior online spatiotemporal feature acquisition and recognition model, and train the model with the classification regression module loss function;
[0010] Step 5: Use a depth sensor to collect the skeleton data of the actor in real time, build a behavior recognition system, output the data stream to the deep learning model for recognition, and store the recognition results in the database.
[0011] Preferably, the continuous collection of information on persons in the area in the form of video streams in step 1 includes: designing and identifying behavior categories according to requirements, and continuously collecting multi-stream data on persons.
[0012] Preferably, the preprocessing of the skeleton joint coordinate information in the risky behavior data set in step 2 includes:
[0013] Step 2.1: Translate the origin of the skeleton data coordinates collected by the depth sensor to the skeleton hip coordinate point to eliminate the influence of the human body position;
[0014] Step 2.2: Rotate the skeleton data collected by the depth sensor to eliminate the influence of the human body's orientation and posture;
[0015] Step 2.3: Considering the symmetry problem, the human skeleton data is flipped to achieve data enhancement and improve the robustness of the model;
[0016] Step 2.4: Model and express the human skeleton model based on the form of a topological graph.
[0017] Preferably, the temporal feature extraction module uses one-dimensional dilated convolution to input sequences of arbitrary length, and obtains temporal receptive fields of different sizes by using different numbers of layers.
[0018] Preferably, the behavior time bounding box regression module is composed of two graph convolution layers and one fully connected layer, and the final network output data size is 1×1, representing the duration of the current node behavior.
[0019] Preferably, the frame-level behavior classification module consists of two graph convolutional layers, one fully connected layer and one Softmax layer. The final output of the network is N×1, where N is the number of risky behavior categories. The output represents the confidence of each behavior category, and the sum of the confidences is 1.
[0020] Preferably, step 4 also includes: importing the preprocessed skeleton joint coordinate information into the risky behavior online spatiotemporal feature acquisition and recognition model, performing specified rounds of training, lowering the learning rate during training to achieve excellent training results, and ending the model training and saving the model when the accuracy rate increases and the loss function decreases to a flat level.
[0021] Preferably, the real-time acquisition of skeleton data of the actor using the depth sensor in step 5 includes: designing a dynamic link library file through C++ to provide an interface for Python to obtain the real-time skeleton of the depth sensor.
[0022] Compared with the prior art, the present invention has the following significant advantages:
[0023] 1. The present invention is based on computer vision and artificial intelligence technology to achieve virtual expression of the actor, extract the actor's temporal and spatial characteristics, and then intelligently perceive and identify the actor's behavior, so as to achieve the purpose of intuitive, transparent, and real-time intelligent behavior monitoring of the actor, reduce safety hazards, and realize multi-dimensional and multi-level monitoring and perception of the actor's behavior.
[0024] 2. The present invention collects the behavioral data of the actors and trains the online spatiotemporal feature acquisition and recognition model, so that predefined behavior categories can be identified without the participation of security management actors to achieve real-time intelligent monitoring.
[0025] 3. The present invention uses a temporal convolutional network to extract temporal features, and uses a graph neural network to extract the spatial features of the actor, which can well obtain behavioral features; the risky behavior data set collects the skeleton joint position flow information of each behavior of the actor, which can effectively reduce the impact of environmental factors such as light; through the online spatiotemporal feature acquisition and recognition model, the real-time skeleton data stream can be input, thereby realizing the discrimination of the actor's behavior at the current time point and the regression of the behavior time boundary box. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a flow chart for the implementation of the method for acquiring and identifying spatiotemporal features of risky behaviors based on three-dimensional depth vision in the present invention.
[0027] Figure 2 A human skeleton diagram represented by a topological diagram.
[0028] Figure 3 This is the overall structure diagram of the spatial-temporal feature extraction and recognition network for risky behavior.
[0029] Figure 4 for Figure 3 Network structure diagram of the sequential convolutional layer.
[0030] Figure 5 for Figure 4 Detailed structure diagram of each layer of the network.
[0031] Figure 6 for Figure 3 Diagram of the neural network layer structure. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] The implementation of the present invention is described in detail below in conjunction with specific embodiments.
[0034] The present invention provides a method for acquiring and identifying spatiotemporal features of risky behaviors based on three-dimensional depth vision, the method comprising the following steps:
[0035] Step 1: For the area where risky behaviors need to be monitored, a visual sensing network composed of depth sensors is constructed to continuously collect information about the people in the area in the form of video streams. The coordinate information of the skeleton joints of each behavior of the person is obtained through a human posture estimation algorithm based on depth sensors, so as to construct a risky behavior dataset;
[0036] Step 2: Preprocess the coordinate information of the skeleton joint points in the risky behavior dataset;
[0037] Step 3: Design an online spatiotemporal feature acquisition and recognition model for risky behaviors based on deep learning. The model includes a temporal feature extraction module and a spatial feature extraction module. The temporal feature extraction module based on the temporal convolutional network is responsible for extracting temporal features, and the spatial feature extraction module based on the graph neural network is responsible for extracting spatial features. The spatial feature extraction module includes a behavior temporal bounding box regression module and a frame-level behavior classification module. The behavior temporal bounding box regression module is used to predict the duration of the current behavior, and the frame-level behavior classification module is used to determine the behavior category at the current time point.
[0038] Step 4: Build a deep learning model, import the preprocessed skeleton joint point coordinate information into the risk-causing behavior online spatiotemporal feature acquisition and recognition model, and train the model with the classification regression module loss function;
[0039] Step 5: Use a depth sensor to collect the skeleton data of the actor in real time, build a behavior recognition system, output the data stream to the deep learning model for recognition, and store the recognition results in the database.
[0040] Preferably, the continuous collection of information on persons in the area in the form of video streams in step 1 includes: designing and identifying behavior categories according to requirements, and continuously collecting multi-stream data on persons.
[0041] Preferably, the preprocessing of the skeleton joint coordinate information in the risky behavior data set in step 2 includes:
[0042] Step 2.1: The coordinate origin of the skeleton data collected by the depth sensor is translated to the coordinate point of the skeleton hip to eliminate the influence of the human body position; Step 2.2: The skeleton data collected by the depth sensor is rotated to eliminate the influence of the human body orientation and posture;
[0043] Step 2.3: Considering the symmetry problem, the human skeleton data is flipped to achieve data enhancement and improve the robustness of the model;
[0044] Step 2.4: Model and express the human skeleton model based on the form of a topological graph.
[0045] Preferably, the temporal feature extraction module uses one-dimensional dilated convolution to input sequences of arbitrary length, and obtains temporal receptive fields of different sizes by using different numbers of layers.
[0046] Preferably, the behavior time bounding box regression module is composed of two graph convolution layers and one fully connected layer, and the final network output data size is 1×1, representing the duration of the current node behavior.
[0047] Preferably, the frame-level behavior classification module consists of two graph convolutional layers, one fully connected layer and one Softmax layer. The final output of the network is N×1, where N is the number of risky behavior categories. The output represents the confidence of each behavior category, and the sum of the confidences is 1.
[0048] Preferably, step 4 also includes: importing the preprocessed skeleton joint coordinate information into the risky behavior online spatiotemporal feature acquisition and recognition model, performing specified rounds of training, lowering the learning rate during training to achieve excellent training results, and ending the model training and saving the model when the accuracy rate increases and the loss function decreases to a flat level.
[0049] Preferably, the real-time acquisition of skeleton data of the actor using the depth sensor in step 5 includes: designing a dynamic link library file through C++ to provide an interface for Python to obtain the real-time skeleton of the depth sensor. Specific embodiment:
[0051] Combination Figure 1 The method for acquiring and identifying spatiotemporal features of risky behaviors based on three-dimensional depth vision includes the following steps:
[0052] Step 1: Build a 3D visual sensing network consisting of Kinect v2 depth sensors at key workstations in the final assembly workshop, and design dangerous behavior categories according to workshop requirements.
[0053] Kinect v2 depth sensors are installed at key workstations in the final assembly workshop, with two depth sensors installed in each of the two docking areas, to achieve all-round monitoring of personnel at the workstations. Dangerous behavior categories are designed according to workshop requirements, including leaning on products, running, jumping, using mobile phones, smoking, and entering the equipment range.
[0054] Step 2: Collect data on personnel behavior in the form of video streams, and obtain the skeleton joint coordinate information of each behavior of workshop production personnel through the human posture estimation algorithm based on depth sensors; use depth sensors to obtain multi-stream data of assembly process stations, among which the skeleton joint data of production personnel is obtained based on the personnel posture estimation algorithm to establish a data set. The specific steps are as follows:
[0055] 2.1. Design and identify behavioral categories based on workshop requirements, and continuously collect multi-stream data of personnel, including color image data and depth data.
[0056] 2.2. Obtain the coordinate information of the skeleton joint points of the workshop production personnel through the method of human posture estimation, and store it in the form of a data set for subsequent use: Detect the position, scale, direction and other information of the key parts of the human body from the image data, and use the key points instead of the image to express the human posture. Since the key points are mainly located at the positions of the main joints of the human body, they are similar to the human skeleton after connection. The collected data is streamData, which consists of frameData data of indefinite length, where frameData={x1,y1,z1,...,x i ,y i ,z i}; frameData represents the skeleton joint position data of a single frame image. m represents the number of frames of this set of data. Since the continuous acquisition time is long enough, m is large enough. The acquisition rate of the data set is 30fps. The occurrence time and duration of the actions in the data set are randomly generated. The background action in the middle of each action instance is also random and the duration is uncertain. We have collected 7 long sequences in total, totaling 463,750 frames.
[0057] Step 3: Preprocess the skeleton joint point coordinate information in the behavior data set, and model the preprocessed skeleton joint point data in the form of a topological graph. The specific steps are as follows:
[0058] 3.1. Coordinate translation: The coordinate origin of the skeleton data collected by the depth sensor is translated to the skeleton hip coordinate point to eliminate the influence of the human body position. For example:
[0059]
[0060] 3.2 Coordinate rotation: Rotate the skeleton data collected by the depth sensor to eliminate the influence of the orientation and posture of the workshop personnel. For example, the skeleton between joints 0 and 1 is parallel to the z-axis, and the skeleton between joints 8 and 4 is parallel to the x-axis. Taking the skeleton between joints 0 and 1 as an example to be parallel to the z-axis, the rotation axis is the common perpendicular line between the skeleton and the z-axis, and the rotation angle is the angle between them:
[0061]
[0062] The rotation axis is obtained by the cross product of the two vectors. The rotation angle of the two vectors can be obtained by the arc cosine, as shown in the following formula:
[0063]
[0064] 3.3. Dataset enhancement: Considering the symmetry problem, the skeleton data of the person is flipped to achieve data enhancement and improve the robustness of the model.
[0065] Data enhancement is achieved based on the symmetry of the human body structure, where the left and right upper limb joint data are swapped, and the left and right lower limb data are swapped. Preferably, when 25 joint nodes are selected, 20 joint points are calculated in turn for pairwise replacement ( Figure 2 In the figure, joint 5: left shoulder; joint 6: left elbow; joint 7: left wrist; joint 8: left hand; joint 9: right shoulder; joint 10: right elbow; joint 11: right wrist; joint 12: right hand; joint 13: left hip; joint 14: left knee; joint 15: left ankle; joint 16: left foot; joint 17: right hip; joint 18: right knee; joint 19: right ankle; joint 20: right foot; joint 22: left fingertips; joint 23: left thumb; joint 24: right fingertips; joint 25: right thumb), the replacement principle is left shoulder-right shoulder; left elbow-right elbow; left wrist-right wrist; left hand-right hand; left hip-right hip; left knee-right knee; left ankle-right ankle; left foot-right foot; left fingertips-right fingertips; left thumb-right thumb.
[0066] 3.4. Topological graph representation: The human skeleton model is modeled and expressed in the form of a topological graph.
[0067] We use graphs to form a hierarchical representation of skeleton sequences. For each frame in an action instance, there are 25 joints. This skeleton sequence with body connections can be regarded as an undirected acyclic graph. We construct an undirected space-time graph on it. Similar to the representation method of graphs in graph theory, the human skeleton graph can be represented as G = (joints, bones). Where joints = {V1,…,V n}, n = 25, which is 25 skeleton joints, where V i ={x i ,y i ,z i}, represents the three-dimensional coordinate data of each joint point, so the edge bones are the skeletons of the natural connection of the joint points of the human body. Among them, the meaning of the nodes in the skeleton topology graph and the adjacent nodes are shown in Table 1, and the adjacency matrix is constructed according to this connection method.
[0068] Table 1 Topology map node meaning and adjacent nodes
[0069]
[0070]
[0071] Step 4: Design an online spatiotemporal feature acquisition and recognition model for risky behaviors based on deep learning. The model includes a temporal feature extraction module and a spatial feature extraction module. The temporal feature extraction module based on the temporal convolutional network is responsible for extracting temporal features, and the spatial feature extraction module based on the graph neural network is responsible for extracting spatial features. The spatial feature extraction module includes a behavior temporal bounding box regression module and a frame-level behavior classification module. The behavior temporal bounding box regression module is used to predict the duration of the current behavior, and the frame-level behavior classification module is used to determine the behavior category at the current time point.
[0072] Based on the method of time series convolutional network and graph neural network, a model for online spatiotemporal feature acquisition and recognition of workshop personnel is designed. The time information of skeleton data is extracted through the time series convolutional network. The time series convolutional network module is as follows: Figure 4 As shown in the figure, in this module, the residual is used in each network layer. In each layer of the network, the dilated convolution is first performed, and then the residual structure is used to solve the gradient vanishing and gradient exploding problems, ensuring good information transfer. This accelerates convergence, allowing deeper network models to be stacked for training. Figure 5 It is a node in the network. It is stacked multiple times in the network. The number of stacking layers depends on the receptive field requirements mentioned above. The gated activation function is used in the network, which can better alleviate the problem of gradient disappearance.
[0073] z=tanh(ω f,k *x)⊙σ(ωg,k *x)
[0074] Where * represents the convolution operation, ⊙ represents the matrix dot multiplication, σ represents the sigmoid activation function, and ω is the learned convolution kernel.
[0075] The spatial feature extraction module is integrated into Figure 6 The graph convolution layer implementation shown in the figure, where the behavior temporal bounding box regression module consists of two graph convolution layers and one fully connected layer, and the frame-level behavior classification module consists of two graph convolution layers, one fully connected layer and one Softmax layer. The graph convolution layer formula is as follows:
[0076]
[0077] Where A is the joint point adjacency matrix, W is the graph convolution operation, and f in is the input feature, f out To output features, Softmax extracts the importance of different joints, thereby realizing the attention mechanism.
[0078] Step 5: Build a deep learning model, import the preprocessed joint point data into the online spatiotemporal feature acquisition and recognition model, and train the model. The model is developed using the pytorch1.1 deep learning framework in Python language.
[0079] The objective function of this classification task is to minimize the cross entropy loss function:
[0080]
[0081] where z t,k Indicates whether the true label of the tth frame belongs to the kth category; z t,k =1 means the true label is the kth class, P(c t,k |v0,...v t ) represents the estimated probability of being action class k in the tth frame.
[0082] The regression task uses the mean square error loss function, which is the mean of the sum of squares of the errors between the predicted data and the original data:
[0083]
[0084] where s t,k is the distance from the actual behavior starting point, The distance from the starting point of the model prediction behavior. Reduce the learning rate during training to achieve good training results. When the accuracy rate increases and the loss function decreases to a flat level, end the model training and save the model.
[0085] During training, we selected sequence number 1 for testing and sequences 2, 3, 4, 5, 6, and 7 for training. We performed 100 rounds of model training, using a learning rate of 0.01 for the first 50 rounds, 0.005 for the 50th to 80th rounds, and 0.002 for the 80th to 100th rounds to achieve good training results. When the accuracy rate increased and the loss function decreased to a flat level after 100 rounds, we ended the model training and saved the model.
[0086] Step 6: Use a depth sensor to collect skeleton data of people in real time, build a behavior recognition system, output the data stream to the deep learning model for recognition, and store the recognition results in a database. The specific steps are as follows:
[0087] 6.1 Design a dynamic link library file through C++ to provide an interface for Python to obtain depth sensor data. Based on the visual studio compiler, use the extern "C" __declspec (dllexport) function to export the dynamic link library file and configure the project type as a dynamic link library.
[0088] 6.2 Based on the preprocessing schemes of step 3.1, step 3.2 and step 3.4, a preprocessing module is designed to preprocess the skeleton data input by the sensor in real time.
[0089] 6.3 Based on the deep learning model, build an online behavior detection module, close the model for gradient calculation and back propagation functions, input the preprocessed data into the network in real time, and obtain the behavior recognition result output. The output result is:
[0090] output={duration,class,confidence}.
[0091] Among them, duration is the duration of the behavior, class is the behavior category, and confidence is the confidence of the behavior category.
[0092] The above different modules use different threads, and the running speeds are shown in Table 2.
[0093] Table 2 Running speed
[0094] Model TGCN-5 TGCN-6 TGCN-7 TGCN-8 TGCN-9 Speed (FPS) 29.4 27.0 25.6 24.4 20.8
[0095] 6.4 Use MySQL database to save the recognition results.
[0096] The database tables related to personnel behavior recognition are shown in Table 3 and Table 4.
[0097] Table 3. Example 1 of the data table for identifying the behavior of personnel at key workstations
[0098]
[0099] Table 4. Example 2 of the data table for identifying the behavior of personnel at key workstations
[0100]
[0101] The above shows and describes the basic principles, main features and advantages of the present invention. Technical actors in this industry should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A method for acquiring and identifying spatiotemporal features of risky behaviors based on three-dimensional depth vision, characterized in that: The method comprises the following steps: Step 1: For the area where risky behaviors need to be monitored, a visual sensing network composed of depth sensors is constructed to continuously collect information about the people in the area in the form of video streams. The coordinate information of the skeleton joints of each behavior of the person is obtained through a human posture estimation algorithm based on depth sensors, so as to construct a risky behavior dataset; Step 2: Preprocess the coordinate information of the skeleton joint points in the risky behavior dataset; Step 3: Design an online spatiotemporal feature acquisition and recognition model for risky behaviors based on deep learning, wherein the model includes a temporal feature extraction module and a spatial feature extraction module. The temporal feature extraction module based on a temporal convolutional network is responsible for extracting temporal features, and the spatial feature extraction module based on a graph neural network is responsible for extracting spatial features. The spatial feature extraction module includes a behavior temporal bounding box regression module and a frame-level behavior classification module. The output of the temporal feature extraction module based on a temporal convolutional network is used as the input of the behavior temporal bounding box regression module and the frame-level behavior classification module. The behavior temporal bounding box regression module is used to predict the duration of the current behavior, and the frame-level behavior classification module is used to determine the behavior category at the current time point. Step 4: Build a deep learning model, import the preprocessed skeleton joint point coordinate information into the risk-causing behavior online spatiotemporal feature acquisition and recognition model, and train the model with the classification regression module loss function; Step 5: Use a depth sensor to collect the skeleton data of the actor in real time, build a behavior recognition system, output the data stream to the deep learning model for recognition, and store the recognition results in the database.
2. The method according to claim 1, characterized in that: The continuous collection of information on the actors in the area in the form of video streams described in step 1 includes: designing and identifying the behavior categories according to requirements, and continuously collecting multi-stream data on the actors.
3. The method according to claim 1, characterized in that: The preprocessing of the skeleton joint coordinate information in the risky behavior data set in step 2 includes: Step 2.1: Translate the origin of the skeleton data coordinates collected by the depth sensor to the skeleton hip coordinate point to eliminate the influence of the human body position; Step 2.2: Rotate the skeleton data collected by the depth sensor to eliminate the influence of the human body's orientation and posture; Step 2.3: Considering the symmetry problem, the human skeleton data is flipped to achieve data enhancement and improve the robustness of the model; Step 2.4: Model and express the human skeleton model based on the form of a topological graph.
4. The method according to claim 1, characterized in that: The temporal feature extraction module uses one-dimensional dilated convolution to input sequences of arbitrary length and obtains temporal receptive fields of different sizes by using different numbers of layers.
5. The method according to claim 1, characterized in that: The behavior time bounding box regression module consists of two graph convolutional layers and one fully connected layer. The final network output data size is 1×1, representing the duration of the current node behavior.
6. The method according to claim 1, characterized in that: The frame-level behavior classification module consists of two graph convolutional layers, one fully connected layer and one Softmax layer. The final output of the network is N×1, where N is the number of risky behavior categories. The output represents the confidence of each behavior category, and the sum of the confidences is 1.
7. The method according to claim 1, characterized in that: Step 4 also includes: importing the preprocessed skeleton joint coordinate information into the risky behavior online spatiotemporal feature acquisition and recognition model, performing specified rounds of training, reducing the learning rate during training to achieve excellent training results, and ending the model training and saving the model when the accuracy rate increases and the loss function decreases to a flat level.
8. The method according to claim 1, characterized in that: The use of the depth sensor in real time to collect the skeleton data of the actor in step 5 includes: designing a dynamic link library file through C++ to provide an interface for Python to obtain the real-time skeleton of the depth sensor.