Method and system for identifying dangerous behaviors of passengers in elevator car
By repairing key points of the human skeleton in elevator surveillance videos and utilizing a front-to-back dual-fusion graph convolutional network to integrate multi-stream features and spatiotemporal channel attention mechanisms, the accuracy of identifying dangerous behaviors of passengers in elevator cars was solved, and the recognition effect was improved.
Patent Information
- Application Number
- CN202511473658.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-01-09
AI Technical Summary
Existing technologies for identifying dangerous behaviors of passengers inside elevator cars have poor motion recognition accuracy, especially when objects obstruct or people overlap, resulting in the loss of skeletal information and affecting the recognition effect.
This method extracts the human skeletal key point sequence of passengers from elevator surveillance videos, repairs missing or erroneous skeletal key points through nearest neighbor interpolation and linear interpolation algorithms, combines a front-to-back dual-fusion graph convolutional network to fuse multi-flow features of joint flow, skeletal flow and motion flow, and uses a spatiotemporal channel parallel attention mechanism to extract discriminative features to identify violent behavior in elevator cars.
It improves the accuracy of identifying dangerous behaviors of passengers in elevator cars, compensates for information loss caused by occlusion, and enhances the performance of the model.
Smart Images

Figure CN121305673A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for identifying dangerous behaviors of passengers in an elevator car, belonging to the field of behavior recognition technology. Background Technology
[0002] As an indispensable vertical transportation tool in modern buildings, the safety of elevator operation is directly related to the personal safety of passengers. In recent years, safety accidents caused by dangerous behavior of passengers in elevator cars have occurred frequently. Therefore, developing an intelligent monitoring system that can automatically and in real time identify dangerous behavior of passengers in elevator cars is of great significance for improving the level of elevator safety operation and preventing safety accidents.
[0003] In existing technologies, passenger dangerous behavior recognition mainly uses skeletal data extracted from video frames as input data through pose estimation algorithms to analyze the appearance and motion information in the video frames, including key point coordinates and confidence information. However, existing technologies only use key point coordinates to extract spatiotemporal feature maps, which results in insufficient representation information. Due to object occlusion or personnel overlap, the recognized skeleton information is also lost, and the skeletal sequence input is incomplete, affecting the accuracy of action recognition. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for identifying dangerous behaviors of passengers in an elevator car, so as to solve the problem of poor accuracy in action recognition in the prior art.
[0005] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution.
[0006] A method for identifying dangerous behaviors of passengers in an elevator car, comprising:
[0007] Specifically, the sequence of key points of the human skeleton of passengers was extracted from the elevator surveillance video;
[0008] Based on the passenger head detection and tracking results, a continuous skeleton sequence with identification is constructed, and missing or incorrect skeletal key points in the sequence are repaired by nearest neighbor frame interpolation and linear interpolation algorithms.
[0009] Based on the repaired skeletal sequence, the elbow flexion angle and the angle between the arm and shoulder are calculated, and passenger door-opening behavior is identified according to a preset angle threshold.
[0010] The skeleton sequence is input into a front-to-back dual-fusion graph convolutional network. By fusing the multi-flow features of joint flow, skeleton flow and motion flow, and using a spatiotemporal channel parallel attention mechanism to extract discriminative features, the identification result of violent behavior in the elevator car is obtained.
[0011] Specifically, extract the key point sequence of the human skeleton of passengers from the elevator surveillance video, including:
[0012] The OpenPose network with a single-dual structure is adjusted to the input branch structure of the surveillance video frame, and the image features are extracted by the MobileNetV3 network with the convolutional block attention mechanism.
[0013] The image features are input into a parameter-shared single-branch dual-output prediction network to generate a human keypoint confidence map and a partial affinity field, respectively.
[0014] Based on the key point confidence map and partial affinity field, the detected key points are assembled into a complete skeleton map of multiple people using a bipartite graph matching algorithm.
[0015] Based on the spatial relationship between the passenger head detection box and the key points on the head, the skeleton sequence is matched with the passenger identity identifier by calculating the Euclidean distance to construct a skeleton sequence with identity continuity.
[0016] Specifically, based on the passenger head detection and tracking results, a continuous skeleton sequence with identification is constructed, and missing or erroneous skeletal key points in the sequence are repaired using nearest neighbor frame interpolation and linear interpolation algorithms, including:
[0017] Based on the visibility of five key points on the passenger's head, the arithmetic mean is used when all key points are visible, the dynamic weight allocation is used when some key points are occluded, and the key point coordinates are used directly when only one key point is visible.
[0018] Calculate the Euclidean distance between the head center point and the passenger head detection frame center point, and assign the identity of the tracking frame to the corresponding skeleton sequence by minimum distance matching;
[0019] The integrity of key points in the skeleton sequence is checked. Key points with a coordinate value of 0 are identified as missing key points, and key points whose position changes between adjacent frames exceed a preset threshold are identified as erroneous key points.
[0020] Missing key points are repaired using the nearest neighbor frame interpolation algorithm, which is based on the weighted calculation of the coordinate values of the nearest valid frames before and after. Erroneous key points are corrected using a linear interpolation algorithm, which replaces them with the average value of the corresponding key point coordinates of the preceding and following frames.
[0021] Specifically, based on the repaired skeletal sequence, the elbow flexion angle and the angle between the arm and shoulder are calculated. Passenger door-opening behavior is identified according to preset angle thresholds, including:
[0022] Extract the coordinates of key points related to the door-opening behavior from the skeleton sequence, including the two-dimensional coordinates of key points of the left hand, right hand, left elbow, right elbow, left shoulder, right shoulder and neck;
[0023] Based on the coordinates of the key points, a limb triangle is constructed, and the bending angles of the left and right elbows are calculated using the inverse cosine function according to the relationship between the lengths of the three sides of the triangle.
[0024] Based on the coordinates of the key points, a shoulder-neck-elbow triangle is constructed, and the angle between the arm and shoulder is calculated using the inverse cosine function according to the relationship between the lengths of the three sides of the triangle.
[0025] When the four angles of the left and right arms, namely the elbow angle and the angle between the arm and shoulder, are detected to be within the corresponding range, it is judged as a door-opening behavior.
[0026] Specifically, the skeleton sequence is input into a front-to-back dual-fusion graph convolutional network. By fusing multi-flow features of joint flow, skeletal flow, and motion flow, and utilizing a spatiotemporal parallel attention mechanism to extract discriminative features, the identification results of violent behavior inside the elevator car are obtained, including:
[0027] The joint flow, skeletal flow, joint motion flow, and skeletal motion flow in the skeleton sequence are extracted as input. The joint flow and skeletal flow are fused according to the channel dimension to form a spatial flow, and the joint motion flow and skeletal motion flow are fused to form a motion flow.
[0028] Based on the spatiotemporal channel parallel attention mechanism, spatiotemporal graph convolution is performed on spatial flow and motion flow, and attention weights are calculated from the spatial dimension, temporal dimension and channel dimension respectively. The spatiotemporal attention features and channel attention features are obtained by performing outer product operation based on the attention weights.
[0029] Adaptive optimization processing of spatial flow and motion flow is performed using spatiotemporal attention features and channel attention features. Spatial flow uses a physical topological adjacency matrix to focus on the connection relationship between joints, while motion flow uses a normalized adjacency matrix containing self-connectivity to focus on the relative motion relationship between joints.
[0030] The processed spatial flow and motion flow features are used to calculate behavior discrimination scores through global average pooling layers and fully connected layers, respectively. The violent behavior recognition result is determined based on the category probability distribution of the behavior discrimination scores.
[0031] Specifically, a spatiotemporal channel parallel attention mechanism is introduced into the spatiotemporal graph convolution, calculating attention weights from the spatial, temporal, and channel dimensions respectively. Spatiotemporal attention features are fused through outer product operations and combined with channel attention features to enhance attention to key features, including:
[0032] Global average pooling is performed on the input feature map along the joint dimension to generate spatial feature vectors, and spatial attention weights for each joint are calculated through a fully connected layer and a non-linear activation function.
[0033] Global average pooling is performed on the input feature map along the time frame dimension to generate a time feature vector, and the time attention weights of each time frame are calculated through a fully connected layer and a non-linear activation function.
[0034] The spatial attention weights and temporal attention weights are subjected to a channel-level outer product operation to generate a joint spatiotemporal attention feature map.
[0035] The input feature map is compressed by global average pooling. The non-linear relationship between channels is learned through a bottleneck structure containing two fully connected layers, and channel attention weight coefficients are generated.
[0036] Specifically, global average pooling is used to compress the feature dimensions of spatial flow and motion flow. The nonlinear relationships between channels are learned through the bottleneck structure of the fully connected layer to generate channel attention features, including:
[0037] Global average pooling is performed on spatial and motion flows in both spatial and temporal dimensions to compress the feature map of each channel into a single feature value and generate a channel description vector.
[0038] The channel description vector is input into the first fully connected layer and reduced in dimensionality proportionally to the original number of channels. Then, it is transformed nonlinearly by the Hardswish activation function to obtain the activated channel description vector.
[0039] The activated channel description vector is input into the second fully connected layer to restore the original channel dimension, and the initial channel attention weight coefficients between 0 and 1 are generated by the Sigmoid activation function.
[0040] By multiplying the initial channel attention weight coefficients with the original spatial flow and motion flow, the important feature channels and unimportant channels are adjusted to obtain the channel attention weight coefficients.
[0041] A system for identifying dangerous behaviors of passengers inside an elevator car, comprising a key point extraction module, a skeleton sequence construction module, a door-pushing behavior recognition module, and a violent behavior recognition module:
[0042] The key point extraction module is used to extract the sequence of key points of the human skeleton of passengers in the elevator monitoring video.
[0043] The skeleton sequence construction module is used to construct a continuous skeleton sequence with identity identification based on the passenger head detection and tracking results, and to repair missing or incorrect skeletal key points in the sequence through nearest neighbor frame interpolation and linear interpolation algorithms.
[0044] The door-opening behavior recognition module is used to calculate the elbow bending angle and the angle between the arm and shoulder based on the repaired skeleton sequence, and to identify the passenger's door-opening behavior according to a preset angle threshold.
[0045] The violent behavior recognition module is used to input the skeleton sequence into a front-to-back dual-fusion graph convolutional network, and obtain the recognition result of violent behavior in the elevator car by fusing multi-flow features of joint flow, skeleton flow and motion flow, and extracting discriminative features using a spatiotemporal channel parallel attention mechanism.
[0046] The door-opening behavior recognition module includes a joint angle calculation unit and an angle threshold matching unit:
[0047] The joint angle calculation unit is used for key joint screening and angle calculation;
[0048] The angle threshold matching unit is used for threshold calibration and behavior determination.
[0049] The violent behavior recognition module includes a multi-stream feature extraction unit and an attention unit:
[0050] The multi-stream feature extraction unit is used to extract complementary feature streams;
[0051] The attention unit is used to optimize features through spatial attention, temporal attention, and channel attention.
[0052] Compared with existing technologies, the beneficial effects achieved by this invention are as follows: human skeleton data is obtained based on the improved OpenPose algorithm, and skeleton sequence information is obtained by combining the head target tracking algorithm. Violent behavior recognition is performed by using a front-to-back dual-fusion graph convolutional network, which fuses multiple information streams of the skeleton and captures richer and more diverse feature information from different information streams, thereby compensating for information loss caused by occlusion and improving recognition accuracy. Attention mechanisms are constructed from three dimensions: space, time, and channel. Attention features in each dimension are extracted and fused, enabling the model to deeply learn the action features in the skeleton sequence, thereby improving the model's performance.
[0053] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description
[0054] Figure 1 A flowchart of a method for identifying dangerous behaviors of passengers in an elevator car provided by the present invention;
[0055] Figure 2 This is a schematic diagram of a passenger prying open the door provided by the present invention;
[0056] Figure 3 This is a schematic diagram of the elbow angle when opening a door, provided by the present invention.
[0057] Figure 4 This invention provides a schematic diagram illustrating the angle between the shoulder and arm when opening a door.
[0058] Figure 5 This is a schematic diagram of multiple input feature flows provided by the present invention;
[0059] Figure 6 A schematic diagram of the network loss function curve provided by this invention;
[0060] Figure 7 This invention provides a structural diagram of a passenger dangerous behavior recognition system in an elevator car. Detailed Implementation
[0061] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0062] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0063] Example 1
[0064] Please see Figures 1-6 This invention provides an embodiment of a method for identifying dangerous behaviors of passengers in an elevator car, comprising the following specific steps:
[0065] Step S1: Extract the sequence of key points of the human skeleton of passengers from the elevator monitoring video.
[0066] The specific steps of step S1 are as follows:
[0067] Step S101: Adjust the input branch structure of the monitoring video frame to a single-dual structure OpenPose network, and extract image features through a MobileNetV3 network that integrates convolutional block attention mechanism.
[0068] In this embodiment, the first half of the parallel branch structure of each stage in the OpenPose network is changed to a single-branch structure. At the end of the stage, it is divided into two branches to predict the key point confidence map and part of the affinity field, respectively, to achieve parameter sharing of the convolution operation. The input video frame is preprocessed, and the image size is uniformly adjusted to the input size specified by the network and standardized. The preprocessed image data is sent to the feature extraction network. During the feature extraction process, a convolutional block attention mechanism is introduced after each bottleneck residual block. Global max pooling and global average pooling operations are performed on the input feature map at the same time. The two spatial feature descriptors obtained are input into a shared multilayer perceptron. Channel attention weights are generated by summation and sigmoid activation function. The weighted feature map is max pooled and average pooled along the channel dimension respectively. The concatenated feature map is compressed by the convolutional layer and spatial attention weights are generated by sigmoid function. The channel attention weights and spatial attention weights are applied to the feature map to enhance key features and suppress irrelevant features, and output the optimized image feature map.
[0069] Step S102: Input the image features into a parameter-shared single-branch-dual-output prediction network to generate a human keypoint confidence map and a partial affinity field, respectively.
[0070] In this embodiment, the image features are input into a parameter-shared single-branch dual-output prediction network to generate a human keypoint confidence map and a partial affinity field. The input feature map is processed through multi-level convolution through a shared single-branch structure. At each processing stage, the original feature map is concatenated with the output features of the previous stage to combine contextual information and progressively optimize the feature representation. At the end of each stage, the processed feature map is input into two independent prediction branches. The first branch generates the human keypoint confidence map through convolution, and the second branch generates the partial affinity field through convolution. The outputs of the two branches are then jointly passed to the next processing stage as additional input for that stage. This iterative optimization mechanism continuously improves the accuracy of keypoint localization and association.
[0071] Step S103: Based on the key point confidence map and partial affinity field, the detected key points are assembled into a complete skeleton map of multiple people using a bipartite graph matching algorithm.
[0072] In this embodiment, based on the keypoint confidence map and partial affinity field, the detected keypoints are assembled into a complete skeleton map of multiple people using a bipartite graph matching algorithm. Specifically, all candidate keypoints with a confidence level higher than a preset threshold are extracted from the keypoint confidence map and categorized according to body parts. Simultaneously, vector field information connecting keypoint pairs of different body parts is obtained from the partial affinity field to construct a bipartite graph model. Candidate keypoints of adjacent body parts are used as two vertex sets of the bipartite graph, and the association score between the keypoint pairs connecting these two parts is calculated as the edge weight based on the partial affinity field. Based on the edge weights, the maximum weight matching of the bipartite graph is solved using optimal allocation logic. The connection relationship between adjacent key points belonging to the same person is determined. All adjacent human body part pairs are traversed, and the bipartite graph construction and matching process is repeated to form multiple initial human skeletons composed of key point connections. The initial human skeletons are integrated and conflict-handled. Skeletons sharing the same key points are merged to eliminate duplicate connections. Unassigned key points are supplemented and associated according to the principle of spatial proximity. Finally, the complete human skeleton diagram of each independent individual is output. It should be noted that the confidence threshold is set by those skilled in the art according to the actual situation.
[0073] Step S104: Based on the spatial relationship between the passenger head detection box and the head key points, the skeleton sequence is matched with the passenger identity identifier by calculating the Euclidean distance to construct a skeleton sequence with identity continuity.
[0074] In this embodiment, based on the visibility status of the five key points on the passenger's head detected by OpenPose, a case-by-case calculation strategy is adopted to determine the coordinates of the head center point. When all key points are visible, the arithmetic mean of the horizontal and vertical coordinates of the five key points is calculated as the coordinates of the head center point. When some key points are not visible due to occlusion, the weight ratio of each visible key point is reassigned, and the head center point is calculated by weighted averaging of the visible key point coordinates according to the adjusted weights. When only one key point is visible, the coordinates of the key point are used as the coordinates of the head center point. The calculated coordinates of the head center point are spatially matched with the coordinates of the center point of the passenger head detection box. The Euclidean distance between each head center point and the center point of each detection box is calculated, and a distance matrix is established. Through the nearest neighbor matching principle, the detection box identity identifier with the smallest Euclidean distance is assigned to each head center point, thus establishing the correspondence between the skeleton sequence and the passenger identity. Based on the established identity correspondence, the passenger identity identifier obtained by the head target tracking algorithm is assigned to the corresponding skeleton sequence. By implementing the above matching process frame by frame, a passenger skeleton sequence with temporal continuity and identity consistency is constructed, providing a complete data foundation for subsequent behavior analysis.
[0075] Step S2: Based on the passenger head detection and tracking results, construct a continuous skeleton sequence with identification, and repair missing or incorrect skeletal key points in the sequence using nearest neighbor frame interpolation and linear interpolation algorithms.
[0076] The specific steps of step S2 are as follows:
[0077] Step S201: Based on the visibility of the five key points of the passenger's head, the arithmetic mean is used when all key points are visible, the dynamic weight allocation is used when some key points are occluded, and the key point coordinates are used directly when only one key point is visible.
[0078] In this embodiment, it is determined whether all head key points have been detected and the confidence level is higher than a set threshold. If this condition is met, the first processing flow is entered, where the arithmetic mean of the sum of the horizontal and vertical coordinates of the five key points is taken and directly output as the coordinates of the head center point. If some key points are not detected due to occlusion or insufficient confidence, the second processing flow is entered. Through a dynamic weight allocation mechanism and preset initial weights for each key point, when some key points are detected as invisible, the effective weights of each point are recalculated based on the set of currently visible key points to ensure that the sum of all weights remains constant. The weighted average is calculated based on the coordinates of the visible key points and their effective weights to obtain the coordinates of the head center point. If only a single key point is successfully detected and the confidence level meets the standard, the third processing flow is entered, where the coordinates of the key point are output as the coordinates of the head center point. It should be noted that the initial weights are set by those skilled in the art according to the actual situation.
[0079] Step S202: Calculate the Euclidean distance between the head center point and the passenger head detection frame center point, and assign the identity identifier of the tracking frame to the corresponding skeleton sequence by minimum distance matching.
[0080] In this embodiment, the spatial relationship between the center point of the passenger's head and the center point of the passenger's head detection box is calculated. Using the Euclidean distance formula between two points in a two-dimensional coordinate system, the straight-line distance between each passenger's head center point and the center points of all head detection boxes in the current frame is calculated, and a distance matrix is established. For each head center point, the detection box center point with the smallest distance is searched. When the minimum distance is lower than a preset matching threshold, it is determined that the head center point and the corresponding detection box belong to the same passenger target. The identity of the detection box is assigned to the passenger's corresponding complete skeleton sequence. For skeleton sequences that fail to match, their historical identity is retained and matching is attempted again in subsequent frames. It should be noted that the matching threshold is set by those skilled in the art according to the actual situation.
[0081] Step S203: Perform integrity detection on the key points in the skeleton sequence, determine the key points with a coordinate value of 0 as missing key points, and determine the key points whose position changes more than a preset threshold between adjacent frames as erroneous key points.
[0082] In this embodiment, integrity detection and error determination are performed on key points in the skeleton sequence. The continuous frame sequence under each passenger identity identifier is traversed, and the coordinate data of each skeleton key point is checked frame by frame. For each key point, when its horizontal and vertical coordinate values are both zero, the key point is determined to be in a missing state. For key points with non-zero coordinates, the coordinate displacement between the key point and the same key point in the adjacent previous frame is calculated. When the coordinate displacement exceeds a preset threshold, the coordinate displacement between the key point and the adjacent subsequent frame is further calculated. If the displacement before and after both exceed the displacement tolerance, the key point is determined to be in an error detection state in the current frame. A key point state set containing missing state markers and error detection state markers is output. It should be noted that the threshold is set by those skilled in the art according to the actual situation.
[0083] Step S204: Repair missing key points using the nearest neighbor frame interpolation algorithm. The algorithm calculates the coordinates of the nearest valid frames before and after the missing key points using a weighted average. Correct erroneous key points using a linear interpolation algorithm and replace them with the average coordinates of the corresponding key points in the preceding and following frames.
[0084] In this embodiment, the skeleton sequence of each passenger is traversed, and the coordinate values of each key point are checked frame by frame. Key points with coordinate values of zero are identified as missing key points. The change in coordinate position of each key point between adjacent frames is calculated. When the change exceeds a preset motion continuity threshold, the key point is identified as an erroneous key point. For the identified missing key points, nearest neighbor frame interpolation repair is performed. Using the current missing frame as a reference, the frame with the closest distance and valid coordinates of the key point in the passenger skeleton sequence is searched forward and backward, respectively, as the forward reference frame and the backward reference frame. The time interval between the current missing frame and these two reference frames is used as the basis for the repair. The distance is calculated separately, and the time weighting coefficient is calculated. Based on the time weighting coefficient, the coordinate values of the key point in the forward reference frame and the backward reference frame are weighted and fused to generate the repaired key point coordinates and update them to the current missing frame. For the identified erroneous key points, linear interpolation correction is performed to obtain the coordinate values of the same key point in the adjacent frames of the frame where the erroneous key point is located. The arithmetic mean of the coordinates of the corresponding key points in the preceding and following frames is calculated. This mean is used as the corrected coordinate value to replace the original coordinate value of the erroneous key point. It should be noted that the motion continuity threshold is set by those skilled in the art according to the actual situation.
[0085] Step S3: Based on the repaired skeletal sequence, calculate the elbow flexion angle and the angle between the arm and shoulder, and identify the passenger's door-opening behavior according to the preset angle threshold.
[0086] The specific steps of step S3 are as follows:
[0087] Step S301: Extract the coordinates of key points related to the door-opening behavior from the skeleton sequence, including the two-dimensional coordinates of key points of the left hand, right hand, left elbow, right elbow, left shoulder, right shoulder and neck.
[0088] In this embodiment, all human skeleton instances output by the improved OpenPose algorithm in the current frame are traversed. Key point data parsing is performed on each independent skeleton instance. According to the predefined human key point index mapping relationship, the two-dimensional pixel coordinates of seven specific key points—left wrist, right wrist, left elbow, right elbow, left shoulder, right shoulder, and neck—are located and separated from the coordinate data of each skeleton instance. The validity of each extracted key point coordinate is verified, and coordinate points with a confidence level higher than a set threshold are selected as valid data. The valid key point coordinates that have passed the verification are grouped and integrated according to the skeleton instance to form a set of key point coordinates related to the door-opening behavior analysis of each passenger in the current frame.
[0089] exist Figure 2 In the image, lines are used to represent the passenger's arms and shoulders working together, which are positioned at a relatively high level, forming a certain posture of bending or extending the arms.
[0090] Step S302: Construct a limb triangle based on the coordinates of the key points, and calculate the bending angles of the left and right elbows respectively using the inverse cosine function according to the relationship between the lengths of the three sides of the triangle.
[0091] In this embodiment, based on the two-dimensional coordinates of the key points of the left elbow, left shoulder, and left hand, the Euclidean distances from the left elbow to the left shoulder, from the left elbow to the left hand, and from the left shoulder to the left hand are calculated respectively. Based on these three distance values, a triangle of the left elbow joint is constructed, and the bending angle of the left elbow is calculated using the inverse cosine function. Based on the two-dimensional coordinates of the key points of the right elbow, right shoulder, and right hand, the Euclidean distances from the right elbow to the right shoulder, from the right elbow to the right hand, and from the right shoulder to the right hand are calculated respectively. Then, based on these three distance values, a triangle of the right elbow joint is constructed, and the bending angle of the right elbow is calculated using the inverse cosine function. The calculated bending angles of the left and right elbows are output as bending feature quantities of the left and right arms, respectively.
[0092] Step S303: Construct a shoulder-neck-elbow triangle based on the coordinates of the key points, and calculate the angle between the arm and shoulder using the inverse cosine function according to the relationship between the lengths of the three sides of the triangle.
[0093] In this embodiment, two-dimensional coordinate data of the neck key points, the target side shoulder key points, and the ipsilateral elbow key points are extracted from the skeleton sequence. The lengths of the line segments from the neck to the shoulder, from the neck to the elbow, and from the shoulder to the elbow are calculated based on the coordinates of the three points, forming the three sides of a triangle. An angle calculation model is constructed based on the law of cosines, with the line segment from the shoulder to the elbow as the opposite side and the line segments from the neck to the shoulder and from the neck to the elbow as the adjacent sides. An angle calculation formula is established, and the angle calculation formula is solved by the inverse cosine function to obtain the angle between the arm and the shoulder, thus completing the quantitative extraction of upper limb posture features.
[0094] Step S304: When the four angles of the left and right arms, namely the elbow angle and the angle between the arm and the shoulder, are detected to be within the corresponding range, it is determined to be a door-opening behavior.
[0095] In this embodiment, door-opening behavior recognition is performed based on the joint angle features of the left and right arms. Key point coordinate data of the left and right limbs are extracted from the passenger skeleton sequence of the current frame. The left limb involves key points of the left hand, left elbow, and left shoulder; the right limb involves key points of the right hand, right elbow, and right shoulder. For the left arm, based on the coordinate positions of the three key points (left hand, left elbow, and left shoulder), the Euclidean distances from the left elbow to the left shoulder, from the left elbow to the left hand, and from the left shoulder to the left hand are calculated. Based on the corresponding distance values, the elbow flexion angle of the left arm is calculated using trigonometric functions. Combining the coordinate positions of the three key points (left shoulder, neck, and left elbow), the Euclidean distances from the left shoulder to the neck and from the left elbow to the neck are calculated. The Euclidean distance and the Euclidean distance from the left elbow to the left shoulder are calculated using trigonometric functions based on the corresponding distance values. The same calculation process is used to process the key points of the right hand, right elbow, and right shoulder of the right arm, and calculate the elbow flexion angle and the angle between the right arm and the shoulder. The four angle values are then combined and logically judged. When the elbow flexion angle, the angle between the left arm and the shoulder, the elbow flexion angle, and the angle between the right arm and the shoulder are all within the preset door-opening behavior angle threshold range, it is determined that the current passenger is performing door-opening behavior. It should be noted that the door-opening behavior angle threshold range is set by those skilled in the art according to the actual situation.
[0096] exist Figure 3 and Figure 4 In the diagram, the dots represent the angle between the elbow and the shoulder when exerting force during the process of prying open the door.
[0097] Step S4: Input the skeleton sequence into the front and rear dual-fusion graph convolutional network, fuse the multi-flow features of joint flow, skeleton flow and motion flow, and use the spatiotemporal channel parallel attention mechanism to extract discriminative features to obtain the recognition result of violent behavior in the elevator car.
[0098] The specific steps of step S4 are as follows:
[0099] Step S401: Extract the joint flow, bone flow, joint motion flow and bone motion flow from the skeleton sequence as input, fuse the joint flow and bone flow according to the channel dimension to form a spatial flow, and fuse the joint motion flow and bone motion flow to form a motion flow.
[0100] In this embodiment, the joint flow, bone flow, joint motion flow, and bone motion flow in the skeleton sequence are extracted as input. For the joint flow, the two-dimensional coordinate data of each joint point in each frame of the skeleton sequence are directly read as input features. For the bone flow, the source joint and target joint of each bone segment are determined according to the physical connection relationship of the human skeleton, and the vector representation from the source joint to the target joint is calculated to form bone vector features. For the joint motion flow, the skeleton sequence of two consecutive frames is selected, and the coordinate displacement of the same joint point between adjacent frames is calculated to obtain the joint motion vector. For the bone motion flow, based on the bone flow data of two consecutive frames, the vector change of the same bone segment in adjacent frames is calculated to obtain bone motion features. The joint flow and bone flow are spliced and fused in the channel dimension to form a spatial flow containing absolute position information. The joint motion flow and bone motion flow are spliced and fused in the channel dimension to form a motion flow containing relative motion information.
[0101] exist Figure 5 In the image, from left to right, multiple input feature flows of the skeleton data are displayed sequentially. The first graphic represents the joint flow, the second graphic represents the bone flow, the third graphic represents the joint motion flow, and the fourth graphic represents the bone motion flow.
[0102] Step S402: Process the spatial flow and motion flow through spatiotemporal graph convolution, where the spatial flow uses a physical topological adjacency matrix to focus on the connection relationship between joints, and the motion flow uses a normalized adjacency matrix containing self-connectivity to focus on the relative motion relationship between joints.
[0103] In this embodiment, the fused spatial flow and motion flow are respectively input into spatiotemporal graph convolution for feature processing. For spatial flow features, a physical topological adjacency matrix is constructed based on the physical connection structure of the human skeleton. The elements in the adjacency matrix represent whether there are natural physiological connections between human joints and the connection strength. This adjacency matrix guides the graph convolution operation, enabling the network to learn the inherent spatial structural relationships between human joints. For motion flow features, a normalized adjacency matrix containing self-connectivity is constructed, and graph convolution is performed through the normalized adjacency matrix.
[0104] Step S403: Introduce a spatiotemporal channel parallel attention mechanism into the spatiotemporal graph convolution, calculate attention weights from the spatial dimension, temporal dimension and channel dimension respectively, fuse spatiotemporal attention features through outer product operation, and combine them with channel attention features to enhance attention to key features.
[0105] The specific steps of step S403 are as follows:
[0106] Step S4031: Perform global average pooling on the input feature map along the joint dimension to generate spatial feature vectors, and calculate the spatial attention weights of each joint through a fully connected layer and a non-linear activation function.
[0107] In this embodiment, global average pooling is performed on the input feature map along the time and channel dimensions. The feature values of each joint across all frames and channels are aggregated into a single scalar to form an initial spatial feature vector. The initial spatial feature vector is input into a fully connected layer for feature dimension enhancement, and an intermediate spatial feature representation is obtained by mapping through a non-linear activation function. The intermediate spatial feature representation is input into a fully connected layer for dimension restoration to generate spatial attention logits corresponding to the number of joints. The spatial attention logits are normalized, and the weight coefficient of each joint is constrained to the range of 0 to 1 to generate the final spatial attention weight vector.
[0108] Step S4032: Perform global average pooling on the input feature map in the time frame dimension to generate a time feature vector, and calculate the time attention weights for each time frame through a fully connected layer and a non-linear activation function.
[0109] In this embodiment, global average pooling is performed on the input feature map along the time frame dimension, compressing the feature information of each time frame into a single statistic to form a time feature description vector. This time feature vector is then input into a fully connected layer for nonlinear transformation. The feature representation capability is enhanced by processing with an activation function. The activated features are then input into the fully connected layer, restoring their dimension to the original number of time frames. The output value is mapped to the attention weight of each time frame through a normalized exponential function, thus completing the importance assessment along the time dimension.
[0110] Step S4033: Perform channel-level outer product operation on the spatial attention weights and temporal attention weights to generate a joint spatiotemporal attention feature map.
[0111] In this embodiment, the spatial dimension attention weight represents the importance distribution of each joint node, and the temporal dimension attention weight represents the importance distribution of each time frame. The spatial dimension attention weight and the temporal dimension attention weight are expanded into a two-dimensional spatiotemporal joint weight matrix through outer product operation. The element values of the two-dimensional spatiotemporal joint weight matrix reflect the comprehensive importance of a specific joint at a specific time. The two-dimensional spatiotemporal joint weight matrix is broadcast and aligned to complete the synchronous assignment of the spatial and temporal dimensions of the spatiotemporal attention weight, thereby obtaining an attention feature map with spatiotemporal joint perception capability.
[0112] Step S4034: Perform global average pooling to compress the feature dimensions of the spatial flow and motion flow, learn the nonlinear relationship between channels through the bottleneck structure of the fully connected layer, and generate channel attention features.
[0113] The specific steps of step S4034 are as follows:
[0114] Step S40341: Perform global average pooling operations on the spatial flow and motion flow in the spatial and temporal dimensions to compress the feature map of each channel into a single feature value and generate a channel description vector.
[0115] In this embodiment, the spatial flow and motion flow are traversed according to the channel dimension. For each independent feature channel, the arithmetic mean of the feature values is calculated at all spatial locations and all time frames contained therein. The feature information originally distributed in the spatial and temporal dimensions of each channel is integrated into a scalar value with global representativeness. All channels are traversed and scalar values are collected in sequence. The channel description vector is assembled according to the original channel order. By constructing the channel description vector, the global information of the feature channel can be compressed.
[0116] Step S40342: Input the channel description vector into the first fully connected layer to perform feature dimensionality reduction proportional to the original number of channels, and perform nonlinear transformation through the Hardswish activation function to obtain the activated channel description vector.
[0117] In this embodiment, the channel description vector is input to the first fully connected layer for feature dimension compression and nonlinear feature transformation. The high-dimensional channel description vector is projected to the low-dimensional feature space through a learnable weight matrix. During feature compression, the input vector is multiplied by the weight matrix through matrix multiplication and a bias vector is superimposed to complete the linear transformation. The linearly transformed feature vector is then subjected to element-wise nonlinear mapping through the Hardswish activation function. This achieves a smooth nonlinear transformation of the feature values while maintaining gradient propagation efficiency. The nonlinearly activated feature vector is then output to the next processing layer to obtain a channel description vector with a compact representation and rich nonlinear characteristics. Feature dimensionality reduction can reduce model complexity, while nonlinear transformation can enhance feature representation capabilities.
[0118] Step S40343: Input the activated channel description vector into the second fully connected layer to restore the original channel dimension, and generate initial channel attention weight coefficients between 0 and 1 through the Sigmoid activation function.
[0119] In this embodiment, the activated channel description vector is input to the second fully connected layer for feature dimension recovery and weight coefficient generation. The compressed feature vector is reprojected onto the original channel dimension using a learnable weight matrix to achieve complete reconstruction of the feature space. During the dimension recovery process, matrix multiplication is performed to multiply the input feature vector with the weight matrix and superimpose the corresponding bias vector to complete the linear reconstruction of the features. The reconstructed feature vector is then subjected to element-wise nonlinear transformation using the Sigmoid activation function. Exponential and normalization calculations are performed using the activation function to map the feature value of each channel to a continuous numerical space between 0 and 1. The nonlinearly transformed feature vector is output as the initial channel attention weight coefficient, resulting in initial channel attention weight coefficients that are channel-specific and have a normalized numerical range.
[0120] Step S40344: Adjust the important feature channels and unimportant channels by multiplying the initial channel attention weight coefficients with the original spatial flow and motion flow to obtain the channel attention weight coefficients.
[0121] In this embodiment, the initial channel attention weight coefficients are multiplied channel by channel with the original spatial flow and motion flow. The initial channel attention weight coefficients are broadcast expanded in both spatial and temporal dimensions to align their dimensions perfectly with the original input feature map. Element-wise multiplication is performed on each channel, and the initial weight coefficient of each channel is multiplied by all feature values of the corresponding channel. Important feature channels with weight coefficients greater than a preset weight threshold will have their corresponding feature values enhanced, while non-important feature channels with weight coefficients less than the preset weight threshold will have their corresponding feature values suppressed. The channel-adjusted feature map is then residually connected to the original input feature map, preserving the original feature information while highlighting the feature responses of important channels, thus generating the final optimized channel attention weight coefficients. It should be noted that the weight thresholds are set by those skilled in the art based on actual conditions.
[0122] Step S404: Calculate the behavior discrimination score by passing the processed spatial flow and motion flow features through a global average pooling layer and a fully connected layer, respectively, and determine the violent behavior recognition result based on the category probability distribution of the behavior discrimination score.
[0123] In this embodiment, the processed spatial flow features are input to a global average pooling layer, where feature compression operations are performed in both spatial and temporal dimensions. An arithmetic mean of all feature values for each channel is calculated to generate a spatial dimension channel feature vector. The processed motion flow feature map is then input to the same global average pooling layer, where the same feature compression operation is performed to generate a motion dimension channel feature vector. The spatial dimension channel feature vector is input to a fully connected layer, where a linear transformation involving matrix multiplication and bias addition is performed to map the feature vector to the dimension space of the number of behavior categories, generating a spatial flow behavior discrimination score. The motion dimension channel feature vector is then input to the same fully connected layer, where the same linear transformation operation is performed to generate a motion flow behavior discrimination score. The spatial flow behavior discrimination score and the motion flow behavior discrimination score are linearly combined according to a preset weighting coefficient, and the probability value of each behavior category is calculated using exponential normalization to obtain a behavior category probability distribution. Based on the behavior category probability distribution, the category with the highest probability value is selected as the output of the violent behavior recognition result. It should be noted that the weighting coefficients are set by those skilled in the art according to the actual situation.
[0124] exist Figure 6 In the graph, the horizontal and vertical axes represent the number of training iterations and the loss value on the training or validation set, respectively. Each colored curve represents the training loss for different inputs. The convergence curve of the loss function drops rapidly at the beginning of training, indicating that the experiment has set an appropriate learning rate. As training progresses, the convergence curve of the loss function tends to stabilize, indicating that the model converges stably during training.
[0125] Example 2
[0126] Please see Figure 7 The present invention provides an embodiment of a passenger dangerous behavior recognition system in an elevator car, comprising a key point extraction module, a skeleton sequence construction module, a door-opening behavior recognition module, and a violent behavior recognition module.
[0127] The key point extraction module is used to extract the sequence of key points of the human skeleton of passengers in elevator monitoring videos.
[0128] The skeleton sequence construction module is used to construct a continuous skeleton sequence with identification based on the passenger head detection and tracking results, and to repair missing or incorrect skeletal key points in the sequence through nearest neighbor frame interpolation and linear interpolation algorithms.
[0129] The door-opening behavior recognition module is used to calculate the elbow bending angle and the angle between the arm and shoulder based on the repaired skeleton sequence, and to identify the passenger's door-opening behavior according to a preset angle threshold.
[0130] The violent behavior recognition module is used to input the skeleton sequence into a front-to-back dual-fusion graph convolutional network, and obtain the recognition result of violent behavior in the elevator car by fusing multi-flow features of joint flow, skeleton flow and motion flow, and extracting discriminative features using a spatiotemporal channel parallel attention mechanism.
[0131] The door-opening behavior recognition module includes a joint angle calculation unit and an angle threshold matching unit.
[0132] The joint angle calculation unit is used for key joint screening and angle calculation.
[0133] The angle threshold matching unit is used for threshold calibration and behavior determination.
[0134] The violent behavior recognition module includes a multi-stream feature extraction unit and an attention unit.
[0135] The multi-stream feature extraction unit is used to extract complementary feature streams.
[0136] The attention unit is used to optimize features through spatial attention, temporal attention, and channel attention.
[0137] In addition, the parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of the corresponding technical solutions in the prior art have not been described in detail, so as to avoid excessive elaboration.
[0138] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.
Claims
1. A method for identifying dangerous behaviors of passengers in an elevator car, characterized in that, include: Extract the sequence of key points of the human skeleton of passengers from elevator surveillance video; Based on the passenger head detection and tracking results, a continuous skeleton sequence with identification is constructed, and missing or incorrect skeletal key points in the sequence are repaired by nearest neighbor frame interpolation and linear interpolation algorithms. Based on the repaired skeletal sequence, the elbow flexion angle and the angle between the arm and shoulder are calculated, and passenger door-opening behavior is identified according to a preset angle threshold. The skeleton sequence is input into a front-to-back dual-fusion graph convolutional network. By fusing the multi-flow features of joint flow, skeleton flow and motion flow, and using a spatiotemporal channel parallel attention mechanism to extract discriminative features, the identification result of violent behavior in the elevator car is obtained.
2. The method for identifying dangerous passenger behavior in an elevator car according to claim 1, characterized in that, The extraction of key skeletal point sequences from passengers in elevator surveillance video includes: The OpenPose network with a single-dual structure is adjusted to the input branch structure of the surveillance video frame, and the image features are extracted by the MobileNetV3 network with the convolutional block attention mechanism. The image features are input into a parameter-shared single-branch dual-output prediction network to generate a human keypoint confidence map and a partial affinity field, respectively. Based on the key point confidence map and partial affinity field, the detected key points are assembled into a complete skeleton map of multiple people using a bipartite graph matching algorithm. Based on the spatial relationship between the passenger head detection box and the key points on the head, the skeleton sequence is matched with the passenger identity identifier by calculating the Euclidean distance to construct a skeleton sequence with identity continuity.
3. The method for identifying dangerous passenger behavior in an elevator car according to claim 2, characterized in that, The process of constructing a continuous skeleton sequence with identification based on passenger head detection and tracking results, and repairing missing or erroneous skeletal key points in the sequence using nearest neighbor frame interpolation and linear interpolation algorithms, includes: Based on the visibility of five key points on the passenger's head, the arithmetic mean is used when all key points are visible, the dynamic weight allocation is used when some key points are occluded, and the key point coordinates are used directly when only one key point is visible. Calculate the Euclidean distance between the head center point and the passenger head detection frame center point, and assign the identity of the tracking frame to the corresponding skeleton sequence by minimum distance matching; The integrity of key points in the skeleton sequence is checked. Key points with a coordinate value of 0 are identified as missing key points, and key points whose position changes between adjacent frames exceed a preset threshold are identified as erroneous key points. Missing key points are repaired using the nearest neighbor frame interpolation algorithm, which is based on the weighted calculation of the coordinate values of the nearest valid frames before and after. Erroneous key points are corrected using a linear interpolation algorithm, which replaces them with the average value of the corresponding key point coordinates of the preceding and following frames.
4. The method for identifying dangerous passenger behavior in an elevator car according to claim 3, characterized in that, Based on the repaired skeletal sequence, the elbow flexion angle and the angle between the arm and shoulder are calculated, and passenger door-opening behavior is identified according to a preset angle threshold, including: Extract the coordinates of key points related to the door-opening behavior from the skeleton sequence, including the two-dimensional coordinates of key points of the left hand, right hand, left elbow, right elbow, left shoulder, right shoulder and neck; Based on the coordinates of the key points, a limb triangle is constructed, and the bending angles of the left and right elbows are calculated using the inverse cosine function according to the relationship between the lengths of the three sides of the triangle. Based on the coordinates of the key points, a shoulder-neck-elbow triangle is constructed, and the angle between the arm and shoulder is calculated using the inverse cosine function according to the relationship between the lengths of the three sides of the triangle. When the four angles of the left and right arms, namely the elbow angle and the angle between the arm and shoulder, are detected to be within the corresponding range, it is judged as a door-opening behavior.
5. The method for identifying dangerous passenger behavior in an elevator car according to claim 4, characterized in that, The process involves inputting the skeleton sequence into a front-to-back dual-fusion graph convolutional network. By fusing multi-flow features of joint flow, skeletal flow, and motion flow, and utilizing a spatiotemporal parallel attention mechanism to extract discriminative features, the identification result of violent behavior inside the elevator car is obtained, including: The joint flow, skeletal flow, joint motion flow, and skeletal motion flow in the skeleton sequence are extracted as input. The joint flow and skeletal flow are fused according to the channel dimension to form a spatial flow, and the joint motion flow and skeletal motion flow are fused to form a motion flow. Based on the spatiotemporal channel parallel attention mechanism, spatiotemporal graph convolution is performed on spatial flow and motion flow, and attention weights are calculated from the spatial dimension, temporal dimension and channel dimension respectively. The spatiotemporal attention features and channel attention features are obtained by performing outer product operation based on the attention weights. Adaptive optimization processing of spatial flow and motion flow is performed using spatiotemporal attention features and channel attention features. Spatial flow uses a physical topological adjacency matrix to focus on the connection relationship between joints, while motion flow uses a normalized adjacency matrix containing self-connectivity to focus on the relative motion relationship between joints. The processed spatial flow and motion flow features are used to calculate behavior discrimination scores through global average pooling layers and fully connected layers, respectively. The violent behavior recognition result is determined based on the category probability distribution of the behavior discrimination scores.
6. The method for identifying dangerous passenger behavior in an elevator car according to claim 5, characterized in that, The spatiotemporal channel-parallel attention mechanism performs spatiotemporal graph convolution on spatial flow and motion flow, calculates attention weights from the spatial, temporal, and channel dimensions respectively, and fuses them through outer product operations based on the attention weights to obtain spatiotemporal attention features and channel attention features, including: Global average pooling is performed on spatial flow and motion flow at the joint dimension to generate spatial feature vectors, and spatial attention weights of each joint are calculated through fully connected layers and nonlinear activation functions. Global average pooling is performed on spatial flow and motion flow in the temporal frame dimension to generate temporal feature vectors, and the temporal attention weights of each time frame are calculated through a fully connected layer and a non-linear activation function. The spatial attention weights and temporal attention weights are subjected to a channel-level outer product operation to generate joint spatiotemporal attention features; Global average pooling is used to compress the feature dimensions of spatial flow and motion flow. The nonlinear relationship between channels is learned through the bottleneck structure of the fully connected layer to generate channel attention features.
7. The method for identifying dangerous passenger behavior in an elevator car according to claim 6, characterized in that, The global average pooling of spatial flow and motion flow is used to compress feature dimensions. The nonlinear relationships between channels are learned through the bottleneck structure of the fully connected layer to generate channel attention features, including: Global average pooling is performed on spatial and motion flows in both spatial and temporal dimensions to compress the feature map of each channel into a single feature value and generate a channel description vector. The channel description vector is input into the first fully connected layer and reduced in dimensionality proportionally to the original number of channels. Then, it is transformed nonlinearly by the Hardswish activation function to obtain the activated channel description vector. The activated channel description vector is input into the second fully connected layer to restore the original channel dimension, and the initial channel attention weight coefficients between 0 and 1 are generated by the Sigmoid activation function. By multiplying the initial channel attention weight coefficients with the original spatial flow and motion flow, the important feature channels and unimportant channels are adjusted to obtain the channel attention weight coefficients.
8. A passenger dangerous behavior recognition system in an elevator car, used to implement the passenger dangerous behavior recognition method in an elevator car according to any one of claims 1-7, characterized in that, It includes a key point extraction module, a skeleton sequence construction module, a door-opening behavior recognition module, and a violent behavior recognition module: The key point extraction module is used to extract the sequence of key points of the human skeleton of passengers in the elevator monitoring video. The skeleton sequence construction module is used to construct a continuous skeleton sequence with identity identification based on the passenger head detection and tracking results, and to repair missing or incorrect skeletal key points in the sequence through nearest neighbor frame interpolation and linear interpolation algorithms. The door-opening behavior recognition module is used to calculate the elbow bending angle and the angle between the arm and shoulder based on the repaired skeleton sequence, and to identify the passenger's door-opening behavior according to a preset angle threshold. The violent behavior recognition module is used to input the skeleton sequence into a front-to-back dual-fusion graph convolutional network, and obtain the recognition result of violent behavior in the elevator car by fusing multi-flow features of joint flow, skeleton flow and motion flow, and extracting discriminative features using a spatiotemporal channel parallel attention mechanism.
9. A passenger dangerous behavior recognition system in an elevator car according to claim 8, characterized in that, The door-opening behavior recognition module includes a joint angle calculation unit and an angle threshold matching unit: The joint angle calculation unit is used for key joint screening and angle calculation; The angle threshold matching unit is used for threshold calibration and behavior determination.
10. A passenger dangerous behavior recognition system in an elevator car according to claim 9, characterized in that, The violent behavior recognition module includes a multi-stream feature extraction unit and an attention unit: The multi-stream feature extraction unit is used to extract complementary feature streams; The attention unit is used to optimize features through spatial attention, temporal attention, and channel attention.