Human Action Recognition Method Based on Multi-Stream Skeletal Features and Spatiotemporal Self-Attention
By using multi-flow bone features and space-time self-attention mechanism in human behavior recognition, the long-term motion characteristics of each node in the action and their mutual dependence characteristics are extracted, and sub-poses are aggregated in the time dimension, the problem of insufficient extraction of joint node information in the prior art is solved, and more accurate behavior classification is achieved.
Patent Information
- Application Number
- CN202310306295.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-03-27
AI Technical Summary
The existing human behavior recognition methods have shortcomings in extracting useful information of the joint nodes in space and time, and it is difficult to effectively identify complex actions.
The human behavior recognition method based on multi-flow bone characteristics and space-time self-attention is adopted. By obtaining the original joint node coordinate matrix and transforming it into multiple modes as different input streams, the long-term motion characteristics of each node in the action and its mutual dependence characteristics are extracted using the space-time self-attention mechanism, and the sub-poses in the complete action are aggregated between frames in the time dimension, and finally the behavior classification results are obtained through the full connection layer.
More semantic human behavior characteristics are obtained through multi-stream input information, the self-attention mechanism obtains the interdependence between nodes in the space, and the weighted fusion mechanism of the self-attention and inter-frame aggregation module in the time dimension effectively extracts long-term motion information and short-term information. Finally, the action characteristics under different semantics are aggregated through the multi-stream feature aggregation module to obtain the optimal behavior classification results.
Smart Images

Figure CN116311527B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human behavior recognition, and particularly to a human behavior recognition method based on multi-stream skeleton features and spatio-temporal self-attention. Background Art
[0002] Behavior recognition plays a crucial role in video understanding and is widely applied in fields such as video surveillance, human-computer interaction, virtual reality, etc. Based on various data representation methods, such as visual appearance, skeleton, depth, optical flow, etc., great progress has been made in this field. Among them, skeleton-based action recognition has received extensive attention from researchers due to its strong adaptability to dynamic environments and complex backgrounds.
[0003] The human body behavior recognition method analyzes and studies based on skeleton joint point data, aiming to learn the relationship between joint points and frames in an action to determine the action category. The general idea of existing methods is to judge the action category by learning the information of joint points in time and space. Therefore, how to reasonably extract the useful information of joint points in space and time is the key to solving the problem. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a human behavior recognition method based on multi-stream skeleton features and spatio-temporal self-attention, aiming to solve the above problems.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A human behavior recognition method based on multi-stream skeleton features and spatio-temporal self-attention includes the following steps:
[0007] Obtain the original joint point coordinate matrix and convert it into multiple modalities as different input streams;
[0008] For different input streams, use the spatio-temporal self-attention mechanism to extract the long-term motion features of each node in the action and the dependence features between them, and aggregate each sub-pose in the complete action frame by frame in the time dimension; aggregate the optimal proportion of long-term and short-term features, and obtain the behavior classification result through the fully connected layer;
[0009] Aggregate the motion feature classification results obtained from different input streams to generate the final behavior classification result.
[0010] Further, converting the original joint point coordinate matrix into multiple modalities as different input streams is specifically as follows:
[0011] Convert the joint point coordinate matrix into a human skeleton vector matrix, a joint point motion trend matrix, a human skeleton motion trend matrix, a human joint angle matrix, and the original joint point coordinate matrix;
[0012] Define the joint close to the center of gravity of the bone as the source joint and the joint far from the center of gravity as the target joint. Each bone is represented as a vector pointing from the source joint to the target joint, which contains not only length information but also direction information. Given a bone, its source joint v 1 =(x 1 , y 1 , z 1 ), and the target joint v 2 =(x 2 , y 2 , z 2 ). The vector calculation of the human skeleton is:
[0013] e v1,v2 =(x 2 - x 1 , y 2 - y 1 , z 2 - z 1 )
[0014] For the same node V in different frames t1 and t2, where V is V t1 =(x t1 , y t1 , z t1 ) at time t1 and V is V t2 =(x t2 , y t2 , z t2 ) at time t2, then its motion trend vector is expressed as:
[0015] m v1,v2 =(x t1 - x t2 , y t1 - y t2 , z t1 - z t2 )
[0016] The angle between two bone segments is defined by the included angle between the bone vectors. Let the two bone vectors e 1 , e 2 be e 1 =(x 3 , y 3 , z 3 ) and e 2 =(x 4 , y 4 , z 4 ) in the same frame respectively. Then the cosine of the included angle between the bone vectors is:
[0017]
[0018] If the joint vector J only contains angular information, then the joint vector J = (cosθ, cosθ, cosθ).
[0019] Since the number of bone vectors and the number of joint vectors are less than the number of joint points, the missing vectors are filled with 0 vectors to keep the dimension of the matrix unchanged.
[0020] Furthermore, the spatio-temporal self-attention mechanism includes a spatial self-attention mechanism module and a temporal self-attention mechanism module.
[0021] Furthermore, the spatial self-attention mechanism module performs self-attention calculation on all node vectors in the same frame as follows:[[]]END]]
[0022] Project the encoded sequence into query Q, key K, and value V, and use a convolutional layer with a kernel size of 1×1 to project the encoded sequence X in :
[0023] Q, K, V = Conv 2D(1×1) (X in )
[0024] By calculating the similarity between the query Q and the transpose of the key k, weights are obtained, and the dot product is used as the similarity function;
[0025] Then, the obtained weights are normalized using the Softmax function, and the weighted sum of the weights and the corresponding values V is calculated to obtain the self-attention map.
[0026]
[0027] Furthermore, the spatial self-attention mechanism module adopts the multi-head self-attention mechanism. If h groups are used, the final spatial self-attention map is:[[]]END]]
[0028]
[0029] Furthermore, the temporal self-attention mechanism module transforms the original dimension (N, C, T, V) into (N, C, V, T), where N is the total number of samples, C is the number of channels, T is the number of frames in a sample, and V is the number of nodes in a single frame. After the dimension transformation, the self-attention map for calculating all node vectors in the same frame is changed to calculating the self-attention map for the same node vector in different frames. Its subsequent calculation method is analogous to that of the spatial self-attention module, through the following formula:
[0030] Q t , K t , V t = Conv 2D(1×1) (X Atten )
[0031]
[0032]
[0033] Finally, the output XST of the spatio-temporal self-attention mechanism module can be obtained. Atten , where Q t , K t , V t represent the query vector, key vector, and value vector of the key points with the same semantic information in space at the t-th frame.
[0034] Furthermore, inter-frame aggregation aggregates discrete sub-poses in the same action through convolution. Specifically, for the output of the spatial self-attention mechanism module, it is regarded as a set of a series of sub-poses in different frames in the time dimension, and the convolution is used to obtain the correlation between them. The convolution calculation formula is:
[0035] X conv = Conv 2D(1×1) (XS Atten ).
[0036] Furthermore, the action features of the spatio-temporal self-attention mechanism and inter-frame aggregation are weighted and fused. Through the learnable parameters α and β, the optimal weights are calculated through continuous iterative learning, and finally the final classification result is output through the fully connected layer. The weighted fusion formula is:
[0037] X out = αXST Atten + βX conv
[0038] The loss function uses the cross-entropy loss function:
[0039]
[0040] where n is the number of samples, y is the output classification result X out , y true is the predicted label.
[0041] Furthermore, the motion feature classification results obtained by aggregating different input streams are used to generate the final behavior classification result. Specifically:
[0042] The behavior classification results obtained by inputting different input streams into the model are weighted and aggregated, and the weights are adjusted through variable parameters to obtain the optimal behavior classification result. The weighted aggregation formula is:
[0043] X Final_out = aX join_out + bX join_motion_out + cX bone_out + dX bone_motion_out + eX angle_out
[0044] Among them, X join_out is the output of the original joint point coordinate matrix, X bone_out is the human body bone vector matrix, X join_motion_out is the joint point motion trend matrix, X bone_motion_out is the human body bone motion trend matrix, X angle_out is the human body joint angle matrix, and a, b, c, d, e are variable weights.
[0045] A human behavior recognition system based on multi-stream bone features and spatio-temporal self-attention includes:
[0046] A multi-stream preprocessing module for converting the original joint point coordinate matrix into multiple modalities as different input streams;
[0047] A spatio-temporal self-attention mechanism module for extracting the long-term motion features of each node in the action and the dependence features between them;
[0048] An inter-frame aggregation module for aggregating each sub-pose in the complete action in the time dimension;
[0049] A human behavior classification output module for aggregating the long-term and short-term features with the optimal ratio and obtaining the behavior classification result through a fully connected layer.
[0050] A multi-stream motion feature aggregation module for aggregating the motion feature classification results obtained from multi-stream inputs and generating the final behavior classification result.
[0051] The present invention has the following beneficial effects compared with the prior art:
[0052] The present invention can obtain the features of human behavior from more semantics using multi-stream input information, and use the self-attention mechanism to obtain the mutual dependence relationship between each node in space. In the time dimension, a weighted fusion mechanism of self-attention and inter-frame aggregation module is adopted to effectively extract the long-term motion information and short-term motion information. Finally, the action features under different semantics are aggregated through the multi-stream feature aggregation module to obtain the optimal behavior classification result. Description of the Drawings
[0053] Figure 1 is a schematic diagram of the method flow of the present invention. Detailed Embodiments
[0054] The present invention will be further described below with reference to the drawings and embodiments.
[0055] Please refer to Figure 1 , the present invention provides a human behavior recognition system based on multi-stream bone features and spatio-temporal self-attention, including:
[0056] Multi-stream preprocessing module, used to convert the original joint point coordinate matrix into multiple modalities as different input streams;
[0057] Spatio-temporal self-attention mechanism module, used to extract the long-term motion features of each node in the action and the dependence features between them;
[0058] Inter-frame aggregation module, used to aggregate each sub-pose in the complete action in the time dimension;
[0059] Human behavior classification output module, used to aggregate the optimal proportion of long-term and short-term features and obtain the behavior classification result through the fully connected layer.
[0060] Multi-stream motion feature aggregation module, used to aggregate the motion feature classification results obtained from multi-stream inputs and generate the final behavior classification result.
[0061] In this embodiment, preferably, the multi-stream preprocessing module converts the original joint point coordinate matrix into input matrices with different semantics. Specifically, it converts the joint point coordinate matrix into a human body bone vector matrix, a joint point motion trend matrix, a human body bone motion trend matrix, a human body joint angle matrix, and the original joint point coordinate matrix.
[0062] Since each bone has two joints combined, the joint closer to the bone's center of gravity is defined as the source joint, and the joint farther from the center of gravity is defined as the target joint. Each bone is represented as a vector pointing from the source joint to the target joint, which contains not only length information but also direction information.
[0063] For example, given a bone, its source joint v 1 =(x 1 ,y 1 ,z 1 ), and the target joint v 2 =(x 2 ,y 2 ,z 2 ), the vector calculation of the human body bone is:
[0064] e v1,v2 =(x 2 -x 1 ,y 2 -y 1 ,z 2 -z 1 )
[0065] For the same node V (V can be either a joint point coordinate or a bone vector) in different frames t1 and t2, where V at time t1 is V t1 =(x t1 ,y t1 ,z t1 ), and V at time t2 is V t2=(x t2 ,y t2 ,z t2 ), then its motion trend vector can be expressed as:
[0066] m v1,v2 =(x t1 -x t2 ,y t1 -y t2 ,z t1 -z t2 )
[0067] By encoding each frame of the video sequence as a set of angles to describe the relative positions of different body parts, these angles can be derived from the human body bone data. The angle between two bone segments can be defined by the included angle between the bone vectors. Let the two bone vectors e 1 , e 2 in the same frame be e 1 =(x 3 ,y 3 ,z 3 ), e 2 =(x 4 ,t 4 ,z 4 ), then the cosine of the included angle between the bone vectors is:
[0068]
[0069] The joint vector J only contains angle information, so the joint vector J=(cosθ,cosθ,cosθ).
[0070] Since the number of bone vectors and the number of joint vectors are less than the number of joint points, the missing vectors are filled with 0 vectors to keep the dimension of the matrix unchanged.
[0071] In this embodiment, preferably, the spatio-temporal self-attention mechanism module is divided into a spatial self-attention mechanism module and a temporal self-attention mechanism module. In the spatial self-attention mechanism module, self-attention calculation is performed on all node vectors in the same frame. Specifically, when calculating self-attention, not only the influence of all other nodes on node n i should be considered, but also the influence of node n i on other nodes must be considered. Therefore, the encoded sequence is usually projected into query Q, key K, and value V.
[0072] This paper uses a convolutional layer with a 1×1 kernel size to project the encoded sequence X in :
[0073] Q, K, V = Conv 2D(1×1) (X in )
[0074] Then, by calculating the similarity between the query Q and the transpose of the key k, the weights can be obtained, and the dot product is used as the similarity function. Then, the obtained weights are normalized using the Softmax function. Then, the weighted sum of the weights and the corresponding values V is calculated to obtain the self-attention map.
[0075]
[0076] Preferably, the multi-head self-attention mechanism is adopted to enable the model to learn relevant information in different representation subspaces. Specifically, the self-attention operation projects multiple groups of Q, K, and V by different learnable parameters, and then concatenates multiple groups of self-attention matrices. If h groups are adopted, the final spatial self-attention map is:
[0077]
[0078] The self-attention map output by the spatial self-attention mechanism module is used as the input of the temporal self-attention mechanism module after changing the dimension. Specifically, the original dimension (N, C, T, V) is transformed into (N, C, V, T), where N is the total number of samples, C is the number of channels, T is the number of frames in a sample, and V is the number of nodes in a single frame. It can be seen that after the dimension transformation, the self-attention map for calculating all node vectors in the same frame is changed to calculating the self-attention map for the same node vector in different frames. The subsequent calculation method is analogous to the spatial self-attention module and is calculated through the following formula:
[0079] Q t ,K t ,V t =Conv 2D(1×1) (X Atten )
[0080]
[0081]
[0082] Finally, the output XST of the spatio-temporal self-attention mechanism module can be obtained Atten .
[0083] In this embodiment, preferably, the inter-frame aggregation module mainly aggregates the discrete sub-poses in the same action through convolution. Specifically, for the output of the spatial self-attention mechanism module, it can be regarded as a set of a series of sub-poses in different frames in the time dimension. Using convolution to obtain the correlation between them can more effectively capture the short-term features of the action and complement the spatio-temporal self-attention mechanism module 2. The convolution calculation formula is:
[0084] X conv =Conv2D(1×1) (XS Atten )
[0085] In this embodiment, preferably, the human behavior classification output module performs weighted fusion on the action features in the spatio-temporal self-attention mechanism module and the inter-frame aggregation module. Through continuous iterative learning with learnable parameters α and β, the optimal weights are calculated, and finally the final classification result is output through a fully connected layer. The weighted fusion formula is:
[0086] X out = αXST Atten + βX conv
[0087] The loss function adopts the cross-entropy loss function:
[0088]
[0089] where n is the number of samples, y is the output classification result X out , y true is the predicted label.
[0090] In this embodiment, preferably, the multi-stream motion feature aggregation module
[0091] In this part, the behavior classification results obtained by inputting different input streams into the model are weighted and aggregated, and the weights are adjusted through variable parameters to obtain the optimal behavior classification result. The weighted aggregation formula is:
[0092] X Final_out = aX join_out + bX join_motion_out + cX bone_out + dX bone_motion_out + eX angle_out
[0093] where X join_out is the output of the original joint point coordinate matrix, X bone_out is the human body bone vector matrix, X join_motion_out is the joint point motion trend matrix, X bone_motion_out is the human body bone motion trend matrix, X angle_out is the human joint angle matrix, and a, b, c, d, e are variable weights.
[0094] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope of the present invention.
Claims
1. A human behavior recognition method based on multi-stream skeletal features and spatio-temporal self-attention, characterized in that, it includes the following steps: Obtain the original joint point coordinate matrix and transform it into multiple modalities as different input streams; For different input streams, use the spatio-temporal self-attention mechanism to extract the long-term motion features of each node in the action and the dependence features between them respectively, and aggregate each sub-pose in the complete action frame by frame in the time dimension; Aggregate the optimal proportion of long-term and short-term features, and obtain the behavior classification result through the fully connected layer; Aggregate the motion feature classification results obtained from different input streams to generate the final behavior classification result; Transform the joint point coordinate matrix into a human skeletal vector matrix, a joint point motion trend matrix, a human skeletal motion trend matrix, a human joint angle matrix, and the original joint point coordinate matrix; The spatio-temporal self-attention mechanism includes a spatial self-attention mechanism module and a temporal self-attention mechanism module; Frame aggregation aggregates discrete sub-poses in the same action through convolution. Specifically: regarding the output of the spatial self-attention mechanism module as a set of a series of sub-poses in different frames in the time dimension, use convolution to obtain the correlation between them. The convolution calculation formula is: X conv = Conv 2D(1×1) (XS Atten ) Weightedly fuse the action features of the spatio-temporal self-attention mechanism and frame aggregation. Through the learnable parameters α and β, continuously iterate and learn to calculate the optimal weights, and finally output the final classification result through the fully connected layer. The weighted fusion formula is: X out = αXST Atten + βX conv wherein, XST Atten is the output of the spatio-temporal self-attention mechanism module; The loss function uses the cross-entropy loss function: where n is the number of samples, y is the output classification result X out , y true is the predicted label; The aggregating the motion feature classification results obtained from different input streams to generate the final behavior classification result is specifically: Weightedly aggregate the behavior classification results obtained by inputting different input streams into the model, and adjust each weight through variable parameters to obtain the optimal behavior classification result; The weighted aggregation formula is: X Final_out = aX join_out + bX join_motion_out + cX bone_out + dX bone_motion_out + eX angle_out Among them, X join_out is the output of the original joint point coordinate matrix, X bone_out is the human body bone vector matrix, X join_motion_out is the joint point movement trend matrix, X bone_motion_out is the human body bone movement trend matrix, X angle_out is the human body joint angle matrix, and a, b, c, d, e are variable weights.
2. The human behavior recognition method based on multi-stream skeletal features and spatio-temporal self-attention according to claim 1, characterized in that, transform the original joint point coordinate matrix into multiple modalities as different input streams, specifically as follows: Define the joint close to the center of gravity of the bone as the source joint, and the joint far from the center of gravity as the target joint. Each bone is represented as a vector pointing from the source joint to the target joint, containing length information and direction information; given a bone, its source joint v 1 =(x 1 , y 1 , z 1 ), and the target joint v 2 =(x 2 , y 2 , z 2 ). The vector calculation of the human body skeleton is as follows: e v1,v2 = (x 2 - x 1 , y 2 - y 1 , z 2 - z 1 ) For the same node V in different frames t1 and t2, where V is V at time t1 t1 =(x t1 , y t1 , z t1 ), and V is V at time t2 t2 =(x t2 , y t2 , z t2 ), then its motion trend vector is expressed as: m v1,v2 = (x t1 - x t2 , y t1 - y t2 , z t1 - z t2 ) The angle between two bone segments is defined by the included angle between bone vectors. Let two bone vectors e 1 , e 2 in the same frame be e 1 = (x 3 , y 3 , z 3 ), e 2 = (x 4 , y 4 , z 4 ), respectively. Then the cosine of the included angle between the bone vectors is: The joint vector J only contains angle information, then the joint vector J = (cosθ, cosθ, cosθ).
3. The human behavior recognition method based on multi-stream skeletal features and spatio-temporal self-attention according to claim 2, characterized in that, the spatial self-attention mechanism module performs self-attention calculation on all node vectors in the same frame, specifically as follows: Project the encoded sequence into query Q, key K, and value V', using a convolutional layer with a kernel size of 1×1 to project the encoded sequence X in : Q, K, V' = Conv 2D(1×1) (X in ) Calculate the similarity between the query Q and the transpose of the key K to obtain the weight, and use the dot product as the similarity function; Then use the Softmax function to normalize the obtained weight, and perform weighted sum on the weight and the corresponding value V' to obtain the self-attention map; In the formula, Softmax represents the Softmax function.
4. The human behavior recognition method based on multi-stream skeletal features and spatio-temporal self-attention according to claim 3, characterized in that, the spatial self-attention mechanism module adopts the multi-head self-attention mechanism, with h groups, then the final spatial self-attention map is:
5. The human behavior recognition method based on multi-stream skeletal features and spatio-temporal self-attention according to claim 4, characterized in that, The time self-attention mechanism module transforms the original dimension (N, C, T', V”) into (N, C, V, T), where N is the total number of samples, C is the number of channels, T' is the number of frames in a sample, and V” is the number of nodes in a single frame. After the dimension transformation, the self-attention map that originally calculates all node vectors in the same frame becomes the self-attention map that calculates the same node vector in different frames. Its subsequent calculation method is analogous to that of the spatial self-attention module, through the following formula: Q t , K t , V t = Conv 2D(1×1) (X Atten ) Finally, the output XST of the spatio-temporal self-attention mechanism module can be obtained. Atten , where Q t , K t , V t represent the query vector, key vector, and value vector of the key points with the same semantic information in space at the t-th frame.
6. A system for implementing the human behavior recognition method based on multi-stream skeleton features and spatio-temporal self-attention according to any one of claims 1-5, characterized in that, it includes: A multi-stream preprocessing module for converting the original joint point coordinate matrix into multiple modalities as different input streams; A spatio-temporal self-attention mechanism module for extracting the long-term motion features of each node in the action and the dependence features between them; An inter-frame aggregation module for aggregating each sub-pose in the complete action in the time dimension; A human behavior classification output module for aggregating the long-term and short-term features with the optimal ratio and obtaining the behavior classification result through a fully connected layer; A multi-stream motion feature aggregation module for aggregating the motion feature classification results obtained from multi-stream inputs and generating the final behavior classification result.
Citation Information
Patent Citations
Abnormal gait recognition method and system based on space-time attention enhancement graph convolution
CN113887486A
Human body action recognition method, human body action recognition system, and device
WO2022000420A1