Human body action recognition method based on combination of space-time diagram convolution and multi-head cross attention

By combining the spatiotemporal graph convolution and multi-head cross-attention mechanism, the problem of dynamic coupling of joint nodes and bone features in human body movement recognition is solved, achieving higher recognition accuracy and lower computing complexity.

CN120340131APending Publication Date: 2025-07-18XIAN UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510408342.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing human body movement recognition method based on graph convolution networks is difficult to fully consider the dynamic coupling of spatiotemporal features between joint nodes and bones, resulting in an increase in parameters and an increase in computational volume, and traditional attention mechanisms are difficult to capture the complex dependence between multimodal features.

Method used

Combining the spatial and temporal graph convolution and the multi-head cross-attention mechanism, deep feature extraction is performed on the joint node position flow and velocity flow of the human body action sequence, and feature fusion is performed using the temporal multi-head cross-attention mechanism and the spatial multi-head cross-attention mechanism, and the fusion of spatial and temporal features is achieved through the dual-path gate mechanism.

Benefits of technology

It improves the accuracy of human body movement recognition, solves the problem of dynamic coupling of spatial and temporal features in GCN in modeling, and improves the model's representation ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340131A_ABST
    Figure CN120340131A_ABST
Patent Text Reader

Abstract

The invention discloses a human body action recognition method based on combination of space-time diagram convolution and multi-head cross attention, and the method specifically comprises the steps: carrying out the preprocessing of the data of a human body action sequence, and obtaining a joint point position flow and a velocity flow; extracting depth features of the human body joint point position flow; extracting depth features of the velocity flow of the human body articulation points; using a time multi-head cross attention mechanism and a space multi-head cross attention mechanism to fuse the features of the position flow and the velocity flow of the joint points; fusing a joint point position flow and a speed flow passing through a time and space multi-head cross attention mechanism by using double-path gating; and performing feature extraction by using the serial mainstream space-time diagram convolution blocks to complete action recognition. According to the method provided by the invention, through combination of the GCN and the multi-head cross attention mechanism, the human body action characterization capability of the deep network model can be improved, so that the action recognition precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and machine learning methods, and particularly relates to a human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention. Background Art

[0002] At present, extensive research progress has been made in behavior recognition based on human skeleton action sequences. In these studies, human skeleton behaviors are usually composed of limb positions in space and limb changes in time. When performing behavior recognition, it is necessary to extract spatial features and temporal features simultaneously, that is, to obtain the spatio-temporal features of human skeleton data. In the behavior recognition methods of human skeleton action sequences, there are mainly traditional methods and deep learning-based methods; traditional methods extract the spatio-temporal information of skeleton data through manually designed features, such as features like the positions of joint points, the lengths and angles of bones, etc., and then use pattern recognition techniques for classification and recognition. However, since the manually designed features rely on human understanding of behaviors, it is often difficult to fully mine the complex information hidden in the data directly using these features. Different from this, deep learning-based methods can automatically extract more informative features from the data through training on a large amount of data, and have achieved remarkable results in classification and recognition tasks.

[0003] In recent years, methods based on the Graph Convolutional Network (GCN) have achieved relatively good results in human skeleton action recognition. The graph data structure is used to represent the temporal and spatial adjacency relationships between the joint points and bones of the human skeleton, and then convolution is used to extract features from the adjacent joint points and bones of the human action sequence. A classic algorithm is the Spatial Temporal Graph Convolutional Networks (ST-GCN). Currently, most methods based on spatio-temporal graph convolutional networks and their improved methods have achieved high recognition accuracies on public datasets (such as NTU-RGB+D, etc.). However, methods based on graph convolutional networks need to consider features such as the static physical connections (hard connections) between the joint points or bones of the human body and the dynamic connection relationships between the joint points and bones in the action sequence. Usually, features of different modalities are directly concatenated together, such as features like the positions of joint points, the speeds of joint points, and bone vectors. The movements of joint points and the changes in the skeleton structure are dynamic and sequential. Simple concatenation may not fully consider the complex dependencies of these features in time and space, and will lead to an increase in parameters and thus an increase in the computational amount.

[0004] The attention mechanism is currently widely used in the field of deep learning. However, traditional attention mechanisms can only capture features of one data modality, while cross-attention can fuse features of two different data modalities. In addition, ordinary single-head attention mechanisms are difficult to capture all important features of the input sequence. Introducing the multi-head attention mechanism can focus on different parts of the input sequence from multiple perspectives. The multi-head cross-attention mechanism (MHCA) combines the advantages of the multi-head attention mechanism and the cross-attention mechanism and is used for the interaction between two different sequences, that is, to capture the correlation between different modalities or different sequences, which can avoid the situation where the model processes directly concatenated multi-modal features and cannot fully consider the relationship between features. And using multiple attention heads can learn different attention distributions in parallel, enabling the model to simultaneously focus on different key points or feature patterns, thereby obtaining richer feature representations. In human action recognition based on skeletal points, the position of human joints in the next frame is affected by the direction of velocity. The multi-head cross-attention mechanism can fuse the position of joints and the direction of velocity (soft connection), avoiding the situation where the relationship between these features cannot be fully considered due to direct concatenation. Summary of the Invention

[0005] The object of the present invention is to provide a human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention, which can improve the representation ability of actions and thus improve the recognition accuracy of actions.

[0006] The technical solution adopted by the present invention is a human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention, which is specifically implemented according to the following steps:

[0007] Step 1: Preprocess the data of the human action sequence to obtain the joint position stream and the velocity stream;

[0008] Step 2: Extract the depth features of the human joint position stream;

[0009] Step 3: Extract the depth features of the human joint velocity stream;

[0010] Step 4: Use the temporal multi-head cross-attention mechanism and the spatial multi-head cross-attention mechanism to fuse the features of the joint position stream processed in Step 2 and the velocity stream processed in Step 3;

[0011] Step 5: Use dual-path gating to fuse the joint position stream and the velocity stream that have passed through the temporal and spatial multi-head cross-attention mechanisms;

[0012] Step 6: Use a series of mainstream spatio-temporal graph convolution blocks to extract features from the result obtained in Step 5;

[0013] Step 7: Complete action recognition.

[0014] The features of the present invention also lie in that

[0015] In step 1, the human body motion sequence is represented as X, T represents the number of frames of the motion sequence, V represents the number of human body joint points, and C represents the number of features of each joint point; the preprocessing process is specifically as follows:

[0016] Step 1.1: Calculate the relative positions of the joint points, that is, the joint point position stream;

[0017] Calculate the set of relative positions of the joint points Q = {q i |i = 1, 2,..., V in}, and the calculation formula of q i is shown in Equation (1);

[0018] q i = x[:, :, i] - x[:, :, c] (1)

[0019] Step 1.2: Calculate the velocities of the joint points

[0020] Divide the velocity information into two groups. The fast velocity is the time difference between adjacent frames, representing the short-term changes of the motion, that is, the fast motion, which is represented by F = {f t |t = 1, 2,..., T in}; the slow velocity is the time difference between spaced frames, representing the long-term changes of the motion, that is, the slow motion, which is represented by S = {s t |i = 1, 2,..., T in}, as shown in Equations (2) and (3);

[0021] f t = x[:, t + 1, :] - x[:, t, :] (2)

[0022] S t = x[:, t + 2, :] - x[:, t, :] (3)

[0023] In step 2, it is specifically as follows:

[0024] Step 2.1: Use cascaded spatio-temporal graph convolutional blocks for feature extraction;

[0025] Use 1 initial spatio-temporal graph convolutional block GCN for graph convolution, and the specific calculation formula is shown in Equation (4):

[0026]

[0027] Connect the output feature maps of each frame to obtain

[0028] The joint point position stream is processed through two cascaded spatio-temporal graph convolution blocks; one spatio-temporal graph convolution block consists of 1 GCN block, 2 TCN blocks, and 1 ST-Att connected in sequence; the GCN module is the same as the initial GCN module, and the calculation formula of the TCN module is as shown in Equation (5):

[0029] X STGCN = X GCN * W(5)

[0030] where * represents convolution in the time dimension, and the size of the convolution kernel is

[0031] The input features are divided into two streams for separate processing. These two streams of data are the same. One stream considers the time frames as a whole and pools the features of the joint points, that is The other stream processes the joint points as a whole and pools the features of the time frames, that is Then, they are concatenated along the channel dimension to obtain After the concatenated features are processed by the fully connected layer, the data stream is then separated along the channel dimension. One stream of data extracts features at the frame level, and the other stream of data extracts features at the joint point level. The outer product of these two sets of data is taken, and then the data after feature extraction is applied to the original feature map, that is, multiplied element-wise with the original input feature map, to obtain the key frames and key joint points in the action sequence The specific calculation steps are as shown in Equation (6) and Equation (7):

[0032] X FC = σ((pool t (X STGCN ) ⊕ pool v (X STGCN )) · W)(6)

[0033]

[0034] where pool t (X STGCN ) and pool v (X STGCN ) respectively represent pooling the joint point features with the time frames as a whole; pooling the time frame features with the joint points as a whole; ⊕ represents the concatenation operation, σ is the activation function, X FC is the feature map after being processed by the fully connected layer; X FC · W t and X FC · W v respectively represent feature extraction at the frame level and feature extraction at the joint point level, ω is the activation function, represents the feature-level outer product, ⊙ represents element-by-element multiplication;

[0035] Stack two spatiotemporal graph convolution blocks to complete the calculation of the spatiotemporal graph convolution of the position stream, and finally output X out_j .

[0036] In step 3, specifically:

[0037] First, use an initial GCN block to perform graph convolution. The specific calculation formula is shown in formula (8):

[0038]

[0039] Connect the output feature maps of each frame to get

[0040] The joint velocity stream is processed by two serially connected spatiotemporal graph convolution blocks. A spatiotemporal graph convolution block is composed of a sequentially connected GCN block, two TCN blocks, and one ST-Att. The GCN module is the same as the initial GCN module, and the formula of the TCN module is shown in formula (9):

[0041] X′ STGCN =X′ GCN *W (9)

[0042] The input features are divided into two streams for processing. The two streams have exactly the same data. One stream considers the time frame as a whole and pools the features of the joints, that is, The other stream processes the joint point as a whole and pools the features of the time frame, i.e. Then concatenate them along the channel dimension to obtain After the concatenated features are processed by the fully connected layer, the data stream is separated along the channel dimension. Feature extraction is performed at the frame level, and another stream of data Feature extraction is performed at the joint point level, the two sets of data are outer-producted, and the feature-extracted data is applied to the original feature map, that is, element-by-element multiplication with the original input feature map to obtain the key frames and key joint points in the action sequence. The specific calculation steps are shown in formula (10) and formula (11):

[0043] X′ Fc =σ((pool t (X′ STGCN )⊕pool v (X′ STGCN ))·W) (10)

[0044]

[0045] Stack two spatio-temporal graph convolutional blocks. Thus, the calculation of the spatio-temporal graph convolution of the velocity flow is completed, and the final output is X. out_v 。

[0046] In step 4, specifically:

[0047] Step 4.1: Dynamic window-controlled temporal multi-head cross-attention mechanism;

[0048] First, perform dimensionality reorganization on the input data X out_j and X out_v , and then use the multi-head cross-attention mechanism to model the relationship between X out_j and X out_v . Specifically, generate the query Q t by passing the input position flow through a linear layer, generate the key K t and the value V t . The specific formulas are shown in Equations (12) and (13):

[0049] Q t =Linear Q (X out_j ) (12)

[0050] K t , V t =Linear KV (X out_v ) (13)

[0051] Calculate the dot product between the query Q t and the transpose of the key K t to obtain the attention score matrix, as shown in Equation (14):

[0052]

[0053] Introduce the time window mask matrix M ∈ {0, 1} T×T , defined as Equation (15):

[0054]

[0055] where W is the size of the time window. Apply the mask to the attention score matrix to generate the locally constrained attention scores, as shown in Equation (16):

[0056]

[0057] where A masked [t, t′] are the locally constrained attention scores. Use the softmax function to convert the attention scores into a probability distribution, representing the attention weights of the query to each key. Apply the locally constrained attention weights to the value V tabove, as shown in Equation (17):

[0058] X Attn_heads = softmax(A masked [t,t′])·V (17)

[0059] Finally, the subspaces processed by each head are concatenated to obtain the final output X of the dynamic window-controlled temporal multi-head cross-attention mechanism module Att_T , as shown in Equation (18):

[0060] X Att_T = Concat(X att_1 ,X att_2 ,…,X att_heads ) (18)

[0061] where X att_1 ,X att_2 ,…,X att_heads are the number of heads of attention for heads;

[0062] Step 4.2: Multimodal dynamic weighted spatial multi-head cross-attention;

[0063] First, the input data X out_j and X out_v are dimensionally reorganized, and then the multi-head cross-attention mechanism is used to model the relationship between X out_j and X out_v . Specifically, the input position stream is used to generate the query Q v through a linear layer, the velocity stream generates the key K v and the value V v , as shown in Equations (19) and (20):

[0064] Q v = Linear Q (X out_j ) (19)

[0065] K v ,V v = Linear KV (X out_v ) (20)

[0066] Calculate the similarity between them to obtain the attention score matrix, and the specific formula is shown in Equation (21):

[0067]

[0068] Introduce sparse attention constraints, that is, the original attention scores are simultaneously normalized by the Softmax function and non-linearly screened by the Sigmoid function, and the two are balanced with a coefficient of 0.1, as shown in Equation (22):

[0069] A enhanced = Softmax(A v ) + 0.1 * σ(A v ) (22)

[0070] Introduce learnable weight parameters for each attention head, and then normalize them through the Softmax function and weight them to the attention matrix, as shown in Equation (23):

[0071] A weighted = A enhanced ⊙ Softmax(W heads ) (23)

[0072] Then, convert the attention matrix into a probability distribution through the Softmax function, and then apply it to the value V v as shown in Equation (24):

[0073] X Attn_heads = Softmax(A weighted ) · V v (24)

[0074] Concatenate the subspaces processed by each head to obtain the output of the total space, as shown in Equation (21):

[0075] X Att = Concat(X att_1 , X att_2 , …, X att_heads ) (21)

[0076] where X att_1 , X att_2 , …, X att_heads is the number of heads of heads attention.

[0077] Introduce a gating mechanism to mix the new features and the original input features according to the gating ratio, as shown in Equations (25) and (26):

[0078] G = σ(W g * X Att_V + b g ) (25)

[0079]

[0080] where σ is the Sigmoid function, b g is the bias term of the gating linear layer, is the initial value of the multi-head cross-attention in the input space.

[0081] In step 5, specifically:

[0082] Perform multi-scale convolution on the data stream output by the spatial multi-head cross-attention mechanism, as shown in equations (27) and (28):

[0083] MSConv(X) = Conv 3×3 (X Att_V ) ⊕ Conv 5×5 (X Att_V ) (27)

[0084] X enhanced = X Att_V + W fuse ·MSConv(X) (28)

[0085] Among them, ⊕ is to splice the features after 3×3 and 5×5 convolutions, and W fuse is the weight of the convolutional layer;

[0086] Perform dual-path gating fusion on X Att_T and X enhanced . They are channel-level gating and spatial-level gating respectively. Before that, in order to process global information, first splice X Att_T and X enhanced , as shown in equation (29):

[0087] X fused = Concat(X Att_T , X enhanced ) (29)

[0088] Perform channel-level gating fusion on the two groups of data, as shown in equations (30) and (31):

[0089] G C = σ(W C2 · RELU(W C1 · GAP(X fused ))) (30)

[0090] X channel = G C · X Att_T + G C · X enhanced (31)

[0091] Among them, W C1 and W C2 are both channel gating convolution parameters, RELU is the activation function, GAP is global average pooling, and G C is the generated gating value;

[0092] Send the data stream after channel-level gating into spatial-level gating processing, as shown in equations (32) and (33):

[0093] G S = σ (W S · X fused ) (32)

[0094] X CS_Fused = G S ⊙ X channel + (1 - G S ) ⊙ Mean(X fused ) (33)

[0095] To retain the information of the original spatio-temporal features, cross-residual connections are introduced, as shown in formula (34);

[0096]

[0097] In step 6, specifically:

[0098] The data stream is processed through two cascaded spatio-temporal graph convolutional blocks; one spatio-temporal graph convolutional block consists of 1 GCN block, 2 TCN blocks, and 1 ST-Att connected in sequence; 1 GCN module is as shown in formula (35):

[0099]

[0100] Among them, and respectively represent the input and output feature maps;

[0101] Connecting the output feature maps of each frame gives

[0102] The calculation formula of 1 TCN module is as shown in formula (36):

[0103]

[0104] 1 ST-Att module divides the input features into two streams for separate processing. These two streams of data are exactly the same. One stream considers the time frames as a whole and pools the features of the joint points, that is The other stream processes the joint points as a whole and pools the features of the time frames, that is Then they are concatenated along the channel dimension to obtain After processing the concatenated features through a fully connected layer, the data stream is then separated along the channel dimension. One stream of data extracts features at the frame level, and the other stream of data extracts features at the joint point level. The outer product of these two sets of data is taken, and then the data after feature extraction is applied to the original feature map, that is, multiplied element-wise with the original input feature map, to obtain the key frames and key joint points in the action sequence The specific calculation steps are shown in Equation (37) and Equation (38):

[0105]

[0106] Among them, and respectively represent pooling joint point features with the time frame as a whole, and pooling time frame features with the joint point as a whole; is the feature map after the processing of the fully connected layer; and respectively represent feature extraction at the frame level and feature extraction at the joint point level.

[0107] The beneficial effects of the present invention are as follows: The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention of the present invention. The spatio-temporal graph convolution GCN can represent the joint points of the human skeleton and the connections of the bones. Through the feature propagation mechanism of GCN, it can display and model the physical topological connection relationship of the joint points in the human skeleton. Its hierarchical convolution operation captures the bone topological features in the spatial dimension; on this basis, a multi-head cross-attention mechanism is introduced to perform feature decoupling and interactive fusion on the spatial coordinates (position stream) and motion speeds (speed stream) of the joint points respectively, solving the problem that it is difficult for GCN to model the dynamic coupling of spatio-temporal features. At the same time, the multi-head cross-attention mechanism strengthens the action continuity through a local adjacent frame sliding window in the time dimension, and restricts the attention weights to be distributed only in the temporal neighborhood of adjacent frames; in the spatial dimension, the weight of important positions is enhanced through sparse attention constraints. Therefore, the combination of GCN and the multi-head cross-attention mechanism can improve the human action representation ability of the deep network model, and thus improve the accuracy of action recognition. Description of the Drawings

[0108] Figure 1 is a flowchart of the human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention of the present invention;

[0109] Figure 2 is a schematic structural diagram of the spatio-temporal joint attention (ST-Att) module in the method of the present invention. Specific Embodiments

[0110] The present invention will be described in detail below in conjunction with the specific embodiments and the drawings.

[0111] Embodiment 1

[0112] The present invention aims at the human bone joint sequence obtained based on the motion capture technology, and realizes the feature modeling and fusion of the position and speed direction of human joint points by combining stacked spatio-temporal graph convolution blocks with a multi-head cross-attention mechanism, and then realizes the classification and recognition of human actions. As Figure 1As shown in the figure, first, preprocess the input data, that is, the 3D coordinate positions of the joint points in the 3D space, to obtain two types of input data, namely the original positions of the joint points, the relative positions relative to the central joint point, and the directions of the joint point velocities. Subsequently, input the two streams (joint stream, velocity direction stream) into 1 initial GCN block and 2 spatio-temporal graph convolutional blocks connected in series respectively. After that, both streams of data pass through the temporal multi-head cross-attention layer and the spatial multi-head cross-attention layer, are fused in the fusion layer, and then the fused data stream is processed by the main stream composed of 2 spatio-temporal graph convolutional blocks, and finally enters the linear layer for classification and recognition.

[0113] Embodiment 2

[0114] The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention of the present invention is specifically implemented according to the following steps:

[0115] Step 1: Preprocess the data of the human action sequence;

[0116] Represent the human action sequence obtained by the motion capture technology as X, T represents the number of frames of the action sequence, V represents the number of human joint points, and C represents the number of features (number of channels) of each joint point. Usually C = 3, representing the coordinates of a joint point in the three-dimensional space; the preprocessing of the human action sequence data is as follows:

[0117] Step 1.1: Calculate the relative positions of the joint points;

[0118] Calculate the set of relative positions of the joint points Q = {q i |i = 1, 2,..., V in}, where the calculation formula of q i is shown in Equation (1);

[0119] q i = x[:, :, i] - x[:, :, c] (1)

[0120] where c is the index of the central spine joint, and the position of the joint point is represented by the original position X and the relative position Q together;

[0121] Step 1.2: Calculate the velocities of the joint points

[0122] Divide the velocity information into two groups. The fast velocity is the time difference between adjacent frames, representing the short-term change of the action and capable of capturing the dynamic characteristics of fast actions, that is, fast actions, which is represented by F = {f t |t = 1, 2,..., T in}; the slow velocity is the time difference between spaced frames, representing the long-term change of the action and capable of capturing the overall trend of slow actions, that is, slow actions, which is represented by S = {s t|i = 1, 2, …, T in} is expressed as shown in Equation (2) and Equation (3);

[0123] f t = x[:, t + 1, :] - x[:, t, :] (2)

[0124] S t = x[:, t + 2, :] - x[:, t, :] (3)

[0125] Among them, f t represents the difference between adjacent frames, and s t represents the difference between spaced frames. Both f t and s t are vectors, containing the magnitude and direction information of the velocity. The input of the velocity information is jointly composed of F and S.

[0126] Step 2: Extract the depth features of the human joint point position flow;

[0127] After obtaining the positions of the joint points (position flow) after the preprocessing in Step 1.1, the position flow is sent into the stacked initial spatio-temporal graph convolution block for processing, as Figure 1 shown. Specifically, it is divided into four steps, and the detailed steps are as follows:

[0128] Step 2.1: Use the concatenated spatio-temporal graph convolution block for feature extraction;

[0129] First, use 1 initial spatio-temporal graph convolution block GCN for graph convolution. The specific calculation formula is as shown in Equation (4):

[0130]

[0131] Among them, v ti represents the i-th joint point in the t-th frame, N(v tj ) represents all the neighborhood nodes of v ti , Z ti is used to balance the contributions of different neighborhood nodes, d(v ti , v tj ) represents the distance between v ti and v tj , f in (·) and f out (·) respectively represent the input and output feature maps. w() is the weight function, which assigns different weights to the nodes according to the different distances d(v ti , v tj );

[0132] Connect the output feature maps of each frame to obtain (indicating the size).

[0133] After the initial GCN block acts, the position stream is then processed through two cascaded spatio-temporal graph convolutional blocks. One spatio-temporal graph convolutional block consists of 1 graph convolutional (GCN) block, 2 temporal convolutional (TCN) blocks, and 1 spatio-temporal joint attention module (ST-Att) connected in sequence. The GCN module is the same as the initial GCN module, and the calculation formula of the TCN module is shown in Equation (5):

[0134] X STGCN = X GCN * W(5)

[0135] where, * represents convolution in the time dimension, and the size of the convolution kernel is that is, it slides 5 frames in the time dimension and has no effect in the spatial dimension.

[0136] To extract the most informative key frames and key joints in an action sequence, the data stream is input into the spatio-temporal joint attention (ST-Att) module, as Figure 2 shown, the input features are divided into two streams for separate processing. These two streams of data are exactly the same. One stream considers the time frames as a whole and pools the features of the joints, that is The other stream processes the joints as a whole and pools the features of the time frames, that is Then they are concatenated along the channel dimension to obtain After the concatenated features are processed by the fully connected layer (FC), the data stream is then separated along the channel dimension. One stream of data extracts features at the frame level, and the other stream of data extracts features at the joint level. The outer product of these two sets of data is taken, and then the data after feature extraction is applied to the original feature map, that is, multiplied element-wise with the original input feature map, to obtain the key frames and key joints in the action sequence The specific calculation steps are shown in Equations (6) and (7):

[0137] X FC = σ((pool t (X STGCN ) ⊕ pool v (X STGCN )) · W)(6)

[0138]

[0139] where, pool t (X STGCN ) and pool v (X STGCNrespectively represent pooling joint point features with the time frame as a whole and pooling time frame features with the joint point as a whole. ⊕ represents the concatenation operation, σ is the activation function, and X FC is the feature map after the fully connected layer processing. X FC ·W t and X FC ·W v respectively represent feature extraction at the frame level and feature extraction at the joint point level, ω is the activation function, represents the feature-level outer product, and ⊙ represents element-wise multiplication.

[0140] Stack two spatio-temporal graph convolution blocks. Thus, the calculation of the spatio-temporal graph convolution of the position flow is completed, and the final output is X out_j .

[0141] Step 3: Deep feature extraction of human joint point velocity flow;

[0142] Send the velocity (velocity flow) of the joint points obtained after the preprocessing in Step 1.2 into the stacked initial spatio-temporal graph convolution blocks for processing, which is specifically divided into four steps. The detailed steps are as follows:

[0143] First, use 1 initial GCN block to perform a basic graph convolution. The specific calculation formula is shown in Equation (8):

[0144]

[0145] where v ti represents the i-th joint point in the t-th frame, N(v tj ) represents all the neighborhood nodes of v ti , Z ti is used to balance the contributions of different neighborhood nodes, d(v ti , v tj ) represents the distance between v ti and v tj , f i ′ n (·) and f o ′ ut (·) respectively represent the input and output feature maps, and w() is the weight function, which assigns different weights to the nodes according to the different distances d(v ti , v tj ).

[0146] Connect the output feature maps of each frame to obtain

[0147] After the initial GCN block acts, the velocity flow is then processed through two cascaded spatio-temporal graph convolutional blocks. A spatio-temporal graph convolutional block consists of 1 graph convolutional (GCN) block, 2 temporal convolutional (TCN) blocks, and 1 spatio-temporal joint attention module (ST-Att) connected in sequence. The GCN module is the same as the initial GCN module, and the formula of the TCN module is shown in Equation (9):

[0148] X′ STGCN = X′ GCN *W (9)

[0149] where * represents convolution in the time dimension, and the size of the convolution kernel is that is, it slides 5 frames in the time dimension and has no effect in the spatial dimension.

[0150] To extract the most informative key frames and key joints in an action sequence, the velocity flow is input into the spatio-temporal joint attention module. The input features are divided into two streams for separate processing. These two streams of data are exactly the same. One stream considers the time frames as a whole and pools the features of the joints, that is The other stream processes the joints as a whole and pools the features of the time frames, that is Then they are concatenated along the channel dimension to obtain After the concatenated features are processed by the fully connected layer (FC), the data stream is then separated along the channel dimension. One stream of data extracts features at the frame level, and the other stream of data extracts features at the joint level. The outer product of these two sets of data is taken, and then the data after feature extraction is applied to the original feature map, that is, multiplied element-wise with the original input feature map, to obtain the key frames and key joints in the action sequence The specific calculation steps are shown in Equations (10) and (11):

[0151] X′ FC = σ((pool t (X′ STGCN ) ⊕ pool v (X′ SvGCN )) · W) (10)

[0152]

[0153] where pool t (X′ STGCN ) and pool v (X′ STGcN ) respectively represent pooling the joint features with the time frames as a whole; pooling the time frame features with the joints as a whole. ⊕ represents the concatenation operation, σ is the activation function, X FCThat is the feature map after the fully connected layer processing. X' FC ·W t and X′ FC ·W v respectively represent feature extraction at the frame level and feature extraction at the joint level. ω is the activation function, represents the outer product at the feature level, and ⊙ represents element-wise multiplication.

[0154] Stack two spatio-temporal graph convolution blocks. Thus, the calculation of the spatio-temporal graph convolution of the velocity flow is completed, and the final output is X out_v .

[0155] Step 4: Use the temporal multi-head cross-attention mechanism and the spatial multi-head cross-attention mechanism to achieve the feature fusion of the joint position flow and the velocity flow;

[0156] This step mainly focuses on the joint position flow X out_j and the velocity flow X out_v after being processed by the spatio-temporal graph convolution blocks in Steps 2 and 3, and captures the correlations in time and space; the specific process is as follows:

[0157] Step 4.1: The dynamic window-controlled temporal multi-head cross-attention mechanism;

[0158] In the temporal module, in order to focus on important time frames, first perform dimension reorganization on the input data X out_j and X out_v , focus on the time frame T, and ignore the dimension V of the number of joints, that is, integrate the dimension of the number of joints into the batch dimension. The batch dimension is how many samples the model processes at one time during training. Then, use the multi-head cross-attention mechanism to model the relationship between X out_j and X out_v . Specifically, generate a query (Q t ) from the input position flow through a linear layer, and generate a key (K t ) and a value (V t ) from the velocity flow. The specific formulas are shown in Equations (12) and (13) as follows:

[0159] Q t = Linear Q (X out_j ) (12)

[0160] K t ,V t = Linear KV (X out_v ) (13)

[0161] where, W t , K t and V tThe dimensions are first all expanded to heads * C, and then the high-dimensional channels are split into heads, and each head learns relevant information in its respective subspace.

[0162] Then, by calculating the dot product between the query (Q t ) and the transpose of the key (K t ), the similarity between them is calculated to obtain the attention score matrix, as specifically shown in Equation (14):

[0163]

[0164] where τ is a learnable temperature coefficient, which is a fixed scaling factor with an initial value of i.e., the dimension of the feature, which can be adaptively optimized and adjusted during the model training process, enabling the model to be trained more stably.

[0165] Introduce a time window mask matrix M ∈ {0, 1} T×T , defined as Equation (15):

[0166]

[0167] where M is the size of the time window. The mask is applied to the attention score matrix to generate locally constrained attention scores, as specifically shown in Equation (16):

[0168]

[0169] where A masked [t, t′] is the locally constrained attention score. The softmax function is used to convert the attention scores into a probability distribution, representing the attention weights of the query to each key. The locally constrained attention weights are applied to the value (V t ), as shown in Equation (17):

[0170] X Attn_heads = softmax(A masked [t, t′]) · V (17)

[0171] Finally, the subspaces processed by each head are concatenated to obtain the final output X Att_T of the dynamic window-controlled temporal multi-head cross-attention mechanism module, as shown in Equation (18):

[0172] X Att_T = Concat(X att_1 , X att_2 , …, X att_heads ) (18)

[0173] where X att_1 , X att_2,…,X att_heads is the number of heads of attention.

[0174] Step 4.2: Multimodal dynamic weighted spatial multi-head cross-attention;

[0175] In spatial cross-attention, to focus on important joints, similar to temporal multi-head cross-attention, first, the input data X out_j and X out_v are dimensionally reorganized, focusing on the number of joints dimension V and ignoring the time frame T, that is, integrating the time frame dimension into the batch dimension. Then, the multi-head cross-attention mechanism is used to model the relationship between X out_j and X out_v . Specifically, the input position stream is used to generate queries (Q v ), the velocity stream is used to generate keys (K v ), and values (V v ) through a linear layer, as shown in equations (19) and (20):

[0176] Q v = Linear Q (X out_j ) (19)

[0177] K v ,V v = Linear KV (X out_v ) (20)

[0178] where the dimensions of Q v , K v and V v are all expanded to heads*C. Next, the high-dimensional channels are split into heads, and each head learns relevant information in its own subspace.

[0179] Furthermore, the similarity between them is also calculated using the dot product to obtain the attention score matrix, and the specific formula is shown in equation (21):

[0180]

[0181] where C is the dimension of the feature, which can make the model training more stable during the model training process.

[0182] Next, to enhance the sparsity of the attention matrix and retain key connections, a sparse attention constraint is introduced, that is, the original attention scores are simultaneously normalized by the Softmax function and non-linearly screened by the Sigmoid function, and balanced with a coefficient of 0.1 for both, as shown in equation (22):

[0183] A enhanced= Softmax(A v ) + 0.1 * σ(A v ) (22)

[0184] Meanwhile, to adaptively adjust the importance of different attention heads, learnable weight parameters are introduced for each attention head, and after being normalized by the Softmax function, they are weighted to the attention matrix, as shown in Equation (23):

[0185] A weighted = A enhanced ⊙ Softmax(W heads ) (23)

[0186] That is the learnable head weight parameter matrix.

[0187] Then, the attention matrix is converted into a probability distribution through the Softmax function, and then it is applied to the values (V), as shown in Equation (24):

[0188] X Attn_heads = Softmax(A weighted ) · V v (24)

[0189] Furthermore, the subspaces processed by each head are concatenated to obtain the output of the total space. The total space is the joint feature space formed by concatenating the subspace features of all attention heads (Heads), as shown in Equation (21):

[0190] X Att = Concat(X att_1 , X att_2 , …, X att_heads ) (21)

[0191] where X att_1 , X att_2 , …, X att_heads is the number of heads of the heads attention.

[0192] Finally, to control the fusion ratio of the processed features and the original features, a gating mechanism is introduced to mix the new features and the original input features according to the gating ratio, as shown in Equations (25) and (26):

[0193] G = σ(W g * X Att_V + b g ) (25)

[0194]

[0195] Among them, a gating value ranging from 0 to 1 is generated through the Sigmoid function, where σ is the Sigmoid function, controlling the generated gating value within the range of 0 to 1, and b g is the bias term of the gated linear layer and is a learnable parameter. is the initial value of the multi-head cross-attention in the input space.

[0196] So far, the time multi-head cross-attention mechanism and the space multi-head cross-attention mechanism have completed the feature fusion module for the joint point flow and velocity flow.

[0197] Step 5: Use dual-path gating to achieve spatio-temporal feature fusion;

[0198] This step mainly fuses the two data streams X Att_T and X Att_V that have passed through the time and space multi-head cross-attention mechanisms. First, in order to enhance the spatial features before fusion, the data stream output by the space multi-head cross-attention mechanism is subjected to multi-scale convolution, as shown in Equations (27) and (28):

[0199] MSConv(X) = Conv 3×3 (X Att_V ) ⊕ Conv 5×5 (X Att_V ) (27)

[0200] X enhanved = X Att_V + W fuse ·MSConv(X) (28)

[0201] Among them, ⊕ is to splice the features after 3×3 and 5×5 convolutions, and W fuse is the weight of the convolutional layer.

[0202] Next, X ttt_T and X enhanced are fused through dual-path gating, namely channel-level gating and spatial-level gating. Before that, in order to process global information, first X Att_T and X enhanced are concatenated, as shown in Equation (29):

[0203] X fused = Concat(X Att_T , X enhanced ) (29)

[0204] Furthermore, in order to dynamically allocate the channel importance of time or spatial features, channel-level gating fusion is first performed on the two groups of data, as shown in Equations (30) and (31):

[0205] G C = σ(WC2 ·RELU(W C1 ·GAP(X fused ))) (30)

[0206] X channel =G C ·X Att_T +G C ·X enhanced (31)

[0207] where W C1 and W C2 are both channel gating convolution parameters, RELU is an activation function to alleviate the occurrence of overfitting, GAP is global average pooling, σ is the Sigmoid function to control the generated gating value within the range of 0 to 1, and G C is the generated gating value.

[0208] Then, the data stream after channel-level gating is fed into spatial-level gating processing, as shown in Eqs. (32) and (33):

[0209] G S =σ(W e ·X fused ) (32)

[0210] X CS_Fused =G S ⊙X channel +(1 - G S )⊙Mean(X fused ) (33)

[0211] where W S is the spatial gating convolution parameter, G S is the generated gating value, G S ⊙X cnannel focuses on important channels to strengthen local features, and Mean(X fused ) means averaging the concatenated temporal and spatial features along the channel dimension to provide context background information.

[0212] Next, to retain the information of the original spatio-temporal features, cross-residual connections are introduced, as shown in formula (34);

[0213]

[0214] where DWConv is depthwise separable convolution, which can maintain the model performance while reducing the computational cost and the number of parameters compared with ordinary convolution.

[0215] Thus, the dual-path gating spatio-temporal feature fusion module is completed.

[0216] Step 6: Feature extraction is performed using cascaded mainstream spatio-temporal graph convolutional blocks;

[0217] In this step, the data stream after double-gated fusion is subjected to spatio-temporal feature extraction again. It consists of 2 stacked spatio-temporal graph convolutional blocks. The structure of 1 spatio-temporal graph convolutional block is composed of 1 GCN block, 3 TCN blocks, and 1 spatio-temporal joint attention (ST-Att) module in series.

[0218] First, 1 spatio-temporal graph convolutional block GCN is used for graph convolution. The specific calculation formula is shown in Equation (35):

[0219]

[0220] where, v ti represents the i-th joint point in the t-th frame, N(v tj ) represents all the neighborhood nodes of v ti , Z ti is used to balance the contributions of different neighborhood nodes, d(v ti , v tj ) represents the distance between v ti and v tj , and represent the input and output feature maps respectively, and w() is the weight function. Different weights are assigned to nodes according to the different distances d(v ti , v tj ).

[0221] Connecting the output feature maps of each frame gives

[0222] After the action of 1 GCN block, the data stream is then processed through two cascaded spatio-temporal graph convolutional blocks. One spatio-temporal graph convolutional block consists of 1 GCN block, 2 TCN blocks, and 1 ST-Att connected in sequence; the calculation formula of 1 TCN module is shown in Equation (36):

[0223]

[0224] where, * represents convolution in the time dimension, and the size of the convolution kernel is i.e., sliding 5 frames in the time dimension without acting in the spatial dimension.

[0225] To extract the key frames and key joint points with the richest information in an action sequence, the data stream is input into the spatio-temporal joint attention (ST-Att) module. The input features are divided into two streams for separate processing. These two streams of data are exactly the same. One stream considers the time frames as a whole and pools the features of the joint points, i.e., The other stream processes the joint points as a whole and pools the features of the time frames, i.e., Then, splice them along the channel dimension to obtain After processing the spliced features through a fully connected layer (FC), separate the data stream along the channel dimension. One stream of data Performs feature extraction at the frame level, and the other stream of data Performs feature extraction at the joint level. Take the outer product of these two sets of data, and then apply the data after feature extraction to the original feature map, that is, multiply it element-wise with the original input feature map to obtain the key frames and key joints in the action sequence The specific calculation steps are shown in Equations (37) and (38):

[0226]

[0227] Among them, and respectively represent pooling joint features with the time frame as a whole and pooling time frame features with the joint as a whole. ⊕ represents the splicing operation, and σ is the activation function is the feature map after being processed by the fully connected layer and respectively represent feature extraction at the frame level and feature extraction at the joint level, ω is the activation function represents the outer product at the feature level, and ⊙ represents element-wise multiplication

[0228] So far, the feature extraction of the mainstream spatio-temporal graph convolution block is completed

[0229] Step 7: Complete action recognition;

[0230] Example 3

[0231] Furthermore, in Step 7, specifically: after completing the feature extraction of the mainstream spatio-temporal graph convolution block, normalize the processed data stream and send it to the global average pooling layer for spatio-temporal dimension averaging to focus on the overall action pattern, as shown in Equation (39):

[0232]

[0233] Among them, LN represents the normalization operation, and GAP represents the global average pooling

[0234] Finally, send X pool to the fully connected layer FC for the action classification task, as shown in Equation (40):

[0235]

[0236] Among them, the output dimension of the fully connected layer is determined by the total number of classes Classes of the classification task is the predicted value.

[0237] The action classification task is trained using cross - entropy loss as the loss function Loss, as shown in Equation (41):

[0238]

[0239] where y is the actual value, i.e., the label.

[0240] Finally, a predicted vector is output Classes represents the number of predicted action classes. Each component is the probability corresponding to each action class. The class represented by the maximum probability is the predicted action class. Thus, the action recognition task ends.

[0241] Example 4

[0242] For the human action recognition method based on the combination of spatio - temporal graph convolution and multi - head cross - attention of the present invention, GCN can represent the joint points and bone connections of the human skeleton. Through the feature propagation mechanism of GCN, it can explicitly model the physical topological connection relationship of the joint points in the human skeleton. Its hierarchical convolution operation captures the bone topological features in the spatial dimension. On this basis, a multi - head cross - attention mechanism is introduced to decouple and interactively fuse the features of the spatial coordinates (position stream) and motion speeds (velocity stream) of the joint points respectively, solving the problem that GCN is difficult to model the dynamic coupling of spatio - temporal features. At the same time, the multi - head cross - attention mechanism strengthens the action continuity in the time dimension through a local adjacent frame sliding window, and constrains the attention weights to be distributed only within the temporal neighborhood of adjacent frames; in the spatial dimension, it enhances the weights of important positions through sparse attention constraints. Therefore, the combination of GCN and the multi - head cross - attention mechanism can improve the human action representation ability of the deep network model, and thus improve the accuracy of action recognition.

[0243] Example 5

[0244] Hardware environment: The GPU is an NVIDIA vGPU (32GB video memory), and 12 CPUs are allocated to each GPU. The CPU model is Xeon(R) Platinum 8352V, with a main frequency of 2.10GHz

[0245] Software environment: PyTorch 2.1.2 deep learning framework.

[0246] Datasets: NTU - RGB+D, NTU - RGB+D120

[0247] The size of the time window W is set to 160 frames - 240 frames respectively, and experiments are carried out on the X-sub and X-view of the NTU-RGB+D dataset and the X-sub120 and X-set120 metrics of the NTU-RGB+D 120 dataset. The recognition accuracy results are shown in Table 1. The best experimental result on the X-sub metric of the NTU-RGB+D dataset is 90.81% accuracy at 180 frames. The best experimental result on the X-view metric of the NTU-RGB+D dataset is 94.72% accuracy at 220 frames. The best experimental result on the X-sub120 metric of the NTU-RGB+D120 dataset is 86.44% accuracy at 240 frames. The best experimental result on the X-set120 metric of the NTU-RGB+D 120 dataset is 87.79% accuracy at 200 frames.

[0248] Table 1 Comparison of the accuracy (%) of different window mask sizes

[0249]

[0250]

[0251] Example 6

[0252] On the basis of calculating the accuracy, the parameter quantities generated during the training of some single-stream and two-stream models on the NTU-RGB+D 120 dataset are also analyzed. As shown in Table 2, the parameter quantity of the MHCA-GCN of the present invention (0.20MB) is the least, which is 14.6 times less than the parameter quantity (2.92MB) of the model STC-Net (2 ensembles) with the highest accuracy of 89.3%.

[0253] Table 2 Comparison of the parameter quantities (MB) of MHCA-GCN and other models

[0254] Model Published Journal Number of Parameters (MB) ST-GCN AAAI2018 3.10 TE-GCN CVPR2020 5.80 SGN CVPR2020 0.69 MST-GCN AAAI2021 12.00 SF-GCN(J) ELSEVIER2021 3.50 Dynamic-GCN IEEE2022 14.40 EfficientGCN-B4(JV) IEEE2022 0.26 STC-Net(2 ensembles) ICCV2023 2.92 MHCA-GCN This invention 0.20

[0255] The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention of the present invention, on the basis of modeling the human skeleton action sequence with GCN, uses the multi-head cross-attention mechanism to efficiently fuse the position information and speed information of the action sequence, and improves the accuracy of action recognition.

Claims

1. A human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention, characterized in that, The implementation is specifically carried out according to the following steps: Step 1: Preprocess the data of the human action sequence to obtain the joint point position stream and velocity stream; Step 2: Extract the depth features of the human joint point position stream; Step 3: Extract the depth features of the human joint point velocity stream; Step 4: Use the temporal multi-head cross-attention mechanism and the spatial multi-head cross-attention mechanism to fuse the features of the joint point position stream processed in Step 2 and the velocity stream processed in Step 3; Step 5: Use dual-path gating to fuse the joint point position stream and velocity stream that have passed through the temporal and spatial multi-head cross-attention mechanisms; Step 6: Use a series of mainstream spatio-temporal graph convolutional blocks to extract features from the result obtained in Step 5; Step 7: Complete action recognition.

2. The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention according to claim 1, wherein In the above step 1, the human action sequence is represented as X, T represents the number of frames of the action sequence, V represents the number of human joint points, and C represents the number of features of each joint point; the specific preprocessing process is as follows: Step 1.1: Calculate the relative positions of the joint points, that is, the joint point position stream; Calculate the set of relative positions of the joint points \(Q = \{q i | i = 1, 2, \ldots, V in \}\), and the calculation formula of \(q i is shown in Equation (1); q i = x[:,:,i] - x[:,:,c] (1) Among them, c is the index of the central spine joint, and the position of the joint point is jointly represented by the original position X and the relative position Q; Step 1.2: Calculate the velocity of the joint points The speed information is divided into two groups. The fast speed is the time difference between adjacent frames, representing the short-term change of the action, that is, the fast action, which is represented by F = {f t | t = 1, 2, …, T in}; The slow speed is the time difference between the interval frames, representing the long-term change of the action, that is, the slow action, which is represented by S = {s t | i = 1, 2, …, T in}, as shown in Equation (2) and Equation (3); f t = x[:, t+1, :] - x[:, t, :] (2) S t = x[:, t+2, :] - x[:, t, :] (3) Among them, f t represents the difference between adjacent frames, and s t represents the difference between spaced frames.

3. The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention according to claim 2, wherein In the said Step 2, specifically: Step 2.1: Use a series of spatio-temporal graph convolutional blocks to extract features; Use 1 initial spatio-temporal graph convolutional block GCN for graph convolution, and the specific calculation formula is as shown in Equation (4): Among them, v ti represents the i-th joint point of the t-th frame, and N(v tj ) represents all the neighborhood nodes of v ti . Z ti is used to balance the contributions of different neighborhood nodes. d(v ti , v tj ) represents the distance between v ti and v tj . f in (·) and f out (·) represent the input and output feature maps respectively; w() is a weight function that assigns different weights to nodes according to different distances d(v ti , v tj ); Connecting the output feature maps of each frame gives Process the joint point position stream through two series-connected spatio-temporal graph convolutional blocks; one spatio-temporal graph convolutional block is composed of 1 GCN block, 2 TCN blocks and 1 ST-Att connected in sequence; the GCN module is the same as the initial GCN module, and the calculation formula of the TCN module is as shown in Equation (5): X STGCN = X GCN * W(5) Among them, * represents convolution in the time dimension, and the size of the convolution kernel is The input features are divided into two streams for separate processing. The data in these two streams is the same. One stream considers the time frames as a whole and pools the features of the joint points, that is The other stream processes the joint points as a whole and pools the features of the time frames, that is Then, they are concatenated along the channel dimension to obtain After the concatenated features are processed by the fully connected layer, the data stream is then separated along the channel dimension. One stream of data Performs feature extraction at the frame level, and the other stream of data Performs feature extraction at the joint point level. The outer product of these two sets of data is taken, and then the data after feature extraction is applied to the original feature map, that is, multiplied element-wise with the original input feature map, to obtain the key frames and key joint points in the action sequence The specific calculation steps are shown in Equations (6) and (7): X FC = σ((pool t (X STGCN )) ⊕ pool v (X STGCN )) · W) (6) Among them, pool t (X STGCN ) and pool v (X STGCN ) respectively represent pooling joint point features with the time frame as a whole and pooling time frame features with the joint point as a whole; ⊕ represents the concatenation operation, σ is the activation function, and X FC is the feature map after the fully connected layer processing; X FC ·W t and X FC ·W v respectively represent feature extraction at the frame level and feature extraction at the joint point level, ω is the activation function, represents the feature-level outer product, and ⊙ represents element-wise multiplication; Stack two spatio-temporal graph convolutional blocks. Thus, the spatio-temporal graph convolution calculation of the position flow is completed, and the final output is X out_j .

4. The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention as claimed in claim 3, wherein In the said Step 3, specifically: First use 1 initial GCN block for graph convolution, and the specific calculation formula is as shown in Equation (8): Among them, f′ in (·) and f′ out (·) represent the input and output feature maps respectively; Connecting the output feature maps of each frame gives Process the joint point velocity stream through two series-connected spatio-temporal graph convolutional blocks; one spatio-temporal graph convolutional block is composed of 1 GCN block, 2 TCN blocks and 1 ST-Att connected in sequence; the GCN module is the same as the initial GCN module, and the formula of the TCN module is as shown in Equation (9): X′ STGCN = X′ GCN * W(9) Among them, * represents convolution in the time dimension, and the size of the convolution kernel is The input features are divided into two streams for separate processing. These two streams of data are exactly the same. One stream considers the time frames as a whole and pools the features of the joint points, that is The other stream processes the joint points as a whole and pools the features of the time frames, that is Then they are concatenated along the channel dimension to obtain After the concatenated features are processed by the fully connected layer, the data stream is then separated along the channel dimension. One stream of data Extracts features at the frame level, and the other stream of data Extracts features at the joint point level. The outer product of these two sets of data is taken, and then the data after feature extraction is applied to the original feature map, that is, multiplied element by element with the original input feature map, and the key frames and key joint points in the action sequence can be obtained The specific calculation steps are shown in Equations (10) and (11): X′ FC = σ((pool t (X′ STGCN ) ⊕ pool v (X′ SvGCN )) · W) (10) Among them, pool t (X′ STGCN ) and pool v (X′ STGcN ) respectively represent pooling joint feature points with the time frame as a whole and pooling the features of the time frame with the joint points as a whole; ⊕ represents the concatenation operation, σ is the activation function, and X FC is the feature map after the fully connected layer processing; X' FC ·W t and X′ FC ·W v respectively represent feature extraction at the frame level and feature extraction at the joint point level; Stack two spatio-temporal graph convolution blocks. Thus, the calculation of the spatio-temporal graph convolution of the velocity flow is completed, and the final output is X out_v .

5. The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention according to claim 4, wherein In the said Step 4, specifically: Step 4.1: Dynamic window-controlled temporal multi-head cross-attention mechanism; First, perform dimensionality reorganization on the input data X out_j and X out_v Then, use the multi-head cross-attention mechanism to model the relationship between X out_j and X out_v Specifically, generate the query Q t by passing the input position stream through a linear layer, generate the key K t and the value V t , and the specific formulas are shown in Equations (12) and (13): Q t = Linear Q (X out_j ) (12) K t ,V t = Linear KV (X out_v ) (13) Compute the dot product between the query Q t and the transpose of the key K t to obtain the attention score matrix, as shown specifically in Equation (14): Among them, τ is a learnable temperature coefficient; Introduce the time window mask matrix M ∈ {0, 1} T×T , defined as Equation (15): Among them, W is the size of the time window, apply the mask to the attention score matrix to generate locally constrained attention scores, specifically as shown in Equation (16): Among them, A masked [t, t′] is the attention score of local constraint. The softmax function is used to convert the attention score into a probability distribution, representing the attention weight of the query to each key. The attention weight of local constraint is applied to the value V t as shown in Equation (17): X Attn_heads = softmax(A masked [t,t′])·V (17) Finally, the subspaces processed by each head are concatenated to obtain the final output X of the dynamic window-controlled temporal multi-head cross-attention mechanism module Att_T , as shown in Equation (18): X Att_T = Concat(X att_1 , X att_2 , …, X att_heads ) (18) Among them, X att_1 , X att_2 , …, X att_heads is the number of heads of heads attention; Step 4.2: Multimodal dynamic weighted spatial multi-head cross-attention; First, the input data X out_j and X out_v are dimensionally reorganized, and then the multi-head cross-attention mechanism is used to model the relationship between X out_j and X out_v . Specifically, the input position stream is used to generate the query Q v , the velocity stream is used to generate the key K v and the value V v , as shown in Equations (19) and (20) specifically: Q v = Linear Q (X out_j ) (19) K v ,V v = Linear KV (X out_v )(20) Calculate the similarity between them to obtain the attention score matrix, and the specific formula is as shown in Equation (21): Among them, C is the dimension of the feature; Introduce sparse attention constraints, that is, normalize the original attention scores through the Softmax function and perform non-linear screening through the Sigmoid function, and balance the two with a coefficient of 0.1, as shown in Equation (22): A enhanced = Softmax(A v ) + 0.1 * σ(A v ) (22) Introduce learnable weight parameters for each attention head, and then weight them to the attention matrix after normalization through the Softmax function, as shown in Equation (23): A weighted = A enhanced ⊙ Softmax(W heads ) (23) Then, the attention matrix is converted into a probability distribution through the Softmax function, and then applied to the value V v as shown in Equation (24): X Attn_heads = Softmax(A weighted )·V v (24) Stitch the subspaces processed by each head to obtain the output of the total space, as shown in Equation (21): X Att = Concat(X att_1 , X att_2 , …, X att_heads ) (21) Among them, X att_1 , X att_2 , …, X att_heads is the number of heads of heads attention; Introduce a gating mechanism to mix the new features and the original input features according to the gating ratio, as shown in Equations (25) and (26): G = σ(W g * X Att_V + b g ) (25) where σ is the Sigmoid function, b g is the bias term of the gated linear layer, is the initial value of the multi-head cross-attention in the input space.

6. The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention according to claim 5, characterized in that In step 5, specifically: Perform multi-scale convolution on the data stream output by the spatial multi-head cross-attention mechanism, as shown in equations (27) and (28): MSConv(X) = Conv 3×3 (X Att_V ) ⊕ Conv 5×5 (X Att_V ) (27) X enhanced = X Att_V + W fuse · MSConv(X) (28) Among them, ⊕ is used to splice the features after 3×3 and 5×5 convolutions, and W fuse is the weight of the convolutional layer; For X Att_T With X enhanced Perform dual-path gating fusion, namely channel-level gating and spatial-level gating. Before that, in order to process global information, first X Att_T And X enhanced Are concatenated as shown in Equation (29): X fused = Concat(X Att_T , X enhanced ) (29) Perform channel-level gated fusion on the two sets of data, as shown in equations (30) and (31): G C = σ(W C2 ·RELU(W C1 ·GAP(X fused ))) (30) X channel = G C · X Att_T + G C · X enhanced (31) Among them, W C1 and W C2 are both channel gating convolution parameters, RELU is the activation function, GAP is global average pooling, and G C is the generated gating value; Send the data stream after channel-level gating into spatial-level gating processing, as shown in equations (32) and (33): G S = σ(W S · X fused ) (32) X CS_Fused = G S ⊙X channel + (1 - G S ) ⊙ Mean(X fused ) (33) Among them, W S is the spatial gating convolution parameter, and G S is the generated gating value. G S ⊙X channel that is, focusing on important channels to strengthen local features. Mean(X fused ) means averaging the concatenated temporal and spatial features along the channel dimension; To retain the information of the original spatio-temporal features, introduce cross-residual connections, as shown in formula (34); 7. The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention according to claim 6, characterized in that In step 6, specifically: Process the data stream through two cascaded spatio-temporal graph convolution blocks; one spatio-temporal graph convolution block consists of 1 GCN block, 2 TCN blocks, and 1 ST-Att connected in sequence; 1 GCN block performs graph convolution, and the specific calculation formula is as shown in equation (35): Among them, and respectively represent the input and output feature maps; Connecting the output feature maps of each frame gives The calculation formula of 1 TCN block is as shown in equation (36): The input features are divided into two streams for separate processing. These two streams of data are exactly the same. One stream considers the time frames as a whole and pools the features of the joint points, that is The other stream processes the joint points as a whole and pools the features of the time frames, that is Then they are concatenated along the channel dimension to obtain After the concatenated features are processed by the fully connected layer, the data stream is then separated along the channel dimension. One stream of data Extracts features at the frame level, and the other stream of data Extracts features at the joint point level. The outer product of these two sets of data is taken, and then the data after feature extraction is applied to the original feature map, that is, multiplied element-wise with the original input feature map, to obtain the key frames and key joint points in the action sequence The specific calculation steps are shown in Equations (37) and (38): Among them, and respectively represent pooling joint point features with the time frame as a whole and pooling time frame features with the joint point as a whole; is the feature map after the fully connected layer processing; and respectively represent feature extraction at the frame level and feature extraction at the joint point level.

8. The human action recognition method based on the combination of spatio-temporal graph convolution and multi-head cross-attention as claimed in claim 7, wherein In step 7, specifically: The processed data stream is normalized and then fed into the global average pooling layer, as shown in Equation (39): Among them, LN represents the normalization operation, and GAP represents global average pooling; Send X pool to the fully connected layer FC for action classification tasks, as shown in Equation (40): Among them, the output dimension of the fully connected layer is determined by the total number of classes Classes of the classification task; is the predicted value; For the action classification task, use cross-entropy loss as the loss function Loss for training, as shown in equation (41): Among them, y is the actual value, that is, the label; Finally output a prediction vector Classes represents the number of predicted action classes. Each component is the probability corresponding to each action class. The class represented by the maximum probability is the predicted action class.

Citation Information

Cited By

  • Hierarchical dynamic fusion focus recognition system based on multi-source data driving

    CN122156885A