Transform fusion-based multi-stream space-time diagram convolution behavior identification method and system
By building a multi-stream spatiotemporal graph convolution behavior recognition network based on Transformer fusion, the problem of insufficient constraints and fusion strategies of spatiotemporal graph convolution unit in the existing technology is solved, the accuracy and computing efficiency of behavior recognition are improved, and different information characteristics are adapted.
Patent Information
- Application Number
- CN202510280842.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
The domain constraints and temporal convolutional receptive field of space-time graph convolution units in existing graph convolution behavior recognition algorithms based on skeleton sequences are limited, and the late fusion strategy for multi-stream behavior recognition is poor, resulting in the recognition accuracy and calculation amount that cannot meet the needs.
A multi-stream spatiotemporal graph convolution behavior recognition network based on Transformer fusion is constructed, and the long-distance dependence between nodes is enhanced through spatiotemporal graph convolution unit and spatiotemporal self-attention unit, skeleton information is extracted in combination with a lightweight backbone network, and a full connection layer and softmax output behavior category scores are used to realize feature fusion and behavior recognition.
It improves the recognition accuracy of long-term actions, reduces the calculation amount, improves the model's adaptability to different information characteristics, and improves the output results of behavior recognition.
Smart Images

Figure CN120260115A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of behavior recognition, and particularly relates to a multi-stream spatio-temporal graph convolutional behavior recognition method based on Transformer fusion. Background Art
[0002] With the improvement of computer performance in recent years, the field of computer vision has developed rapidly, and behavior recognition, as a popular field in computer vision, has also attracted the attention of many researchers. Behavior recognition plays an important role in the fields of national defense, autonomous driving, and video surveillance. Behavior recognition algorithms can, with their efficient and accurate recognition capabilities, promptly recognize relevant dangerous actions, thereby timely preventing the occurrence of dangerous behaviors. Traditional methods for recognizing human behaviors using RGB videos have large computational requirements while failing to achieve the expected recognition accuracy.
[0003] The behavior recognition method based on skeleton sequences is very suitable for human behavior recognition due to its computational resource savings and low complexity. Among them, the human behavior recognition algorithm represented by the spatio-temporal graph convolutional network (ST-GCN) has higher recognition accuracy and smaller computational requirements, making it very suitable for the field of human behavior recognition. ST-GCN is composed of spatial graph convolution (GCN) and temporal convolution (TCN). However, there are problems in the current graph convolutional behavior recognition algorithm based on skeleton sequences, such as the domain constraint of the spatio-temporal graph convolutional unit and the limited receptive field of temporal convolution. Therefore, there is an urgent need to design a new graph convolutional unit to solve this problem. In addition, the late fusion strategy of multi-stream behavior recognition has poor practicality. Therefore, there is an urgent need to design a new fusion strategy to enable the model to better adapt to the characteristics of each type of information.
[0004] Therefore, there is an urgent need to propose a multi-stream spatio-temporal graph convolutional behavior recognition method based on Transformer fusion to solve the above problems. Summary of the Invention
[0005] Aiming at the defects existing in the above-mentioned prior art, the purpose of the present invention is to provide a multi-stream spatio-temporal graph convolution behavior recognition method based on Transformer fusion. The present invention constructs a multi-stream spatio-temporal graph convolution behavior recognition network based on Transformer fusion. The first 5 layers of the network are composed of spatio-temporal graph convolution units, and the last 3 layers of the network are constructed by spatio-temporal self-attention units. Finally, the behavior category scores are output through a fully connected layer and softmax. The present invention enhances the modeling of long-range dependencies between all nodes in the skeleton sequence, reduces the dependence of the behavior recognition algorithm on the accuracy of joint point positioning, and at the same time enables the model to better adapt to the characteristics of each type of information, improves the output results of the existing behavior recognition, and greatly reduces the computational complexity compared with the prior art.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A multi-stream spatio-temporal graph convolution behavior recognition method based on Transformer fusion, comprising:
[0008] S1. Use a lightweight backbone network to extract the skeleton information of the human body graph to be recognized, and obtain bone information and joint information;
[0009] S2. Obtain a joint skeleton sequence based on the joint information;
[0010] S3. Obtain a bone skeleton sequence based on the bone information;
[0011] S4. Based on the joint skeleton sequence and the bone skeleton sequence, obtain a motion skeleton sequence of the human body motion information;
[0012] S5. Input the motion skeleton sequence into the behavior recognition network for feature extraction, and stack the features extracted by each branch of the behavior recognition network according to the channel dimension to achieve feature fusion, and obtain a fused feature map;
[0013] S6. Input the fused feature map in step S5 into the spatio-temporal self-attention unit for global feature learning;
[0014] S7. Input the learned fused feature map into a fully connected layer and softmax to calculate the behavior category scores, and perform human behavior recognition output based on the score values.
[0015] Preferably, in step S1, the process of using a lightweight backbone network to extract the skeleton information of the human body graph to be recognized includes:
[0016] S101. Change the number of channels of the human body graph to be recognized to 32 channels, and perform a convolution operation on the changed human body graph to be recognized to obtain a first branch input feature map and a second branch input feature map;
[0017] S102. Use two branches with different resolution characteristics to extract the features of the first-branch input feature map and the second-branch input feature map multiple times respectively, obtaining the first-branch multi-scale feature map and the second-branch multi-scale feature map;
[0018] S103. Use the multi-scale feature fusion module to redefine and fuse the sizes of the first-branch multi-scale feature map and the second-branch multi-scale feature map respectively, obtaining the first-branch fused feature map and the second-branch fused feature map;
[0019] S104. Input the first-branch fused feature map and the second-branch fused feature map into two identity mapping layers respectively, obtaining the first-branch mapped feature map and the second-branch mapped feature map. At the same time, perform a convolution operation on the second-branch fused feature map to obtain a high-dimensional convolution map;
[0020] S105. Use 4 cascaded basic feature extraction units to extract and fuse the first-branch mapped feature map, the second-branch mapped feature map, and the high-dimensional convolution map respectively, obtaining the first-branch feature map, the second-branch feature map, and the third-branch feature map;
[0021] S106. Input the first-branch feature map, the second-branch feature map, and the third-branch feature map into 3 identity mapping layers respectively. At the same time, input the third-branch feature map into a 3×3 convolution layer with an output channel number of 192 and a stride of 2, obtaining the first-branch mapped map, the second-branch mapped map, the third-branch mapped map, and the fourth-branch high-dimensional feature map;
[0022] S107. Input the first-branch mapped map, the second-branch mapped map, the third-branch mapped map, and the fourth-branch high-dimensional feature map into 3 cascaded basic feature extraction units to obtain the bone information and joint information.
[0023] Preferably, in step S102, the two branches with different resolution characteristics use the lightweight depth convolution transformation module to extract the features of the first-branch input feature map and the second-branch input feature map respectively, and the lightweight depth convolution transformation module includes a feature extraction layer; wherein,
[0024] The extraction process of the feature extraction layer includes:
[0025] S1021. Use layer normalization to eliminate the differences between the first-branch input feature map and the second-branch input feature map;
[0026] S1022. Use depthwise separable convolution with a stride of 3×3 to extract the context information of the pixel local regions of the first-branch input feature map and the second-branch input feature map after normalization in step S1021, obtaining the first-branch input feature map after feature extraction and the second-branch input feature map after feature extraction.
[0027] Preferably, in step S103, the multi-scale feature fusion module includes a feature size redefinition layer and a feature information fusion layer; wherein,
[0028] The feature size redefinition layer redefines the size through upsampling and downsampling operations, and superimposes them, and uses the superimposed feature map as the input of the feature information fusion layer;
[0029] The feature information fusion layer sequentially passes the feature map after superimposing the two branches output by the feature size redefinition layer through 1×1 convolution, 3×3 depthwise separable convolution, and cascaded channel attention to obtain the first branch fusion feature map and the second branch fusion feature map.
[0030] Preferably, in step S105, the basic feature extraction unit includes a depth convolution transformation module and a feature fusion module;
[0031] Among them, the depth convolution transformation module is used to extract the deep features of the first branch mapping feature map, the second branch mapping feature map, and the high-dimensional convolution map respectively, and input the extracted deep features into the feature fusion module for multi-scale feature fusion.
[0032] Preferably,
[0033] The expression of the joint skeleton sequence is:
[0034]
[0035] In the formula, T is the length of the skeleton sequence, V is the number of nodes in a single-frame skeleton, x tv、 y tv represents the specific position of the joint in space, and s tv represents the confidence of the key point detection;
[0036] The expression of the bone skeleton sequence is:
[0037]
[0038] In the formula, i represents different key points in the same frame, j represents the same key point in different frames, and Es represents the skeleton connection relationship of each frame;
[0039] The expression of the motion skeleton sequence is:
[0040]
[0041] Preferably, in step S5, the behavior recognition network includes 4 branches, and each branch uses a spatio-temporal graph convolutional network to extract features from the motion skeleton sequence. The first five layers of the spatio-temporal graph convolutional network are spatio-temporal graph convolutional units, and the spatio-temporal graph convolutional unit includes a spatial graph convolutional unit and a temporal convolutional unit.
[0042] Preferably, in step S6, the process of global feature learning includes:
[0043] S601. Perform positional encoding on the skeleton sequence nodes in the fused feature map of step S5;
[0044] S602. Map the encoded skeleton sequence nodes to the query matrix key matrix and value matrix
[0045] S603. Based on the mapped query matrix key matrix and value matrix of the skeleton sequence nodes, calculate the self-attention values of the current node in the skeleton sequence with other nodes, and based on the self-attention values, combined with the parameter matrix W O perform a linear transformation to obtain the learned fused feature map.
[0046] Preferably, in step S602, the mapping formula for mapping the encoded skeleton sequence nodes to the query matrix key matrix and value matrix is:
[0047] Q = G × W q 、K = G × W k 、V2 = G × W v (11)
[0048] where is the skeleton sequence, and w is the learned weight matrix;
[0049] In step S603, the calculation formula for the self-attention value of the current node with other nodes is:
[0050]
[0051] where K T is the transpose of the key matrix K.
[0052] The second object of the present invention is to provide a multi-stream spatio-temporal graph convolutional behavior recognition system based on Transformer fusion, including:
[0053] A skeleton information and joint information acquisition module, configured to use a lightweight backbone network to extract the skeleton information of the human body graph to be recognized, and obtain the skeleton information and joint information;
[0054] A joint skeleton sequence acquisition module, configured to obtain a joint skeleton sequence based on the joint information;
[0055] The bone skeleton column sequence acquisition module is used to obtain the bone skeleton column sequence based on the bone information;
[0056] The motion skeleton sequence acquisition module is used to obtain the motion skeleton sequence of the human body motion information based on the joint skeleton sequence and the bone skeleton column sequence;
[0057] The fused feature map acquisition module is used to input the motion skeleton sequence into the behavior recognition network for feature extraction, and stack the features extracted by each branch of the behavior recognition network according to the channel dimension to achieve feature fusion, so as to obtain the fused feature map;
[0058] The learning module is used to input the fused feature map of the fused feature map acquisition module in the previous step into the spatio-temporal self-attention unit for global feature learning;
[0059] The output module is used to input the learned fused feature map into the fully connected layer and softmax to calculate the behavior category score, and perform human behavior recognition output based on the score value.
[0060] The beneficial effect of the present invention is that the present invention discloses a multi-stream spatio-temporal graph convolutional behavior recognition method based on Transformer fusion. Compared with the prior art, the improvement of the present invention lies in:
[0061] (1) The present invention designs a spatio-temporal self-attention unit for modeling the long-range dependencies between all nodes of the skeleton sequence, and combines components such as multi-head attention, feed-forward neural network, and residual connection included in the encoder of Transformer, so that the network can globally capture the complex relationships between nodes, and can improve the recognition accuracy of the network for long-term actions such as "running" and "clapping".
[0062] (2) The present invention adopts a mid-term fusion strategy to realize feature fusion between joints, bones, and motion information in the skeleton sequence, thereby constructing a new multi-stream behavior recognition network, avoiding the doubled workload of the late fusion strategy, and improving the practicability of the multi-stream behavior recognition method. Description of the Drawings
[0063] Figure 1 It is the overall flowchart of the multi-stream spatio-temporal graph convolutional behavior recognition method based on Transformer fusion of the present invention;
[0064] Figure 2 It is the block diagram of the multi-stream spatio-temporal graph convolutional behavior recognition method based on Transformer fusion of the present invention;
[0065] Figure 3 It is the visualization result diagram of the multi-stream spatio-temporal graph convolutional behavior recognition method based on Transformer fusion of the present invention. Detailed Embodiments
[0066] To enable those of ordinary skill in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0067] Embodiment 1:
[0068] Referring to Figures 1-3 A multi-stream spatio-temporal graph convolutional behavior recognition method based on Transformer fusion as shown, comprising:
[0069] S1. Use a lightweight backbone network to extract the skeleton information of the human body graph to be recognized, obtaining joint information and bone information;
[0070] Specifically, the skeleton information includes bone information and joint information. Joint information is one of the basic elements in the skeleton sequence, providing a static description of the human body posture. The present invention uses a self-constructed lightweight backbone network to extract the skeleton information. The specific method is as follows;
[0071] S101. Change the number of channels of the human body graph to be recognized to 32 channels, and perform convolution operations on the image with the changed channels to obtain a first-branch input feature map and a second-branch input feature map;
[0072] Specifically, use two 3×3 convolutions with a stride of 2 to downsample the resolution of the human body graph to be recognized to 1 / 4, and use 4 bottleneck modules to extract the shallow features of the human body graph to be recognized, changing the number of channels of the human body graph to be recognized from 3 channels to 32 channels;
[0073] Then perform convolution operations on the human body graph to be recognized with the changed channels to obtain a first-branch input feature map and a second-branch input feature map; Specifically, pass the human body graph with the changed channels through two 3×3 convolutions with the number of output channels being 24 and 48 respectively and the strides being 1 and 2 respectively. The former is used to generate the first-branch input feature map, and the latter is used to generate the second-branch input feature map.
[0074] S102. Use two branches with features of different resolutions to extract the features of the first-branch input feature map and the second-branch input feature map respectively multiple times to obtain a first-branch multi-scale feature map and a second-branch multi-scale feature map;
[0075] Specifically, the output images of step S101 are respectively input into two branches with characteristics of resolutions of 640×640 and 320×320 to extract features. Lightweight depth convolution transformation modules are set on the two branches. The input feature maps of the first branch and the input feature maps of the second branch are respectively input into the lightweight depth convolution transformation modules on the corresponding two branches. The output features of the lightweight depth convolution transformation modules have the same dimension as the input features. The input feature maps of the first branch and the input feature maps of the second branch undergo information extraction by the lightweight depth convolution transformation modules twice to output multi-scale feature maps;
[0076] The lightweight depth convolution transformation module consists of a feature extraction layer and a channel mapping layer with a residual structure. The specific extraction process is as follows:
[0077] The extraction process of the feature extraction layer includes:
[0078] S1021. Layer normalization is used to eliminate the differences between the input feature maps of the first branch and the input feature maps of the second branch. Layer normalization helps to ensure that the distribution of the input data is relatively consistent among different samples and provides a more stable input for subsequent processing;
[0079] S1022. Depthwise separable convolutions with a stride of 3×3 are used to extract the context information of the pixel local regions of the input feature maps of the first branch and the input feature maps of the second branch after normalization in step S1021, obtaining the input feature maps of the first branch after feature extraction and the input feature maps of the second branch after feature extraction;
[0080] The mapping process of the channel mapping layer is as follows:
[0081] The channel mapping layer adopts a residual structure to alleviate the problems of gradient vanishing and explosion during the training process. The input of the residual structure is the feature output by step S1022. The specific mapping process of the channel mapping layer is as follows: The input feature maps of the first branch and the second branch after feature extraction are respectively input into the corresponding channel mapping layers. The channel mapping layer structurally includes two 1×1 convolution operations, aiming to enhance the non-linear expression ability of the model. The role of the first 1×1 convolution is to expand the number of channels of the input feature map of the first branch after feature extraction by 2 times. Through this operation, the model can learn the information in the input features more fully, increasing the understanding of rich feature representations. Next, the second 1×1 convolution compresses the number of channels of the input feature map of the second branch after feature extraction to the original number of channels, aiming to retain important feature information through non-linear transformation while reducing the computational burden of the model. When outputting the feature map of the channel mapping layer, the channel mapping of the residual structure is adopted to reduce the risk of gradient explosion during the training process of the model. Finally, the feature map output by the channel mapping layer is the multi-scale feature map output by the lightweight depth convolution transformation module, that is, the multi-scale feature map of the first branch and the multi-scale feature map of the second branch are obtained.
[0082] S103. Use the multi-scale feature fusion module to redefine and fuse the sizes of the multi-scale feature maps of the first branch and the second branch respectively, to obtain the fused feature maps of the two branches, that is, the fused feature map of the first branch and the fused feature map of the second branch;
[0083] Specifically, the multi-scale feature fusion module is mainly composed of a feature size redefinition layer and a feature information fusion layer; among them, the feature size redefinition layer redefines the sizes of the multi-scale feature maps from different branches through a 1×1 convolution under the condition of keeping the number of channels unchanged by means of upsampling and downsampling operations, and stacks the feature maps with unified sizes in the channel dimension, and takes the stacked feature map as the input of the feature information fusion layer;
[0084] The feature information fusion layer consists of three key components, namely 1×1 convolution, 3×3 depthwise separable convolution, and cascaded channel attention. The stacked feature map output by the feature size redefinition layer is successively passed through these three operations: 1×1 convolution, 3×3 depthwise separable convolution, and cascaded channel attention, to obtain the fused feature maps of the two branches respectively; the design of the feature information fusion layer aims to fuse the feature information from different branches through these three operations to improve the fusion accuracy, and the output feature of the multi-scale feature fusion module has the same scale as the input feature of the corresponding branch of the multi-scale feature fusion module and contains the input multi-scale information, thus improving the performance of the network.
[0085] S104. Input the first-branch fusion feature map and the second-branch fusion feature map of step S103 into two identity mapping layers respectively to obtain a first-branch mapped feature map and a second-branch mapped feature map. At the same time, perform a convolution operation on the second-branch fusion feature map to obtain a high-dimensional convolution map;
[0086] Specifically, input the first-branch fusion feature map and the second-branch fusion feature map of step S103 into two identity mapping layers respectively. The input and output of the identity mapping layer have a completely mapped relationship to obtain a first-branch mapped feature map and a second-branch mapped feature map. At the same time, the second-branch fusion feature map is input into a 3×3 convolution layer with an output channel number of 96 and a stride of 2, thereby outputting a high-dimensional convolution map. The high-dimensional convolution map is a feature map with a higher dimension and a lower resolution;
[0087] S105. Use 4 cascaded basic feature extraction units to extract and fuse the first-branch mapped feature map, the second-branch mapped feature map, and the high-dimensional convolution map respectively to obtain a first-branch feature map, a second-branch feature map, and a third-branch feature map;
[0088] Specifically, in order to deepen the network depth and improve the model performance, four cascaded basic feature extraction units are used to sequentially extract features from the three feature images in step 104. The basic feature extraction unit consists of two depth convolution transformation modules and one feature fusion module. The specific extraction process is as follows: First, the first-branch mapped feature map, the second-branch mapped feature map, and the high-dimensional convolution map are respectively input into the first basic feature extraction unit. Deep features are extracted through the convolution module. Specifically, first, layer normalization is used to normalize by solving the variance of each feature map to eliminate the differences between input samples. Layer normalization helps ensure that the input data has a relatively consistent distribution among different samples, providing a more stable input for subsequent processing. Second, a convolution with a stride of 3×3 is used to extract context information. Then, the extracted deep features are input into the feature fusion module for multi-scale feature fusion. First, these features are stacked in the channel dimension. Second, the stacked features are operated on through a 1×1 convolution to compress their number of channels to be equal to the number of feature channels of the current branch. Finally, through a 3×3 depthwise separable convolution, the spatial information of the features of each branch is fused. So that the features of each branch have the feature information of other sizes, and the feature maps of the first feature extraction of the three branches are obtained respectively. Then, the feature maps of the three branches of the first feature extraction are input into the second basic feature extraction unit to obtain the feature maps of the second feature extraction of the three branches respectively, and so on. Finally, the fourth feature maps of the three branches are obtained for input to the input map of step S106. The first-branch mapped feature map corresponds to the first-branch feature map, the second-branch mapped feature map corresponds to the second-branch feature map, and the high-dimensional convolution map corresponds to the third-branch mapped feature map.
[0089] S106. Input the first-branch feature map, the second-branch feature map, and the third-branch feature map into 3 identity mapping layers respectively, and at the same time input the third-branch feature map into a convolution layer to obtain the first-branch mapped map, the second-branch mapped map, the third-branch mapped map, and the fourth-branch high-dimensional feature map respectively;
[0090] Input the first-branch feature map, the second-branch feature map, and the third-branch feature map of step S105 into 3 identity mapping layers respectively, and then output the corresponding first-branch mapped map, second-branch mapped map, and third-branch mapped map. At the same time, input the third-branch feature map into a 3×3 convolution layer with an output channel number of 192 and a stride of 2 to obtain the fourth-branch high-dimensional feature map;
[0091] S107. Input the first-branch mapped map, the second-branch mapped map, the third-branch mapped map, and the fourth-branch high-dimensional feature map into 3 cascaded basic feature extraction units to obtain bone information and joint information;
[0092] Specifically, first, the feature maps of the above four branches are input into 3 serially connected basic feature extraction units, and each basic feature extraction unit is composed of 2 depth convolution transformation modules and 1 feature fusion module. The specific extraction process is as follows: First, the first branch mapping feature map, the second branch mapping feature map, the third branch mapping feature map, and the high-dimensional convolution map are respectively input into the first basic feature extraction unit. Deep features are extracted through the convolution module, and then the extracted deep features are input into the feature fusion module for multi-scale feature fusion, so that the features of each branch have feature information of other sizes, and the feature maps of the first feature extraction of the four branches are obtained respectively. Then, the feature maps of the four branches of the first feature extraction are input into the second basic feature extraction unit to obtain the feature maps of the second feature extraction of the four branches respectively, and so on, until the fourth feature maps of the three branches are finally obtained. Among them, the last basic unit outputs a high-resolution feature map of the first branch after fusing the features of other branches, and uses a 1×1 convolution with the number of output channels equal to the number of key points to calculate the probability that each pixel point in the feature map is a key point, that is, the key point heat map. The point with the highest probability in the heat map is the human joint information, that is, the joint information such as the nose, eyes, shoulders, etc. Then, according to the recognized human joint information, the connection method is defined according to the human body structure, such as the left shoulder connecting to the left elbow and the left elbow connecting to the left wrist. Finally, according to the defined connection method, each joint is used as a node, and the connection between joints is used as an edge to connect, constructing a graph structure, which is the bone information.
[0093] S2. Obtain the joint skeleton sequence based on the joint information;
[0094] The joint information includes the accurate position coordinates and corresponding scores of each joint. These position coordinates (x, y) reflect the specific position of the joint in space, and the score s represents the confidence of the key point detection, that is, the credibility of the human pose estimation algorithm for each joint. Therefore, the joint skeleton sequence based on the joint information can be characterized as follows:
[0095]
[0096] In the formula, T is the skeleton sequence length, V is the number of nodes in a single-frame skeleton, x tv、 y tv represents the specific position of the joint in space, and the score s tv represents the confidence of the key point detection.
[0097] S3. Obtain the bone skeleton sequence based on the bone information;
[0098] Skeletal information involves the connection relationships between joints and the structure of the entire skeleton. This information captures the morphological characteristics of the human body by describing the connection methods of the bone chains and the bone lengths. The connection relationships between joints specify the relationships between each joint and its adjacent joints, forming the bone chains of the human body. At the same time, the bone lengths provide information on relative scales, which helps to maintain a consistent pose representation in different scenarios. In the COCO dataset format, the connection relationship of the skeletal information is E S ={(0,0),(1,0),(2,0),(3,1),(4,2),(5,0),(6,0),(7,5),(8,6),(9,7),(10,8),(11,0),(12,0),(13,11),(14,12),(15,13),(16,14)}. The bone skeleton column order based on the skeletal information can be characterized as follows:
[0099]
[0100] where i represents different key points in the same frame, j represents the same key point in different frames, and Es represents the skeleton connection relationship of each frame.
[0101] S4. Based on the joint skeleton sequence and the bone skeleton column order, obtain the motion skeleton sequence of the human motion information;
[0102] Human motion information involves the changes of joints or bones in the time dimension, including motion trajectories, speeds, and accelerations, etc. The motion trajectory describes the specific motion paths of joints or bones over a period of time, while the speed and acceleration provide the dynamic characteristics of these motions. This type of information is a key capture of the temporality and dynamics of human actions. Through motion information, the evolution process of human actions can be understood more meticulously, so as to express and recognize different behaviors more comprehensively. In the specific implementation process, the differences between joints or bones in the front and back frames are used to describe the motion information. Then, the motion skeleton sequence based on the motion information can be characterized as follows:
[0103]
[0104] where
[0105] S5. Use the behavior recognition network to extract features from the motion skeleton sequence obtained in step S4, and stack the channel dimensions of the extracted feature maps to achieve feature fusion, obtaining a fused feature map;
[0106] Specifically, the behavior recognition network has 4 branches. The skeleton sequence is input into each branch. Each branch uses a spatio-temporal graph convolutional network to extract features from the motion skeleton sequence. The first five layers of the spatio-temporal graph convolutional network are spatio-temporal graph convolutional units. This unit adopts a multi-stream network structure and a mid-term fusion strategy to fuse the extracted data.
[0107] The spatio-temporal graph convolutional unit consists of a spatial graph convolutional unit and a temporal convolutional unit, which are responsible for processing the node features of a single frame of the skeleton and paying attention to the changes in the node features between frames respectively. The complex relationships between joints in the motion skeleton sequence are captured through the spatial graph convolution (GCN), while the temporal graph convolution (TCN) emphasizes the evolution of actions in the time dimension. Overall, the spatio-temporal graph convolutional unit is a powerful tool for processing the skeleton sequence, providing an efficient means for spatio-temporal feature extraction and modeling of graph convolutional behavior recognition.
[0108] Specifically, the essence of the spatial graph convolutional unit is to update its own information by aggregating the features of other nodes in the neighborhood of the source node. And the graph convolutional unit can process data with irregular structures using the adjacency matrix. The process of the graph convolutional unit extracting features from the skeleton sequence can be summarized by the following formula:
[0109]
[0110] Among them,
[0111]
[0112] f in and conv_s(f in ) represent the input and output features of the graph convolutional unit respectively. A k is the symmetric Laplacian matrix, W k represents the learnable weight, K s represents the size of the graph convolutional kernel in the spatial dimension. According to the spatial allocation strategy in ST-GCN, K s = 3, is the adjacency matrix of the spatial graph of the skeleton sequence, used to characterize the connection relationship between nodes in the spatial domain. I represents the identity matrix, and D k represents the degree matrix, and the values on its diagonal are the sum of the elements in each row of the adjacency matrix, and the elements in the remaining positions are 0.
[0113] Among them, the purpose of introducing I into the adjacency matrix in formula (4) is that when transmitting information, the information of the source node itself also needs to be considered. And using D k to normalize the adjacency matrix can prevent the model from focusing on some nodes with more "edges".
[0114] The process of obtaining the fused feature map using the behavior recognition network includes:
[0115] S501. First, use the spatial graph convolutional unit to extract the feature map:
[0116] The spatial graph convolutional unit extracts features from the motion skeleton sequence to obtain the node features of a single-frame skeleton. Specifically, the spatial graph convolutional unit first uses a spatial graph convolutional layer conv_s with a kernel size of 1×1 and an output channel number multiplied by K s to perform parametric learning on each node of the motion skeleton sequence, and inputs the learned samples into the bn layer for normalization to reduce the difference in feature distribution. Then, the adjacency matrix A is used to perform matrix multiplication with the normalized features to transfer information between nodes. Among them, the relu function is used to non-linearly activate the normalized features so that the model can fit more complex sample relationships. Finally, the output single-frame skeleton node features are connected through the residual structure res to reduce the risk of gradient explosion during training. The entire process of the spatial graph convolutional unit can be expressed as follows:
[0117] f s-out = relu(bn(conv_s(f in )) × A) + res(f in ) (7)
[0118] where conv_s(f in ) is the output feature of the spatial graph convolutional layer conv_s, A is the adjacency matrix, and res(f in ) is the output of the residual structure res;
[0119] S502. Then, use the output of the spatial graph convolutional unit as the input of the temporal convolutional unit to further extract features. The process is as follows:
[0120] Use the temporal convolutional unit to extract features from the features output by the spatial graph convolutional unit. The temporal convolutional unit uses a temporal convolutional layer conv_t with a kernel size of 9×1, a stride of 2×1 or 1×1 and an unchanged number of channels to learn the features between nodes in the time domain, inputs the learned samples into the bn layer for normalization, and its implementation principle is similar to traditional convolutional operations. Then, the normalized features are input into the dropout layer to reduce the risk of model overfitting, and finally, features with time and space dimensions are output. The entire process can be expressed as follows:
[0121] f t-out = dropout(bn(conv_t(f s-out ))) (8)
[0122] S503. Connect the input of the spatial graph convolution and the output of the temporal graph convolution using the residual structure to obtain the fused feature map;
[0123] The outermost layer of the spatio-temporal graph convolution unit uses a residual structure for connection, adding the input of the spatial graph convolution and the output of the temporal graph convolution. The purpose is to reduce the risk of gradient explosion during the training process through the residual connection output, further improving the training process of the deep neural network and enhancing the training effect and performance of the network. The operation process formula of the spatio-temporal graph convolution unit is as follows:
[0124] f out = relu(f t-out + res(f in )) (9)
[0125] S6. Input the fused feature map obtained in step S5 into the spatio-temporal self-attention unit for global feature learning to obtain the learned fused feature map; specifically including:
[0126] S601. Perform position encoding on the skeleton sequence nodes in the fused feature map obtained in step S5;
[0127] For the skeleton sequence in the fused feature map, position encoding is used to provide the relative position information of the nodes in the skeleton sequence, which enables the model to better consider the spatio-temporal relationship between key points during the self-attention operation, helping to capture the overall coordination and evolution process of the action. For each fused feature key point position pos and each dimension i in the skeleton sequence, the calculation of the position encoding is as follows:
[0128]
[0129] where dim represents the dimension of the node embedding, and the skeleton sequence can be expressed as T is the length of the skeleton sequence, V is the number of nodes in the skeleton, C represents the specific dimension of the embedding, is the set of real numbers, indicating that the elements of this tensor are real values, which is suitable for representing the features of each joint in each time step of the skeleton sequence.
[0130] We need to calculate the dependency relationship between all nodes in the skeleton sequence of the fused feature map, so through operations such as dimension swapping and spreading, the skeleton sequence is changed to where L = T * V1 represents the total number of all nodes in the skeleton sequence, V1 is the number of nodes in a single-frame skeleton, then pos = (1,..., L) and dim = C in the position encoding.
[0131] S602. Map the encoded skeleton sequence node embedding vectors to the query matrix key matrix and value matrix
[0132] Specifically, three different learnable weight matrices are utilized to perform a linear transformation on the skeleton sequence encoded in step S601 and map the node embedding vectors in the skeleton sequence to the query matrix key matrix and value matrix respectively. The mapping formulas are as follows:
[0133] Q = G × W q 、K = G × W k 、V2 = G × W v (11)
[0134] S603. Based on the skeleton sequence nodes mapped to the query matrix key matrix and value matrix calculate the self-attention values of the current node in the skeleton sequence with other nodes, and based on the self-attention values, adopt the multi-head attention mechanism to learn the weights of different skeleton sequence nodes, and perform a linear transformation through the parameter matrix W O to obtain the learned fused feature map.
[0135] The process is as follows:
[0136] First, calculate the product of the Q and K T matrices of the current node to obtain the attention weight matrix, where K T is the transpose of the key matrix K, and to prevent the calculated value from being too large, the attention weight needs to be divided by At the same time, use the softmax function to output the attention coefficients of the current node with other nodes, and finally weight the attention coefficients with the matrix V2 to calculate the attention values of the current node with other nodes. The calculation process is as follows:
[0137]
[0138] To allow the model to focus on different aspects of information in the skeleton sequence and improve the model's representation ability, the multi-head attention mechanism is adopted to learn different weights to obtain multiple sets of attention outputs, that is, repeat the above process multiple times, each time using different W q 、W k and W v for mapping to form multiple self-attention heads, and perform a linear transformation through a learnable parameter matrix W O to obtain the final multi-head attention output. The calculation process is as follows:
[0139] MultiHead(Q,K,V) = Concat(head1,…,head h )W O (13)
[0140] This step is based on the self-attention values between the current node and other nodes, and uses the multi-head attention mechanism to learn the weights of nodes in different skeleton sequence nodes, and performs a linear transformation through a learnable parameter matrix W O to obtain the final multi-head attention output, that is, the learning is completed, and the fused feature map after learning is obtained.
[0141] S7. Input the fused feature map after learning into the fully connected layer and softmax to calculate the behavior category scores, and perform human behavior recognition output based on the score values.
[0142] First, the output of step S6 is used as the input of step S7 and input into the fully connected layer. Inside the fully connected layer, a convolution operation is performed through a convolution kernel with the same size and number of channels as the input feature map to integrate the output feature map of step S6 into a value, and then this is repeated N times to obtain N values, which are the scores for different behavior categories. This helps to reduce the influence of feature positions on the classification results and improves the robustness of the entire network.
[0143] After that, the scores output by the fully connected layer are input into the softmax function. Softmax maps the N values output by the fully connected layer to real numbers between 0 and 1 respectively, and performs normalization processing, ensuring that the sum of the output scores for each behavior category is 1 while also ensuring the non-negativity of the scores. In this way, the scores can be represented as probabilities.
[0144] The calculation formula of the Softmax function is:
[0145]
[0146] In the formula, s i is the output behavior category score, e i is the score for the i-th behavior category to be recognized, n represents the number of all behaviors to be recognized, e j represents the sum of all behavior category scores, and i = 1…n;
[0147] The Softmax function outputs multiple behavior category scores. The category with the highest probability, that is, the classification with the highest score, is the final output behavior category we obtain.
[0148] The second object of the present invention is to provide a multi-stream spatio-temporal graph convolutional behavior recognition system based on Transformer fusion, including:
[0149] A bone information and joint information acquisition module, which is used to extract the skeleton information of the human body to be recognized using a lightweight backbone network to obtain bone information and joint information;
[0150] A joint skeleton sequence acquisition module, which is used to obtain a joint skeleton sequence based on joint information;
[0151] A bone skeleton column sequence acquisition module, which is used to obtain a bone skeleton column sequence based on bone information;
[0152] A motion skeleton sequence acquisition module, which is used to obtain a motion skeleton sequence of human motion information based on the joint skeleton sequence and the bone skeleton column sequence;
[0153] A fused feature map acquisition module, which is used to input the motion skeleton sequence into a behavior recognition network for feature extraction, and stack the features extracted by each branch of the behavior recognition network according to the channel dimension to achieve feature fusion, and obtain a fused feature map;
[0154] A learning module, which is used to input the fused feature map of the fused feature map acquisition module into a spatio-temporal self-attention unit for global feature learning;
[0155] An output module, which is used to input the learned fused feature map into a fully connected layer and softmax to calculate the behavior category score, and perform human behavior recognition output based on the score value.
[0156] Embodiment 2:
[0157] Embodiment 2 is an offline experiment on the method of Embodiment 1. In this invention patent, experiments are carried out on the NTU-RGB+D60 and NTU-RGB+D120 skeleton sequence behavior recognition datasets. The Stochastic Gradient Descent (SGD) algorithm with a momentum of 0.9 and a weight decay of 0.0005 is used to update the model parameters. The initial learning rate is 0.01, and the learning rate is dynamically updated by cosine annealing. The training batch is 32, and the maximum training step is 16. The number of sampled frames for each video is 100, and it is fixed at two entities. When there is only one entity in the video, it is filled with 0.
[0158] As Figure 3 shown, it is a visualization display of the experimental results, showing the actual effect diagram of key point extraction from the input video to the human pose estimation model and action classification of behavior recognition. It can be seen from the figure that the human pose estimation model has high accuracy for key points, and behavior recognition can accurately classify the actions of pedestrian targets.
[0159] Embodiment 3:
[0160] Example 3 is a detailed comparison of Example 1 with its representative algorithms of the same kind, covering several different algorithms, including the baseline network (Spatial temporal graph convolutional networks for skeleton-based action recognition, ST-GCN), the two-stream (joint, skeleton) adaptive graph convolutional network (Two-Stream Adaptive Graph Convolutional Networks for Skeleton-Based Action Recognition, 2S-AGCN), the shifted graph convolutional network that combines joint, skeleton, joint motion, and skeleton motion information (Skeleton-Based Action Recognition with Shift Graph Convolutional Network, 4S-ShiftGCN), and the ST-TR network that also uses the Transformer self-attention mechanism (Skeleton-based Action Recognition via Spatial and Temporal Transformer Networks, ST-TR). The comparison results are shown in Table 1. The experimental results show that the multi-stream behavior recognition algorithm 4S-ST-GCNFormer proposed in this invention patent performs excellently on the NTU-RGB+D60 and NTU-RGB+D120 datasets. Compared with the representative algorithms of the same kind, its recognition accuracy reaches an excellent level.
[0161] The metrics used for the analysis of the comparison results in this example are as follows:
[0162] X-Sub (%) : Accuracy of the Cross-Subject test. That is to say, the test set and the training set come from data of different people. This metric measures the performance of the model when facing actions of people it has never seen before and is an important criterion for testing the generalization ability of the model.
[0163] X-View (%) : Accuracy of the Cross-View test. Here, it refers to the performance of the model under different perspectives. For example, if the training set is data collected from one camera perspective, the test set may come from cameras of other perspectives. This metric reflects the model's ability to recognize the same action under different perspectives.
[0164] X-Set (%) : Accuracy of the Cross-Set test. It means training and testing on different datasets or data subsets. Here, it may refer to cross-validation on different parts of the dataset to evaluate the performance of the model under different combinations of training sets and test sets.
[0165] In summary, these metrics evaluate the generalization ability of the model, i.e., whether it can maintain a high accuracy rate across different people (X-Sub), different perspectives (X-View), and different data combinations (X-Set).
[0166] Table 1 Comparison of recognition accuracies of different action networks on the NTU-RGB+D60 / 120 dataset
[0167]
[0168]
[0169] The above has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and all these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A multi-stream spatio-temporal graph convolutional behavior recognition method based on Transformer fusion, characterized in that Including: S1. Use a lightweight backbone network to extract the skeleton information of the human body graph to be recognized, obtaining bone information and joint information; S2. Obtain the joint skeleton sequence based on the joint information; S3. Obtain the bone skeleton sequence based on the bone information; S4. Based on the joint skeleton sequence and the bone skeleton sequence, obtain the motion skeleton sequence of the human body motion information; S5. Input the motion skeleton sequence into the behavior recognition network for feature extraction, and stack the features extracted by each branch of the behavior recognition network according to the channel dimension to achieve feature fusion, obtaining a fused feature map; S6. Input the fused feature map in step S5 into the spatio-temporal self-attention unit for global feature learning; S7. Input the learned fused feature map into the fully connected layer and softmax to calculate the behavior category score, and perform human behavior recognition output based on the score value.
2. The method according to claim 1, wherein: In step S1, the process of using a lightweight backbone network to extract the skeleton information of the human body graph to be recognized includes: S101. Change the number of channels of the human body graph to be recognized to 32 channels, and perform a convolution operation on the changed human body graph to be recognized, obtaining the first branch input feature map and the second branch input feature map; S102. Use two branches with different resolution features to extract the features of the first branch input feature map and the second branch input feature map respectively for multiple times, obtaining the first branch multi-scale feature map and the second branch multi-scale feature map; S103. Use the multi-scale feature fusion module to redefine and fuse the sizes of the first branch multi-scale feature map and the second branch multi-scale feature map respectively, obtaining the first branch fused feature map and the second branch fused feature map; S104. Input the first branch fused feature map and the second branch fused feature map into two identity mapping layers respectively, obtaining the first branch mapped feature map and the second branch mapped feature map. At the same time, perform a convolution operation on the second branch fused feature map to obtain a high-dimensional convolution map; S105. Use 4 cascaded basic feature extraction units to extract and fuse the first branch mapped feature map, the second branch mapped feature map, and the high-dimensional convolution map respectively, obtaining the first branch feature map, the second branch feature map, and the third branch feature map; S106. Input the first branch feature map, the second branch feature map, and the third branch feature map into 3 identity mapping layers respectively, and input the third branch feature map into a 3×3 convolution layer with an output channel number of 192 and a stride of 2, obtaining the first branch mapped map, the second branch mapped map, the third branch mapped map, and the fourth branch high-dimensional feature map; S107. Input the first branch mapped map, the second branch mapped map, the third branch mapped map, and the fourth branch high-dimensional feature map into 3 cascaded basic feature extraction units to obtain the bone information and the joint information.
3. The method according to claim 2, wherein: In step S102, two branches with different resolution features use the lightweight depth convolution transformation module to extract the features of the first branch input feature map and the second branch input feature map respectively, and the lightweight depth convolution transformation module includes a feature extraction layer; among them, The extraction process of the feature extraction layer includes: S1021. Use layer normalization to eliminate the difference between the input feature map of the first branch and the input feature map of the second branch; S1022. Use depthwise separable convolution with a stride of 3×3 to extract the context information of the local regions of the pixels of the input feature map of the first branch and the input feature map of the second branch after normalization in step S1021, and obtain the input feature map of the first branch after feature extraction and the input feature map of the second branch after feature extraction.
4. The method according to claim 2, wherein: In step S103, the multi-scale feature fusion module includes a feature size redefinition layer and a feature information fusion layer; among them, The feature size redefinition layer redefines the size through upsampling and downsampling operations, and superimposes them, and uses the superimposed feature map as the input of the feature information fusion layer; The feature information fusion layer sequentially passes the superimposed feature map of the two branches output by the feature size redefinition layer through a 1×1 convolution, a 3×3 depthwise separable convolution, and a cascaded channel attention to obtain the first branch fusion feature map and the second branch fusion feature map.
5. The method according to claim 2, characterized in that: In step S105, the basic feature extraction unit includes a depth convolution transformation module and a feature fusion module; Among them, the depth convolution transformation module is used to extract the deep features of the first branch mapping feature map, the second branch mapping feature map, and the high-dimensional convolution map respectively, and input the extracted deep features into the feature fusion module for multi-scale feature fusion.
6. The method according to claim 1, wherein: The expression of the joint skeleton sequence is: where T is the length of the skeleton sequence, V is the number of nodes in a single skeleton, and x tv、 y tv represent the specific positions of joints in space, and s tv represents the confidence of key point detection; The expression of the bone skeleton sequence is: In the formula, i represents different key points in the same frame, j represents the same key point in different frames, and Es represents the skeleton connection relationship of each frame; The expression of the motion skeleton sequence is:
7. The method according to claim 1, wherein: In step S5, the behavior recognition network includes 4 branches, and each branch uses a spatio-temporal graph convolutional network to extract features from the motion skeleton sequence. The first five layers of the spatio-temporal graph convolutional network are spatio-temporal graph convolutional units, and the spatio-temporal graph convolutional unit includes a spatial graph convolutional unit and a temporal convolutional unit.
8. The method according to claim 1, characterized in that: In step S6, the process of global feature learning includes: S601. Perform position encoding on the skeleton sequence nodes in the fusion feature map of step S5; S602. Map the encoded skeleton sequence nodes to the query matrix respectively Key matrix And value matrix Based on the mapped query matrix key matrix and value matrix of the skeleton sequence nodes, calculate the self-attention values between the current node and other nodes in the skeleton sequence, and based on the self-attention values, combined with the parameter matrix W O perform a linear transformation to obtain the learned fused feature map.
9. The method according to claim 8, characterized in that: In step S602, map the encoded skeleton sequence nodes to the query matrix key matrix and value matrix The mapping formula in is: Q = G × W q 、K = G × W k 、V2 = G × W v (11) Among them, is the skeleton sequence, and w is the learning weight matrix; In step S603, the calculation formula for the self-attention value between the current node and other nodes is: where K T is the transpose of the key matrix K.
10. A multi-stream spatio-temporal graph convolutional behavior recognition system based on Transformer fusion, characterized in that, Includes: A bone information and joint information acquisition module, which is used to use a lightweight backbone network to extract the skeleton information of the human body map to be recognized, and obtain bone information and joint information; A joint skeleton sequence acquisition module, which is used to obtain a joint skeleton sequence based on the joint information; A bone skeleton sequence acquisition module, which is used to obtain a bone skeleton sequence based on the bone information; A motion skeleton sequence acquisition module, which is used to obtain a motion skeleton sequence of human motion information based on the joint skeleton sequence and the bone skeleton sequence; A fusion feature map acquisition module, which is used to input the motion skeleton sequence into the behavior recognition network for feature extraction, and superimpose the features extracted by each branch of the behavior recognition network according to the channel dimension to achieve feature fusion, and obtain a fusion feature map; A learning module, which is used to input the fusion feature map of the fusion feature map acquisition module in the step into a spatio-temporal self-attention unit for global feature learning; An output module, configured to input the learned fused feature map into a fully connected layer and perform softmax to calculate the behavior category scores, and perform human behavior recognition output based on the score values.