Dual-stream action recognition system and method based on spatiotemporal GT
By adopting a dual-flow structure of space-time G-T in the human body motion recognition system, combined with graph convolution and Transformer flow, the problem of not connecting but low recognition rate of highly correlated bone characteristics in the prior art is solved, and a higher accuracy of motion recognition is achieved.
Patent Information
- Application Number
- CN202410729004.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-06-06
AI Technical Summary
The existing human body motion recognition method has low motion recognition rate when dealing with two-part bone characteristics that are not connected but highly correlated.
The dual-flow action recognition system based on space-time G-T is adopted, combining space-time graph convolutional flow and space-time Transformer flow, and the global space-time features of the action are extracted through the adaptive space-time graph convolution network, time-convolution network, space transformer module and time-transformer module, and the global space-time characteristics of the action are improved through information fusion.
Through information fusion and global spatiotemporal feature extraction, the accuracy of action recognition is improved, especially when dealing with unconnected but highly relevant bone characteristics, it is significantly improved.
Smart Images

Figure CN118675230B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and in particular to a dual-stream action recognition system and method based on spatiotemporal GT. Background Art
[0002] Human action recognition has become a core task in video understanding. Existing research has covered various feature representations of human actions, such as RGB frames, optical flow, sound waves, and human skeletons. Among these modes, the deep skeleton data used by human skeleton features can be compared with other information extraction to ignore background elements during the training process of the recognition model. Combining time and space to analyze the human skeleton, spatial features can be effectively extracted in the topological map for learning, which has broad application prospects in the fields of behavior recognition, action prediction, video understanding, intelligent monitoring, pedestrian tracking, human-computer interaction, etc.
[0003] In 2018, Sijie Yan et al. first proposed the ST-GCN model. ST-GCN overcomes the limitations of traditional manual model methods and traversal rules. By constructing a human skeleton model, it extracts the temporal features of continuous frames and the spatial features of skeleton joints within the frame for action recognition, and has achieved excellent results in two typical action datasets Kinetics and NTU-RGB+D. Later, many GCN-based methods were expanded on this. Although GCN has been proven to be effective in the field of action recognition, it still has certain limitations, such as insufficient flexibility of the attention mechanism, failure to connect nodes with large correlations in motion, and failure to utilize the human characteristics of bone length and direction. In order to solve the above defects, Lei Shi et al. proposed a two-stream adaptive graph convolutional network (2S-AGCN). The two shortcomings of ST-GCN were improved, and two innovative points, two-stream network and adaptive structure, were proposed. This data-driven method increases the flexibility of the graph construction model, making it more versatile to adapt to different data samples. However, due to the limitation of the convolution kernel size, only neighborhood information exchange can be performed in time and space. Zhang et al. proposed a simple and effective semantically guided neural network (SGN) for skeleton-based action recognition. By explicitly introducing the high-level semantics of joints (joint types and frame indices) into the network, the feature representation capability of the network is enhanced. Chiara Plizzari et al. proposed a new spatiotemporal transformer network (ST-TR) that uses the transformer's self-attention operator to simulate the dependencies between nodes. Ignoring the natural connections in space, each skeleton is set as an independent point in the graph, thereby improving the accuracy of action recognition of some unconnected but highly correlated skeletal features in reality.
[0004] Most of the above-mentioned human action recognition methods combine graph convolution with RNN or CNN to describe temporal correlation, artificially construct the skeleton as a joint coordinate vector sequence or pseudo image, and then feed the pseudo image into RNNs or CNNs to generate predictions. However, representing skeleton data as a vector sequence or pseudo image cannot fully utilize the graph structure of skeleton data, and it is difficult to generalize to any form of skeleton. Skeletons are naturally constructed as graphs in non-Euclidean space, with joints as vertices and their natural connections in the human body as edges. They regard the human skeleton as an unrelated and complete graph, resulting in low recognition rates for actions with two unconnected but highly correlated skeletal features. Summary of the invention
[0005] The purpose of the present invention is to provide a dual-stream action recognition system and method based on spatiotemporal GT, aiming to solve the problem that the existing human action recognition method has a low action recognition rate for two parts of skeleton features that are not connected but highly correlated.
[0006] To achieve the above objectives, in a first aspect, the present invention provides a dual-stream action recognition system based on spatiotemporal GT, comprising the following steps:
[0007] Input feature graph data, train the spatial adaptive spatial graph convolution network module based on the spatiotemporal graph convolution flow, and obtain training data;
[0008] Extract features from the training data using the temporal convolutional network module of the spatiotemporal graph convolutional flow, make category predictions, and obtain a first score;
[0009] The spatial transformer module of the spatiotemporal Transformer stream is used to initialize the connections between all relevant nodes of the body to the same strength, and then a trainable linear transformation is applied to the node features to extract the query vector, index vector, and value vector of all nodes;
[0010] The single frame data enters the spatiotemporal Transformer stream temporal transformer module, applies the trainable linear transformation to the inter-frame features, and extracts the query vector, index vector, and value vector of all frames of a single node;
[0011] The output of the spatial transformer module is used as the input of the temporal transformer module, features are extracted, and a category prediction is made to obtain a second score;
[0012] The first score and the second score are added to obtain a fusion score, and an action label is predicted based on the fusion score.
[0013] Among them, the training of the adaptive spatial graph convolutional network module requires adding two batch normalization layers and two ReLU layers.
[0014] Among them, the number of the spatiotemporal graph convolution flow of the adaptive spatial graph convolution network module and the spatiotemporal Transformer flow of the spatiotemporal Transformer flow time transformer module is 9.
[0015] Among them, label smoothing is introduced in the training of the adaptive spatial graph convolutional network module to optimize the traditional cross entropy loss function.
[0016] In a second aspect, the present invention also provides a dual-stream action recognition system based on spatiotemporal GT, including an adaptive spatial graph convolutional network module, a temporal convolutional network module, a spatial transformer module and a temporal transformer module;
[0017] The adaptive spatial graph convolutional network module is used to learn the topological structure of graphs of different layers and skeleton samples;
[0018] The temporal convolutional network module is used to model the temporal connections between adjacent or several frames of graphs of different layers and skeleton samples in the temporal dimension;
[0019] The spatial transformer module is used to capture the relationship between any joints in the image frame;
[0020] The temporal transformer module is used to capture the relationship between any joints between image frames.
[0021] The dual-stream action recognition method based on spatiotemporal GT of the present invention inputs feature map data, and trains a spatial adaptive spatial graph convolution network module based on a spatiotemporal graph convolution flow to obtain training data; extracts features from the training data using a temporal convolution network module of the spatiotemporal graph convolution flow, makes a category prediction, and obtains a first score; initializes the connections between all relevant nodes of the body to the same strength using a spatial transformer module of the spatiotemporal Transformer flow, applies a trainable linear transformation to the node features, and extracts query vectors, index vectors, and value vectors of all nodes; single-frame data enters the temporal transformer module of the spatiotemporal Transformer flow, applies a trainable linear transformation to inter-frame features, and extracts query vectors, index vectors, and value vectors of all frames of a single node; uses the output of the spatial transformer module as the input of the temporal transformer module, extracts features, makes a category prediction, and obtains a second Softmax score; compares the first Softmax score and the The second Softmax score is added to obtain a fusion score, and the action label is predicted based on the fusion score. This method uses information fusion, and they can learn feature information from each other, and enrich the action features by maximizing the mutual information between the two types of action representations. This method uses a new dilated convolution design module that is different from traditional convolution to learn action information containing local and global joint relationships in a longer time frame. A skeletal feature learning model composed entirely of Transformers is designed, which can more accurately capture information between any joints. The GCN-Transformer (space-time graph convolution flow and space-time Transformer flow) parallel network can maintain the natural topological structure of the human skeleton graph while more accurately capturing information between any joints. The accuracy is improved by the complementary information between the two stream action representations, and the existing human action recognition method solves the problem of low action recognition rate for two parts of skeletal characteristics that are not connected but highly correlated. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0023] Figure 1 This is the overall structure of the GCN-TransformerNetwork of the present invention.
[0024] Figure 2 It is a skeletal space-time diagram.
[0025] Figure 3It is a common convolution and an improved convolution form.
[0026] Figure 4 This is a comparison chart of the accuracy of 2S-AGCN and this method.
[0027] Figure 5 This is the confusion matrix diagram of some categories of this method.
[0028] Figure 6 Schematic diagram of the effect of this method on (a) NTU-60, (b) NTU-120 and (c) kinetics.
[0029] Figure 7 It is a flow chart of the dual-stream action recognition method based on spatiotemporal GT provided by the present invention.
[0030] Figure 8 It is a schematic diagram of the dual-stream action recognition system based on spatiotemporal GT provided by the present invention, wherein (C) is a schematic diagram of the spatiotemporal graph convolution flow, and (D) is a schematic diagram of the spatiotemporal Transformer flow.
[0031] In the figure: 1-adaptive spatial graph convolutional network module, 2-temporal convolutional network module, 3-spatial transformer module, 4-temporal transformer module. DETAILED DESCRIPTION
[0032] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0033] See also Figures 1 to 7 In a first aspect, the present invention provides a dual-stream action recognition system based on spatiotemporal GT, comprising the following steps:
[0034] S1 inputs feature graph data and trains the spatial adaptive spatial graph convolution network module based on the spatiotemporal graph convolution flow to obtain training data;
[0035] Specifically, we introduce label smoothing in model training to optimize the traditional cross entropy loss function, add noise through soft one-hot, reduce the weight of the category of the real sample label when calculating the loss function, and ultimately suppress overfitting.
[0036] Suppose the spatiotemporal graph of the skeleton sequence with a length of T frames is G = (V, E), V is a point set, |V| = n represents the joint points at all times, and the edge set E is divided into two subsets, one is the skeleton edge E between each joint in each frame s, reflects the spatial properties in a single frame. Another subset E F is the connecting edge of the corresponding joint between each frame, reflecting the time attribute between multiple frames. Figure 2 As shown in Figure 2. The vertex feature in G = (V, E) is the joint point coordinate vector F (v ti ), v ti Represents the coordinates of the i-th joint in the t-th frame.
[0037] The form of the spatial graph convolutional neural network can be obtained:
[0038]
[0039] Among them, the input of the l-th layer network is H (l) ∈R N×D , the input signal is initially H (0) =X, N is the number of nodes in the graph, is the degree matrix of the natural human skeleton, is the adjacency matrix of the natural human skeleton, each node has a D-dimensional feature vector, W (l) ∈R D×D is the weight of the lth layer, and σ is the activation function.
[0040] The adaptive graph convolution layer optimizes the graph topology together with other parameters of the network in an end-to-end learning manner. The graph is unique for different layers and samples, which greatly increases the flexibility of the model. At the same time, it is designed as a residual branch to ensure the stability of the original model. Among them, the topology of the graph is actually determined by the adjacency matrix and the mask. In order to make the structure of the graph adaptable, the adaptive graph convolution layer can be expressed as:
[0041]
[0042] In addition to the traditional adjacency matrix In addition, a matrix L obtained entirely by training is added to represent the connection strength between two points, and E is a matrix related to the input data. A unique matrix is learned for each sample to reflect the connection strength between any two points.
[0043] Unlike spatial topology, the temporal topology of joints is linear. Therefore, temporal relationships are usually captured by using ordinary convolution operations instead of graph convolution. Different from the traditional ordinary convolution operation, we propose a new type of temporal convolution operation called TDCN using dilated convolution, which can be expressed as:
[0044] H l+1 =TDCN(AGCN(H (l) ))
[0045] Among them, TDCN (*) is a dilated convolution kernel with a kernel size of K and multi-scale dilation, see Figure 3 Compared with ordinary convolution, dilated convolution increases the receptive field of the convolution kernel while keeping the number of parameters unchanged, so that the convolution output of each step contains a wider range of information; at the same time, the size of the output feature map remains unchanged. By improving the overall network architecture and introducing residual connections, the recognition accuracy is further improved.
[0046] After the spatial graph convolution, a Batch normalization (BN) layer and a ReLU layer are added to speed up the training and convergence of the network.
[0047] S2 uses the temporal convolutional network module of the spatiotemporal graph convolutional flow to extract features from the training data, make category predictions, and obtain a first score;
[0048] Specifically, the data obtained after the ReLU layer enters the temporal convolution TDCN to extract the temporal information. Since TDCN is a dilated convolution kernel with a kernel size of K and multi-scale dilation, we can extract temporal information of different frames, merge temporal information of different scales, increase the number of temporal information features, and retain the original information through the residual link to enter the next layer.
[0049] After the temporal graph convolution, a Batch normalization (BN) layer and a ReLU layer are added to speed up the training and convergence of the network.
[0050] The above model is called a spatiotemporal graph convolution module, such as Figure 1 As shown in the figure, our data passes through nine spatiotemporal convolution modules to extract features layer by layer, and finally enters Softmax to make category predictions.
[0051] S3 uses the spatial transformer module 3 of the spatiotemporal Transformer stream to initialize the connections between all relevant nodes of the body to the same strength, and then applies the trainable linear transformation to the node features to extract the query vector, index vector and value vector of all nodes;
[0052] Specifically, the spatial Transformer flow abandons the original naturally connected joints and initializes the connections between all relevant nodes of the body to the same strength. The skeleton data passes through the frame at a given time t. For each node Vti of the skeleton, the query vector q, index vector k and value vector v of all nodes are first extracted by applying a trainable linear transformation to the node features.
[0053] Apply the query-index vector dot product to get the correlation strength between two nodes i and j. The formula is as follows:
[0054]
[0055] Multiply the intensity by the matrix of the value vector v, as follows:
[0056]
[0057] In the same frame, each joint point contains the features of other joint points, thus forming a correlation between all the related nodes in a single frame, allocating reasonable joint point attention for the input data, and providing a high correlation strength for two joint points in a certain action that are highly correlated but not naturally connected.
[0058] S4 single frame data enters the spatiotemporal Transformer stream temporal transformer module 4, applies trainable linear transformation to inter-frame features, and extracts query vectors, index vectors, and value vectors for all frames of a single node;
[0059] Specifically, the single frame data enters the temporal Transformer flow, at this time we discuss the correlation strength of a single joint in different frames, each single joint is considered independent, and the correlation between frames is calculated by comparing the changes in the embedding of the same body joint along the time dimension. Given a skeleton point v, for different frames of the time axis, first extract the query vector q, index vector k and value vector v of all frames of a single node by applying a trainable linear transformation to the inter-frame features.
[0060] Apply the query-index vector dot product to obtain the correlation strength between two frames t and s The formula is as follows:
[0061]
[0062] Multiply the intensity by the matrix of the value vector v, as follows:
[0063]
[0064] S5 uses the output of the spatial transformer module 3 as the input of the temporal transformer module 4, extracts features, makes a category prediction, and obtains a second score;
[0065] Specifically, the spatial Transformer stream and the temporal Transformer stream are arranged in sequence, and the output of the spatial Transformer stream is used as the input of the temporal Transformer stream. The spatiotemporal Transformer stream can be expressed as:
[0066] H l+1 =TT(ST(H (l) ))
[0067] Here, H represents the output of the current spatiotemporal Transformer layer.
[0068] In order to maintain structural symmetry with the spatiotemporal graph convolution flow (spatial-temporal GCN flow), we replicate the spatiotemporal Transformer layer nine times, extract features layer by layer, and finally enter Softmax to make category predictions.
[0069] S6 adds the first score and the second score to obtain a fusion score, and predicts an action label based on the fusion score.
[0070] Specifically, the softmax scores of the two streams (the first score and the second score) are added together to obtain the fusion score and predict the action label. We appropriately weaken the score value of the spatiotemporal Transformer stream, and increase the proportion of the spatiotemporal Transformer stream to supplement the features when the recognition rate of the spatiotemporal graph convolution stream is not high for some actions, and obtain the final classification result.
[0071] To verify the effectiveness of the proposed method, we conducted extensive experiments on three widely used datasets: NTU-RGBD 60, NTU-RGBD 120, and KineticsSkeleton400.
[0072] NTU RGB+D60 and NTU RGB+D120. The NTU RGB+D60 (NTU-60) dataset is a large-scale benchmark for 3D human action recognition collected using Microsoft Kinect v2. The skeleton information consists of 3D coordinates of 25 human joints and a total of 60 different action classes. The NTU-60 dataset follows two different evaluation criteria. The first one is called Cross-View Evaluation (X-View), which uses 37,920 training samples and 18,960 test samples, split according to the camera view the action comes from. The second one is Cross-Subject Evaluation (X-Sub), which consists of 40,320 training samples and 26,560 test samples, collected from 40 different subjects and divided into two groups, one for training and the other for testing. NTU RGB+D120 (Liu et al. (2019)) (NTU-120) is an extension of NTU-60, which adds 57,367 new skeleton sequences representing 60 new actions. To perform evaluation, the extended dataset follows two standards: the first one is cross-subject evaluation (X-Sub), which is the same as NTU-60, while the second one is called cross-setting evaluation (X-Set), which replaces cross-view by splitting the training and testing samples based on the names of camera setting ids.
[0073] The Kinetics Skeleton Dataset (Yan et al. (2018)) was obtained by extracting skeleton annotations from videos constituting the Kinetics 400 dataset (Kay et al. (2017)) using the OpenPose toolbox (Cao et al. (2019)). It consists of 240,436 training samples and 19,796 test samples, representing a total of 400 action classes. Each skeleton consists of 18 joints, each with 2D coordinates and confidence. For each frame, up to 2 people are selected based on the highest confidence score.
[0074] Implementation details
[0075] All experiments were conducted on the PyTorch deep learning framework. Stochastic gradient descent (SGD) with Nesterov momentum (0.9) was used as the optimization strategy. The batch size of the spatiotemporal GCN flow was 32. The traditional cross entropy loss function was selected as the loss function by introducing label smoothing optimization. The weight decay was set to 0.0001. For the NTU-RGBD dataset, there are at most two people in each sample of the dataset. If the number of objects in the sample is less than 2, we fill the second object with 0. The maximum number of frames for each sample is 300. For samples with less than 300 frames, we repeatedly sample until it reaches 300 frames. The learning rate is set to 0.1 and divided by 10 at the 30th, 40th, and 70th epochs respectively. The training process ends at the 10th round. The spatiotemporal Transformer flow trained the model for a total of 100 epochs on NTU-60 and NTU-120 with a batch size of 64 and SGD as the optimizer. The learning rate was set to 0.1 at the beginning and then decreased by 10 times at {60, 80} rounds respectively. For the Kinetics-Skeleton dataset, the spatiotemporal GCN stream and the spatiotemporal Transformer stream are set the same, with a batch size of 32, containing 150 frames, each with 2 bodies. The learning rate is set to 0.1 and divided by 10 at the 45th and 55th epochs. The training process ends at the 80th epoch.
[0076] Comparison and analysis of experimental results
[0077] To verify the effectiveness of our network, we compared our prediction accuracy with some of the most advanced methods on the NTU-60, NTU-120 and NW-UCLA datasets under two evaluation protocols. The experimental results of the two-stream and the whole on NTURGBD are shown in Table 1, where * refers to the results obtained by experiments using four-stream fusion (joints, skeletons, joint motion and skeleton motion).
[0078] Table 1
[0079]
[0080] Table 1 shows the comparison of our work with previous related methods on NTU-60. In comparison, our method improves the performance of the baseline work 2S-AGCN by 1.7% on X-View and 4.2% on X-Sub. Our model achieves excellent performance, and compared with some methods that use four-stream data to improve accuracy, we only use two-stream data to achieve the same effect. Compared with the classic four-stream fusion method Shift-GCN, it exceeds 0.3% on X-View and 2.0% on X-Sub.
[0081] Table 2 shows our performance on NTURGBD120, where * refers to the results obtained by experiments using four-stream fusion.
[0082] Table 2
[0083]
[0084] In Table 2, we conducted comparative experiments on the NTU-120 dataset on the X-Sub and X-Set benchmarks. The competitive results in Table 2 verify that our proposed method outperforms all methods on the X-Set of the NTU-120 dataset and is close to the current sota level on X-Sub. Table 3 is the performance of our work on kinetics:
[0085] Table 3
[0086]
[0087] As shown in Table 3, our proposed GTnet achieves the best accuracy of 39.0% on the kinetics dataset, surpassing the most advanced methods in the table. However, due to the large amount of data and high complexity of the kinetics dataset, the traditional model no longer has an advantage, and the kinetics dataset is still a difficult dataset for work in this direction.
[0088] Figure 4 The accuracy comparison of our method and the traditional model 2S-AGCN network in NTU-RGBD60 for individual categories. Figure 4 As shown in Figure 2, in the comparison between our method and the 2S-AGCN method, our method has higher accuracy in most categories, and only in a few categories our accuracy is slightly lower than the baseline. The disadvantages are the same. Our accuracy is significantly lower in categories 11 (reading) and 29 (playing with mobile phones). We will Figure 5 Through Figure 5In Figures A and B, the categories that are misclassified with categories 11 and 29 are 10 (writing) and 28 (talking). This shows that our model cannot describe in detail the specific actions that the skeleton points cannot describe, such as writing and reading, where only the objects in the hands change. The actions themselves are very similar and cannot be accurately judged by skeleton points alone. Then the dual-mode data format is born. It can be said that dual-mode data to improve the distinction between similar actions is a more advantageous work direction in the future.
[0089] Figure 6 For the performance of our work on NTU-60, NTU-120 and kinetics datasets, we achieve the best results on xview in NTU60.
[0090] Secondly, we conducted an ablation experiment to explore the contribution of each module to the entire algorithm model. The ablation experiment results are shown in Table 4:
[0091] Table 4
[0092]
[0093] We conducted an ablation experiment to explore the contribution of each module to the entire algorithm model. The ablation experiment results are shown in Table 5:
[0094] Table 5
[0095]
[0096] Table 6 shows the impact of different numbers of convolutional layers on the Gnet results:
[0097] Table 6
[0098]
[0099] The method is based on the GCN parallel Transformer network, which fuses spatial and temporal modules in a parallel manner to extract global spatiotemporal features of action information. The model includes two parallel streams: the spatiotemporal GCN stream and the spatiotemporal Transformer stream. The spatiotemporal GCN stream is designed to obtain action representations with natural topological structures of human skeletons. The spatiotemporal Transformer stream is designed to obtain action representations containing global relationships between joints. Since the action representations produced by these two streams contain different features and each stream knows little about each other, we fuse the output results of these two parts. Through information fusion, the spatiotemporal GCN stream and the spatiotemporal Transformer stream can learn information from each other and enrich the action features by maximizing the mutual information between the two types of action representations. Ablation studies are conducted in this work, which verify the effectiveness of our method. Experiments on three public datasets demonstrate the superiority of this method.
[0100] See also Figure 8 ,In a second aspect, the present invention also provides a dual-stream action recognition system based on ,spatial GT, including an adaptive spatial graph convolutional network module 1, a temporal convolutional network module 2, a spatial transformer module 3 and a temporal transformer module 4;
[0101] The adaptive spatial graph convolutional network module 1 is used to learn the topological structure of the graph of different skeleton points in the same frame;
[0102] The temporal convolutional network module 2 is used to model the temporal connections between adjacent or several frames of graphs of different layers and skeleton samples in the temporal dimension;
[0103] The spatial transformer module 3 is used to capture the relationship between any joints in the image frame;
[0104] The time transformer module 4 is used to capture the relationship between any joints between image frames.
[0105] In this embodiment, the adaptive spatial graph convolutional network module 1 learns the topological structure of graphs of different layers and skeleton samples, the temporal convolutional network module 2 models the temporal connection between adjacent frames of graphs of different layers and skeleton samples in the time dimension, and the spatial transformer module 3 and the temporal transformer module 4 capture the relationship between any joints within and between image frames.
[0106] What is disclosed above is only a preferred embodiment of the dual-stream action recognition system and method based on spatiotemporal GT of the present invention. Of course, this cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiments are implemented, and equivalent changes made according to the claims of the present invention are still within the scope of the invention.
Claims
1. A dual-stream action recognition method based on spatiotemporal GT, characterized in that: The following steps are involved: Input feature graph data, train the adaptive spatial graph convolution network module based on the spatiotemporal graph convolution flow, and obtain training data; Extract features from the training data using the temporal convolutional network module of the spatiotemporal graph convolutional flow, make category predictions, and obtain a first score; The spatial transformer module of the spatiotemporal Transformer stream is used to initialize the connections between all relevant nodes of the body to the same strength, and then a trainable linear transformation is applied to the node features to extract the query vector, index vector, and value vector of all nodes; The single frame data enters the spatiotemporal Transformer stream temporal transformer module, applies the trainable linear transformation to the inter-frame features, and extracts the query vector, index vector, and value vector of all frames of a single node; The output of the spatial transformer module is used as the input of the temporal transformer module, features are extracted, and a category prediction is made to obtain a second score; The first score and the second score are added to obtain a fusion score, and an action label is predicted based on the fusion score.
2. The dual-stream action recognition method based on spatiotemporal GT as claimed in claim 1, characterized in that: The training of the adaptive spatial graph convolutional network module requires adding two batch normalization layers and a ReLU layer.
3. The dual-stream action recognition method based on spatiotemporal GT as claimed in claim 1, characterized in that: There are 9 adaptive spatial graph convolutional network modules and 9 spatiotemporal Transformer stream-time transformer modules.
4. The dual-stream action recognition method based on spatiotemporal GT as claimed in claim 1, characterized in that: Label smoothing is introduced in the training of the adaptive spatial graph convolutional network module to optimize the traditional cross entropy loss function.
5. A dual-stream action recognition system based on spatiotemporal GT, applied to the dual-stream action recognition method based on spatiotemporal GT as claimed in claim 1, characterized in that: It includes an adaptive spatial graph convolutional network module, a temporal convolutional network module, a spatial transformer module and a temporal transformer module; The adaptive spatial graph convolutional network module is used to learn the topological structure of the graph of different skeleton point samples in the same frame; The temporal convolutional network module is used to model the temporal connections between adjacent or several frames of graphs of different layers and skeleton samples in the temporal dimension; The spatial transformer module is used to capture the relationship between any joints in the image frame; The temporal transformer module is used to capture the relationship between any joints between image frames.