Action recognition method and device, electronic equipment and storage medium
The multi-stage spatio-temporal fusion network addresses limitations in skeletal data action recognition by integrating local spatial and global temporal features, improving accuracy and robustness for complex actions.
Patent Information
- Application Number
- CN202510811538.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-15
AI Technical Summary
The existing skeleton-based action recognition methods have shortcomings in local spatio-temporal feature fusion and timing information utilization, and it is difficult to accurately identify human movements. Especially in long-term action sequences, key features at different stages are ignored, and timing modeling lacks global perception capabilities.
A multi-stage spatiotemporal fusion network is adopted, including an action structure diagram convolution module, a time convolution module, a spatiotemporal tuple attention module and an inter-frame feature aggregation module. By adaptively modeling the dynamic relationship and timing characteristics between the nodes, a multi-stage spatiotemporal fusion model is constructed.
It improves the accuracy and robustness of recognition of complex actions, is suitable for offline batch processing and real-time recognition, is suitable for embedded device deployment, has good real-time recognition capabilities, and can effectively handle noise and missing in skeleton data.
Smart Images

Figure CN120318921A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of deep learning and computer vision, and particularly relates to an action recognition method, apparatus, electronic device and storage medium. Background Art
[0002] With the development of deep learning, human behavior recognition has been widely applied in various industries, especially playing an important role in video understanding. In recent years, it has become an active research field. For example, behavior recognition plays a very important role in fields such as healthcare, intelligent monitoring, and autonomous driving. Human actions can be recognized from multiple modalities, such as appearance, depth optical flow, and skeleton. In recent years, more and more people are interested in skeleton-based action recognition. The core idea of skeleton-based behavior recognition is to capture the spatio-temporal dynamic changes of human key joints, and use the positions and movement trajectories of these joints to identify and distinguish different actions or behaviors. Skeleton data has strong robustness, and usually models the spatio-temporal structure of the human body through methods such as graph convolutional networks, so as to perform behavior classification efficiently and accurately.
[0003] Deep learning has been widely applied to the spatio-temporal evolution model of skeleton sequences. Various network structures have been developed, such as Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM) network, Convolutional Neural Network (CNN), and Graph Convolutional Neural Network (GCN). In the early stage, RNN / LSTM was the preferred network for developing short-term and long-term time dynamics. Recently, there has been a trend to use feed-forward (i.e., non-recurrent) convolutional neural networks and skeletons for speech and language sequence modeling. Most skeleton-based methods organize the coordinates of joints into a 2D map and resize the map to fit the input size of the CNN, where its rows / columns correspond to different types of joints / frame indices. In these methods, long-term dependencies and semantic information are expected to be captured by the large receptive field of the deep network, which usually leads to a high complexity of the model. Moreover, the skeleton is in the form of a graph rather than a 2D or 3D grid, which makes it difficult to use proven models such as convolutional networks. Recently, graph neural networks (GCNs) that generalize convolutional neural networks to arbitrary structured graphs have received increasing attention and have been successfully applied to many applications.
[0004] At present, skeleton-based human action recognition methods still have drawbacks. Firstly, there is insufficient fusion of local spatio-temporal features. Currently, most GCN-based methods use fixed human skeleton graphs for spatial modeling, but this static topological structure ignores the dynamic relationships between human joint points in different action phases. For example, in the starting phase of the "waving" action, the association between the hand and the torso is strong, while in the later stage of waving, the hand may interact more closely with the head or shoulder. Existing methods are difficult to adaptively adjust these dynamic relationships, resulting in limited feature expression ability. Some improved methods attempt to optimize the representation of the skeleton graph through self-attention mechanism or learnable adjacency matrix, but still have difficulty in accurately modeling the feature changes in different phases, especially in long action sequences, where the key features in different phases may be ignored. Secondly, the utilization of temporal information is insufficient. Traditional GCN-based methods mainly model the dependency relationships of joint points in the spatial domain, while temporal information is usually modeled through simple temporal convolution or recurrent neural networks (such as LSTM, Gated Recurrent Unit (GRU)). This strategy may lead to strong short-term dependencies and insufficient long-term dependencies, being sensitive to action patterns within a short time range but difficult to capture long-term dependency information across multiple time steps; and the lack of global perception ability in temporal modeling. Based on time convolution, models usually use convolutional kernels of a fixed size, resulting in limited receptive fields and being unable to comprehensively capture the global features of the entire action sequence. All the above reasons lead to the inability to accurately recognize human actions based on skeleton data. Summary of the Invention
[0005] Based on this, in view of the technical problem that the prior art cannot accurately recognize human actions based on skeleton data, it is necessary to provide an action recognition method, device, electronic device and storage medium.
[0006] The present invention adopts the following technical solutions: In a first aspect, the present invention provides an action recognition method, the method comprising: Obtain multiple segments of skeleton data sequences, label the action types of each segment of skeleton sequence data to obtain a training set; Build a multi-stage spatio-temporal fusion network, use the training set as input and the action type as output, and train the multi-stage spatio-temporal fusion network to obtain a human action recognition model; Input the skeleton data sequence to be recognized into the human action recognition model to obtain the action type corresponding to the skeleton data sequence to be recognized; Among them, the multi-stage spatio-temporal fusion network includes an action structure graph convolutional module, a temporal convolutional module, a spatio-temporal tuple attention module, an inter-frame feature aggregation module, and a classification module connected in sequence. The action structure graph convolutional module is used to identify each joint point in the skeleton data of each frame and construct the skeletal topological relationship between each joint point to obtain the spatial features between each joint point. The temporal convolutional module is used to extract the temporal features of each joint point in the skeleton data of different frames. The spatio-temporal tuple attention module is used to enhance the features of the spatial features and the temporal features to obtain enhanced features. The inter-frame feature fusion module is used to integrate the dynamic information of each joint point in the skeleton data of different frames according to the enhanced features. The classification module is used to output the behavior type according to the dynamic information.
[0007] Further, the action structure graph convolutional module includes an action graph convolutional module and a behavior graph convolutional module; The action graph convolutional module includes: An encoder, which is used to learn the connection probability between any two joint points according to the input skeleton data sequence and generate an action link probability matrix according to the connection probability between any two joint points; A decoder, which is used to predict the position of the joint point at the next time point according to the action link probability matrix; The behavior graph convolutional module is used to diffuse the information of the joint point to non-neighboring joint points in a high-order power diffusion manner, and then perform structure graph convolution on the skeleton data sequence to obtain a structure link probability matrix. The information of the joint point includes the spatial coordinates of the joint point and the skeletal topological structure.
[0008] Further, the function of the encoder is: ; Among them, x is the coordinate sequence of the joint point, C is the number of types of action links, A represents the action link probability matrix, and the dimension is n n C , n represents the number of joint points, encode (·) is the encoder function.
[0009] Further, the function of the decoder is: ; Among them, refers to t the joint point feature at time refers to the joint point feature at the initial time, Refers to the joint point coordinates at the next moment. A Refers to the action link probability matrix.
[0010] Furthermore, the function of the structural graph convolution is: ; Wherein, X in Refers to the input feature matrix. L Is the number of diffusion times. P Is the skeleton partition set. p ∈ P Is the partition index. Is the p Power of the l adjacency matrix of the Is the trainable weight matrix. Is the feature transformation matrix.
[0011] Furthermore, the temporal features include the velocity change features and acceleration change features of each joint point in the skeleton data of different frames.
[0012] In a second aspect, the present invention provides an action recognition device, including: A dataset construction module, configured to obtain multiple segments of skeleton data sequences, label the behavior types of each segment of skeleton sequence data, and obtain a training set. A model training module, configured to build a multi-stage spatio-temporal fusion network, use the training set as input and the behavior type as output, and train the multi-stage spatio-temporal fusion network to obtain a human action recognition model. An action recognition module, configured to input the skeleton data sequence to be recognized into the human action recognition model to obtain the behavior type corresponding to the skeleton data sequence to be recognized. Wherein, the multi-stage spatio-temporal fusion network includes an action structure graph convolution module, a temporal convolution module, a spatio-temporal tuple attention module, an inter-frame feature aggregation module, and a classification module connected in sequence. The action structure graph convolution module is configured to identify each joint point in each frame of skeleton data and construct the skeletal topological relationship between each joint point to obtain the spatial features between each joint point. The temporal convolution module is configured to extract the temporal features of each joint point in the skeleton data of different frames. The spatio-temporal tuple attention module is configured to perform feature enhancement on the spatial features and the temporal features to obtain enhanced features. The inter-frame feature fusion module is configured to integrate the dynamic information of each joint point in the skeleton data of different frames according to the enhanced features. The classification module is configured to output the behavior type according to the dynamic information.
[0013] In a third aspect, the present invention provides an electronic device, characterized by comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the above-mentioned action recognition method is implemented.
[0014] In a fourth aspect, the present invention provides a computer-readable storage medium, characterized in that the storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned action recognition method is implemented.
[0015] At least one technical solution adopted by the present invention can achieve the following beneficial effects: By introducing a multi-stage processing strategy, the present invention can model the spatio-temporal relationship in the skeleton sequence at both local and global levels, capture the dynamic structural changes between local joint points, and extract global temporal dependencies, effectively improving the recognition accuracy of complex actions. By introducing a multi-stage processing strategy, the model performs spatio-temporal feature fusion at different scales and different abstraction levels, making the action feature expression more complete and clear, especially suitable for long-term action sequences with phased characteristics. While ensuring the model's expression ability, the phased modeling method is adopted to effectively control the consumption of computing resources and the number of parameters, making this method not only applicable to offline batch processing but also having good real-time recognition ability, suitable for deployment in embedded devices or edge computing scenarios. By adopting a structure adaptive mechanism and an attention guidance strategy, the adaptability of the model to action sequences of different individuals, different speeds, and different angles is effectively enhanced, significantly improving the recognition performance of complex and fine-grained actions, and having strong robustness to noise or missing in the skeleton data. Therefore, through the above solutions, the accuracy and timeliness of action recognition for skeleton data can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 It is a flowchart of an action recognition method provided by the present invention; Figure 2 It is a schematic framework diagram of an action recognition method provided by the present invention; Figure 3 It is a schematic structural diagram of an action graph convolutional module provided by the present invention; Figure 4 It is a schematic structural diagram of a spatio-temporal tuple attention module provided by the present invention; Figure 5 It is a schematic structural diagram of an inter-frame feature aggregation module provided by the present invention; Figure 6Schematic diagram of an action recognition device provided by the present invention; Figure 7 Schematic diagram of an electronic device for implementing an action recognition method provided by the present invention. Specific implementation manners
[0017] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Currently, the server mentioned in the present invention can be a server set up on a service platform, or a device such as a desktop computer or a laptop computer that can execute the solution of the present invention. For the convenience of description, the following will only be described with the server as the execution subject. The technical solutions provided by each embodiment of the present invention will be described in detail below with reference to the drawings.
[0019] Refer to Figure 1 , an action recognition method in the present invention specifically includes the following steps: S10: Obtain multiple segments of skeleton data sequences, label the action types of each segment of skeleton sequence data, and obtain a training set.
[0020] In this embodiment, the skeleton data sequence is composed of multiple frames of skeleton data, and each frame of skeleton data includes: The spatial coordinates of the joint points (the three-dimensional coordinates of the key joint points of the human body), usually represented as (x, y, z). Common joint points include: head, shoulder, elbow, wrist, hip, knee, ankle, etc.
[0021] Timestamp: The time mark of each frame of data, used to align the time sequence of the action sequence. The data format is in matrix form, usually represented as R T×V×D , where: T is the number of frames (time dimension), V is the number of joint points, and D is the coordinate dimension. Example: A 100-frame 3D skeleton sequence can be represented as a 100×17×3 matrix (17 joint points, 3D coordinates for each joint point).
[0022] In this embodiment, the acquisition methods of the skeleton data sequence include: 1. Acquisition through sensor devices: Depth cameras (such as Kinect, RealSense), which directly obtain the 3D coordinates of the human skeleton through infrared structured light.
[0023] Inertial Measurement Unit: Sensors worn on various parts of the body that calculate the motion trajectories of joint points through the data fusion of accelerometers, gyroscopes, and magnetometers.
[0024] 2. Acquisition through computer vision algorithms: Such as OpenPose, AlphaPose. Based on monocular / binocular camera videos, convolutional neural networks are used to detect joint points and reconstruct the skeleton. Annotation and preprocessing: Manual annotation: For unannotated video data, the joint point coordinates are manually annotated through tools (such as CrowdPose); Automatic generation: The skeleton data is extracted from the video using a pre-trained model (such as OpenPose). Preprocessing steps: Normalization: Normalize the joint point coordinates relative to the torso (such as the hip) to eliminate individual size differences (Smoothing filtering: Remove noise through Kalman filtering or median filtering; Missing value filling: For missing joint points caused by occlusion, use interpolation methods (such as linear interpolation, prediction based on adjacent frames) to complete).
[0025] S20: Build a multi-stage spatio-temporal fusion network, use the training set as the input, and the behavior type as the output, train the multi-stage spatio-temporal fusion network to obtain a human action recognition model.
[0026] In this embodiment, referring to Figure 2 , the multi-stage spatio-temporal fusion network includes a convolutional module for action structure diagrams, a temporal convolutional module, a spatio-temporal tuple attention module, an inter-frame feature aggregation module, and a classification module connected in sequence. The convolutional module for action structure diagrams is used to identify each joint point in the skeleton data of each frame and construct the skeletal topological relationship between each joint point to obtain the spatial features between each joint point; the temporal convolutional module is used to extract the temporal features of each joint point in the skeleton data of different frames; the spatio-temporal tuple attention module is used to enhance the features of the spatial features and temporal features to obtain enhanced features; the inter-frame feature fusion module is used to integrate the dynamic information of each joint point in the skeleton data of different frames according to the enhanced features; the classification module is used to output the behavior type according to the dynamic information. By constructing a multi-stage spatio-temporal modeling framework, the local spatial structure and global temporal dynamics in the skeleton sequence are respectively extracted and fused for feature extraction.
[0027] Specifically, the skeletal topological structure refers to: the natural connection relationship between joint points (such as "shoulder - elbow - wrist" constitutes the arm bone chain), usually represented by an adjacency matrix. The multi-stage spatio-temporal fusion network includes a local adaptive spatial relationship modeling module and a temporal global dependence modeling module. The former is used to mine the dynamic connection relationship between key nodes, and the latter is used to capture the action evolution law within a long temporal range. Finally, the fused action features are input into the classifier to achieve accurate recognition of skeleton actions, which can effectively improve the recognition accuracy in multi-category, complex background, and fine-grained action recognition scenarios.
[0028] In this embodiment, the action structure graph convolution module includes an action graph convolution module and a behavior graph convolution module.
[0029] The action graph convolution module includes: An encoder for learning the connection probability between any two joint points according to the input skeleton data sequence, and generating an action link probability matrix according to the connection probability between any two joint points; a decoder for predicting the position of the joint point at the next time point according to the action link probability matrix.
[0030] The behavior graph convolution module is used to diffuse the information of the joint points to non-adjacent joint points in a high-power diffusion manner, and then perform structure graph convolution on the skeleton data sequence to obtain a structure link probability matrix. The information of the joint points includes the spatial coordinates of the joint points and the bone topology structure.
[0031] It should be noted that in one or more embodiments of the present application, the meanings of "node" and "joint point" are the same.
[0032] Specifically, referring to Figure 3 , the encoder (Encoder) will learn the connection probability between any two nodes according to the input action sequence (Input ActionSequence) during training, and generate an A-link directly related to the behavior. The decoder (Decoder) infers the node position at the next time point (i.e., the future action Future Action) based on the features generated by the encoder (i.e., the action link A-link).
[0033] Among them, the function of the encoder (Encoder) is as follows: ; Among them, x is the coordinate sequence of the joint points, C is the number of types of action links, A represents the action link probability matrix, and the dimension is n n C , n represents the number of joint points, encode (·) is the encoder function. encode (·) will first extract the features from the 3D joint point information and then convert them into the connection probability between the nodes. In order to extract the connection features, the nodes alternately transmit information between the connections. Represent the i-th node of T frames as: ; Among them, vecDenotes a vectorization operation, that is, flattening a matrix into a column vector in column-major or row-major order. T Denotes the total number of frames of the input skeleton sequence, d Represents the feature dimension (in this embodiment d take 3). At the k th iteration, the feature vector of the i th node is represented as , and the initial node feature p can be represented as: .
[0034] At the k th iteration, the information transfer between Link Features and Joint Features can be expressed as: ; ; Among them, the link feature is a representation of the interaction information between two joint points, used to describe whether there is a potential action association or structural relationship between them. The node feature is the semantic expression of each skeleton joint point itself, including its motion state, local spatial information, and information aggregated from the connections. and are both multi-layer perceptrons, is the concatenated vector, is the operation of aggregating link features and obtaining joint features, such as averaging and element maximization, is the link feature between node k and i at the j +1th iteration, calculated by the multi-layer perceptron. After propagating k times, the encoder outputs the link probability as: ; Among them, r is a random variable, and its elements are i.i.d. sampled from the Gumbel(0, 1) distribution, controls 's discretization, and finally obtains an approximate categorical probability result through the Gumbel softmax.
[0035] The function of the decoder of the action graph convolution module is as follows: ; Among them, refers to the joint point feature at t time, refers to the joint point feature at the initial time, Refers to the joint coordinates at the next moment, A Refers to the action link probability matrix.
[0036] Convert the features of A-link to the joint positions at the next time point ( t +1). If the features of the i th node in the t th frame are defined as: .
[0037] Then decode (.) can be expressed as: ; ; ; ; Among them, , , are all multi-layer perceptrons. Step (a) is to generate features based on A-link, step (b) is to aggregate the features of the previous step onto the corresponding node features, step (c) updates the node features through GRU, and step (d) is to predict the node coordinates at the next time point. Finally, the result is obtained through a Gaussian distribution. Refers to the aggregated features of the i-th node in the t-th frame, Refers to the node state updated through GRU, Refers to the predicted node coordinates in the next frame,
[0038] Step 2, construct the action graph convolution module. The function of the action graph convolution module is as follows: ; Among them, X in Refers to the input feature matrix, Is the graph convolution kernel of the c th type of action, Is the trainable weight used to capture the importance of this feature, Refers to the output feature dimension, Refers to the output features of the action graph convolution.
[0039] Step 3, the function of the structure graph convolution module is as follows: ; Among them, X in Refers to the input feature matrix, L Is the number of diffusion times, P Is the set of skeleton partitions,p ∈ P , is the partition index, is the p power of the partition adjacency matrix of the l th power, is a trainable weight matrix, is a feature transformation matrix.
[0040] Step 4, combine the action graph convolution and behavior graph convolution modules. The action graph convolution module uses the method of high - power diffusion to spread the information of the joint points to non - adjacent joint points, and then performs structural graph convolution (SGC) on the skeleton data sequence to obtain a structure link probability matrix. The function of the action - structure graph convolution module is as follows: ;
[0041] Among them, is a hyperparameter that can control the influence degree of using behavior features.
[0042] Preferably, after obtaining the human action recognition model, perform accuracy detection on the multi - stage spatio - temporal fusion network and evaluate the detection results.
[0043] In this embodiment, the time convolution module uses 1D convolution to extract features from the changes of each joint point in different time frames, such as acceleration, speed change, etc. And further improve the model's ability to capture long - term dependence relationships through dilation convolution.
[0044] Step 6, construct the Figure 4 shown spatio - temporal tuple attention module, Input Map is the input mapping: perform a linear mapping transformation on the original skeleton sequence; Sequence Division is the sequence division: slice the time series into multiple subsequences; Reshape is the deformation: adjust the subsequence structure to a tuple structure; Spatio - Temporal Tuples Encoding is the spatio - temporal tuple encoding: combine multiple joint points at multiple moments into a high - dimensional tuple vector for modeling. The essence of self - attention can be described as a mapping from a query vector to a series of key - value pairs. After completing spatio - temporal tuple encoding and position encoding for the skeleton sequence, use the multi - head self - attention mechanism to model the relationship between input features. Project the encoded sequence through a 1×1 convolutional layer into three groups of vectors: ; Among them, represents the input spatio - temporal feature matrix, ∈RC×T×V ( C is the number of channels, T is the number of frames, V is the number of nodes), calculate the query Q and the similarity of the key K after transposition, and normalize it with Tanh to obtain the normalized attention-weighted feature: ; where C is the number of channels of the key vector, which is used to prevent the gradient from becoming unstable due to an overly large inner product.
[0045] To enhance the representation ability, the multi-head attention mechanism (Multi-headed Self-Attention) is used to enhance the projection effect: Use multiple groups of Q, K, and V with independent parameters for parallel attention. Among them, Q (Query): query vector: used to represent the information to be focused on currently, representing the query vector at each position in the current input sequence, used to find other elements related to itself from the entire sequence. Each attention head will generate its own query vector according to the input sequence, and the Q vector is used to calculate the similarity with the K vector to determine which values (V) should be focused on. K (Key): key vector: represents the features of each element in the input sequence, used to match the query (Q) and calculate the attention weights. The similarity between the key and the query determines the importance of other positions to the current position. Each input element has a corresponding K vector, and the similarity between Q and K determines the importance of this element to the current query. V (Value): value vector: actually contains the information, including the actual information content. Each K vector has a corresponding V vector, and the final output is obtained by weighted summing the V vectors, where the weights are determined by the similarity between Q and K.
[0046] Concatenate the outputs of each group of attention to obtain the concatenated feature as: ; Among them, refers to the output feature of the first attention head. Through 1× convolution, map the concatenated result to the output space ( corresponding to the flattened dimension of the spatio-temporal tuple), and the output of the spatio-temporal tuple attention module is: ; Use 1×1 two-dimensional convolution to implement the feed-forward layer for fusion output.
[0047] Step 7, construct Figure 5The Inter-Frame Feature Aggregation module shown, where Xin and Xout are the input feature and output feature respectively, and Projection to Attention is the input of attention projection, which is used to map the original input feature into three forms: query (Q), key (K), and value (V) for the multi-head self-attention mechanism to perform information matching and weighted calculation. Multi-headed Self-Attention is the multi-head self-attention mechanism, which is used to divide the input feature into multiple sub-spaces and perform self-attention calculations respectively to capture the feature correlations in different sub-spaces. Finally, the outputs of multiple attention heads are fused to enhance the model's ability to express complex spatio-temporal relationships. Projection outAttention is the output of attention projection, which is used to splice the feature of each sub-space output by the multi-head attention mechanism and then map it back to the original feature dimension to integrate the information of different attention heads and facilitate subsequent processing.
[0048] Feed Forward is the feed-forward network, which is used to perform non-linear transformation and dimension mapping on the features of each time step or node to enhance the feature expression ability and the non-linear modeling ability of the model. Inter-Frame Feature Aggregation is the inter-frame feature aggregation module, which is used to fuse the feature information at different times and enhance the model's ability to model the action continuity and global dynamics in the time dimension. Specifically, the inter-frame feature aggregation module is implemented by applying a two-dimensional convolution operation of the convolution kernel in the time dimension, and the inter-frame aggregated feature is: ; ; S30: Input the skeleton data sequence to be recognized into the human action recognition model to obtain the behavior type corresponding to the skeleton data sequence to be recognized.
[0049] Based on Figure 1An action recognition method as shown can model the spatio-temporal relationships in the skeleton sequence at both local and global levels by introducing a multi-stage processing strategy. It can capture the dynamic structural changes between local joint points and extract global temporal dependencies, effectively improving the recognition accuracy of complex actions. By introducing the multi-stage processing strategy, the model performs spatio-temporal feature fusion at different scales and different abstraction levels, making the action feature expression more complete and clear. It is particularly suitable for long-term action sequences with phased characteristics. While ensuring the model's expressive ability, it adopts a phased modeling method to effectively control the consumption of computing resources and the number of parameters, making this method not only applicable to offline batch processing but also having good real-time recognition ability, suitable for deployment in embedded devices or edge computing scenarios. By adopting a structure adaptive mechanism and an attention guidance strategy, it effectively enhances the model's adaptability to action sequences of different individuals, different speeds, and different angles, significantly improving the recognition performance of complex and fine-grained actions and having strong robustness to noise or missing in the skeleton data. Therefore, through the above scheme, the accuracy and timeliness of action recognition for skeleton data can be improved.
[0050] When applying an action recognition method provided by the present invention, it is not necessary to execute according to Figure 1 the order of the steps shown. The specific execution order of each step can be determined as needed, and the present invention does not limit this.
[0051] The above is an action recognition method provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding action recognition device, as Figure 6 shown, including: A dataset construction module, used to obtain multiple segments of skeleton data sequences, label the behavior types of each segment of skeleton sequence data, and obtain a training set.
[0052] A model training module, used to build a multi-stage spatio-temporal fusion network, use the training set as the input, use the behavior type as the output, train the multi-stage spatio-temporal fusion network, and obtain a human action recognition model.
[0053] An action recognition module, used to input the skeleton data sequence to be recognized into the human action recognition model to obtain the behavior type corresponding to the skeleton data sequence to be recognized.
[0054] Among them, the multi-stage spatio-temporal fusion network includes an action structure graph convolution module, a temporal convolution module, a spatio-temporal tuple attention module, an inter-frame feature aggregation module, and a classification module connected in sequence. The action structure graph convolution module is used to identify each joint point in the skeleton data of each frame and construct the skeletal topological relationship between each joint point to obtain the spatial features between each joint point. The temporal convolution module is used to extract the temporal features of each joint point in the skeleton data of different frames. The spatio-temporal tuple attention module is used to enhance the features of the spatial features and the temporal features to obtain enhanced features. The inter-frame feature fusion module is used to integrate the dynamic information of each joint point in the skeleton data of different frames according to the enhanced features. The classification module is used to output the behavior type according to the dynamic information, and this module is composed of a global average pooling (Global Average Pooling, GAP) and a fully connected layer (Fully Connected Layer, FC).
[0055] The present invention also provides a computer-readable storage medium, which stores a computer program that can be used to execute the Figure 1 action recognition method provided above.
[0056] The present invention also provides Figure 7 a schematic structural diagram of the electronic device shown in, as Figure 7 shown. At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the Figure 1 message passing method for the multi-commodity flow problem in the communication network provided above.
[0057] For the specific definition of an action recognition device, reference can be made to the definition of an action recognition method in the above text, which will not be elaborated here. Each module in the action recognition device can be implemented in whole or in part by software, hardware, and their combination. Each module can be embedded in or independent of the processor in the electronic device in the form of hardware, or stored in the memory of the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0058] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the described embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0059] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope recorded by the present invention.
Claims
1. An action recognition method, characterized in that, Including: Obtain multiple segments of skeleton data sequences, label the behavior types of each segment of skeleton sequence data to obtain a training set; Build a multi-stage spatio-temporal fusion network, use the training set as input and the behavior type as output, and train the multi-stage spatio-temporal fusion network to obtain a human action recognition model; Input the skeleton data sequence to be recognized into the human action recognition model to obtain the behavior type corresponding to the skeleton data sequence to be recognized; Wherein, the multi-stage spatio-temporal fusion network includes an action structure graph convolution module, a temporal convolution module, a spatio-temporal tuple attention module, an inter-frame feature aggregation module, and a classification module connected in sequence. The action structure graph convolution module is used to identify each joint point in each frame of skeleton data and construct the skeletal topological relationship between each joint point to obtain the spatial features between each joint point; the temporal convolution module is used to extract the temporal features of each joint point in different frames of skeleton data; the spatio-temporal tuple attention module is used to enhance the features of the spatial features and the temporal features to obtain enhanced features; the inter-frame feature fusion module is used to integrate the dynamic information of each joint point in different frames of skeleton data according to the enhanced features; the classification module is used to output the behavior type according to the dynamic information.
2. The action recognition method according to claim 1, wherein The action structure graph convolution module includes an action graph convolution module and a behavior graph convolution module; The action graph convolution module includes: An encoder, which is used to learn the connection probability between any two joint points according to the input skeleton data sequence and generate an action link probability matrix according to the connection probability between any two joint points; a decoder, which is used to predict the position of the joint point at the next time point according to the action link probability matrix; The behavior graph convolution module is used to diffuse the information of the joint points to non-adjacent joint points in the form of high-power diffusion, and then perform structure graph convolution on the skeleton data sequence to obtain a structure link probability matrix. The information of the joint points includes the spatial coordinates of the joint points and the skeletal topological structure.
3. The action recognition method according to claim 2, characterized in that, The function of the encoder is: ; Among them, x is the coordinate sequence of the joint points, C is the number of types of action links, A represents the action link probability matrix, and the dimension is n n C , n represents the number of joint points, encode (·) is the encoder function.
4. The action recognition method according to claim 2, characterized in that The function of the decoder is: ; Among them, refers to t the joint point feature at time refers to the joint point feature at the initial time, refers to the joint point coordinates at the next time, A refers to the action link probability matrix.
5. The action recognition method according to claim 2, characterized in that, The function of the structure graph convolution is: ; Among them, X refers to the input feature matrix, L is the number of diffusion times, P is the set of skeleton partitions, p ∈ P is the partition index, is the p th power of the l adjacency matrix of the th partition, is the trainable weight matrix, is the feature transformation matrix.
6. The action recognition method according to claim 1, characterized in that, The temporal features include the speed change feature and the acceleration change feature of each joint point in different frames of skeleton data.
7. An action recognition device, characterized in that, Including: A dataset construction module, which is used to obtain multiple segments of skeleton data sequences, label the behavior types of each segment of skeleton sequence data to obtain a training set; A model training module, which is used to build a multi-stage spatio-temporal fusion network, use the training set as input and the behavior type as output, and train the multi-stage spatio-temporal fusion network to obtain a human action recognition model; An action recognition module, which is used to input the skeleton data sequence to be recognized into the human action recognition model to obtain the behavior type corresponding to the skeleton data sequence to be recognized; Among them, the multi-stage spatio-temporal fusion network includes an action structure graph convolution module, a temporal convolution module, a spatio-temporal tuple attention module, an inter-frame feature aggregation module, and a classification module connected in sequence. The action structure graph convolution module is used to identify each joint point in the skeleton data of each frame and construct the bone topology relationship between each joint point to obtain the spatial features between each joint point; the temporal convolution module is used to extract the temporal features of each joint point in the skeleton data of different frames; the spatio-temporal tuple attention module is used to enhance the features of the spatial features and the temporal features to obtain enhanced features; the inter-frame feature fusion module is used to integrate the dynamic information of each joint point in the skeleton data of different frames according to the enhanced features; the classification module is used to output the behavior type according to the dynamic information.
8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements an action recognition method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by the processor, it implements an action recognition method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Human skeleton action recognition method and system and medium
CN110490035A
Motion recognition method and system based on fusion graph convolutional network and Transform network
CN115100574A
Human body behavior recognition method based on GCN-Transform and multi-scale time convolution
CN120148117A
Cited By
Skeleton identification method and system
CN121148012A
Daily life activity ability assessment method and system based on multi-branch space-time fusion network
CN121171479A
Method for constructing, updating and retrieving action memory bank
CN121764980A