Human body action recognition method and system based on skeleton spatio-temporal feature fusion graph convolutional network
By adopting a skeleton spatiotemporal feature fusion graph convolution network in human body motion recognition technology, breaking through physical connection constraints, designing a multi-scale time convolution network, and introducing a dual-stream feature fusion mechanism, the problems of non-neighbor joint coordination, motion difference adaptation and information utilization in the existing technology are solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510224984.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
The existing human motion recognition technology has limitations in capturing functional coordination of non-adjacent joints, adapting to short-term motion differences, and making full use of second-order motion information, resulting in insufficient recognition accuracy and robustness.
The method of skeleton spatiotemporal and spatial feature fusion graph convolution network is adopted to break through physical connection constraints through global adaptive spatial graph convolution, and enhance cross-joint functional correlation modeling capabilities; design multi-scale time convolution network to achieve differentiated feature extraction of short-term micro-actions and long-term cycle actions; and innovatively introduce a dual-flow feature fusion mechanism to jointly mine the complementary information of joint coordinates and bone vectors.
It significantly improves the robustness and generalization ability of action recognition, especially when dealing with complex actions, and shows better recognition accuracy and real-time processing performance compared to traditional methods.
Smart Images

Figure CN120164256A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision, and in particular to a method and system for human motion recognition based on a skeleton-based spatiotemporal feature fusion graph convolutional network. Background Art
[0002] Human action recognition technology is a key research area in the field of computer vision. This technology achieves behavior understanding by analyzing the spatiotemporal motion patterns of human joints, and has significant application value in many scenarios such as intelligent monitoring and human-computer interaction. Action recognition technology based on human skeleton features has attracted much attention due to its strong robustness to background changes and lighting interference. Traditional methods usually adopt the spatiotemporal graph convolutional network architecture, the core concept of which is to model the human skeleton as a graph structure, capture the relationship between joints through spatial graph convolution, and combine temporal convolution to extract the temporal characteristics of the action. However, in practical applications, existing methods still have some obvious limitations:
[0003] (1) Limitations of spatial modeling and feature expression: Existing methods rely on predefined physical connections of the skeleton to construct the adjacency matrix, which makes it difficult for them to capture the functional synergy between non-adjacent joints.
[0004] (2) Insufficient adaptability of temporal modeling: Current technologies generally use fixed-scale temporal convolution kernels, which makes it difficult for them to simultaneously meet the differentiated needs of short-term and long-term actions.
[0005] (3) Existing methods have the problem of insufficient information utilization at the feature expression level. Mainstream solutions usually only use joint coordinates as input features, ignoring the supplementary role of second-order motion information such as bone vectors and joint motion speed. Summary of the invention
[0006] The purpose of this application is to provide a human action recognition method and system based on skeleton spatiotemporal feature fusion graph convolutional network, which can effectively improve the recognition accuracy and robustness of complex actions.
[0007] To achieve the above objectives, this application provides the following solutions:
[0008] In a first aspect, the present application provides a method for human action recognition based on a skeleton spatiotemporal feature fusion graph convolutional network, comprising:
[0009] Obtain a video to be detected; the video to be detected is a video containing human body movements.
[0010] The video to be detected is input into the OpenPose posture estimation algorithm to obtain a human skeleton data set of the video to be detected; the human skeleton data set is composed of human skeleton data in each video frame.
[0011] Data preprocessing is performed on the human skeleton data set to obtain skeleton joint information flow, skeleton joint motion information flow, bone information flow and bone motion information flow, and the skeleton joint information flow, skeleton joint motion information flow, bone information flow and bone motion information flow are fused to form a dual-stream network branch; the dual-stream network branch includes a skeleton joint information flow-skeleton joint motion information flow branch and a bone information flow-skeleton motion information flow branch.
[0012] The skeleton joint information flow-skeleton joint motion information flow branch and the skeleton information flow-skeleton motion information flow branch in the dual-stream network branch are respectively input into the skeleton spatiotemporal feature fusion graph convolutional network model to obtain the dual-stream network branch output; the dual-stream network branch output includes the probability distribution of human motion categories under the skeleton joint information flow-skeleton joint motion information flow branch and the probability distribution of human motion categories under the skeleton information flow-skeleton motion information flow branch; the skeleton spatiotemporal feature fusion graph convolutional network model is composed of a batch normalization layer connected in sequence, a number of skeleton spatiotemporal feature fusion graph convolutional networks, a global average pooling layer, a fully connected layer and a Softmax function layer.
[0013] Based on the output of the dual-stream network branches, a weighted fusion method is used to obtain the human action prediction results of the video to be detected.
[0014] Optionally, the skeleton spatiotemporal feature fusion graph convolutional network model includes a 9-layer skeleton spatiotemporal feature fusion graph convolutional network.
[0015] Optionally, for each skeleton spatiotemporal feature fusion graph convolutional network, the skeleton spatiotemporal feature fusion graph convolutional network is composed of a global adaptive spatial graph convolutional network layer, a BN layer, a ReLU activation function layer, a Dropout layer, a multi-scale temporal convolutional network layer, a BN layer, a ReLU activation function and a residual connection structure connected in sequence.
[0016] Optionally, the skeleton joint information flow-skeleton joint motion information flow branch and the skeleton information flow-skeleton motion information flow branch in the dual-stream network branch are respectively input into the skeleton spatiotemporal feature fusion graph convolutional network model to obtain the dual-stream network branch output, specifically including:
[0017] When the input data is the skeleton joint information flow-skeleton joint motion information flow branch, the input data is standardized through a batch normalization layer.
[0018] The standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the first spatiotemporal dimension information; the first spatiotemporal dimension information is the skeleton spatiotemporal features in the spatial dimension and the temporal dimension.
[0019] The first spatiotemporal dimension information is compressed through a global average pooling layer.
[0020] Through the fully connected layer, the compressed first spatiotemporal dimension information is mapped to the action category space.
[0021] The probability distribution of each action category is calculated through the normalized exponential operation of the Softmax function.
[0022] Optionally, the skeleton joint information flow-skeleton joint motion information flow branch and the skeleton information flow-skeleton motion information flow branch in the dual-stream network branch are respectively input into the skeleton spatiotemporal feature fusion graph convolutional network model to obtain the dual-stream network branch output, specifically including:
[0023] When the input data is the skeleton information flow-skeletal motion information flow branch, the input data is standardized through the batch normalization layer.
[0024] The standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the second spatiotemporal dimension information; the second spatiotemporal dimension information is the skeleton spatiotemporal features in the spatial dimension and the temporal dimension.
[0025] The second spatiotemporal dimension information is compressed through a global average pooling layer.
[0026] Through the fully connected layer, the compressed second spatiotemporal dimension information is mapped to the action category space.
[0027] The probability distribution of each action category is calculated through the normalized exponential operation of the Softmax function.
[0028] Optionally, the input data after the normalization process is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the first spatiotemporal dimension information, specifically including:
[0029] According to the skeleton joint information flow-skeleton joint motion information flow branch, based on the global adaptive spatial graph convolutional network, an adaptive mechanism is used to learn the physical and non-physical connection relationships between joints, thereby enhancing the flexibility of spatial modeling of standardized input data.
[0030] The feature distribution is stabilized through the BN layer, the ReLU activation function introduces nonlinear transformation, and the Dropout layer randomly removes features with a probability of 0.5.
[0031] A multi-scale temporal convolutional network layer is used to expand the receptive field through multi-scale dilated convolution technology to extract multi-granularity temporal features.
[0032] BN layer and ReLU layer are used to optimize feature expression.
[0033] The processed features are superimposed on the original input through the residual connection structure.
[0034] Optionally, the standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the second spatiotemporal dimension information, specifically including:
[0035] According to the skeleton information flow-skeleton motion information flow branch, based on the global adaptive spatial graph convolutional network, an adaptive mechanism is used to learn the physical and non-physical connection relationships between joints, thereby enhancing the flexibility of spatial modeling of standardized input data.
[0036] The feature distribution is stabilized through the BN layer, the ReLU activation function introduces nonlinear transformation, and the Dropout layer randomly removes features with a probability of 0.5.
[0037] A multi-scale temporal convolutional network layer is used to expand the receptive field through multi-scale dilated convolution technology to extract multi-granularity temporal features.
[0038] BN layer and ReLU layer are used to optimize feature expression.
[0039] The processed features are superimposed on the original input through the residual connection structure.
[0040] Optionally, the processing steps of the global adaptive spatial graph convolutional network layer specifically include:
[0041] According to formula A k =A physical +A loop , calculate the skeleton physical connection matrix; where A physical is the physical connection adjacency matrix, A loop is the identity matrix.
[0042] Initialize an enhanced mask matrix; the enhanced mask matrix is a learnable common correlation matrix used to capture stable feature patterns across samples through global parameter optimization.
[0043] According to the formula Compute the spatially adaptive adjacency matrix.
[0044] The output features of the global adaptive spatial graph convolutional network are determined based on the skeleton physical connection matrix, the enhanced mask matrix and the spatial domain adaptive adjacency matrix.
[0045] Optionally, the output features of the global adaptive spatial graph convolutional network are determined according to the skeleton physical connection matrix, the enhanced mask matrix and the spatial adaptive adjacency matrix, specifically including:
[0046] According to the formula Determine the output features of the nth layer of the globally adaptive spatial graph convolutional network.
[0047] Among them, f in is the input feature, W k is the spatial convolution kernel weight, K is the number of convolution kernels, λ is the dynamic proportional coefficient that gradually decreases with the training rounds, and A k is the skeleton physical connection matrix, B k is the enhanced mask matrix, C k is the spatial domain adaptive adjacency matrix.
[0048] In the second aspect, the present application provides a human action recognition system based on a skeleton spatiotemporal feature fusion graph convolutional network, comprising:
[0049] The video acquisition module is used to acquire the video to be detected; the video to be detected is a video containing human body movements.
[0050] The data extraction module is used to input the video to be detected into the OpenPose posture estimation algorithm to obtain a human skeleton data set of the video to be detected; the human skeleton data set is composed of human skeleton data in each video frame.
[0051] The data processing module is used to perform data preprocessing on the human skeleton data set to obtain skeleton joint information flow, skeleton joint motion information flow, bone information flow and bone motion information flow, and fuse the skeleton joint information flow, skeleton joint motion information flow, bone information flow and bone motion information flow to form a dual-stream network branch; the dual-stream network branch includes a skeleton joint information flow-skeleton joint motion information flow branch and a bone information flow-skeleton motion information flow branch.
[0052] An input module is used to input the skeleton joint information flow-skeleton joint movement information flow branch and the skeleton information flow-skeleton movement information flow branch in the dual-stream network branch into the skeleton spatiotemporal feature fusion graph convolutional network model respectively to obtain a dual-stream network branch output; the dual-stream network branch output includes the probability distribution of human motion categories under the skeleton joint information flow-skeleton joint movement information flow branch and the probability distribution of human motion categories under the skeleton information flow-skeleton movement information flow branch; the skeleton spatiotemporal feature fusion graph convolutional network model is composed of a batch normalization layer connected in sequence, a plurality of skeleton spatiotemporal feature fusion graph convolutional networks, a global average pooling layer, a fully connected layer and a Softmax function layer.
[0053] The output module is used to obtain the human action prediction result of the video to be detected based on the dual-stream network branch output and the weighted fusion method.
[0054] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0055] This application provides a method and system for human motion recognition based on a skeleton spatiotemporal feature fusion graph convolution network. It breaks through physical connection constraints by constructing a global adaptive spatial graph convolution and enhances the ability to model cross-joint functional associations. It designs a multi-scale temporal convolution network to achieve differentiated feature extraction of short-term micro-motions and long-term periodic motions. It innovatively introduces a dual-stream feature fusion mechanism to jointly mine the complementary information of joint coordinates and skeleton vectors. This application significantly improves the robustness and generalization ability of motion recognition, especially when processing complex motions, and shows better recognition accuracy and real-time processing performance than traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0057] Figure 1 A flowchart of a method for human motion recognition based on a skeleton-based spatiotemporal feature fusion graph convolutional network is provided for one embodiment of the present application.
[0058] Figure 2 A schematic diagram of the overall structure of the skeleton spatiotemporal feature fusion graph convolutional network model provided in one embodiment of the present application.
[0059] Figure 3 A schematic diagram of the internal structure of a skeleton spatiotemporal feature fusion graph convolutional network layer provided in one embodiment of the present application.
[0060] Figure 4 A schematic diagram of the global adaptive spatial graph convolutional network structure provided in one embodiment of the present application.
[0061] Figure 5 A schematic diagram of the functional modules of a human motion recognition device based on a skeleton-based spatiotemporal feature fusion graph convolutional network provided in one embodiment of the present application. DETAILED DESCRIPTION
[0062] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0063] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0064] Embodiment 1
[0065] like Figure 1 As shown, this embodiment provides a method for human action recognition based on a skeleton spatiotemporal feature fusion graph convolutional network, including:
[0066] Step 101: Obtain a video to be detected; the video to be detected is a video containing human body movements.
[0067] Step 102: Input the video to be detected into the OpenPose posture estimation algorithm to obtain a human skeleton data set of the video to be detected; the human skeleton data set is composed of human skeleton data in each video frame.
[0068] Step 103: Preprocess the human skeleton data set to obtain skeleton joint information flow, skeleton joint motion information flow, bone information flow and bone motion information flow, and fuse the skeleton joint information flow, skeleton joint motion information flow, bone information flow and bone motion information flow to form a dual-stream network branch; the dual-stream network branch includes a skeleton joint information flow-skeleton joint motion information flow branch and a bone information flow-skeleton motion information flow branch.
[0069] Step 104: The skeleton joint information flow-skeleton joint motion information flow branch and the skeleton information flow-skeleton motion information flow branch in the dual-stream network branch are respectively input into the skeleton spatiotemporal feature fusion graph convolutional network model to obtain the dual-stream network branch output; the dual-stream network branch output includes the probability distribution of human motion categories under the skeleton joint information flow-skeleton joint motion information flow branch and the probability distribution of human motion categories under the skeleton information flow-skeleton motion information flow branch; the skeleton spatiotemporal feature fusion graph convolutional network model is composed of a batch normalization layer connected in sequence, a plurality of skeleton spatiotemporal feature fusion graph convolutional networks, a global average pooling layer, a fully connected layer and a Softmax function layer.
[0070] Step 105: Based on the output of the dual-stream network branches, a weighted fusion method is used to obtain a human motion prediction result of the video to be detected.
[0071] In some embodiments, when executing step 103, the specific steps may be as follows:
[0072] The processing steps of the skeleton spatiotemporal feature fusion graph convolutional network are as follows (the specific structure diagram is as follows Figure 2 shown):
[0073] 1) The two-stream network branch is used as the input data of the network. First, the input data is standardized through the batch normalization layer (BatchNormalization, BN).
[0074] 2) The data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the skeleton spatiotemporal features in the spatial and temporal dimensions.
[0075] 3) Compress the spatiotemporal dimension information through the global average pooling layer (GAP).
[0076] 4) It is further mapped to the action category space through the fully connected layer (FC).
[0077] 5) The probability distribution of each action category is calculated through the normalized exponential operation of the Softmax function to complete the classification decision.
[0078] Specifically, when the input data is a skeleton joint information flow-skeleton joint motion information flow branch, the input data is standardized through a batch normalization layer.
[0079] The standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the first spatiotemporal dimension information; the first spatiotemporal dimension information is the skeleton spatiotemporal features in the spatial dimension and the temporal dimension.
[0080] The first spatiotemporal dimension information is compressed through a global average pooling layer.
[0081] Through the fully connected layer, the compressed first spatiotemporal dimension information is mapped to the action category space.
[0082] The probability distribution of each action category is calculated through the normalized exponential operation of the Softmax function.
[0083] Specifically, when the input data is the skeleton information flow-skeletal motion information flow branch, the input data is standardized through a batch normalization layer.
[0084] The standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the second spatiotemporal dimension information; the second spatiotemporal dimension information is the skeleton spatiotemporal features in the spatial dimension and the temporal dimension.
[0085] The second spatiotemporal dimension information is compressed through a global average pooling layer.
[0086] Through the fully connected layer, the compressed second spatiotemporal dimension information is mapped to the action category space.
[0087] The probability distribution of each action category is calculated through the normalized exponential operation of the Softmax function.
[0088] Wherein, in some embodiments, Figure 3As shown, for each skeleton spatiotemporal feature fusion graph convolutional network, the skeleton spatiotemporal feature fusion graph convolutional network is composed of a global adaptive spatial graph convolutional network layer, a BN layer, a ReLU activation function layer, a Dropout layer, a multi-scale time convolutional network layer, a BN layer, a ReLU activation function and a residual connection structure connected in sequence.
[0089] The processing steps inside the skeleton spatiotemporal feature fusion graph convolutional network layer are as follows:
[0090] 1) The input features pass through the global adaptive spatial graph convolutional network, i.e., the improved GCN module, which uses an adaptive mechanism to learn the physical and non-physical connection relationships between joints, breaking through the fixed topology limitations and enhancing the flexibility of spatial modeling.
[0091] 2) The BN layer is used to stabilize the feature distribution, the ReLU activation function introduces nonlinear transformation, and the Dropout layer randomly removes features with a probability of 0.5 to avoid overfitting.
[0092] 3) Through the multi-scale temporal convolutional network, that is, the improved TCN module, multi-scale dilated convolution is used to expand the receptive field and extract multi-granularity temporal features.
[0093] 4) The feature expression is optimized again through BN and ReLU layers.
[0094] 5) The processed features are superimposed on the original input through the residual connection structure, which retains the low-order features while promoting gradient propagation.
[0095] Specifically, according to the skeleton joint information flow-skeleton joint motion information flow branch, based on the global adaptive spatial graph convolutional network, an adaptive mechanism is used to learn the physical and non-physical connection relationships between joints, thereby enhancing the flexibility of spatial modeling of standardized input data.
[0096] The feature distribution is stabilized through the BN layer, the ReLU activation function introduces nonlinear transformation, and the Dropout layer randomly removes features with a probability of 0.5.
[0097] A multi-scale temporal convolutional network layer is used to expand the receptive field through multi-scale dilated convolution technology to extract multi-granularity temporal features.
[0098] BN layer and ReLU layer are used to optimize feature expression.
[0099] The processed features are superimposed on the original input through the residual connection structure.
[0100] Specifically, the standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the second spatiotemporal dimension information, including:
[0101] According to the skeleton information flow-skeleton motion information flow branch, based on the global adaptive spatial graph convolutional network, an adaptive mechanism is used to learn the physical and non-physical connection relationships between joints, thereby enhancing the flexibility of spatial modeling of standardized input data.
[0102] The feature distribution is stabilized through the BN layer, the ReLU activation function introduces nonlinear transformation, and the Dropout layer randomly removes features with a probability of 0.5.
[0103] A multi-scale temporal convolutional network layer is used to expand the receptive field through multi-scale dilated convolution technology to extract multi-granularity temporal features.
[0104] BN layer and ReLU layer are used to optimize feature expression.
[0105] The processed features are superimposed on the original input through the residual connection structure.
[0106] In some embodiments, the processing steps of the global adaptive spatial graph convolutional network layer are as follows (the specific structure diagram is as follows Figure 4 shown):
[0107] 1) Calculate the skeleton physical connection matrix A k:
[0108] Skeleton physical connection matrix A k It is the topological basis of the global adaptive spatial graph convolutional network. Its essence is an undirected graph adjacency matrix based on the physical structure of the human body. The dimension of this matrix is determined by the number of joints N in the data set, and the parameters are kept consistent between different time frames, reflecting the stability of the physical connection relationship of the skeleton joints during movement. Specifically, the skeleton graph constructs an undirected graph structure by abstracting the joints as vertices and the bones as edges. Its adjacency matrix is defined as:
[0109]
[0110] Since the physical connection structure of the skeleton joints does not change in the action sequence, the parameters of the skeleton graph between different frames are exactly the same. This matrix can clearly reflect the basic physical connection relationship of the human skeleton, which is consistent with the N×N adjacency matrix defined in ST-GCN, and is not affected by sample differences or training processes, and strictly follows the physical connection rules of the human skeleton. In order to enhance the ability to retain the node's own features during the propagation process, the self-connection matrix Aloop=I is further introduced. N×N , forming a complete skeleton physical connection matrix:
[0111] A k =A physical +A loop .
[0112] Among them, A physical is the physical connection adjacency matrix, A loop is the identity matrix.
[0113] 2) Initialize the enhanced mask matrix B k :
[0114] Enhanced mask matrix B k As a learnable common correlation matrix, it captures stable feature patterns across samples through global parameter optimization and can dynamically adjust the strength of physical connections. k Initialized to A k Same value.
[0115] 3) Calculate the spatial adaptive adjacency matrix C k :
[0116] Spatial adaptive adjacency matrix C k As a sample-specific association matrix, it is generated by calculating sample data and focuses on mining the dynamic functional associations between joints in the current data. The similarity between each pair of joints is calculated using a normalized Gaussian function:
[0117]
[0118] 4) The three matrices are weighted and fused to obtain the output features of the n-th layer global adaptive spatial graph convolutional network:
[0119]
[0120] where f in is the input feature, W k is the spatial convolution kernel weight, K is the number of convolution kernels, and λ is the dynamic proportional coefficient that gradually decreases with the training rounds.
[0121] 5) Through the residual connection structure, the input features are directly superimposed on the output features.
[0122] Among them, the processing steps of the multi-scale temporal convolutional network layer are as follows:
[0123] 1) Divide the input features into six parallel branches.
[0124] 2) Each branch is processed through a 1×1 convolution kernel to adjust the channel dimension and reduce computational complexity.
[0125] 3) The first four branches use a 3×1 dilated convolution kernel with dilation rates set to 1, 2, 3, and 4 respectively, and extract multi-scale temporal features from the skeleton data at different sampling rates; the fifth branch uses a 3×1 maximum pooling operation to compress redundant frame information and enhance sensitivity to significant motion features; the sixth branch, as an identity mapping branch, retains the original feature information and prevents information loss during training.
[0126] 4) By splicing the outputs of multiple branches in the channel dimension, the fusion of multi-scale temporal feature information is achieved:
[0127] Y = Concat(Y d1 ,Y d2 ,Y d3 ,Y d4 ,Y MaxPool ,Y Identity ).
[0128] Among them, Concat represents the concatenation operation, Y represents the multi-scale feature output, and Y d1~Yd4 represents the output of the dilated convolution branch with a dilation rate of 1 to 4, Y MaxPool is the maximum pooling branch output, Y Identity It is the output of the identity mapping branch.
[0129] 5) The multi-scale feature output is superimposed on the original input through the residual connection mechanism to alleviate the gradient vanishing problem in deep networks.
[0130] Embodiment 2
[0131] like Figure 5 As shown, this embodiment provides a human action recognition system based on skeleton spatiotemporal feature fusion graph convolutional network, including:
[0132] The video acquisition module 501 is used to acquire a video to be detected; the video to be detected is a video containing human body movements.
[0133] The data extraction module 502 is used to input the video to be detected into the OpenPose posture estimation algorithm to obtain a human skeleton data set of the video to be detected; the human skeleton data set is composed of human skeleton data in each video frame.
[0134] The data processing module 503 is used to perform data preprocessing on the human skeleton data set to obtain the skeleton joint information flow, the skeleton joint motion information flow, the bone information flow and the bone motion information flow, and fuse the skeleton joint information flow, the skeleton joint motion information flow, the bone information flow and the bone motion information flow to form a dual-stream network branch; the dual-stream network branch includes the skeleton joint information flow-skeleton joint motion information flow branch and the bone information flow-skeleton motion information flow branch.
[0135] The input module 504 is used to input the skeleton joint information flow-skeleton joint motion information flow branch and the skeleton information flow-skeleton motion information flow branch in the dual-stream network branch into the skeleton spatiotemporal feature fusion graph convolutional network model respectively to obtain the dual-stream network branch output; the dual-stream network branch output includes the probability distribution of human motion categories under the skeleton joint information flow-skeleton joint motion information flow branch and the probability distribution of human motion categories under the skeleton information flow-skeleton motion information flow branch; the skeleton spatiotemporal feature fusion graph convolutional network model is composed of a batch normalization layer connected in sequence, a plurality of skeleton spatiotemporal feature fusion graph convolutional networks, a global average pooling layer, a fully connected layer and a Softmax function layer.
[0136] The output module 505 is used to obtain the human motion prediction result of the video to be detected by adopting a weighted fusion method based on the dual-stream network branch output.
[0137] In summary, this application has the following technical effects:
[0138] This application provides a method and system for human motion recognition based on a skeleton-based spatiotemporal feature fusion graph convolutional network. With the goal of improving the accuracy of motion recognition in complex scenarios, this application makes systematic improvements to the limitations of existing spatiotemporal graph convolutional networks. By constructing a global adaptive spatial graph convolution, it breaks through physical connection constraints and enhances the ability to model cross-joint functional associations; designs a multi-scale temporal convolutional network to achieve differentiated feature extraction of short-term micro-motions and long-term periodic motions; and innovatively introduces a dual-stream feature fusion mechanism to jointly mine the complementary information of joint coordinates and skeleton vectors. This method significantly improves the robustness and generalization ability of motion recognition, especially when dealing with complex motions, and exhibits better recognition accuracy and real-time processing performance than traditional methods.
[0139] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0140] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A human action recognition method based on skeleton spatiotemporal feature fusion graph convolutional network, characterized in that: The human motion recognition method comprises: Obtain a video to be detected; the video to be detected is a video containing human body movements; Input the video to be detected into the OpenPose posture estimation algorithm to obtain a human skeleton data set of the video to be detected; the human skeleton data set is composed of human skeleton data in each video frame; Data preprocessing is performed on a human skeleton data set to obtain a skeleton joint information flow, a skeleton joint motion information flow, a bone information flow, and a bone motion information flow, and the skeleton joint information flow, the skeleton joint motion information flow, the bone information flow, and the bone motion information flow are fused to form a dual-stream network branch; the dual-stream network branch includes a skeleton joint information flow-skeleton joint motion information flow branch and a bone information flow-skeleton motion information flow branch; The skeleton joint information flow-skeleton joint motion information flow branch and the skeleton information flow-skeleton motion information flow branch in the dual-stream network branch are respectively input into the skeleton spatiotemporal feature fusion graph convolutional network model to obtain the dual-stream network branch output; the dual-stream network branch output includes the probability distribution of human motion categories under the skeleton joint information flow-skeleton joint motion information flow branch and the probability distribution of human motion categories under the skeleton information flow-skeleton motion information flow branch; the skeleton spatiotemporal feature fusion graph convolutional network model is composed of a batch normalization layer, a plurality of skeleton spatiotemporal feature fusion graph convolutional networks, a global average pooling layer, a fully connected layer and a Softmax function layer connected in sequence; Based on the output of the dual-stream network branches, a weighted fusion method is used to obtain the human action prediction results of the video to be detected.
2. According to claim 1, a human action recognition method based on skeleton spatiotemporal feature fusion graph convolutional network is characterized in that: The skeleton spatiotemporal feature fusion graph convolutional network model includes 9 layers of skeleton spatiotemporal feature fusion graph convolutional networks.
3. According to claim 2, a human action recognition method based on skeleton spatiotemporal feature fusion graph convolutional network is characterized in that: For each skeleton spatiotemporal feature fusion graph convolutional network, the skeleton spatiotemporal feature fusion graph convolutional network is composed of a global adaptive spatial graph convolutional network layer, a BN layer, a ReLU activation function layer, a Dropout layer, a multi-scale temporal convolutional network layer, a BN layer, a ReLU activation function and a residual connection structure connected in sequence.
4. According to claim 3, a human action recognition method based on skeleton spatiotemporal feature fusion graph convolutional network is characterized in that: The skeleton joint information flow-skeleton joint motion information flow branch and the skeleton information flow-skeleton motion information flow branch in the two-stream network branch are respectively input into the skeleton spatiotemporal feature fusion graph convolutional network model to obtain the two-stream network branch output, which specifically includes: When the input data is the skeleton joint information flow-skeleton joint motion information flow branch, the input data is standardized through the batch normalization layer; The standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the first spatiotemporal dimension information; the first spatiotemporal dimension information is the skeleton spatiotemporal features in the spatial dimension and the temporal dimension; Compress the first spatiotemporal dimension information through a global average pooling layer; Through the fully connected layer, the compressed first spatiotemporal dimension information is mapped to the action category space; The probability distribution of each action category is calculated through the normalized exponential operation of the Softmax function.
5. According to claim 4, a method for human action recognition based on skeleton spatiotemporal feature fusion graph convolutional network is characterized in that: The skeleton joint information flow-skeleton joint motion information flow branch and the skeleton information flow-skeleton motion information flow branch in the two-stream network branch are respectively input into the skeleton spatiotemporal feature fusion graph convolutional network model to obtain the two-stream network branch output, which specifically includes: When the input data is the skeleton information flow-skeletal motion information flow branch, the input data is standardized through the batch normalization layer; The standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the second spatiotemporal dimension information; the second spatiotemporal dimension information is the skeleton spatiotemporal features in the spatial dimension and the temporal dimension; Compress the second spatiotemporal dimension information through a global average pooling layer; Through the fully connected layer, the compressed second spatiotemporal dimension information is mapped to the action category space; The probability distribution of each action category is calculated through the normalized exponential operation of the Softmax function.
6. The method for human action recognition based on skeleton spatiotemporal feature fusion graph convolutional network according to claim 5 is characterized in that: The standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the first spatiotemporal dimension information, including: According to the skeleton joint information flow-skeleton joint motion information flow branch, based on the global adaptive spatial graph convolutional network, an adaptive mechanism is used to learn the physical and non-physical connection relationships between joints, enhancing the flexibility of spatial modeling of standardized input data; The feature distribution is stabilized through the BN layer, the ReLU activation function introduces nonlinear transformation, and the Dropout layer randomly removes features with a probability of 0.5; Adopting multi-scale temporal convolutional network layers, the multi-scale dilated convolution technique is used to expand the receptive field and extract multi-granularity temporal features. Use BN layer and ReLU layer to optimize feature expression; The processed features are superimposed on the original input through the residual connection structure.
7. The method for human action recognition based on skeleton spatiotemporal feature fusion graph convolutional network according to claim 6 is characterized in that: The standardized input data is sequentially passed through a 9-layer skeleton spatiotemporal feature fusion graph convolutional network to extract the second spatiotemporal dimension information, including: According to the skeleton information flow-skeleton motion information flow branch, based on the global adaptive spatial graph convolutional network, an adaptive mechanism is used to learn the physical and non-physical connection relationships between joints, enhancing the flexibility of spatial modeling of standardized input data; The feature distribution is stabilized through the BN layer, the ReLU activation function introduces nonlinear transformation, and the Dropout layer randomly removes features with a probability of 0.5; Adopting multi-scale temporal convolutional network layers, the multi-scale dilated convolution technique is used to expand the receptive field and extract multi-granularity temporal features. Use BN layer and ReLU layer to optimize feature expression; The processed features are superimposed on the original input through the residual connection structure.
8. The method for human action recognition based on skeleton spatiotemporal feature fusion graph convolutional network according to claim 7 is characterized in that: The processing steps of the global adaptive spatial graph convolutional network layer specifically include: According to formula A k =A physical +A loop , calculate the skeleton physical connection matrix; where A physical is the physical connection adjacency matrix, A loop is the identity matrix; Initializing an enhanced mask matrix; the enhanced mask matrix is a learnable common correlation matrix used to capture stable feature patterns across samples through global parameter optimization; According to the formula Compute the spatial adaptive adjacency matrix; The output features of the global adaptive spatial graph convolutional network are determined based on the skeleton physical connection matrix, the enhanced mask matrix and the spatial domain adaptive adjacency matrix.
9. The method for human action recognition based on skeleton spatiotemporal feature fusion graph convolutional network according to claim 8, characterized in that: According to the skeleton physical connection matrix, the enhanced mask matrix and the spatial domain adaptive adjacency matrix, the output features of the global adaptive spatial graph convolutional network are determined, including: According to the formula Determine the output features of the n-th layer of the global adaptive spatial graph convolutional network; Among them, f in is the input feature, W k is the spatial convolution kernel weight, K is the number of convolution kernels, λ is the dynamic proportional coefficient that gradually decreases with the training rounds, and A k is the skeleton physical connection matrix, B k is the enhanced mask matrix, C k is the spatial domain adaptive adjacency matrix.
10. A human action recognition system based on skeleton spatiotemporal feature fusion graph convolutional network, characterized in that: include: A video acquisition module is used to acquire a video to be detected; the video to be detected is a video containing human body movements; A data extraction module is used to input the video to be detected into the OpenPose posture estimation algorithm to obtain a human skeleton data set of the video to be detected; the human skeleton data set is composed of human skeleton data in each video frame; A data processing module is used to perform data preprocessing on a human skeleton data set to obtain a skeleton joint information flow, a skeleton joint motion information flow, a bone information flow, and a bone motion information flow, and fuse the skeleton joint information flow, the skeleton joint motion information flow, the bone information flow, and the bone motion information flow to form a dual-stream network branch; The dual-stream network branch includes a skeleton joint information flow-skeleton joint motion information flow branch and a skeleton information flow-skeleton motion information flow branch; An input module is used to input the skeleton joint information flow-skeleton joint motion information flow branch and the skeleton information flow-skeleton motion information flow branch in the dual-stream network branch into the skeleton spatiotemporal feature fusion graph convolutional network model respectively to obtain a dual-stream network branch output; the dual-stream network branch output includes the probability distribution of human motion categories under the skeleton joint information flow-skeleton joint motion information flow branch and the probability distribution of human motion categories under the skeleton information flow-skeleton motion information flow branch; the skeleton spatiotemporal feature fusion graph convolutional network model is composed of a batch normalization layer, a plurality of skeleton spatiotemporal feature fusion graph convolutional networks, a global average pooling layer, a fully connected layer and a Softmax function layer connected in sequence; The output module is used to obtain the human action prediction result of the video to be detected based on the dual-stream network branch output and the weighted fusion method.
Citation Information
Patent Citations
Method for constructing human body behavior recognition model based on graph convolution network
CN111652124A
Human body behavior recognition method based on double-flow dynamic characteristics
CN116386131A
Skeleton action recognition method based on space-time adaptive feature fusion graph convolutional network
CN116665300A
Fitness action recognition method based on multi-branch fusion graph convolutional network
CN119296177A
Human-robot collaboration method based on multi-scale graph convolutional neural network
US12159486B1
Cited By
Water supply network global water quality prediction method based on double flow-graph convolutional network
CN120873478A