An action recognition method based on graph-guided selective scanning
By constructing a learnable adjacency matrix and multi-branch feature extraction structure, combined with graph convolution and multi-scale time convolution, the problem of difficulty in capturing complex joint relationships and spatial structures in the prior art is solved, and higher accuracy of action recognition is achieved.
Patent Information
- Application Number
- CN202411424744.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-12
AI Technical Summary
When existing motion recognition methods deal with human movements, it is difficult to effectively capture complex joint relationships and spatial structure information, resulting in low accuracy.
By constructing a learnable adjacency matrix, combining graph convolution, 2D selective scanning and multi-scale time convolution, multi-level features of joints, bones, and motion are extracted, and feature fusion is performed to achieve action recognition.
Effectively capture dynamic topological relationships, improving the accuracy of recognition of complex actions and the accuracy of feature extraction.
Smart Images

Figure CN119360445B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of action recognition, and in particular to an action recognition method based on graphic-guided selective scanning. Background Art
[0002] Traditional action recognition methods rely on manual feature extraction, mainly representing human motion through the relative positions of joints and translations between body parts. However, such methods are not only time-consuming and labor-intensive, but also have low accuracy. With the development of deep learning, methods based on deep neural networks have gradually replaced manual feature extraction, treating the coordinates of human joints as images or vector sequences. For example, both recurrent neural networks (RNNs) and convolutional neural networks (CNNs) have achieved remarkable results in action recognition, but they are only suitable for processing regular data in Euclidean space, not graph data in non-Euclidean space. Therefore, they tend to ignore the spatial structure information of the human body, resulting in low accuracy. Summary of the invention
[0003] The purpose of the present invention is to overcome the deficiencies of the prior art and provide an action recognition method based on graphic-guided selective scanning.
[0004] The objective of the present invention is achieved through the following technical solutions:
[0005] The present invention provides an action recognition method based on graphic-guided selective scanning, comprising the following steps:
[0006] S1. Data collection and preprocessing: converting the connections between joints and bones in human skeleton data into an adjacency matrix, where each element of the adjacency matrix represents the connection relationship between joint points. By calculating the shortest distance between joints and assigning weights from a trainable parameter table, a learnable adjacency matrix is generated, and its structure is dynamically adjusted during the training process.
[0007] S2, parallel multi-branch feature extraction, dividing the human skeleton data into feature streams, the feature streams include joints, bones and movements, and inputting the joint feature stream, the bone feature stream and the movement feature stream into the model to form three branches, each branch is composed of L basic blocks, each basic block combines graph convolution, 2D selective scanning and multi-scale time convolution to achieve multi-level feature extraction; after being processed by L basic blocks, the features are reduced in dimension through the average pooling layer, and then connected to the fully connected layer for classification and recognition;
[0008] S3, feature fusion, after the features of each branch are extracted, the action recognition results of each branch are fused, and the information extracted by each branch is processed through a fully connected layer or other classifier to output the final classification result.
[0009] Furthermore, the step S1 specifically includes the following sub-steps:
[0010] S11, define the connection between joints and bones, the joint connection includes a joint list , the bone connection relationship includes ;
[0011] S12, construct an adjacency matrix, construct a size of The adjacency matrix A of n is the number of joints, the initial value of which is 1, and then the adjacency matrix A traverses the bone connection relationship , based on the bone connection relationship, mark the corresponding elements in the adjacency matrix as 1;
[0012] S13. Calculate the shortest distance between joints using the Euclidean distance Compute the shortest paths between all joints in the adjacency matrix A, where For joints On the coordinate axis The coordinates on the axis, For joints On the coordinate axis The coordinates on the axis, For joints The coordinate on the z-axis, For joints On the coordinate axis The coordinates on the axis, For joints On the coordinate axis The coordinates on the axis, For joints The coordinate on the z-axis, For joints To the joint The distance between them is based on the shortest path, and the shortest distance matrix D is generated through the adjacency matrix;
[0013] S14. Generate a learnable adjacency matrix and randomly initialize a trainable weight matrix , the weight matrix The size of is the same as the size of the adjacency matrix A, and the weight matrix During the training process, the weight matrix is back-propagated through the loss function. The weights in will be updated according to the gradient, and the weight matrix will be linearly transformed Combined with the shortest distance matrix D, and then normalized to obtain a learnable adjacency matrix , ,in is the first nonlinear activation function.
[0014] Preferably, the step S2 specifically includes the following sub-steps:
[0015] S21. Graph convolution extracts features. Use a learnable adjacency matrix to extract features through graph convolution. Specifically, the learnable adjacency matrix Multiply it with the input feature flow, propagate the feature, and then pass it through the activation function Generate new node features, and the output features after graph convolution are ,in is the input feature, is the learnable weight matrix, is the second nonlinear activation function;
[0016] S22, 2D selective scanning, after graph convolution, the output features Enter the 2D selective scanning module and perform four-way scanning on each joint point in the order of bone number. The four-way scanning includes four directions: up, down, left, and right. For each node , the corresponding initial feature after graph convolution is , in After a 2D selective scan, the aggregate node features are updated as follows: ,in For Node In the The corresponding features after scanning, Respectively represent nodes In the The adjacent node features in four directions in the scan are obtained by the function Aggregate the adjacent node features and then pass them through a nonlinear activation function Improve the model's expressiveness, including is the third nonlinear activation function, It is the output feature, and then the output feature is integrated with the initial feature through the residual connection ,in This is the final output node of the SS2D module. The final output feature matrix is ;
[0017] S23, multi-scale time volume, uses multiple convolution kernels of different sizes to simultaneously extract features in different time ranges, and the size of the convolution kernel is set to Perform temporal convolution on the feature matrix of the final SSD output: , and then fuse the features of different time scales. ,in is the weighting coefficient for each scale.
[0018] Preferably, the step S3 specifically includes: the feature representation of the joint feature stream is , the feature representation of the skeleton feature stream is , the feature representation of the motion feature flow is , and then fuse the features of the three feature streams in a weighted manner ,in is the weight parameter of each feature flow, It is the fused feature representation. After feature fusion, it enters the fully connected layer for action classification ,in is the weight matrix of the classification layer; is bias; is the final classification output, through The function obtains the probability distribution of action categories.
[0019] Preferably, the fusion method in step S3 also includes splicing.
[0020] The beneficial effects of the present invention are:
[0021] 1) The present invention effectively captures dynamic topological relationships by constructing a learnable adjacency matrix, overcoming the limitation of traditional static topology that it is difficult to reflect multiple joint relationships when processing complex movements.
[0022] 2) In the feature extraction process, the present invention adopts a multi-branch structure to process joint, bone and motion features, and extends the convolution operation to non-Euclidean space through graph convolution to capture the complex spatial relationships in the skeleton data.
[0023] 3) The 2D selective scanning of the present invention fully utilizes the graph-guided state-space modeling features and comprehensively considers global and local features to improve the performance and accuracy of feature extraction.
[0024] 4) The application of multi-scale temporal convolution in the present invention enhances the understanding of complex motion patterns and makes the analysis of dynamically changing time series data more effective. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 The figure is a flow chart of a method for motion recognition based on graphic-guided selective scanning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0027] The present invention provides an action recognition method based on graph-guided selective scanning. In the present invention, data preprocessing does not adopt the traditional method of static adjacency matrix, but allocates weights by calculating the shortest distance to construct a learnable adjacency matrix, thereby retaining the topological structure between joints and enhancing the representation ability of joint distance information; in addition, Vmamba's 2D selective scanning module is introduced in each basic block to ensure that each element can integrate information from all other positions to form a global receptive field, which solves the limitation that graph convolution cannot effectively capture long-distance action dependencies, and then captures fast and slow motion information through multi-scale time convolution.
[0028] Illustratively, the schematic flow diagram of the present invention is as follows Figure 1 As shown, the method comprises the following steps: S1, data collection and preprocessing, converting the connection between joints and bones in human skeleton data into an adjacency matrix, wherein each element of the adjacency matrix represents the connection relationship between joint points, calculating the shortest distance between joints, assigning weights from a trainable parameter table, generating a learnable adjacency matrix, and dynamically adjusting its structure during training; step S1 specifically comprises the following sub-steps S11, defining the connection between joints and bones, wherein the joint connection comprises a joint list , the bone connection relationship includes ; S12, construct an adjacency matrix, construct a size of The adjacency matrix A of n is the number of joints, the initial value of which is 1, and then the adjacency matrix A traverses the bone connection relationship , based on the bone connection relationship, mark the corresponding elements in the adjacency matrix as 1; S13, calculate the shortest distance between joints, using the Euclidean distance Compute the shortest paths between all joints in the adjacency matrix A, where For joints On the coordinate axis The coordinates on the axis, For joints On the coordinate axis The coordinates on the axis, For joints The coordinate on the z-axis, For joints On the coordinate axis The coordinates on the axis, For joints On the coordinate axis The coordinates on the axis, For joints The coordinate on the z-axis, For joints To the joint The distance between them is based on the shortest path, and the shortest distance matrix D is generated through the adjacency matrix; S14, generating a learnable adjacency matrix, and randomly initializing a trainable weight matrix , the weight matrix The size of is the same as the size of the adjacency matrix A, and the weight matrix During the training process, the weight matrix is back-propagated through the loss function. The weights in will be updated according to the gradient, and the weight matrix will be linearly transformed Combined with the shortest distance matrix D, and then normalized to obtain a learnable adjacency matrix , ,in is the first nonlinear activation function.
[0029] S2, parallel multi-branch feature extraction, divides the human skeleton data into feature streams, the feature streams include joints, bones and movements, and inputs the joint feature stream, the bone feature stream and the movement feature stream into the model to form three branches respectively. The model is a graph-guided selective scanning model composed of L basic blocks, including graph convolution, 2D selective scanning and multi-scale time convolution, which can process human skeleton data energy-efficiently and extract multiple features; each branch is composed of L basic blocks, and each basic block combines graph convolution, 2D selective scanning and multi-scale time convolution to achieve multi-level feature extraction; after being processed by L basic blocks, the features are reduced in dimension through the average pooling layer, and then connected to the fully connected layer for classification and recognition; specifically includes the following sub-steps: S21, graph convolution extracts features, uses a learnable adjacency matrix, and extracts features through graph convolution, specifically including: converting the learnable adjacency matrix Multiply it with the input feature flow, propagate the feature, and then pass it through the activation function Generate new node features, and the output features after graph convolution are ,in is the input feature, is the learnable weight matrix, is the second nonlinear activation function; S22, 2D selective scanning, after graph convolution, the output features Enter the 2D selective scanning module and perform four-way scanning on each joint point in the order of bone number. The four-way scanning includes four directions: up, down, left, and right. For each node , the corresponding initial feature after graph convolution is , in After a 2D selective scan, the aggregate node features are updated as follows: ,in For Node In the The corresponding features after scanning, Respectively represent nodes In the The adjacent node features in four directions in the scan are obtained by the function Aggregate the adjacent node features to gradually capture more extensive node information in the process of continuous iteration, and then use a nonlinear activation function Improve the model's expressiveness, including is the third nonlinear activation function, It is the output feature, and then the output feature is integrated with the initial feature through the residual connection ,in This is the final output node of the SS2D module. The final output feature matrix is ; S23, multi-scale time volume, using multiple convolution kernels of different sizes to simultaneously extract features within different time ranges, the size of the convolution kernel is set to Perform temporal convolution on the feature matrix of the final SSD output: , and then fuse the features of different time scales. ,in is the weighting coefficient for each scale.
[0030] S3, feature fusion, after the features of each branch are extracted, the action recognition results of each branch are fused, and the information extracted by each branch is processed through a fully connected layer or other classifier to output the final classification result. Specifically, the feature representation of the joint feature flow is: , the feature representation of the skeleton feature stream is , the feature representation of the motion feature flow is , and then fuse the features of the three feature streams in a weighted manner ,in is the weight parameter of each feature flow, It is the fused feature representation. After feature fusion, it enters the fully connected layer for action classification ,in is the weight matrix of the classification layer; is bias; is the final classification output, through The function obtains the probability distribution of the action category. Other ways to perform fusion include splicing.
[0031] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.
Claims
1. An action recognition method based on graphic-guided selective scanning, characterized in that: The following steps are involved: S1. Data collection and preprocessing: converting the connections between joints and bones in human skeleton data into an adjacency matrix, where each element of the adjacency matrix represents the connection relationship between joint points. By calculating the shortest distance between joints and assigning weights from a trainable parameter table, a learnable adjacency matrix is generated, and its structure is dynamically adjusted during the training process. S2, parallel multi-branch feature extraction, dividing the human skeleton data into feature streams, the feature streams include joints, bones and movements, and inputting the joint feature stream, the bone feature stream and the movement feature stream into the model to form three branches respectively. The model includes graph convolution, 2D selective scanning and multi-scale time convolution for feature extraction. Each branch is composed of L basic blocks, and each basic block combines graph convolution, 2D selective scanning and multi-scale time convolution to achieve multi-level feature extraction. After being processed by L basic blocks, the features are reduced in dimension through the average pooling layer, and then connected to the fully connected layer for classification and recognition; S3, feature fusion, after the features of each branch are extracted, the action recognition results of each branch are fused, and the information extracted by each branch is processed through a fully connected layer or other classifier to output the final classification result; The step S2 specifically includes the following sub-steps: S21. Graph convolution extracts features. Use a learnable adjacency matrix to extract features through graph convolution. Specifically, the learnable adjacency matrix A L Multiply it with the input feature flow to propagate the feature, and then generate new node features through the activation function. The output feature after graph convolution is X GCN =σ2(A L XW b ), where X is the input feature, W b is the learnable weight matrix, σ2(·) is the second nonlinear activation function; S22, 2D selective scanning, after graph convolution, the output feature X GCN Enter the 2D selective scanning module, and perform four-directional scanning on each joint point in the order of bone number. The four-directional scanning includes four directions: up, down, left, and right. For each node i, the initial feature corresponding to it after graph convolution is hi. After the lth 2D selective scanning, the aggregate node feature is updated to: hi (l) =f(hi (l-1) ,hleft (l-1) ,hright (l-1) ,hup (l-1) ,hdown (l-1) ), where hi (l-1) is the feature corresponding to node i after the l-1th scan, hleft (l-1) ,hright (l-1) ,hup (l-1) ,hdown (l-1) They represent the neighboring node features of node i in the four directions in the l-1th scan, aggregate the neighboring node features through function f, and then pass the nonlinear activation function hi (l)′ =σ3(hi (l) ) improves the model’s expressiveness, where σ3(·) is the third nonlinear activation function, hi (l)′ It is the output feature, and then the output feature is integrated with the initial feature through the residual connection hifinal= hi (l)′ +hi, where hifinal is the feature of node i finally output by the SS2D module, and the feature matrix H is finally output SS2D =[h1final,h2final,h3final,...,hnfinal]; S23, multi-scale time volume, uses multiple convolution kernels of different sizes to simultaneously extract features in different time ranges, and the size of the convolution kernel is set to k i , perform temporal convolution on the feature matrix of the final output of SSD: Then the features of different time scales are fused. where αi is the weighting coefficient for each scale.
2. The method for motion recognition based on graphic-guided selective scanning according to claim 1, characterized in that: The step S1 specifically includes the following sub-steps: S11, define the connection between joints and bones, the joint connection includes a joint list join = [1, 2, 3 ..., n], and the bone connection relationship includes connections = [(1, 2), (2, 3), (4, 3), (5, 3) ...]; S12, construct an adjacency matrix, construct an adjacency matrix A of size n×n, where n is the number of joints, and the initial value of the number of joints is 1, and then traverse the bone connection relationships connections in the adjacency matrix A, and mark the corresponding elements in the adjacency matrix as 1 based on the bone connection relationships; S13. Calculate the shortest distance between joints using the Euclidean distance Compute the shortest paths between all joints in the adjacency matrix A, where is the coordinate of joint j on the x-axis, is the coordinate of joint j on the y-axis, is the coordinate of joint j on the coordinate axis z; is the coordinate of joint i on the x-axis, is the coordinate of joint i on the y-axis, is the coordinate of joint i on the z-axis; d j,i is the distance between joint j and joint i. Based on the shortest path, the shortest distance matrix D is generated through the adjacency matrix. S14, generate a learnable adjacency matrix and randomly initialize the trainable weight matrix W a , the weight matrix W a The size of is the same as the size of the adjacency matrix A, and the weight matrix W a During the training process, the weight matrix W is back-propagated through the loss function. a The weights in will be updated according to the gradient, and the weight matrix W will be linearly transformed. a Combined with the shortest distance matrix D, and then normalized to obtain the learnable adjacency matrix A L , A L =σ1(HW a ), where σ1(·) is the first nonlinear activation function.
3. The method for motion recognition based on graphic-guided selective scanning according to claim 2, characterized in that: The step S3 specifically includes: the feature representation of the joint feature stream is H joint , the feature representation of the skeleton feature stream is H bone , the feature representation of the motion feature flow is H motion Then, the features of the three feature streams are fused in a weighted manner. fused =W1H joint +W2H bone +W3H motion , where W1, W2, W3 are the weight parameters of each feature flow, H fused It is the fused feature representation. After feature fusion, it enters the fully connected layer for action classification y = Softmax (W c H fused +b), where W c is the weight matrix of the classification layer; b is the bias; y is the final classification output, and the probability distribution of the action category is obtained through the Softmax function.
4. The method for motion recognition based on graphic-guided selective scanning according to claim 3, characterized in that: The fusion method in step S3 also includes splicing.
Citation Information
Patent Citations
Action recognition method and system based on multi-scale space-time theme map convolutional network
CN116453218A
Complex long-range action recognition method based on skeleton space-time diagram convolution
CN117037285A