A Fine-grained Skeleton Action Recognition Method Based on Action-responsive Contrastive Network
By applying an action-responsive comparison network on the skeleton data, a multi-channel cross-time dynamic skeleton joint attention topology is built, and feature comparison learning is used to solve the problem of extracting skeleton data features in complex backgrounds in the existing technology, and high-accuracy recognition of fine actions is achieved.
Patent Information
- Application Number
- CN202411112171.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-08-14
AI Technical Summary
The prior art is difficult to extract rich features of skeleton data in complex contexts, and graph convolutional networks have problems of information loss and expression ability limitations when identifying fine actions.
A skeleton fine motion recognition method based on action-responsive comparison network is proposed. Through action-responsive topology graph convolution network and fine motion contrast device, a multi-channel cross-time dynamic skeleton joint attention topology is constructed, and a feature comparison learning is used to explore the potential space of fine-grained actions.
It effectively overcomes the limitations of a single topological structure, improves the ability to recognize fine-grained movements in the human body, alleviates the problems of excessive discrete and blurred boundaries of fine-motion feature representations, and significantly improves the accuracy of fine-motion recognition.
Smart Images

Figure CN119131889B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly relates to a fine-grained action recognition method for skeletons based on an action-responsive contrast network. Background Art
[0002] Human communication is the cornerstone of society, and daily behaviors convey emotions and intentions that affect individuals and society. With the development of the intelligent society, action recognition is increasingly widely applied in fields such as human-computer interaction, video surveillance, and abnormal behavior detection, and the requirements for accuracy and details are also constantly increasing. One of the main challenges in action recognition is to extract rich features in complex backgrounds. Skeleton data has received extensive attention due to its robust and stable action features.
[0003] Early methods mainly used recurrent neural networks (RNNs) or convolutional neural networks (CNNs) to analyze joint features, but these methods often ignored the structured representation of skeletons. Skeleton data naturally forms a topological graph structure. Therefore, graph convolutional networks (GCNs) were introduced to model skeleton sequences. GCNs can effectively capture key skeleton action features by using the topological structure of skeletons, and this graph-based model has been widely accepted and applied.
[0004] However, everything has two sides. The graph convolution technology can make full use of the advantages of skeleton data, which has rich motion features and strong robustness to complex background information. However, the price is the loss of strongly correlated background information of actions and the limitation of the predefined single fixed topological structure. It only establishes the correlation between naturally connected nodes, while the implicit relationship between non-naturally connected nodes during the occurrence of actions is ignored, which will lead to the loss of key information during graph convolution, thus restricting the expressive power of the graph convolution network and making it more difficult to recognize fine-grained actions. Skeleton actions with very small motion pattern differences, such as "reading" and "writing", are difficult to be perceived and utilized by general graph convolution models based on skeleton data alone. Their features are almost mixed together and cannot be distinguished. Given that each channel represents different types of dynamic features and the interaction between joints is not constant under different action characteristics, it is not an ideal choice to uniformly apply a single topology. Some researchers have proposed to establish a topology refinement model to reflect the motion patterns of different actions. Although building a dynamic topology in the graph convolution layer alleviates the limitation of the single topological structure to a certain extent, it only explores the overall spatial structure of different motion patterns, which is very effective for actions with significant motion pattern differences, but the expression of the temporal features of different motion patterns learned from continuous actions is weak, and the recognition effect for fine-grained actions is lacking. Secondly, for different action patterns, especially when facing fine-grained actions, it becomes more important to focus on the multi-dimensional feature fusion of important joints. For example, using a spatio-temporal transformer network, the transformer self-attention operator is used through a two-stream network to model the dependencies between joints. However, if the joint spatial and temporal features are separately processed and then feature fusion is performed, the interactivity of the spatio-temporal information of some joints may be reduced. Taking "reading" and "writing" as an example again, the joint features they learn independently in space and time are extremely similar, but the temporal feature expression under the similar spatial structure can reflect the subtle differences in their joint features. In addition, due to the over-discrete and blurred boundary of fine-grained action feature representation, although there are some research works on fine-grained action features, the research in this area is still limited. In order to fully explore the tiny feature differences between fine-grained actions, it is crucial to explore a motion implicit space that is intra-class compact and inter-class dispersed for refined motion features. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for skeleton fine-grained action recognition based on an action-responsive contrast network. The present invention provides an action-responsive contrast network for the task of skeleton fine-grained action recognition. Through feature fusion at different spatio-temporal scales, the invention can learn the unique bone topologies of different fine-grained actions, and use feature contrast learning to construct a learnable latent space to capture the subtle feature differences of fine-grained actions, thereby improving the ability to recognize human fine-grained actions and providing strong support for downstream tasks such as human-computer interaction and intelligent monitoring.
[0006] The inventive concept of the present invention is as follows: The present invention proposes an Action Response Contrast Network (ARCN), which establishes a multi-channel cross-time dynamic skeleton joint attention topology and uses feature contrast learning to construct the latent space of fine-grained actions. This method aims to overcome the limitations of a single topology structure and make up for the lack of appearance information in skeleton data from different perspectives. Specifically, the present invention first designs an action response topology graph convolutional network, which consists of two parts: 1) The Action Response Topology Module (ART) enhances multi-scale temporal feature learning based on an improved dynamic skeleton topology. It emphasizes integrating unique action patterns between consecutive frames, focuses on the spatio-temporal topology features of fine-grained actions, and establishes a multi-channel cross-time dynamic skeleton topology structure; 2) The Action Response Attention Module (ARA), which embeds decoupled spatio-temporal features into the same space for specific joint spatio-temporal attention feature learning. This method better captures the unique joint feature differences of fine-grained operations from both global and local perspectives. Secondly, the present invention applies the feature contrast learning method to explore the latent space of fine-grained actions, and proposes a Fine-grained Action Comparator (FAC); by classifying sample encodings into similar and different clusters, the global feature representation is updated, and multi-level feature contrast learning is performed at different stages of the model.
[0007] To achieve the above-mentioned inventive purpose, the technical solution adopted by the present invention is specifically as follows: A skeleton fine-grained action recognition method based on an action response contrast network, comprising the following steps:
[0008] S1. Perform data preprocessing on the collected human skeleton data to form four streams of features: joint data stream, bone data stream, joint velocity data stream, and bone velocity data stream;
[0009] S2. Construct an action-responsive contrast network framework ARCN for the fine-grained action recognition method. Among them, ARCN is mainly composed of two modules: an action-responsive graph convolutional network (Action-Responsive Graph Convolutional Network, ARGCN) and a fine-grained action comparator (Fine-Grained Action Comparator, FAC). ARGCN establishes a multi-channel cross-temporal dynamic skeleton joint attention topology to predict the action category, which is composed of an action-responsive topology module (Action-Responsive Topology, ART) and an action-responsive attention module (Action-Responsive Attention, ARA). While ARGCN is running, the discriminative features are periodically passed into FAC. FAC respectively conducts feature comparison and update in the spatial and temporal dimensions to perform more detailed discrimination on fine-grained actions;
[0010] S3. The preprocessed multi-stream data is input into the network framework ARCN for processing. Each type of skeleton data will finally generate two types of loss values: one comes from the cross-entropy loss of the action-responsive graph convolutional network, and the other is the action comparison loss obtained from the fine-grained action comparator based on feature comparison. Considering these two types of loss values corresponding to the four-stream input data jointly constitutes the basis for finally estimating the accuracy of fine-grained action recognition.
[0011] Furthermore, the S1 step includes the following steps:
[0012] S11. Symbol definition: In skeleton-based action recognition, a skeleton sequence is a collection of a series of action video frames. In each frame, the skeleton is represented as a graph G=(V, E), where V={v 1 , v 2 , …, v N} represents the vertex set composed of N joints, and E={e 1 , e 2 , …, e M} represents the edge set composed of the bones between the joints. The structure of the skeleton graph can be described by the unique adjacency matrix A of the graph, where the value of A ij is 0 or 1, indicating whether there is an edge connection between the joints v i and v j . The entire skeleton sequence can be simply represented by a coordinate set X={x∈R C×T×V}, where C is the coordinate dimension, generally C = 3, T is the total number of frames in the skeleton sequence time, and V represents the total number of joints of the current skeleton;
[0013] S12. For the joint data stream input, fix the joint dimension, use the defined human spine center node as the unified standardized zero node, and obtain the normalized position feature P = {p i |(i = 1, 2, …, V in )} through subtracting other joint points from this standard zero node. The formula is:
[0014] p i = x[:, :, i] - x[:, :, c] (1)
[0015] where p i is the normalized position feature of each joint index of the skeleton sequence X, x is the input feature of the skeleton sequence X, c is the index of the human spine center node, and concatenating X and P can directly obtain the joint stream input;
[0016] S13. The bone data stream input is obtained by calculating the relative positions between adjacent joints in each frame, that is, the direction vector of each bone;
[0017] S14. The joint velocity data stream input is obtained by calculating the position change amount of each joint in each frame;
[0018] S15. The bone velocity data stream input is obtained by calculating the change amount of the direction vector of each bone in each frame.
[0019] Furthermore, the S2 step includes the following steps:
[0020] S21. Construct the Action-Responsive Graph Convolutional Network (ARGCN) of the framework ARCN. ARGCN completes the learning of the multi-channel cross-temporal dynamic skeleton joint attention topology, which mainly includes the Action-Responsive Topology Module (ART) and the Action-Responsive Attention Module (ARA). ART learns the multi-channel cross-temporal dynamic skeleton topology structure of different channels and aggregated action temporal features, and ARA models the importance of skeleton joints in the spatio-temporal dimension.
[0021] S22. Construct the Fine Action Comparator (FAC) of the framework ARCN. FAC constructs a learnable fine action latent space to explore the distinguishable hidden information between fine actions.
[0022] Furthermore, in the S21 step, constructing the Action-Responsive Graph Convolutional Network (ARGCN) of the framework ARCN includes the following steps:
[0023] S211. The input skeleton data first passes through the Action-Responsive Topology Module (ART) to construct a topology graph of a specific motion category;
[0024] S212, the features are continuously transmitted to the action responsive attention module (ARA) to strengthen the feature expression of the main motion joints. After each ARA, the feature map is transmitted to the fine action comparator (FAC) again to strengthen the judgment;
[0025] S213. Obtain the probability distribution of the required motion category through a pooling layer and a fully connected layer with a softmax activation function.
[0026] Furthermore, in the step S211, the input skeleton data is first processed by an action-responsive topology module (ART) including the following steps:
[0027] S2111, normalizing the input data to stabilize the value;
[0028] S2112, through three parallel action-responsive graph convolution blocks to extract some hidden topological connections that are most relevant to the action in space;
[0029] S2113. Perform multi-scale temporal convolution to learn the strength of joint connections when the duration of an action changes.
[0030] Furthermore, in the step S2112, normalizing the input data standard to stabilize the value includes the following steps:
[0031] S21121, input feature map x∈R C×T×V Transformed into a more refined feature representation through a bottleneck layer (1x1 convolution layer);
[0032] S21122, performing pooling along the time dimension to aggregate the time features, thereby obtaining the spatial features after dimensionality reduction processing;
[0033] S21123. Through the designed correlation function, we can deeply explore the relationship between non-adjacent joints and mine the hidden topological correlation matrix A′∈R specific to non-connected joints. C′×V×V :
[0034]
[0035] where a′ ij ∈R C′ is the skeleton vertex (v i , v y ) The corresponding relationship vector in the hidden topological association matrix A′, x i 、x j ∈R C×T is the corresponding skeleton vertex v i 、v j The input features, M 1(·) Simulate the correlation relationship between two vertices by performing bottleneck convolution after calculating the distance between vertex features after feature transformation. φ(·) and φ(·) transform the input features, and here bottleneck convolution and temporal dimension pooling are both used sequentially. σ(·) represents the activation function Tanh, and Conv represents the bottleneck layer convolution.
[0036] S21124. Use A′ to improve the original natural connection skeleton topology connection matrix A ∈ R V×V to obtain the specific action topology structure matrix H ∈ R C′×V×V :
[0037] H = M 2 (A, A′) = A + α · A′ (3)
[0038] where, M 2 (·) constructs the specific action topology structure matrix, α is a learnable parameter that can adjust the importance of hidden topology associations. Each specific action topology structure matrix reveals the association pattern of vertex interconnections under specific motion attributes. To amplify the differences in different action features and aggregate global skeleton topology information, the representation x′ C,:,: ∈ R 1×T×V after the original features are transformed by the bottleneck layer is c ∈ R 1×V×V aggregated channel by channel with the specific action topology structure matrix H C′×T×V .
[0039] Furthermore, in the S2113 step, performing multi-scale temporal convolution includes the following steps:
[0040] S21131. Introduce multi-scale temporal convolution to study the different joint connection strengths across time frames. This module contains four branches, and each branch contains a bottleneck convolution to compress the channels. The first three branches also include two temporal convolutions with different dilations and a max pooling layer to capture more extensive action-related information. The results of these four branches are pooled to obtain the multi-channel cross-temporal domain skeleton topology, and its representation feature is F ∈ R C′×T′×V . The stability of the network is enhanced through residual connections during the training process.
[0041] Furthermore, in the S212 step, passing the features into the action-responsive attention module (ARA) includes the following steps:
[0042] S2121. The input multi-channel cross-temporal domain dynamic skeleton topology representation feature F in ∈ R C×T×VPooling is performed separately in the time and space dimensions to extract specific action features as corresponding spatial joint features and time frame features. Then, the two feature vectors are concatenated:
[0043] F mid1 = Cat[Pool t (F in ), Pool s (F in )] (4)
[0044] Among them, F mid1 represents the features of intermediate transformation, Pool t (·) and Pool s (·) are average pooling operations on the time frame and spatial joints respectively, and Cat represents the concatenation operation.
[0045] S2122. Information is compressed through a bottleneck layer, and hidden features are mined through normalization and non-linear transformation of the activation function:
[0046] F mid2 = ξ(BN(Conv(F mid1 ))) (5)
[0047] Among them, F mid2 represents the features of intermediate transformation, Conv represents the convolution of the bottleneck layer, BN represents the normalization operation, and ξ(·) represents the HardSwish activation function.
[0048] S2123. After that, the updated spatial joint features and time frame features are separated one by one, and spatial joint and time joint feature scores are extracted independently:
[0049] F t , F s = Split(F mid2 , dim = 2) (6)
[0050] Among them, Split(·) re-separates the spatio-temporal embedding features into spatial features F s and time features F t .
[0051] S2124. Through the outer product operation of the spatial joint feature score and the time feature score, a score feature F att ∈R C×T×V is generated to represent the joint attention strength of the entire action sequence.
[0052]
[0053] Among them, σ(·) represents the Sigmoid activation function.
[0054] Finally, through the original input feature F in and the obtained score feature F att perform outer product and normalization processing to complete feature mapping, and combine the obtained feature with the original input feature F in perform residual connection again to obtain the required multi-channel cross-temporal dynamic skeleton joint attention topology F out ∈R C×T×V :
[0055]
[0056] wherein, and represent channel outer product and element aggregation, and ψ(·) represents the Swish activation function. It should be noted that the shapes of the output feature and the input feature remain unchanged.
[0057] Furthermore, in the step S22, constructing the Fine-grained Action Comparator (FAC) of the framework ARCN includes the following steps:
[0058] S221. Decouple the feature in space and time to deeply explore the distinguishable hidden features that fine-grained actions may contain across different dimensions for further research;
[0059] S222. Input the space-time decoupled features into the fine-grained action contrast learning respectively to complete the feature contrast learning of the action hidden space in the independent space and time of the fine-grained action.
[0060] Furthermore, in the step S221, decoupling the feature in space and time includes the following steps:
[0061] S2211. The input feature sequence undergoes spatial and temporal pooling, and then bottleneck convolution and feature flattening to respectively capture a deep understanding of temporal and spatial changes without interfering with each other:
[0062] F t ′ = ReLU(BN(Conv(Pool s (F in )))) (9)
[0063] F′ s = ReLU(BN(Conv(Pool t (F in )))) (10)
[0064] wherein, F in represents the incoming feature, F t ′ represents the temporal feature extracted from the sample, and F s ′ represents the spatial feature extracted from the sample.
[0065] Further, in step S222, passing the spatio-temporal decoupled features into the fine-grained action contrastive learning respectively includes the following steps:
[0066] S2221. For sample i, the time feature F t ′ and the space feature F′ s extracted after spatio-temporal decoupling are subjected to contrastive learning by the same method. Hereinafter, F i is used instead.
[0067] S2222. Given an action m, the predicted label results of the features F i extracted from a sample i are respectively one-hot encoded, and then compared with the true label m encoding to determine whether they are classified as true positives (TP), false negatives (FN), or false positives (FP). TP refers to the positive examples of actions correctly determined by the model. By performing element-wise multiplication on the one-hot encoding of the action label and the predicted one-hot encoding, we obtain the TP for each sample for each category, which is called the same cluster; FN and FP are the results of incorrect determination by the model. However, FN means that the model incorrectly predicts the current action as other actions, that is, misjudgment. And FP means that the model incorrectly predicts other actions as the current action, but in fact, they are samples of other actions. FN and FP together constitute the overall different clusters.
[0068] S2223. Generally speaking, there are more samples within each same cluster and their features have better feature consistency. The exponential moving average is used for each same cluster to maintain the update of the global representation (i.e., the class prototype feature):
[0069]
[0070] where J m is the global representation of the features of action m, is the same-cluster samples of action m, is the number of existing same-cluster samples, and α is the momentum term, which is generally set to 0.9 according to experience. By using the global representation, the sample features belonging to the same cluster can be concentrated together to reduce the within-class difference.
[0071] S2224. Calculate the average value of their features for different clusters to optimize these misclassified samples in the subsequent contrastive learning. The formula is as follows:
[0072]
[0073] where F fn and F fp respectively represent the average feature representations of the FN sample set and the FP set, and they together constitute the central representation of different clusters. is the FN sample of different clusters, is the FP sample in the heterogeneous cluster, and are the quantities of the corresponding samples respectively.
[0074] S2225. To quantify the similarity between the sample features and the average feature vector, calculate the cosine similarity between the sample features and the prototype features of each class and the heterogeneous cluster average feature representation, obtain the contrastive learning score and use it for subsequent loss calculation:
[0075]
[0076] where Score m is the contrastive learning score between the global representation of action m (class prototype feature) and the current sample i, which is the basis of contrastive learning. Score fn and Score fp are the contrastive learning scores of the FN sample and the FP sample with the current sample i respectively.
[0077] S2226. Although the FN sample and the FP sample belong to the same heterogeneous cluster, they are different. For the FN sample, it should be made closer to the sample features of the same cluster in the feature space, which may calibrate the model's incorrect prediction of fine action features. For the FP sample, it should be made farther from the sample features of the same cluster in the feature space, which may more effectively distinguish the subtle differences between fine features. For this reason, a compensation term φ i is introduced for the FN sample, and a penalty term Ψ i is introduced for the FP sample:
[0078]
[0079] Adjusting the compensation term φ i and the penalty term Ψ i helps to correct the misclassification of FN samples in different clusters and prevent the misjudgment of FP samples.
[0080] S2227. Finally, use the various contrastive learning scores obtained above to define the action contrast loss
[0081]
[0082] where p in is the action contrast loss for the features extracted from the current action m, is the action contrast i of the features F
[0083] extracted from the current action m.
[0084] S2228. The complete FAC obtains the action contrast loss L in the time dimensiontac and the spatial dimension action contrast loss L sac Adding the two-dimensional action contrast losses together can form the action contrast loss L AC :
[0085] L AC = L tac + L sac (20)
[0086] Furthermore, S3 includes the following steps:
[0087] S31. In the ARGCN, use the cross-entropy loss to train the network:
[0088]
[0089] where N is the number of samples in the batch. y ij is the one-hot encoded representation of action sample i. y ij = 1 if and only if j is the target action class of sample i. p ij is the probability score that the network predicts sample i belongs to class m.
[0090] S32. In order to widely learn the action latent space of fine actions, every time an ART is passed, the feature map of the multi-channel cross-temporal dynamic skeleton joint attention topology is passed into the FAC, forming a multi-stage action contrast loss. The multi-stage action contrast loss can be defined as a weighted average sum:
[0091]
[0092] where L MAC represents the multi-stage action contrast loss finally obtained by the FAC of the entire action-responsive contrast network. i represents the number of stages, a total of 4 stages, respectively at the second, seventh, eleventh, and fourteenth layers of the ARGCN. L AC is the action contrast loss obtained by each stage of the FAC, and λ i is a learnable parameter to adjust the proportion of each stage's loss.
[0093] S33. Combine the cross-entropy loss with the multi-stage action contrast loss to form a complete learning objective function:
[0094]
[0095] where, is a learnable hyperparameter that can be used to adjust the importance of the multi-stage action contrast loss in the ARCN.
[0096] Compared with the prior art, the beneficial effects of the present invention are:
[0097] 1. The present invention proposes an Action Response Contrast Network (ARCN), which establishes a multi-channel cross-time dynamic skeletal joint attention topology and uses feature contrast learning to construct the latent space of fine-grained actions. This method aims to overcome the limitations of a single topological structure and make up for the lack of appearance information in skeletal data from different perspectives.
[0098] 2. The present invention first designs an action response topological graph convolutional network, which consists of two parts: 1) The Action Response Topology Module (ART) enhances multi-scale temporal feature learning based on the improved dynamic skeleton topology. It emphasizes integrating unique action patterns between consecutive frames, focuses on the spatio-temporal topological features of fine-grained actions, and establishes a multi-channel cross-time dynamic skeleton topology structure; 2) The Action Response Attention Module (ARA), which embeds decoupled spatio-temporal features into the same space for specific joint spatio-temporal attention feature learning. This method better captures the unique joint feature differences of fine-grained operations from both global and local perspectives.
[0099] 3. The present invention applies the feature contrast learning method to explore the latent space of fine-grained actions, and proposes a Fine-grained Action Comparator (FAC). By classifying sample encodings into similar and different clusters, it updates the global feature representation and performs multi-level feature contrast learning at different stages of the model. This method alleviates the problem of fuzzy boundaries between fine-grained behaviors and effectively improves the network's ability to distinguish fine-grained skeletal behaviors. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention.
[0101] Figure 1 It is the overall framework diagram of the Action Response Contrast Network (ARCN) in the present invention.
[0102] Figure 2 It is the framework diagram of the Action Response Graph Convolutional Network (ARGCN) in the present invention.
[0103] Figure 3 It is the exploration graph of the classification action accuracy of the Action Response Topological Graph Convolutional Network in the present invention on the X-Sub only joint input branch of the NTU RGB+D dataset; among them, (a) represents the recognition performance of ARGCN (with 80% as the benchmark), blue represents actions with a recognition accuracy greater than 80%, and red represents actions with a recognition accuracy less than 80%; (b) is the error analysis of the low recognition rate of the true label "writing", that is, the error probability that "writing" may be misjudged as other actions.
[0104] Figure 4This is the specific structure diagram of the fine action comparator FAC in the present invention.
[0105] Figure 5 This is the dynamic skeleton topology diagram of the action random samples in the present invention under multi-channel cross-time domain in the deep layer of the model.
[0106] Figure 6 This is the joint feature attention map of different action random samples in the present invention.
[0107] Figure 7 This is the radar chart of whether FAC has an impact on the model performance on six datasets in the present invention.
[0108] Figure 8 This is the confusion matrix diagram of the related failure actions in the present invention, where the accuracy of the baseline model is less than 80% under the input of only joint modality in the NTU RGB+D X-Sub benchmark; among them, (a) and (b) respectively represent the confusion matrices of the failure actions of the baseline model and the model of the present invention, where the coordinate axes represent each action category, and the red rectangles represent similar actions.
[0109] Figure 9 This is the schematic diagram of the classification accuracy comparison results of different methods for six easily confused actions in the embodiment of the present invention.
[0110] Figure 10 This is the t-SNE visualization result diagram of different models for several fine hand action random samples in the present invention.
[0111] Figure 11 This is the visualized skeleton diagram of the model on 10 frames of some example actions in the present invention. Detailed implementation manners
[0112] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0113] Embodiment 1
[0114] Refer to Figure 1 And Figure 11 In this embodiment, its technical solution is provided as a skeleton fine action recognition method based on an action-responsive contrast network. The invention discovers the differential expression of the hidden features of fine actions by establishing a multi-channel cross-time domain dynamic skeleton joint attention topology and performing feature contrast update in a low-dimensional space to improve the accuracy of fine action recognition. The proposed framework is the Action-Responsive Contrastive Network (ARCN) architecture, as Figure 1As shown. The entire architecture adopts the joint evaluation of multi-stream data and is mainly composed of two modules: the Action-Responsive Graph Convolutional Network (ARGCN) and the Fine-Grained Action Comparator (FAC). Among them, ARGCN establishes a multi-channel cross-temporal dynamic skeleton joint attention topology to predict the action category, which is composed of the Action-Responsive Topology (ART) and the Action-Responsive Attention (ARA). ART learns the multi-channel cross-temporal dynamic skeleton topology that aggregates action temporal features across different channels, and ARA models the importance of skeleton joints in the spatio-temporal dimension. While ARGCN is running, the discriminative features are periodically passed into FAC. FAC performs feature comparison and update in the spatial and temporal dimensions respectively to conduct a more detailed discrimination of fine-grained actions. Through the joint action of ARGCN and FAC, each type of skeleton data will finally generate two types of loss values: one is the cross-entropy loss from the classification task, and the other is the action contrast loss based on feature comparison and update. The joint consideration of these two types of loss values corresponding to the four-stream input data constitutes the basis for finally estimating the accuracy of fine-grained action recognition.
[0115] Action-Responsive Graph Convolutional Network (ARGCN): Although the natural human bone connections are reliable for action recognition, their fixed nature limits the ability to detect hidden information in actions and affects the recognition of complex and fine-grained actions. By establishing a dynamic skeleton topology for specific actions and focusing on key parts, this method can capture global and local action information. Focusing on key action areas can enhance the feature representation beyond bone connections, thus improving action understanding and analysis. For these reasons, the present invention designs the Action-Responsive Topology Graph Convolutional Network (ARGCN), as Figure 2As shown in the figure. Specifically, the ARGCN backbone consists of 14 basic blocks, including 10 action-responsive topological modules (ARTs) and 4 action-responsive attention modules (ARAs). The input skeleton data first passes through ART to construct a topological map of a specific motion category, and then passes through ARA to strengthen the feature expression of the main motion joints. After each ARA, the feature map will be passed back to the fine action comparator (FAC) to strengthen the judgment. Finally, the probability distribution of the required motion category is obtained through the pooling layer and the fully connected layer with a softmax activation function. It is particularly important to note that multi-scale feature matching helps in skeleton action recognition. The 6th and 10th basic blocks are implemented by cross-row multi-scale temporal convolutions with a step size of 2. The corresponding module increases the channel dimension while reducing the time dimension, generating multi-scale features.
[0116] Action-responsive topology module (ART) in ARGCN: In order to fully explore the global and local information of the skeleton topology, the present invention designs ART by taking advantage of the different unnatural topologies of the skeletons of different motion types. It first normalizes the input data to stabilize the values, then uses three parallel action-responsive graph convolution blocks to spatially extract some hidden topological connections that are most relevant to the action, and finally performs multi-scale temporal convolution to learn the strength of joint connections when the duration of the action changes. Figure 2 The upper right shows the specific structure of ART.
[0117] Specifically, in order to effectively infer the hidden skeleton connection representation between non-adjacent joints and reduce the amount of computation, the input feature map x∈R C×T×V It is converted into a more refined feature representation through the bottleneck layer (1x1 convolution layer). Then, pooling is performed along the time dimension to aggregate the time features, thereby obtaining the spatial features after dimensionality reduction. The designed correlation function is then used to deeply explore the relationship between non-adjacent joints and to mine the hidden topological association matrix A′∈R specific to the unconnected joints. C′×V×V :
[0118]
[0119] where a′ ij ∈R C′ is the skeleton vertex (v i , v y ) The corresponding relationship vector in the hidden topological association matrix A′, x i 、x j ∈R C×T is the corresponding skeleton vertex v i 、v i The input features, M 1(·) Simulate the correlation relationship between two vertices by performing bottleneck convolution after calculating the distance between vertex features after feature transformation. φ(·) and φ(·) transform the input features, and here bottleneck convolution and temporal dimension pooling are adopted successively. σ(·) represents the activation function Tanh, and Conv represents the bottleneck layer convolution. Then, use A′ to improve the original natural connection skeleton topology connection matrix A∈R V×V to obtain the specific action topology matrix H∈R C′×V×V :
[0120] H = M 2 (A, A′) = A + α·A′ (3)
[0121] where M 2 (·) establishes the specific action topology matrix, and α is a learnable parameter that can adjust the importance of hidden topological associations. Each specific action topology matrix reveals the association pattern of vertex interconnections under specific motion attributes. To amplify the differences in different action features and aggregate global skeleton topology information, the representation x′ C,:,: ∈R 1×T×V after the original feature is transformed by the bottleneck layer is c ∈R 1×V×V aggregated channel by channel with the specific action topology matrix H C′×T×V ∈R
[0122] For fine-grained action recognition, since the spatial structures of some actions are extremely similar, it is not enough to only explore the spatial topology of the skeleton. Considering that temporal features can be used as a differentiating factor for fine-grained actions, this module introduces multi-scale temporal convolution to study the different joint connection strengths across time frames, as Figure 2 shown in the upper right. This module contains four branches, and each branch contains a bottleneck convolution to compress the channels. The first three branches also include two temporal convolutions with different dilations and a max pooling layer to capture more extensive action-related information. The results of these four branches are pooled to obtain the multi-channel cross-temporal-domain skeleton topology, and its representation feature is F∈R C′×T′×V . The stability of the network is enhanced through residual connections during the training process.
[0123] Action-responsive attention module (ARA) in ARGCN: Multidimensional information interaction is very effective for generating better attention maps. Considering the mutual relationship between spatial and temporal information in the human action skeleton sequence, it is crucial to identify the joints with the most information in a specific frame of the entire sequence, especially for identifying fine-grained actions with subtle differences. For this reason, the present invention designs ARA, as outlined Figure 2 shown in the lower right.
[0124] First, the input multi-channel dynamic skeleton topological representation features F across time domains in ∈R C×T×V are respectively pooled in the time and space dimensions to extract specific action features as corresponding spatial joint features and time frame features. Then, the two feature vectors are concatenated, information is compressed through a bottleneck layer, and hidden features are mined through normalization and non-linear transformation of the activation function. After that, the updated spatial joint features and time frame features are separated one by one, spatial joint and time joint feature scores are independently extracted, and through the outer product operation of the spatial joint feature score and the time feature score, a score feature F representing the joint attention strength of the entire action sequence is generated att ∈R C×T×V . Finally, through the original input feature F in and the obtained score feature F att , an outer product and normalization process are performed to complete feature mapping, and the obtained feature and the original input feature F in are subjected to residual connection again to obtain the required multi-channel dynamic skeleton joint attention topology F across time domains out ∈R C×T×V . It should be noted that the shapes of the output feature and the input feature remain unchanged. The proposed action-responsive attention module can be expressed as:
[0125] F mid1 = Cat[Pool t (F in ), Pool s (F in )] (4)
[0126] F min2 = ξ(BN(Conv(F mid1 ))) (5)
[0127] F t , F s = Split(F mid2 , dim = 2) (6)
[0128]
[0129] where F mid1 and F mid2 represent the features of intermediate transformations, Pool t (·) and Pool s (·) are average pooling operations on the time frame and spatial joints respectively, Cat represents the concatenation operation, Conv represents the bottleneck layer convolution, BN represents the normalization operation, and Split(·) re-separates the spatio-temporal embedded features into spatial features Fs With the time feature F t 。 and represent the channel outer product and element aggregation, and ξ(·), σ(·) and ψ(·) represent the HardSwish, Sigmoid and Swish activation functions respectively.
[0130] Fine Action Comparator (FAC): Even though a multi-channel cross-temporal dynamic skeleton joint attention topology that strongly represents action features is established, there are still problems similar to other models: it is still difficult to identify some fine actions that only differ in minor parts, such as Figure 3 。 Figure 3 (a) shows the classification accuracy of ARGCN with only joint input on the NTU RGB+D dataset X-Sub. Taking 80% as the cut-off point, the red part on the right represents the actions with an action recognition rate less than 80%. It can be seen from the figure that some actions are still difficult to identify, especially "writing". To analyze the reason for the low recognition rate of "writing", an error analysis of "writing" is carried out, as shown in Figure 3 (b). The results show that for samples with the true label of "writing", there is a 38.8% probability of misjudging as "typing on the keyboard", a 30.1% probability of misjudging as "reading", and a small probability of misjudging as other actions. The motion patterns between these actions are very similar, and generally, appearance information can be used to assist in recognition. However, due to the lack of appearance information in the skeleton, these actions are extremely easy to be confused. To strengthen the recognition of such fine actions, the present invention introduces a method of feature contrast learning, designs a fine action comparator, and constructs a learnable fine action latent space from another angle to explore the distinguishable hidden information between fine actions. FAC consists of two parts, as shown in Figure 4 shown.
[0131] Spatio-temporal decoupling in FAC: Fine-grained actions may contain distinguishable hidden features across different dimensions. To deeply explore such hidden information for further research, the input feature sequence undergoes spatial and temporal pooling, and then bottleneck convolution and feature flattening are performed to capture a profound understanding of temporal and spatial changes without interfering with each other:
[0132] F t ′ = ReLU(BN(Conv(Pool s (F in )))) (9)
[0133] F′ s = ReLU(BN(Conv(Pool t (F in )))) (10)
[0134] where Fin Represents the incoming feature, F t ′ represents the time feature extracted from the sample, F s ′ represents the spatial feature extracted from the sample.
[0135] Fine-grained action contrast learning in FAC: In skeleton-based action recognition, feature contrast learning can optimize feature representation and improve the model's ability to distinguish fine-grained actions by comparing sample features and class prototype features. For sample i, the time feature F t ′ and the spatial feature F s ′ are subjected to contrast learning using the same method, and hereinafter represented by F i instead.
[0136] Given an action m, the predicted label results of the features F i extracted from a sample i are respectively one-hot encoded, and then compared with the true label m encoding to determine true positives (TP), false negatives (FN), and false positives (FP). TP refers to the positive examples of actions correctly determined by the model. By performing element-wise multiplication on the one-hot encoding of the action label and the predicted one-hot encoding, we obtain the TP for each sample and each category, which is called the same cluster; FN and FP are the results of incorrect determination by the model, but FN means that the model incorrectly predicts the current action as other actions, that is, misjudgment. While FP means that the model incorrectly predicts other actions as the current action, but in fact it is a sample of other actions. FN and FP together constitute the overall different clusters. Generally speaking, there are more samples within each same cluster and their features have better feature consistency. The exponential moving average is used for each same cluster to keep the global representation (i.e., class prototype feature) updated:
[0137]
[0138] Among them, I m is the global representation of the features of action m, is the same-cluster sample of action m, is the number of existing same-cluster samples, and α is the momentum term, which is generally set to 0.9 according to experience. By using the global representation, the sample features belonging to the same cluster can be concentrated together to reduce the within-class difference. Then, the average value of their features is calculated for different clusters in order to optimize these misclassified samples in the subsequent contrast learning. The formula is as follows:
[0139]
[0140] Among them, F fn and F fp respectively represent the average feature representations of the FN sample set and the FP set, and they together constitute the central representation of different clusters. is a FN sample of a different cluster, is a FP sample in the different cluster, and are the numbers of the corresponding samples respectively.
[0141] To quantify the similarity between the sample features and the average feature vector, the cosine similarity between the sample features and the prototype features of each class and the average feature representation of different clusters is calculated to obtain the contrastive learning score and used for subsequent loss calculation:
[0142]
[0143] where Score m is the contrastive learning score between the global representation (class prototype feature) of action m and the current sample i, which is the basis of contrastive learning. Score fn and Score fp are the contrastive learning scores of FN samples and FP samples with the current sample i respectively. Although FN samples and FP samples belong to the same different cluster, they are not the same. For FN samples, they should be made closer to the sample features of the same cluster in the feature space, which may calibrate the model's incorrect prediction of fine action features. For FP samples, they should be made farther away from the sample features of the same cluster in the feature space, which may more effectively distinguish the subtle differences between fine features. For this reason, a compensation term φ i is introduced for FN samples, and a penalty term Ψ i is introduced for FP samples:
[0144]
[0145] Adjusting the compensation term φ i and the penalty term Ψ i helps to correct the misclassification of FN samples in different clusters and prevent the misjudgment of FP samples.
[0146] Finally, various contrastive learning scores obtained above are used to define the action contrast loss
[0147]
[0148] where p im is the action contrast loss of the features extracted by the current action m, is the action contrast loss of the features F i extracted by the current action m. The complete FAC obtains the action contrast loss L tac in the time dimension and the action contrast loss L sac in the space dimension. Adding the two-dimensional action contrast losses can constitute the action contrast loss L AC :
[0149] L AC = L tac + L sac (20)
[0150] In ARGCN, cross-entropy loss is used to train the network:
[0151]
[0152] where N is the number of samples in the batch. y ij is the one-hot encoded representation of action sample i. y ij = 1 if and only if j is the target action class of sample i. p ij is the probability score that the network predicts sample i belongs to class m.
[0153] To widely learn the action latent space of fine actions, every time an ART is performed, the feature map of the multi-channel cross-temporal dynamic skeleton joint attention topology is fed into the FAC, forming a multi-stage action contrast loss. The multi-stage action contrast loss can be defined as a weighted average sum:
[0154]
[0155] where L MAC represents the multi-stage action contrast loss finally obtained by the FAC of the entire action-responsive contrast network. i represents the number of stages, a total of 4 stages, respectively at the second, seventh, eleventh, and fourteenth layers of ARGCN. L AC is the action contrast loss obtained by each stage of the FAC, and λ i is a learnable parameter that adjusts the proportion of the loss of each stage. Combining the cross-entropy loss and the multi-stage action contrast loss forms a complete learning objective function:
[0156]
[0157] where is a learnable hyperparameter that can be used to adjust the importance of the multi-stage action contrast loss in ARCN.
[0158] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0159] Furthermore, the effectiveness of the method proposed by the present invention is verified through simulation experiments.
[0160] Table 1. Comparison of the performance (%) of the proposed model with existing methods on NTU RGB+D, NTU RGB+D 120, NW-UCLA, and UAV-Human datasets. The top part consists of several models without GCN technology, while the other part contains some graph-based models. The best results are marked in bold, and the second-best are underline marked.
[0161]
[0162] The model is compared with the current state-of-the-art methods on NTU RGB+D, NTU RGB+D 120, NW-UCLA, UAV-Human datasets, as well as two fine-grained datasets, FineGYM and Diving48. All reported results are trained and evaluated on a server with Nvidia GeForce RTX 4090 and experiments are conducted using PyTorch. We adopt standard evaluation metrics to assess the performance of the model for fine-grained action recognition of skeletons.
[0163] Table 1 shows the comparison with previous methods on NTU RGB+D, NTU RGB+D 120, NW-UCLA, and UAV-Human datasets. In the early stage, some CNN- or RNN-based methods were more accurate than GCN-based models. For example, VA-fusion had an accuracy 7.9% and 6.7% higher than the classic GCN-based model ST-GCN on the NTU RGB+D dataset. This higher accuracy stems from the fact that VA-fusion combines the mature technologies of RNN and CNN, fully extracting action features, while GCN technology was just emerging at that time and the advantages of skeletal data were not fully explored. With the continuous development of GCN technology, the task of using GCN for skeleton action recognition has been getting better and better. It is worth noting that most current state-of-the-art methods adopt a multi-stream fusion framework. For fair comparison, the proposed model reports the fusion results of four modes: joints, bones, joint motion, and bone motion as the final results. It can be observed that the proposed model outperforms most existing methods on these four datasets. On the NTU RGB+D 120X-Sub, NTU RGB+D 120X-Set, NW-UCLA, and UAV-Human datasets, the accuracies of the proposed model reached 89.7%, 91.2%, 97.2%, 44.6%, and 72.0% respectively. This is because the proposed model conducts comprehensive spatio-temporal feature interaction extraction and global information aggregation, effectively utilizing the spatio-temporal features of the skeleton action sequence and achieving more fine-grained human action recognition. On the NTU RGB+D dataset, the proposed model achieved state-of-the-art results with a reasonable gap of 0.1% from the best model.
[0164] Table 2. Performance (%) comparison of the model of the present invention with existing methods on the FineGYM and Diving48 fine-grained datasets. The best is marked in bold, and the second best is underline marked.
[0165]
[0166] To further verify the effectiveness of the model of the invention for the fine action recognition task, two fine-grained action datasets, FineGYM and Diving48, were selected for experiments. The results are shown in Table 2. The model of the invention achieved the highest results on the FineGYM dataset, improving by 8.6%, 3.3%, 3.4%, and 1.2% compared to models such as ST-GCN, MS-G3D, CTR-GCN, and PoseC3D respectively. On the Diving48 dataset, there is only a 0.2% gap with the current best PoseC3D. The results show that the model of the invention can effectively learn the rich action representations of fine actions.
[0167] Table 3. Comparison of model performance (%) using different stream data inputs and multi-stream data fusion on six datasets.
[0168]
[0169] Table 4. Performance (%) comparison of different model components using only joint data stream input on the NTU RGB+D dataset X-Sub, NTU RGB+D 120 dataset X-Sub, NW-UCLA, UAV-Human dataset CSv1, FineGYM, and Diving48 datasets.
[0170]
[0171] ART: Action Responsive Topology Module; ARA: Action Responsive Attention Module; FAC: Fine Action Comparator.
[0172] After that, the present invention designed ablation experiments to study the effectiveness of the invention model. The present invention first verified the impact of multi-stream data fusion on the final result. Then, the effectiveness of each component in the model was verified. Next, the present invention respectively demonstrated the dynamic skeleton topology across multi-channels and time domains, spatio-temporal joint attention maps, confusion matrices, t-SNE visual feature representations for some fine actions, and model comparison diagrams to further illustrate the advantages of the model in fine action recognition. Finally, the process of visualizing the working of the model was used to intuitively show the role of the model of the present invention.
[0173] 1) Influence of multi-stream data fusion on model results: To explore the importance of multi-stream data fusion, Table 3 was designed in this invention. The results show that, compared with the traditional method of only using the joint stream for action recognition tasks, as the number of data branches increases, the performance of the model is significantly improved on six datasets, namely NTU RGB+D, NTU RGB+D 120, NW-UCLA, UAV-Human, FineGYM, and Diving48. The performance of the model is the best when four-stream data fusion is used. Specifically, compared with using a single joint stream, the performance of the model on X-Sub and X-View of the NTU RGB+D dataset is improved by 2.1% and 1.6% respectively, the performance of the model on X-Sub and X-Set of the NTU RGB+D 120 dataset is improved by 4.1% and 3.7% respectively, the performance of the model on the NW-UCLA dataset is improved by 2.6%, the performance of the model on CSv1 and CSv2 of the UAV-Human dataset is improved by 7.1% and 4.5% respectively, the performance of the model on the FineGYM dataset is improved by 3.8%, and the performance of the model on the Diving48 dataset is improved by 9.1%. This result further verifies the effectiveness of training the model using multi-stream data fusion.
[0174] 2) Influence of each component in the Action Response-based Contrastive Network (ARCN): The Action Response-based Topology Module (ART), Action Response-based Attention Module (ARA), and Fine Action Comparator (FAC) are the main components of ARCN. To evaluate their respective contributions to the effectiveness of the model, CTR-GCN was set as the baseline model, and Table 4 was designed to verify using only the joint data stream input on six datasets (only X-Sub for NTU RGB+D and NTU RGB+D 120). It can be seen that all these components contribute to improving the performance of the baseline. The specific roles of each component are described as follows:
[0175] a) Effectiveness of the Action Response-based Topology Module (ART): As can be seen from Table 4, compared with the baseline model that only uses the multi-channel dynamic skeleton topology, ART, which combines action duration information to establish a multi-channel cross-time-domain dynamic skeleton topology, improves the model effect by 0.3%, 0.3%, 0.8%, 0.3%, 0.3%, 0.5%, and 0.7% on the six datasets respectively. This shows that considering the action patterns of the whole human body by combining spatio-temporal information to establish the skeleton topology is more effective for fine action recognition. To verify this statement, in Figure 5The multi-channel cross-temporal dynamic skeleton topology generated by the 13th layer of the visualization model, where the input action random sample labels are the easily recognizable "jump" and the easily confused actions "reading" and "writing". It can be observed that: (1) The joint correlation of the generated topological structure tends to be rough and dense, which can capture the global features during the action duration and facilitate fine action recognition; (2) The differential expression of the joint correlation of the topological structure shows that the joint correlation of the lower part of the human body is stronger during "jumping", and the joint correlation of the hands is stronger during "reading" and "writing", indicating that the model has successfully simulated different joint relationships under different motion types; (3) For easily distinguishable actions, such as the obvious differences in the joint correlation of the topological structures of "jump" and "reading", "writing" respectively. For the easily confused "reading" and "writing", the model can learn global and subtle differences. For example, when "reading" and "writing" are in progress, the action patterns are different, and the correlation between the spine and other joints is different (the joint correlation in the green box). For the jointly important hand joints, the correlation of the tip of the left hand is stronger during "reading" (the joint correlation in the red box). The results show that ART is effective for the discrimination of fine actions.
[0176] b) Effectiveness of the action-responsive attention module (ARA): To illustrate the characteristics of ARA, the attention maps of three action random samples are depicted in four layers of the model. The selected actions are "jump", "reading", and "writing", as Figure 6 shown. For each subfigure, in the primary stage (such as the 2nd layer), the spatio-temporal joint attention feature map obtained by ARA shows stronger selectivity in spatial joint connections, while in the advanced stage (such as the 14th layer), ARA shows additional selectivity in action time. Taking "jump" as an example, in the primary stage, ARA learns the joints spatially associated with the action in almost all frames, such as the left wrist, right wrist of the waving hands, and the right hip exerting force during "jump". In the advanced stage, ARA mainly focuses on expressing the importance of the most important joints related to "jump" within a certain range of frames. It should be noted that the model uses cross-row time convolution in the 6th and 10th layers, and the frames outside the later action duration are filled with zeros to be consistent with the action time frames of the initial input. Generally speaking, ARA pays more attention to information joints in the early stage of the model, and differentiates information frames in the later stage of the model. It is beneficial for mining fine action recognition that is similar in space but different in time, such as Figure 6Attention maps for the two easily confused actions of "reading" and "writing" shown. At the primary stage, the joint feature attention maps of "reading" and "writing" are very similar. At the advanced stage, the joint importance related to "reading" tends to be expressed in the early time frames, while the joint importance related to "writing" tends to be expressed in the later time frames. In Table 4, using ARA can improve the model performance by approximately 0.1%-0.5% respectively on six datasets. Thus, it can be concluded that ARA is effective for fine-grained action recognition.
[0177] c) Effectiveness of the Fine-grained Action Comparator (FAC): From the results in Table 4, the great role of FAC for the model can be observed. On the four conventional action recognition datasets (NTU RGB+D, NTU RGB+D 120, NW-UCLA, UAV-Human), the performance of the model with FAC is improved by 0.6%, 0.5%, 1.6%, and 0.8% respectively compared to without FAC. On the two fine-grained action recognition datasets (FineGYM, Diving48), the performance is improved by 1% and 1.3%. To more intuitively show the role of FAC, Figure 7 . Figure 7 shows the impact of the presence or absence of FAC on the classification performance of the model on six datasets. Among them, the red range represents the classification performance of the baseline model, and the blue range represents the classification performance of the baseline model after adding FAC. The blue completely surrounds the red, indicating that for each dataset, the classification performance of the baseline model after adding FAC is better than the original.
[0178] 3) Effectiveness of the final model (ARCN):
[0179] The results in Table 4 show that on the six datasets, the performance of the model of the present invention is improved to varying degrees compared to the baseline model. The best is on the NW-UCLA dataset, with a 1.9% performance improvement.
[0180] To further analyze the performance of the model, we compared the model of the present invention with the baseline model in recognizing some fine-grained actions. Figure 8 (a) and Figure 8 (b) highlight the actions where the accuracy of the baseline model using joint-modal input on NTU RGB+D X-Sub is less than 80%. From Figure 8From (a) and (b), we can see three groups of similar actions, marked with red rectangles, namely, "reading" and "writing", "playing with mobile phone" and "typing on keyboard", "taking off shoes" and "putting on shoes". The first four actions are mainly completed by slight shaking of the two hands. They are extremely similar in spatial configuration and temporal dynamics, and the probability of misjudgment is relatively high. The latter two actions have similar spatial structures but different temporal dynamics. From the results of (a) and (b), compared with the baseline model, the recognition accuracy of the model proposed in this paper has been improved to varying degrees in these fine actions. The recognition accuracy of actions such as "reading", "writing", "playing with mobile phone", and "typing on keyboard" has increased by 7%, 3%, 9%, and 7% respectively, and the recognition degree of "putting on shoes" and "taking off shoes" has reached more than 80%. In addition, the probability of misjudgment of fine actions has decreased to varying degrees. For example, the probability of misjudging "writing" as "typing on the keyboard" has dropped by 5%, the probability of misjudging "reading" as "writing" has dropped by 6%, and the probability of misjudging "wearing shoes" as "taking off shoes" has dropped by 8%. This shows that the model can mine the subtle differences between fine movements from multiple perspectives of time and space.
[0181] Figure 9 The above six fine movements are compared between the method of the present invention and other methods. The results show that our method has achieved excellent performance, and the accuracy of identifying the six movements has reached the highest level. In particular, the performance of "reading", "playing with mobile phones", and "typing on keyboards" is improved by at least 4.3%, 4.0%, and 4.7% respectively compared with other methods, indicating that our method can classify fine movements more accurately than other methods.
[0182] In addition, in order to further explore the degree of distinction of fine movements by the model in this paper, the present invention uses t-SNE on the more complex NTURGB+D120 dataset X-Sub to visualize the distribution of the features finally learned by the four models for random samples of five random fine hand movements in space, such as Figure 10 As shown. These actions mainly involve subtle interactions between fingers, which are very similar in spatial structure. In order to more intuitively feel the effect of the model on these difficult-to-distinguish actions, we calculated the normalized mutual information (NMI) to quantitatively support the discrimination of the learned representations, where the closer the NMI is to 1, the better the clustering effect. From the results, the model of the present invention can make the clustering of easily confused action features more compact and more distinguishable.
[0183] 4) Model visualization: To show how the model works, Figure 11Show the skeleton sequence diagrams of five actions after being processed by the model. It can be found from the diagrams that different actions have different dynamic skeleton topologies. For the fine actions that are easily confused, such as "reading" and "writing", the number and strength of their non-natural joint topological connections (blue lines) are different. For the actions that are easily distinguishable, such as "reading" and "kicking something", the strength of their natural joint topological connections (red lines) can be distinguished. For example, the connection strength of the hand joints is higher during "reading", while the connection strength of the leg joints is higher during "kicking something". In addition, the model has successfully focused on the most informative joints when different movements occur, that is, the arm joints are the most important during "reading" and "writing", the leg joints are the most important during "kicking something", and the arm and leg joints are both relatively important during "throwing something" and "jumping". For the easily confused fine actions such as "reading" and "writing", based on learning their different dynamic skeleton topologies, the model has also learned that the importance expressions of their joint points are different at different frames. This means that the proposed model works well and can effectively perform the fine action recognition task.
[0184] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A skeleton fine motion recognition method based on an action-responsive contrast network, characterized in that: The following steps are involved: S1. Preprocess the collected human skeleton data to form four stream features: joint data stream, bone data stream, joint velocity data stream and bone velocity data stream; S2. Construct an action-responsive contrast network framework ARCN for skeleton fine action recognition method, where ARCN consists of two modules: action-responsive graph convolutional network ARGCN and fine action contrastor FAC; ARGCN establishes a multi-channel cross-temporal dynamic skeleton joint attention topology to predict the category of action, which consists of an action-responsive topology module ART and an action-responsive attention module ARA. While ARGCN is running, the discriminant features are periodically passed into FAC, and FAC performs feature comparison and update in spatial and temporal dimensions respectively; The S2 step includes the following steps: S21. Build the action-responsive graph convolutional network ARGCN of the ARCN framework. ARGCN completes the learning of multi-channel cross-temporal dynamic skeleton joint attention topology, which includes the action-responsive topology module ART and the action-responsive attention module ARA. ART learns the multi-channel cross-temporal dynamic skeleton topology structure of different channels and aggregated action temporal features, and ARA models the importance of skeleton joints in the spatiotemporal dimension. In the step S21, constructing the action-responsive graph convolutional network ARGCN of the framework ARCN includes the following steps: S211, the input skeleton data first passes through the action-responsive topological module ART to construct a topological graph of a specific motion category; In the step S211, the input skeleton data is first processed by the action-responsive topology module ART, which includes the following steps: S2111, normalizing the input data to stabilize the value; S2112, through three parallel action-responsive graph convolution blocks to extract some hidden topological connections that are most relevant to the action in space; S2113, perform multi-scale temporal convolution to learn the strength of joint connections when the duration of the action changes; S212, the features are continuously transmitted to the action responsive attention module ARA to strengthen the feature expression of the motion joints. After each ARA, the feature map is transmitted to the fine action comparator FAC again to strengthen the judgment; S213, obtaining the probability distribution of the required motion category through a pooling layer and a fully connected layer with a softmax activation function; S22, construct the fine motion contrastor FAC of the ARCN framework, FAC constructs a learned fine motion latent space to explore the distinguishable hidden information between fine motions; S3. The preprocessed multi-stream data is input into the network framework ARCN for processing. Each skeleton data will eventually produce two types of loss values: one is the cross entropy loss from the action-responsive graph convolutional network, and the other is the motion contrast loss obtained by the fine motion comparator based on feature contrast. The two types of loss values corresponding to the four-stream input data are jointly considered to form the basis for the final estimation of the accuracy of fine motion recognition.
2. The skeleton fine motion recognition method based on the action-responsive contrast network according to claim 1 is characterized in that: The S1 step includes the following steps: S11. In skeleton-based action recognition, a skeleton sequence is a collection of action video frames. In each frame, the skeleton is represented as a graph G = (V, E), where V = {v1, v2, ..., v N } represents a vertex set consisting of N joints, E = {e1, e2, …, e M } represents the edge set consisting of bones between joints. The structure of the skeleton graph is described using the graph-specific adjacency matrix A, where A ij The value is 0 or 1, indicating that the joint v i and v j Is there an edge connection between them? The entire skeleton sequence is represented by a coordinate set X = {x∈R C×T×V } represents, where C is the coordinate dimension, C=3, T is the total number of frames in the skeleton sequence, and V represents the total number of joints in the current skeleton; S12, for the joint data stream input, the joint dimension is fixed, the defined human spine center node is used as the unified standardized zero node, and the normalized position feature P is obtained by subtracting other joint points from the standard zero node. i |(i=1,2,…,V in )}, the formula is: p i =x[:,:,i]-x[:,:,c] (1) Among them, p i is the normalized position feature of each joint index of the skeleton sequence X, x is the input feature of the skeleton sequence X, c is the index of the central node of the human spine, and the joint stream input is directly obtained by connecting X and P in series; S13, the skeleton data stream input is obtained by calculating the relative positions between adjacent joints in each frame, that is, the direction vector of each bone; S14, the joint velocity data stream input is obtained by calculating the position change of each joint in each frame; S15. The bone velocity data stream input is obtained by calculating the change in the direction vector of each bone in each frame.
3. The skeleton fine motion recognition method based on the action-responsive contrast network according to claim 2 is characterized in that: In the step S2112, the input data is normalized to stabilize the value, including the following steps: S21121, input feature map x∈R C×T×V Transformed into refined feature representation through bottleneck 1x1 convolution layer; S21122, performing pooling along the time dimension to aggregate the time features, thereby obtaining the spatial features after dimensionality reduction processing; S21123. Through the designed correlation function, we can deeply explore the relationship between non-adjacent joints and mine the hidden topological correlation matrix A′∈R specific to non-connected joints. C′×V×V : where a′ ij ∈R C′ is the skeleton vertex (v i ,v j ) The corresponding relationship vector in the hidden topological association matrix A′, x i 、x j ∈R C×T is the corresponding skeleton vertex (v i ,v j ), M1(·) simulates the correlation between two vertices by calculating the distance between the vertex features after feature transformation and performing bottleneck convolution. (·)and Transform the input features. Bottleneck convolution and time dimension pooling are used in sequence. σ(·) represents the activation function Tanh, and Conv represents the bottleneck layer convolution. S21124. Use A′ to improve the original natural connection skeleton topology connection matrix A∈R V×V To obtain the specific action topology matrix H∈R associated with the action C′×V×V : H=M2(A,A')=A+α·A' (3) Among them, M2(·) establishes a specific action topological structure matrix, α is a learnable parameter that adjusts the importance of hidden topological associations. Each specific action topological structure matrix reveals the association pattern of vertex interconnection under specific motion attributes. In order to amplify the differences in different action features and aggregate the global skeleton topological information, the original features are transformed through the bottleneck layer to represent x′ C,:,: ∈R 1×T×V The specific action topology matrix H obtained by each channel c ∈R 1×V×V (c∈{1,…,c′}) and perform channel-by-channel aggregation to obtain the skeleton topology information output G∈R C′×T×V .
4. The skeleton fine motion recognition method based on the action-responsive contrast network according to claim 2 is characterized in that: In the step S2113, performing multi-scale time convolution includes the following steps: S21131. Multi-scale temporal convolution is introduced to study the strength of different joint connections across time frames. The multi-scale temporal convolution contains four branches, each of which contains a bottleneck convolution to compress the channel. The first three branches also include two temporal convolutions with different expansions and a maximum pooling layer to capture a wider range of action-related information. The results of these four branches are pooled to obtain a multi-channel cross-temporal skeleton topology, which is represented by F∈R C′×T′×V , the stability of the network is enhanced through residual connections during training.
5. The skeleton fine motion recognition method based on the action-responsive contrast network according to claim 4 is characterized in that: In the step S212, the feature is continuously transferred to the action-responsive attention module ARA, including the following steps: S2121, input multi-channel cross-time domain dynamic skeleton topology representation feature F in ∈R C×T×V After pooling in the time and space dimensions, the specific action features are extracted into corresponding spatial joint features and time frame features; the two feature vectors are connected in series: F mid1 =Cat[Pool t (F in ),Pool s (F in )] (4) Among them, F mid1 Represents the characteristics of the intermediate transformation, Pool t (·) and Pool s (·) are the average pooling operations on the time frame and the spatial joint, respectively, and Cat represents the concatenation operation; S2122, compress information through a bottleneck layer, and mine hidden features through standardization and nonlinear transformation of activation functions: F mid2 =ξ(BN(Conv(F mid1 ))) (5) Among them, F mid2 represents the features of the intermediate transformation, Conv represents the bottleneck layer convolution, BN represents the normalization operation, and ξ(·) represents the HardSwish activation function; S2123. Separate the updated spatial joint features and temporal frame features one by one, and independently extract the spatial joint and temporal joint feature scores: F t ,F s =Split(F mid2 ,dim=2) (6) Among them, Split(·) re-divides the spatiotemporal embedding features into spatial features F s With time feature F t ; S2124, generate a score feature F representing the joint attention intensity of the entire action sequence through the outer product operation of the spatial joint feature score and the temporal feature score att ∈R C×T×V ; Where σ(·) represents the Sigmoid activation function; S2125, through the original input feature F in And the obtained score feature F att Perform outer product and standardization processing to complete feature mapping, and compare the obtained features with the original input features F in Perform residual connection again to obtain the required multi-channel cross-temporal dynamic skeleton joint attention topology F out ∈R C×T×V : (8) in, and represents the channel outer product and element aggregation, ψ(·) represents the Swish activation function, and the shape of the output feature remains unchanged from the input feature.
6. The skeleton fine motion recognition method based on the action-responsive contrast network according to claim 1 is characterized in that: In the step S22, constructing the fine motion comparator FAC of the framework ARCN includes the following steps: S221, decouple features in time and space, and deeply explore the distinguishable hidden features of fine-grained actions across different dimensions; In the step S221, decoupling the features in time and space includes the following steps: S2211, the input feature sequence is spatially and temporally pooled, followed by bottleneck convolution and feature flattening to capture a deep understanding of temporal and spatial variations respectively without interfering with each other: F t ′=ReLU(BN(Conv(Pool s (F in )))) (9) F s ′=ReLU(BN(Conv(Pool t (F in )))) (10) Among them, F in Represents the incoming features, F t ′ represents the temporal features extracted from the sample, F s ′ represents the spatial features extracted from the sample; S222, respectively passing the spatiotemporal decoupled features into the fine action contrast learning to complete the feature contrast learning of the action latent space in the independent spatiotemporal fine action; In the step S222, the steps of respectively transferring the spatiotemporal decoupled features into the fine motion contrast learning include the following steps: S2221, for sample i, the temporal feature F extracted after time-space decoupling t ′ and spatial feature F s ′After the same method of contrast learning, the following uses F i replace; S2222, given an action m, extract the feature F from a sample i i The predicted label results are encoded with one-hot encoding, and then compared with the true label m encoding to determine whether they are true positive TP, false negative FN, and false positive FP. TP refers to the positive example of the action correctly judged by the model. The action label one-hot encoding and the predicted one-hot encoding are element-wise multiplied to obtain the TP of each category of each sample, which is called the same cluster; FN and FP are the results of the model's incorrect judgment. FN means that the model mistakenly predicts the current action as other actions, that is, misjudgment; and FP means that the model mistakenly predicts other actions as the current action. FN and FP together constitute the overall different clusters; S2223. Use exponential moving average to maintain global representation for each cluster, i.e., prototype feature update: Among them, J m is the global representation of the action m feature, is the same cluster sample of action m, is the number of samples in the same cluster, α is the momentum term, and by using the global representation, the features of samples belonging to the same cluster are gathered together to reduce the intra-class differences; S2224. Calculate the feature averages of different clusters and optimize these misclassified samples in the subsequent contrastive learning. The formula is as follows: Among them, F fn 、F fp Represent the average feature representations of the FN sample set and the FP set, respectively, which together constitute the central representation of the heterogeneous clusters. is a heterogeneous FN sample, is a FP sample in a different cluster, and are the number of corresponding samples, respectively; S2225. In order to quantify the similarity between the sample features and the average feature vector, the cosine similarity between the sample features and various prototype features and the average feature representation of different clusters is calculated to obtain the comparative learning score and use it for subsequent loss calculation: Among them, Score m It is the comparative learning score between the prototype feature of the global representation class of action m and the current sample i, which is the basis of contrastive learning. fn 、Score fp is the comparative learning score of FN sample and FP sample with the current sample i respectively; S2226. Although FN samples and FP samples belong to different clusters, for FN samples, they should be made closer to the sample features of the same cluster in the feature space, and the calibration model will make incorrect predictions of fine motion features. For FP samples, they should be made farther away from the sample features of the same cluster in the feature space, and a compensation term φ is introduced for FN samples. i , introduce a penalty term ψ for FP samples i : Adjust the compensation term φ i and the penalty term ψ i It helps to correct the misclassification of FN samples in different clusters and prevent the misclassification of FP samples; S2227, using the various contrastive learning scores obtained above to define the action contrast loss : Among them, p im is the action contrast loss for extracting features from the current action m, is the feature F extracted from the current action m i Action contrast loss; S2228, the complete FAC obtains the time dimension action contrast loss L tac And the spatial dimension action contrast loss L sac , add the two-dimensional action contrast loss to form the action contrast loss L AC : L AC =L tac +L sac (20) Among them, Pool t (·) is the average pooling operation over the time frame; Pool s (·) Average pooling operation on spatial joints; BN represents normalization operation.
7. The skeleton fine motion recognition method based on the action-responsive contrast network according to claim 1 is characterized in that: The step S3 comprises the following steps: S31. In ARGCN, cross entropy loss is used to train the network: Where N is the number of samples in the batch, y ij is the one-hot encoding representation of action sample i, if and only if j is the target action class of sample i, y ij =1, p ij The probability score that the network predicts that sample i belongs to class m; S32. In order to extensively learn the latent space of fine movements, each time through ART, the feature map of the dynamic skeleton joint attention topology of the multi-channel cross-time domain is passed to FAC, forming a multi-stage action contrast loss. The multi-stage action contrast loss is defined as the weighted average sum: Among them, L MAC represents the multi-stage action contrast loss finally obtained by FAC of the entire action-responsive contrast network, i represents the number of stages, a total of 4 stages, respectively in the second layer, the seventh layer, the eleventh layer and the fourteenth layer of ARGCN, L AC is the action contrast loss obtained at each stage of FAC, λ i For learnable parameters, adjust the loss ratio at each stage; S33. Combine the cross entropy loss with the multi-stage action contrast loss to form a complete learning objective function: in, is a learnable hyperparameter used to adjust the importance of multi-stage action contrast loss in ARCN.
Citation Information
Patent Citations
Action recognition method and system based on multi-stream information enhancement graph convolutional network, and storage medium
CN113408455A
Bone action recognition method based on learnable PL-GCN and ECLSTM
CN114529984A