Semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning

Through the semi-supervised skeleton action recognition method of adaptive enhancement and representation learning, combined with push-pull comparison loss and calculation recognition loss, the problem of low recognition accuracy of semi-supervised data in the prior art is solved, and efficient action recognition under a small amount of label data is achieved.

CN114511924BActive Publication Date: 2025-05-06NANJING UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111676997.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-05-06
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The existing deep learning-based skeletal action recognition method is difficult to effectively utilize a small amount of labeled data when processing semi-supervised data, resulting in low recognition accuracy.

Method used

The semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning is adopted. The original skeleton sequence is converted into an enhanced skeleton sequence through the adaptive enhancement module, and the representation features are learned using the residual graph convolution network. At the same time, push-pull comparison loss and calculation identification loss are introduced, and the model is trained through X-CAR loss.

Benefits of technology

This method can effectively learn feature representation under a small amount of labeled data, significantly improving the accuracy of semi-supervised skeleton action recognition, and avoiding the problem of instability in comparison loss in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511924B_ABST
    Figure CN114511924B_ABST
Patent Text Reader

Abstract

The present invention proposes a semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning, comprising: inputting an original skeleton sequence containing joints, bones and differences, converting it into an enhanced skeleton sequence through an enhanced transformation matrix of an adaptive augmentation module; learning the representations of the three types of data in the enhanced skeleton sequence and the original skeleton sequence respectively through a residual graph convolution network in a residual convolution encoder, then fusing the three types of data in the two groups of skeleton sequences and performing unified representation learning to obtain enhanced and original spatiotemporal action feature representations; inputting the enhanced and original spatiotemporal action features into a multi-layer perceptron MLP to obtain a push-pull contrast loss, and then obtaining the calculation recognition loss through a fully connected layer; combining the push-pull contrast loss and the recognition loss as an X-CAR loss, and learning and training the adaptive augmentation module and the residual graph convolution encoder. This method can effectively improve the accuracy of action recognition based on semi-supervised skeletons.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of skeleton-based action recognition in the field of computer vision, and in particular to a semi-supervised skeleton action recognition method based on adaptive augmentation and representation learning. Background Art

[0002] Human action recognition is a fundamental but challenging task in computer vision, playing an important role in many real-world applications such as video surveillance, human-computer interaction, video retrieval, virtual reality, etc. Compared with raw RGB videos, human skeleton sequences consisting of accurate joint positions are more robust to human scale, dynamic environment, camera viewpoint, and interfering background. In addition, human skeleton sequences can be regarded as a high-level abstract representation of an action, which has more advantages in computing and storage. Therefore, skeleton-based action recognition has attracted widespread attention.

[0003] With the development of deep learning, action recognition methods based on deep learning have achieved remarkable performance by using convolutional neural networks (CNN) or recurrent neural networks (RNN) to learn the representation of skeleton sequences based on specific recognition tasks, but they still cannot meet practical needs. The main challenge is that existing models rarely explore the interdependence between body joints and ignore a large amount of motion information, which leads to methods that regard skeleton sequences as spatiotemporal graphs and extract features through graph convolutional networks. However, most of these existing methods are fully supervised and heavily rely on a large amount of labeled training data, but a large amount of labeled data is difficult to obtain in many cases. Therefore, how to effectively learn feature representations from a small amount of labeled skeleton data remains a problem worthy of attention. Summary of the invention

[0004] The purpose of the present invention is to provide a semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning, which can effectively learn features from a small amount of labeled skeleton data and effectively improve the accuracy of action recognition based on semi-supervised skeleton.

[0005] The technical solution to achieve the purpose of the present invention is: a semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning, comprising the following steps:

[0006] Step 1, input the original skeleton sequence containing joints, bones and differences, and convert it into an enhanced skeleton sequence through the enhanced transformation matrix of the adaptive enhancement module;

[0007] Step 2: The residual graph convolution network in the residual convolution encoder is used to learn the representations of the three types of data, namely, joints, bones, and differences, in the enhanced skeleton sequence and the original skeleton sequence, respectively. Then, the three types of data in the two sets of skeleton sequences are fused and unified representation learning is performed to obtain the enhanced spatiotemporal motion feature and the original spatiotemporal motion feature representation.

[0008] Step 3: Input the enhanced spatiotemporal motion features and the original spatiotemporal motion features into the multi-layer perceptron MLP to obtain the push-pull contrast loss, and then obtain the calculation recognition loss through the fully connected layer;

[0009] In step 4, the push-pull contrast loss and the recognition loss are combined as the X-CAR loss, and the adaptive enhancement module and the residual graph convolution encoder are trained at the same time.

[0010] Compared with the prior art, the present invention has the following beneficial effects: the present invention replaces the traditional two-stage method with a one-stage method, obtains a learnable enhancement method to promote the consistency of enhancement and representation, and avoids the problem of unstable results caused by blurred samples caused by pre-defined positive and negative samples through a new push-pull contrast loss, thereby effectively improving the accuracy of action recognition based on semi-supervised skeleton. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 The flowchart of the semi-supervised skeletal action recognition method of the present invention.

[0012] Figure 2 Schematic diagram of rotation, shearing and scaling enhancements in the method of the present invention.

[0013] Figure 3 (a) is a schematic diagram of the original skeleton features on the NW-UCLA dataset. Figure 3 (b) is a schematic diagram of the skeleton features learned by the traditional contrast learning method. Figure 3 (c) is a t-SNE visualization diagram of the skeleton features learned by X-CAR of the present invention.

[0014] Figure 4 This is a visualization result of some samples learned by adaptively combining enhanced AA. DETAILED DESCRIPTION

[0015] The present invention will be further described in detail below in conjunction with the accompanying drawings:

[0016] A semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning includes four steps: obtaining a learnable augmented skeleton sequence, obtaining spatiotemporal action features, calculating the push-pull contrast loss and action recognition loss, and combining PPCL and recognition loss to simultaneously learn and train the augmentation and representation. The specific process is as follows Figure 1 As shown. Among them:

[0017] Obtaining the enhanced skeleton sequence includes the following steps:

[0018] Step 1) Input the original skeleton sequence S containing three types of data: joints, bones and differences into the adaptive combination enhancement module (AA).

[0019] Step 2) According to the characteristics of the skeleton, three types of enhanced transformations are defined, namely rotation, shearing and scaling. Examples of the three types of transformations are shown in Figure 2 .

[0020] a. Perform rotation enhancement on all types of data in step 1). The rotation enhancement transformation matrix is ​​defined as follows:

[0021]

[0022] where r ro Indicates the basic enhancement value of rotation enhancement, R ro is the rotation enhancement transformation matrix, Enhance the transformation matrix for rotation around the X axis, Enhance the transformation matrix for rotation around the Y axis, Augment the transformation matrix for rotation around the Z axis.

[0023] The coordinate transformation for rotation enhancement is as follows:

[0024]

[0025] in represents the learnable control factor of rotation enhancement, i is the i-th skeleton sequence in each training, and [x, y, z] is the coordinate before rotation enhancement transformation.

[0026] b. Use the shear enhancement transformation matrix to perform shear enhancement on the skeleton sequence obtained in step a. The shear enhancement transformation matrix is ​​defined as follows:

[0027]

[0028] where r sh Indicates the base enhancement value for shear enhancement.

[0029] The coordinate transformation of shear enhancement is as follows:

[0030]

[0031] in Represents the learnable control factor of shear enhancement, and [x, y, z] is the coordinate before shear enhancement transformation.

[0032] c. Use the scaling enhancement transformation matrix to scale and enhance the skeleton sequence obtained in step b. The scaling enhancement transformation matrix is ​​defined as follows:

[0033]

[0034] where r sc Indicates the base boost value for the scaling boost.

[0035] The coordinate transformation for zoom enhancement is as follows:

[0036]

[0037] in Represents the learnable control factor of scaling enhancement, and [x, y, z] is the coordinate before scaling enhancement transformation.

[0038] Examples of skeleton motions obtained after three types of transformations are as follows: Figure 4 , Figure 4 The first column in Fig. 1 is three skeleton actions from the NTU RGB+D dataset, namely, using a fan, putting the palms together, and checking the time (from a watch). Figure 4 The second column in Figure 1 shows three skeleton actions (i.e., standing up, picking up with one hand, and putting on) from the NW-UCLA dataset. Each group includes the transformation from the original skeleton to the enhanced skeleton, as well as the learned enhancement control factors.

[0039] Step 3), after scaling enhancement, an enhanced skeleton sequence S' containing three types of data: joints, bones, and differences is obtained. In this step, the rotation, shear, and scaling transformations are defined in a parametric way, so that the three types of enhancement transformations can be learned, and the one-stage method is used to simultaneously learn adaptive combined enhancement and feature representation, replacing the traditional two-stage method.

[0040] Obtaining spatiotemporal action features includes the following steps:

[0041] Step 4) A spatial graph convolutional network (SGCN), a temporal graph convolutional network (TGCN) and some dense residual connections are combined to form a residual graph convolution module, namely the residual graph convolution R-GCN module. The formula of SGCN is defined as follows:

[0042]

[0043] Where D is the predefined maximum graph distance, A is the skeleton matrix, Λ j is a diagonal matrix, For normalization of α j , i, k are the i-th row and k-th column of the skeleton matrix A, f in Represents the input of the graph, W j Represents the learnable weight parameters.

[0044] TGCN is a normal L×1 convolutional layer that aggregates contextual representations embedded in adjacent frames. R-GCN can not only capture the spatiotemporal context of skeletal actions, but also benefit from its dense residual connections, making the training process more convergent.

[0045] Step 5), connect the BatchNorm layer (BN) with the R-GCN network obtained in step 4) as a branch for processing three types of data.

[0046] Step 6), respectively input the enhanced skeleton data obtained in step 3) into the branches obtained in step 5), perform early fusion on the contents of each branch, and obtain the representation of the enhanced data; input the original skeleton data into another branch obtained in step 5), perform early fusion on the contents of each branch, and obtain the representation of the original data.

[0047] Step 7), connect the branch obtained in step 5) with 6 R-GCN modules. Input the representation of the enhanced skeleton sequence and the original skeleton sequence obtained in step 6) into the backbone R-GCN network block respectively to obtain the feature representation of the two sets of sequences.

[0048] The interconnected BN, R-GCN, fusion module and backbone R-GCN blocks in step 8), step 4), step 5), step 6) and step 7) constitute the residual graph convolution encoder, denoted as R-GCE F(·). R-GCE retains the specific features of the three types of data and increases the complementarity of the final representation. The R-GCE for processing the enhanced sequence and the R-GCE for processing the original sequence are defined as branch A and branch B respectively, and the features obtained after processing by branch A and branch B are denoted as f a and f b .

[0049] Calculating the push-pull contrast loss and action recognition loss involves the following steps:

[0050] Step 9), the MLP module is obtained by connecting the linear layer, the BatchNorm (BN) layer and the nonlinear activation function (ReLU). The present invention uses two MLP modules, which are respectively denoted as MLP H(·) and MLP G(·). The two modules have the same architecture, but their parameters are not shared.

[0051] Step 10), respectively inputting the two sets of sequence representations obtained in step 7) into the MLP H(·) and MLP G(·) modules obtained in step 9) to obtain representations;

[0052] Step 11), Definition As does not include s i The representation of other skeleton sequences of The distance between representations is calculated using the mean square error, and the stopping gradient strategy is used to learn to promote the consistency between enhancement and representation. A new push-pull contrast loss (PPCL) is formulated as follows:

[0053]

[0054] where h a =H(F(s' i )),g b =G(F(s i )), N is the number of skeleton sequences for each training. r j ∈Ω( / s i ), which is the number of S that does not include s i The representation of other skeleton sequences of , i is the i-th skeleton sequence in each training, s i ∈S,s' i ∈S', H(·) and G(·) represent the MLP functions corresponding to the enhanced spatiotemporal action features and the original spatiotemporal action features, respectively, and F(·) represents the graph convolutional encoder function, which is composed of the above-mentioned BatchNorm, residual graph convolutional network, fusion module and multiple residual graph convolutional blocks.

[0055] PPCL encourages enhanced samples to move closer to their same samples, but pushes other samples away.

[0056] Step 12), the enhanced data and the original data obtained in step 3) are input into branch B and branch A in step 8) respectively, and the above steps are repeated to obtain the symmetric contrast loss, and the push-pull contrast loss is optimized, that is:

[0057]

[0058] where h b =H(F(s i )), g a =G(F(s' i )), r j ∈Ω( / s i ), which is S does not include s i The representation of other skeleton sequences, N is the number of skeleton sequences in each training, i is the i-th skeleton sequence in each training, s i ∈S,s' i ∈S'.

[0059] Step 13), sum the push-pull contrast loss obtained in step 11) and step 12) to obtain the final push-pull contrast loss PPCL, that is:

[0060]

[0061] Step 14), the two sets of features f obtained in step 7) are a and f b Input the fully connected layer (FC) to obtain the final classification representation feature vector p a =FC(f a ), p b =FC(f b ) and action recognition loss. The action recognition loss is defined as follows:

[0062]

[0063] Where y is the label of the action, AP(·) is the average pooling operation, p a =FC(f a ), p b =FC(f b ) is the classification representation feature vector.

[0064] And integrate PPCL and identification loss X-CAR loss as follows:

[0065]

[0066] Combining PPCL and recognition loss to simultaneously learn enhancements and representations includes the following steps:

[0067] Step 15), in this invention, we train the model by minimizing L using SGD and stopping gradient strategy to update all parameters of X-CAR.

[0068] This method uses a learnable enhancement method to promote the consistency of enhancement and representation, and avoids the problem of unstable results caused by blurred samples due to pre-defined positive and negative samples through a new contrast loss. The effect is as follows Figure 3 As stated, Figure 3 (a) is a schematic diagram of the original skeleton features on the NW-UCLA dataset. Figure 3 (b) is a schematic diagram of the skeleton features learned by the traditional contrast learning method. Figure 3(c) is a t-SNE visualization diagram of the skeleton features learned by X-CAR of the present invention; the recognition accuracy of different methods is shown in Tables 1 and 2, Table 1 shows the recognition accuracy (%) obtained by different methods on the NTU RGB+D dataset (Cross-Subject (CS) and Cross-View (CV)), and the training set labeled data are 5%, 10%, 20% and 40% respectively; Table 2 shows the recognition accuracy (%) obtained by different methods on the NW-UCLA dataset, and the training set labeled data are 5%, 15%, 30% and 40% respectively. It can be seen from Tables 1 and 2 that the present invention can effectively improve the accuracy of action recognition based on semi-supervised skeleton.

[0069] Table 1

[0070]

[0071] Table 2

[0072]

Claims

1. A semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning, characterized in that: The following steps are involved: Step 1, input the original skeleton sequence containing joints, bones and differences, and convert it into an enhanced skeleton sequence through the enhanced transformation matrix of the adaptive enhancement module; Step 2: The residual graph convolution network in the residual convolution encoder is used to learn the representations of the three types of data, namely, joints, bones, and differences, in the enhanced skeleton sequence and the original skeleton sequence, respectively. Then, the three types of data in the two sets of skeleton sequences are fused and unified representation learning is performed to obtain the enhanced spatiotemporal motion feature and the original spatiotemporal motion feature representation. Step 3: Input the enhanced spatiotemporal motion features and the original spatiotemporal motion features into the MLP module to obtain the push-pull contrast loss, and then obtain the calculation recognition loss through the fully connected layer; Step 4: Combine the push-pull contrast loss and the recognition loss as the X-CAR loss, and simultaneously train the adaptive enhancement module and the residual graph convolution encoder; The residual graph convolution encoder is formed by connecting BatchNorm, the residual graph convolution network, a fusion module and multiple residual graph convolution blocks, and the step 2 specifically includes: The enhanced skeleton sequence and the original skeleton sequence are respectively input into the BatchNorm layer and the residual graph convolutional network to process the three types of data, and the representations of the three types of data, namely, joints, bones and differences, are obtained, and early fusion is performed through the fusion module to obtain the data representations of the enhanced and original skeleton sequences; The data representations of the enhanced and original skeleton sequences are input into multiple residual graph convolution blocks to obtain the spatiotemporal action feature representations of the two sequences f a and f b , the number of residual graph convolution blocks is 6; The X-CAR loss is: is the push-pull contrast loss, To identify losses.

2. The semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning according to claim 1, characterized in that: The step of converting the enhanced skeleton sequence into an enhanced skeleton sequence by using the enhanced transformation matrix of the adaptive enhancement module specifically includes: The original skeleton sequence S is rotationally enhanced by using a rotation enhancement transformation matrix to obtain a rotation enhanced skeleton sequence; Perform shear enhancement on the rotation enhanced skeleton sequence through the shear enhancement transformation matrix to obtain a shear enhanced skeleton sequence; The shear enhanced skeleton sequence is scaled and enhanced by using a scaling enhancement transformation matrix to obtain a final enhanced skeleton sequence S'.

3. The semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning according to claim 2, characterized in that: The rotation enhancement transformation matrix is: where r ro Indicates the basic enhancement value of rotation enhancement, R ro is the rotation enhancement transformation matrix, Enhance the transformation matrix for rotation around the X axis, Enhance the transformation matrix for rotation around the Y axis, Enhance the transformation matrix for rotation around the Z axis; The coordinates after rotation enhancement transformation are: in represents the learnable control factor for the rotation enhancement around the coordinate axis, i is the i-th skeleton sequence in each training, and [x, y, z] is the coordinate before the rotation enhancement transformation.

4. The semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning according to claim 3, characterized in that: The shear enhancement transformation matrix R sh for: where r sh represents the base enhancement value of shear enhancement; The coordinates after shear enhancement transformation are: in Represents the learnable control factor of shear enhancement, and [x, y, z] is the coordinate before shear enhancement transformation.

5. The semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning according to claim 4, characterized in that: The scaling enhancement transformation matrix is: where r sc Indicates the base enhancement value of the scaling enhancement; The coordinates after scaling and enhancement transformation are: in Represents the learnable control factor of scaling enhancement, and [x, y, z] is the coordinate before scaling enhancement transformation.

6. The semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning according to claim 1, characterized in that: The residual graph convolution network is a residual graph convolution R-GCN module consisting of a spatial graph convolution network SGCN, a temporal graph convolution network TGCN and multiple dense residual connections, where the temporal graph convolution network TGCN is an L×1 convolution layer that aggregates the context representation embedded in adjacent frames, and its output is: Where D is the predefined maximum graph distance, A is the skeleton matrix, Λ j is a diagonal matrix, For normalization of α j , i, k are the i-th row and k-th column of the skeleton matrix A, f in Represents the input of the graph, W j Represents the learnable weight parameters.

7. According to the semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning according to claim 1, the MLP module is obtained by connecting two groups of linear layers, BatchNorm and nonlinear activation functions.

8. The semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning according to claim 7, characterized in that: The enhanced spatiotemporal motion features and the original spatiotemporal motion features are respectively input into the MLP module to obtain the push-pull contrast loss, and then the recognition loss is calculated through the fully connected layer, specifically including: The enhanced and original spatiotemporal action feature representations are input into the corresponding MLP modules to obtain representations; The distance between the representations is calculated using the mean square error, and the push-pull contrast loss is obtained by stopping the gradient strategy: where h a =H(F(s' i )),g b =G(F(s i )), N is the number of skeleton sequences for each training. r j ∈Ω( / s i ), which is S does not include s i The representation of other skeleton sequences of , i is the i-th skeleton sequence in each training, s i ∈S,s' i ∈S', H(·) and G(·) represent the MLP functions corresponding to the enhanced spatiotemporal action features and the original spatiotemporal action features, respectively, and F(·) represents the graph convolutional encoder function; The original skeleton sequence and the enhanced skeleton sequence obtained in step 1 are input into the residual convolution encoder respectively, and the above steps are repeated to obtain the symmetrical push-pull contrast loss: where h b =H(F(s i )), g a =G(F(s' i )), r j ∈Ω( / s i ), which is S does not include s i The representation of other skeleton sequences, N is the number of skeleton sequences in each training, i is the i-th skeleton sequence in each training, s i ∈S,s' i ∈S'; Will and Sum, the final push-pull contrast loss is: The enhanced spatiotemporal action feature obtained in step 2 is represented as f a and the original spatiotemporal action feature representation f b Input to the fully connected layer to obtain the final classification representation feature vector and action recognition loss. The action recognition loss is: Where y is the label of the action, AP(·) is the average pooling operation, p a =FC(f a ), p b =FC(f b ) is the classification representation feature vector.

9. The semi-supervised skeletal action recognition method based on adaptive augmentation and representation learning according to claim 1, characterized in that: The learning and training of the adaptive enhancement module and the residual graph convolution encoder is specifically as follows: minimizing the X-CAR loss and updating all parameters of X-CAR to learn and train the model through SGD and stopping gradient strategy.

Citation Information

Patent Citations

  • Graph convolution action recognition method, device and equipment based on 2S-AGCN

    CN113642400A