A Human Skeleton Action Recognition Method Based on Cross-Scale Graph Contrastive Learning
Through the cross-scale graph comparison learning method, the graph data enhancement and consistency knowledge mining is used to use labeled skeleton data to build a self-supervised action recognition network, solving the problem of limited recognition performance caused by relying on labeled data in the existing technology, and achieving more efficient human skeleton action recognition.
Patent Information
- Application Number
- CN202211385326.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-11-07
AI Technical Summary
The existing human skeleton action recognition model relies on labeled RGB video data, which is susceptible to occlusion, environmental changes and shadow interference, resulting in limited recognition performance, and labeling data are cumbersome and expensive, and it is not possible to effectively use labelless data for training.
The cross-scale graph comparison learning method is adopted to build the final model through the acquisition of label-free skeleton sequences, graph data enhancement, encoder network coding, graph comparison self-supervised action recognition network and cross-scale consistency knowledge mining, and the final model recognition performance is constructed using labeled data fine-tuning.
It improves the accuracy and generalization ability of human skeleton motion recognition, solves the problem of limited model recognition performance under label-free data, and achieves more efficient motion recognition effect.
Smart Images

Figure CN115661718B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of self-supervised learning, and particularly to a method for human skeleton action recognition based on cross-scale graph contrast learning. Background Art
[0002] Action recognition plays a crucial role in many application scenarios such as video surveillance, human-computer interaction, video understanding, and so on.
[0003] Existing human skeleton action recognition models generally process and analyze labeled RGB video data to obtain human action information in the video.
[0004] However, when extracting RGB video data, it is vulnerable to occlusion, environmental changes, and shadow interference, resulting in the easy loss of color and texture features in the depth map, and it is relatively time-consuming to process. And the human action recognition model composed of labeled data belongs to the full-supervised learning framework, which requires a large amount of manually labeled data, and the labeled data is cumbersome and expensive. Therefore, the recognition performance of the model based on supervised learning will be limited. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for human skeleton action recognition based on cross-scale graph contrast learning, aiming to solve the problem of low accuracy of human skeleton action recognition in the absence of labeled data.
[0006] To achieve the above purpose, the present invention provides a method for human skeleton action recognition based on cross-scale graph contrast learning, including the following steps:
[0007] Collect and obtain unlabeled skeleton sequences;
[0008] Uniformly plan the continuous frame numbers of the unlabeled skeleton sequences;
[0009] Use the idea of graph data augmentation to obtain different instances of the unlabeled skeleton sequences;
[0010] Encode different instances through an encoder network to obtain encoded features, and establish a graph contrast self-supervised action recognition network;
[0011] Combine the cross-scale consistency knowledge mining method for interaction between multi-scale information;
[0012] Based on the graph contrast self-supervised action recognition network and the cross-scale consistency knowledge mining method, obtain the final model;
[0013] Fine-tune the parameters of the final model using labeled training data, and obtain the recognition performance of the final model based on the linear evaluation protocol.
[0014] Among them, the unlabeled skeleton sequence is obtained by collecting a human skeleton dataset from different perspectives using a camera.
[0015] Among them, in the process of uniformly planning the continuous number of frames of the unlabeled skeleton sequence, in order to avoid data redundancy and reduce computational complexity, in the human skeleton dataset, the continuous number of frames of the skeleton sequence is uniformly planned.
[0016] Among them, the process of obtaining different instances of the unlabeled skeleton sequence using the idea of graph data augmentation is specifically to use data augmentation methods including graph data augmentation on the same skeleton sequence to obtain different instances of the unlabeled skeleton sequence, and use the different instances as the positive sample set.
[0017] Among them, the encoder network is a graph convolutional neural network for extracting the encoded features of the different instances; the graph contrast self-supervised action recognition network consists of two paths, namely the original path and the graph contrast path, and each path can be further divided into a data augmentation module, an encoder module, and a projection layer module, where the projection layer module is composed of a fully connected layer and a linear layer.
[0018] Among them, in the process of combining the cross-scale consistency knowledge mining method for the interaction between multi-scale information, multi-scale graph modeling is used to represent three-dimensional bone features, and the key relevant features of bone joints are aggregated. The specific way to achieve the interaction between multi-scale information is: convert the original skeleton sequence into a multi-scale skeleton graph sequence, combine the cross-scale consistency knowledge mining module, and use the similarity of feature information in one scale graph to promote the effective clustering of similar features in another scale graph.
[0019] Among them, based on the graph contrast self-supervised action recognition network and the cross-scale consistency knowledge mining method, the specific way to obtain the final model is to input the multi-scale skeleton graph sequence into the constructed network model, and through the collaborative association mode between individual multi-scale mappings, achieve the intra-scale association and the semantic collaborative interaction between multi-scales.
[0020] Among them, the process of fine-tuning is specifically to attach a linear classifier to the pre-trained encoder.
[0021] Among them, the linear evaluation protocol is to append a classifier after the pre-trained encoder. The classifier consists of a fully connected layer and a non-linear layer, and the network of the fixed encoder is trained.
[0022] Using the fine-tuning method, through the use of labeled sample data, the entire feature encoder and the linear layer are trained end-to-end to obtain a well-performing action recognition model; using the linear evaluation technique, by fixing the parameters of the pre-trained encoder and training a linear classifier on it, the learned action representation is evaluated.
[0023] The present invention provides a human skeleton action recognition method based on cross-scale graph contrast learning. Based on a graph contrast self-supervised action recognition network and a cross-scale consistency knowledge mining method, a final model is obtained. Then, the parameters of the final model are fine-tuned using labeled training data, and the recognition performance of the final model is obtained based on a linear evaluation protocol. By fully utilizing the graph contrast learning method, when augmenting unlabeled skeleton data, the edges constituting the skeleton structure are randomly cropped, enhancing the expression of high-level semantic information and enabling the model to obtain better generalization ability. Secondly, by using the method of aggregating bone joint points, skeleton graphs of multiple scales are constructed. Through cross-scale perceptual consistency, the nearest neighbor mining strategy is further improved, making the learning process more reasonable and thus enhancing the recognition performance, solving the problem that the existing human skeleton action recognition methods do not use unlabeled data to train the model, resulting in limited recognition performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0025] Figure 1 It is a schematic flowchart of a human skeleton action recognition method based on cross-scale graph contrast learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as a limitation of the present invention.
[0027] Please refer to Figure 1 , the present invention provides a human skeleton action recognition method based on cross-scale graph contrast learning, including the following steps:
[0028] S1: Collect and obtain unlabeled skeleton sequences;
[0029] S2: Uniformly plan the continuous frame numbers of the unlabeled skeleton sequences;
[0030] S3: Use the idea of graph data augmentation to obtain different instances of the unlabeled skeleton sequences;
[0031] S4: Encode different instances through an encoder network to obtain encoded features, and establish a graph contrast self-supervised action recognition network;
[0032] S5: Combine the cross-scale consistency knowledge mining method to perform the interaction between multi-scale information;
[0033] S6: Based on the graph contrast self-supervised action recognition network and the cross-scale consistency knowledge mining method, obtain the final model;
[0034] S7: Fine-tune the parameters of the final model using the labeled training data, and obtain the recognition performance of the final model based on the linear evaluation protocol.
[0035] The following is an illustration in combination with specific implementation steps:
[0036] S1: Collect and obtain unlabeled skeleton sequences;
[0037] In step S1, use a camera system to obtain unlabeled 3D skeleton sequences, and capture a human skeleton dataset from different perspectives based on the camera to obtain unlabeled 3D skeleton sequences.
[0038] Specifically, given a continuous 3D skeleton sequence X = (X1, …, X l ), where X i ∈ R W×J×D , W is the total number of people, J is the number of bone joints, D is the dimension of the position vector (the position vector dimension of X i is 3). The training set contains skeleton sequences of B different actions collected from multiple views and multiple people. Each skeleton sequence X i corresponds to a label y i , where y i ∈ {a1, …, a c}, a i represents the i-th action category, c represents the total number of action categories, and the size of the sample data batch size input into the network each time is N.
[0039] S2: Uniformly plan the continuous frame numbers of the unlabeled skeleton sequences;
[0040] Specifically, obtain the 3D skeleton sequence X (i.e., the coordinate data matrix of the bone points), and the dimension of this matrix is [N, D, l, J, W]. To avoid data redundancy and reduce the computational complexity, in the human skeleton dataset, uniformly set the continuous frame number l of the skeleton sequence to 50, the batch size N = 128, the position vector dimension D = 3, the number of bone joints J = 25, and the total number of people W = 2.
[0041] S3: Obtain different instances of the unlabeled skeleton sequence using the idea of graph data augmentation;
[0042] In step S3, it specifically further includes the following two steps:
[0043] S31: Use the data augmentation module τ to obtain different instances Q;
[0044] Specifically, on the basis of obtaining the skeleton data, introduce the data augmentation methods of Shear and Temporal Crop on the original path and the graph comparison path respectively to obtain different views Q and The specific method is outlined as follows.
[0045] Shear: Shearing transformation is to make the three-dimensional coordinate shape of human joints tilt at an arbitrary angle by constructing the corresponding affine matrix. The formula of the affine matrix is:
[0046]
[0047] where are 6 shearing factors, and the value range is between -β and β.
[0048] Temporal Crop: It is data augmentation in the time dimension. It symmetrically fills some frames into the sequence and then randomly shears it to the original length. The filling length is defined as l / r, and r is the filling ratio (taking positive integer values).
[0049] S32: Use the data augmentation module τ + mask_edg to obtain different instances K.
[0050] Specifically, perform random edge cropping (mask_edg) on the view to obtain different graph representation vectors K. Specific idea: Use the random mask [0~ξ] to remove the connection edges between joint points to form a new skeleton graph structure.
[0051] S4: Encode different instances through the encoder network to obtain encoded features, and establish a graph contrast self-supervised action recognition network;
[0052] The specific method of step S4 is as follows:
[0053] S41: Use the encoder module to obtain encoded features;
[0054] Specifically, embed different instances Q and K into the encoders f θ and respectively to obtain the encoded features h and where θ and are the parameters required for the two encoders, following momentum update: α is the momentum coefficient, h, Use a graph convolutional neural network as the encoder network.
[0055] S42 inputs the encoded features into the projection layer module to obtain a lower-dimensional spatial feature vector;
[0056] Input the obtained encoded feature h and into the projection layer g and respectively, to obtain a lower-dimensional spatial vector: z = g(h), where z, The projection layer is composed of a fully connected (FC) layer and a linear (ReLU) layer.
[0057] S43 performs network splicing based on the above modules to construct a graph contrast self-supervised action recognition framework.
[0058] Specifically, in the graph contrast learning process, when a skeleton sequence is input into two different paths with different instances, the output features are similar. During the model training process, the following loss needs to be minimized:
[0059]
[0060] Among them, the repository stores a large number of negative samples, avoiding redundant calculations of embeddings. It is a first-in-first-out queue and is updated by each time. Specifically, after each update iteration, will enter the queue to become a new negative sample, while the feature of the earliest embedded sample in M will exit the queue. M i ∈M is the negative sample set in the repository, t is a hyperparameter, represents the dot product of two vectors, and its result indicates the similarity degree between two instances, where z and have been normalized.
[0061] S5: Combine the cross-scale consistency knowledge mining method to perform the interaction between multi-scale information;
[0062] In step S5, it also includes using multi-scale graph modeling to represent three-dimensional skeletal features and aggregating the key relevant features of skeletal joint points.
[0063] The specific process of realizing the interaction between multi-scale information is as follows: Given a skeleton sequence X containing l frames, which is called the joint point scale (i.e., body joints as nodes), denoted as Θ 1, construct a coarse-grained proportional graph, divide the skeleton structure of the mover into 10 parts (including the torso, head, upper right arm, lower right arm, upper left arm, lower left arm, upper right leg, lower right leg, upper left leg, and lower left leg) and 5 parts (including the torso, right upper limb, left upper limb, right lower limb, and left lower limb), and average the position coordinates of the bone joint points included in each part, merge them into new bone joint points, and name them the coarse joint point scale (i.e., the body part as the joint node), denoted as Θ 2 and Θ 3 . Finally, based on the above operations, skeleton graphs Θ of different scales are obtained m (V m , ε m )(m ∈ {1, 2, 3}), where represents the set of joint points corresponding to different skeleton scale graphs, represents the set of edge structure relationships of different skeleton scale graphs, and n m is the number of joint points of the m-th scale graph Θ m . Given that the multi-scale data of the skeleton is composed of merged joint points, resulting in a change in the graph structure and cannot be directly input into the graph contrast learning module established above. Therefore, corresponding graph structures will be constructed for different scale skeleton graphs
[0064] S6: Based on the graph contrast self-supervised action recognition network and the cross-scale consistency knowledge mining method, obtain the final model
[0065] Specifically, as an extension of the previous method, after the training of the graph contrast learning network is completed, cross-scale graph contrast learning is carried out to obtain stronger learning representation ability and avoid misclassification when the network is trained from scratch. Specifically, given a skeleton sequence X, two different scale graphs need to be obtained where Θ 1 and Θ 2 are different scale skeleton graphs composed of 25 and 10 bone joint points respectively, and represent the skeleton sequences corresponding to the respective scales. The purpose of cross-scale graph contrast representation learning is to learn and with good generalization ability, where is the feature representation of Θ 1 and Θ 2 , which can effectively perform various downstream tasks. The main idea is different from the graph contrast learning method in that it is necessary to reconstruct the positive sample set and negative sample set, that is, positive samples that are difficult to find in the Θ 1 scale can be found in Θ 2 . The multi-scale data Θ m (V m , εm (m ∈ {1, 2}) is input into the multi-scale graph contrastive learning network, and graph encoding features are obtained through two graph contrastive learning network modules with different graph structures. and the corresponding repository As the training progresses, the representation learning ability of the model is gradually enhanced.
[0066] During the model training process, a contrastive loss function is used to update the parameters, and the formula is as follows:
[0067]
[0068] where t is a hyperparameter, is the negative sample set in the repository. The numerator contains 1 + k positive samples, and the denominator contains a total of N + 1 samples including 1 + k positive samples and N - k negative samples. k is the index of the similar sample feature embedding, which is selected by the topk(·) function. In the experiment, the value of k is taken as 1.
[0069] Similarly, instances similar in the Θ 1 scale feature space can also be used as pseudo-labels to help the network at the Θ 2 scale perform better representation learning. Its loss function is as follows:
[0070]
[0071] The meaning of its parameters is the same as that of The two networks sample positive samples for each other to enhance the performance of the network model and obtain a better clustering effect.
[0072]
[0073] Compared with the single-scale loss function L1, the multi-scale loss function L2 pulls closer more high-confidence positive samples, making it easier for the features of the same-class samples in the feature space to aggregate.
[0074] S7: Use the labeled training data to fine-tune the parameters of the final model, and obtain the recognition performance of the final model based on the linear evaluation protocol.
[0075] Specifically, fine-tuning is to add a linear classifier to the learnable encoder, and then train the entire model to complete the action recognition task; the linear evaluation protocol is to attach the frozen encoder to a linear classifier (a fully connected (FC) layer and a non-linear (softmax) layer), and then perform supervised training on the classifier to verify the recognition performance of the final model.
[0076] Furthermore, the present invention intends to construct a human skeleton action recognition network framework based on self-supervised learning starting from a human skeleton dataset and a self-supervised learning method. However, it is noted that most current self-supervised models based on skeleton data use contrastive learning methods for modeling, without considering that skeleton data is a discrete data structure that requires graph structure learning, and the idea of using data augmentation to obtain positive samples is too single, and the cross-scale information joint method is less applied to self-supervised models, making it difficult to overcome the defect of insufficient single-scale feature information and being unfavorable to the clustering effect of the model. Therefore, the present invention proposes a self-supervised action recognition method based on graph contrast learning and cross-scale consistency knowledge mining, which realizes the intra-scale association and inter-scale semantic collaborative interaction through the collaborative association pattern between individual multi-scale mappings.
[0077] The above-disclosed is only a preferred embodiment of the present invention, and of course, it cannot be used to limit the scope of the rights of the present invention. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A human skeleton action recognition method based on cross-scale graph contrastive learning, characterized in that It includes the following steps: Collect and obtain unlabeled skeleton sequences; Uniformly plan the consecutive frame numbers of the unlabeled skeleton sequences; Use the idea of graph data augmentation to obtain different instances of the unlabeled skeleton sequences; Encode different instances through an encoder network to obtain encoded features, and establish a graph contrast self-supervised action recognition network; The encoder network is a graph convolutional neural network for extracting the encoded features of different instances; the graph contrast self-supervised action recognition network consists of two paths, namely the original path and the graph contrast path, and each path includes a data augmentation module, an encoder module, and a projection layer module, where the projection layer module is composed of a fully connected layer and a linear layer; Combine the cross-scale consistency knowledge mining method for interaction between multi-scale information; In the process of combining the cross-scale consistency knowledge mining method for interaction between multi-scale information, use multi-scale graph modeling to represent three-dimensional skeletal features, aggregate key relevant features of skeletal joints, and the specific way to achieve interaction between multi-scale information is: convert the original skeleton sequence into a multi-scale skeleton graph sequence, combine the cross-scale consistency knowledge mining module, and use the similarity of feature information in one scale graph to promote the effective clustering of similar features in another scale graph; Based on the graph contrast self-supervised action recognition network and the cross-scale consistency knowledge mining method, obtain the final model; Fine-tune the parameters of the final model using labeled training data, and obtain the recognition performance of the final model based on the linear evaluation protocol.
2. The method for human skeleton action recognition based on cross-scale graph contrast learning according to claim 1, wherein The unlabeled skeleton sequences are obtained by collecting a human skeleton dataset from different perspectives using a camera.
3. The method for human skeleton action recognition based on cross-scale graph contrast learning according to claim 2, wherein In the process of uniformly planning the consecutive frame numbers of the unlabeled skeleton sequences, to avoid data redundancy and reduce computational complexity, in the human skeleton dataset, uniformly plan the consecutive frame numbers of the skeleton sequences.
4. The method for human skeleton action recognition based on cross-scale graph contrast learning according to claim 3, wherein The process of using the idea of graph data augmentation to obtain different instances of the unlabeled skeleton sequences is specifically to use data augmentation methods including graph data augmentation on the same skeleton sequence to obtain different instances of the unlabeled skeleton sequences, and use these different instances as the positive sample set.
5. The method for human skeleton action recognition based on cross-scale graph contrast learning according to claim 4, wherein The specific way to obtain the final model based on the graph contrast self-supervised action recognition network and the cross-scale consistency knowledge mining method is to input the multi-scale skeleton graph sequence into the constructed final model, and through the collaborative association mode between individual multi-scale mappings, achieve intra-scale association and inter-scale semantic collaborative interaction.
6. The method for human skeleton action recognition based on cross-scale graph contrast learning according to claim 5, wherein The process of fine-tuning is specifically to attach a linear classifier to the pre-trained encoder.
7. The human skeleton action recognition method based on cross-scale graph contrast learning according to claim 6, characterized in that the linear evaluation protocol adds a classifier after the pre-trained encoder, the classifier consists of a fully connected layer and a non-linear layer, and the network of the fixed encoder is trained.
Citation Information
Patent Citations
Target detection method and system based on self-supervised contrast learning
CN114549985A
Self-supervised action recognition method and device based on hierarchical multi-view
CN115147676A