Action recognition method based on skeleton and knowledge graph comparative learning
By introducing knowledge graph comparison learning and the fusion of attention of space-time channels in action recognition, the problem of insufficient feature representation in the prior art is solved, and the accuracy of action recognition and the ability to distinguish similar actions are significantly improved.
Patent Information
- Application Number
- CN202510008818.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art ignores channel information in human body movement recognition and fails to effectively combine space-time and channel attention, resulting in insufficient richness and uniqueness of features and difficulty in distinguishing similar behaviors.
A method of action recognition based on skeleton and knowledge graph comparison learning is proposed. By constructing an action recognition model, a skeleton encoder and a skeleton encoder are used to obtain action features, and a standard learning template is obtained by combining a knowledge graph encoder to perform instance-level and feature-level comparison learning, and integrating space-time channel attention.
By using prior knowledge to construct knowledge graph templates and standardize feature learning, the model's ability to discriminate similar actions is improved and the action recognition accuracy is improved.
Smart Images

Figure CN120088849A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of action recognition. Specifically, it relates to an action recognition method based on contrastive learning of skeleton and knowledge graph. Background Art
[0002] Human action recognition is a popular research topic in the field of computer vision and has extensive applications in fields such as human-computer interaction, monitoring systems, virtual reality, and medical diagnosis. Compared with RGB data and optical flow data, skeleton data can provide high-level semantic information, has strong environmental robustness, can simplify data representation, and reduce computational complexity. Therefore, it has been widely used in human action recognition.
[0003] A typical method for using skeletons for action recognition is to construct graph convolutional networks (GCNs). Graph convolutional networks extend convolution from images to graphs and have been successfully applied in many fields. Considering that GCNs can more effectively process the graph structure formed by joints and bones and mine their structural information, in the prior art, GCNs are applied to human action recognition, and the bone data is represented as a directed acyclic graph based on the motion dependence relationship between natural human joints and bones to extract information about joints, bones, and their mutual relationships. There is also a method in the prior art that transplants the shift operation idea in CNNs into GCNs and proposes a shifted graph convolutional network to overcome these two shortcomings. However, these methods rarely consider the optimization of temporal edges. The prior art further proposes a method ST-GCN that combines GCN and CNN, which uses spatio-temporal graph convolution technology to obtain the spatio-temporal distribution and structural information of bone sequence data. In addition, it is proposed to obtain spatial and temporal dimension information through a spatio-temporal graph convolutional network to obtain the spatio-temporal distribution of actions. However, these methods only focus on spatio-temporal information and ignore channel information. And the prior art fails to combine spatio-temporal and channel attention and ignores the differences in the feature representation capabilities of different dimensions. It results in insufficient richness and uniqueness of features and is difficult to distinguish similar behaviors. In addition, these GCN-based models adaptively learn features and obviously lack a standard template to regulate feature learning.
[0004] In recent years, models for learning text and understanding human language have been developed, thus giving rise to natural language processing (NLP). To adapt to different tasks, the Prompt Learning model has been proposed. It adds specific text parameters to the input of large language models (LLMs) according to task definitions to adapt to new downstream tasks and improve the efficiency of knowledge utilization in large language models. Soon, these methods that utilize language knowledge were applied to the visual field.
[0005] Although these studies have provided some insights into the field of action recognition, there are still some deficiencies in the process of feature learning, which are specifically reflected in the following three aspects: 1) These GCN-based models adaptively learn features from the input skeleton data, which obviously lacks a standard template to regulate the feature learning process. 2) The current learning process using prior knowledge only maps the knowledge to the instance-level joint topology. These processes lack effective constraints and guidance and do not utilize the semantic information at the feature level. This may lead to insufficient information learned by the model. 3) These methods fail to combine spatio-temporal attention and channel attention simultaneously and ignore the differences in the feature representation capabilities of different dimensions, resulting in insufficient richness and uniqueness of features. Summary of the Invention
[0006] To overcome at least one deficiency in the prior art, the present application provides an action recognition method based on contrastive learning of skeleton and knowledge graph.
[0007] In a first aspect, an action recognition method based on contrastive learning of skeleton and knowledge graph is provided, including:
[0008] Construct an action recognition model. The action recognition model includes a first action recognition module and a second action recognition module. The first action recognition module includes a skeleton encoder, a first pooling layer, a first fully connected layer, and a first Softmax layer. The second action recognition module includes a spatio-temporal channel skeleton encoder, a second pooling layer, a second fully connected layer, and a second Softmax layer. After the skeleton information passes through the skeleton encoder, the first pooling layer, and the first fully connected layer, a first skeleton feature is obtained. After the first skeleton feature passes through the first Softmax layer, a first action recognition result is output. After the skeleton information passes through the spatio-temporal channel skeleton encoder, the second pooling layer, and the second fully connected layer, a second skeleton feature is obtained. The first skeleton feature and the second skeleton feature are fused to obtain a fused feature, and the fused feature is input into the second Softmax layer to output the final action recognition result.
[0009] Obtain a training data set for obtaining the training data set. The samples in the training data set are skeleton information, and the samples have label text data. Feature extraction is performed on the label text data using a knowledge graph encoder to obtain feature vectors.
[0010] Train the action recognition model based on the training data set to obtain a trained action recognition module. During the training process, a standard cross-entropy loss is constructed based on the first action recognition result, an instance-level contrastive loss is constructed based on the feature vector and the first skeleton feature, and a feature-level contrastive loss is constructed based on the feature vector and the second skeleton feature. The standard cross-entropy loss, the instance-level contrastive loss, and the feature-level contrastive loss are accumulated to obtain the total training loss.
[0011] Input the skeleton information to be recognized into the trained action recognition module to obtain the final action recognition result; the final action recognition result includes the probabilities of each action category; select the action category corresponding to the maximum probability as the action category of the skeleton information to be recognized.
[0012] In one embodiment, the skeleton encoder includes: four skeleton encoding units, namely Stage1, Stage2, Stage3, and Stage4. Among them, Stage1 includes 1 TGN, Stage2 includes 4 TGNs, Stage3 includes 3 TGNs, and Stage4 includes 2 TGNs.
[0013] In one embodiment, the spatio-temporal channel skeleton encoder includes: four spatio-temporal channel skeleton encoding units, namely Stage1, Stage2, Stage3, and Stage4. Among them, Stage1 includes 1 TGN and 1 STC-AFFN, Stage2 includes 4 TGNs and 1 STC-AFFN, Stage3 includes 3 TGNs and 1 STC-AFFN, and Stage4 includes 2 TGNs and 1 STC-AFFN.
[0014] In one embodiment, STC-AFFN includes: 2 STC-NETs and 1 AFFN. After the original input of STC-AFFN passes through one STC-NET, the output of STC-NET is input into AFFN, and the output of STC-NET is multiplied by the original input to obtain the multiplication result;
[0015] After the original input of STC-AFFN passes through another STC-NET, the output of STC-NET is input into AFFN, and the output of STC-NET is added to the original input to obtain the addition result;
[0016] The multiplication result and the addition result are added to obtain the output of STC-AFFN.
[0017] In one embodiment, STC-NET includes STA and CA. After the input of STC-NET passes through STA, the first output is obtained;
[0018] After the input of STC-NET passes through CA, the second output is obtained;
[0019] After weighting the first output and the second output respectively and then adding them, the output of STC-NET is obtained.
[0020] In one embodiment, the standard cross-entropy loss is:
[0021]
[0022] Among them, is the standard cross-entropy loss, is the predicted value of the action category corresponding to sample i, y i is the true value of the action category corresponding to sample i.
[0023] In one embodiment, the instance-level contrastive loss is:
[0024]
[0025] where is the instance-level contrastive loss, i + is the positive sample set in the instance library is an element in, and the instance library includes the first skeleton feature of the sample, is the feature vector, sim represents the cosine distance function, i - is the negative sample set in the instance library is an element in, and τ is the temperature hyperparameter.
[0026] In one embodiment, the feature-level contrastive loss is:
[0027]
[0028] where is the feature-level contrastive loss, f + is the positive sample set in the feature library is an element in, and the feature library includes the second skeleton feature of the sample, is the feature vector, sim represents the cosine distance function, f - is the negative sample set in the feature library is an element in, and τ is the temperature hyperparameter.
[0029] Second, a device for action recognition based on contrastive learning of skeleton and knowledge graph is provided, including:
[0030] A model construction module, configured to construct an action recognition model. The action recognition model includes a first action recognition module and a second action recognition module. The first action recognition module includes a skeleton encoder, a first pooling layer, a first fully connected layer, and a first Softmax layer. The second action recognition module includes a spatio-temporal channel skeleton encoder, a second pooling layer, a second fully connected layer, and a second Softmax layer. After the skeleton information passes through the skeleton encoder, the first pooling layer, and the first fully connected layer, the first skeleton feature is obtained. After the first skeleton feature passes through the first Softmax layer, the first action recognition result is output. After the skeleton information passes through the spatio-temporal channel skeleton encoder, the second pooling layer, and the second fully connected layer, the second skeleton feature is obtained. The first skeleton feature and the second skeleton feature are fused to obtain the fused feature, and the fused feature is input into the second Softmax layer to output the final action recognition result;
[0031] A feature vector acquisition module, configured to obtain a training data set, where the samples in the training data set are skeleton information and the samples have label text data; perform feature extraction on the label text data using a knowledge graph encoder to obtain feature vectors;
[0032] A model training module, configured to train an action recognition model based on the training data set to obtain a trained action recognition module; during the training process, construct a standard cross-entropy loss based on the first action recognition result, construct an instance-level contrast loss based on the feature vector and the first skeleton feature, and construct a feature-level contrast loss based on the feature vector and the second skeleton feature; accumulate the standard cross-entropy loss, the instance-level contrast loss, and the feature-level contrast loss to obtain a total training loss;
[0033] A prediction module, configured to input the skeleton information to be recognized into the trained action recognition module to obtain a final action recognition result; the final action recognition result includes the probabilities of each action category; select the action category corresponding to the maximum probability as the action category of the skeleton information to be recognized.
[0034] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned action recognition method based on contrastive learning between skeleton and knowledge graph.
[0035] Compared with the prior art, the present application has the following beneficial effects: In the action recognition method based on contrastive learning between skeleton and knowledge graph of the present application, prior knowledge is obtained through a large language model, which is further transformed into feature vectors for contrastive learning by a knowledge graph encoder; at the same time, an instance library and a feature library are constructed for double-branch contrastive learning; then a spatio-temporal channel adaptive feature fusion module is inserted into the internal of the skeleton encoder to enrich the feature expression in the feature library; the feature information of the action is obtained through the skeleton encoder and the spatio-temporal channel skeleton encoder, and a standard learning template is obtained through the knowledge graph encoder; finally, positive examples are pulled closer and negative examples are pushed farther in the instance library and the feature library to standardize feature learning. The present application uses prior knowledge to construct a knowledge graph as a template to standardize the learning of features, improves the discriminative ability of the model for similar actions, and further improves the action recognition accuracy. Description of the Drawings
[0036] The present application can be better understood by referring to the description given below in conjunction with the accompanying drawings. The drawings, together with the following detailed description, are included in this specification and form a part of this specification. In the drawings:
[0037] Figure 1 Shows the schematic diagram of the action recognition method based on contrastive learning between skeleton and knowledge graph;
[0038] Figure 2 Shows a schematic structural diagram of the skeleton encoder;
[0039] Figure 3 Shows a schematic structural diagram of the spatio-temporal channel skeleton encoder;
[0040] Figure 4 Shows a schematic structural diagram of STC-AFFN;
[0041] Figure 5 Shows a schematic structural diagram of STC-NET;
[0042] Figure 6 Shows a comparison chart of the efficiency of the combination of the STC-AFFN module and different GCN-based backbones;
[0043] Figure 7 Shows a comparison chart of the action recognition rates between the KCL-GCN model and the baseline under the X-Sub setting of the NTU-RGB+D dataset. Detailed implementation manners
[0044] In the following, exemplary embodiments of the present application will be described in conjunction with the accompanying drawings. For clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many embodiment-specific decisions may be made during the development of any such actual embodiment to achieve the specific goals of the developer, and these decisions may vary with different embodiments.
[0045] Here, it should also be noted that in order to avoid obscuring the present application with unnecessary details, only the device structures closely related to the solution of the present application are shown in the drawings, and other details less related to the present application are omitted.
[0046] It should be understood that the present application is not limited to the described embodiments only due to the following description with reference to the drawings. In this document, where feasible, embodiments can be combined with each other, features can be replaced or borrowed between different embodiments, and one or more features can be omitted in one embodiment.
[0047] An embodiment of the present application provides an action recognition method based on contrastive learning of skeleton and knowledge graph, Figure 1 Shows the principle diagram of the action recognition method based on contrastive learning of skeleton and knowledge graph, see Figure 1 , and the method mainly includes:
[0048] Step S1, construct an action recognition model. The action recognition model includes a first action recognition module and a second action recognition module. The first action recognition module includes a skeleton encoder, a first pooling layer, a first fully connected layer, and a first Softmax layer. The second action recognition module includes a spatio-temporal channel skeleton encoder, a second pooling layer, a second fully connected layer, and a second Softmax layer. After the skeleton information passes through the skeleton encoder, the first pooling layer, and the first fully connected layer, a first skeleton feature is obtained. After the first skeleton feature passes through the first Softmax layer, a first action recognition result is output. After the skeleton information passes through the spatio-temporal channel skeleton encoder, the second pooling layer, and the second fully connected layer, a second skeleton feature is obtained. The first skeleton feature and the second skeleton feature are fused to obtain a fused feature, and the fused feature is input into the second Softmax layer to output the final action recognition result. Here, the first skeleton feature and the second skeleton feature can be fused by weighting and then accumulating.
[0049] Here, a pooling layer is used to obtain a one-dimensional high-level feature vector. Then, a fully connected layer activated by Softmax maps the features to a probability distribution of C candidate classes, where C represents the number of classes.
[0050] Figure 2 The structural schematic diagram of the skeleton encoder is shown. Refer to Figure 2 , the skeleton encoder SE includes: four skeleton encoding units, namely Stage1, Stage2, Stage3, and Stage4. Among them, Stage1 includes 1 TGN (Temporal Graph Network), Stage2 includes 4 TGNs, Stage3 includes 3 TGNs, and Stage4 includes 2 TGNs. TGN is an existing network structure. In order to learn more discriminative representations, the backbone network is divided into four stages according to the traditional method: the first, fifth, eighth, and last layers of the TGN. Skip operations are used for the fifth and eighth layers, as Figure 2 shown.
[0051] Figure 3 The structural schematic diagram of the spatio-temporal channel skeleton encoder is shown. Refer to Figure 3 , the spatio-temporal channel skeleton encoder STC-SE includes: four spatio-temporal channel skeleton encoding units, namely Stage1, Stage2, Stage3, and Stage4. Among them, Stage1 includes 1 TGN and 1 STC-AFFN, Stage2 includes 4 TGNs and 1 STC-AFFN, Stage3 includes 3 TGNs and 1 STC-AFFN, and Stage4 includes 2 TGNs and 1 STC-AFFN.
[0052] Here, on the basis of constructing the SE, a plug-and-play spatio-temporal channel adaptive feature fusion module STC-AFFN is additionally designed and inserted into the interior of the SE module to enrich the feature expression of the feature library. Specifically, the STC-AFFN module is composed of the STC-NET and the AFFN module. Through the STC-SE, more discriminative features can be obtained for subsequent contrast learning with the standard template.
[0053] To simultaneously consider spatio-temporal and channel attention and also consider the differences in feature representation capabilities between dimensions, the STC-AFFN module composed of the STC-NET and the adaptive weighted fusion network AFFN is designed. Figure 4 The structural schematic diagram of the STC-AFFN is shown, see Figure 4 , the STC-AFFN includes: 2 STC-NETs and 1 AFFN. After the original input of the STC-AFFN passes through one STC-NET, the output of the STC-NET is input into the AFFN, and the output of the STC-NET is multiplied by the original input to obtain the multiplication result;
[0054] After the original input of the STC-AFFN passes through another STC-NET, the output of the STC-NET is input into the AFFN, and the output of the STC-NET is added to the original input to obtain the addition result;
[0055] The multiplication result and the addition result are added to obtain the output of the STC-AFFN.
[0056] The adaptive weighted fusion network AFFN uses the weighted fusion method to achieve the weighted sum of attention features by using the result output by the STC-NET module together with the original features. The original features are directly fused with the attention features of the AFFN as an additional constraint, and the attention features of the AFFN are weighted and fused with the original features as detail features; further, the weighted sum introduces the output A of the STC-NET STC as the weight, and then performs the weighted sum on the original feature f to obtain f p , which can be expressed as:
[0057] f p = f * A STC
[0058] Introduce the output of the STC-NET as the feature f′, and directly add it to the original feature f to obtain f a , which can be expressed as:
[0059] f a = f'+ f
[0060] Figure 5 The structural schematic diagram of the STC-NET is shown, see Figure 5, STC-NET includes a spatio-temporal feature attention module STA and a channel feature attention module CA. After the input of STC-NET passes through STA, a first output is obtained;
[0061] After the input of STC-NET passes through CA, a second output is obtained;
[0062] After weighting the first output and the second output respectively and then adding them together, the output of STC-NET is obtained.
[0063] A spatio-temporal feature attention module (STA) and a channel feature attention module (CA) are designed. The spatio-temporal attention is combined with the channel joint attention, and STC-NET is designed. A weighting strategy is introduced to effectively adjust the contribution rates of the spatio-temporal dimension and the channel dimension to action recognition. Here, the spatio-temporal feature attention module (STA) focuses on the temporal relationship between different frames and captures the motion patterns and trajectories in the action; the channel feature attention module (CA) weights the respective channel features of the input tensor to capture the correlation between different features and improve the model's perception ability of key channels.
[0064] STC-NET takes the spatio-temporal feature attention and channel feature attention mined by the spatio-temporal feature attention module and the channel feature attention module as inputs, performs weighted fusion, adjusts the contribution rates of the spatio-temporal dimension and the channel dimension to action recognition, and introduces hyperparameters λ1 and λ2 to perform weighted summation on the two attention features. The formula is as follows:
[0065] A STC =λ 1 *A ST +λ 2 *A C *
[0066] Among them, A STC is the output of STC-NET, A ST is the first output, is the second output.
[0067] Step S2, obtain a training data set. The samples in the training data set are skeleton information, and the samples have label text data; the knowledge graph encoder is used to extract features from the label text data to obtain feature vectors.
[0068] Here, the data set is preprocessed first, which mainly includes: loading the original skeleton data in the data set, and each sample contains 3D joint positions at a series of time steps. Next, analyze the skeleton data of each sample and extract the skeleton information of all performers. Then, check whether there are some joint data missing or cannot be correctly recognized in these skeleton information, and remove the skeleton frames of these incomplete data.
[0069] Secondly, to unify the coordinate system and remove perspective or scale differences, the coordinates of each skeleton sample are transformed to the center point of the first frame. The first frame is used as a reference point through translation transformation to ensure that all subsequent frames are based on this. The input to the final model is a series of skeleton information, with dimensions X ∈ R H×W×3 , which represents H frames of W joints in 3D space.
[0070] Given the operation label l, first use the engineered prompt function to create appropriate prompts. The content of is to provide a detailed description of the l action, including fine-grained operation information for differentiating similar operations. The result l p is provided to the LLM, such as GPT-3, and then the operation descriptions generated by GPT-3 are manually checked to ensure their accuracy. Finally, the prior knowledge Prior knowledge The prior knowledge is input into the Knowledge Graph Encoder KGE. The prior knowledge is tokenized in sentence form to facilitate subsequent indexing of the prior knowledge in the vocabulary by the model. Then, these data are indexed in the embedding layer according to the previous tokens and converted into continuous vectors. Finally, by processing the Transformer module, a feature vector representing the prior knowledge is output which is further used as a standard template for contrastive learning.
[0071] Use a pre-trained language model BERT as the Knowledge Graph Encoder KGE. Therefore, KGE is a language model based on the Transformer architecture that can generate context-related word embeddings. Commonly used language models include CLIP and BERT. Based on previous experimental results, a pre-trained text transformer model from BERT is used as KGE. This context-based model can capture global information, thus better obtaining the standard template.
[0072] Step S3, train the action recognition model based on the training dataset to obtain the trained action recognition module; during the training process, construct the standard cross-entropy loss based on the first action recognition result, construct the instance-level contrast loss based on the feature vector and the first skeleton feature, and construct the feature-level contrast loss based on the feature vector and the second skeleton feature; the standard cross-entropy loss, the instance-level contrast loss, and the feature-level contrast loss are accumulated to obtain the total training loss.
[0073] Here, an instance library is constructed based on the first skeleton features of the samples, and a feature library is constructed based on the second skeleton features of the samples for dual-branch contrastive learning. The instance library provides complete time operation information to ensure the time consistency of operations, thus helping the model locate operations and understand the differences between various operations. In contrast, the feature library provides action features to help the model capture action details, identify and compare the similarities between different actions to distinguish similar actions. In the subsequent contrastive learning, the feature vector (knowledge graph) serves as a template to standardize the learning process of these instance information and feature information in the instance library and the feature library, and finally strengthens the constraints of the model by pulling positive examples closer and pushing negative examples further apart.
[0074] During the model training process, the contrastive learning method is adopted. The feature vector serves as a template to standardize the learning process of these instance information and feature information in the instance library and the feature library, and finally strengthens the constraints of the model by pulling positive examples closer and pushing negative examples further apart.
[0075] Positive and negative samples are selected from the instance library and the feature library respectively, and the output of the knowledge graph encoder is used as the standard template for comparison. Here, based on the similarity metric (cosine similarity) in the feature space, the feature vector most similar to the template feature is selected as the positive sample. Vice versa, the feature vector least similar to the template feature is selected as the negative sample. In the instance library, represents the positive sample set, represents the negative sample set. The cosine distance function is used to measure the distance between the positive sample and the standard template and make them gradually approach, while is used to measure the distance between the negative sample and the standard template and make them gradually move away. Then the same concept is applied in the feature library.
[0076] Specifically, the standard cross-entropy loss is:
[0077]
[0078] where, is the standard cross-entropy loss, is the predicted value of the action category corresponding to sample i, y i is the true value of the action category corresponding to sample i.
[0079] The instance-level contrastive loss is:
[0080]
[0081] where, is the instance-level contrastive loss, i + is the positive sample set in the instance library The elements in it, the instance library includes the first skeleton feature of the sample, is the feature vector, sim represents the cosine distance function, i - is the negative sample set in the instance library The elements in it, τ is the temperature hyperparameter.
[0082] The feature-level contrast loss is:
[0083]
[0084] Among them, is the feature-level contrast loss, f + is the positive sample set in the feature library The elements in it, the feature library includes the second skeleton feature of the sample, is the feature vector, sim represents the cosine distance function, f - is the negative sample set in the feature library The elements in it, τ is the temperature hyperparameter.
[0085] Step S4, input the skeleton information to be recognized into the trained action recognition module to obtain the final action recognition result. The final action recognition result includes the probabilities of each action category; select the action category corresponding to the maximum probability as the action category of the skeleton information to be recognized.
[0086] In this embodiment, prior knowledge is obtained through the large language model, which is further transformed into a feature vector for contrast learning by the knowledge graph encoder; at the same time, an instance library and a feature library are constructed for dual-branch contrast learning; then the spatio-temporal channel adaptive feature fusion module is inserted into the internal of the skeleton encoder to enrich the feature expression in the feature library; the feature information of the action is obtained through the skeleton encoder and the spatio-temporal channel skeleton encoder, and the standard learning template is obtained through the knowledge graph encoder; finally, positive examples are pulled closer and negative examples are pushed farther in the instance library and the feature library to standardize feature learning. This embodiment uses prior knowledge to construct a knowledge graph as a template to standardize the learning of features, improves the discriminative ability of the model for similar actions, and further improves the action recognition accuracy.
[0087] Adopting the same inventive concept as the action recognition method based on skeleton and knowledge graph contrast learning, this embodiment also provides a corresponding action recognition device based on skeleton and knowledge graph contrast learning, including:
[0088] A model construction module for constructing an action recognition model, which includes a first action recognition module and a second action recognition module. The first action recognition module includes a skeleton encoder, a first pooling layer, a first fully-connected layer, and a first Softmax layer. The second action recognition module includes a spatio-temporal channel skeleton encoder, a second pooling layer, a second fully-connected layer, and a second Softmax layer. After the skeleton information passes through the skeleton encoder, the first pooling layer, and the first fully-connected layer, a first skeleton feature is obtained. After the first skeleton feature passes through the first Softmax layer, a first action recognition result is output. After the skeleton information passes through the spatio-temporal channel skeleton encoder, the second pooling layer, and the second fully-connected layer, a second skeleton feature is obtained. The first skeleton feature and the second skeleton feature are fused to obtain a fused feature, and the fused feature is input into the second Softmax layer to output the final action recognition result.
[0089] A feature vector acquisition module for obtaining a training data set. The samples in the training data set are skeleton information, and the samples have label text data. The knowledge graph encoder is used to extract features from the label text data to obtain feature vectors.
[0090] A model training module for training the action recognition model based on the training data set to obtain a trained action recognition module. During the training process, a standard cross-entropy loss is constructed based on the first action recognition result, an instance-level contrast loss is constructed based on the feature vector and the first skeleton feature, and a feature-level contrast loss is constructed based on the feature vector and the second skeleton feature. The standard cross-entropy loss, the instance-level contrast loss, and the feature-level contrast loss are accumulated to obtain the total training loss.
[0091] A prediction module for inputting the skeleton information to be recognized into the trained action recognition module to obtain the final action recognition result. The final action recognition result includes the probabilities of each action category. The action category corresponding to the maximum probability is selected as the action category of the skeleton information to be recognized.
[0092] The action recognition device based on skeleton and knowledge graph contrast learning in this embodiment has the same inventive concept as the above-mentioned action recognition method based on skeleton and knowledge graph contrast learning. Therefore, the specific implementation of this device can be seen in the embodiment part of the above-mentioned action recognition method based on skeleton and knowledge graph contrast learning, and its technical effects correspond to those of the above method, which will not be elaborated here.
[0093] This application embodiment provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it realizes the above-mentioned action recognition method based on skeleton and knowledge graph contrast learning.
[0094] To further verify the effectiveness of the method of this application, the following experimental analysis was carried out.
[0095] I. Dataset Selection
[0096] The NTU - RGB+D dataset is one of the most widely used datasets for skeleton - based action recognition research. It contains a total of 56,880 action videos, including actions performed by 40 volunteers aged between 10 and 35 years old, covering 60 different action categories. The dataset provides four different modalities of data, including RGB videos, depth map sequences, 3D skeleton data, and infrared videos, but this application focuses on using skeleton data in the research. These data were captured by a Microsoft Kinect V2 device at a speed of 30 frames per second, and each action was captured using three different camera angles, namely - 45°, 0°, and 45°. When using the NTU - RGB+D dataset, this application followed the convention and used two common benchmark settings: 1), cross - subject (X - sub): In this setting, the volunteers in the training set and the validation set are different. The training data comes from 20 different volunteers, containing 40,320 videos, while the test data comes from another 20 volunteers, containing 16,560 videos. 2), cross - view (X - view): In this setting, the training set and the validation set use data from different camera perspectives. The training data comes from camera views Figure 2 and 3 , containing 37,920 videos (0° and 45° perspectives), while the test data comes from camera views Figure 1 , containing 18,960 videos (- 45° perspective).
[0097] The NW - UCLA dataset was captured by three Kinect cameras simultaneously from multiple angles. It contains 1,494 video clips, covering 10 categories. Each action was performed by 10 actors. This application adopted the same evaluation protocol: This application used the samples from the first two cameras as training data and the samples from the other camera as test data.
[0098] The NTU RGB+D 120 dataset is currently the largest skeleton-based action recognition dataset. It extends NTU RGB+D by introducing 60 new action classes, adding 57,367 samples. It has collected a total of 113,945 skeleton sequences performed by 106 participants in 120 different categories. It also increases the number of camera setups to 32 by using different locations and backgrounds. The authors of this dataset recommend two benchmarks: 1) cross-subject (X-sub) benchmark: The training data comes from 53 subjects, and the test data comes from the other 53 tested subjects. 2) cross-setup (X-set) benchmark: The training data comes from samples with even setup IDs, and the test data comes from samples with odd setup IDs.
[0099] II. Experimental Details
[0100] In the experiments of this application, all experiments were conducted under the PyTorch deep learning framework. The Stochastic Gradient Descent (SGD) optimizer was used to train the model, with a momentum of 0.9 and a weight decay of 0.0004. When training the model on three datasets, this application applied a warm-up strategy in the first 5 epochs for stable training. This application set the initial learning rate to 0.1 and reduced it by a factor of 0.1 at epochs 35 and 55. This application trained all models for 80 epochs and selected the best performance. For NTU RGB+D and NTU RGB+D 120, the batch size was 64, and the size of each sample was adjusted to 64 frames. For NW-UCLA, this application set the batch size to 16. All experiments were conducted on an RTX 3080TI GPU.
[0101] III. Analysis of Experimental Results
[0102] 1. Verification of the effectiveness of the prompting function
[0103] To verify the effectiveness of the prompting function, taking "reading" as an example, the prior knowledge generated by different prompting types is shown in Table 1. Obviously, the behavior descriptions obtained through the prompting function of this application contain richer behavior details, which is beneficial for distinguishing different behaviors.
[0104] Table 1: Prior knowledge corresponding to different types of prompting functions
[0105]
[0106] 2. Performance comparison of different algorithms on three datasets
[0107] The modules of this application have been evaluated on three widely used datasets with state-of-the-art methods: NTU-RGB+D 120, NTU-RGB+D, and NW-UCLA. To make a fair comparison, this application adopts a multi-stream fusion strategy. As shown in the following table, the accuracy of action recognition is significantly improved by using the KCL-GCN model of this application. In addition, this application also provides a solution for similar actions, which is of great significance in skeleton-based action recognition.
[0108] Table 2: Performance comparison of different algorithms on three datasets
[0109]
[0110] 3. Effectiveness analysis of the Spatiotemporal Channel Skeleton Encoder (STC-SE)
[0111] The effectiveness of the Spatiotemporal Channel Skeleton Encoder (STC-SE) is crucial for the quality of subsequent dual-branch contrastive learning. Therefore, this application first focuses on studying its effectiveness to prove that the model can effectively extract semantic information, thus enriching the feature representation in the feature library. SE is a module that combines the STC-AFFN module and SE in a plug-and-play manner. This part studies the design of STC-AFFN from three aspects: (1) the effectiveness of each STC-AFFN module, (2) the setting of hyperparameters, (3) the impact of the number of STC-AFFN modules, and (4) the combination with other skeletons. Then STC-AFFN is combined with SE to form STC-SE. This application uses the NTU-RGB+D and NW-UCLA datasets as inputs for ablation experiments.
[0112] (1) The effectiveness of each STC-AFFN module. To evaluate the contribution of each sub-module, this application compares the effectiveness of each sub-module separately. This application divides the module into three sub-modules and introduces five variants. Table 3 shows that adding each module can improve the accuracy of the baseline. The main contribution comes from STC-NET, which adopts a weighted strategy to capture more unique attention features. Combining all modules can improve the result AFFN to enhance the feature representation. It is worth noting that each module can improve the performance while limiting the increase in computational cost.
[0113] Table 3: Recognition accuracy when inserting different modules
[0114]
[0115] (2) The setting of hyperparameters. As shown in Table 4, this application analyzes the hyperparameter configuration of the module. This application explores λ iDifferent combinations. Spatiotemporal information is the key to distinguishing behavior categories, while channel information can provide more details or background information. A C Adjust the channel information of the skeleton, which plays a role in supplementing spatiotemporal features in the whole framework. Finally, configure λ 1 = 0.8, λ 2 = 0.2.
[0116] Table 4: Recognition accuracy corresponding to different hyperparameter configurations
[0117]
[0118]
[0119] (3) Influence of the number of STC-AFFN modules. The SE structure can be divided into 10 stages, and STC-AFFN can be inserted into any of these stages. This application applies striding operations at the 5th and 8th layers. Considering the importance of the final features, this application inserts the STC-AFFN module into the 5th, 8th, and 10th layers. The results of this application are given in Table 5. The experimental results show that the effect is improved after STC-AFFN is added to different layers. However, extracting features from the early stage may mislead subsequent recognition. This means that high-level features dominate in the later stage, while low-level features serve as an auxiliary. To obtain higher accuracy with the least increase in parameters, choose to insert the STC-AFFN module into the 1st, 5th, 8th, and 10th layers. It should be noted that no matter how the STC-AFFN module is inserted, there are significant improvements.
[0120] Table 5: Recognition accuracy corresponding to different numbers of STC-AFFN modules
[0121] Phase Top-1(%) Baseline 92.9 +5,7,8,10 93.5 +5,8,9,10 94.0 +1,5,8,9,10 94.2 +1,3,5,7,8,10 93.8 +1,5,8,10 95.3
[0122] (4) Combination with other skeletons. The STC-SE module proposed in this application is plug-and-play and can be combined with most GCN-based backbones. To test its generality, it is applied to 3 widely used GCN-based backbones (CTR-GCN, ST-GCN, and 2s-AGCN) and evaluated on the joint stream of the NW-UCLA dataset.
[0123] Figure 6 Shows the efficiency comparison diagram of the STC-AFFN module combined with different GCN-based backbones, and the specific data is shown in the following table:
[0124] Table 6: Performance of different GCN backbone methods on the NW-UCLA dataset under the joint input method
[0125]
[0126] 4. Ablation experiments of dual-branch contrastive learning. The memory bank is designed to store complementary instance-level and feature-level information for dual-branch contrastive learning. Therefore, this application combines the memory bank with dual-branch contrastive learning for experiments. The results are shown in Table 7. After introducing the instance bank into the model and performing single-branch contrastive learning, the accuracy rates are increased by 0.2% and 0.4% respectively. Different from previous methods, this application combines the designed instance bank and feature bank simultaneously. Through dual-branch contrastive learning, the distance between the feature and the standard template is significantly reduced from two complementary aspects, resulting in a significant improvement in recognition performance, with an increase of 0.8% in NTU RGB+D under the X-Sub setting and an increase of 2.4% in the NW-UCLA joint input mode.
[0127] Table 7: Recognition accuracy rates (%) corresponding to different memory banks
[0128]
[0129] 5. Ablation experiments of the knowledge graph encoder. This application carefully designs a text prompt function and inputs the result into the knowledge graph encoder to construct the knowledge graph, which is a key part of the knowledge graph. As shown in Table 8, this application compares the results of using only labels as prompts and using the text prompt function. Using only labels as prompts contains less information, but due to dual-branch contrastive learning, the performance improvement is also significant, with increases of 0.3% and 1.1% respectively. Using the designed allows the knowledge graph encoder to obtain fine-grained information to generate detailed features of the operations. Under the joint input mode, under the X-Sub setting, the improvement rate of NTU RGB+D is 0.8%, and under NW-UCLA, the improvement rate is 2.4%.
[0130] Table 8: Recognition accuracies (%) of different text prompt types on the dataset
[0131]
[0132] 6. Analysis of the recognition performance of similar actions in this application. To evaluate whether KCL-GCN can effectively classify similar actions, this application analyzes its confusion matrix. This application selects some representative and easily confused similar operations, such as 29->11, where the correct action is A29 (playing with mobile phone / tablet), but is misrecognized as A11 (reading). As shown in Table 9, in the KCL-GCN model, the error rates of these operations have been significantly reduced. By using dual-branch contrastive learning to reduce the distance between positive examples and the template, while increasing the distance from negative examples, thereby regularizing feature learning, KCL-GCN solves the challenge of distinguishing similar behaviors.
[0133] Table 9: Error rate (%) of NTU-RGB+D action recognition under X-Sub setting in combined input mode
[0134]
[0135] This application further compares the recognition performance of KCL-GCN and the baseline in the combined mode of NTU-RGB+D X-Sub setting, Figure 7 shows the comparison chart of action recognition rate between the KCL-GCN model and the baseline under the X-Sub setting of the NTU-RGB+D dataset, and visualizes the parts with a difference greater than 3% in Figure 7 . Since this method effectively reduces the error rate of recognizing similar actions, the recognition accuracy has been significantly improved.
[0136] In summary, this application proposes an action recognition method based on contrastive learning of skeleton and knowledge graph, encodes the knowledge of the large language model into action-specific templates to regulate the feature learning process. The idea of this model is to pull the positive examples closer to the template and push the negative examples away. To maximize the constraint ability of the template, this application also constructs a two-branch contrastive learning framework and designs a spatio-temporal channel skeleton encoder to further improve the feature representation. It is worth noting that this application provides a novel solution for action recognition and achieves state-of-the-art performance on three widely used benchmarks. This application is applied to the fields of computer vision and artificial intelligence, especially suitable for action recognition of human skeleton data, and has broad application potential, such as autonomous driving, intelligent monitoring, human action analysis, health monitoring, etc.
[0137] The above are only various embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.
Claims
1. An action recognition method based on comparative learning of skeleton and knowledge graph, characterized in that: include: Constructing an action recognition model, the action recognition model includes a first action recognition module and a second action recognition module, the first action recognition module includes a skeleton encoder, a first pooling layer, a first fully connected layer, and a first Softmax layer, and the second action recognition module includes a spatiotemporal channel skeleton encoder, a second pooling layer, a second fully connected layer, and a second Softmax layer; after the skeleton information passes through the skeleton encoder, the first pooling layer, and the first fully connected layer, a first skeleton feature is obtained, and after the first skeleton feature passes through the first Softmax layer, a first action recognition result is output; after the skeleton information passes through the spatiotemporal channel skeleton encoder, the second pooling layer, and the second fully connected layer, a second skeleton feature is obtained; the first skeleton feature and the second skeleton feature are subjected to feature fusion to obtain fused features, and the fused features are input into the second Softmax layer to output a final action recognition result; Obtain a training data set, wherein the samples in the training data set are skeleton information, and the samples have label text data; perform feature extraction on the label text data using a knowledge graph encoder to obtain a feature vector; Training the action recognition model based on the training data set to obtain a trained action recognition module; During the training process, a standard cross entropy loss is constructed based on the first action recognition result, an instance-level contrast loss is constructed based on the feature vector and the first skeleton feature, and a feature-level contrast loss is constructed based on the feature vector and the second skeleton feature; the standard cross entropy loss, the instance-level contrast loss, and the feature-level contrast loss are accumulated to obtain a total training loss; Inputting the skeleton information to be identified into the trained action recognition module to obtain a final action recognition result; the final action recognition result includes the probability of each action category; The action category corresponding to the maximum probability is selected as the action category of the skeleton information to be identified.
2. The method according to claim 1, characterized in that The skeleton encoder includes: four skeleton encoding units, namely Stage 1, Stage 2, Stage 3, and Stage 4, wherein Stage 1 includes 1 TGN, Stage 2 includes 4 TGNs, Stage 3 includes 3 TGNs, and Stage 4 includes 2 TGNs.
3. The method according to claim 1, characterized in that The spatiotemporal channel skeleton encoder includes: four spatiotemporal channel skeleton encoding units, namely Stage 1, Stage 2, Stage 3, and Stage 4, wherein Stage 1 includes 1 TGN and 1 STC-AFFN, Stage 2 includes 4 TGNs and 1 STC-AFFN, Stage 3 includes 3 TGNs and 1 STC-AFFN, and Stage 4 includes 2 TGNs and 1 STC-AFFN.
4. The method according to claim 3, characterized in that The STC-AFFN includes: 2 STC-NETs and 1 AFFN. After the original input of the STC-AFFN passes through an STC-NET, the output of the STC-NET is input into the AFFN. The output of the STC-NET is multiplied by the original input to obtain a multiplication result. After the original input of the STC-AFFN passes through another STC-NET, the output of the STC-NET is input into the AFFN, and the output of the STC-NET is added to the original input to obtain an addition result; The multiplication result and the addition result are added to obtain the output of the STC-AFFN.
5. The method according to claim 3, characterized in that The STC-NET includes a STA and a CA, and the input of the STC-NET passes through the STA to obtain a first output; The input of the STC-NET is passed through the CA to obtain a second output; The first output and the second output are weighted respectively and then added to obtain the output of the STC-NET.
6. The method according to claim 1, characterized in that in, The standard cross entropy loss is: in, is the standard cross entropy loss, is the predicted value of the action category corresponding to sample i, y i is the true value of the action category corresponding to sample i.
7. The method according to claim 1, characterized in that in, The instance-level contrast loss is: in, is the instance-level contrast loss, i + is the positive sample set in the instance library The elements in the instance library include the first skeleton features of the samples. is the feature vector, sim represents the cosine distance function, i - is the negative sample set in the instance library The elements in , τ is the temperature hyperparameter.
8. The method according to claim 1, characterized in that The feature-level contrast loss is: in, is the feature-level contrast loss, f + is the positive sample set in the feature library The elements in the feature library include the second skeleton features of the samples. is the feature vector, sim represents the cosine distance function, f - is the negative sample set in the feature library The elements in , τ is the temperature hyperparameter.
9. An action recognition device based on comparative learning of skeleton and knowledge graph, characterized in that: include: A model construction module is used to construct an action recognition model, wherein the action recognition model includes a first action recognition module and a second action recognition module, wherein the first action recognition module includes a skeleton encoder, a first pooling layer, a first fully connected layer, and a first Softmax layer, and the second action recognition module includes a spatiotemporal channel skeleton encoder, a second pooling layer, a second fully connected layer, and a second Softmax layer; after the skeleton information passes through the skeleton encoder, the first pooling layer, and the first fully connected layer, a first skeleton feature is obtained, and after the first skeleton feature passes through the first Softmax layer, a first action recognition result is output; after the skeleton information passes through the spatiotemporal channel skeleton encoder, the second pooling layer, and the second fully connected layer, a second skeleton feature is obtained; the first skeleton feature and the second skeleton feature are subjected to feature fusion to obtain fused features, and after the fused features are input into the second Softmax layer, a final action recognition result is output; A feature vector acquisition module is used to acquire a training data set, wherein the samples in the training data set are skeleton information, and the samples have label text data; a knowledge graph encoder is used to perform feature extraction on the label text data to obtain a feature vector; A model training module, used to train the action recognition model based on the training data set to obtain a trained action recognition module; During the training process, a standard cross entropy loss is constructed based on the first action recognition result, an instance-level contrast loss is constructed based on the feature vector and the first skeleton feature, and a feature-level contrast loss is constructed based on the feature vector and the second skeleton feature; the standard cross entropy loss, the instance-level contrast loss, and the feature-level contrast loss are accumulated to obtain a total training loss; A prediction module, used to input the skeleton information to be recognized into the trained action recognition module to obtain a final action recognition result; the final action recognition result includes the probability of each action category; The action category corresponding to the maximum probability is selected as the action category of the skeleton information to be identified.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the action recognition method based on skeleton and knowledge graph comparative learning as described in any one of claims 1-8.