Open vocabulary action recognition method and system based on heterogeneous skeleton

By constructing a unified open vocabulary skeleton dataset and a multi-granularity motion-text alignment strategy based on the Transformer architecture, we solved the recognition problem of heterogeneous skeleton data and improved the accuracy and generalization ability of action recognition, especially the recognition performance of rare actions in open vocabulary scenarios.

CN120673478APending Publication Date: 2025-09-19SOUTHEAST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510822581.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing action recognition methods have insufficient generalization ability when processing heterogeneous skeletal data, especially in open vocabulary scenarios, where it is difficult to effectively identify new action categories. Existing methods usually ignore the heterogeneity of skeletal data, resulting in poor recognition results in practical applications.

Method used

By integrating multiple heterogeneous skeleton datasets, a unified open vocabulary skeleton dataset is constructed, and a skeleton motion encoder model based on the Transformer architecture is adopted, combined with a multi-granularity motion-text alignment strategy, cross-modal alignment and unified representation learning are achieved, thereby improving the effectiveness and generalization ability of action recognition.

Benefits of technology

It significantly improves the accuracy of action recognition and the generalization of the model, and can effectively identify complex human actions, especially the recognition performance of rare action categories in long-tail distribution data and open vocabulary scenarios, and enhances the model's representation ability and robustness for skeletal actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673478A_ABST
    Figure CN120673478A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous skeleton-based open vocabulary action recognition method and system. The method comprises the following steps of: constructing a heterogeneous open vocabulary skeleton data set; unifying the skeleton representation in the heterogeneous open vocabulary skeleton data set, and defining a unified space structure by establishing the maximum joint number and member number; a skeleton motion encoder model based on a Transform architecture is constructed, the skeleton motion encoder model comprises feature embedding, space-time encoding and a projection layer used for cross-modal alignment, and global time features, global space features and global visual features output by space-time encoding are subjected to cross-modal alignment through three parallel projection networks of the projection layer. Mapping to a semantic space aligned with text embedding from the pre-trained language model; training loss is constructed based on a multi-granularity motion-text alignment strategy including global instance alignment, stream specific alignment and fine granularity alignment, and a motion encoder model is trained. A wide range of experiments on a popular benchmark with heterogeneous skeleton data prove the effectiveness and generalization ability of the proposed method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning framework for heterogeneous skeleton data, which is intended to be able to process heterogeneous skeleton data and achieve the effect of open vocabulary, and belongs to the technical field of human action recognition. Background Art

[0002] With the development of robotics and embodied intelligence, motion recognition has become increasingly critical. It enhances human-robot interaction, enabling robots to understand human gestures, postures, and intentions. It also improves the adaptability of intelligent agents to diverse scenarios, promoting their intelligence by enabling them to analyze human behavior and make autonomous decisions. Motion recognition also plays a vital role in other fields, including medical rehabilitation, healthcare, smart homes, smart education, and smart sports.

[0003] Skeletons are a commonly used modality in action recognition tasks. Unlike other modalities such as video sequences and depth maps, skeletons focus on key joint locations and are highly abstract. This results in low data complexity and reduces computational burden. Another key advantage is their robustness. Unlike video and depth map data, which can be degraded by lighting changes, occlusions, or background clutter, skeleton data remains reliable. Furthermore, the inherent alignment of skeleton data with human kinematics facilitates a more precise representation of actions. These properties make skeletons an ideal choice for action recognition.

[0004] Although deep learning methods have developed rapidly in skeleton-based human action recognition, these methods mainly focus on proposing different network architectures for single homogeneous data. Based on different types of neural networks, they can be roughly divided into three categories: recurrent neural network (RNN)-based methods, graph convolutional network (GCN)-based methods, and transformer-based methods. These methods ignore the data heterogeneity of skeletons, which comes from the various motion capture technologies, wearable sensors, and pose estimation algorithms used in data collection. These different sources produce skeleton data with different numbers of joints, different skeleton topologies, and different coordinate dimensions. Figure 1 As shown, the Kinect v2 depth sensor captures 25 3D joints, while 2D pose estimation algorithms applied to RGB video typically produce 17 2D joints. Furthermore, motion capture systems following models such as SMPL generate motions with 22 3D joints. Therefore, models trained on data from a specific source cannot be directly applied to skeletons with different structural properties and may exhibit poor generalization.

[0005] Furthermore, most existing action recognition paradigms are designed under the closed-set assumption. Although self-supervised skeletal action recognition and zero-shot skeletal action recognition can be applied to open-set scenarios, the former requires an additional fine-tuning step, while the latter is limited to a limited number of unseen categories. Large-scale skeletal-based action recognition with an open vocabulary beyond the explicitly trained categories is an important problem to be addressed. However, in practical applications, the challenge of data heterogeneity further complicates the goal of integrative skeletal-based action recognition. In this context, open-vocabulary systems must generalize not only to new semantic concepts but also to the diverse and heterogeneous skeletal data associated with them. Summary of the Invention

[0006] Purpose of the invention: In view of the shortcomings of the existing technology, the purpose of the present invention is to provide a method and system for open vocabulary action recognition based on heterogeneous skeletons, to study the challenging problems of action recognition based on heterogeneous skeletons with an open vocabulary, and to improve the effectiveness and generalization ability of action recognition.

[0007] Technical solution: To achieve the above-mentioned purpose, the present invention adopts the following technical solution:

[0008] In a first aspect, the present invention provides an open vocabulary action recognition method based on heterogeneous skeletons, comprising the following steps:

[0009] Integrate multiple heterogeneous skeleton datasets to construct a heterogeneous open vocabulary skeleton dataset;

[0010] Unify the skeleton representation in heterogeneous open vocabulary skeleton datasets and define a unified spatial structure by establishing the maximum number of joints and members;

[0011] A skeletal motion encoder model based on the Transformer architecture is constructed. The model architecture includes feature embedding, spatiotemporal encoding, and a projection layer for cross-modal alignment. The global temporal features, global spatial features, and global visual features output by the spatiotemporal encoding are mapped to a semantic space aligned with the text embeddings from the pre-trained language model through three parallel projection networks in the projection layer.

[0012] The motion encoder model is trained based on a multi-granularity motion-text alignment strategy that constructs a training loss. The multi-granularity motion-text alignment strategy includes: global instance alignment, which aims to align the final global visual representation with the corresponding action label embedding at the instance level; stream-specific alignment, which aims to align the projected global temporal features and spatial features with the action label embedding respectively, and to align the global temporal features with the spatial features; fine-grained alignment, which divides the time series features into multiple non-overlapping segments and the spatial sequence features into multiple semantic parts. For each time period / spatial part, representative features are obtained, with the goal of aligning local features with the action label embedding.

[0013] In some embodiments, the multiple heterogeneous skeletal datasets include one or more of NW-UCLA, 3D and 2D versions of NTU-60, NTU-120, and HumanML3D; for HumanML3D, a multi-label classification benchmark is established, including: using a large language model to extract core action verbs from text descriptions, and generating multiple preliminary action labels for each sequence; applying a clustering algorithm to cluster the original labels into semantically coherent categories, and the final multi-label annotation of each sequence is defined as the union of its refined label candidates; using stratified sampling to divide the dataset; dividing the labels into three subsets: head, middle, and tail.

[0014] In some embodiments, the skeleton representation is unified, and the input bone sequence is represented as X orig ∈R N×C×T×K×M , where N is the batch size, C, T, K, and M represent the number of coordinate dimensions, the number of frames, the number of joints, and the number of members, respectively. The maximum number of joints is determined based on the cardinality of the joint and member sets of all datasets D in the training corpus. Maximum number of members in represents the unique set of joints present in the dataset d, Represents an individual set; for input sequences with fewer joints or members than the uniform number, zero padding is performed on the corresponding dimension to ensure that all bone sequences meet K unified ×M unified consistent spatial structure.

[0015] In some embodiments, the feature embedding in the model architecture adopts multimodal embedding and fusion, deriving multiple modalities including joint J, skeleton B and motion M from the unified joint coordinates, mapping each modality to a common latent space, and generating a fused multimodal representation for the temporal stream and a fused multimodal representation for the spatial stream through weighted average fusion and linear projection.

[0016] The spatiotemporal encoding in the model architecture is based on the dual-stream Transformer backbone, which processes the fusion features through parallel temporal and spatial encoders to obtain global temporal features. and global space features By connecting these two global stream representations to form a comprehensive global visual feature N is the batch size, D h is the feature dimension; in the projection layer, three parallel projection networks transform v g , Mapping to text embeddings from pre-trained language models Aligned semantic space:

[0017]

[0018] in D a Represents the dimension of text embedding, each represents the corresponding projection network.

[0019] In some embodiments, the loss based on the global instance alignment objective is expressed as in represents the symmetric contrast loss function, ensuring that the alignment objectives are considered from both visual-to-text and text-to-visual perspectives; the loss based on the flow-specific alignment objective is expressed as The loss based on the fine-grained alignment objective is expressed as where N seg The number of segments into which the time series features are divided, N part The number of groups divided by spatial sequence characteristics; v t,j 、v s,k are the local features corresponding to time period j and spatial part k respectively; the final total loss where λ ts ,λ consis and λ part is a hyperparameter that balances the contribution of each component.

[0020] In some embodiments, the symmetric contrast loss function is:

[0021]

[0022] in represents the cosine similarity and τ is the temperature hyperparameter.

[0023] In a second aspect, the present invention provides an open vocabulary action recognition system based on heterogeneous skeletons, comprising:

[0024] A dataset construction module, which is used to integrate multiple heterogeneous skeleton datasets and construct a heterogeneous open vocabulary skeleton dataset;

[0025] Skeleton representation unification module, which is used to unify the skeleton representations in heterogeneous open vocabulary skeleton datasets and define a unified spatial structure by establishing the maximum number of joints and members;

[0026] The model building module is used to build a skeletal motion encoder model based on the Transformer architecture. The model architecture includes feature embedding, spatiotemporal encoding, and a projection layer for cross-modal alignment. The three parallel projection networks in the projection layer map the global temporal features, global spatial features, and global visual features output by the spatiotemporal encoding to a semantic space aligned with the text embeddings from the pre-trained language model.

[0027] The model training module is used to construct training losses based on the multi-granularity motion-text alignment strategy to train the motion encoder model; the multi-granularity motion-text alignment strategy includes: global instance alignment, the goal is to align the final global visual representation with the corresponding action label embedding at the instance level; stream-specific alignment, the goal is to align the projected global temporal features and spatial features with the action label embedding respectively, and to align the global temporal features with the spatial features; fine-grained alignment, by dividing the time series features into multiple non-overlapping segments and the spatial sequence features into multiple semantic parts, for each time period / spatial part, representative features are obtained, the goal is to align the local features with the action label embedding.

[0028] In a third aspect, the present invention provides a computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the heterogeneous skeleton-based open vocabulary action recognition method are implemented.

[0029] In a fourth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the heterogeneous skeleton-based open vocabulary action recognition method.

[0030] Beneficial effects: The open vocabulary action recognition method based on heterogeneous skeletons provided by the present invention can process heterogeneous skeletons in an open set setting by designing a unified action recognition model. The present invention constructs a large-scale heterogeneous open vocabulary (HOV) skeleton dataset by integrating and improving multiple representative large-scale skeleton-based action datasets. In order to solve the problem of integrated skeleton-based action recognition, the present invention proposes a Transformer-based architecture, which realizes unified representation learning across various heterogeneous skeleton data through multi-granularity motion-text alignment. By capturing information at multiple granularities, the present invention enhances the recognition of rare actions, thereby promoting robust heterogeneous skeleton recognition in open vocabulary scenarios. Compared with the prior art, the present invention has the following advantages:

[0031] 1. Improving Action Recognition Accuracy: Experiments on the HumanML3D dataset demonstrate that the proposed method significantly outperforms existing advanced methods (such as RelationNet, CADA-VAE, and SMIE) in terms of overall accuracy and multi-shot, medium-shot, and few-shot categories. The experiments validate the proposed method's ability to recognize complex human actions in open vocabulary scenarios. In particular, the multi-granularity alignment strategy effectively improves recognition performance for rare action categories when processing long-tail distributed data. Furthermore, model performance is further enhanced after incorporating multimodal inputs, demonstrating that multimodal feature fusion can enhance the model's ability to represent skeletal actions.

[0032] 2. Improved model generalization and robustness: Performance comparison experiments on heterogeneous skeletal action recognition on the NW-UCLA, NTU-60, and NTU-120 datasets show that our method significantly outperforms other baseline methods, demonstrating that the representations learned through multi-granularity alignment have stronger transferability and can adapt to heterogeneous skeletal data from different sources. In the more challenging zero-shot setting, the model achieves 43.5% accuracy through direct inference without a classifier head, highlighting its inherent ability to achieve open vocabulary recognition through cross-domain semantic alignment and verifying the model's generalization and robustness in heterogeneous skeletal scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 Schematic diagram of heterogeneous skeletons from different sources in an embodiment of the present invention.

[0034] Figure 2 This is a diagram of the overall structure of an embodiment of the present invention.

[0035] Figure 3 This is a joint number definition diagram for the HOV skeleton dataset in an embodiment of the present invention.

[0036] Figure 4 This is a long-tail distribution diagram of the action categories of the HumanML3D dataset in an embodiment of the present invention. DETAILED DESCRIPTION

[0037] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0038] like Figure 2 As shown, the embodiment of the present invention discloses an open vocabulary action recognition method based on heterogeneous skeletons, which mainly includes the following steps:

[0039] Step S1, constructing the HOV dataset: integrating multiple heterogeneous skeleton datasets to construct a heterogeneous open vocabulary skeleton dataset.

[0040] Human skeleton data acquired from different sources exhibits inherent structural heterogeneity in terms of the number of joints, topology, and dimensionality. For example, early depth sensors such as Kinect v1 typically provide 20 3D joints, while its successor, Kinect v2, captures 25 3D joints. In contrast, 2D pose estimation from RGB video typically produces 17 2D joints. Motion capture (MoCap) systems may follow a biomechanical model such as SMPL, which contains approximately 22 3D joints. This inherent heterogeneity highlights the need for a unified input processing strategy to facilitate learning across datasets.

[0041] To solve the problem of heterogeneous skeletons, in this embodiment, the HOV skeleton dataset integrates multiple representative heterogeneous skeleton datasets, including NW-UCLA, 3D and 2D versions of NTU-60, NTU-120, and HumanML3D. For the HOV skeleton dataset, the human joint numbers are as follows: Figure 3 However, the raw data and annotations of HumanML3D, a large-scale motion capture dataset for action synthesis, cannot be directly applied to action recognition. Therefore, we manually processed the dataset, and the detailed process will be presented in the subsequent description.

[0042] Open-vocabulary action recognition enables models to recognize new action categories beyond those for which they were trained. Advancing this field requires large-scale datasets covering diverse actions and exhibiting realistic complexity, including data heterogeneity. However, there is currently a lack of large-scale benchmark datasets suitable for open-vocabulary skeleton-based action recognition.

[0043] Since the original annotations contain multiple sub-actions in a single motion sequence, we manually processed these annotations to adapt them to the action recognition task and then established a new multi-label classification benchmark on HumanML3D. By leveraging the rich motion data of HumanML3D, we aim to facilitate action recognition on MoCap data. The construction of this benchmark involves several key steps:

[0044] Description-to-label conversion: A large language model is used to extract core action verbs from text descriptions. For example, "a person squats very low, then jumps in an unsatisfactory manner" is converted to "squat and jump" and then labeled as a list of labels: ["squat", "jump"]. Each description goes through this process, resulting in 3 to 4 preliminary action labels for each sequence.

[0045] Label Space Clustering: We collected all unique action labels in the dataset, generating an initial label space of 8,801 categories. However, this label space suffers from semantic redundancy (e.g., “walk,” “walks,” “walking”) and an extreme long-tail distribution. To alleviate these issues, we applied the Balanced K-means algorithm to cluster the original labels into 400 semantically coherent categories. This step merges rare and similar categories into a more compact and balanced label set. The final multi-label annotation for each sequence is defined as the union of its refined label candidates.

[0046] Stratified Dataset Split: The original split (80% training, 15% testing, 5% validation) was designed for action synthesis. We reorganize the dataset into 70% training and 30% testing to align with classification benchmarks, using stratified sampling to keep the class distribution consistent.

[0047] Long-tail evaluation: Frequency analysis shows that there is a long-tail distribution among the 400 categories, such as Figure 4 As shown in Figure 3. Specifically, the minority class (head) contains a large number of samples, while the majority (tail) has only a few samples. To systematically evaluate the model's performance on classes with different frequency, we divide the labels into three subsets: head (top 10%), middle (middle 30%), and tail (bottom 60%). During evaluation, we report the performance of the entire set and each subset to demonstrate the model's ability to handle rare classes.

[0048] Based on this process, we establish a new multi-label action classification benchmark based on HumanML3D with 400 long-tailed categories. This benchmark serves as a valuable testbed for modeling and understanding complex human actions in an open-set setting.

[0049] Step S2, unified skeleton representation: unify the skeleton representation of the heterogeneous open vocabulary skeleton dataset and define a unified spatial structure by establishing the maximum number of joints and members.

[0050] To process heterogeneous skeletons, the first key step is to standardize these different inputs into a unified format. In this embodiment, let the input skeleton sequence be represented as X orig ∈R N×C×T×K×M , where N is the batch size, C, T, K, and M represent the number of coordinate dimensions, the number of frames, the number of joints, and the number of members, respectively. We define a unified spatial structure by establishing the maximum number of joints and members, denoted by K unified and M unified These values ​​are determined based on the cardinality of the set of joints and members of all datasets D in the training corpus:

[0051]

[0052] in represents the unique set of joints present in the dataset d, Represents a set of individuals. The number of joints or members is less than the uniform number (K′ <K unified or M′ <M unified ) is zero-padded in the corresponding dimension. This adaptive expansion ensures that all bones conform to K unified ×M unified In this embodiment, according to Figure 3For node numbers, zero values ​​are used to fill in the corresponding positions of non-existent nodes. In the member dimension, all joints of newly added blank members are filled with zero values.

[0053] Step S3, motion encoder: Construct a skeletal motion encoder model based on the Transformer architecture. The model architecture includes feature embedding, spatiotemporal encoding, and a projection layer for cross-modal alignment. The three parallel projection networks of the projection layer map the global temporal features, global spatial features, and global visual features output by the spatiotemporal encoding to a semantic space aligned with the text embedding from the pre-trained language model.

[0054] Specifically, in this embodiment, the model architecture consists of three core stages: multimodal input embedding and fusion, spatiotemporal feature extraction, and a projection layer for cross-modal alignment.

[0055] Multimodal embedding and fusion: To capture different aspects of human motion, we derive multiple modalities from the unified joint coordinates: joint J, skeleton B, and motion M. Each derived modality is subjected to modality-specific processing before fusion. We use different embedding modules implemented as MLPs to map the flattened representation of each modality to a vector of dimension D. h These representations are ready for temporal and spatial processing flows. The embedding feature H J 、H B 、H M Generate a fused multimodal representation for temporal streams by weighted averaging fusion followed by linear projection and for spatial streams

[0056] Spatiotemporal Encoding: The dual-stream Transformer-based backbone processes fused features through parallel temporal and spatial encoders. Temporal Encoder The time of fusion is expressed as H mmt Add standard sinusoidal position encoding E pos to inject temporal order information. The sequence is then passed through an L-layer Transformer encoder stack designed to capture the temporal evolution and dynamics of the action. In parallel, the spatial encoder Processing fusion spatial representation H mms Embed the learnable space E spe is added to the input to encode spatial position information. The sequence then passes through an L-layer Transformer encoder to model the complex interrelationships and configurations between joints in a unified skeletal structure. The backbone generates sequence-level feature representations for two streams:

[0057]

[0058] Each Transformer layer uses a multi-head self-attention (MHSA) mechanism and feed-forward networks (FFNs), combined with residual connections and layer normalization to achieve stable training. Global representations are derived by applying maximum pooling along the sequence dimension, thereby obtaining global temporal features. and global space features By connecting these two global stream representations to form a comprehensive global visual feature

[0059] Decoupled Projection: To facilitate multi-granular cross-modal alignment, we propose a decoupled projection strategy. Three parallel projection networks map the encoder output to text embeddings that are consistent with those from the pre-trained language model. Aligned semantic space:

[0060]

[0061] where v,v t , D a Represents the dimension of text embedding. Each represents the corresponding projection network.

[0062] 4. Multi-granularity motion-text alignment

[0063] To make the learned visual representations semantically meaningful, it is crucial to effectively align them with the corresponding textual action labels. We achieve this through a multi-granularity alignment strategy based on contrastive learning principles.

[0064] Global Instance Alignment: The main goal is to align the final global visual representation v with the corresponding action label embedding a at the instance level:

[0065]

[0066] in Denotes a symmetric contrast loss function that ensures alignment objectives from both vision-to-text and text-to-vision perspectives:

[0067]

[0068] Core contrastive loss component Adapted from InfoNCE, the expression is:

[0069]

[0070] here Denotes cosine similarity, τ is a temperature hyperparameter. The denominator aggregates the similarity of positive and negative sample pairs within the batch, including cross-modal y k and the same modality x kThis objective pulls the representations of corresponding visual-text pairs closer together while pushing away non-corresponding pairs. It also encourages the model to learn discriminative visual features by explicitly pushing away the representations of different visual instances.

[0071] Stream-Specific Alignment: To explicitly guide the temporal and spatial processing streams to capture label-related information independently, we introduce a stream-specific alignment (SSA) module. This module enforces two key alignment objectives. First, the spatiotemporal alignment loss Encourage the global temporal features x of the projection t and spatial features v s Aligned with the action label embedding a respectively:

[0072]

[0073] In addition, to promote the consistency between the high-level semantics extracted by the two streams for the same action instance, the SSA module uses the loss Enforce spatiotemporal consistency:

[0074]

[0075] Fine-grained alignment: Relying solely on global features may mask subtle spatiotemporal cues that are crucial for recognizing rare actions. To capture more local details, we propose a fine-grained alignment (FGA) module. In this module, time series features Divided into N seg non-overlapping segments, spatial sequence features Grouped into N part semantic parts (e.g., head, arms, legs based on joint index). For each time segment j and spatial part k, we and Perform maximum pooling within the segment / part features to obtain the representative feature h t,j and h s,k Then use the corresponding projector to project these local features and get and Fine-grained alignment loss These local features are encouraged to align with the action label embedding a:

[0076]

[0077] Training loss: The final training objective function is a weighted sum of these multi-granularity alignment losses:

[0078]

[0079] Here λ ts ,λ consis and λ partis a hyperparameter that balances the contribution of each component. This multi-granularity optimization learns a representation that preserves global action semantics while capturing discriminative spatiotemporal details, thereby achieving robust performance on heterogeneous skeletons. During inference, the most relevant action category is identified by computing the cosine similarity θ between the global visual embedding v and each action label embedding a, achieving open vocabulary evaluation. When performing Zero-Shot experiments, the set of unseen classes is denoted as C novel , the visible class set is recorded as C seen , used to train the model.

[0080] The effects of the present invention compared with the existing methods are described below with reference to specific experiments.

[0081] Table 1: Action recognition results on the HumanML3D dataset

[0082]

[0083] Table 1 presents the action recognition results on the HumanML3D dataset, comparing our method (J modality alone) with zero-shot skeletal action recognition methods such as RelationNet, CADA-VAE, and SMIE. The results show that our method significantly outperforms existing state-of-the-art methods in terms of overall accuracy and performance across multi-shot, medium-shot, and few-shot categories. For example, our method achieves an overall accuracy of 59.34%, compared to SMIE's 49.21%. In the few-shot category, our method achieves 47.36%, significantly exceeding RelationNet's 5.65%. This demonstrates our method's ability to recognize complex human actions in open vocabulary scenarios. In particular, when processing long-tailed data, the multi-granularity alignment strategy effectively improves recognition performance for rare action categories. Furthermore, when fusing multimodal inputs (J+B+M), model performance is further improved, reaching an overall accuracy of 61.23%, demonstrating that multimodal feature fusion can enhance the model's ability to represent skeletal actions.

[0084] Table 2: Performance comparison of heterogeneous skeletal action recognition on the NW-UCLA, NTU-60, and NTU-120 datasets

[0085]

[0086] Table 2 shows the performance comparison of heterogeneous skeleton-based action recognition on the NW-UCLA, NTU-60 (x-sub) and NTU-120 (x-sub) datasets. "Training" indicates whether the encoder or fully connected classification head of the method has been trained / fine-tuned on the corresponding dataset. The method of the present invention is compared with self-supervised learning methods such as LongT GAN, P&C, and GL-Transformer. The results show that when the encoder is frozen and the linear classifier is trained, the method of the present invention (w / FCclassifier) ​​achieves an accuracy of 93.5% on NW-UCLA and 79.5% on NTU-120, significantly outperforming other baseline methods such as GL-Transformer on NW-UCLA with 90.4% and 3s-CF Colorization on NTU-120 with 69.2%. This proves that the representation learned through multi-granularity alignment has stronger transferability and can adapt to heterogeneous skeleton data from different sources. In the more challenging zero-shot setting, the model achieves an accuracy of 43.5% by direct inference without a classifier head, highlighting its inherent ability to achieve open vocabulary recognition through cross-domain semantic alignment and verifying the generalization and robustness of the model in heterogeneous skeleton scenarios.

[0087] Table 3: Effects of multi-granular motion-text alignment on NTU-60 and HumanML3D

[0088]

[0089] Table 3 evaluates the contribution of each component in the proposed multi-granular motion-text alignment framework. Starting with a baseline model, we gradually introduce other alignment losses. The models are jointly trained on NTU-60 (3D and 2D versions) and HumanML3D datasets and then evaluated separately. The results show that and The introduction of continues to improve the classification performance, indicating that both SSA (stream-specific alignment) and FGA (fine-grained alignment) modules help the visual encoder learn more effective heterogeneous skeletal representations. SSA improves recognition accuracy on all evaluation datasets by promoting semantic learning of temporal and spatial streams. Notably, the integration of FGA brings significant performance improvements in the medium-shot and few-shot categories of HumanML3D, indicating that it is effective in identifying rare actions and alleviating biases in the head category. When fusing multimodal inputs (J+B+M, joints+skeletons+motion), the model achieves the best performance, verifying the importance of multimodal feature complementarity in enhancing action representation.

[0090] Table 4: Effect of fine-grained alignment (FGA) partitioning strategy on HumanML3D

[0091]

[0092] The FGA module captures local action semantics by dividing the temporal and spatial feature sequences. To evaluate the impact of different division strategies, we conducted ablation experiments on the HumanML3D dataset for medium-shot and few-shot categories. Table 4 shows that when using N seg = 4 time periods and N part = 4 spatial body parts (head, arms, spine, legs). The results show that FGA significantly improves the recognition of rare actions by leveraging fine-grained local cues. Effectively decomposing actions into meaningful spatiotemporal components helps the model achieve better performance in the long-tail distribution scenario of HumanML3D. Excessive partitioning (such as N seg =6, N part =6) may lead to feature fragmentation, which in turn reduces performance, indicating that a reasonable division granularity is the key.

[0093] Table 5: Cross-dataset generalization performance on the NTU-120 dataset

[0094]

[0095] To further investigate the transferability and generalization capabilities of the representations learned by our model, especially when trained on datasets with different skeletal structures and action vocabularies, we performed a cross-dataset evaluation on the NTU-120 benchmark. Table 5 compares the performance of our method with two state-of-the-art self-supervised representation learning methods, UmURL and USDRL, under different training schemes. We considered two variants of our method: one pre-trained solely on HumanML3D (“Our method (HumanML3D)”) and the other pre-trained on a combination of NTU-60 and HumanML3D (“Our method (NTU60+HumanML3D)”). During evaluation, we freeze the pre-trained backbone and train a linear classifier on the NTU-120 dataset. Results show that our model achieves significant zero-shot transfer performance on NTU-120 even when pre-trained solely on HumanML3D, which differs significantly from NTU-120 in both skeletal structure (22 joints based on SMPL vs. 25 joints based on Kinectv2) and action semantics. Furthermore, incorporating NTU-60 into pre-training (“Our method (NTU60+HumanML3D)”) significantly improves generalization to NTU-120. When using joint-only input (J), our method achieves 82.9% accuracy under cross-subject agreement, outperforming USDRL. When using multimodal input (J+B+M), our model achieves the best performance among all compared methods. These findings further highlight the effectiveness of our method in learning generalizable representations for heterogeneous action recognition.

[0096] To evaluate the open vocabulary recognition capability of our proposed method under more stringent conditions, we conduct zero-shot learning (ZSL) experiments on the HumanML3D multi-label classification benchmark. Specifically, we select the 8 least frequent (tail) categories from the HumanML3D classes to form the unseen class set, denoted as C novel The remaining categories constitute the visible class set C seen , used to train the model. For evaluation, we report the performance under two standard ZSL settings:

[0097] ● Traditional ZSL (ZSL): The test sample is known to belong to C novel ,Accuracy is measured by correctly classifying in the unseen classes.

[0098] ● Generalized ZSL (GZSL): The test sample can belong to C seen or C novel , the model must predict the correct class from the set of all 400 classes. When making predictions in the GZSL setting, given the projected global visual features v of a sample and all action label text features We adopt a calibration scoring mechanism inspired by Calibrated Stacking

[20] . Predicted class c * Determined as:

[0099]

[0100] where θ is the cosine similarity, γ ≥ 0 is a calibration factor that penalizes the visible class, and I[·] is the indicator function. In the GZSL setting, C test =C seen ∪C novel ,γ helps balance the predictions between seen and new categories. We report seen The visible accuracy S of the test sample from C novel The unseen accuracy U of the test samples and their harmonic mean H are defined as

[0101] Table 6: Zero-shot learning (ZSL) and generalized zero-shot learning (GZSL) performance on the HumanML3D benchmark

[0102]

[0103] We compared our method with several skeleton-based zero-shot action recognition methods. To ensure a fair comparison, the visual backbones of all baselines were normalized to match the architecture used in our method. As shown in Table 6, our method consistently outperformed all evaluated baselines on both ZSL and GZSL metrics. Under the traditional ZSL setting, our model achieved an accuracy of 43.26% on the unseen tail category, exceeding SMIE by 3.37%. This demonstrates that our model is able to generalize the learned action-text alignment to new action classes that share no instances with the training set. Our method also showed significant advantages in the more challenging and realistic GZSL scenario. It achieved a harmonic mean (H) score of 31.49%, a balance metric for GZSL performance. This H score is significantly higher than SMIE, highlighting our model's stronger ability to correctly classify seen and unseen instances when the label space contains all categories. These results emphasize the effectiveness of our proposed architecture and alignment mechanism in the challenging task of zero-shot action recognition from heterogeneous skeleton data, particularly in leveraging textual semantics to bridge the gap with unseen action classes.

[0104] The references involved in the above specific embodiments are as follows:

[0105] [1]Jasani B,Mazagonwalla A.Skeleton based zero shot actionrecognition in joint pose-language semantic space[J].arXiv preprint arXiv:1911.11344,2019.

[0106] [2]Schonfeld E,Ebrahimi S,Sinha S,et al.Generalized zero-and few-shotlearning via aligned variational autoencoders[C] / / Proceedings of the IEEE / CVFconference on computer vision and pattern recognition.2019:8247-8255.

[0107] [3]Zhou Y,Qiang W,Rao A,et al.Zero-shot skeleton-based actionrecognition via mutual information estimation and maximization[C] / / Proceedings of the 31st ACM International Conference on Multimedia.2023:5302-5310.

[0108] [4]Zheng N,Wen J,Liu R,et al.Unsupervised representation learningwith long-term dynamics for skeleton based action recognition[C] / / Proceedingsof the AAAI conference on artificial intelligence.2018,32(1).

[0109] [5]Su K,Liu X,Shlizerman E.Predict&cluster:Unsupervised skeletonbased action recognition[C] / / Proceedings of the IEEE / CVF conference oncomputer vision and pattern recognition.2020:9631-9640.

[0110] [6]Lin L,Song S,Yang W,et al.Ms2l:Multi-task self-supervised learningfor skeleton based action recognition[C] / / Proceedings of the 28th ACMinternational conference on multimedia.2020:2490-2498.

[0111] [7]Xu S,Rao H,Hu X,et al.Prototypical contrast and reverseprediction:Unsupervised skeleton based action recognition[J].IEEETransactions on Multimedia,2021,25:624-634.

[0112] [8]Xu Z,Shen X,Wong Y,et al.Unsupervised motion representationlearning with capsule autoencoders[J].Advances in Neural InformationProcessing Systems,2021,34:3205-3217.

[0113] [9]Wang P,Wen J,Si C,et al.Contrast-reconstruction representationlearning for self-supervised skeleton-based action recognition[J].IEEETransactions on Image Processing,2022,31:6224-6238.

[0114]

[10] Cheng Y B,Chen X,Chen J,et al.Hierarchical transformer:Unsupervised representation learningfor skeleton-based human actionrecognition[C] / / 2021 IEEE International Conference on Multimedia and Expo(ICME).IEEE,2021:1-6.

[0115]

[11] Kim B,Chang H J,Kim J,et al.Global-local motion transformer forunsupervised skeleton-based action learning[C] / / European conference oncomputer vision.Cham:Springer Nature Switzerland,2022:209-225.

[0116]

[12] Li L,Wang M,Ni B,et al.3d human action representation learningvia cross-view consistency pursuit[C] / / Proceedings of the IEEE / CVF conferenceon computer vision and pattern recognition.2021:4741-4750.

[0117]

[13] Gao X,Yang Y,Zhang Y,et al.Efficient spatio-temporal contrastivelearning for skeleton-based 3-d action recognition[J].IEEE Transactions onMultimedia,2021,25:405-417.

[0118]

[14] Yang Y,Liu G,Gao X.Motion guided attention learning for self-supervised 3D human action recognition[J].IEEE Transactions on Circuits andSystems for Video Technology,2022,32(12):8623-8634.

[0119]

[15] Zeng Q,Liu C,Liu M,et al.Contrastive 3d human skeleton actionrepresentation learning via crossmoco with spatiotemporal occlusion mask dataaugmentation[J].IEEE Transactions on Multimedia,2023,25:1564-1574.

[0120]

[16] Yang S,Liu J,Lu S,et al.Skeleton cloud colorization forunsupervised 3d action representation learning[C] / / Proceedings of the IEEE / CVF International Conference on Computer Vision.2021:13423-13433.

[0121]

[17] Yang S,Liu J,Lu S,et al.Self-supervised 3D action representationlearning with skeleton cloud colorization[J].IEEE Transactions on PatternAnalysis and Machine Intelligence,2023,46(1):509-524.

[0122]

[18] Sun S,Liu D,Dong J,et al.Unified multi-modal unsupervisedrepresentation learning for skeleton-based action understanding[C] / / Proceedings of the 31st ACM International Conference on Multimedia.2023:2973-2984.

[0123]

[19] Weng W,Wang H,Wang J,et al.Usdrl:Unified skeleton-based denserepresentation learning with multi-grained feature decorrelation[C] / / Proceedings of the AAAI Conference on Artificial Intelligence.2025,39(8):8332-8340.

[0124]

[20] Chao WL, Changpinyo S, Gong B, et al. An empirical study and analysis of generalized zero-shot learning for object recognition in the wild[C] / / Computer Vision-ECCV 2016:14th European Conference,Amsterdam,TheNetherlands,October 11-14,2016,Proceedings,Part II 14.Springer InternationalPublishing,2016:52-68.

[0125] Based on the same inventive concept, an embodiment of the present invention also discloses an open vocabulary action recognition system based on heterogeneous skeletons, including: a data set construction module for integrating multiple heterogeneous skeleton data sets to construct a heterogeneous open vocabulary skeleton data set; a skeleton representation unification module for unifying the skeleton representation in the heterogeneous open vocabulary skeleton data set, and defining a unified spatial structure by establishing a maximum number of joints and members; a model construction module for constructing a skeleton motion encoder model based on the Transformer architecture, the model architecture including feature embedding, spatiotemporal encoding and a projection layer for cross-modal alignment; the global temporal features, global spatial features and global visual features output by the spatiotemporal encoding are mapped to the global temporal features, global spatial features and global visual features from the prediction layer through three parallel projection networks of the projection layer. A semantic space for text embedding alignment of training language models; a model training module for constructing training losses based on a multi-granularity motion-text alignment strategy to train the motion encoder model; the multi-granularity motion-text alignment strategy includes: global instance alignment, the goal is to align the final global visual representation with the corresponding action label embedding at the instance level; stream-specific alignment, the goal is to align the projected global temporal features and spatial features with the action label embedding respectively, and to align the global temporal features with the spatial features; fine-grained alignment, by dividing the time series features into multiple non-overlapping segments and the spatial sequence features into multiple semantic parts, for each time period / spatial part, representative features are obtained, the goal is to align the local features with the action label embedding.

[0126] The specific implementation of each module can be found in the above method embodiment and will not be repeated here.

[0127] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the heterogeneous skeleton-based open vocabulary action recognition method are implemented.

[0128] An embodiment of the present invention further discloses a computer program product, including a computer program, which implements the steps of the heterogeneous skeleton-based open vocabulary action recognition method when executed by a processor.

[0129] The program code for implementing the inventive method can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that the program code, when executed by the processor or controller, causes the steps of the inventive method to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine, or completely on a remote machine or server. The present invention is not described in detail herein, and all of these are known techniques to those skilled in the art.

Claims

1. A heterogeneous skeleton-based open vocabulary action recognition method, characterized in that The steps include: Integrate multiple heterogeneous skeleton datasets to construct a heterogeneous open vocabulary skeleton dataset; Unify the skeleton representation in heterogeneous open vocabulary skeleton datasets and define a unified spatial structure by establishing the maximum number of joints and members; A skeletal motion encoder model based on the Transformer architecture is constructed. The model architecture includes feature embedding, spatiotemporal encoding, and a projection layer for cross-modal alignment. The global temporal features, global spatial features, and global visual features output by the spatiotemporal encoding are mapped to a semantic space aligned with the text embeddings from the pre-trained language model through three parallel projection networks in the projection layer. Construct training loss based on multi-granularity motion-text alignment strategy to train the motion encoder model; The multi-granularity motion-text alignment strategies include: global instance alignment, which aims to align the final global visual representation with the corresponding action label embedding at the instance level; stream-specific alignment, which aims to align the projected global temporal features and spatial features with the action label embedding respectively, and to align the global temporal features with the spatial features; fine-grained alignment, which obtains representative features for each time segment / spatial part by dividing the time series features into multiple non-overlapping segments and grouping the spatial sequence features into multiple semantic parts, with the goal of aligning the local features with the action label embedding.

2. The method for open vocabulary action recognition based on heterogeneous skeletons according to claim 1, characterized in that The multiple heterogeneous skeletal datasets include one or more of NW-UCLA, 3D and 2D versions of NTU-60, NTU-120, and HumanML3D; for HumanML3D, a multi-label classification benchmark is established, including: A large language model is used to extract core action verbs from text descriptions, generating multiple preliminary action labels for each sequence. A clustering algorithm is applied to cluster the original labels into semantically coherent categories, and the final multi-label annotation of each sequence is defined as the union of its refined label candidates; Use stratified sampling to partition the data set; The tags are divided into three subsets: head, middle and tail.

3. The method for open vocabulary action recognition based on heterogeneous skeletons according to claim 1, characterized in that Unified skeleton representation, let the input bone sequence be represented as X orig ∈R N×C×T×K×M , where N is the batch size, C, T, K, and M represent the number of coordinate dimensions, the number of frames, the number of joints, and the number of members, respectively. The maximum number of joints is determined based on the cardinality of the joint and member sets of all datasets D in the training corpus. Maximum number of members in represents the unique set of joints present in the dataset d, Represents an individual set; for input sequences with fewer joints or members than the uniform number, zero padding is performed on the corresponding dimension to ensure that all bone sequences meet K unified ×M unified consistent spatial structure.

4. The method for open vocabulary action recognition based on heterogeneous skeletons according to claim 1, characterized in that The feature embedding in the model architecture adopts multimodal embedding and fusion, deriving multiple modalities including joint J, skeleton B and motion M from the unified joint coordinates, mapping each modality to a common latent space, and generating a fused multimodal representation for the temporal stream and a fused multimodal representation for the spatial stream through weighted average fusion and linear projection.

5. The method for open vocabulary action recognition based on heterogeneous skeletons according to claim 1, characterized in that The spatiotemporal encoding in the model architecture is based on the dual-stream Transformer backbone, which processes the fused features through parallel temporal and spatial encoders to obtain global temporal features. and global space features By connecting these two global stream representations to form a comprehensive global visual feature N is the batch size, D h is the feature dimension; in the projection layer, three parallel projection networks transform v g , Mapping to text embeddings from pre-trained language models Aligned semantic space: where v,v t , D a Represents the dimension of text embedding, each represents the corresponding projection network.

6. The method for open vocabulary action recognition based on heterogeneous skeletons according to claim 5, characterized in that The loss based on the global instance alignment objective is expressed as in represents a symmetric contrast loss function, ensuring that the alignment goal is considered from both vision-to-text and text-to-vision perspectives; The loss based on the flow-specific alignment objective is expressed as The loss based on the fine-grained alignment objective is expressed as where N seg The number of segments into which the time series features are divided, N part The number of groups divided by spatial sequence characteristics; v t,j 、v s,k are the local features corresponding to time period j and spatial part k respectively; the final total loss where λ ts ,λ consis and λ part is a hyperparameter that balances the contribution of each component.

7. The method for open vocabulary action recognition based on heterogeneous skeletons according to claim 6, characterized in that The symmetric contrast loss function is: in represents the cosine similarity and τ is the temperature hyperparameter.

8. An open vocabulary action recognition system based on heterogeneous skeletons, characterized by: include: A dataset construction module, which is used to integrate multiple heterogeneous skeleton datasets and construct a heterogeneous open vocabulary skeleton dataset; Skeleton representation unification module, which is used to unify the skeleton representations in heterogeneous open vocabulary skeleton datasets and define a unified spatial structure by establishing the maximum number of joints and members; The model building module is used to build a skeletal motion encoder model based on the Transformer architecture. The model architecture includes feature embedding, spatiotemporal encoding, and a projection layer for cross-modal alignment. The three parallel projection networks in the projection layer map the global temporal features, global spatial features, and global visual features output by the spatiotemporal encoding to a semantic space aligned with the text embeddings from the pre-trained language model. A model training module, which is used to construct training losses based on a multi-granularity motion-text alignment strategy to train the motion encoder model; The multi-granularity motion-text alignment strategies include: global instance alignment, which aims to align the final global visual representation with the corresponding action label embedding at the instance level; stream-specific alignment, which aims to align the projected global temporal features and spatial features with the action label embedding respectively, and to align the global temporal features with the spatial features; fine-grained alignment, which obtains representative features for each time segment / spatial part by dividing the time series features into multiple non-overlapping segments and grouping the spatial sequence features into multiple semantic parts, with the goal of aligning the local features with the action label embedding.

9. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by a processor, the steps of the heterogeneous skeleton-based open vocabulary action recognition method according to any one of claims 1 to 7 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the heterogeneous skeleton-based open vocabulary action recognition method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Multimodal assembly action recognition method for comparing semantic query

    CN120997911A