Single sample action recognition method and recognition system oriented to shielded skeleton sequence
By adopting parallel fusion module, adaptive component embedding and dynamic relationship learning mechanisms in single-sample action recognition, the problem of low-quality skeleton sequence processing caused by occlusion is solved, and the robustness and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510129578.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
AI Technical Summary
In the context of single-sample learning, the prior art is difficult to effectively deal with low-quality bone sequences caused by occlusion, resulting in insufficient robustness of the model to data containing a large amount of noise.
The network structure parallel to time convolution and spatial convolution is used to obtain joint-level features, and the adaptive component embedding module and dynamic relationship learning mechanism are gradually updated, and the embedded space is optimized by combining component-based comparison learning and measurement-based small sample learning.
The model's robustness to occluded bone sequences is improved, and key motion patterns of visible joints are effectively captured and represented, and the ability to generalize new movements is enhanced.
Smart Images

Figure CN120067840A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a one-shot action recognition method and recognition system for occluded skeleton sequences. Background Art
[0002] Human actions usually occur in three-dimensional space, and three-dimensional skeleton sequences provide more comprehensive information than RGB videos captured by two-dimensional cameras. Nowadays, with the rapid development of low-cost depth sensors such as Microsoft Kinect and ASUS Xtion, it is possible to directly generate accurate skeleton data in real time, making it more accurate, simple, and convenient to capture human skeleton actions. However, in practical applications, it is difficult to obtain enough training samples in a short time to train newly emerging categories. The few-shot learning methods proposed in recent years can effectively solve the problem of data scarcity. When there is only one training sample in each category, few-shot learning evolves into more difficult one-shot learning. In addition, action recognition faces multiple challenges in real-world scenarios, such as changes in the visual environment and occlusion of objects, which can lead to incomplete information of the acquired human skeleton sequences. Therefore, in the context of one-shot learning, improving the robustness of the model to low-quality skeleton sequences with a large amount of noise has become an urgent problem to be solved. Summary of the Invention
[0003] In view of this, the present invention proposes a one-shot action recognition method and recognition system for occluded skeleton sequences to solve the problems existing in the prior art.
[0004] The technical solution adopted by the present invention to solve its technical problems is: a one-shot action recognition method for occluded skeleton sequences, including:
[0005] S1: Obtain an occluded skeleton sequence sample and preprocess the occluded skeleton sequence sample to obtain a skeleton sequence sample with a unified data length. Then, divide the preprocessed skeleton sequence sample into a training set and a test set;
[0006] S2: Construct a one-shot action recognition model and use the training set and the test set for training and testing to obtain a trained one-shot action recognition model. Among them, model training includes the following steps:
[0007] S21: Adopt a network structure with parallel temporal convolution and spatial convolution to obtain joint-level features of the processed skeleton sequence sample;
[0008] S22: Adopt an adaptive part embedding method, use an instance-specific part template to aggregate the joint-level features into an adaptive part-level embedding, and gradually iteratively update the part template to make the part template gradually fuse the sample feature information;
[0009] S23: Enhance the component-level embedding using a dynamic relationship learning mechanism, emphasizing the key features of the task and suppressing task-irrelevant features;
[0010] S24: Input the enhanced component-level embedding into a learning module for one-shot learning. During the one-shot learning process, combine component-based contrastive learning with metric-based few-shot learning;
[0011] S3: Use the trained one-shot action recognition model to recognize the one-shot action and obtain the recognition result.
[0012] Preferably, in S1, adopt a uniform sampling and zero-padding strategy to unify the data format and make the number of time frames of the samples in the sample dataset uniform.
[0013] More preferably, in S21, use a parallel fusion module to obtain the joint-level features of the skeleton sequence samples. Among them, the parallel fusion module network is composed of temporal convolution and spatial graph convolution in parallel. The temporal convolution and spatial graph convolution have a total of 10 layers. The parallel fusion module network uses a two-stream parallel fusion structure to fuse the motion information and joint point information of the action. One stream represents the motion stream, and the other stream represents the position stream. Among them, the motion information obtained in each layer of the motion stream is fused with the position information of that layer in the position stream. The fused information is sent to the next-layer spatial convolution in the motion stream. Finally, the joint-level feature Z is obtained. The motion stream is composed of a total of four layers of temporal convolution and one layer of spatial convolution. The spatial convolution is located in the third layer. The position stream is composed of a total of five layers of temporal convolution.
[0014] More preferably, in S21, it further includes the step of performing max pooling on the position information features along the time dimension.
[0015] More preferably, in S22, the process of integrating the joint-level feature Z into the component-level embedding is as follows:
[0016] The instance-specific component template starts from a set of fixed parameters and is then iteratively updated. The component template in the i-th iteration is denoted as where N represents the number of components in the component template, and C represents the number of channels of the joint-level feature Z;
[0017] where, during the i-th iteration process, first calculate the similarity matrix S i between the joint-level feature Z and the component template β i , where, calculate the similarity between the corresponding component templates and joint-level features of N components respectively, and then splice them into S i , The calculation formula of
[0018]
[0019] Among them, n represents the nth component, t and u respectively represent the t-th time step and the u-th joint in the joint-level feature Z, (·) T , ||·|| 2 respectively represent the transpose function, matrix multiplication, and l 2 -norm, f ψ and f δ represent two independent FC layers, followed by a BN layer and a Relu activation function, f ψ represents applying a non-linear transformation to the component template, f δ represents calculating the cosine distance between the transformed component template and the joint;
[0020] After that, using the similarity matrix S i between the component template and the joint-level feature Z, integrate the joint-level feature Z into the component-level embedding F i in, where the calculation formula for the nth component-level embedding is:
[0021]
[0022] Further preferably, in S22, the process of gradually iteratively updating the component template is as follows:
[0023] Using the component-level embedding F i dynamically update the component template β i , to obtain the updated component template:
[0024] β i+1 = α·F i +(1 - α)·β i ;
[0025] where α represents the update rate of the component template, α ∈ [0, 1], β i represents the component template before update, β i+1 represents the component template after update. Through iterative update, the component template gradually integrates sample feature information.
[0026] Further preferably, in S23, the method of enhancing the component-level embedding using a dynamic relationship learning mechanism is as follows:
[0027] Use the multi-head attention mechanism to calculate the attention scores between components and generate new embeddings for each semantic component;
[0028] Transfer the original features with high attention score semantic components to other components. Conversely, the embeddings of low attention score components usually contain information unrelated to the action and are adjusted by fusing high attention context features;
[0029] For each head j in the self-attention mechanism, the original part-level embeddings of the input skeleton sequence are linearly transformed to produce the query matrix Q j , the key matrix K j and the value matrix V j :
[0030]
[0031] where represent three different parameter matrices learned during training;
[0032] Subsequently, calculate the attention weights A of each part relative to all other parts through the dot product between the query Q j and the linearly transformed features of the key K j : j :
[0033]
[0034] Then, apply the attention weights to the value V j to produce a weighted sum, obtaining the output of the attention mechanism:
[0035]
[0036] Each head focuses on different aspects of the original part-level embedding F, exploring different dependencies between parts. Assuming there are M heads, the outputs of all heads are concatenated and then linearly transformed by W O to produce the enhanced part-level embedding H, with the formula as follows:
[0037]
[0038] Further preferably, in S24, contrastive learning is performed using the enhanced part-level embedding H and the part prototype. A meta-learning strategy is adopted and a K-way One-shot action recognition task is simulated. In one episode, the support set S consists of K randomly sampled classes, with each class having one skeleton sequence. Similarly, the query set Q also contains skeleton sequences from these K classes, with each class containing one sequence. Therefore, B = 2 represents the number of sequences. The enhanced part embeddings of the sequences are divided into support sequences and query sequences for one-shot learning;
[0039] where the part prototype P is the average of the part embeddings of various skeleton sequences in the current episode, and the calculation formula is as follows:
[0040]
[0041] Among them, refers to the enhanced part-level embedding extracted from the bone sequence;
[0042] Compare the enhanced part-level embedding with the part prototypes to obtain the similarity between the embedding and all part prototypes
[0043]
[0044] Among them, H k,b,n refers to the embedding corresponding to the nth part, f φ refers to the non-linear transformation implemented by the FC-BN-Relu structure, and the objective representation of contrastive learning based on parts is the cross-entropy between and the true label:
[0045]
[0046] In metric-based one-shot learning, the goal is to correctly classify the sequences in the query set Q with reference to the sequences of different actions in the support set S. The enhanced part-level embedding H of the sequences in the current episode is divided into two parts, corresponding to the support sequence, corresponding to the query sequence, and the loss function for action classification is expressed as:
[0047]
[0048] Among them, and represent the query sequence and the reference sequence respectively, is the reference sequence of the same class as dis(·, ·) refers to the Euclidean distance between two sequences;
[0049] The overall learning objective is expressed as follows:
[0050] L total = L c + λL p ;
[0051] Among them, λ is a predefined parameter used to balance the two terms.
[0052] Further preferably, in S23, during the model testing process, use the test set to test the trained model, calculate the category of the test sample using the K-nearest neighbor classifier, and predict the classification result of the test data.
[0053] The present invention also provides a single-sample action recognition system for occluded skeleton sequences, which is used to execute the above-mentioned single-sample action recognition method for occluded skeleton sequences.
[0054] The technical solution of the present invention has the following advantages compared with the prior art:
[0055] 1. The present invention adopts parallel fusion of multi-layer temporal convolution and spatial convolution, transmits key information from the motion stream to the joint stream to generate joint-level features, which is more efficient and accurate than previous methods;
[0056] 2. An adaptive part embedding module is proposed. For each data sample, a specific part-level embedding can be adaptively learned. Even in the presence of occlusion, the adaptive part embedding module can capture and represent the key motion patterns of visible joints. This method effectively solves the influence brought by the occlusion phenomenon by using the adaptive part-level embedding method, fully extracts the features of occluded data, and solves the problem of data scarcity.
[0057] 3. A dynamic learning mechanism is adopted to highlight key motion features and at the same time minimize the influence of irrelevant features. This method is particularly obvious for the recognition effect of occluded data. The joint features related to the action are selected according to the attention score, reducing the influence of partial occlusion on action recognition, and combining single-sample learning and contrast learning to optimize the embedding space, improving its generalization ability for new actions. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The present invention will be further described in detail below with reference to the drawings and embodiments:
[0059] Figure 1 is a schematic structural diagram of the parallel fusion module;
[0060] Figure 2 is a schematic diagram of the adaptive part embedding module;
[0061] Figure 3 is a schematic diagram of the dynamic relationship learning mechanism;
[0062] Figure 4 is the entire training flow chart;
[0063] Figure 5 is the test flow chart. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0065] On the one hand, the present invention provides a single-sample action recognition method for occluded skeleton sequences, including:
[0066] S1: Obtain an occluded skeleton sequence sample and preprocess the occluded skeleton sequence sample to obtain a skeleton sequence sample with a unified data length. Then, divide the preprocessed skeleton sequence sample into a training set and a test set;
[0067] Since there are differences in the number of time frames of different samples in the sample dataset, an even sampling and zero-padding strategy is adopted to unify the data format and unify the number of time frames of the samples in the sample dataset;
[0068] Specifically: For an action sequence with a length of T, it is equally divided into R time periods, and then a frame is randomly selected from each time period to represent the action characteristics of that period. In this way, each action sequence is simplified to a representation of R frames. In the case where T is less than R, in order to retain the complete action information, all the frames in the original sequence are directly used and padded with zeros to R frames to meet the frame number requirement.
[0069] The occluded skeleton sequence samples used in this embodiment come from the NTU-RGB+D 60 real occlusion (RE) dataset. The skeleton sequence samples are divided into a training set and a test set. The training set includes 50 categories, and the test set includes 10 categories; The present invention adopts the episode training method. In each episode, the limited labeled data is defined as the support set S, and the unlabeled dataset is defined as the query set Q. If there are K categories in the support set S and 1 sample in each category, then this task is called a K-way One-shot task; Specifically, in each episode, K different categories are randomly selected from the training set, and 1 sample is drawn from each category. These K samples form the support set in the training stage; In addition, K samples different from the support set are further drawn from these K categories to form the query set. The support set samples carry category labels, that is, during the training process, it is necessary to predict the category of the query set based on the support set with category labels to obtain the ability to distinguish these K categories. Because multiple episodes will be constructed during the training process, the purpose is to enable it to learn the ability to predict new categories;
[0070] In this embodiment, the entire dataset is converted into a unified 60-frame length format as follows:
[0071] S11: For samples with a data length exceeding 60 frames, the sample is equally divided into 60 time periods by using the even sampling method, and a frame is randomly selected from each time period to represent the action characteristics of that time period;
[0072] S12: For samples with a data length less than 60 frames, retain all frames of the sample and pad it with zeros to 60 frames;
[0073] S2: Build a single-sample action recognition model and use the training set and test set for training and testing to obtain a trained single-sample action recognition model;
[0074] Among them, as Figure 4 shown, the process of model training is as follows:
[0075] S21: Use a network structure with parallel temporal convolution and spatial convolution to obtain the joint-level features of the processed skeletal sequence samples;
[0076] Use a parallel fusion module to obtain the joint-level features of the skeletal sequence samples. Among them, as Figure 1 shown, the parallel fusion module network consists of parallel temporal convolution and spatial graph convolution. The temporal convolution and spatial graph convolution have a total of 10 layers. The parallel fusion module network uses a two-stream parallel fusion structure to fuse the motion information and joint point information of the action. The left stream represents the motion stream, and the right stream represents the position stream. Among them, the motion information obtained in each layer of the motion stream is fused with the position information of this layer in the position stream, and the fused information is sent to the next-layer spatial convolution in the motion stream, and finally the joint-level feature Z is obtained;
[0077] Among them, as Figure 1 shown, the motion stream consists of a total of four layers of temporal convolution and one layer of spatial convolution, and the spatial convolution is located in the third layer;
[0078] The role of including a spatial convolution layer in the motion stream is to fuse the spatial relationship of the motion information between joints. At the same time, in order to capture long-term dependencies, the temporal resolution of the parallel fusion module is reduced by using strides of 2 and 3 for the second and fourth temporal convolutions in the motion stream;
[0079] Among them, the position stream consists of a total of five layers of temporal convolution. Since different strides of temporal convolution are used in the motion stream, resulting in a mismatch in the feature dimensions of the motion information and position information, the position information features are max-pooled along the temporal dimension so that the dimensions match in each fusion step.
[0080] Among them, the formula for fusing the motion information obtained in each layer of the motion stream with the position information of this layer in the position stream is:
[0081] M s = TCN(M s-1 );
[0082] X s = GCN(Concat(M s-1 , Max(X s-1 )));
[0083] In the formula, M s represents the motion feature of the s-th temporal convolutional output, Xs represents the position feature of the s-th spatial convolutional output, TCN represents the temporal convolutional operation, GCN represents the spatial convolutional operation, Concat represents concatenating the motion information and the spatial information along the temporal dimension, and Max represents the max pooling operation;
[0084] S22: Adopt an adaptive part embedding method, such as Figure 2 shown, use an instance-specific part template to aggregate the joint-level features into an adaptive part-level embedding, and gradually iteratively update the part template so that the part template gradually integrates the sample feature information;
[0085] wherein, the instance-specific part template starts from a set of fixed parameters and is then iteratively updated. The part template in the i-th iteration is denoted as wherein, N represents the number of parts in the part template, and C represents the number of channels of the joint-level feature Z;
[0086] wherein, in the i-th iteration process, first calculate the similarity matrix S i between the joint-level feature Z and the part template β i , wherein, calculate the similarity between the corresponding part template and the joint-level feature of N parts respectively, and then concatenate them into S i , The calculation formula of is as follows:
[0087]
[0088] wherein, n represents the n-th part, t and u respectively represent the t-th time step and the u-th joint in the joint-level feature Z, (·) T , ||·|| 2 respectively represent the transpose function, matrix multiplication and l 2 -norm, f ψ and f δ represent two independent FC layers, followed by a Batch Normalization (BN) layer and a Relu activation function, f ψ represents applying a non-linear transformation to the part template, and f δ represents calculating the cosine distance between the transformed part template and the joint;
[0089] After that, use the similarity matrix S i between the part template and the joint-level feature Z to integrate the joint-level feature Z into the part-level embedding F i ; Among them, the calculation formula for the nth component-level embedding is as follows:
[0090]
[0091] Among them, the process of gradually iteratively updating the component template is as follows:
[0092] Utilize the component-level embedding F i Dynamically update the component template β i , to obtain an updated component template, making it better able to capture the current action sample and discriminate feature information:
[0093] β i+1 = α·F i +(1 - α)·β i ;
[0094] Among them, α represents the update rate of the component template, α ∈ [0, 1], β i represents the component template before update, β i+1 represents the component template after update. Through iterative update, the component template gradually integrates sample feature information, enabling it to adapt to the current bone sequence and express personalized motion patterns.
[0095] S23: Utilize Figure 3 the dynamic relationship learning mechanism shown to enhance the component-level embedding, emphasizing the key features of the task and suppressing the features irrelevant to the task;
[0096] Utilize the multi-head attention mechanism to calculate the attention scores between components and generate new embeddings for each semantic component;
[0097] Transfer the original features of the semantic components with high attention scores as context information to other components. On the contrary, the embeddings of the components with low attention scores usually contain information irrelevant to the action and are adjusted by fusing the high-attention context features;
[0098] For each head j in the self-attention mechanism, the original component-level embedding of the input bone sequence undergoes a linear transformation to generate a query matrix Q j , key matrix K j and value matrix V j :
[0099]
[0100] Among them, represents three different parameter matrices learned during training;
[0101] Subsequently, through the query Q j and the key K jCalculate the attention weight A of each component relative to all other components using the dot product between the features after linear transformation j :
[0102]
[0103] Then, apply the attention weights to the values V j to produce a weighted sum, obtaining the output of the attention mechanism:
[0104]
[0105] Each head focuses on different aspects of the original component-level embedding F, exploring different dependencies between components. Assuming there are M heads, concatenate the outputs of all heads and then perform a linear transformation W O to produce an enhanced component-level embedding H, as follows:
[0106]
[0107] S24: Input the enhanced component-level embedding into the learning module for one-shot learning, where, during the one-shot learning process, combine part-based contrastive learning with metric-based few-shot learning;
[0108] Among them, use the enhanced component-level embedding H and component prototypes for contrastive learning, adopt a meta-learning strategy and simulate a K-way One-shot action recognition task. In one episode, the support set S consists of K randomly sampled classes, each class having one skeleton sequence. Similarly, the query set Q also contains skeleton sequences from these K classes, with each class containing one sequence. Therefore, B = 2 represents the number of sequences. Divide the enhanced component embeddings of the sequences into support sequences and query sequences for one-shot learning;
[0109] Among them, the component prototype P is the average of the component embeddings of various skeleton sequences in the current episode, and the calculation formula is as follows:
[0110]
[0111] Among them, refers to the enhanced component-level embedding extracted from the skeleton sequence;
[0112] Compare the enhanced component-level embedding with the component prototype to obtain the similarity between the embedding and all component prototypes
[0113]
[0114] Among them, H k,b,n refers to the embedding corresponding to the nth component, f φRefers to the non - linear transformation implemented by the FC - BN - Relu structure, and the objective of contrastive learning based on components is expressed as The cross - entropy between
[0115]
[0116] In metric - based one - shot learning, the goal is to correctly classify the sequences in the query set Q with reference to the sequences of different actions in the support set S. Specifically, the augmented part - level embedding H of the sequences in the current episode is divided into two parts, Corresponding to the support sequences, Corresponding to the query sequences, the loss function for action classification is expressed as:
[0117]
[0118] Where, And Represent the query sequence and the reference sequence respectively, Is the reference sequence of the same class as , and dis(·, ·) refers to the Euclidean distance between two sequences;
[0119] The overall learning objective is expressed as follows:
[0120] L total =L c +λL p ;
[0121] Where, λ is a predefined parameter used to balance the two terms.
[0122] During the model testing process, the trained model is tested using the test set. The K - nearest neighbor classifier is used to calculate the class of the test samples, and the classification results of the test data are predicted. The test process is as Figure 5 Shown;
[0123] After training is completed, as long as there is only one labeled sample for each class in the test set, the labeled samples in the test set are used as reference samples, and the unlabeled samples are used as test samples. They are input into the pre - trained recognition model, and the K - nearest neighbor classifier is used to calculate the class of the test samples, and the classification results of the test data are predicted.
[0124] S3: Use the trained one - shot action recognition model to recognize one - shot actions and obtain the recognition results.
[0125] Using the above recognition model, when facing unseen classes, only one labeled sample is needed as a reference sample to obtain good classification results.
[0126] The present invention also provides a single-sample action recognition system for occluded skeleton sequences, which is used to execute the above-mentioned single-sample action recognition method for occluded skeleton sequences.
[0127] The above technical solutions illustrate the technical concept of the present invention. The protection scope of the present invention cannot be limited thereby. Any modification and decoration made to the above technical solutions based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the protection scope of the technical solutions of the present invention.
Claims
1. A single-sample action recognition method for occluded skeleton sequences, characterized in that: include: S1: Obtain occluded skeleton sequence samples and preprocess the occluded skeleton sequence samples to obtain skeleton sequence samples with uniform data length, and then divide the preprocessed skeleton sequence samples into a training set and a test set; S2: constructing a single-sample action recognition model and using the training set and the test set for training and testing to obtain a trained single-sample action recognition model, wherein the model training includes the following steps: S21: A network structure with parallel temporal convolution and spatial convolution is used to obtain joint-level features of the processed skeleton sequence samples; S22: adopting an adaptive component embedding method, aggregating the joint-level features into an adaptive component-level embedding using an instance-specific component template, and gradually iteratively updating the component template so that the component template gradually integrates the sample feature information; S23: Enhance the component-level embedding using a dynamic relational learning mechanism to emphasize key features of the task and suppress features irrelevant to the task; S24: inputting the enhanced component-level embedding into the learning module for single-shot learning, wherein in the single-shot learning process, component-based contrastive learning is combined with metric-based small-shot learning; S3: Use the trained single-sample action recognition model to recognize the single-sample action and obtain the recognition result.
2. The single-sample action recognition method for occluded skeleton sequences according to claim 1, characterized in that: In S1, uniform sampling and zero-filling strategies are adopted to unify the data format and unify the number of time frames of samples in the sample data set.
3. The single-sample action recognition method for occluded skeleton sequences according to claim 1, characterized in that: In S21, a parallel fusion module is used to obtain joint-level features of skeleton sequence samples, wherein the parallel fusion module network is composed of temporal convolution and spatial graph convolution in parallel, and the temporal convolution and spatial graph convolution have a total of 10 layers. The parallel fusion module network uses a two-stream parallel fusion structure to fuse the motion information of the action with the joint point information, one stream represents the motion stream, and the other stream represents the position stream, wherein the motion information obtained in each layer of the motion stream is fused with the position information of the layer in the position stream, and the fused information is sent to the next layer of spatial convolution of the motion stream, and finally the joint-level feature Z is obtained. The motion stream is composed of four layers of temporal convolution and one layer of spatial convolution, and the spatial convolution is located in the third layer. The position stream is composed of five layers of temporal convolution.
4. The single-sample action recognition method for occluded skeleton sequences according to claim 3, characterized in that: S21 also includes a step of performing maximum pooling on the location information features along the time dimension.
5. The single-sample action recognition method for occluded skeleton sequences according to claim 1, characterized in that: In S22, the process of integrating the joint-level feature Z into the part-level embedding is as follows: The instance-specific component template starts with a set of fixed parameters and is then iteratively updated. The component template in the i-th iteration is represented as Where N represents the number of parts in the part template, and C represents the number of channels of the joint-level feature Z; In the i-th iteration process, the joint-level feature Z is first calculated and compared with the component template β in the i-th iteration. i The similarity matrix S between i , Among them, the similarities of the corresponding component templates and joint-level features of N components are calculated separately, and then spliced into S i , The calculation formula is as follows: in, n represents the nth part, t and u represent the tth time step and uth joint in the joint-level feature Z, respectively. T , ||·||2 represent the transpose function, matrix multiplication and l2-norm, respectively, f ψ and f δ represents two independent FC layers, followed by a BN layer and a Relu activation function, f ψ represents the application of nonlinear transformation to the component template, f δ represents the cosine distance between the transformed component template and the joint; Afterwards, the similarity matrix S between the part template and the joint-level feature Z is used i Integrate joint-level features Z into part-level embedding F i middle, The calculation formula for the nth component-level embedding is:
6. The single-sample action recognition method for occluded skeleton sequences according to claim 5, characterized in that: In S22, the process of gradually iterating and updating the component template is as follows: Using the component-level embedding F i Dynamically update the widget template β i , get the updated component template: β i+1 =α·F i +(1-α)·β i ; Among them, α represents the update rate of the component template, α∈[0,1], β i represents the component template before updating, β i+1 Represents the updated component template. Through iterative updating, the component template gradually integrates the sample feature information.
7. The single-sample action recognition method for occluded skeleton sequences according to claim 1, characterized in that: In S23, the method of enhancing the component-level embedding using the dynamic relationship learning mechanism is as follows: Use the multi-head attention mechanism to calculate the attention scores between components and generate new embeddings for each semantic component; The original features of semantic parts with high attention scores are transferred as contextual information to other parts. In contrast, the embeddings of parts with low attention scores usually contain information irrelevant to the action and are adjusted by fusing high-attention contextual features. For each head j in the self-attention mechanism, the original part-level embedding of the input skeleton sequence After linear transformation, the query matrix Q is generated j , key matrix K j Sum value matrix V j : in, represents the three different parameter matrices learned during training; Then, by querying Q j and key K j The dot product between the linearly transformed features is used to calculate the attention weight A of each component relative to all other components. j : Then, the attention weight is applied to the value V j Produce a weighted sum and get the output of the attention mechanism: Each head focuses on different aspects of the original part-level embedding F and explores different dependencies between parts. Assuming there are M heads, the outputs of all heads are concatenated and then transformed by the linear transformation W. O To generate the enhanced component-level embedding H, the formula is as follows:
8. The single-sample action recognition method for occluded skeleton sequences according to claim 1, characterized in that: In S24, the enhanced component-level embedding H and component prototypes are used for comparative learning, a meta-learning strategy is adopted and the K-way One-shot action recognition task is simulated. In one episode, the support set S consists of K randomly sampled classes, each class has a skeleton sequence. Similarly, the query set Q also contains skeleton sequences from these K classes, each class contains one sequence. Therefore, B = 2 represents the number of sequences. The enhanced component embedding of the sequence is divided into support sequence and query sequence for single-sample learning; Among them, the component prototype P is the average of the component embeddings of various skeleton sequences in the current episode, and the calculation formula is as follows: in, refers to the enhanced part-level embeddings extracted from skeleton sequences; Compare the enhanced part-level embeddings with part prototypes to get the similarity between the embeddings and all part prototypes Among them, H k,b,n refers to the embedding corresponding to the nth component, f φ It refers to the nonlinear transformation implemented by the FC-BN-Relu structure. The objective of component-based contrastive learning is expressed as The cross entropy between and the true label: In metric-based single-shot learning, the goal is to correctly classify the sequences in the query set Q using the sequences of different actions in the support set S as reference. The enhanced component-level embedding H of the sequence in the current episode is divided into two parts: corresponds to the support sequence, Corresponding to the query sequence, the loss function of action classification is expressed as: in, and represent the query sequence and the reference sequence, respectively. is with For reference sequences of the same category, dis(·,·) refers to the Euclidean distance between the two sequences; The overall learning objectives are stated as follows: THE total =L c +λL p ; Here, λ is a predefined parameter used to balance the two terms.
9. The single-sample action recognition method for occluded skeleton sequences according to claim 1, characterized in that: In S23, during the model testing process, the trained model is tested using the test set, the category of the test sample is calculated using the K nearest neighbor classifier, and the classification result of the test data is predicted.
10. A single-sample action recognition system for occluded skeleton sequences, characterized in that: Used to execute the single-sample action recognition method for occluded skeleton sequences as described in any one of claims 1-9.
Citation Information
Cited By
Behavior recognition method and system based on artificial intelligence
CN122049981A
An artificial intelligence-based behavior recognition method and system
CN122049981B