Skeleton action recognition method based on channel feature fusion

By adopting a network structure of channel feature fusion in the skeleton action recognition method, including a channel dynamic modeling graph convolution network, a multi-scale fusion time convolution network and a channel attention module, the shortcomings of the existing technology in handling similar actions and improving recognition accuracy are solved, and a more efficient skeleton action recognition effect is achieved.

CN119992659APending Publication Date: 2025-05-13XINJIANG UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510130100.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing skeleton action recognition methods are difficult to distinguish when dealing with similar or fuzzy actions, and rely on hand-designed feature extraction technology, which has limitations in feature representation ability and generalization ability.

Method used

The skeleton action recognition method based on channel feature fusion is adopted. The network includes 10 layers of spatiotemporal convolution modules. Each layer includes a channel dynamic modeling graph convolution network, a multi-scale fusion time convolution network and a channel attention module. These modules extract the spatial structure characteristics of the skeleton diagram and the dynamic features of joints changing over time in the action, and improve the ability to distinguish different actions through the channel attention module.

Benefits of technology

It improves the accuracy of skeleton action recognition and the classification performance of complex actions, can more effectively distinguish similar actions, and improves the robustness and computing efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992659A_ABST
    Figure CN119992659A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of behavior recognition, and relates to a skeleton action recognition method based on channel feature fusion, which comprises the following steps: (1) acquiring three-dimensional skeleton data of a human body, (2) constructing a skeleton action recognition network based on channel feature fusion, (3) constructing a spatial-temporal feature refinement contrast learning training framework, (4) constructing a training target loss function, and (5) constructing a skeleton action recognition network based on channel feature fusion. And (5) model optimization and training. Experimental results of the skeleton action recognition network based on channel feature fusion provided by the invention on two large skeleton data sets NTU-RGB + D and NTU-RGB + D120 show that the skeleton action recognition network based on channel feature fusion shows excellent performance in a skeleton action recognition task. The action recognition accuracy rates of 92.9% and 97.1% are respectively obtained under the CS and CV division modes of the NTURGB + D60 data set, and the action recognition accuracy rates of 89.9% and 91.5% are respectively obtained under the CS and CT division modes of the NTURGB + D120 data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a skeleton action recognition method based on channel feature fusion, belonging to the technical field of behavior recognition. Background Art

[0002] Human action recognition is one of the key technologies in the field of computer vision and artificial intelligence, and is widely used in many fields such as human-computer interaction, intelligent monitoring, health care, virtual reality, etc. In these applications, skeleton action recognition, as a technology for action recognition based on human skeleton data, has been proven to be an effective and efficient solution.

[0003] Usually, a depth sensor is used to directly capture the 3D skeleton data of the characters in the target scene or a posture estimation algorithm is used to extract the 3D skeleton data of the human body from the RGB video, which is expressed as the 2D or 3D coordinates of the human body joints in each frame. In traditional action recognition methods, the spatiotemporal feature extraction of skeleton data is a key step in the recognition task. The advantage of 3D skeleton data is that it is more robust to background interference and has higher computational efficiency because the joint structure of the skeleton is usually simpler.

[0004] However, existing graph convolutional network-based methods, while showing good performance in spatiotemporal feature extraction, still face many challenges in dealing with similar or ambiguous actions. A major problem is that skeleton data lacks interactive objects and contextual information, making it difficult to distinguish between similar actions. For example, in skeleton data, the actions "writing", "reading", and "typing" may show similar joint movement patterns, which are difficult to distinguish in the absence of other contextual information.

[0005] In addition, most existing skeleton action recognition methods rely on manually designed feature extraction techniques, which have limitations in feature representation and generalization capabilities. Especially when faced with complex or fine-grained actions, traditional methods often find it difficult to fully capture the detailed information of the action, resulting in reduced recognition accuracy.

[0006] With the development of deep learning technology, skeleton action recognition methods based on deep learning have gradually become a hot topic of research, especially models based on graph convolutional networks. Compared with traditional methods, graph convolutional networks can better process spatiotemporal features and capture the spatial dependencies between skeleton joints. As a relatively pioneering skeleton recognition method, spatiotemporal graph convolutional networks have made significant progress. In addition, recent studies have proposed more complex network structures, such as multi-scale feature extraction, spatiotemporal adjacency matrix optimization methods, which aim to improve the limitations of traditional methods and further improve the accuracy of skeleton action recognition. Despite this, existing methods still face certain challenges in processing similar actions, reducing computational complexity and improving recognition accuracy, especially when faced with the distinction between complex and ambiguous actions, the performance of the model is still insufficient.

[0007] Therefore, how to enhance the model's ability to capture spatiotemporal features through effective network structure design and solve the shortcomings of existing technologies in processing similar actions and improving recognition accuracy has become an urgent problem to be solved in the field of skeleton action recognition. Summary of the invention

[0008] In order to solve the deficiencies in the prior art, the purpose of the present invention is to provide a skeleton action recognition method based on channel feature fusion, the network contains 10 layers of spatiotemporal convolution modules, each layer includes a channel dynamic modeling graph convolution network, a multi-scale fusion temporal convolution network and a channel attention module; the channel dynamic modeling graph convolution network is used to extract the spatial structural features of the skeleton graph and describe the topological relationship between human joints, the multi-scale fusion temporal convolution network is used to capture the dynamic features of the joints changing over time in the action, and the channel attention module improves the network's ability to distinguish different actions by weighting the feature channels.

[0009] In order to achieve the above-mentioned invention object and solve the problems existing in the prior art, the technical solution adopted by the present invention is: a skeleton action recognition method based on channel feature fusion, comprising the following steps:

[0010] Step 1: Obtain the 3D skeleton data of the human body. Use a depth sensor to directly capture the 3D skeleton data of the person in the target scene or use a posture estimation algorithm to extract the 3D skeleton data of the human body from the RGB video. The spatial three-dimensional coordinates represent the spatial position of each skeleton node, and a three-dimensional skeleton sequence data of each sample is collected. , where T represents the number of skeleton graphs contained in the sequence, and V represents the number of nodes contained in each skeleton graph;

[0011] Step 2: Construct a skeleton action recognition network based on channel feature fusion. The network contains 10 layers of spatiotemporal convolution modules, each of which contains a channel dynamic modeling graph convolution network, a multi-scale fusion time convolution network and a channel attention module; the channel dynamic modeling graph convolution network is used to extract the spatial structural features of the skeleton graph and describe the topological relationship between human joints, the multi-scale fusion time convolution network is used to capture the dynamic features of the joints in the action over time, and the channel attention module improves the network's ability to distinguish different actions by weighting the feature channels. It specifically includes the following sub-steps:

[0012] (a) Channel dynamic modeling graph convolutional network is used to model the spatial relationship between joints in the skeleton graph and use dynamic channel topology to enhance feature representation. Each spatial modeling module contains three parallel channel topology optimization graph convolution modules (CTRGC for short). This module is a variant of the graph convolutional network (GCN) for skeleton data. It focuses on improving the feature modeling ability by improving the graph topology of each channel. Processing to obtain high-dimensional features , described by formula (1):

[0013]

[0014] in , Represents one of the three parallel modules, which outputs three features Add and use the ReLU activation function for nonlinear transformation, which is described by formula (2):

[0015]

[0016] in In order to maintain training stability and efficiency, the initial input data The convolution changes the channel dimension to match The size of the final skeleton space feature is obtained , described by formula (3):

[0017]

[0018] in It is a graph feature representation of skeleton data, which contains the geometric structure and spatial dependency between joints.

[0019] (b) The multi-scale fusion temporal convolutional network is used to capture the changing characteristics of joints in action sequences over time. The network contains four branches. The first branch contains a Convolution and Temporal convolution; the second branch contains a Convolution and Temporal convolution; the third branch contains a convolution and a max pooling layer; the fourth branch has only one Convolution; the time convolution of the first two branches has different dilation rates, The input is sent to four branches for processing, including the following steps:

[0020] 1) The first temporal convolution branch unit, dilation rate Temporal convolution increases the size of the receptive field by inserting gaps between elements in the convolution kernel, so that the convolution can cover a longer time range, thereby more effectively capturing long-term dependencies in the time series. The convolution operation is performed on all channels along the time dimension, which is described by formula (4):

[0021]

[0022] in, Represents the first temporal convolution branch The time characteristic output of Expansion rate ;

[0023] 2) The second temporal convolution branch unit, the expansion rate , expand the receptive field of the convolution, capture features of a longer time range, and the output features of the second branch , described by formula (5):

[0024]

[0025] 3) The third maximum pooling branch unit compresses the data volume by taking the maximum value of the local area while retaining the significant features. , described by formula (6):

[0026]

[0027] 4) The fourth 1×1 convolution unit first uses a convolution kernel size of The convolutional layer reduces the channel dimension to obtain , described by formula (7):

[0028]

[0029] The four branch features are concatenated and fused to obtain multi-scale fusion features :

[0030]

[0031] In order to maintain training stability and efficiency, the initial input data After convolution, the channel dimension is changed and added to , and obtain the final time feature , described by formula (9):

[0032]

[0033] in, Contains the temporal dynamic information of skeleton nodes;

[0034] (c) The channel attention module is used to dynamically adjust the importance of the channel, highlight important channel features, and Global average pooling is performed on each channel of to compress the spatial dimension, which is described by formula (10):

[0035]

[0036] in, It is a channel The output feature vector of is then dynamically convolved to capture the local channel relationship, which is described by formula (11):

[0037]

[0038] in, Contains the correlation between channels, and then uses the sigmoid activation function Normalized, the weight range is , get the features is the weight of the channel, Expand the dimension to match The channel weights are applied to the input features to match the size of , and obtain the final feature representation , described by formula (12):

[0039]

[0040] in, represents the output features, Represents element-by-element multiplication, weighting the channels with channel weights to enhance useful features and suppress irrelevant features;

[0041] After the feature extraction is completed through the network, the fully connected layer is used to map the features and perform average pooling operations. Finally, the prediction results are output through the softmax function to obtain the probability distribution of each skeleton action.

[0042] Step 3: Construct a spatiotemporal feature refinement contrastive learning training framework, which includes a spatiotemporal feature refinement module and a confident sample guided contrastive learning module; the spatiotemporal feature refinement module is used to process the skeleton features extracted by the backbone network, and refine the skeleton features from the time and space dimensions respectively to obtain the time features and space features of the skeleton features; the confident sample guided contrastive learning module optimizes the feature distribution to make the sample closer to its category, which specifically includes the following sub-steps:

[0043] (a) Spatiotemporal feature refinement module. When processing skeleton sequence data, the separation of spatiotemporal features helps to better capture the structure and motion information of the skeleton, which is beneficial to the skeleton features obtained by the backbone network. Processing from the time and space dimensions includes the following processes:

[0044] 1) Input features Perform an average operation in the time dimension to eliminate the influence of time dynamic information and obtain the feature Spatial representation of , and then use Convolution transforms the feature dimension Compress to , in order to reduce the feature dimension, reduce the computational complexity, and adapt to the processing of high-dimensional skeleton data:

[0045]

[0046] in, ,Will Features are flattened into one-dimensional vectors ;

[0047] 2) Input features Perform an average operation in the spatial dimension to remove the influence of the structural position and retain only the temporal dynamic features to obtain the feature Spatial representation of ,use Convolution transforms the feature dimension Compress to :

[0048]

[0049] in, ,Will Features are flattened into one-dimensional vectors ;

[0050] (b) Confident sample guided contrastive learning module. The goal of this module is to refine the skeleton action feature representation through confident sample guided contrastive learning, optimize the temporal features and spatial features respectively, so that samples of the same type are clustered in the feature space and samples of different types are separated from each other. The core of the module is to use the clustering information of confident samples and fuzzy samples to improve the feature discrimination ability through the contrastive learning module, and use the contrast relationship between the basic action truth value and the fuzzy action to improve the sample representation, including the following processes:

[0051] 1) Credible sample clustering. In each training process, the output of the model includes the feature vector of each sample and the corresponding category label. For an action category , if the sample Correctly classified into categories , then it is considered to be a credible sample. The feature prototype is the global representation of the category and represents the typical features of the category.

[0052] Use the exponential moving average method to update the prototype of each category , ensuring that the updating process of feature prototypes is smooth and stable, It's action The size of the credible sample set is , the exponential moving average operation is described by formula (15):

[0053]

[0054] in, It's action The prototype, represents the number of training rounds, It is a sample The characteristic representation of Momentum coefficient, controls the update speed, usually set to 0.9, for smooth transition and avoid large prototype fluctuations. During training, the prototype becomes the action A stable estimate of the cluster center, which can learn the characteristics of new samples;

[0055] 2) Fuzzy sample clustering: During the training process, the model will encounter some samples that are difficult to classify. These samples are easily confused with samples of other categories in the feature space, which are called fuzzy samples. These samples include two categories: false negative (False Negative, FN), that is, a sample actually belongs to a category , but is misclassified as other categories, a false positive (FalsePositive, FP) means that a sample does not actually belong to a category , but was misclassified as category , , For Action The false negative and false positive sample sets are of size , , using the average value of the sample set as the sample center, described by formula (16):

[0056]

[0057] in , Representative Action The center of the false negative samples and false positive samples is represented by In order to represent the fuzzy samples and distinguish the two fuzzy samples, a compensation term is introduced for the false negative samples, aiming to bring the feature vector of the false negative samples closer to the center of its correct category. For the false negative samples, the compensation term is defined as:

[0058]

[0059] By minimizing the compensation term , making the false negative samples closer to the credible samples in the feature space. When there are no false negative samples or the cosine distance converges to 1, Reaching the minimum value of 0 makes the model treat these blurred samples as actions Similarly, for false positive samples, the penalty term is defined as:

[0060]

[0061] By minimizing the penalty , so that the false positive samples are far away from the credible samples in the feature space. When there are no false positive samples or the cosine distance converges to -1, Reaching a minimum value of 0 prevents the model from identifying these ambiguous samples as actions ;

[0062] Finally, the sample As the anchor point, the proposed contrast loss function can be defined as:

[0063]

[0064] in The function represents the similarity of the calculated vector. Here, cosine similarity is used to measure the similarity between two sample features. For sample right The predicted probability score of the class, is the temperature parameter used to adjust the sensitivity of contrastive learning and the spatial characteristics of each sample in the training batch. and time characteristics Apply formula (19) to calculate the loss respectively, and add the two losses as the contrastive learning loss of each sample , described by formula (20):

[0065]

[0066] By refining the spatial and temporal features of the skeleton, and optimizing the comparative learning of credible samples and fuzzy samples, the representation capability of the model skeleton action features is refined, and the model's classification performance for complex and similar actions is improved.

[0067] Step 4: Construct the training target loss function. The final training target combines the cross entropy loss and contrastive learning loss. The contrastive learning loss is calculated using formula (20). The cross entropy loss is It is used to supervise classification tasks and optimize the classification accuracy of the model, which is described by formula (21):

[0068]

[0069] in is the number of samples in the batch, is the number of action classification categories, It is a sample Corresponding category The label of It is a sample When the target class is ,otherwise , Samples predicted by the network belong The probability score of the class is combined with the cross entropy loss and the contrastive learning loss to form the complete objective function:

[0070]

[0071] Step 5: Model optimization and training. Use the stochastic gradient descent optimizer to train the entire model. Adjust the model parameters to minimize the cross entropy loss and contrastive learning loss. During the training process, adaptively adjust the learning rate and momentum coefficient to ensure network convergence.

[0072] Beneficial effects of the present invention: A skeleton action recognition method based on channel feature fusion, comprising the following steps: (1) obtaining human three-dimensional skeleton data, (2) constructing a skeleton action recognition network based on channel feature fusion, (3) constructing a spatiotemporal feature refinement contrast learning training framework, (4) constructing a training target loss function, and (5) model optimization and training. The skeleton action recognition network based on channel feature fusion constructed in the present invention comprises 10 layers of spatiotemporal convolution modules, each layer comprising a channel dynamic modeling graph convolution network, a multi-scale fusion temporal convolution network and a channel attention module; the channel dynamic modeling graph convolution network is used to extract the spatial structural features of the skeleton graph and describe the topological relationship between human joints, the multi-scale fusion temporal convolution network is used to capture the dynamic features of the joints changing over time in the action, and the channel attention module improves the network's ability to distinguish different actions by weighting the feature channels. The experimental results of the skeleton action recognition network based on channel feature fusion proposed in the present invention on two large skeleton datasets, NTU-RGB+D and NTU-RGB+D120, show that the present invention exhibits excellent performance in the skeleton action recognition task. In the CS and CV partition modes of the NTURGB+D60 dataset, the action recognition accuracy rates were 92.9% and 97.1% respectively, and in the CS and CT partition modes of the NTURGB+D120 dataset, the action recognition accuracy rates were 89.9% and 91.5% respectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 The figure is a flow chart of the steps of the method of the present invention.

[0074] Figure 2 Schematic diagram of the network model and comparative learning framework of the present invention.

[0075] Figure 3 Schematic diagram of the spatiotemporal convolution module of the present invention.

[0076] Figure 4 Schematic diagram of the spatiotemporal feature refinement and comparative learning of the present invention. DETAILED DESCRIPTION

[0077] like Figure 1 As shown, a skeleton action recognition method based on channel feature fusion includes the following steps:

[0078] Step 1: Obtain the 3D skeleton data of the human body. Use a depth sensor to directly capture the 3D skeleton data of the person in the target scene or use a posture estimation algorithm to extract the 3D skeleton data of the human body from the RGB video. The spatial three-dimensional coordinates represent the spatial position of each skeleton node, and a three-dimensional skeleton sequence data of each sample is collected. , where T represents the number of skeleton graphs contained in the sequence, and V represents the number of nodes contained in each skeleton graph;

[0079] Step 2: Construct a skeleton action recognition network based on channel feature fusion. The network model and contrastive learning framework are shown in the figure below. Figure 2 As shown in the figure, the network contains 10 layers of spatiotemporal convolution modules, each layer contains a channel dynamic modeling graph convolution network, a multi-scale fusion time convolution network and a channel attention module; the channel dynamic modeling graph convolution network is used to extract the spatial structural features of the skeleton graph and describe the topological relationship between the human body joints, the multi-scale fusion time convolution network is used to capture the dynamic features of the joints changing over time in the action, and the channel attention module improves the network's ability to distinguish different actions by weighting the feature channels. The schematic diagram of the spatiotemporal convolution module is shown in the figure. Figure 3 As shown, it specifically includes the following sub-steps:

[0080] (a) Channel dynamic modeling graph convolutional network is used to model the spatial relationship between joints in the skeleton graph and use dynamic channel topology to enhance feature representation, such as Figure 3 As shown in (a), each spatial modeling module contains three parallel channel topology optimization graph convolution modules (CTRGC for short), which is a variant of the graph convolution network (GCN) for skeleton data. It focuses on improving the feature modeling ability by improving the graph topology of each channel. Processing to obtain high-dimensional features , described by formula (1):

[0081]

[0082] in , Represents one of the three parallel modules, which outputs three features Add and use the ReLU activation function for nonlinear transformation, which is described by formula (2):

[0083]

[0084] in In order to maintain training stability and efficiency, the initial input data The convolution changes the channel dimension to match The size of the final skeleton space feature is obtained , described by formula (3):

[0085]

[0086] in It is a graph feature representation of skeleton data, which contains the geometric structure and spatial dependency between joints.

[0087] (b) Multi-scale fusion temporal convolutional network is used to capture the changing characteristics of joints in action sequences over time, such as Figure 3 As shown in (b), the network contains 4 branches, the first branch contains a Convolution and Temporal convolution; the second branch contains a Convolution and Temporal convolution; the third branch contains a convolution and a max pooling layer; the fourth branch has only one Convolution; the time convolution of the first two branches has different dilation rates, The input is sent to four branches for processing, including the following steps:

[0088] 1) The first temporal convolution branch unit, dilation rate Temporal convolution increases the size of the receptive field by inserting gaps between elements in the convolution kernel, so that the convolution can cover a longer time range, thereby more effectively capturing long-term dependencies in the time series. The convolution operation is performed on all channels along the time dimension, which is described by formula (4):

[0089]

[0090] in, Represents the first temporal convolution branch The time characteristic output of Expansion rate ;

[0091] 2) The second temporal convolution branch unit, the expansion rate , expand the receptive field of the convolution, capture features of a longer time range, and the output features of the second branch , described by formula (5):

[0092]

[0093] 3) The third maximum pooling branch unit compresses the data volume by taking the maximum value of the local area while retaining the significant features. , described by formula (6):

[0094]

[0095] 4) The fourth 1×1 convolution unit first uses a convolution kernel size of The convolutional layer reduces the channel dimension to obtain , described by formula (7):

[0096]

[0097] The four branch features are concatenated and fused to obtain multi-scale fusion features :

[0098]

[0099] In order to maintain training stability and efficiency, the initial input data After convolution, the channel dimension is changed and added to , and obtain the final time feature , described by formula (9):

[0100]

[0101] in, Contains the temporal dynamic information of skeleton nodes;

[0102] (c) The channel attention module is used to dynamically adjust the importance of channels and highlight important channel features, such as Figure 3 As shown in (c), for the input features Global average pooling is performed on each channel of to compress the spatial dimension, which is described by formula (10):

[0103]

[0104] in, It is a channel The output feature vector of is then dynamically convolved to capture the local channel relationship, which is described by formula (11):

[0105]

[0106] in, Contains the correlation between channels, and then uses the sigmoid activation function Normalized, the weight range is , get the features is the weight of the channel, Expand the dimension to match The channel weights are applied to the input features to match the size of , and obtain the final feature representation It is described by formula (12):

[0107]

[0108] in represents the output features, Represents element-by-element multiplication, weighting the channels with channel weights to enhance useful features and suppress irrelevant features;

[0109] After the feature extraction is completed through the network, the fully connected layer is used to map the features and perform average pooling operations. Finally, the prediction results are output through the softmax function to obtain the probability distribution of each skeleton action.

[0110] Step 3: Construct a spatiotemporal feature refinement contrastive learning training framework, which includes a spatiotemporal feature refinement module and a confident sample guided contrastive learning module. The spatiotemporal feature refinement contrastive learning diagram is shown in the figure below: Figure 4 As shown in the figure, the spatiotemporal feature refinement module is used to process the skeleton features extracted by the backbone network, and refine the skeleton features from the time and space dimensions respectively to obtain the time features and space features of the skeleton features; the confident sample guided contrastive learning module optimizes the feature distribution to make the sample closer to its category, which specifically includes the following sub-steps:

[0111] (a) Spatiotemporal feature refinement module. When processing skeleton sequence data, the separation of spatiotemporal features helps to better capture the structure and motion information of the skeleton, which is beneficial to the skeleton features obtained by the backbone network. Processing from the time and space dimensions includes the following processes:

[0112] 1) Input features Perform an average operation in the time dimension to eliminate the influence of time dynamic information and obtain the feature Spatial representation of , and then use Convolution transforms the feature dimension Compress to , in order to reduce the feature dimension, reduce the computational complexity, and adapt to the processing of high-dimensional skeleton data:

[0113]

[0114] in, ,Will Features are flattened into one-dimensional vectors ;

[0115] 2) Input features Perform an average operation in the spatial dimension to remove the influence of the structural position and retain only the temporal dynamic features to obtain the feature Spatial representation of ,use Convolution transforms the feature dimension Compress to :

[0116]

[0117] in, ,Will Features are flattened into one-dimensional vectors ;

[0118] (b) Confident sample guided contrastive learning module. The goal of this module is to refine the skeleton action feature representation through confident sample guided contrastive learning, optimize the temporal features and spatial features respectively, so that samples of the same type are clustered in the feature space and samples of different types are separated from each other. The core of the module is to use the clustering information of confident samples and fuzzy samples to improve the feature discrimination ability through the contrastive learning module, and use the contrast relationship between the basic action truth value and the fuzzy action to improve the sample representation, including the following processes:

[0119] 1) Credible sample clustering. In each training process, the output of the model includes the feature vector of each sample and the corresponding category label. For an action category , if the sample Correctly classified into categories , then it is considered to be a credible sample. The feature prototype is the global representation of the category and represents the typical features of the category.

[0120] Use the exponential moving average method to update the prototype of each category , ensuring that the updating process of feature prototypes is smooth and stable, It's action The size of the credible sample set is , the exponential moving average operation is described by formula (15):

[0121]

[0122] in, It's action The prototype, represents the number of training rounds, It is a sample The characteristic representation of Momentum coefficient, controls the update speed, usually set to 0.9, for smooth transition and avoid large prototype fluctuations. During training, the prototype becomes the action A stable estimate of the cluster center, which can learn the characteristics of new samples;

[0123] 2) Fuzzy sample clustering: During the training process, the model will encounter some samples that are difficult to classify. These samples are easily confused with samples of other categories in the feature space, which are called fuzzy samples. These samples include two categories: false negative (False Negative, FN), that is, a sample actually belongs to a category , but is misclassified as other categories, a false positive (FalsePositive, FP) means that a sample does not actually belong to a category , but was misclassified as category , , For Action The false negative and false positive sample sets are of size , , using the average value of the sample set as the sample center, described by formula (16):

[0124]

[0125] in , Representative Action The center of the false negative samples and false positive samples is represented by In order to represent the fuzzy samples and distinguish the two fuzzy samples, a compensation term is introduced for the false negative samples, aiming to bring the feature vector of the false negative samples closer to the center of its correct category. For the false negative samples, the compensation term is defined as:

[0126]

[0127] By minimizing the compensation term , making the false negative samples closer to the credible samples in the feature space. When there are no false negative samples or the cosine distance converges to 1, Reaching the minimum value of 0 makes the model treat these blurred samples as actions Similarly, for false positive samples, the penalty term is defined as:

[0128]

[0129] By minimizing the penalty , so that the false positive samples are far away from the credible samples in the feature space. When there are no false positive samples or the cosine distance converges to -1, Reaching a minimum value of 0 prevents the model from identifying these ambiguous samples as actions ;

[0130] Finally, the sample As the anchor point, the proposed contrast loss function can be defined as:

[0131]

[0132] in The function represents the similarity of the calculated vector. Here, cosine similarity is used to measure the similarity between two sample features. For sample right The predicted probability score of the class, is the temperature parameter used to adjust the sensitivity of contrastive learning and the spatial characteristics of each sample in the training batch. and time characteristics Apply formula (19) to calculate the loss respectively, and add the two losses as the contrastive learning loss of each sample , described by formula (20):

[0133]

[0134] By refining the spatial and temporal features of the skeleton, and optimizing the comparative learning of credible samples and fuzzy samples, the representation capability of the model skeleton action features is refined, and the model's classification performance for complex and similar actions is improved.

[0135] Step 4: Construct the training target loss function. The final training target combines the cross entropy loss and contrastive learning loss. The contrastive learning loss is calculated using formula (20). The cross entropy loss is It is used to supervise classification tasks and optimize the classification accuracy of the model, which is described by formula (21):

[0136]

[0137] in is the number of samples in the batch, is the number of action classification categories, It is a sample Corresponding category The label of It is a sample When the target class is ,otherwise , Samples predicted by the network belong The probability score of the class is combined with the cross entropy loss and the contrastive learning loss to form the complete objective function:

[0138]

[0139] Step 5: Model optimization and training. The entire model is trained using a stochastic gradient descent optimizer. The model parameters are adjusted to minimize the cross entropy loss and contrastive learning loss. During the training process, the learning rate and momentum coefficient are adaptively adjusted to ensure network convergence. The performance and accuracy of the model are verified on datasets such as NTU-RGB+D60 and NTU-RGB+D120. The experimental results are shown in Tables 1 and 2.

[0140] Table 1 Comparison of recognition accuracy of different models on NTU-RGB+D60 dataset

[0141]

[0142] Table 2 Comparison of recognition accuracy of different models on NTU-RGB+D120 dataset

[0143]

Claims

1. A skeleton action recognition method based on channel feature fusion, characterized in that: The following steps are involved: Step 1: Obtain the 3D skeleton data of the human body. Use a depth sensor to directly capture the 3D skeleton data of the person in the target scene or use a posture estimation algorithm to extract the 3D skeleton data of the human body from the RGB video. The spatial three-dimensional coordinates represent the spatial position of each skeleton node, and a three-dimensional skeleton sequence data of each sample is collected. , where T represents the number of skeleton graphs contained in the sequence, and V represents the number of nodes contained in each skeleton graph; Step 2: Construct a skeleton action recognition network based on channel feature fusion. The network contains 10 layers of spatiotemporal convolution modules, each of which contains a channel dynamic modeling graph convolution network, a multi-scale fusion time convolution network and a channel attention module; the channel dynamic modeling graph convolution network is used to extract the spatial structural features of the skeleton graph and describe the topological relationship between human joints, the multi-scale fusion time convolution network is used to capture the dynamic features of the joints in the action over time, and the channel attention module improves the network's ability to distinguish different actions by weighting the feature channels. It specifically includes the following sub-steps: (a) Channel dynamic modeling graph convolutional network is used to model the spatial relationship between joints in the skeleton graph and use dynamic channel topology to enhance feature representation. Each spatial modeling module contains three parallel channel topology optimization graph convolution modules (CTRGC for short). This module is a variant of the graph convolutional network (GCN) for skeleton data. It focuses on improving the feature modeling ability by improving the graph topology of each channel. Processing to obtain high-dimensional features , described by formula (1): ; in , Represents one of the three parallel modules, which outputs three features Add and use the ReLU activation function for nonlinear transformation, which is described by formula (2): ; in In order to maintain training stability and efficiency, the initial input data The convolution changes the channel dimension to match The size of the final skeleton space feature is obtained , described by formula (3): ; in It is a graph feature representation of skeleton data, which contains the geometric structure and spatial dependency between joints. (b) The multi-scale fusion temporal convolutional network is used to capture the changing characteristics of joints in action sequences over time. The network contains four branches. The first branch contains a Convolution and Temporal convolution; the second branch contains a Convolution and Temporal convolution; the third branch contains a convolution and a max pooling layer; the fourth branch has only one Convolution; the time convolution of the first two branches has different dilation rates, Input to four branches for processing respectively, The process includes: 1) The first temporal convolution branch unit, dilation rate Temporal convolution increases the size of the receptive field by inserting gaps between elements in the convolution kernel, so that the convolution can cover a longer time range, thereby more effectively capturing long-term dependencies in the time series. The convolution operation is performed on all channels along the time dimension, which is described by formula (4): ; in, Represents the first temporal convolution branch The time characteristic output, Expansion rate ; 2) The second temporal convolution branch unit, the expansion rate , expand the receptive field of the convolution, capture features of a longer time range, and the output features of the second branch , described by formula (5): ; 3) The third maximum pooling branch unit compresses the data volume by taking the maximum value of the local area while retaining the significant features. , described by formula (6): ; 4) The fourth 1×1 convolution unit first uses a convolution kernel size of The convolutional layer reduces the channel dimension to obtain , described by formula (7): ; The four branch features are concatenated and fused to obtain multi-scale fusion features : ; In order to maintain training stability and efficiency, the initial input data After convolution, the channel dimension is changed and added to , and obtain the final time feature , described by formula (9): ; in, Contains the temporal dynamic information of skeleton nodes; (c) The channel attention module is used to dynamically adjust the importance of the channel, highlight important channel features, and Global average pooling is performed on each channel of to compress the spatial dimension, which is described by formula (10): ; in, It is a channel The output feature vector of is then dynamically convolved to capture the local channel relationship, which is described by formula (11): ; in, Contains the correlation between channels, and then uses the sigmoid activation function Normalized, the weight range is , get the features is the weight of the channel, Expand the dimension to match The channel weights are applied to the input features to match the size of , and obtain the final feature representation , described by formula (12): ; in, represents the output features, Represents element-by-element multiplication, weighting the channels with channel weights to enhance useful features and suppress irrelevant features; After the feature extraction is completed through the network, the fully connected layer is used to map the features and perform average pooling operations. Finally, the prediction results are output through the softmax function to obtain the probability distribution of each skeleton action. Step 3: Construct a spatiotemporal feature refinement contrastive learning training framework, which includes a spatiotemporal feature refinement module and a confident sample guided contrastive learning module; the spatiotemporal feature refinement module is used to process the skeleton features extracted by the backbone network, and refine the skeleton features from the time and space dimensions respectively to obtain the time features and space features of the skeleton features; the confident sample guided contrastive learning module optimizes the feature distribution to make the sample closer to its category, which specifically includes the following sub-steps: (a) Spatiotemporal feature refinement module. When processing skeleton sequence data, the separation of spatiotemporal features helps to better capture the structure and motion information of the skeleton, and improves the skeleton features obtained by the backbone network. Processing from the time and space dimensions includes the following processes: 1) Input features Perform an average operation in the time dimension to eliminate the influence of time dynamic information and obtain the feature Spatial representation of , and then use Convolution transforms the feature dimension Compress to , in order to reduce the feature dimension, reduce the computational complexity, and adapt to the processing of high-dimensional skeleton data: ; in, ,Will Features are flattened into one-dimensional vectors ; 2) Input features Perform an average operation in the spatial dimension to remove the influence of the structural position and retain only the temporal dynamic features to obtain the feature Spatial representation of ,use Convolution transforms the feature dimension Compress to : ; in, ,Will Features are flattened into one-dimensional vectors ; (b) Confident sample guided contrastive learning module. The goal of this module is to refine the skeleton action feature representation through confident sample guided contrastive learning, optimize the temporal features and spatial features respectively, so that samples of the same type are clustered in the feature space and samples of different types are separated from each other. The core of the module is to use the clustering information of confident samples and fuzzy samples to improve the feature discrimination ability through the contrastive learning module, and use the contrast relationship between the basic action truth value and the fuzzy action to improve the sample representation, including the following processes: 1) Credible sample clustering. In each training process, the output of the model includes the feature vector of each sample and the corresponding category label. For an action category , if the sample Correctly classified into categories , then it is considered to be a credible sample. The feature prototype is the global representation of the category and represents the typical features of the category. Use the exponential moving average method to update the prototype of each category , ensuring that the updating process of feature prototypes is smooth and stable, It's action The size of the credible sample set is , the exponential moving average operation is described by formula (15): ; in, It's action The prototype, represents the number of training rounds, It is a sample The characteristic representation of Momentum coefficient, controls the update speed, usually set to 0.9, for smooth transition and avoid large prototype fluctuations. During training, the prototype becomes the action A stable estimate of the cluster center, which can learn the characteristics of new samples; 2) Fuzzy sample clustering: During the training process, the model will encounter some samples that are difficult to classify. These samples are easily confused with samples of other categories in the feature space, which are called fuzzy samples. These samples include two categories: false negative (False Negative, FN), that is, a sample actually belongs to a category , but is misclassified as other categories, a false positive (FalsePositive, FP) means that a sample does not actually belong to a category , but was misclassified as category , , For Action The false negative and false positive sample sets are of size , , using the average value of the sample set as the sample center, described by formula (16): ; in , Representative Action The center of the false negative samples and false positive samples is represented by In order to represent the fuzzy samples and distinguish the two fuzzy samples, a compensation term is introduced for the false negative samples, aiming to bring the feature vector of the false negative samples closer to the center of its correct category. For the false negative samples, the compensation term is defined as: ; By minimizing the compensation term , making the false negative samples closer to the credible samples in the feature space. When there are no false negative samples or the cosine distance converges to 1, Reaching the minimum value of 0 makes the model treat these blurred samples as actions Similarly, for false positive samples, the penalty term is defined as: ; By minimizing the penalty , so that the false positive samples are far away from the credible samples in the feature space. When there are no false positive samples or the cosine distance converges to -1, Reaching a minimum value of 0 prevents the model from identifying these ambiguous samples as actions ; Finally, the sample As the anchor point, the proposed contrast loss function can be defined as: ; in The function represents the similarity of the calculated vector. Here, cosine similarity is used to measure the similarity between two sample features. For sample right The predicted probability score of the class, is the temperature parameter used to adjust the sensitivity of contrastive learning and the spatial characteristics of each sample in the training batch. and time characteristics Apply formula (19) to calculate the loss respectively, and add the two losses as the contrastive learning loss of each sample , described by formula (20): ; By refining the spatial and temporal features of the skeleton, and optimizing the comparative learning of credible samples and fuzzy samples, the representation capability of the model skeleton action features is refined, and the model's classification performance for complex and similar actions is improved. Step 4: Construct the training target loss function. The final training target combines the cross entropy loss and contrastive learning loss. The contrastive learning loss is calculated using formula (20). The cross entropy loss is It is used to supervise classification tasks and optimize the classification accuracy of the model, which is described by formula (21): ; in is the number of samples in the batch, is the number of action classification categories, It is a sample Corresponding category The label of It is a sample When the target class is ,otherwise , Samples predicted by the network belong The probability score of the class is combined with the cross entropy loss and the contrastive learning loss to form the complete objective function: ; Step 5: Model optimization and training. Use the stochastic gradient descent optimizer to train the entire model. Adjust the model parameters to minimize the cross entropy loss and contrastive learning loss. During the training process, adaptively adjust the learning rate and momentum coefficient to ensure network convergence.

Citation Information

Cited By

  • Open set skeleton action recognition method and device based on outlier prototype learning

    CN121564806A

  • An open-set skeleton action recognition method and device based on outlier prototype learning

    CN121564806B