Class balance space-time federated feature learning-based skeleton action recognition method and related device
Through the class-balanced space-time federated feature learning method, the global feature extractor and debiased classifier are used to optimize the skeleton action recognition model, which solves the data privacy and classifier deviation problems, and achieves high accuracy and robust skeleton action recognition.
Patent Information
- Application Number
- CN202510615670.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-04
AI Technical Summary
The existing human activity recognition method based on RGB videos has data privacy leakage problems, and federated learning faces data heterogeneity and classifier deviation problems in skeleton action recognition, resulting in a decline in model performance.
The class-balanced space-time federated feature learning method is adopted, and the global feature extractor and debiased classifier model is combined with cross-entropy loss, comparison and alignment loss and prototype Gaussian sampling technology to generate space-time federated features, and the classifier is trained using knowledge matching strategies to optimize the action recognition model.
It improves the accuracy and robustness of skeleton video action recognition, enhances the generalization ability of the model, and solves the problems of data privacy protection and classifier deviation.
Smart Images

Figure CN120260133A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a method and related device for skeletal action recognition based on class-balanced spatio-temporal federated feature learning. Background Art
[0002] With the wide popularization of edge devices, the collection of video data has become increasingly convenient, which has greatly promoted the application of human activity recognition (HAR) technology in many fields such as intelligent security, video surveillance, and health management. In this process, the problem of data privacy has become increasingly prominent. Traditional HAR methods based on RGB (Red, Green, Blue) videos have serious defects because they are prone to leaking personal identity information such as appearance, clothing, and behavior habits when processing data.
[0003] To solve the privacy problem, researchers have explored from two directions: data format and training paradigm. In terms of data format, 3D skeleton data has received attention due to its advantages of unbiased representation and low storage requirements. It can effectively reduce the risk of exposing personal appearance and strongly promote the development of skeleton-based action recognition (SAR) methods. In terms of training paradigm, federated learning (FL), as an emerging decentralized training paradigm, has emerged. It can train a global model without sharing local datasets, which well protects user privacy and is particularly suitable for SAR because skeleton data itself does not involve personal appearance information. The combination of the two provides a new idea for video privacy protection.
[0004] However, applying image-based FL to SAR faces many challenges. First, there are significant differences in the representation forms between skeleton data and image data, which makes it extremely difficult to maintain spatio-temporal motion attributes related to diverse action distributions among different clients. Second, the data heterogeneity problem in FL further exacerbates these challenges. The data sources and distributions collected by different devices are diverse, resulting in differences in data distributions among devices and causing unstable training. In addition, most existing methods ignore the fairness of the classifier, that is, the classifier bias problem. In a non-independent and identically distributed (non-IID) scenario, the class-imbalanced distribution of clients will cause the gradient backpropagation to be unbalanced, making the classifier biased towards heterogeneous local data, resulting in problems such as low feature similarity and large differences in weight L2 norms among local classifiers, seriously affecting the model performance. Summary of the Invention
[0005] The purpose of this application is to provide a method and related device for skeletal action recognition based on class-balanced spatio-temporal federated feature learning, which can achieve accurate action recognition of skeleton videos and improve the generalization ability and robustness of the model at the same time.
[0006] To achieve the above purpose, the present application provides the following solutions: In a first aspect, the present application provides a skeletal action recognition method based on class-balanced spatio-temporal federated feature learning, including: Based on the client, obtain the global feature extractor model, the global classifier model, and the retrained debiased classifier model distributed by the server.
[0007] Based on the received global feature extractor model, global classifier model, and retrained debiased classifier model, perform the following steps on the input skeleton video: According to the skeleton video, generate a reference feature through the global feature extractor model, and use the reference feature as an anchor point.
[0008] Generate a contrast feature of the skeleton video based on the locally updated feature extractor; the locally updated feature extractor is the feature extractor updated by the client after receiving the global model distributed by the server; the global model includes the global feature extractor model, the global classifier model, and the retrained debiased classifier model.
[0009] According to the reference feature and the contrast feature, calculate the tightness of the same-class action features and the similarity of different-class features based on the cross-entropy loss and the contrast alignment loss, and obtain the feature extraction result.
[0010] Based on the feature extraction result, update the gradients of the local prototype, local soft label, and true motion feature on the client respectively.
[0011] Upload the updated feature extractor, local prototype, local soft label, and the gradients of the true motion feature to the server.
[0012] Through the server, respectively adopt model aggregation, prototype aggregation, soft label aggregation, and gradient aggregation for the received updated feature extractor, local prototype, local soft label, and the gradients of the true motion feature to generate a global model, a global prototype, a global soft label, and global gradients.
[0013] Through the server, based on the global model, global prototype, global soft label, and global gradients, adopt the prototype Gaussian sampling technique to generate spatio-temporal federated features.
[0014] Through the server, based on the spatio-temporal federated features, retrain the classifier jointly with the cross-entropy loss and the loss function calculated by the global soft label to generate an optimized action recognition model.
[0015] Through the server, based on the optimized action recognition model, perform action recognition on the skeleton video to obtain the action category label.
[0016] In a second aspect, the present application provides a skeletal action recognition device based on class-balanced spatio-temporal federated feature learning, including: A model acquisition module, configured to obtain, based on a client, a global feature extractor model, a global classifier model, and a re-trained debiased classifier model distributed by a server.
[0017] An execution module, configured to perform the following steps on an input skeleton video based on the received global feature extractor model, global classifier model, and re-trained debiased classifier model: A benchmark feature generation sub-module, configured to generate benchmark features through the global feature extractor model according to the skeleton video, and use the benchmark features as anchor points.
[0018] A contrast feature generation sub-module, configured to generate contrast features of the skeleton video based on a locally updated feature extractor; the locally updated feature extractor is a feature extractor updated by the client after receiving the global models distributed by the server; the global models include the global feature extractor model, the global classifier model, and the re-trained debiased classifier model.
[0019] A calculation sub-module, configured to calculate the tightness of homogeneous action features and the similarity of heterogeneous features based on the benchmark features and the contrast features, based on cross-entropy loss and contrast alignment loss, to obtain a feature extraction result.
[0020] An update sub-module, configured to update the gradients of local prototypes, local soft labels, and true motion features at the client based on the feature extraction result.
[0021] An upload sub-module, configured to upload the updated feature extractor, local prototypes, local soft labels, and gradients of true motion features to the server.
[0022] An aggregation sub-module, configured to, through the server, respectively perform model aggregation, prototype aggregation, soft label aggregation, and gradient aggregation on the received updated feature extractor, local prototypes, local soft labels, and gradients of true motion features to generate a global model, global prototypes, global soft labels, and global gradients.
[0023] A spatio-temporal federated feature generation sub-module, configured to, through the server, generate spatio-temporal federated features based on the global model, global prototypes, global soft labels, and global gradients, using prototype Gaussian sampling technology.
[0024] An optimization sub-module, configured to, through the server, re-train a classifier based on the spatio-temporal federated features, jointly with a loss function calculated by cross-entropy loss and global soft labels, to generate an optimized action recognition model.
[0025] An identification sub-module, configured to, through the server, perform action recognition on the skeleton video based on the optimized action recognition model to obtain an action category label.
[0026] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement a skeleton action recognition method based on class-balanced spatio-temporal federated feature learning described in any one of the above.
[0027] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements a skeleton action recognition method based on class-balanced spatio-temporal federated feature learning described in any one of the above.
[0028] According to the specific embodiments provided by the present application, the following technical effects are disclosed: The present application provides a skeleton action recognition method and related devices based on class-balanced spatio-temporal federated feature learning. In this method, first, the global feature extractor model and the global classifier model are distributed by the server, ensuring that each client has a common knowledge base at the initial stage of training. After receiving these models, the client will update according to local data, generate contrast features, and compare them with the global reference features, which helps to capture the subtle differences in actions and improve the recognition accuracy. By calculating the compactness of the same-class action features and the similarity of the different-class features, the dual loss function acts on the feature extraction process together, making the features more discriminative and reducing the possibility of misclassification. The client uploads the gradients of the updated feature extractor, local prototype, local soft label, and real motion features to the server. The server generates a global model, global prototype, global soft label, and global gradient by aggregating these updates, which helps to eliminate individual data biases and improve the generalization ability of the model. Based on the global model, global prototype, global soft label, and global gradient, the prototype Gaussian sampling technique is used to generate spatio-temporal federated features. These features not only contain global information but also integrate the local information of each client, making the model more robust when recognizing actions. Based on the spatio-temporal federated features, the classifier is retrained with the combined cross-entropy loss and the loss function calculated based on the global soft label to generate an optimized action recognition model. This process helps to eliminate the class imbalance bias and improve the accuracy of the model in recognizing various actions. Finally, based on the optimized action recognition model, the action recognition of the skeleton video is performed to obtain the action category label. Since the model fully considers the global and local information and the constraints of various loss functions during the training process, it can accurately recognize actions and improve the recognition accuracy and robustness. Description of the Drawings
[0029] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0030] Figure 1 It is a schematic flowchart of a skeleton action recognition method based on class-balanced spatio-temporal federated feature learning provided by an embodiment of the present application.
[0031] Figure 2 It is a schematic diagram of the specific recognition process of the skeleton action recognition method based on class-balanced spatio-temporal federated feature learning provided by an embodiment of the present application.
[0032] Figure 3 It is a schematic diagram of the functional modules of a skeleton action recognition device based on class-balanced spatio-temporal federated feature learning provided by an embodiment of the present application.
[0033] Figure 4 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0035] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0036] In an exemplary embodiment, as Figure 1 shown, a skeleton action recognition method based on class-balanced spatio-temporal federated feature learning is provided, including: Step 101, based on the client, obtain the global feature extractor model, the global classifier model, and the re-trained debiased classifier model distributed by the server.
[0037] Step 102, based on the received global feature extractor model, global classifier model, and re-trained debiased classifier model, perform the following steps on the input skeleton video: Step 111, according to the skeleton video, generate a reference feature through the global feature extractor model and use the reference feature as an anchor point.
[0038] Step 112: Generate contrast features of the skeleton video based on the locally updated feature extractor; the locally updated feature extractor is the feature extractor updated by the client after receiving the global model distributed by the server; the global model includes a global feature extractor model, a global classifier model, and a re-trained de-biased classifier model.
[0039] Step 113: Based on the benchmark features and the contrast features, calculate the tightness of the same-class action features and the similarity of different-class features based on the cross-entropy loss and the contrast alignment loss, to obtain the feature extraction result.
[0040] Step 114: Based on the feature extraction result, update the gradients of the local prototype, local soft label, and true motion features on the client side respectively.
[0041] Step 115: Upload the updated feature extractor, local prototype, local soft label, and the gradients of the true motion features to the server.
[0042] Step 116: Through the server, respectively adopt model aggregation, prototype aggregation, soft label aggregation, and gradient aggregation for the received updated feature extractor, local prototype, local soft label, and true motion features to generate a global model, a global prototype, a global soft label, and global gradients.
[0043] Step 117: Based on the global model, global prototype, global soft label, and global gradients, adopt the prototype Gaussian sampling technique to generate spatio-temporal federated features.
[0044] Step 118: Based on the spatio-temporal federated features, re-train the classifier by combining the cross-entropy loss and the loss function calculated by the global soft label to generate an optimized action recognition model.
[0045] Step 119: Based on the optimized action recognition model, perform action recognition on the skeleton video to obtain the action category label.
[0046] Among them, when performing steps 101 - 102, specifically, it can be as follows: During the th communication round, each client will receive three models from the server: the global feature extractor model , the global classifier model , and the re-trained de-biased classifier model .
[0047] Among them, in some embodiments, when performing steps 111 - 119, as Figure 2 shown, specifically, it can be as follows: In the first stage, to ensure effective local task performance, cross-entropy loss is adopted as the main classification loss function. Specifically, for the th action sample belonging to class , the client generates the corresponding ground truth motion feature .
[0048] where is the unique encoded ground truth label of the th sample of class on the client , and represents the normalized predicted probability of class . Specifically, represents the ground truth label, while represents the predicted probability of the model.
[0049] When the client receives the global model from the server, it locally updates the model to . For each input skeleton action , the global feature extractor model generates the feature representation , while the locally updated model generates the feature representation . Taking as the anchor point, the following loss function is defined to enhance the alignment between and while reducing the similarity between and the features of non-homogeneous samples: .
[0050] where is the batch size, is the temperature hyperparameter, is the measure of similarity between features, , is the indicator function, is the similarity loss, i is the i th batch, h i is the feature representation extracted by the global feature extractor for the i th batch, is the feature representation generated by the local feature extractor for the th batch. The client Subsequently, local training iterations are performed to optimize the following objective function: .
[0051] In the formula, is a balanced cross-entropy loss term and the model calibration loss term hyperparameters.
[0052] In the second stage, after updating the local model , the client generates gradients of the local prototype, local soft labels, and true motion features. These outputs enrich the semantic information of the spatio-temporal federated features, enhance the global understanding ability of the debiased classifier, and contribute to updating the spatio-temporal federated features.
[0053] For a client with categories , the prototype of the -th class is defined as . For each action sample belonging to the class, this embodiment uses the first blocks of the feature extraction to generate an intermediate feature representation of . The prototype for each class is calculated as follows: .
[0054] In the formula, is the number of categories of the client , is the action sample and its label.
[0055] The spatio-temporal characteristics of the prototype are retained, i.e., , where and represent the feature dimension and the number of time frames extracted by the first blocks, respectively. Here, represents the local dataset of the -th class on the client , is the total number of samples in the -th class. The aggregated prototype set is then transmitted to the central server for global prototype aggregation.
[0056] For the categories on the client , the average logarithm and the corresponding soft labels are calculated based on its local dataset : .
[0057] Among them represents the category calculated by the client of the average logits. Then apply the softening function to obtain the soft label: . .
[0058] Here is implemented using the last softmax layer of the client model, where represents the temperature parameter. The generated is then uploaded to the server for global aggregation of the soft label.
[0059] The server will broadcast the debiased classifier model to all clients. Each client calculates the gradient of the true action class according to a resampled subset of its local dataset . Using the feature extractor model and the debiased classifier model , the true sample features and their gradients of the category are calculated as follows: .
[0060] In the formula, represents the number of sampling instances of the category in , represents the aggregated gradient of the category on the client . Only the aggregated gradients of the true motion features will be uploaded to the server, which ensures irreversibility and reduces the risk of gradient-based privacy attacks. Then the server uses these gradients to update the spatio-temporal federated features.
[0061] In the third stage, after completing the local update process, each client will , , and upload to the server. The server aggregates this information to train the global debiased classifier.
[0062] Use weighted average to aggregate the client model to obtain the global model , the formula is as follows: .
[0063] The server receives the local prototypes from the clients and aggregates the prototypes for each class. Specifically, for a class , it uses the prototypes of all the clients that contain this class, that is, for of . In the formula, represents the number of clients that contain class . The global prototype of class is calculated as follows: .
[0064] In the formula, represents the total number of samples of class among all the clients, while represents the aggregated global prototype of all classes. These global prototypes are then used to optimize the spatio-temporal federated features, is the gradient of the true motion feature, is the th client, and is the total number of samples of class
[0065] Similar to the prototype aggregation described above, the local class soft labels uploaded by the clients are aggregated to generate the global soft labels. For a given class , the global soft label is calculated as follows: .
[0066] In the formula, represents the aggregated global soft label of all classes. During the classifier re-training, supplements the global knowledge of the debiased classifier, enhances its generalization ability and reduces the bias, is the result of softening the average logits value of the th class calculated on client , is the total number of clients, is the th client.
[0067] The server aggregates the local gradients uploaded by the clients to optimize the spatio-temporal federated features. Before that, the aggregation process is similar to the aggregation process used for prototypes. Specifically, for each class , the global gradient is calculated as follows: .
[0068] In the formula, represents the aggregated global gradient of all classes, The number of clients containing the category . These aggregated gradients guide the optimization of the spatio-temporal federated features, ensuring that the classifier can accurately perceive the true category distribution and effectively perform the classification task.
[0069] In the fourth stage, using the aggregated prototypes , for each category 's prototype generate trainable features for it, denoted as , where . Incorporate Gaussian sampling noise into these features: .
[0070] In the formula, is a mixing parameter that ensures a balance between similarity to the prototype and diversity of federated features. represents the Gaussian distribution. The generated features will be further processed by a neural network parameterized by to improve the training performance. For all categories belonging to the set , the neural network will convert into spatio-temporal federated features of the same shape as follows: .
[0071] Specifically, the model consists of two fully connected (FC) layers with a ReLU activation function between the two fully connected layers.
[0072] To ensure that the debiased classifier can effectively distinguish all categories, a gradient matching loss is used to align the gradients calculated based on spatio-temporal federated features with the gradients calculated based on true features. During the th communication round, the top layer (the module after the th module) shared by the server encodes to generate the final federated feature representation . Using these representations, the spatio-temporal federated feature gradients of category can be calculated by the debiased classifier as follows: .
[0073] The server has aggregated the true feature gradients of each category uploaded by the clients , it optimizes the spatio-temporal federated features by minimizing the gradient matching loss between and . This ensures the consistency between the federated features and the ground-truth features: .
[0074] In the formula, represents the -th row of the -th class gradient of the server.
[0075] Traditional class-aware gradients only consider the direction and magnitude of the gradients, ignoring the actual loss values. This limitation may lead to the lack of key motion characteristics necessary for each class in the federated features. To address this issue, a motion-aware discrepancy loss (MA) is proposed to enhance the motion consistency between the spatio-temporal federated features and their corresponding prototypes. It is achieved by incorporating derivatives from the zero-th order (joint positions) to the first order (joint velocities): .
[0076] In the formula, and . represents the total number of frames, represents the skeleton length vector per frame, and are the -th and -th order derivatives calculated by the forward difference method for is the encoded federated feature representation, is the prototype of the -th class.
[0077] The server uses the following objective function to optimize the spatio-temporal federated features within a preset number of iterations: .
[0078] In the formula, the hyperparameter controls the degree of motion recovery.
[0079] After sufficient optimization iterations, the spatio-temporal federated features approximate the ground-truth motion features and can effectively capture the discriminative features in their respective classes. Then, these optimized features are used to retrain a debiased classifier.
[0080] Specifically, during the -th communication round, the aggregated global model The classifier in may perform poorly on clients with highly imbalanced class distributions. To alleviate this problem, while retaining the ability of the global model to extract global motion features , freeze and retrain a debiased classifier . This retrained classifier aims to achieve performance comparable to that of a classifier trained on the entire dataset without direct access to client data.
[0081] In the fifth stage, randomly initialize a classifier , and freeze the previously trained spatio-temporal federated features . Utilize the top layer shared by the server (the layer after the th module) to encode into a federated feature representation for classification . Then, use the cross-entropy loss to train the debiased classifier: .
[0082] In the formula, represents the learning rate for gradient update.
[0083] Since information loss inevitably occurs during the construction of spatio-temporal federated features, relying solely on cross-entropy loss to train the debiased classifier will limit its convergence performance and weaken 's generalization ability. To address this problem, a knowledge matching strategy is proposed, which uses global soft labels to enhance 's understanding of the overall class distribution and improve its generalization ability.
[0084] Specifically, the server uses to calculate the global class soft labels : .
[0085] To further enhance 's ability to capture the global class distribution, a Jensen-Shannon (JS) divergence-based loss is introduced, which penalizes the difference between output distributions: .
[0086] In the formula, . The overall training objective can be expressed as: .
[0087] In the formula, is a hyperparameter that controls the regularization strength.
[0088] After completing all the updates, the server will distribute the global feature extractor , the global classifier, and the retrained debiased classifier to each participating client for use in the next round of training.
[0089] This application conducts tests on the MBFSAR method in the image domain and other methods under heterogeneous labels. The specific results are shown in Table 1 as follows: Table 1 Results of the MBFSAR method in the image domain and other methods under heterogeneous labels
[0090] This application conducts tests on the MBFSAR method and other methods in the action domain under a natural heterogeneous environment. The specific results are shown in Table 2 as follows: Table 2 Results of the MBFSAR method and other methods in the action domain under a natural heterogeneous environment
[0091] This application also conducts tests on the MBFSAR method and other methods in the action domain under a label heterogeneous environment. The specific results are shown in Table 3 as follows: Table 3 Results of the MBFSAR method and other methods in the action domain under a label heterogeneous environment
[0092] Specifically, Table 1 is a comparison of the results of this application with other methods on two datasets, CIFAR-10 and CIFAR-100, in the image domain, using the TOP-1 accuracy as the evaluation metric. Table 2 is a comparison of the results of this application with other methods on two different versions of two datasets, NTU RGB+D and PKU-MMD, in the natural heterogeneous environment for skeleton action recognition. Table 3 is a comparison of the results of this application with other methods on two different versions of two datasets, NTU RGB+D and PKU-MMD, in the label heterogeneous environment for skeleton action recognition. † represents the results of re-implementation, using the TOP-1 accuracy as the evaluation metric. Bold indicates the best, and underlined indicates the second best.
[0093] Based on the same inventive concept, the embodiments of this application also provide a device for skeleton action recognition for implementing the above-mentioned class-balanced spatio-temporal federated feature learning. The implementation solutions provided by this device to solve problems are similar to those described in the above method. Therefore, the specific limitations in one or more embodiments of the following skeleton action recognition devices can refer to the limitations on the method for skeleton action recognition based on class-balanced spatio-temporal federated feature learning in the above text, and will not be elaborated here.
[0094] In an exemplary embodiment, as Figure 3 shown, a skeleton action recognition device based on class-balanced spatio-temporal federated feature learning is provided, including: A model acquisition module 301, which is used to obtain, based on the client, a global feature extractor model, a global classifier model, and a re-trained debiased classifier model distributed by the server.
[0095] An execution module 302, which is used to perform the following steps on the input skeleton video based on the received global feature extractor model, global classifier model, and re-trained debiased classifier model: A reference feature generation sub-module 311, which is used to generate a reference feature from the skeleton video through the global feature extractor model and use the reference feature as an anchor point.
[0096] A contrast feature generation sub-module 312, which is used to generate a contrast feature of the skeleton video based on a locally updated feature extractor; the locally updated feature extractor is a feature extractor updated by the client after receiving the global model distributed by the server; the global model includes a global feature extractor model, a global classifier model, and a re-trained debiased classifier model.
[0097] A calculation sub-module 313, which is used to calculate the tightness of the same-class action features and the similarity of different-class features based on the cross-entropy loss and the contrast alignment loss according to the reference feature and the contrast feature, and obtain a feature extraction result.
[0098] An update sub-module 314, which is used to update the gradients of the local prototype, local soft label, and true motion feature at the client respectively based on the feature extraction result.
[0099] An upload sub-module 315, which is used to upload the updated feature extractor, local prototype, local soft label, and the gradients of the true motion feature to the server.
[0100] An aggregation sub-module 316, which is used to generate a global model, a global prototype, a global soft label, and a global gradient by respectively performing model aggregation, prototype aggregation, soft label aggregation, and gradient aggregation on the received updated feature extractor, local prototype, local soft label, and the gradients of the true motion feature through the server.
[0101] A spatio-temporal federated feature generation sub-module 317, which is used to generate spatio-temporal federated features by using the prototype Gaussian sampling technique based on the global model, global prototype, global soft label, and global gradient.
[0102] An optimization sub-module 318, which is used to re-train a classifier based on the spatio-temporal federated features and the loss function calculated jointly by the cross-entropy loss and the global soft label, and generate an optimized action recognition model.
[0103] An identification sub-module 319, configured to perform action recognition on the skeleton video based on the optimized action recognition model to obtain action category labels.
[0104] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store processed data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a skeleton action recognition method based on class-balanced spatio-temporal federated feature learning.
[0105] Those skilled in the art can understand that Figure 4 the structure shown in
[0106] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0107] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0108] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0109] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0110] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0111] In summary, the present application has the following technical effects: This application addresses numerous problems faced by federated learning in the application of skeleton-based action recognition, such as data privacy leakage, model performance degradation caused by data heterogeneity, and classifier bias. Through the proposed class-balanced spatio-temporal federated feature learning framework, the training process is optimized using model calibration loss on the client side to reduce the impact of client data heterogeneity on the classifier; on the server side, prototype Gaussian sampling technology is adopted to generate spatio-temporal federated features rich in semantic information, and it is optimized by combining class-aware gradient matching loss and motion-aware difference loss. At the same time, a knowledge matching strategy is used to train a debiased classifier. This method effectively alleviates the above problems and outperforms the prior art in experiments under natural and label heterogeneity scenarios, significantly improving the accuracy and robustness of skeleton-based action recognition in the federated learning environment.
[0112] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0113] Specific examples are used in this article to elaborate on the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, based on the idea of this application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A skeletal action recognition method based on class-balanced spatio-temporal federated feature learning, characterized in that, Including: Based on the client, obtain the global feature extractor model, global classifier model, and retrained debiased classifier model distributed by the server; Based on the received global feature extractor model, global classifier model, and retrained debiased classifier model, perform the following steps on the input skeleton video: According to the skeleton video, generate a reference feature through the global feature extractor model, and use the reference feature as an anchor point; Generate a contrast feature of the skeleton video based on the locally updated feature extractor; the locally updated feature extractor is the feature extractor updated by the client after receiving the global model distributed by the server; the global model includes the global feature extractor model, global classifier model, and retrained debiased classifier model; According to the reference feature and the contrast feature, calculate the tightness of the same-class action features and the similarity of the different-class features based on the cross-entropy loss and contrast alignment loss to obtain the feature extraction result; Based on the feature extraction result, update the gradients of the local prototype, local soft label, and true motion feature on the client respectively; Upload the updated feature extractor, local prototype, local soft label, and gradient of the true motion feature to the server; Through the server, respectively use model aggregation, prototype aggregation, soft label aggregation, and gradient aggregation for the received updated feature extractor, local prototype, local soft label, and true motion feature to generate a global model, global prototype, global soft label, and global gradient; Through the server, based on the global model, global prototype, global soft label, and global gradient, use the prototype Gaussian sampling technique to generate spatio-temporal federated features; Through the server, based on the spatio-temporal federated features, retrain the classifier by combining the cross-entropy loss and the loss function calculated by the global soft label to generate an optimized action recognition model; Through the server, based on the optimized action recognition model, perform action recognition on the skeleton video to obtain an action category label.
2. The bone action recognition method based on class-balanced spatio-temporal federated feature learning according to claim 1, wherein, The formula expression of the similarity loss for calculating the tightness of the same-class action features and the different-class features is: ; Among them, is the batch size, is the temperature hyperparameter, measures the similarity between features, , is the indicator function, is the similarity loss, i is the i th batch, h i is the feature representation extracted by the global feature extractor for the i th batch, is the feature representation generated by the local feature extractor for the th batch.
3. The bone action recognition method based on class-balanced spatio-temporal federated feature learning according to claim 2, wherein Through the server, respectively use model aggregation, prototype aggregation, soft label aggregation, and gradient aggregation for the received updated feature extractor, local prototype, local soft label, and true motion feature to generate a global model, global prototype, global soft label, and global gradient, specifically including: Generate a global model according to the formula Generate a global prototype according to the formula Generate global soft labels according to the formula According to the formula Generate the global gradient; Among them, is the number of clients containing the category , is the total number of samples of the category in all clients, is the aggregated global prototype of all categories, is the aggregated global soft label of all categories, is the aggregated global gradient of all categories, is the th client, is the nd communication round, is the total number of clients, is the th total number of samples on the th client, is the local prototype, is the local soft label, is the gradient of the true motion feature, is the th total number of samples of the category 4. A method for skeletal action recognition based on class-balanced spatio-temporal federated feature learning according to claim 3, characterized in that, The calculation methods of the gradients of the local prototype, local soft label, and true motion feature specifically include: Calculate the local prototype according to the formula Calculate the local soft label according to the formula According to the formula Calculate the gradient of the true motion feature; where, is the number of categories of the client , is the action sample and its label, is the feature extractor, is the classification loss, is the debiased classifier model, is the client 's resampled subset, represents the sample set of the c-th category on the client , is the client 's first blocks of the feature extractor, is the client 's aggregated prototype set, is to perform a softening operation, is the temperature parameter, is the client 's calculated average logits value of the -th category, is the client 's soft label set, is the client 's number of sampling instances of the local dataset for the -th category, is the client 's aggregated gradient of the -th category, is the result of softening the calculated average logits value of the -th category on the client is the prototype of the -th category of the client 5. The bone action recognition method based on class-balanced spatio-temporal federated feature learning according to claim 4, characterized in that, The calculation formula of the classification loss is: ; Among them, is the client on the th unique encoded true label of the class of the sample, indicating the normalized predicted probability of the class.
6. The bone action recognition method based on class-balanced spatio-temporal federated feature learning according to claim 1, wherein After generating the spatio-temporal federated features based on the global model, global prototype, global soft label, and global gradient using the prototype Gaussian sampling technique, it further includes: According to the objective function , optimize the spatio-temporal federated features within a preset number of iterations to obtain the optimized spatio-temporal federated features; Among them, To control the degree of motion recovery; , Is the motion consistency; , Is the consistency between the federated feature and the true feature; Is the total number of frames; ; ; Is the skeleton length vector of each frame, And Are the nth-order derivatives calculated by the forward difference method for And respectively, Represents the th row of the gradient of the nd class aggregated by the server, Is the encoded federated feature representation, Is the prototype of the th class.
7. A skeletal action recognition device based on class-balanced spatio-temporal federated feature learning, characterized in that, Including: A model acquisition module, used to obtain the global feature extractor model, global classifier model, and retrained debiased classifier model distributed by the server based on the client; An execution module, used to perform the following steps on the input skeleton video based on the received global feature extractor model, global classifier model, and retrained debiased classifier model: A reference feature generation sub-module, configured to generate a reference feature through a global feature extractor model according to the skeleton video, and use the reference feature as an anchor point; A comparison feature generation sub-module, configured to generate a comparison feature of the skeleton video based on a locally updated feature extractor; the locally updated feature extractor is a feature extractor updated by the client after receiving the global model distributed by the server; the global model includes a global feature extractor model, a global classifier model, and a re-trained debiased classifier model; A calculation sub-module, configured to calculate the tightness of the same-class action features and the similarity of the different-class features based on the cross-entropy loss and the contrast alignment loss according to the reference feature and the comparison feature, to obtain a feature extraction result; An update sub-module, configured to update the gradients of the local prototype, the local soft label, and the true motion feature at the client respectively based on the feature extraction result; An upload sub-module, configured to upload the updated feature extractor, the local prototype, the local soft label, and the gradients of the true motion feature to the server; An aggregation sub-module, configured to, through the server, respectively perform model aggregation, prototype aggregation, soft label aggregation, and gradient aggregation on the received updated feature extractor, local prototype, local soft label, and true motion feature to generate a global model, a global prototype, a global soft label, and global gradients; A spatio-temporal federated feature generation sub-module, configured to, through the server, generate spatio-temporal federated features by using the prototype Gaussian sampling technique based on the global model, the global prototype, the global soft label, and the global gradients; An optimization sub-module, configured to, through the server, re-train a classifier based on the spatio-temporal federated features and a loss function calculated by the global soft label to generate an optimized action recognition model; An identification sub-module, configured to, through the server, perform action recognition on the skeleton video based on the optimized action recognition model to obtain an action category label.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement a skeleton action recognition method according to any one of claims 1-6, which is based on class-balanced spatio-temporal federated feature learning.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements a skeleton action recognition method according to any one of claims 1-6, which is based on class-balanced spatio-temporal federated feature learning.
Citation Information
Cited By
Personalized global prototype federal learning method and system based on adaptive feature alignment
CN120509507A