A Small-Sample Class-Incremental Gait Recognition Method in Multiple Activity Scenarios
By introducing a hybrid prototype enhancement model framework and selective prototype enhancement module in gait recognition, the catastrophic forgetting and overfitting problems in incremental gait recognition in small sample types in multiple activity scenarios are solved, and accurate identification of new users and new activities is achieved.
Patent Information
- Application Number
- CN202411502428.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-25
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-10-25
AI Technical Summary
The prior art has catastrophic forgetting and overfitting problems in the identification of incremental gaits in small sample types in multi-activity scenarios, making it difficult to maintain the recognition performance of old users and old activities when facing new users and new gait activities.
Using a model framework based on hybrid prototype enhancement, hybrid prototypes are generated by introducing auxiliary activity tags, and the prototype is adjusted using selective prototype enhancement modules to improve its generalization and discrimination capabilities.
It effectively solves the problem of catastrophic forgetting, improves the accuracy and stability of gait recognition in multiple activity scenarios, especially in small samples, which can maintain the recognition performance of new users and new activities.
Smart Images

Figure CN119475068B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of gait recognition, and particularly relates to a small-sample class incremental gait recognition method in multiple activity scenarios. The proposed method is implemented based on a model framework enhanced by hybrid prototypes. Background Art
[0002] Gait recognition is a technology for identifying individuals by analyzing the motion patterns presented during human walking. Many studies have shown that everyone's gait is unique, so the purpose of identity authentication can be achieved through gait data. In recent years, with the rapid development of technologies related to personal devices and wearable devices, gait recognition technology based on intelligent devices has received extensive attention. It usually embeds inertial sensors such as accelerometers and gyroscopes into portable intelligent devices. Due to its portability and low resource consumption, it can conveniently obtain motion data in daily life and be used for continuous user identification or verification. In recent years, deep learning technology has promoted the rapid development of the gait recognition field. Deep learning models can effectively extract and learn gait features, thereby realizing the identification of individuals. However, compared with humans, conventional deep neural networks cannot accumulate experience and retain knowledge, so they are not applicable when dealing with new user and new walking activity data.
[0003] Natural learning systems are progressive. They can continuously absorb new knowledge and, while maintaining existing knowledge, continuously improve and expand their cognitive abilities. In real-world scenarios, gait recognition requires continuous learning ability, that is, the gait recognition system can continuously expand to adapt to new users while maintaining the recognition performance for old users that have been learned. However, traditional deep learning models are usually trained through offline batch learning, which means that the entire model needs to be retrained using all existing samples. This results in high requirements for computing resources and time when learning new gait samples. When facing new user gait data encountered in the continuous gait data stream, conventional methods cannot achieve "continuous" learning by retraining the model on all new and old user data. In this case, simply retraining the model when facing new user gait data in the continuous gait data stream may lead to catastrophic forgetting of existing knowledge. Some research works have explored incremental learning of gait to meet the needs of new user registration. In addition, since collecting and annotating gait data is an expensive and time-consuming task, gait recognition in the scenario of a small number of samples has become the focus of research.
[0004] In addition, in real scenarios, in addition to walking on flat roads, people also include other activities such as walking on slopes, going upstairs, and going downstairs. Although for the same person, the same gait characteristics may be presented in different walking activities, different walking activities more significantly present differences in gait characteristics, which means that a model trained using user gait data in a single activity scenario cannot be generalized to other activity scenarios.
[0005] To be able to process the continuous information flow in the real world, methods based on incremental learning have been proposed, and these methods have the ability to maintain the already learned old knowledge while learning new knowledge. At the same time, incremental learning methods have also been applied to fields such as gait recognition. Although the gait recognition methods using incremental learning can well handle the problem of continuous data streams, most studies still need to label a large number of samples to retrain the model when identifying new data or new categories, which is obviously infeasible in practical applications, especially for personnel recognition on terminal devices with limited computing and storage resources.
[0006] In the process of gait recognition, few-shot incremental learning is based on incremental learning and considers the situation where the number of samples for a new task is very small, that is, few-shot learning. It is more in line with the real-world scenario but also more difficult. Since new categories are introduced, the model can only use the samples of the new categories to update the parameters, which causes the network to easily forget the category information that has been learned during the update process, resulting in the problem of catastrophic forgetting. In addition, the few-shot training strategy is used in the incremental stage, and the number of samples for each category is small. Although the network can well classify these small numbers of new category samples during the training process, on the test set, the classification performance for these categories will significantly decline, indicating that the model has an overfitting problem. Catastrophic forgetting and overfitting are two important problems restricting the development of few-shot class incremental gait recognition in multi-activity scenarios. Summary of the Invention
[0007] Aiming at the above deficiencies in the prior art, the few-shot class incremental gait recognition method provided by the present invention solves the problem that it is difficult for the gait recognition model in a single scenario to accurately recognize users with few-shot gait data volume in multi-activity scenarios during the existing gait recognition process.
[0008] To achieve the above invention purpose, the technical solution adopted by the present invention is: a few-shot class incremental gait recognition method in multi-activity scenarios, including the following steps:
[0009] S1. Collect gait samples in multi-activity scenarios and construct a data set;
[0010] The labels of the gait samples in the data set include user labels and activity labels;
[0011] S2. Construct a model framework based on hybrid prototype enhancement, including a feature extractor and a selective prototype enhancement module connected in sequence;
[0012] S3. Pre-train the feature extractor of the model framework using the data set to obtain a pre-trained model;
[0013] S4. Based on the pre-trained model, perform incremental training on the model framework, including:
[0014] Use the pre-trained model to extract the features of each support sample in the dataset, and form the initial prototype of each class according to the category; the initial prototype is a mixed prototype containing user labels and activity labels;
[0015] Input the initial prototype into the selected prototype enhancement module for incremental training to obtain the enhanced prototype, and then obtain the gait recognition model;
[0016] S5. For the query sample of the gait to be recognized input into the gait recognition model, obtain the corresponding initial prototype through the feature extractor with frozen parameters, merge it with all the prototypes in the previous session and perform prototype enhancement, and then identify the user category of the query sample by calculating the distance between the query sample and the enhanced prototype, completing gait recognition.
[0017] Further, in step S4, the method of forming the initial prototype of each class according to the category is specifically:
[0018] S401. Combine the activity labels of m different categories and the user labels of n different categories in the dataset to form a combined class Y with m×n categories, where Y = {y i} and i ∈ [0, m*n);
[0019] S402. Obtain the initial prototype of each class by calculating the embedding mean of each support sample in the set of all support samples in the dataset.
[0020] Further, in step S4, the representation of the initial prototype is:
[0021]
[0022] In the formula, P b represents the initial prototype of the combined class b, S b represents the set of support samples belonging to category b, |S b | represents the number of support samples, f b (·) is the pre-trained feature extractor, x i represents the support sample, and y i represents the category label.
[0023] Further, in step S4, in each incremental session, the method of performing incremental training based on the mixed prototype is specifically:
[0024] S411. Generate the mixed prototype of the new category support samples through the feature extractor, merge it into the prototype container of the previous session, and form the support prototype of the current incremental session with all the mixed prototypes in the current prototype container;
[0025] S412. Extract the hybrid prototype of the query sample through the feature extractor. Input it into the selective prototype enhancement module for evaluation.
[0026] S413. In the selective prototype enhancement module, process the hybrid prototype through a 1×1 convolutional kernel with a learning parameter q to obtain a new feature map.
[0027] Respectively, through 1×1 convolutional kernels with learning parameters k and v, and using the prior knowledge of the existing labels, merge the hybrid prototypes with the same user or activity label to obtain a new feature map. and
[0028] S414. Construct the maximum attention weight matrix A according to the features in the feature maps and .
[0029] S415. Multiply the maximum attention weight matrix A by the said feature map and add the hybrid prototype as a residual term to obtain the enhanced prototype.
[0030] Furthermore, in the step S413, the feature maps and are expressed as:
[0031]
[0032] where s q (·), s k (·), s v (·) represent linear transformation functions with q, k, and v as parameters, represents the new prototype obtained by taking the mean of the prototypes with label y in the hybrid prototype , y represents a user label or an activity label, and the finally obtained prototype represents one of the user prototypes or activity prototypes, and T represents the transpose operation.
[0033] Furthermore, in the step S414, the weight element α of the maximum attention weight matrix A ij is expressed as:
[0034]
[0035] where and respectively represent and the features at the i-th and j-th positions in denotes obtaining the original attention matrix the index of the maximum attention score on the first dimension, is the scaling factor, and T represents the transpose operation.
[0036] Furthermore, the selection prototype enhancement module includes an activity enhancement module SPEM-SUDA and a user enhancement module SPEM-SADU;
[0037] The enhanced prototype is the prototype obtained by adjusting through the activity enhancement module SPEM-SUDA and the user module SPEM-SADU respectively and performing feature fusion, which is expressed as:
[0038]
[0039] In the formula, and respectively represent the prototype features adjusted through the activity enhancement module SPEM-SUDA and the user enhancement module SPEM-SADU.
[0040] Furthermore, in step S4, the loss function of the selection prototype enhancement module is expressed as:
[0041]
[0042] In the formula, and respectively represent the loss functions of the activity enhancement module SPEM-SUDA and the user enhancement module SPEM-SADU, and respectively represent the decomposed support prototype and query prototype, cos(·,·) represents the cosine classifier, represents the cross-entropy function, and R represents the batch data.
[0043] The beneficial effects of the present invention are:
[0044] Aiming at the problems existing in small-sample class incremental gait recognition in multi-activity scenarios, this paper proposes a gait recognition framework based on hybrid prototype enhancement. First, the framework generates hybrid prototypes by introducing auxiliary activity labels and adjusts the prototypes by enhancing user information and activity information through a selective prototype enhancement module. This prototype has stronger generalization ability than ordinary prototypes, thus improving the representation ability and discrimination ability of the prototypes. Second, the selective prototype enhancement module based on the attention mechanism in the framework is used to adjust the prototypes, thereby improving the representation ability and discrimination ability of the prototypes. Finally, experimental verification is carried out on the public gait dataset USC-HAD and the self-collected gait dataset CDUT-AG. The results show that the proposed method achieves the best results in the problem of small-sample class incremental gait recognition in multi-activity scenarios, effectively solves the problem of catastrophic forgetting, and verifies the effectiveness of the proposed method. Description of the Drawings
[0045] Figure 1 It is a flowchart of the different recognition methods for small-sample class incremental under multi-activity scenarios provided by the present invention.
[0046] Figure 2 It is a structural block diagram of the model framework (FC-MAGR) based on hybrid prototype enhancement provided by the present invention.
[0047] Figure 3 It is a schematic diagram of different prototype representations provided by the present invention.
[0048] Figure 4 It is a schematic diagram of the selective prototype enhancement module provided by the present invention.
[0049] Figure 5 It is the comparison result of the present invention on the USC-HAD dataset and the CDUT-AG dataset.
[0050] Figure 6 It is the comparison result of the present invention on the USC-HAD dataset and the CDUT-AG dataset.
[0051] Figure 7 It is the classification result of the present method and the comparative method CEC after restricting different numbers of activities in the USC-HAD dataset provided by the present invention.
[0052] Figure 8 It is the classification result of the present method and the comparative method CEC after restricting different numbers of activities in the CDUT-AG dataset provided by the present invention. Detailed Implementation Modes
[0053] The specific embodiments of the present invention will be described below to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions created using the concept of the present invention are within the scope of protection.
[0054] An embodiment of the present invention provides a small-sample class incremental gait recognition method in multiple activity scenarios, as Figure 1 shown, which includes the following steps:
[0055] S1. Collect gait samples in multiple activity scenarios and construct a data set;
[0056] The labels of the gait samples in the data set include user labels and activity labels;
[0057] S2. Construct a model framework based on hybrid prototype enhancement, including a feature extractor and a selective prototype enhancement module connected in sequence;
[0058] S3. Pre-train the feature extractor of the model framework using the data set to obtain a pre-trained model;
[0059] S4. On the basis of the pre-trained model, perform incremental training on the model framework, including:
[0060] Use the pre-trained model to extract the features of each support sample in the data set and form an initial prototype for each class according to the categories; the initial prototype is a hybrid prototype containing user labels and activity labels;
[0061] Input the initial prototype into the selective prototype enhancement module for incremental training to obtain an enhanced prototype, and then obtain a gait recognition model;
[0062] S5. For the query sample of the gait to be recognized input into the gait recognition model, obtain the corresponding initial prototype through the feature extractor with frozen parameters, merge it with all the prototypes in the previous session and perform prototype enhancement, and then identify the user category of the query sample by calculating the distance between the query sample and the enhanced prototype, thus completing gait recognition.
[0063] The gait recognition method provided by the embodiment of the present invention mainly realizes identity recognition in different activity scenarios. Its purpose is to recognize the user himself with a small number of gait samples of a user in multiple activity scenarios, and will not forget the old-class users who have been recognized. The model framework based on hybrid prototype augmentation (FC-MAGR) provided by this embodiment consists of multiple sessions, which are carried out in sequence. In each session, only the training data of the current session can be used to adjust the training model, that is, the samples of all previous sessions are unavailable. At the same time, in the evaluation stage after the training of each session, all categories of test samples of this session and all previous sessions will be used for evaluation.
[0064] Based on this, in step S1 of the embodiment of the present invention, the constructed data set is divided into a training set and a test set and then assigned to multiple sessions, where represents the training set data used in each session, represents the test set data used in multiple sessions. and The samples in and both belong to the label space Y l , but the samples do not overlap with each other. Assume that the label space of session l is represented as Y l , and i≠j, Specifically, in the first session , a large-scale data set with a sufficient number of categories is used, and each category has a sufficient number of data samples. In addition, in the subsequent incremental sessions , each session has only a small number of training samples. In this embodiment, the data is organized in the form of N-way-K-shot, that is, the support set of the current incremental session has only N categories, and each category has only K data samples. The query set contains several data samples of these N categories, and the samples of the support set and the query set do not intersect.
[0065] In the embodiment of the present invention, for the constructed model framework based on hybrid prototype augmentation as Figure 2 shown, in Figure 2 (a), through pre-training the model architecture, a feature representation with discriminative ability and generalization ability is learned. By pre-training on a large amount of unlabeled data, the network can learn useful general abstract features of the input data. In Figure 2(b), first, the pre-trained model is used to extract the features of each support sample, and the basic prototypes of each class are formed according to the categories. However, these basic prototypes are obtained through basic learning, so they lack sufficient feature representations to describe the categories. Therefore, in this embodiment, a Selective Prototype Enhancement Module (SPEM) is proposed, aiming to solve the problem of redundancy in the learned features. During the subsequent incremental session training process, the feature extractor is frozen, Figure 2 (c) shows the entire process of the incremental phase. In the new session, the model learns the initial prototypes of the support samples and updates them into the prototype container, only adding the prototypes of the new categories.
[0066] Based on the framework structure as Figure 2 shown, in step S4 of the embodiment of the present invention, the method for forming the initial prototype of each class according to the category is specifically as follows:
[0067] S401. Combine the m different class activity labels and n different class user labels in the dataset to form a combined class Y = {y i} and i ∈ [0, m*n);
[0068] S402. Calculate the embedding mean of each support sample in the set of all support samples in the dataset to obtain the initial prototype of each class.
[0069] In step S4 of this embodiment, the representation of the initial prototype is:
[0070]
[0071] In the formula, P b represents the initial prototype of the combined class b, S b represents the set of support samples belonging to the category b, |S b | represents the number of support samples, f b (·) is the pre-trained feature extractor, x i represents the support sample, and y i represents the category label.
[0072] Specifically, in this embodiment, a pseudo-class incremental learning scheme based on a hybrid prototype of user and activity is designed. In this process, gait samples with both user and activity dual labels are used simultaneously. Assume that the dataset has m different class activity labels, and its set is A = {a1, a2,..., a m}, and at the same time, it has n different class user labels, and its set is U = {u1, u2,..., u n}, the ultimate goal of the above model framework is to classify users. However, during the training process, active labels are used as prior knowledge and added to the training, hoping that the model's identification of different activities can increase the model's recognition of users. Therefore, in this embodiment, the above method is adopted to generate the initial prototype.
[0073] Compared with the prototypes generated by traditional prototype networks, the initial prototypes generated by the method of this embodiment The additional activity information in it makes up for the deficiencies of the basic user prototypes. For gait recognition, the walking place or the category of walking activities is the main factor affecting the gait pattern. The gait patterns of the same user in different activity scenarios vary greatly, while the gaits of different users in the same activity scenario often show a certain degree of similarity. Considering the above problems, using only user category information often cannot obtain satisfactory discriminative prototypes. Therefore, we propose to use a hybrid prototype It uses gait activity information as a supplement and can better serve our final decision-making.
[0074] In step S4 of the embodiment of the present invention, the gait data is divided according to the possible prototypes of different classification targets as Figure 3 shown. For the prototypes of different activities in (a), there are relatively obvious distinctions, but they cannot provide a basis for user classification. In (b), the user prototypes seem distinct, but the support activities between different user prototypes or the support activities of the same category of different user prototypes may partially overlap, and this part of the overlapping samples may lead to decision-making errors. (c) is the hybrid prototype proposed by the present invention. It takes into account both the distinction between different user prototypes and the overlapping problem of different prototypes in the same activity, thus presenting a more obvious decision boundary in the recognition task.
[0075] In the embodiment of the present invention, for the generated hybrid prototype, it is used as the input of the Selective Prototype Enhancement Module (SPEM), and the self-attention mechanism is applied to process it to obtain an attention map that highlights important local features. These attention maps calculate the importance scores between classes and within classes, amplify the feature information of the target class, and finally obtain discriminative prototype features.
[0076] In step S4 of the embodiment of the present invention, based on the structure of the Selective Prototype Enhancement Module as Figure 4 shown, in each incremental session, the method of incremental training based on the hybrid prototype is specifically as follows:
[0077] S411. Generate the hybrid prototype of the new category support samples through the feature extractor, merge it into the prototype container of the previous session, and form the support prototype of the current incremental session with all the hybrid prototypes in the current prototype container;
[0078] S412. Extract the hybrid prototype of the query samples through the feature extractor Input it into the selection prototype enhancement module for evaluation;
[0079] S413. In the selection prototype enhancement module, process the hybrid prototype through a 1×1 convolutional kernel with a learning parameter q to obtain a new feature map
[0080] Respectively, through 1×1 convolutional kernels with learning parameters k and v, and using the prior knowledge of the existing labels, merge the hybrid prototypes with the same user or activity label to obtain a new feature map and
[0081] S414. According to the features in the feature maps and construct a maximum attention weight matrix A;
[0082] S415. Multiply the maximum attention weight matrix A with the said feature map and add the hybrid prototype as a residual term to obtain an enhanced prototype.
[0083] In step S413 of this embodiment, the feature maps and are expressed as:
[0084]
[0085] In the formula, s q (·), s k (·), s v (·) represent linear transformation functions with q, k, v as parameters, represents the new prototype obtained by taking the mean of the prototypes with label y in the hybrid prototype y represents a user label or an activity label, and the finally obtained prototype represents one of the user prototypes or activity prototypes, and T represents the transpose operation.
[0086] In step S414 of this embodiment, the new prototype is a prototype with more compact information, which is more biased towards our target category - the user category, or the auxiliary category - the activity category. Therefore, in order to make full use of the relationship between the hybrid prototype and the new prototype calculate a maximum attention weight matrix A to obtain this bias information, and express the weight element α ij of the maximum attention weight matrix A as:
[0087]
[0088] In the formula, and respectively represent and the features at the i-th and j-th positions in denotes finding the original attention matrix the index of the maximum attention score on the first dimension, is the scaling factor, and T represents the transpose operation.
[0089] In this embodiment, by adding the hybrid prototype as a residual term, the problems of gradient vanishing and gradient explosion are prevented.
[0090] In this embodiment, the selected prototype enhancement module includes an activity enhancement module SPEM-SUDA and a user enhancement module SPEM-SADU;
[0091] The enhanced prototype in this embodiment is the prototype obtained by adjusting through the activity enhancement module SPEM-SUDA and the user enhancement module SPEM-SADU respectively and performing feature fusion, which is expressed as:
[0092]
[0093] In the formula, and respectively represent the prototype features after being adjusted by the activity enhancement module SPEM-SUDA and the user enhancement module SPEM-SADU.
[0094] Specifically, in the above process, there are two types of labels, user and activity. Therefore, after adding the prior knowledge of specific labels, the original hybrid prototype generates corresponding user prototypes or activity prototypes according to the corresponding labels. When selecting the user prototype to adjust our hybrid prototype Under the action of the maximum attention weight matrix A, it is equivalent to the SPEM module aggregating the prototype features of the same user under different activities into the attention map, which can highlight and enhance the prototypes with important features to obtain the first adjusted prototype On the other hand, we can also use the activity prototype to adjust which is equivalent to the SPEM module highlighting the prototype features of different users under the same activity to obtain the second adjusted prototype To improve the discrimination ability of the prototype, the prototype and the prototype are added together, and the fused feature is used as the final prototype.
[0095] In this embodiment, in order to achieve the best effect of prototype enhancement for SPEM, the cross-entropy loss and cosine classifier are used to train the SPEM-SADU and SPEM-SUDA branches respectively. Finally, the respective losses are added up as our final loss. That is, in step S4 of this embodiment, the loss function of the prototype enhancement module is selected. It is expressed as:
[0096]
[0097] In the formula, and respectively represent the loss functions of the active module SPEM-SUDA and the user enhancement module SPEM-SADU. and respectively represent the decomposed support prototype and query prototype. cos(·,·) represents the cosine classifier. represents the cross-entropy function, and R represents the batch data.
[0098] In step S5 of the embodiment of the present invention, the calculated score on the final prototype is used as the basis for user recognition, and its expression is:
[0099]
[0100] In the formula, and respectively represent the support prototype and the query prototype, and both belong to a part of the final prototype of.
[0101] The embodiment of the present invention provides an example for verifying the effect of the above gait recognition method.
[0102] In this embodiment, the proposed method is compared and analyzed with various other methods on the datasets USC-HAD and CDUT-AG, including Finetune, iCarl, NCM, SPPR, and CEC.
[0103] Among them, the USC-HAD activity recognition dataset contains the behavior datasets of 14 individuals (7 males and 7 females) for 12 types of activities. Since the non-periodic gait activities cannot provide sufficient information for user recognition, in order to conduct incremental gait recognition experiments, we only used the samples of the first 6 types of activities, including: Walking forward, Walking left, Walking right, Walking upstairs, Walking downstairs, Running forward. USC-HAD uses a wired inertial sensing acquisition device, MotionNode, which is connected to a micro laptop to collect data. During data collection, MotionNode is worn on the right front hip of the subject (the x-axis of MotionNode points to the bottom surface and is perpendicular to the plane formed by the y-axis and z-axis). The experiment of USC-HAD requires the subjects to conduct 5 experiments for each activity on different days at different indoor and outdoor locations, and each subject has sufficient activity gait samples.
[0104] CDUT-AG contains the data of 60 healthy people aged between 20 and 30. In the CDUT-AG data collection experiment, we used self-developed sensing insoles to collect data. Each sensor insole is integrated with inertial sensors (BMI160, including a three-axis accelerometer and a three-axis gyroscope) and 10 pressure sensors, and all sensors are connected to a control chip nrf52832 with a Bluetooth module. Finally, the collected inertial and pressure data are transmitted to the host computer for data analysis through Bluetooth. In the experiment, we required the subjects to perform three common walking activities: horizontal walking, going upstairs, and going downstairs. Among them, horizontal walking is to walk back and forth in a horizontal corridor about 8 meters long for about 2.5 minutes. Going upstairs and downstairs are carried out between floors, and the time for each going upstairs or downstairs is about 15 seconds. Each activity experiment is carried out 10 times. Some statistical information about the datasets USC-HAD and CDUT-AG is shown in Table 1.
[0105] Table 1: Statistical information of datasets USC-HAD and CDUT-AG
[0106]
[0107] The detailed experimental results are shown in Table 2 and Table 3 and Figure 5 in.
[0108] Table 2: Result comparison of different methods on dataset USC-HAD
[0109]
[0110]
[0111] Table 3: Comparison of results of different methods on the CDUT-AG dataset
[0112]
[0113] All the baseline comparison methods can achieve the highest classification accuracy in their respective Session 0, and the generalization ability of the model decreases due to the addition of new classes in subsequent sessions. Although some studies have shown that the NCM classifier performs better than the softmax classifier in incremental learning tasks, in this work, the iCarl and NCM methods fail to achieve a relatively high classification accuracy in Session 0. This may be due to the particularity of our multi-activity gait dataset. By simply stacking data sample sets of multiple activity scenarios, it is difficult for us to obtain a sample mean that can represent each user prototype, that is, the characteristics of each user, resulting in low classification performance in the initial session. Therefore, we need some special means to construct a reasonable sample set, such as using meta-learning tasks, which is reflected in the SPPR, CEC, and our methods.
[0114] In addition, as can be seen from Table 2 and Table 3, our method has achieved excellent results on both datasets. Specifically, on the USC-HAD dataset, our method has obtained the best results in all sessions and has the best AA score (average accuracy, measuring the average accuracy of the model in different sessions), which is 18.17% higher than the performance of the second-ranked CEC. At the same time, we also obtained the optimal PD score (performance degradation rate, measuring the absolute difference between the accuracy of the model in the last session and the accuracy in the first session), which is 30.13% lower than that of CEC. On the CDUT-AG, except in Session 0, our method has achieved the best results in the remaining sessions. Our method has obtained an AA score of 91.85% and a PD score of 8.21%, which are at least 1.97% and 4.19% higher than those of other methods, respectively. The experimental results show the great advantage of our method on the multi-activity gait dataset, especially in the USC-HAD data, where there are more activity categories and the improvement of our method compared with other methods is more obvious.
[0115] In this embodiment of the present invention, ablation experiments were conducted on the proposed method to verify the contribution of each component of the proposed model framework to the final classification results. HPTS (Hybrid Prototype Training Scheme) and SPEM (Selected Prototype Enhancement Module) are the main components of our proposed method. HTPS generates hybrid prototypes by introducing auxiliary labels, and SPEM selects prototypes to obtain adjusted prototypes. Here, we discuss the impact of different combinations of HTPS, SPEM-SUDA, and SPEM-SADU on model performance on the USC-HAD and CDUT-AG datasets, respectively. The results are shown in Tables 4 and 5.
[0116] Table 4: Ablation test results of key model modules HTPS, SPEM-SUDA and SPEM-SADU on the dataset USC-HAD
[0117]
[0118] Table 5: Ablation test results of the key modules of the model HTPS, SPEM-SUDA and SPEM-SADU on the dataset CDUT-AG
[0119]
[0120] For the baseline model without HPTS and SPEM, only user labels are used throughout the training process. A frozen feature extractor is used to generate traditional user category prototypes, which are then expanded through incremental sessions. No prototype adjustments are made throughout the training process. When HPTS is used, the model's AA score improves by 24.84% and 5.02% on the USC-HAD and CDUT-AG datasets, and its PD score decreases by 36.1% and 9.91%, respectively. HPTS achieves superior performance on both datasets. Because the user prototypes generated by the traditional prototype network only consider user IDs without further segmentation based on activities, they discard much of the key information. In contrast, HTPS uses a more fine-grained prototype that retains this crucial information, achieving satisfactory results. Furthermore, the performance difference between the two datasets (USC-HAD contains six activity categories, and CDUT-AG contains three) shows that HPTS's advantage is more pronounced when more activity scenarios are involved. This requires further experimental verification.
[0121] In addition, this embodiment verifies the improvement of SPEM on HPTS. Compared with HTPS alone, after using SPEM, the AA score of the model on the USC-HAD and CDUT-AG datasets increased by approximately 1.51% and 2.24%, respectively, and the PD score decreased by approximately 1.66% and 2.51%, respectively. This verifies that the prototype adjusted by SPEM has improved its generalization representation ability to a certain extent. In addition, because SPEM uses an adjustment mechanism for selecting prototypes, SPEM-SUDA and SPEM-SADU can complement each other's information to a certain extent, thereby achieving better results after feature fusion. This can be seen by comparing cases 3, 4, and 5 in Tables 4 and 5. Ablation experiments show that HTPS and SPEM have a positive impact on the performance improvement of the model.
[0122] In order to demonstrate the advantages of our proposed method in the field of multi-activity gait recognition, we use t-SNE to compress the optimized embedding prototype into a 2D space for visualization. Figure 6 As shown (in Figure 6 In the figure, dots of different colors represent data points of different categories (user categories), and the black "x" (i, j) represents the prototype of user i under activity j). We visualize the data of the first and last sessions on the USC-HAD dataset. Figure 6 (a) It can be seen that there are 6 user categories in the first session. Figure 6 As shown in (b), the last session has a total of 14 user categories, and each user category contains 6 activity categories. In the figure, we use (i, j) to represent the prototype of user i under activity j. In gait data containing multiple activities, the same user forms multiple clusters in the metric space, and each cluster usually represents a type of activity of a certain user. The same user is scattered into multiple clusters in the metric space, so using traditional prototype networks to deal with user classification tasks is full of uncertainty. In our work, we propose to use HTPS, which regards each cluster as a separate prototype, that is, the original prototype is refined into multiple prototypes representing cluster centers. In incremental tasks, slightly adjusting the position of a cluster prototype in the metric space is much less costly than adjusting the original prototype.
[0123] To evaluate the superiority of our method on multi-activity datasets, we perform experiments sequentially with the number of activities increasing from 1 to the upper limit of the dataset (6 for the USC-HAD dataset and 3 for the CDUT-AG dataset). Given that the CEC method is the most competitive method besides ours, we compare the performance of the two methods with increasing activity numbers. Figure 7 and Figure 8The classification results after restricting different numbers of activities in the USC-HAD dataset and the CDUT-AG dataset using our method and the comparative method CEC are shown respectively. It can be found from the figure that as the number of activities increases, the AA and PD scores of the comparative method are continuously decreasing, while our method always remains relatively stable, which is in line with our expected results.
[0124] Meanwhile, we compare the difference in accuracy between our method and the comparative method CEC under different experiments. Tables 6 and 7 show the degree of improvement in the accuracy of our method relative to CEC. On the one hand, our method basically maintains a positive improvement under different numbers of activities, indicating that our method can be applied to multi-activity scenarios and has a certain generalization ability for different activity scenarios. On the other hand, as the number of activities continues to increase, the degree of improvement of our method relative to CEC is also increasing.
[0125] Table 6: Incremental results of the classification accuracy of our method compared to CEC under different numbers of activities in the USC-HAD dataset
[0126]
[0127] Table 7: Incremental results of the classification accuracy of our method compared to CEC under different numbers of activities in the CDUT-AG dataset
[0128]
[0129] It can be seen from Table 6 that in the USC-HAD dataset, as the number of activities increases from 1 to 6, the ΔAA score (the difference between the AA of our method and CEC) increases from +2.79% to +18.17%, and the ΔPD score (the difference between the PD of our method and CEC) decreases from -3.67% to -30.12%. It can be seen from Table 7 that in the CDUT-AG dataset, as the number of activities increases from 1 to 3, the ΔAA increases from -0.49% to +1.97%, and the ΔPD decreases from +0.28% to -4.19%. This shows that our method performs better in the face of data with more activities, that is, the more activities there are, the more obvious the advantages of our method.
[0130] Specific embodiments are applied in the present invention to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
[0131] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. A small-sample class-incremental gait recognition method under multiple activity scenarios, characterized in that, It includes the following steps: S1. Collect gait samples in multiple activity scenarios and construct a dataset; The labels of the gait samples in the dataset include user labels and activity labels; S2. Construct a model framework based on hybrid prototype enhancement, including a feature extractor and a selective prototype enhancement module connected in sequence; S3. Use the dataset to pre-train the feature extractor of the model framework to obtain a pre-trained model; S4. On the basis of the pre-trained model, perform incremental training on the model framework, including: Use the pre-trained model to extract the features of each support sample in the dataset and form the initial prototype of each class according to the category; The initial prototype is a hybrid prototype containing user labels and activity labels; Input the initial prototype into the selective prototype enhancement module for incremental training to obtain an enhanced prototype, and then obtain a gait recognition model; S5. For the query sample of the gait to be recognized input into the gait recognition model, obtain the corresponding initial prototype through the feature extractor with frozen parameters, merge it with all the prototypes in the previous session and perform prototype enhancement, and then identify the user category of the query sample by calculating the distance between the query sample and the enhanced prototype, completing gait recognition.
2. The small-sample incremental gait recognition method in multiple activity scenarios according to claim 1, wherein In step S4, the method of forming the initial prototype of each class according to the category is specifically: S401. Combine the activity tags of m different categories and the user tags of n different categories in the dataset to form a combined category Y with m×n categories, where Y = {y i}, and i ∈ [0, m*n); S402. Calculate the embedding mean of each support sample in all support sample sets in the dataset to obtain the initial prototype of each class.
3. The small-sample incremental gait recognition method in multiple activity scenarios according to claim 1, wherein In step S4, the representation of the initial prototype is: where, P b represents the initial prototype of combination class b, S b represents the support sample set belonging to class b, |S b | represents the number of support samples, f b (·) is the pre-trained feature extractor, x i represents the support sample, y i represents the class label.
4. The small-sample incremental gait recognition method in multiple activity scenarios according to claim 1, characterized in that, In step S4, in each incremental session, the method of performing incremental training based on the hybrid prototype is specifically: S411. Generate the hybrid prototype of the new category support sample through the feature extractor, merge it into the prototype container of the previous session, and form the support prototype of the current incremental session with all the hybrid prototypes in the current prototype container; S412. Extract the hybrid prototype of the query sample through the feature extractor Input it into the selected prototype enhancement module for evaluation; S413. In the selection prototype enhancement module, the mixed prototype is processed through a 1×1 convolutional kernel with a learning parameter q to obtain a new feature map Respectively, through a 1×1 convolutional kernel with learning parameters k and v, and using the prior knowledge of existing labels, the mixed prototypes with the same user or activity labels are merged to obtain a new feature map and S414. Construct the maximum attention weight matrix A according to the features in the feature map and . S415. Multiply the maximum attention weight matrix A with the feature map and add the mixed prototype as a residual term to obtain an enhanced prototype.
5. The small-sample incremental gait recognition method in multiple activity scenarios according to claim 4, wherein In the step S413, the feature map and are represented as: where s q (·), s k (·), s v (·) represents a linear transformation function with q, k, v as parameters, denotes the new prototype obtained by taking the mean of the prototypes labeled with y in the hybrid prototype, y represents a user label or an activity label, and the finally obtained prototype represents either the user prototype or the activity prototype, and T represents the transpose operation.
6. The small-sample incremental gait recognition method in multiple activity scenarios according to claim 4, characterized in that In the step S414, the weight element α of the maximum attention weight matrix A ij is expressed as: Wherein, and respectively represent and the features at the i-th and j-th positions in denotes finding the index of the maximum attention score on the first dimension of the original attention matrix is the scaling factor, and T represents the transpose operation. 7. The small-sample incremental gait recognition method in multiple activity scenarios according to claim 4, wherein The selective prototype enhancement module includes an activity enhancement module SPEM-SUDA and a user enhancement module SPEM-SADU; The enhanced prototype is a prototype obtained by adjusting and performing feature fusion on the activity enhancement module SPEM-SUDA and the user module SPEM-SADU respectively, and is expressed as: In the formula, and respectively represent the prototype features adjusted by the activity enhancement module SPEM-SUDA and the user enhancement module SPEM-SADU.
8. The small-sample incremental gait recognition method in multiple activity scenarios according to claim 7, wherein In the step S4, the loss function of the selected prototype enhancement module is expressed as: In the formula, and respectively represent the loss functions of the activity enhancement module SPEM-SUDA and the user enhancement module SPEM-SADU, and respectively represent the decomposed support prototype and query prototype, cos(·,·) represents the cosine classifier, represents the cross-entropy function, and R represents the batch data.
Citation Information
Patent Citations
Infrared image human body gait recognition method under complex scene based on improved ViT
CN114708619A
Lightweight prototype container-based small sample class incremental learning system
CN116229151A