Incremental-like gesture recognition method, electronic equipment and medium

Through task-specific combination mechanism and prototype enhancement hybrid regularization module, dynamically aggregate the task isolation parameters of the gesture recognition system, and enhance model parameters through prototype pseudo-sample, the catastrophic forgetting and recognition confusion problems of the gesture recognition system under the condition of no playback data is solved, and dynamic adaptation to new gesture categories and historical knowledge are achieved.

CN120071433AActive Publication Date: 2025-05-30NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510061303.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-30
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

Existing gesture recognition systems are prone to catastrophic forgetting without replay data, and it is difficult to dynamically adapt to new gesture categories and recognition confusion between gestures.

Method used

The class-incremental gesture recognition method is adopted to dynamically aggregate the isolation parameters of multiple tasks through task-specific combinatorial mechanism (TCM) and prototype enhancement hybrid regularization (PHR) modules, and the final features for the current task are constructed, and the model parameters are enhanced through prototype pseudo-samples to reduce the interference of new gesture learning on the learned knowledge.

Benefits of technology

It effectively solves the problem of catastrophic forgetting and confusion between gestures under the condition of no playback data, dynamically adapts to the introduction of new gesture categories, while retaining historical gesture knowledge, significantly improving the stability and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071433A_ABST
    Figure CN120071433A_ABST
Patent Text Reader

Abstract

The invention relates to a similar increment gesture recognition method, electronic equipment and a medium. The method comprises the following steps: receiving representation input data; extracting intermediate features according to the input data and constructing task-specific index embedding, wherein each embedding corresponds to a gesture recognition task; calculating the similarity between the intermediate features and index embedding, and generating a combined weight; dynamically aggregating the isolation parameters of the plurality of tasks by using the weights, and constructing final features of the current task; inputting the final features into a classifier, and predicting gesture categories; training an adapter module according to the prediction result; task index embedding and model parameters are enhanced by calculating a historical category prototype and generating a pseudo sample; forming prototype enhanced mixed regularization loss in combination with an enhancement result, and further optimizing the model; and after all task training is completed, classifying newly input gesture data by using the optimization model. The method is suitable for gesture recognition tasks with confusion characteristics, and endows continuous learning based on parameter isolation with adaptive combination capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and artificial intelligence, and specifically relates to a class-incremental gesture recognition method, an electronic device, and a medium. Background Art

[0002] Non-verbal communication plays a crucial role in human daily interactions, accounting for more than 65% of the total communication volume (Mohamed, Mustafa, and Jomhari, 2021). Among the various forms of non-verbal communication, gestures have attracted much attention due to their unique intuitiveness and naturalness. Humans instinctively use gestures to communicate, perceive the environment, and interact with the environment. With the rapid development of virtual reality (VR) and augmented reality (AR) technologies, gesture recognition, as an important bridge connecting the physical and virtual worlds, has played a key role in promoting the core immersive and real experiences in the metaverse (Duan et al., 2023). Especially in AR / VR wearable devices (such as Meta Quest and VUZIX), gesture recognition technology has become an indispensable interaction method.

[0003] In recent years, thanks to the rapid development of deep neural networks, gesture recognition technology has achieved an identification accuracy rate close to 96% on standard datasets (Min et al., 2020) and provided a highly immersive user experience in actual virtual environments. However, the current mainstream gesture recognition methods usually focus on improving the test accuracy rate on specific datasets through single centralized training. The limitations of this method are obvious: in real applications, gesture recognition systems need to be able to continuously learn new knowledge and dynamically adapt to the needs of different users. For example, when a gesture recognition system is deployed on a VR edge device, users may gradually register new gesture categories according to their needs. If traditional static methods are used to train samples containing new categories, it often leads to the catastrophic forgetting problem of the system's original knowledge. In addition, the acquisition of old-category samples in gesture recognition tasks is often restricted by privacy protection, and historical data cannot be directly accessed. Therefore, building a continuous learning model that supports sequential registration of new-category gestures during the life cycle without replay data has become an important challenge in the AR / VR field.

[0004] Continual Learning (CL) technology has been proposed to solve the catastrophic forgetting problem and has made remarkable progress in recognition tasks involving images, audio, and text (Qu et al., 2021; Biesialska et al., 2020; Parisi et al., 2019). However, the research on Class-Incremental Hand Gesture Recognition (CI-HGR) is still relatively limited. The existing continual learning methods can be roughly divided into the following three categories:

[0005] 1. Regularization-based methods: By introducing explicit regularization terms to retain knowledge from previous tasks (Li and Hoiem, 2017; Kirkpatrick et al., 2017). Such methods perform well in preventing catastrophic forgetting but often suffer from inefficiency and limited performance.

[0006] 2. Replay-based methods: Mitigate the forgetting problem by storing old-class samples or generating pseudo-samples (Rebuffi et al., 2017; Rolnick et al., 2019). However, this method may involve privacy issues, and the generated pseudo-samples may introduce historical model biases, affecting model performance.

[0007] 3. Parameter isolation-based methods: Construct independent parameters for each task to avoid interference (Madotto et al., 2021). Although this method performs excellently in Task-Incremental Learning (TIL), in Class-Incremental Learning (CIL), due to the unknown test tasks, the model performance is limited.

[0008] In skeleton-based action recognition tasks, due to their high robustness and compactness, continual learning methods have been preliminarily applied. For example, some studies have achieved task incremental learning through dynamic network expansion or pseudo-sample replay techniques (Aich et al., 2023). However, these methods either fail to meet the requirements of privacy protection or are less efficient in terms of time and memory than parameter isolation techniques. Summary of the Invention

[0009] The present invention provides a class-incremental gesture recognition method, an electronic device, and a medium, aiming to solve the problems that existing gesture recognition systems are prone to catastrophic forgetting, difficult to dynamically adapt to new gesture categories, and recognition confusion between gestures under the condition of no replay data.

[0010] To achieve the above object, the first aspect of the present invention provides a class-incremental gesture recognition method, including the following steps:

[0011] Receive input data representing a gesture, where the input data is a three-dimensional skeleton data containing a time series, and each frame of three-dimensional skeleton data includes spatial coordinates of multiple hand joints;

[0012] Extract intermediate features based on the received input data, and construct task-specific index embeddings based on the intermediate features, where each index embedding is associated with a specific gesture recognition task;

[0013] Calculate the similarity between the extracted intermediate features and the task-specific index embeddings to generate combined weights;

[0014] Using the generated combined weights, dynamically aggregate the isolation parameters of multiple tasks to construct the final features for the current task;

[0015] Input the constructed final features into a classifier to perform the prediction of gesture categories;

[0016] According to the prediction results, train the adapter module for each task;

[0017] During the training of the adapter module, by calculating the prototypes of historical gesture categories and generating prototype pseudo-samples, use the prototype pseudo-samples to enhance the task-specific index embeddings and the model parameters;

[0018] Combine the index embedding enhancement and the model parameter enhancement to form the overall prototype-enhanced hybrid regularization loss, and continue to adjust the model;

[0019] After completing the training of all tasks, use the adjusted model to classify and recognize the newly input gesture data.

[0020] Furthermore, the method for constructing task-specific index embeddings based on the intermediate features includes:

[0021] Extract the three-dimensional spatial coordinate features of each frame of hand joints from the input data, and the three-dimensional spatial coordinate features include the positions of each hand joint in three-dimensional space;

[0022] Use the time series modeling method to extract the time series features of the three-dimensional spatial coordinate features to generate the time series features representing the dynamic changes of the gesture;

[0023] Fuse the time series features and the three-dimensional spatial features through a feature fusion network to generate the intermediate features of the gesture;

[0024] Based on the intermediate features, generate a task-specific index embedding vector associated with the current task, and the task-specific index embedding vector is constructed through a task-specific trainable module;

[0025] Dynamically combine the task-specific index embedding vector and the intermediate features to form the task-specific index embedding for subsequent task learning.

[0026] Furthermore, the method for calculating the similarity between the extracted intermediate features and the task-specific index embeddings to generate the combined weights includes:

[0027] Calculate the cosine similarity between the generated intermediate features and each task-specific index embedding to obtain the similarity scores;

[0028] Calculate the above similarity scores for all task-specific index embeddings;

[0029] Divide each similarity score by a predefined scaling factor to obtain a scaled similarity score;

[0030] Apply an exponential function to all the scaled similarity scores to obtain the unnormalized values of the weights;

[0031] Normalize all the unnormalized weight values and calculate the combined weight for each task-specific index embedding vector. The calculation formula is as follows:

[0032]

[0033] where, W t is the combined weight, exp is the exponential function, cos(f, K t ) is the cosine similarity between the intermediate feature representation f and the task-specific index embedding vector K t , K is the index embedding pool, which contains the set of all task-specific index embedding vectors, K j is the f-th task-specific index embedding vector in the set K, f represents the intermediate feature vector extracted from the input gesture data by the feature extractor, τ is the scaling factor, is the similarity.

[0034] Furthermore, the dynamic aggregation includes weighted combination of the isolation parameters of the multiple tasks according to the combined weight, specifically including:

[0035] Apply the combined weight to the isolation parameter of each relevant task;

[0036] Apply the weighted isolation parameters to the extracted intermediate features one by one;

[0037] Accumulate the features after being processed by each task to obtain the final feature for the current task.

[0038] Furthermore, the gesture category prediction includes the following specific steps:

[0039] Input the constructed final feature into a pre-trained classifier module;

[0040] In the classifier module, perform a linear transformation on the final feature using a fully connected layer;

[0041] By applying the Softmax function, convert the output after the linear transformation into the probability distribution of each gesture category;

[0042] According to the probability distribution, select the category with the highest probability as the final gesture recognition result.

[0043] Furthermore, during the training process of the adapter module, the method of calculating the prototype of the historical gesture categories and generating prototype pseudo-samples includes:

[0044] Extract the features belonging to the historical gesture category through the feature extractor of the gesture recognition model;

[0045] According to the gesture samples belonging to the same category, calculate the prototype and covariance matrix of the category. The specific formulas are as follows:

[0046]

[0047] where μ c is the prototype of category c, N c is the number of samples in category c, y i = c indicates that the label of sample x i belongs to category c, represents the input gesture x i after passing through the feature extractor to obtain the intermediate features; is the covariance matrix of category c, and T is the transpose operator;

[0048] Based on the calculated prototype and covariance matrix, generate prototype pseudo-samples for each historical gesture category. The prototype pseudo-samples follow a multivariate Gaussian distribution:

[0049]

[0050] where, is the feature vector of the generated prototype pseudo-sample, representing the virtual sample of category c, is the multivariate Gaussian distribution, with a mean of μ c , and a covariance matrix of c ∈ {C K \C U} means that category c belongs to the set C of the learned historical categories K minus the set C of new categories introduced by the current task U .

[0051] Furthermore, the methods for enhancing the task-specific index embedding using the prototype pseudo-samples include:

[0052] Adopt a metric learning strategy, use the prototype pseudo-samples as negative samples, and perform contrastive learning with the intermediate features of the current task to optimize the index embedding;

[0053] By maximizing the similarity between the original sample features and the correct task embedding, and minimizing its similarity with the pseudo-sample features, update the task-specific index embedding. Use the following loss function to optimize the task-specific index embedding:

[0054]

[0055] where, A prototype-driven embedding-enhanced loss function for measuring and optimizing task-specific index embeddings. m is a margin parameter used to control the minimum distance between different classes. Is a prototype pseudo-sample The cosine similarity between the prototype pseudo-sample and the task-specific index embedding K. cos(f, K) is the current gesture feature The cosine similarity between the prototype pseudo-sample and the task-specific index embedding K.

[0056] Furthermore, the method of enhancing model parameters using prototype pseudo-samples includes:

[0057] Transfer knowledge using the pseudo-sample as a mediator based on the prototype pseudo-samples of historical task categories and the training sample features of the current task;

[0058] Design a loss function that aligns the class probabilities of the current task model and the historical task model;

[0059] Optimize the model parameters using the following loss function:

[0060]

[0061] Where and Represent the predictions of the current task model and the previous task model for the pseudo-sample input respectively. Is the classifier of the current task. Is the intermediate feature representation of the input gesture. Is the combined weight calculated from the index embedding. Is the transformed output of the i-th task's adapter module for the feature ; Is the predicted probability distribution after alignment of the current task model. Is the predicted probability distribution of the current task model on the learned class set, C K Is the learned class set, C U Is the newly added class set of the current task. Is the normalization operation, indicating the normalization process of the historical class probabilities. Is a prototype-driven embedding-enhanced loss function for measuring and optimizing task-specific index embeddings. KL represents the Kullback-Leibler divergence function, γ is the weight coefficient of the loss function used to balance the loss value, CE(·) is the task loss function of the gesture recognition task. Is the label corresponding to the class prototype pseudo-sample.

[0062] To achieve the above object, a second aspect of the present invention provides an electronic device, including a processor and a memory, where the processor is configured to implement the steps of the class-incremental gesture recognition method when executing a computer program stored in the memory.

[0063] To achieve the above object, a third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program executes the steps of the class-incremental gesture recognition method when run by a processor.

[0064] Advantages of the present invention:

[0065] Compared with the prior art, a class-incremental gesture recognition method, an electronic device and a medium provided by the present invention effectively solve the problems of catastrophic forgetting and recognition confusion between gestures under the condition of no replay data by introducing a Task-specific Compositional Mechanism (TCM) and a Prototype-enhanced Hybrid Regularization (PHR). TCM constructs a trainable task index embedding, and adaptively combines task-specific parameters by using the similarity between the features of the input gesture and the historical task features, realizes dynamic fine-tuning and inference, and avoids the problem of difficult parameter selection in the case of unknown tasks. At the same time, PHR uses the prototype pseudo-samples of historical gesture categories to improve the discrimination ability of task-specific embeddings through Prototype-driven Embedding Enhancement (PEE), and realizes knowledge transfer and prevents over-adjustment of parameters through Prototype-driven Parameter Enhancement (PPE), thereby significantly improving the stability and accuracy of the model, and reducing the interference of new gesture learning on the learned knowledge. Description of the Drawings

[0066] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.

[0067] Figure 1 is a flowchart of a class-incremental gesture recognition method disclosed in an embodiment of the present invention.

[0068] Figure 2 is an ablation study diagram of a PHR disclosed in an embodiment of the present invention.

[0069] Figure 3 is a basic process and performance comparison diagram of a class-incremental gesture recognition task disclosed in an embodiment of the present invention.

[0070] Figure 4 is a conceptual diagram of the working mechanism of a PENCIL method disclosed in an embodiment of the present invention. Detailed Embodiments

[0071] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the scope of protection of the present invention.

[0072] According to the embodiments of the present invention, it should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the following manufacturing method, in some cases, the steps shown or described can be executed in a different order than here.

[0073] In the prior art, gesture recognition systems usually rely on static data sets for centralized training, resulting in the problem of catastrophic forgetting when facing newly added gesture categories. In addition, existing continual learning methods such as replay-based and regularization-based methods have problems of data privacy and low efficiency in practical applications. The present invention provides a class-incremental gesture recognition method (hereinafter referred to as the "PENCIL method"). Through a task-specific combination mechanism (TCM) and a prototype-enhanced hybrid regularization (PHR) module, it provides an efficient and privacy-friendly class-incremental gesture recognition solution, which can dynamically adapt to the introduction of new gesture categories while retaining historical gesture knowledge and avoiding forgetting.

[0074] The following will gradually and specifically describe the implementation steps of the PENCIL method, as Figure 1 shown, the method includes the following steps:

[0075] Step S100, receiving input data representing a gesture, where the input data is a three-dimensional skeleton data including a time series, and each frame of the three-dimensional skeleton data includes spatial coordinates of multiple hand joints;

[0076] The received gesture input data is a three-dimensional skeleton data including a time series, and each frame of data includes spatial coordinates of multiple hand joints. Specifically, the data is represented as where: T is the time dimension (number of frames), J is the number of hand joints, and C is the coordinate dimension (usually 3, representing the X, Y, and Z axes). Compared with high-dimensional video data, three-dimensional skeleton data has higher robustness and compactness, and is suitable for action recognition tasks with low sample sizes, such as medical action recognition; gestures have obvious dynamic change characteristics, and through time series modeling methods, the time evolution information of gestures can be effectively captured.

[0077] Step S200: Extract intermediate features from the received input data and construct task-specific index embeddings based on the intermediate features, where each index embedding is associated with a specific gesture recognition task;

[0078] Extract the three-dimensional spatial coordinate features of each frame of hand joints from the input data, denoted as J = {j 1 , j 2 ,..., j V}, where V = |J| is the number of joints included in the gesture skeleton. Use time series modeling methods (such as RNN, Transformer, etc.) to extract temporal features from the three-dimensional spatial coordinate features to generate temporal features representing the dynamic changes of the gesture; fuse the temporal features and the three-dimensional spatial features through a feature fusion network to generate the intermediate feature f of the gesture. Based on the intermediate feature f, generate a task-specific index embedding vector K t associated with the current task through a task-specific trainable module. Dynamically combine the task-specific index embedding vector K t with the intermediate feature f to form a task-specific index embedding for subsequent task learning.

[0079] Step S300: Calculate the similarity between the extracted intermediate features and the task-specific index embeddings to generate combined weights;

[0080] Calculate the cosine similarity between the intermediate feature f and each K 0 , K 1 ,..., K n in the index embedding pool K = [K t to obtain the similarity score cos(f, K t ). Divide each similarity score by a predefined scaling factor τ to obtain the scaled similarity score cos(f, K t ) / τ; apply the exponential function to the scaled similarity score to obtain the unnormalized value of the weight exp(cos(f, K t ) / τ); normalize all the unnormalized weight values to calculate the combined weight W t of each task-specific index embedding vector, and its calculation formula is:

[0081]

[0082] where W t is the combined weight, exp is the exponential function, cos(f, K t ) is the cosine similarity between the intermediate feature representation f and the task-specific index embedding vector K t , K is the index embedding pool, which is the set containing all task-specific index embedding vectors, K jis the f-th task-specific index embedding vector in the set K, where f represents the intermediate feature vector extracted from the input gesture data by the feature extractor, and τ is the scaling factor. is the similarity.

[0083] Step S400: Use the generated combined weights to dynamically aggregate the isolation parameters of multiple tasks and construct the final feature for the current task.

[0084] According to the generated combined weight W t , perform weighted combination on the corresponding isolation parameters M = {M 0 , M 1 ,..., M n} in the index embedding pool K, that is:

[0085]

[0086] where W t is the combined weight of task τ t . By applying the weight W t to the task isolation parameter M t , the dynamic participation degree of different task parameters in the current gesture is realized. The task parameters with higher weights will have a greater impact on the generation of the final feature.

[0087] Apply the weighted combined isolation parameter to the intermediate feature f and add it to f itself to obtain the final output feature That is:

[0088]

[0089] where f is the initial intermediate feature, and M i (f) is the result of the action of the isolation parameter module of task τ i on the feature f. W i is the combined weight of task τ i .

[0090] The final feature is generated by superimposing the output of the task isolation parameter and the intermediate feature. This fusion method not only retains the global characteristics of the input gesture but also introduces task-related characteristics. The accumulation process weights and integrates all task features to ensure that the final feature can reflect the comprehensive relationship between the input gesture and each task characteristic. Higher weights are assigned to the strong correlation characteristics between the input gesture and the current task, while effectively retaining the knowledge of historical tasks and avoiding catastrophic forgetting.

[0091] Step S500: Input the constructed final feature into the classifier to perform the prediction of the gesture category.

[0092] Input the constructed final feature Input to the pre-trained classifier module F c Among them, in the classifier module, the fully connected layer is used to perform a linear transformation on the final features to generate class scores. By applying the Softmax function, the output after the linear transformation is converted into a probability distribution P of each gesture class. According to the probability distribution P, the class with the highest probability is selected as the final gesture recognition result, thereby achieving accurate classification of the input gesture. Specifically, the fully connected layer completes this process through matrix multiplication and bias addition, generating scores for each gesture class, and these scores represent the linear activation values of the input gesture belonging to each class. Then, the classifier converts these linear activation values into a probability distribution by applying the Softmax function.

[0093] Step S600: According to the prediction result, train the adapter module for each task;

[0094] According to the prediction result of the classifier, determine the class to which the current input gesture belongs and associate it with the corresponding task, corresponding to the current task τ t , and train its independent adapter module M t to optimize its adaptability to the current task. Specifically, in the current task training stage, the model only updates the adapter module related to the current task, while the adapter parameters of other tasks remain frozen. Ensure that the knowledge of the current task can be efficiently learned through the adapter module, while avoiding interference with the parameters of historical tasks, thereby alleviating the catastrophic forgetting problem. In addition, through task-specific training, the model can strengthen the gesture feature representation of the current task, improve the classification performance and maintain an efficient parameter isolation mechanism.

[0095] Step S700: During the training process of the adapter module, by calculating the prototypes of historical gesture classes and generating prototype pseudo-samples, use the prototype pseudo-samples to enhance the task-specific index embedding and the model parameters;

[0096] Through the feature extractor of the gesture recognition model extract the features belonging to historical gesture classes, and calculate the prototype μ of each class C according to the gesture samples belonging to the same class c and the covariance matrix The specific formula is:

[0097]

[0098] Among them, μ c is the prototype of class c, N c is the number of samples in class c, y i = c means that the label of the sample x i belongs to class c, represents the input gesture xi The intermediate features obtained through the feature extractor ; is the covariance matrix for class c, and T is the transpose operator.

[0099] Based on the calculated prototype μ c and the covariance matrix generate prototype pseudo-samples for each historical gesture class, and the pseudo-samples follow a multivariate Gaussian distribution:

[0100]

[0101] where is the feature vector of the generated prototype pseudo-sample, representing the virtual sample of class c, is the multivariate Gaussian distribution with mean μ c , and the covariance matrix is c ∈ {C K \C U} means that class c belongs to the set C of learned historical classes K minus the set C of new classes introduced by the current task U .

[0102] Utilize the generated prototype pseudo-samples, through the task-specific combination mechanism (TCM) and prototype-enhanced hybrid regularization (PHR), to enhance the task-specific index embedding and model parameters, improve the model's recognition ability for historical classes, and reduce catastrophic forgetting.

[0103] Prototype Embedding Enhancement (PEE): Adopt a metric learning strategy, use the prototype pseudo-samples as negative samples, perform contrastive learning with the intermediate features of the current task, optimize the index embedding that can be indexed, and use the loss function for optimization:

[0104]

[0105] where is the loss function of prototype-driven embedding enhancement, used to measure and optimize the task-specific index embedding, m is the margin parameter, used to control the minimum distance between different classes, is the prototype pseudo-sample and the cosine similarity between the task-specific index embedding K, cos(f, K) is the cosine similarity between the current gesture feature and the task-specific index embedding K.

[0106] By introducing the prototype pseudo-samples as negative samples, PEE improves the robustness of the index embedding and enhances the discrimination ability between tasks.

[0107] Prototype Parameter Enhancement (PPE): Using prototype pseudo-samples as mediators to achieve knowledge transfer from the previous task model to the current task model and prevent excessive adjustment of model parameters. Design a knowledge transfer loss function as follows:

[0108]

[0109] where and represent the predictions of the current task model and the previous task model for the pseudo-sample input respectively, is the classifier of the current task, is the intermediate feature representation of the input gesture, is the combined weight calculated by the indexable embedding, is the transformed output of the adapter module of the i-th task for the feature ; is the predicted probability distribution after alignment of the current task model, is the predicted probability distribution of the current task model on the learned class set, C K is the learned class set, C U is the newly added class set of the current task, is the normalization operation, indicating normalizing the historical class probabilities; is the loss function of prototype-driven embedding enhancement, used to measure and optimize task-specific index embeddings, KL represents the Kullback-Leibler divergence function, γ is the weight coefficient of the loss function, used to balance the loss value, CE(·) is the task loss function of the gesture recognition task, is the label corresponding to the class prototype pseudo-sample.

[0110] Step S800: Combine index embedding enhancement and model parameter enhancement to form an overall prototype enhancement hybrid regularization loss, and continue to adjust the model;

[0111] Integrate the task loss, embedding enhancement loss, and parameter enhancement loss to calculate the total loss. The total loss function integrates the task recognition loss and two regularization losses:

[0112]

[0113] where α and β are the trade-off factors of and respectively, and represents the cross-entropy loss of the gesture recognition task. Through the above combination, the model can effectively retain the knowledge of historical tasks while learning new tasks, improving the overall task robustness.

[0114] Step S900: After completing all task trainings, use the adjusted model to classify and recognize newly input gesture data.

[0115] After completing all task trainings, use the adjusted model to classify and recognize newly input gesture data. After the training is completed, the model integrates the knowledge of all tasks and has the ability of continuous learning. When new gestures are input: According to the input gesture data, extract intermediate features through a feature extractor, and calculate the similarity with the task index embedding to dynamically aggregate task isolation parameters. Input the aggregated features into the classifier module, output the probability distribution of each gesture category, and select the category with the highest probability as the recognition result. The model trained by prototype-enhanced hybrid regularization can accurately recognize all learned gesture categories, avoid catastrophic forgetting, and maintain high performance and task scalability in the learning of new categories.

[0116] In this embodiment, as described in step S100 above, in this step, the input data is three-dimensional skeleton data in the form of a time series, representing the dynamic changes of gestures. The skeleton data of each frame consists of the spatial coordinates of multiple human hand joints, where each joint coordinate is a point in a three-dimensional Cartesian coordinate system. Specifically:

[0117] The three-dimensional skeleton data is usually collected by an action capture device or a depth sensor (such as Microsoft Kinect, Intel RealSense, or Leap Motion). These devices generate a hand skeleton model by capturing depth information, or extract the three-dimensional skeleton data of the hand from video data using a convolutional neural network (such as OpenPose); before inputting into the model, this data needs to go through preprocessing such as normalization (reducing individual differences of gestures, such as the size and movement range of the hand), denoising (using a smoothing filtering method such as Gaussian filtering to reduce acquisition noise), frame alignment (unifying the time series length through interpolation or cropping), and coordinate transformation (aligning the coordinates to the center of the palm to reduce the initial position deviation); the three-dimensional skeleton data refines the movement trajectories and spatial relationships of hand joints, which are important features for gesture recognition. Because it discards irrelevant information such as color and background, retains the core action features, and at the same time has a low dimension, high computational efficiency, and strong privacy protection ability, it shows excellent generalization.

[0118] In this embodiment, as described in step S200 above, the goal of step S200 is to extract dynamic and structured gesture features from the input three-dimensional skeleton data, and combine these features with task-related information to generate task-specific index embeddings for supporting dynamic parameter combinations in class-incremental gesture recognition. The following is a detailed description in combination with the content of steps S201 - S205:

[0119] Step S201: Extract the three-dimensional spatial coordinate features of each frame of hand joints from the input data.

[0120] In this step, the model extracts the spatial coordinate information of joints frame by frame from the input three-dimensional skeleton data, including the position coordinates of the palm and finger joints, etc. i =(x i , y i , z i ), and these data represent the spatial shape and position of the gesture, capturing the position information of the joints in the three-dimensional space and the spatial relationship between the joints, thus describing the basic shape of the gesture (such as the difference between an open palm and a clenched fist); at the same time, the joint coordinates are normalized (such as normalized with the palm as the center), and the noise that may be introduced by the acquisition device is removed through filtering methods.

[0121] Step S202: Use the time series modeling method to extract the temporal features of the three-dimensional spatial coordinate features and generate the temporal features representing the dynamic changes of the gesture.

[0122] It can be understood that a gesture is an embodiment of a dynamic action, and relying solely on the spatial coordinate information of each frame is not sufficient to comprehensively describe its dynamic characteristics. The time series modeling method (such as based on RNN, LSTM or Transformer) is used to extract the change patterns of the gesture action in the time dimension. Using RNN (Recurrent Neural Network) or its improved versions (such as LSTM, GRU), the time dependencies in the skeleton sequence are captured.

[0123] The generated temporal features contain the dynamic change information of the gesture in the time dimension, such as the action feature of a certain finger gradually bending. Through time series modeling, the dynamic features of the gesture are extracted, enabling the model to understand the continuous changes of the action rather than the static state of a single frame.

[0124] Step S203: Fuse the temporal features and the three-dimensional spatial features through a feature fusion network to generate the intermediate features of the gesture.

[0125] Integrate the spatial features extracted in Step S201 and the temporal features extracted in Step S202 through a feature fusion network. The fusion methods can include simple concatenation, weighted combination or deep fusion through a neural network. The fusion network can be designed as a parallel structure, with two paths of networks respectively processing the spatial features and the temporal features, and finally fusing and outputting in the high-dimensional feature space. It should be noted that the intermediate features are a high-dimensional representation that combines the static shape and dynamic action of the gesture, containing both the spatial layout information of the gesture and the dynamic changes in the time dimension; the fused intermediate features have stronger expressive ability and can comprehensively depict the static and dynamic characteristics of the gesture.

[0126] Step S204: Generate a task-specific index embedding vector associated with the current task based on the intermediate feature.

[0127] Each task has an independent index embedding vector to represent the specific features of that task. These embedding vectors are generated by a trainable module (such as a small indexable neural network) and are associated with a task-specific parameter isolation module. For each task τ t , construct an index embedding vector K t . The training of the embedding vector is completed based on the mapping relationship between the gesture intermediate feature f and the task label. These embedding vectors will be stored in the task retrieval pool K = [K 0 , K 1 ,..., K n to support subsequent dynamic parameter combinations. The task-specific index embedding vector provides an independent feature representation for each task, ensuring that the model can distinguish gesture categories between different tasks during the task incremental learning process.

[0128] Step S205: Dynamically combine the task-specific index embedding vector with the intermediate feature to form a task-specific index embedding for subsequent task learning.

[0129] The process of dynamic combination enables the model to automatically select and activate the parameters corresponding to the task according to the features of the input gesture, thus supporting exemplar-free class incremental learning.

[0130] Through S201 - S205 in Step S200, the model extracts static and dynamic features layer by layer from the 3D skeleton data, generates intermediate features through feature fusion, and further constructs task-specific index embedding vectors. The index embedding provides independent parameter support for each task through a dynamic combination mechanism, thus ensuring that the model can effectively learn and reason in the class incremental gesture recognition task while avoiding catastrophic forgetting and confusion problems.

[0131] In this embodiment, as described in Step S300 above, in order to dynamically combine task-specific parameters, it is necessary to generate the combination weight for each task according to the similarity between the gesture intermediate feature and the task-specific index embedding. These combination weights reflect the relevance of the current input gesture to each task and are used to dynamically select the combination of task-specific parameters. The following will be explained in detail in combination with Step S301 - Step S305:

[0132] Step S301: Calculate the cosine similarity between the generated intermediate feature and each task-specific index embedding to obtain a similarity score.

[0133] The intermediate feature vector is extracted from the input 3D skeleton data and represents the features of the gesture. Each task-specific index embedding vector Kt Representation and task τ t Use cosine similarity to measure the relationship between f and each K t The cosine similarity reflects the similarity between f and K. t By calculating the cosine similarity between f and each task index embedding, the association between the current input gesture and each task can be measured. A higher similarity means that the current gesture is more likely to belong to the gesture category of the task.

[0134] Step S302: Calculate the above similarity scores for all task-specific index embeddings.

[0135] Assume that the current model has learned n tasks τ 0 ,τ 1 ,...,τ n , then each task has a task-specific index embedding K 0 ,K 1 ,...,K n The cosine similarity calculation in step S301 is extended to all task-specific index embeddings K i ∈K(index embedding pooling). By calculating the similarity scores for all tasks, the input gestures can be associated at the task level.

[0136] Step S303: Divide each similarity score by a predefined scaling factor to obtain a scaled similarity score.

[0137] The similarity scores are smoothed using a scaling factor τ. Smoothing can avoid numerical instability caused by extreme differences between similarity values ​​(such as close to 1 or close to 0). The scaled similarity scores make the subsequent exponential function calculation more stable.

[0138] Step S304: applying an exponential function to all scaled similarity scores to obtain unnormalized values ​​of weights.

[0139] Here, the exponential function can amplify the weight of high similarity scores while suppressing the weight of low similarity scores. The exponential function amplifies the influence of similarity and further distinguishes between tasks that are more likely to be associated with the current gesture and tasks that are less associated. The unnormalized weight is an intermediate result in generating the final combined weight.

[0140] Step S305: Normalize all unnormalized weight values ​​and calculate the combined weight of each task-specific index embedding vector.

[0141] All unnormalized weights are normalized to obtain the combined weight.

[0142] Normalization enables each weight to be interpreted as a probability value, representing the degree of association between the current input gesture and task τ i The combined weights provide a scaling factor for subsequent dynamic parameter combinations, enabling the model to dynamically activate task-related parameters based on the characteristics of the current gesture.

[0143] In step S300, starting from the intermediate features of the input, the model generates the combined weights for each task using cosine similarity calculation and a dynamic weighting mechanism. These weights accurately reflect the degree of association between the current input gesture and each task, providing a basis for dynamically aggregating the isolated parameters of multiple tasks (step S400) in subsequent steps. This approach avoids the limitations of directly selecting a single task index embedding, fully utilizes the specific features of different tasks, and improves the flexibility and robustness of the model.

[0144] In this embodiment, as described in step S400 above, the core objective of step S400 is to dynamically aggregate the isolated parameters of multiple tasks based on the task combination weights generated in the previous step (step S300), thereby generating the final features for the current task according to the input gesture data. In this way, the model can achieve dynamic adaptation, not only being able to handle the current task but also effectively utilizing and integrating the knowledge of previous tasks. The following details steps S401 - S403.

[0145] Step S401: Apply the combined weights to the isolated parameters of each relevant task.

[0146] Each task has an independent isolated parameter module, and these modules store task-specific parameters through an adapter mechanism. According to the task combination weights calculated in step S300, the isolated parameters of each task are weighted.

[0147] The task weighting operation is formalized as: W i ·M i (f). Where W i is the combined weight of task τ t . By applying the weight W i to the task isolated parameter M i , the dynamic participation degree of different task parameters in the current gesture is achieved. Task parameters with higher weights will have a greater impact on the generation of the final features.

[0148] Step S402: Apply the weighted isolated parameters to the extracted intermediate features one by one.

[0149] For each task, apply the weighted isolated parameter module W i ·M iAct on the extracted intermediate features, and transform the input features through the isolation parameter module of each task to generate task-specific feature representations. The operations of the isolation parameter module can be processed in parallel, that is, the isolation parameters of all tasks act on the intermediate features simultaneously. Project the intermediate features into each task-specific subspace to obtain different task feature representations, ensuring that the model can make full use of the characteristics of the task isolation parameters.

[0150] Step S403: Accumulate the features after the action of each task to obtain the final feature for the current task.

[0151] Weightedly accumulate the output features of all tasks in step S402 to generate the final feature for the current task.

[0152] In this embodiment, as described in step S500 above, in step S500, the model inputs the final feature generated by step S400 into a pre-trained classifier module for predicting the gesture category. The classifier module first performs a linear transformation on the final feature using a fully connected layer, mapping it from the high-dimensional feature space to the dimension corresponding to the number of gesture categories.

[0153] In this embodiment, as described in step S600 above, in step S600, the model trains the adapter module of each task according to the prediction result to further optimize the task-specific isolation parameters. Specifically, in the current task training stage, the model only updates the adapter module related to the current task, while the adapter parameters of other tasks remain frozen. Ensure that the knowledge of the current task can be efficiently learned through the adapter module, while avoiding interference with the historical task parameters, thus alleviating the catastrophic forgetting problem. In addition, through task-specific training, the model can strengthen the gesture feature representation of the current task, improve the classification performance and maintain an efficient parameter isolation mechanism.

[0154] The following elaborates on the specific implementation process of the PENCIL method in detail in combination with experiments, including experimental steps, key technical points, experimental settings, performance comparison, etc.

[0155] I. Experimental Datasets and Benchmarks:

[0156] 1. Experimental Datasets: Two public datasets are used in this experiment: SHREC-2017 and EgoGesture3D.

[0157] SHREC-2017: It contains 14 gesture categories, captured using an Intel RealSense depth camera, including 22 hand key points (1 wrist, 1 palm, and 4 key points for each finger). The training set contains 1980 samples, and the test set contains 840 samples.

[0158] EgoGesture3D: It contains 83 single - hand and two - hand gesture categories, with a total of 14416 training samples, 4768 validation samples, and 4977 test samples. It is derived from the EgoGesture dataset, and 3D skeletons are extracted using MediaPipe.

[0159] 2. Baseline methods: To verify the effectiveness of the PENCIL method, it is compared with the following methods: Oracle, Base, Fine - tuning, Feature Extraction, LWF, LWFMC, DeepInversion, ABD, R - DFCIL, BOAT - MI (one of the current state - of - the - art methods).

[0160] II. Experimental Settings and Task Details:

[0161] 1. Experimental task division: Using the experimental setting of class - incremental learning, the gesture recognition task is divided into a base task (Task 0) and subsequent incremental tasks (Tasks 1 to 6).

[0162] SHREC - 2017: Task 0 contains 8 gesture classes, and 1 new gesture class is added each time for Tasks 1 to 6;

[0163] EgoGesture3D: Task 0 contains 59 gesture classes, and 4 new gesture classes are added each time for Tasks 1 to 6.

[0164] 2. Model training details:

[0165] Optimizer: Consistent with BOAT - MI, the Adam optimizer is used;

[0166] Learning rate: Set to 6e - 5 for SHREC - 2017 and 3e - 5 for EgoGesture3D;

[0167] Hardware environment: All experiments are conducted on an Nvidia 3090 24GB GPU.

[0168] 3. Performance evaluation metrics:

[0169] Global accuracy (ACC↑): Measures the overall accuracy of gesture recognition;

[0170] Performance degradation rate (PD↓): Reflects the impact of new tasks on the recognition accuracy of historical categories.

[0171] III. Experimental Steps of the PENCIL Method:

[0172] 1. Task 0 training: The initial model is trained on the base - class dataset (Task 0), task - specific index embeddings are extracted, an adapter module is established, and an isolation parameter pool is generated.

[0173] 2. Task incremental learning: Each time a new gesture category is introduced, perform the following steps:

[0174] Extract the intermediate features of the new task; use the task-specific combination mechanism (TCM) to calculate the combination weights according to the similarity between the intermediate features and the index embedding pool, and dynamically aggregate the isolation parameters; use prototype-enhanced hybrid regularization (PHR) to optimize the index embedding and model parameters through pseudo-sample generation and knowledge transfer; while predicting the new category in the classifier, maintain the recognition ability for historical categories.

[0175] 3. Ablation experiment: Analyze the roles of TCM and PHR: Remove TCM and use simple parameter averaging. The results show a significant decrease in accuracy; remove PHR, and the results show that the model's ability to distinguish similar categories is weakened.

[0176] IV. Experimental Results and Analysis:

[0177] 1. Comparison of recognition accuracies:

[0178] The experimental results are shown in Tables 1 and 2. Tables 1 and 2 show the performance of PENCIL on the SHREC-2017 and EgoGesture3D datasets;

[0179] SHREC-2017: Surpass the BOAT-MI method by 13.33% in Task 6, and the average accuracy is increased to 79.89%;

[0180] EgoGesture3D: Surpass the BOAT-MI method by 22.43% in Task 6, and the average accuracy reaches 69.33%.

[0181] Table 1: Class incremental learning results for six tasks in SHREC-2017

[0182]

[0183]

[0184] Table 2: Class incremental learning results for six tasks in EgoGesture3D

[0185]

[0186] 2. Computational efficiency: PENCIL is significantly superior to BOAT-MI in computational efficiency. The experimental results are shown in Table 3:

[0187] Table 3: Overall training time and memory overhead

[0188]

[0189] Table 3 shows the comparison of time and memory overhead between PENCIL and BOAT-MI;

[0190] SHREC-2017: The processing time is reduced by 93.9% and the memory usage is reduced by 93.2%;

[0191] EgoGesture3D: The processing time is reduced by 92.8% and the memory usage is reduced by 97%;

[0192] The results show the deployment advantages of PENCIL on edge devices.

[0193] 3. Ablation study:

[0194] Effectiveness of the task-specific combination mechanism (TCM): The task-specific combination mechanism significantly improves the recognition ability of task-irrelevant samples; Table 4 shows the effectiveness of TCM, and the performance drops significantly after removing TCM:

[0195] Improvement of prototype-enhanced hybrid regularization (PHR): By introducing pseudo-samples, the confusion between similar categories is significantly reduced, and the model stability is improved. Figure 2 The ablation experiment results of PHR are shown, and the performance drops after removing the PEE or PPE module.

[0196] Table 4: Ablation study of different aggregation methods

[0197]

[0198] V. Generality and alternatives of the method:

[0199] 1. Compatibility of the adapter module: As shown in Table 5, three adapter structures (convolutional adapter, compact adapter, bottleneck adapter) are experimentally tested, and PENCIL performs excellently in all structures, proving its model-irrelevance;

[0200] 2. Pseudo-sample generation method: The prototype pseudo-sample generation can be extended to a method based on the generative adversarial network (GAN) to further enhance the diversity and authenticity of pseudo-samples;

[0201] 3. Multi-task extension: For more complex task sequences, a hierarchical index embedding pool can be designed to improve the efficiency and accuracy under multi-task conditions.

[0202] Table 5: Ablation study of different adapter methods

[0203]

[0204] Figure 3Shows the basic process and performance comparison of the class-incremental gesture recognition (CI-HGR) task. Part (a) illustrates the characteristics of the CI-HGR task, that is, the model is first initially trained on the basic category gestures (Task 0), and then sequentially learns new gesture categories (Tasks 1 to N), during which samples of previous tasks cannot be accessed. Part (b) experimentally compares the performance of the proposed PENCIL method with the existing state-of-the-art method BOAT-MI on the SHREC-2017 dataset. The results show that the PENCIL method is significantly superior to BOAT-MI in terms of global accuracy, and its performance degradation rate is lower, verifying the significant advantages of the present invention in solving catastrophic forgetting and improving learning efficiency.

[0205] Figure 4 Is a conceptual diagram of the working mechanism of the proposed PENCIL method, including two core components: (1) Task-Specific Compositional Mechanism (TCM), which is used to dynamically combine task-specific isolation parameters to adapt to different gesture inputs, and at the same time solves the problem of unknowable task information in the class-incremental scenario through adaptive parameter adjustment; (2) Prototype-Enhanced Hybrid Regularization (PHR), which uses prototype pseudo-samples to optimize the performance of TCM from two aspects: one is to improve the accuracy of TCM's matching of pseudo-samples through prototype-driven embedding enhancement (PEE), and the other is to achieve the transfer of historical model knowledge to the current model through prototype-driven parameter enhancement (PPE) to avoid excessive parameter adjustment. The entire mechanism significantly improves the recognition accuracy and model stability in the class-incremental gesture recognition task through the synergistic effect of TCM and PHR.

[0206] VI. Summary:

[0207] The experimental results show that the PENCIL method significantly outperforms the existing state-of-the-art technologies on two large datasets, and performs excellently in terms of accuracy, performance degradation rate, and computational efficiency. By combining the task-specific compositional mechanism and prototype-enhanced hybrid regularization, PENCIL effectively solves the problems of catastrophic forgetting and recognition confusion in class-incremental learning, providing an efficient and reliable solution for the gesture recognition task.

[0208] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a processor and a memory, and the processor is used to implement the steps of the method when executing the computer program stored in the memory.

[0209] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0210] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0211] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0212] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. And the aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs and other various media that can store program codes.

[0213] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A quasi-incremental gesture recognition method, characterized in that: The steps include: Receiving input data representing a gesture, wherein the input data is a three-dimensional skeleton data including a time series, each frame of the three-dimensional skeleton data including spatial coordinates of a plurality of hand joints; extracting intermediate features according to the received input data, and constructing task-specific index embeddings based on the intermediate features, wherein each index embedding is associated with a specific gesture recognition task; Calculate the similarity between the extracted intermediate features and the task-specific index embedding to generate the combination weight; Use the generated combined weights to dynamically aggregate the isolated parameters of multiple tasks to construct the final features for the current task; The constructed final features are input into the classifier to perform gesture category prediction; Based on the prediction results, the adapter module is trained for each task; During the adapter module training process, prototypes of historical gesture categories are calculated and prototype pseudo samples are generated. The prototype pseudo samples are used to embed task-specific indexes and enhance model parameters. Combine index embedding enhancement and model parameter enhancement to form an overall prototype enhancement hybrid regularization loss and continue to adjust the model; After all task training is completed, the adjusted model is used to classify and recognize new input gesture data.

2. The incremental gesture recognition method according to claim 1, characterized in that: Methods for constructing task-specific index embeddings based on this intermediate feature include: Extracting three-dimensional spatial coordinate features of each frame of hand joints from the input data, wherein the three-dimensional spatial coordinate features include the position of each hand joint in three-dimensional space; Extracting time series features from the three-dimensional space coordinate features using a time series modeling method to generate time series features representing dynamic changes in gestures; The temporal features and the three-dimensional spatial features are fused through a feature fusion network to generate intermediate features of the gesture; Based on the intermediate features, generating a task-specific index embedding vector associated with the current task, wherein the task-specific index embedding vector is constructed by a task-specific trainable module; The task-specific index embedding vector is dynamically combined with the intermediate features to form a task-specific index embedding for subsequent task learning.

3. The incremental gesture recognition method according to claim 2, characterized in that: The similarity between the extracted intermediate features and the task-specific index embedding is calculated, and the method of generating the combination weights includes: Calculate the cosine similarity between the generated intermediate features and each task-specific index embedding to obtain a similarity score; Compute the above similarity scores for all task-specific index embeddings; Divide each similarity score by a predefined scaling factor to obtain a scaled similarity score; Apply an exponential function to all scaled similarity scores to obtain the unnormalized value of the weight; All unnormalized weight values ​​are normalized and the combined weight of each task-specific index embedding vector is calculated as follows: Among them, W t is the combined weight, exp is the exponential function, cos(f, K t ) is the intermediate feature representation f and the task-specific index embedding vector K t The cosine similarity between them, K is the index embedding pool, which contains the set of all task-specific index embedding vectors, K j is the f-th task-specific index embedding vector in the set K, f represents the intermediate feature vector extracted from the input gesture data by the feature extractor, τ is the scaling factor, For similarity.

4. The incremental gesture recognition method according to claim 1, characterized in that: Dynamic aggregation includes weighted combination of the isolation parameters of the multiple tasks according to the combination weight, specifically including: applying the combined weights to the isolated parameters of each relevant task; Apply the weighted isolation parameters to the extracted intermediate features one by one; The features after each task are added up to obtain the final features for the current task.

5. The incremental gesture recognition method according to claim 1, characterized in that: Gesture category prediction includes the following specific steps: The constructed final features are input into the pre-trained classifier module; In the classifier module, the final features are linearly transformed using a fully connected layer; By applying the Softmax function, the linearly transformed output is converted into the probability distribution of each gesture category; According to the probability distribution, the category with the highest probability is selected as the final gesture recognition result.

6. The incremental gesture recognition method according to claim 1, characterized in that: During the adapter module training process, the method of calculating the prototype of the historical gesture category and generating the prototype pseudo sample includes: Extract features belonging to the historical gesture category through the feature extractor of the gesture recognition model; According to the gesture samples belonging to the same category, the prototype and covariance matrix of the category are calculated. The specific formula is as follows: Among them, μ c is the prototype of category c, N c is the number of samples of category c, y i =c represents sample x i The label belongs to category c, Indicates input gesture x i Through feature extractor The intermediate features obtained later; is the covariance matrix of category c, T is the transpose operator; Based on the calculated prototype and covariance matrix, a prototype pseudo sample is generated for each historical gesture category, and the prototype pseudo sample follows a multivariate Gaussian distribution: in, is the generated prototype pseudo sample feature vector, representing the virtual sample of category c, is a multivariate Gaussian distribution with mean μ c , the covariance matrix is c∈{C K \C U } is the category c that belongs to the learned historical category set C K Subtract the new category set C introduced by the current task U .

7. The incremental gesture recognition method according to claim 1, characterized in that: Methods for augmenting task-specific index embeddings with prototype pseudo samples include: A metric learning strategy is adopted to take the prototype pseudo sample as a negative sample, and compare it with the intermediate features of the current task to optimize the indexable embedding; The task-specific index embedding is updated by maximizing the similarity between the original sample features and the correct task embedding and minimizing its similarity with the pseudo sample features. The task-specific index embedding is optimized using the following loss function: in, is a prototype-driven embedding enhancement loss function used to measure and optimize task-specific index embeddings. m is a boundary parameter used to control the minimum distance between different categories. Prototype pseudo sample The cosine similarity between the task-specific index embedding K, cos(f,K) is the current gesture feature Embedding with task-specific indexes K The cosine similarity between .

8. The incremental gesture recognition method according to claim 1, characterized in that: Methods for enhancing model parameters using prototype pseudo samples include: According to the prototype pseudo samples of the historical task category and the characteristics of the training samples of the current task, the pseudo samples are used as an intermediary to transfer knowledge; Design a loss function that aligns the category probabilities of the current task model and the historical task model; The model parameters are optimized using the following loss function: in, and Represent the predictions of the current task model and the previous task model for the pseudo sample input, is the classifier for the current task, is the intermediate feature representation of the input gesture, The combined weights calculated for the indexable embeddings, is the adapter module pair feature of the i-th task The transformation output of is the predicted probability distribution after the current task model is aligned, is the predicted probability distribution of the current task model on the learned category set, C K is the learned category set, C U Add a new category set to the current task. It is a normalization operation, which means normalizing the historical category probabilities; is the loss function for prototype-driven embedding enhancement, which is used to measure and optimize task-specific index embedding. KL represents the Kullback-Leibler divergence function. γ is the weight coefficient of the loss function, which is used to balance the loss value. CE(·) is the task loss function for gesture recognition tasks. is the label corresponding to the class prototype pseudo sample.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the processor is used to implement the steps of the quasi-incremental gesture recognition method as claimed in any one of claims 1 to 8 when executing a computer program stored in the memory.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the incremental gesture recognition method according to any one of claims 1 to 8 are executed.

Citation Information

Patent Citations

  • Small sample image increment classification method and device based on embedding enhancement and self-adaption

    CN114549894A

  • Non-paradigm class incremental learning action recognition method and device based on self-supervised learning

    CN117912118A

Cited By

  • Fine-grained gesture recognition method

    CN121350761A