Video behavior recognition model training method, video behavior recognition method and device
By adopting the incremental learning method of slow, fast dynamic reparameterization module in the video behavior recognition model training, combined with deep learning and analytical machine learning, the problem of high accuracy of video behavior recognition in the existing technology is solved, and higher recognition accuracy and model adaptability are achieved.
Patent Information
- Application Number
- CN202510284671.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to achieve high-accurate video behavior recognition results while fully protecting data privacy.
A video behavior recognition model training method is adopted, which includes basic training of the initial model on the basic task, retaining the parameters of the backbone network and bottleneck layer, replacing the linear classification head with a multi-layer perceptron parsing head, and iterative training on K incremental tasks. This method performs incremental learning through slow and fast dynamic re-parameterization module, combining deep learning and analytical machine learning, generates fusion weights and outputs prediction results.
On the premise of fully protecting data privacy, higher video behavior recognition accuracy is achieved, reducing forgetting problems, and improving the robustness and adaptability of the model in long-term incremental tasks.
Smart Images

Figure CN120220022A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to a video behavior recognition method, and particularly relates to a video behavior recognition model training method, a video behavior recognition method and a device. Background Art
[0002] Video action recognition is a crucial computer vision task and is widely applicable to various application fields such as human-computer interaction, security, healthcare, social media, and entertainment. As new actions continue to emerge, the model must learn and recognize them without forgetting previous knowledge. The mainstream methods in the prior art attempt to balance the plasticity of learning new knowledge and the stability of retaining old knowledge by storing historical action data. However, due to the increasing requirements for data privacy and concerns about storage limitations, the example-based methods are not feasible in real-world scenarios.
[0003] To solve this problem, example-free continuous video action recognition is introduced, which focuses on mitigating forgetting in the context of model updates. Example-free Continuous Video Action Recognition (EFCVAR) is a technology for video understanding, aiming to automatically identify and classify actions or activities from videos without relying on traditional labels or example data.
[0004] It can be seen that it is difficult for the prior art to obtain relatively accurate video behavior recognition results under the premise of fully protecting data privacy. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that the above-mentioned methods commonly used in the prior art are difficult to obtain relatively accurate video behavior recognition results under the premise of fully protecting data privacy. To solve the above problem, the present invention provides a video behavior recognition model training method, a video behavior recognition method and a device.
[0006] The content of the present invention includes:
[0007] In a first aspect, an embodiment of the present invention provides a video behavior recognition model training method, including:
[0008] Performing basic training on an initial video behavior recognition model for a basic task, and obtaining a trained initial video behavior recognition model by minimizing cross-entropy loss, where the initial video behavior recognition model includes a backbone network, a bottleneck layer, a linear classification head, and a slow-fast dynamic reparameterization module;
[0009] Retaining the model parameters of the backbone network and the bottleneck layer, and replacing the linear classification head with a randomly initialized multi-layer perceptron parsing head to obtain a basic video behavior recognition model;
[0010] Iteratively train the basic video behavior recognition model on K incremental tasks to obtain a target video behavior recognition model, where K is a positive integer;
[0011] Among them, the slow-fast dynamic reparameterization module includes a slow branch, a fast branch, a streaming discriminator, and a reparameterization branch. The slow branch is used for incremental learning based on deep learning, the fast branch is used for incremental learning based on analytical machine learning, the streaming discriminator is used to generate fusion weights, and the reparameterization branch is used to output a prediction result based on target features, where the target features are obtained by fusing the output of the slow branch and the output of the fast branch based on the fusion weights.
[0012] Optionally, the iteratively training the basic video behavior recognition model on K incremental tasks to obtain a target video behavior recognition model includes:
[0013] Train the basic video behavior recognition model on the k-th incremental task to obtain the k-th model, where k = 1, 2,..., K;
[0014] Among them, the training the basic video behavior recognition model on the k-th incremental task to obtain the k-th model includes:
[0015] Input the training data corresponding to the k-th incremental task into the backbone network for processing to obtain a backbone feature set;
[0016] Extract the means and variances corresponding to a preset number of old-class features from the prototype feature library to form an old-class feature set, where the old-class features are obtained by using the slow-fast dynamic reparameterization module to train the basic video behavior recognition model on the previous k - 1 incremental tasks;
[0017] Combine the old-class feature set and the backbone feature set to form joint training data;
[0018] Based on the joint training data, use the slow-fast dynamic reparameterization module to train the (k - 1)-th model to obtain the k-th model.
[0019] Optionally, before the extracting the means and variances corresponding to a preset number of old-class features from the prototype feature library to form an old-class feature set, the method further includes:
[0020] Obtain the mean μ i and variance v i :
[0021]
[0022] Among them, Flib,i is the old class feature of the i-th class, is the backbone feature set obtained by the backbone network through inference on the training data corresponding to the basic task, represents the mean feature corresponding to the old class feature of the i-th class, represents the variance corresponding to the old class feature of the i-th class, C represents the feature dimension, and Index(·) represents the indexing operation;
[0023] Construct the prototype feature library, where the prototype feature library includes the mean set U k and the variance set V k :
[0024]
[0025] where E t represents the number of action categories of the t-th incremental task.
[0026] Optionally, training the (k - 1)-th model using the slow-fast dynamic reparameterization module based on the joint training data to obtain the k-th model includes:
[0027] Input the joint training data into the slow branch and the fast branch respectively to obtain the output of the slow branch and the output of the fast branch;
[0028] Input the joint training data into the streaming discriminator to obtain the fusion weight;
[0029] Fuse the output of the fast branch and the output of the slow branch based on the fusion weight and then input them into the reparameterization branch to obtain the prediction result;
[0030] Calculate the prediction loss based on the prediction result and the class label, and update the parameters of the slow branch based on the prediction loss.
[0031] Optionally, training the basic video behavior recognition model on the k-th incremental task to obtain the k-th model further includes:
[0032] Input the training data corresponding to the k-th incremental task into the backbone network and the bottleneck layer for feature extraction, and then perform upsampling to obtain the upsampled feature
[0033] Based on the upsampled feature calculate the optimal weight of the downsampling layer corresponding to the k-th incremental task
[0034]
[0035] Among them, is the optimal weight of the downsampling layer corresponding to the (k - 1)-th incremental task, and R k-1 is calculated for the (k - 1)-th incremental task, is the action label corresponding to the k-th incremental task, and I is the identity matrix.
[0036] In a second aspect, an embodiment of the present invention further provides a video behavior recognition method, including:
[0037] Obtain a target video;
[0038] Input the target video into a target video behavior recognition model for video behavior recognition processing to obtain an action recognition result, where the target video behavior recognition model is trained by using the video behavior recognition model training method as described in the first aspect.
[0039] In a third aspect, an embodiment of the present invention provides a video behavior recognition model training device, including:
[0040] A basic training module for performing basic training on an initial video behavior recognition model on a basic task, and obtaining a trained initial video behavior recognition model by minimizing the cross-entropy loss. The initial video behavior recognition model includes a backbone network, a bottleneck layer, a linear classification head, and a slow-fast dynamic reparameterization module;
[0041] A processing module for retaining the model parameters of the backbone network and the bottleneck layer, and replacing the linear classification head with a randomly initialized multi-layer perceptron parsing head to obtain a basic video behavior recognition model;
[0042] An incremental training module for performing iterative training on the basic video behavior recognition model on K incremental tasks to obtain a target video behavior recognition model, where K is a positive integer;
[0043] Among them, the slow-fast dynamic reparameterization module includes a slow branch, a fast branch, a streaming discriminator, and a reparameterization branch. The slow branch is used for incremental learning based on deep learning, the fast branch is used for incremental learning based on analytical machine learning, the streaming discriminator is used to generate a fusion weight, and the reparameterization branch is used to output a prediction result based on a target feature, where the target feature is obtained by fusing the output of the slow branch and the output of the fast branch based on the fusion weight.
[0044] In a fourth aspect, an embodiment of the present invention provides a video behavior recognition device, including:
[0045] An acquisition module for acquiring a target video;
[0046] A video behavior recognition module, which is configured to input the target video into a target video behavior recognition model for video behavior recognition processing to obtain an action recognition result, where the target video behavior recognition model is trained by using the video behavior recognition model training method described in the first aspect.
[0047] In a fifth aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a program stored on the memory and executable on the processor; the processor is configured to read the program in the memory to implement the video behavior recognition model training method described in the first aspect, or to implement the steps in the video behavior recognition method described in the second aspect.
[0048] In a sixth aspect, an embodiment of the present invention provides a readable storage medium for storing a program, where the program, when executed by a processor, implements the steps in the video behavior recognition model training method described in the first aspect, or implements the steps in the video behavior recognition method described in the second aspect.
[0049] The beneficial effects of the present invention are as follows. In the embodiments of the present invention, the model is first trained on a basic task, and then trained on an incremental task. A slow-fast dynamic reparameterization module is provided, which includes a slow branch, a fast branch, a streaming discriminator, and a reparameterization branch. The slow branch is used for incremental learning based on deep learning, the fast branch is used for incremental learning based on analytical machine learning, the streaming discriminator is used to generate a fusion weight, and the reparameterization branch is used to output a prediction result based on a target feature, where the target feature is obtained by fusing the output of the slow branch and the output of the fast branch based on the fusion weight. On the one hand, the slow branch uses gradient backpropagation (i.e., deep learning) to gradually learn new knowledge. On the other hand, the fast branch uses analytical machine learning (i.e., non-gradient machine learning) to quickly consolidate old knowledge and provide an analytical (i.e., closed-form) linear solution. The method provided by the embodiments of the present invention can achieve higher accuracy on the premise of fully protecting data privacy. Description of the Drawings
[0050] Appendix Figure 1 is a flowchart of the video behavior recognition model training method provided by an embodiment of the present invention;
[0051] Appendix Figure 2 is a schematic flowchart of training on a basic task;
[0052] Appendix Figure 3 is a schematic flowchart of training on an incremental task;
[0053] Appendix Figure 4 is a specific schematic diagram of the slow-fast dynamic reparameterization module;
[0054] AppendixFigure 5 It is a schematic flow chart of the Gaussian memory synthesis stage;
[0055] Attached Figure 6 It is a schematic flow chart of the dual knowledge distillation stage;
[0056] Attached Figure 7 It is a schematic diagram of the video behavior recognition model training device provided by an embodiment of the present invention;
[0057] Attached Figure 8 It is a flow chart of the video behavior recognition method provided by an embodiment of the present invention;
[0058] Attached Figure 9 It is a schematic diagram of the video behavior recognition device provided by an embodiment of the present invention;
[0059] Attached Figure 10 It is a schematic diagram of the structure of the electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0060] In the embodiments of the present application, the term "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In the embodiments of the present application, the term "plurality" refers to two or more, and other quantifiers are similar. The terms "first", "second", etc. in the description and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same type, and the number of objects is not limited. For example, the first object may be one or multiple.
[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0062] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0063] An embodiment of the present application provides a method for training a video behavior recognition model, a video behavior recognition method, and a device. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the method for training a video behavior recognition model provided by an embodiment of the present invention. The method specifically includes the following steps:
[0064] Step 101: Perform basic training on the initial video behavior recognition model on a basic task, and obtain the trained initial video behavior recognition model by minimizing the cross-entropy loss. The initial video behavior recognition model includes a backbone network, a bottleneck layer, a linear classification head, and a slow-fast dynamic reparameterization module.
[0065] Step 102: Retain the model parameters of the backbone network and the bottleneck layer, and replace the linear classification head with a randomly initialized multi-layer perceptron parsing head to obtain a basic video behavior recognition model.
[0066] Step 103: Perform iterative training on the basic video behavior recognition model on K incremental tasks to obtain a target video behavior recognition model, where K is a positive integer.
[0067] Among them, the slow-fast dynamic reparameterization module includes a slow branch, a fast branch, a streaming discriminator, and a reparameterization branch. The slow branch is used for incremental learning based on deep learning, the fast branch is used for incremental learning based on analytical machine learning, the streaming discriminator is used to generate fusion weights, and the reparameterization branch is used to output a prediction result based on target features. The target features are obtained by fusing the outputs of the slow branch and the fast branch based on the fusion weights.
[0068] Optionally, in some embodiments, the slow branch is used for incremental learning based on Back Propagation (BP), and the fast branch is used for incremental learning based on Analytic Learning (AL). In Back Propagation (BP) (i.e., deep learning), the task recency bias causes the network to prioritize the most recent tasks and adjust the parameters to minimize the loss of the latest data, which leads the network to be more inclined towards new classes rather than old classes, resulting in catastrophic forgetting. In the absence of past data, it becomes difficult to maintain the model accuracy. Analytic Learning (AL) (i.e., analytic machine learning) uses the pseudo-inverse technique to effectively mitigate catastrophic forgetting by maintaining the representational representation of past knowledge without storing past data. It uses a fixed feature extractor and recursive least squares method to preserve previous information. This static approach has significant limitations. The rigidity of the fixed feature extractor restricts the model's ability to adapt to new classes because it cannot dynamically adjust the representation according to new data. This limitation of analytic learning forms a complementary situation with backpropagation-driven deep learning. Analytic learning provides stability, and deep learning provides adaptability. It should be understood that the structures of the fast branch and the slow branch are identical, but the parameters of their respective network layers are not necessarily the same. Therefore, the fast branch and the slow branch can be reparameterized and fused.
[0069] In Video Action Recognition (VAR), Class-Incremental Learning (CIL) aims to address the learning requirements for an increasing number of new category data without retraining the entire model. In this embodiment, a task sequence T = {T0, T1, T2, …, T K} is given, where K represents the number of incremental tasks. Task T0 is the base task, and T 1:K are incremental tasks. Compared with incremental tasks, the base task usually consists of a larger number of action categories to help the model establish a basic understanding of video action patterns.
[0070] For each task consists of all the videos belonging to the action class O k and the corresponding action labels . represents the number of sample tasks, represents the number of samples for each action category, part ∈ {train, test} is the dataset split, and E k = len(O k ) represents the task T kThe number of middle action categories, and the action categories of are disjoint, denoted as
[0071] In the video action recognition category incremental learning method, the backbone network is usually pre-trained on the base task T0, as Figure 2 shown. In this embodiment, no additional samples are required to store the knowledge of the previous categories. Most of the parameters of the initial video behavior recognition model are pre-trained using the training data of the base task, and fine-tuning is avoided in subsequent incremental tasks.
[0072] In this embodiment, the initial video behavior recognition model includes a backbone network, a bottleneck layer, and a linear classification head. In some embodiments, as Figure 2 shown, a Temporal Shift Module (TSM) is set in the backbone network to enhance the temporal modeling ability.
[0073] Let represent the video samples in the training data corresponding to the base task, where T, W, and H represent the number of frames and the spatial resolution respectively. The specific process of pre-training the initial video behavior recognition model on the base task is as follows:
[0074] Input the video sample X in the training data corresponding to the base task into the backbone network for feature extraction to obtain the backbone feature F BA , and the above process can be expressed as: F BA =Backbone(θ BA ,X);
[0075] Input the backbone feature F BA into the bottleneck layer for feature extraction to obtain the bottleneck feature F BN , and the above process can be expressed as: F BN =BottleNeck(θ BN ,F BA );
[0076] Input the bottleneck feature F BN into the linear classification head for classification processing to obtain the predicted classification probability The above process can be expressed as:
[0077] where, is the backbone feature, and the backbone feature can also be called the global feature, C is the number of feature channels, is the output of the bottleneck layer, represents the predicted classification probability. θ BA 、θ BN and θH Parameters for characterizing the backbone network, bottleneck layer, and linear classification head respectively. In step 101, the initial video action recognition model is trained and evaluated using the training data corresponding to the basic task. By minimizing the cross-entropy loss, the final optimal parameters θ BA , θ BN and θ H , where θ H is not used in the subsequent steps.
[0078] In step 102, the model parameters of the backbone network and bottleneck layer in the trained initial video action recognition model are retained, and the linear classification head is replaced with a randomly initialized two-layer Multilayer Perceptron (MLP) parsing head to obtain the basic video action recognition model. Parsing learning requires two layers of MLP, while the existing linear classification head is one layer of MLP. By replacing the linear classification head with a two-layer MLP parsing head, the current network structure can be adjusted to the structural paradigm of parsing learning. The two-layer MLP parsing head includes one layer of MLP for dimension increase and one layer of MLP for dimension decrease. The single layer among them is used for the subsequent calculation process description. Among them, the hidden dimension of the two-layer MLP parsing head is C H , C H > C.
[0079] In some embodiments, after step 102, the method further includes performing a parsing realignment operation, please refer to Figure 2 , and the specific process is as follows:
[0080] Performing inference on all the training data corresponding to the basic tasks to obtain the backbone feature set and the bottleneck feature set
[0081] First, the bottleneck features are upsampled to the hidden dimension through the MLP for dimension increase (i.e., the upsampling layer) in the two-layer MLP parsing head to obtain the upsampled features
[0082]
[0083] Among them, is used to represent the number of samples for each action category in the basic task, Linear(·) is used to represent the linear layer, ReLU(·) represents the activation function, represents the weight of the upsampling layer.
[0084] By optimizing the parameters of the MLP for dimension decrease (i.e., the downsampling layer) in the two-layer MLP parsing head as follows, the parsing classification result can be obtained:
[0085]
[0086] Among them, represents the weights of the downsampling layer, is the optimal weight of the downsampling layer, ‖·‖F is the Frobenius form, and η is the regularization term. is the action label corresponding to the base task.
[0087] The optimal weights of the downsampling layer are as follows:
[0088]
[0089] Among them, I represents the identity matrix.
[0090] According to the above equation, it can be seen that the optimal weights of the downsampling layer still depend on the old-class data. To get rid of the dependence on the old-class data, optionally, in some embodiments, after the training data corresponding to the kth incremental task is input into the backbone network and the bottleneck layer for feature extraction, it is upsampled to obtain the upsampled features. Let:
[0091]
[0092] Therefore, the optimal weights of the downsampling layer corresponding to the kth incremental task can be obtained.
[0093]
[0094] Among them, R k can only be updated by the current task data to:
[0095]
[0096] Among them, is the optimal weight of the downsampling layer corresponding to the (k - 1)th incremental task, and R k-1 is calculated for the (k - 1)th incremental task. is the action label corresponding to the kth incremental task.
[0097] Through the above method, in analytical learning, each increment needs to calculate an optimal parameter of the downsampling linear layer to meet the increasing number of classes. The optimal weights calculated in the k-step increment will be used for the classification in the k-step and the weight initialization in the (k + 1)-step increment. The calculation of the optimal weights of the downsampling layer can get rid of the dependence on the old-class data, and the optimal weights of each task can be obtained only with the training data corresponding to itself, without the need to restore historical data, thus realizing a sample-free method.
[0098] In step 103, the basic video behavior recognition model is iteratively trained on K incremental tasks to obtain the target video behavior recognition model. For details, please refer to Figure 3 and Figure 4 . Optionally, in some embodiments, step 103 includes:
[0099] Training the basic video behavior recognition model on the k-th incremental task to obtain the k-th model, where k is a positive integer less than or equal to K;
[0100] Among them, training the basic video behavior recognition model on the k-th incremental task to obtain the k-th model includes:
[0101] Inputting the training data corresponding to the k-th incremental task into the backbone network for processing to obtain a backbone feature set;
[0102] Extracting the means and variances corresponding to a preset number of old-class features from the prototype feature library to form an old-class feature set, where the old-class features are obtained by training the basic video behavior recognition model on the previous k-1 incremental tasks using the slow-fast dynamic reparameterization module;
[0103] Combining the old-class feature set and the backbone feature set to form joint training data;
[0104] Based on the joint training data, using the slow-fast dynamic reparameterization module to train the (k - 1)-th model to obtain the k-th model.
[0105] It should be understood that in this embodiment, multiple model trainings are performed on a total of K incremental tasks, and the K-th model obtained by training on the K-th incremental task is the target video action recognition model. For the convenience of description, the following will take the k-th incremental task as an example for illustration. As Figure 4 shown, optionally, in some embodiments, the training the (k - 1)-th model using the slow-fast dynamic reparameterization module based on the joint training data to obtain the k-th model includes:
[0106] Inputting the joint training data into the slow branch and the fast branch respectively to obtain the output of the slow branch and the output of the fast branch;
[0107] Inputting the joint training data into the streaming discriminator to obtain a fusion weight;
[0108] Based on the fusion weight, fusing the output of the fast branch and the output of the slow branch and then inputting them into the reparameterization branch to obtain a prediction result;
[0109] Calculate the prediction loss based on the prediction result and the class label, and update the parameters of the slow branch based on the prediction loss.
[0110] During the initialization process, the weights of the slow branch are obtained as follows:
[0111]
[0112] Among them, is the parameter group of the bottleneck layer of the slow branch, is the weight matrix of the bottleneck layer of the slow branch, is the bias matrix of the bottleneck layer of the slow branch, is the parameter group of the upsampling layer of the slow branch, is the weight matrix of the upsampling layer of the slow branch, is the parameter group of the downsampling layer of the slow branch, is the weight matrix of the downsampling layer of the slow branch.
[0113] The streaming discriminator is used to dynamically control the fusion weights of the fast branch and the slow branch. The features output by the backbone network are first fed back into the streaming discriminator to obtain a reparameterized fusion weight to fuse the fast branch and the slow branch. Then, the fused features are sent to the reparameterized branch to obtain a loss to update the parameters of the slow branch. The forward process can be characterized as follows:
[0114] First, input F BA into the streaming discriminator. As Figure 4 shown, the streaming discriminator includes a Linear layer and a Sigmoid layer. The fusion weight can be output through the streaming discriminator. The above process can be characterized as:
[0115] a = Sigmoid(Linear(θ SD , F BA ));
[0116]
[0117] Among them, a is the fusion weight with a value between 0 and 1, θ SD represents the weight of the streaming discriminator, represents the weight matrix of the bottleneck layer of the fast branch, represents the bias matrix of the bottleneck layer of the fast branch, F BA represents the features output by the backbone network, represents the weight matrix of the upsampling layer of the fast branch, represents the output of the reparameterized network. It should be understood that in the embodiment, the backbone feature set is used for training, and for the output of the reparameterized network With action tags The cross entropy loss between them is minimized to obtain the initial streaming discriminator and the slow branch.
[0118] Re-parameterization training can cause differences between the various components of the model, resulting in unstable knowledge integration. To address this problem, a knowledge reflection mechanism is constructed in some embodiments. Gaussian memory synthesis is used to generate pseudo features for balanced data representation, and then knowledge distillation is used to align the outputs for bridging learning. The pseudo features approximate past knowledge and do not require data storage, thereby alleviating catastrophic forgetting.
[0119] The knowledge reflection mechanism is divided into two stages: Gaussian Memory Synthesis (GMS) and Dual Knowledge Distillation (DKD). In the GMS stage, a Gaussian distribution is constructed by extracting the first-order and second-order moments to capture the statistical characteristics of the actions from the current training stage. This distribution serves as an approximation of past knowledge, allowing pseudo features to be generated in subsequent learning stages without storing actual data. Subsequently, the DKD stage utilizes these pseudo features in the historical and current feature extractors, facilitating a two-level knowledge distillation process at the feature and logarithmic levels.
[0120] Specifically, Figure 5 As shown, the specific process of the Gaussian memory synthesis stage is as follows. In some embodiments, the mean μ corresponding to the old class feature of the i-th class is obtained. i and variance v i :
[0121]
[0122] Among them, F lib,i is the old class feature described in the i-th class, is the backbone feature set obtained by the backbone network through inference on the training data corresponding to the basic task, represents the mean feature corresponding to the old class feature of the i-th class, represents the variance corresponding to the old class feature of the i-th class, C represents the feature dimension, and Index(·) represents the index operation.
[0123] For the continuously arriving tasks t, the parameters θ of the downsampling layer AD,t It has been expanded and updated according to the above method to classify new categories, while the parameters of the slow branch are inherited from the previous task. The prototype feature library constructed based on the features of the old class (i.e., the class learned by the previous incremental task) is as follows, which includes U k and V k :
[0124]
[0125] Among them, E t represents the number of action categories included in task t.
[0126] As Figure 6 shown, in the dual knowledge distillation stage, the means and variances corresponding to a certain number of old-class features are extracted from the prototype feature library, and they are combined with the features of the new class to form new joint training data for the new incremental task. Among them, the features of the new class are the backbone feature sets obtained by inputting the training data corresponding to the k-th incremental task into the backbone network for processing
[0127] Extract a certain number of old-class features from the prototype feature library is:
[0128]
[0129] Joint training data is:
[0130]
[0131] Among them, i is the index of the old-class feature, j represents the average number of samples per class of the k-th incremental task, represents the number of training samples of the k-th incremental task, E k represents the total number of classes of the k-th task, and cat(·) represents the concatenation operation.
[0132] In this embodiment, a knowledge reflection mechanism is set up to solve the task-recency bias problem in the slow branch by generating pseudo-features. The pseudo-features approximate past knowledge and do not require data storage, thus reducing catastrophic forgetting. This process ensures the coherence of knowledge integration and reduces the difference between the fast branch and the slow branch.
[0133] As Figure 2 and Figure 3 shown, as a specific embodiment, the loss value Loss in the training process of the training method provided in this embodiment includes the prediction loss Loss Main , the feature-level distillation loss Loss Feat and the probability loss Loss Prob . The following will introduce them separately as follows:
[0134] The prediction loss is the cross-entropy between the predicted class probabilities and the action labels:
[0135]
[0136] Among them, is the prediction class probability, and Y is the action label.
[0137] As Figure 6 shown, in the knowledge reflection mechanism of this embodiment, a cross-task knowledge distillation strategy is introduced. The slow branch of the previous task is frozen as the teacher network, and the learnable reparameterized branch in the new task is guided to retain the old class knowledge by comparing the differences in bottleneck features and classification probabilities. The feature-level distillation loss Loss Feat is as follows:
[0138]
[0139] Among them, MSE(·,·) represents the mean square error, and F r is the feature sampled from , and w is the fusion weight.
[0140] The output of the teacher network only contains the old class probabilities, which does not match the shape of predicting new and old classes simultaneously with the output of the student network. To solve this problem, the probabilities of the new classes of the student network are normalized with the probabilities of the old classes to obtain the probability loss Loss Prob :
[0141]
[0142] Among them, KL(·) represents the KL divergence, [·:·] represents the list subset, and Prob k-1 is the classification probability predicted using the k-1 model, and Prob k is the classification probability predicted using the k model.
[0143] Therefore, the final loss is:
[0144] Loss = Loss Main + αLoss Feat + βLoss Prod ;
[0145] Among them, α is the weight corresponding to Loss Feat , and β is the weight corresponding to Loss Prob .
[0146] Taking a specific experiment as an example, the beneficial effects of the method provided by the embodiments of the present invention will be described. This experiment was conducted on three video action recognition datasets: UCF10, HMDB51, and SomethingSomethingV2. All classes were randomized using fixed random seeds (three for UCF101 and HMDB51, one for Something-SomethingV2), and then the class list was split into base classes and incremental classes. The base classes were used to initialize the model, while the incremental classes evaluated the performance of incremental learning.
[0147] In the UCF101 dataset, 51 classes were selected as the base classes, and the remaining 50 classes were divided into incremental tasks, with each task containing 10, 5, and 2 classes. In the HMDB51 dataset, 26 classes were used as the base classes, and the remaining 25 classes were divided into incremental tasks, with each task containing 5 and 1 class. For the Something-SomethingV2 dataset, 84 classes were used as the base classes, and the remaining 90 classes were divided into incremental tasks, with each task having 9, 3, and 1 class. This incremental setting is longer than that of UCF101 and HMDB51, providing a challenge for learning new classes while retaining the memory of the old classes.
[0148] To evaluate the performance of the training method provided by the embodiments of the present invention, the following three metrics were used: Average Incremental Accuracy (accuracy rate), forgetting rate, and Performance Dropping Rate (PD). The average incremental accuracy of all tasks was calculated in the experiment as the accuracy rate of the model on the test samples from task 0 to K at the end of each task t. The forgetting rate is the difference in the correct rate of the base tasks and the current task on the base classes. The performance dropping rate is defined as: PD = Acc1 - Acc, where Acc1 is the correct rate of task 1 and Acc is the correct rate of all tasks. Previous methods used samples for memory retention and the nearest mean of exemplars (NME) accuracy test during the test process. Since the training method provided by the embodiments of the present invention does not require samples, only the accuracy rate of the CNN in all tasks is reported.
[0149] The baseline methods in this experiment include the following commonly used existing technologies: LwFMC, LwM, iCaRL, UCIR, PODNet, TCD, SNRO, FrameMaker, HCE, and DBK. We also provided the lower bounds of the accuracy rate by training only using the new class data (Finetuning) and using all the previous data (Oracle). The experimental results are shown in Tables 1 and 2:
[0150] Experimental results on the UCF101 and HMDB51 datasets in Table 1
[0151]
[0152]
[0153] Experimental results on SomethingSomethingV2 in Table 2
[0154]
[0155] According to the results in Table 1 and Table 2, it can be seen that the accuracy of CNN and the accuracy of NME may vary on different datasets, and NME generally performs better on HMDB51. The method provided in this embodiment achieves the optimal results in all settings, demonstrating the ability of the method provided in this embodiment to learn new classes and retain the knowledge of old classes without relying on samples. In particular, in the more challenging long-term incremental tasks, the method provided in this embodiment has achieved more significant improvement effects.
[0156] The training method provided in this embodiment includes a slow-fast dynamic reparameterization module and a knowledge reflection mechanism. The slow-fast dynamic reparameterization module innovatively combines gradient-based and analytical learning techniques in a collaborative setting, and the knowledge reflection mechanism effectively unifies the knowledge integration process between the slow learning branch and the fast learning branch. The training method provided in the embodiment of the present invention, as a paradigm-free method, achieves an accuracy improvement of up to 11.52% compared with the existing methods in the case of using data replay, and reaches the state-of-the-art performance in all settings. In addition, in a long-term continuous task, the framework exceeds the baseline method by 30.68%, highlighting its robustness and adaptability.
[0157] As Figure 7 shown, the embodiment of the present invention also provides a video behavior recognition model training device 700, including:
[0158] A basic training module 701, configured to perform basic training on an initial video behavior recognition model for a basic task, and obtain a trained initial video behavior recognition model by minimizing the cross-entropy loss. The initial video behavior recognition model includes a backbone network, a bottleneck layer, a linear classification head, and a slow-fast dynamic reparameterization module;
[0159] A processing module 702, configured to retain the model parameters of the backbone network and the bottleneck layer, and replace the linear classification head with a randomly initialized multi-layer perceptron parsing head to obtain a basic video behavior recognition model;
[0160] An incremental training module 703, which is used to iteratively train the basic video behavior recognition model on K incremental tasks to obtain a target video behavior recognition model, where K is a positive integer;
[0161] Among them, the slow-fast dynamic reparameterization module includes a slow branch, a fast branch, a streaming discriminator, and a reparameterization branch. The slow branch is used for incremental learning based on deep learning, the fast branch is used for incremental learning based on analytical machine learning, the streaming discriminator is used to generate fusion weights, and the reparameterization branch is used to output a prediction result based on target features, where the target features are obtained by fusing the output of the slow branch and the output of the fast branch based on the fusion weights.
[0162] Optionally, the incremental training module 703 is specifically configured to:
[0163] Train the basic video behavior recognition model on the k-th incremental task to obtain the k-th model, where k = 1, 2,..., K;
[0164] Among them, training the basic video behavior recognition model on the k-th incremental task to obtain the k-th model includes:
[0165] Input the training data corresponding to the k-th incremental task into the backbone network for processing to obtain a backbone feature set;
[0166] Extract the means and variances corresponding to a preset number of old-class features from the prototype feature library to form an old-class feature set, where the old-class features are obtained by using the slow-fast dynamic reparameterization module to train the basic video behavior recognition model on the previous k-1 incremental tasks;
[0167] Form joint training data by combining the old-class feature set and the backbone feature set;
[0168] Based on the joint training data, use the slow-fast dynamic reparameterization module to train the (k - 1)-th model to obtain the k-th model.
[0169] Optionally, before extracting the means and variances corresponding to a preset number of old-class features from the prototype feature library to form an old-class feature set, the method further includes:
[0170] Obtain the mean μ i and variance v i :
[0171]
[0172] Among them, F lib,i is the i-th type of old-class feature, The backbone feature set obtained by the backbone network through inference on the training data corresponding to the basic task represents the mean feature corresponding to the i-th type of the old class features represents the variance corresponding to the i-th type of the old class features, C represents the feature dimension, and Index(·) represents the indexing operation;
[0173] Construct the prototype feature library, where the prototype feature library includes a mean set U k and a variance set V k :
[0174]
[0175] where, E t represents the number of action categories of the t-th incremental task
[0176] Optionally, training the (k - 1)-th model using the slow-fast dynamic reparameterization module based on the joint training data to obtain the k-th model includes:
[0177] Input the joint training data into the slow branch and the fast branch respectively to obtain the output of the slow branch and the output of the fast branch;
[0178] Input the joint training data into the flow discriminator to obtain the fusion weight;
[0179] Based on the fusion weight, fuse the output of the fast branch and the output of the slow branch and then input them into the reparameterization branch to obtain the prediction result;
[0180] Calculate the prediction loss based on the prediction result and the class label, and update the parameters of the slow branch based on the prediction loss
[0181] Optionally, training the basic video behavior recognition model on the k-th incremental task to obtain the k-th model further includes:
[0182] After inputting the training data corresponding to the k-th incremental task into the backbone network and the bottleneck layer for feature extraction, perform upsampling to obtain the upsampled feature
[0183] Based on the upsampled feature calculate the optimal weight of the downsampling layer corresponding to the k-th incremental task
[0184]
[0185] where, is the optimal weight of the downsampling layer corresponding to the (k - 1)-th incremental task, and R k-1 is calculated for the (k - 1)-th incremental task, is the action label corresponding to the k-th incremental task, and I is the identity matrix.
[0186] The video behavior recognition device 700 provided by the embodiments of the present application can execute the embodiments of the above-mentioned video behavior recognition model training method, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0187] As Figure 8 shown, the embodiments of the present invention also provide a video behavior recognition method, including:
[0188] Step 801, obtain a target video;
[0189] Step 802, input the target video into a target video behavior recognition model for video behavior recognition processing to obtain an action recognition result, where the target video behavior recognition model is trained by using the above-mentioned video behavior recognition model training method.
[0190] It should be understood that the target video behavior recognition model in this embodiment is trained by using the above-mentioned video behavior recognition model training method. Therefore, the video behavior recognition method provided by this embodiment has the same beneficial effects as the above-mentioned video behavior recognition model training method, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0191] As Figure 9 shown, the embodiments of the present invention also provide a video behavior recognition device 900, including:
[0192] An acquisition module 901, configured to obtain a target video;
[0193] A video behavior recognition module 902, configured to input the target video into a target video behavior recognition model for video behavior recognition processing to obtain an action recognition result, where the target video behavior recognition model is trained by using the above-mentioned video behavior recognition model training method.
[0194] The video behavior recognition device 900 provided by the embodiments of the present application can execute the embodiments of the above-mentioned video behavior recognition model training method, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0195] It should be noted that the division of units in the embodiments of the present application is illustrative, merely a logical function division, and there may be other division methods in actual implementation. In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0196] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a processor-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0197] As Figure 10 shown, an embodiment of the present application provides an electronic device 1000, including: a memory 1002, a processor 1001, and a program stored on the memory 1002 and executable on the processor 1001; the processor 1001 is configured to read the program in the memory 1002 to implement the steps in the video behavior recognition model training method or the video behavior recognition method as described above.
[0198] The embodiments of the present application further provide a readable storage medium, on which a program is stored. When the program is executed by a processor, it implements each process of the above video behavior recognition model training method or the embodiments of the video behavior recognition method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the readable storage medium can be any available medium or data storage device accessible by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memories (such as compact discs (CD), digital versatile discs (DVD), Blu-ray discs (BD), high-definition versatile discs (HVD), etc.), and semiconductor memories (such as read-only memories (ROM), erasable programmable read-only memories (EPROM), electrically erasable programmable read-only memories (EEPROM), non-volatile memories (NAND FLASH), solid state disks (SSD), etc.).
[0199] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0200] Through the description of the above embodiments, those skilled in the art can clearly understand that the above method of the embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, disk, optical disc), and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0201] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.
Claims
1. A video behavior recognition model training method, characterized in that: include: Performing basic training on the initial video behavior recognition model on the basic task, and obtaining the trained initial video behavior recognition model by minimizing the cross entropy loss, wherein the initial video behavior recognition model includes a backbone network, a bottleneck layer, a linear classification head, and a slow-fast dynamic reparameterization module; The model parameters of the backbone network and the bottleneck layer are retained, and the linear classification head is replaced with a randomly initialized multi-layer perceptron parsing head to obtain a basic video behavior recognition model; Iteratively training the basic video behavior recognition model on K incremental tasks to obtain a target video behavior recognition model, where K is a positive integer; Among them, the slow-fast dynamic reparameterization module includes a slow branch, a fast branch, a streaming discriminator and a reparameterization branch, the slow branch is used for incremental learning based on deep learning, the fast branch is used for incremental learning based on analytical machine learning, the streaming discriminator is used to generate fusion weights, and the reparameterization branch is used to output prediction results based on target features, and the target feature is obtained by fusion of the output of the slow branch and the output of the fast branch based on the fusion weights.
2. The method according to claim 1, characterized in that: The iterative training of the basic video behavior recognition model on the K incremental tasks to obtain the target video behavior recognition model includes: Training the basic video behavior recognition model on the kth incremental task to obtain a kth model, k=1, 2, ..., K; The step of training the basic video behavior recognition model on the kth incremental task to obtain the kth model includes: Inputting the training data corresponding to the kth incremental task into the backbone network for processing to obtain a backbone feature set; Extracting means and variances corresponding to a preset number of old class features from the prototype feature library to form an old class feature set, wherein the old class features are obtained by training the basic video behavior recognition model using the slow-fast dynamic reparameterization module in the first k-1 incremental tasks; The old class feature set and the main feature set constitute joint training data; Based on the joint training data, the k-1th model is trained using the slow-fast dynamic reparameterization module to obtain the kth model.
3. The method according to claim 2, characterized in that: Before extracting means and variances corresponding to a preset number of old class features from the prototype feature library to form an old class feature set, the method further includes: Get the mean μ corresponding to the old class features of the i-th class i and variance v i : Among them, F lib,i is the old class feature described in the i-th class, is the backbone feature set obtained by the backbone network through inference on the training data corresponding to the basic task, represents the mean feature corresponding to the old class feature of the i-th class, represents the variance corresponding to the old class feature of the i-th class, C represents the feature dimension, and Index(·) represents the index operation; Construct the prototype feature library, which includes the mean set U k and variance set V k : Among them, E t Represents the number of action categories of the tth incremental task.
4. The method according to claim 2, characterized in that: The method of training the k-1th model based on the joint training data using the slow-fast dynamic reparameterization module to obtain the kth model includes: Inputting the joint training data into the slow branch and the fast branch respectively to obtain an output of the slow branch and an output of the fast branch; Inputting the joint training data into the streaming discriminator to obtain a fusion weight; After fusing the output of the fast branch and the output of the slow branch based on the fusion weight, the output is input into the re-parameterization branch to obtain a prediction result; A prediction loss is calculated based on the prediction result and the category label, and a parameter of the slow branch is updated based on the prediction loss.
5. The method according to claim 2, characterized in that: The training of the basic video behavior recognition model on the kth incremental task to obtain the kth model further includes: The training data corresponding to the kth incremental task is input into the backbone network and the bottleneck layer for feature extraction, and then upsampled to obtain upsampled features. Based on the up-sampled features Calculate the optimal weight of the downsampling layer corresponding to the kth incremental task in, is the optimal weight of the downsampling layer corresponding to the k-1th incremental task, R k-1 is calculated for the k-1th incremental task, is the action label corresponding to the kth incremental task, and I is the unit matrix.
6. A video behavior recognition method, characterized in that: include: Get the target video; The target video is input into a target video behavior recognition model for video behavior recognition processing to obtain an action recognition result, wherein the target video behavior recognition model is trained using the video behavior recognition model training method according to any one of claims 1 to 5.
7. A video behavior recognition model training device, characterized in that: include: A basic training module, used to perform basic training on the initial video behavior recognition model on the basic task, and obtain the trained initial video behavior recognition model by minimizing the cross entropy loss, wherein the initial video behavior recognition model includes a backbone network, a bottleneck layer, a linear classification head, and a slow-fast dynamic reparameterization module; A processing module, used for retaining the model parameters of the backbone network and the bottleneck layer, replacing the linear classification head with a randomly initialized multi-layer perceptron parsing head, and obtaining a basic video behavior recognition model; An incremental training module, used for iteratively training the basic video behavior recognition model on K incremental tasks to obtain a target video behavior recognition model, where K is a positive integer; Among them, the slow-fast dynamic reparameterization module includes a slow branch, a fast branch, a streaming discriminator and a reparameterization branch, the slow branch is used for incremental learning based on deep learning, the fast branch is used for incremental learning based on analytical machine learning, the streaming discriminator is used to generate fusion weights, and the reparameterization branch is used to output prediction results based on target features, and the target feature is obtained by fusion of the output of the slow branch and the output of the fast branch based on the fusion weights.
8. A video behavior recognition device, characterized in that: include: An acquisition module, used to acquire a target video; The video behavior recognition module is used to input the target video into a target video behavior recognition model for video behavior recognition processing to obtain an action recognition result, wherein the target video behavior recognition model is trained using the video behavior recognition model training method described in any one of claims 1-5.
9. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; wherein the processor is used to read the program in the memory to implement the video behavior recognition model training method as described in any one of claims 1 to 5, or to implement the steps in the video behavior recognition method as described in claim 6.
10. A readable storage medium for storing a program, characterized in that: When the program is executed by a processor, the video behavior recognition model training method as described in any one of claims 1 to 5 is implemented, or the steps in the video behavior recognition method as described in claim 6 are implemented.
Citation Information
Cited By
Video action recognition method and device, computer equipment, readable storage medium and program product
CN120635995A