A behavior recognition method based on new class feature space to alleviate the forgetting of old classes
The memory set is constructed through the keyframe selection and interpolation algorithm, and the feature space of the new class is slightly moved, which solves the balance between the memory set size and the number of old class samples in class incremental behavior recognition, reduces the overfitting of new class and alleviates old class forgetting, and achieves higher recognition performance.
Patent Information
- Application Number
- CN202311241281.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-09-25
AI Technical Summary
Existing class incremental behavior recognition algorithms cannot balance memory set size and old class samples, resulting in the model overfitting new classes and losing old class features.
The memory set is constructed through the keyframe selection algorithm, and the multi-frame model is maintained using the interpolation algorithm, slightly moving the feature space of new classes to alleviate the forgetting of old classes.
The balance between memory set size and replay sample number is achieved, reducing the model's overfitting of new classes, significantly alleviating the catastrophic forgetting of old classes, and improving the overall recognition performance.
Smart Images

Figure CN117292294B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of deep learning, action recognition and class-incremental learning, and designs a class-incremental action recognition algorithm which alleviates the forgetting of old classes by slightly moving the feature space of new classes. Background Art
[0002] Action recognition algorithms have been widely studied in recent years, with the goal of judging a person's actions based on a video of an action. Traditional action recognition algorithms usually complete the training of the model in only one stage, which will provide training samples of all categories. However, in many real-world application scenarios, due to technical limitations or actual needs, model training often involves multiple stages, and training samples of different categories will appear in sequence in multiple stages. In this process, after the model has learned the task of one stage, it will continue to learn the task of the next stage, without using or using very little sample data from the previous stage, but will use all categories for testing in the test stage. The goal of the incremental continuous learning algorithm is to efficiently transform and utilize the knowledge that has been learned to complete the learning of new tasks, and to minimize the catastrophic forgetting of the knowledge that has been learned.
[0003] At present, the existing incremental learning algorithms have achieved remarkable performance in the field of images. These algorithms prove that building a memory set of old categories can alleviate catastrophic forgetting, and the larger the size of the memory set, the less likely it is to occur. Some incremental continuous learning algorithms used in the field of behavior recognition have also verified this. At the end of each stage of the task, they try to select a small number of representative videos and then build a memory set based on these videos. However, storing more old class samples and minimizing the size of the memory set are two conflicting factors. Some works directly store the original representative videos, which leads to a rapid expansion of the size of the memory set and brings considerable memory overhead; other works try to extract a prototype frame from each representative video to represent the overall distribution of the video. The memory set only saves the prototype frame, and the behavior recognition network is also changed to a single-frame-based model. Although the size of the memory set is reduced, the single-frame-based model cannot guarantee the recognition accuracy. At the same time, these methods over-pursue the performance of the new class in the current stage of model learning, resulting in serious overfitting of the model to the new class, thereby losing a large number of old class features. Summary of the invention
[0004] The present invention proposes a class incremental behavior recognition algorithm that slightly moves the feature space of new classes to alleviate the forgetting of old classes, aiming to solve the problem that the existing class incremental behavior recognition algorithm cannot balance the size of the memory set and the number of old class samples, and the problem that the model overfits the new class and causes catastrophic forgetting of the old class.
[0005] The present invention first extracts some key frames of representative videos through a key frame selection algorithm to construct a memory set, thereby avoiding excessive expansion of the memory set size. At the same time, an interpolation algorithm is used to supplement the key frames to the same size as the network input, so that the multi-frame-based behavior recognition model can continue to be used. Secondly, during the training process, by means of a strategy of randomly freezing some model parameters and jumping out of training in advance, the feature space of the new class is slightly moved, which greatly alleviates the model's forgetting of the old class at the expense of a very small amount of new class performance.
[0006] In order to achieve the above object, the present invention adopts the following specific technical scheme: a behavior recognition method based on a new class feature space to alleviate the forgetting of old classes, the method comprising:
[0007] Step 1: Normalize the video frame images and scale the frame images proportionally; initialize the memory set, extract the key frames of the training sample video, and store them in the memory set together with their corresponding behavior labels;
[0008] Step 2: Cut the video, and the cut length is the input length of the feature extraction network;
[0009] Step 3: Use the behavior recognition model trained with behavior to perform behavior recognition on the video captured in step 2;
[0010] The behavior recognition model includes: a feature extraction module and a behavior recognition module, wherein the feature extraction module is a deep residual convolutional neural network;
[0011] Step 4: When adding a new behavior to be recognized, the original trained recognition model Φ K-1 Based on this, further training is performed to obtain a new recognition model Φ K ; The training sample data are the original training sample data and the new behavior sample data, and the training loss function Loss is:
[0012] Loss = αLoss 1 +βLoss 2 +γLoss 3
[0013] Among them, α, β, and γ are weight parameters.
[0014]
[0015]
[0016]
[0017] Loss 1 Represents the recognition model Φ T The cross entropy loss between the predicted result and the true label, Y represents the number of categories, q y (x) is a binary symbol function. When the predicted label x is the same as the true label y, q y (x) is 1, otherwise it is 0; p y (x) represents the recognition model Φ T The confidence that the predicted label x is the true label y, Loss 2 It means that under the constraint of the importance matrix Im, the identification model Φ K-1 With the identification model Φ K The sum of the two norms of the corresponding neuron weight differences; L represents the recognition model Φ K The number of neural network layers, T represents Φ K The time length of the input frame sequence, C represents Φ T The number of channels per frame of the input frame sequence. is the importance matrix, representing the model Φ K The importance of neurons at the lth layer of the neural network with time dimension t and channel dimension c for recognizing old actions; Indicates that the input frame sequence is in Φ K The neuron weights at layer l, with time dimension t and channel dimension c, are similarly Indicates that the input frame sequence is in Φ K-1 The neuron weights at layer l, with time dimension t and channel dimension c; Loss 3 Represents Φ K Inter-frame orthogonality loss of the input frame sequence; Indicates that The splicing is performed on the time scale, and I is the identity matrix with only diagonal values of 1;
[0018] Step 5: After completing the training in step 4, use the trained model Φ K Calculate the average of the fully connected feature matrix of all videos in the new behavior l final Representative model Φ K The last layer; remove The closest videos are extracted, and the key frames of these videos are stored in the memory set along with their corresponding behavior labels.
[0019] Step 6: Identify the model Φ T Make fine adjustments;
[0020] According to the importance matrix ImK Determine the first m most important neurons, freeze the parameters of these m most important neurons, and train the parameters of the remaining neurons.
[0021] The input data is the memory set, the output is the behavior label, and the loss function is:
[0022]
[0023] Step 7: Use the recognition model Φ trained in step 6 K Perform action recognition on new videos.
[0024] Furthermore, in step 4 and step 6, if the data lengths of the data in the memory set and the input data of the recognition model are different, the data in the memory set is upsampled or downsampled to make the lengths of the two data the same;
[0025] Suppose the frame sequence of video k in the memory set is The upsampling method is;
[0026] Simple repeated interpolation algorithm:
[0027]
[0028] Average interpolation algorithm:
[0029]
[0030] Inter-frame prediction model based on convolutional neural network Interpolation algorithm:
[0031]
[0032] Represents the frame sequence of video k after the inter-frame prediction model After that, the i-th predicted frame is output, where i≤8.
[0033] Furthermore, the importance matrix Im K It is represented as model Φ K-1 The sum of the gradient norms of the cross entropy loss for all samples. Specifically, for The calculation method is as follows:
[0034]
[0035] in Represented by the model Φ K-1 The calculated loss value loss is the reverse gradient propagation value at the lth layer of the neural network, the time dimension is t, and the channel dimension is c; Represents the memory set, which contains all memory set samples from category 1 to category K-1.
[0036] Beneficial effects:
[0037] The class incremental behavior recognition algorithm proposed in the present invention alleviates the forgetting of old classes by slightly moving the feature space of new classes. First, the memory set constructed by the key frame extraction algorithm and the preprocessing interpolation algorithm effectively achieves a balance between the size of the memory set and the number of replay samples. While consuming the same storage space as other methods, the present invention can store more replay samples in the memory set, thereby effectively alleviating the catastrophic forgetting of old classes in class incremental continuous learning. Secondly, when the present invention reversely updates the model in the new stage, it randomly freezes the neurons that are sensitive to the old class and jumps out of training in advance, effectively preventing the model from excessively losing old class features. By slightly moving the model's feature space for the new class, more old class feature distributions are retained. Although the present invention slightly sacrifices the model's recognition performance for the new class, it greatly improves the model's recognition performance for the old class. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 The overall process framework diagram of the class incremental behavior recognition algorithm of the present invention which slightly moves the feature space of the new class to alleviate the forgetting of the old class. DETAILED DESCRIPTION
[0039] The present invention is implemented on the 2D convolutional network behavior recognition algorithm TSM based on the Pytorch deep learning platform. The basic steps are as follows:
[0040] Step 1: Select a suitable action recognition dataset, such as UCF101, HMDB51, or the larger Something-SomethingV2 dataset. All action categories in the dataset will be divided into multiple non-overlapping stage tasks and given in turn during the incremental learning process of the model. Specifically, if the model Φ K-1 It can recognize actions of categories 1 to K-1 and build memory sets of categories 1 to K-1. The current stage gives the training samples and memory set of the K-th action Train a model Φ K Realize the identification of categories 1 to K;
[0041] Step 2: Combine the training samples and memory set of the K-th action Input video preprocessing module, the processing results of the two input data are the same, both are 8-frame sequence, each frame is a 224*224 three-channel RGB image, that is, the final processing result size is 8*224*224*3. After many comparative experiments, for the memory set, the inter-frame prediction model based on convolutional neural network is used The interpolation algorithm has the highest performance.
[0042] Since a network based on 8 frames is used for behavior recognition, the module will extract 8 frames from the input video. It should be noted that the input of the module includes two categories, one is the new class video of the current stage, and the other is the old class memory set made in the previous stage. Since the memory set does not store the complete video in the present invention, but stores a frame sequence of 2 or 4 frames in length made by the key frame extraction module, the module will insert the frame sequence in the memory set and upsample it to a frame sequence of the same length as the network input. Specifically, for the frame sequence of video k in the memory set, Assume that the key frame extraction module extracts 4 key frames, and the video preprocessing module There are a variety of interpolation algorithms that can be used to upsample it, including a naive repeated interpolation algorithm:
[0043]
[0044] Average interpolation algorithm:
[0045]
[0046] And the inter-frame prediction model based on convolutional neural network Interpolation algorithm:
[0047]
[0048] Step 3: Input the frame sequence processed in step 2 into the frame sequence feature extraction module based on the knowledge distillation structure. The feature extraction network is designed based on ResNet. Since different datasets have different sensitivities to spatiotemporal information and require different network parameters, ResNet34 is used as the feature extraction network for the UCF101 dataset, while ResNet50 is used as the feature extraction network for the HMDB51 and Something-SomethingV2 datasets. In this module, in addition to the reverse update of the cross entropy loss for classification prediction, there is also a reverse update of the knowledge distillation loss and the neuron importance loss, and the three losses are added according to different weights. A deep residual convolutional neural network extracts the features of the frame image and fuses them between multiple frames, and transfers the learned old class features based on knowledge distillation. Knowledge distillation aims to guide the training of the new model through a model that has been trained; specifically, for the model Φ that has been trained in the previous stage, K-1 , which can realize the recognition of actions of categories 1 to K-1, and construct the memory set of categories 1 to K-1 At this stage, the training samples and memory sets of the K-th action are given Train a model Φ KRealize the recognition of categories 1 to K; for any training sample, it will enter the frame sequence feature extraction module for feature extraction. After the output feature passes through the Sigmoid normalization layer, the cross entropy loss with the true label is calculated, recorded as Loss 1 ;
[0049]
[0050] For the training samples of the 1st to K-1st action types from the memory set, they also enter the old model Φ K-1 Inference is made in Φ K-1 The training has been completed, so in the process of reasoning Φ K-1 The neurons in each layer, time scale, and channel will generate corresponding weights. K Calculate Φ under the constraint K-1 With Φ K The sum of the two norms of the corresponding neuron weight differences is denoted as Loss 2 , which will be used to guide the new model Φ T The neurons mimic Φ T-1 The neurons in , thereby realizing the recognition of old categories;
[0051]
[0052] In order to further improve the effectiveness of continuous learning and the discrimination of the importance matrix for different neurons, the orthogonal similarity loss between frames is calculated at the same time, denoted as Loss 3 By After splicing, orthogonal multiplication is performed on the time scale, and the identity matrix is used to maximize the difference between frames, ensuring that the model Φ K Pay more attention to the difference features between frames rather than the similar features.
[0053]
[0054] Finally, the loss function of the entire frame sequence feature extraction module is the sum of the above three parts, which realizes the new model Φ K While completing training on new categories, the memory of old categories is maintained.
[0055] Loss = αLoss 1 +βLoss 2 +γLoss 3
[0056] Step 4: After training in step 3, Φ K The recognition of categories 1 to K has been realized, and the key frame extraction module produces replay samples for the Kth category action to form the latest memory set After many comparative experiments, a video key frame extraction model based on convolutional neural network was used. The frame extraction algorithm has the highest performance. After the training of each stage, a memory set is made for the new class of the current stage. Specifically, in the model Φ K After training, the model has the ability to recognize actions of categories 1 to K, of which categories 1 to K-1 have been made into memory sets. Now we need to extract some representative videos from the training set of the K-th action and also make memory samples to add to the memory set. , forming a memory set Use ModelΦ K Calculate the average value of the fully connected feature matrix of all videos of the K-th action Take out and The closest partial video is used as the representative video. In addition, in our multiple verifications, we found a notable feature of action videos, that is, there is a lot of redundancy between frames. For a video with a duration of 4 seconds and a total of 120 frames, the number of key frames used to calibrate its action category can be limited to 2 to 4 frames. Therefore, this module will perform key frame extraction on the selected representative video and save its key frame sequence. Specifically, for a selected representative video k, the key frame extraction module δ can use a variety of algorithms to downsample it, including an equally spaced frame extraction algorithm based on the preprocessing results:
[0057]
[0058] Video key frame extraction model based on convolutional neural network Frame extraction algorithm:
[0059]
[0060] Through the key frame extraction module, the present invention can save 2 or even 4 times more replay samples than other methods under the premise of consuming the same memory set storage space. At the same time, based on the key frame extraction algorithm of the neural network, the extracted frames can better represent the overall feature distribution of the video, which is more efficient than other methods.
[0061] Step 5: The model Φ obtained through training in step 3 K and the memory set created in step 4 After the video preprocessing module, the fine-tuning module based on parameter freezing and early exit is input. By slightly moving the model Φ K The feature space of the K-th action in the model Φ K More feature distribution of the 1st to K-1th categories is retained. After each stage of training, in order to better prevent the model Φ KTo overfit the new class with enough samples in the current stage, the present invention deploys an additional fine-tuning module, in which some parameters are appropriately frozen and an early exit strategy is used to slightly shift the feature space of the current stage to compensate for the feature forgetting of the old class; specifically, after completing step 4, the memory set The replay samples of the 1st to Kth types of actions are already included in the fine-tuning module. K Fine-tune and memorize As input, all inputs need to be processed by the video preprocessing module. Since the number of replay samples of each type in the memory set is equal, the uniformly distributed training data can prevent the model from overfitting. K Describes the model Φ K The importance of neurons in each layer, time scale, and channel to the previous task. The fine-tuning module will randomly select Im K Some important neurons in the model Φ T In the fine-tuning of Φ, the reverse update of these neurons is frozen, that is, the model Φ is slightly blocked. K The extension to the feature space of the new class retains its distribution in the feature space of the old class. At the same time, the present invention finds that the model only needs a small number of training iterations (usually 5-10) to fall into overfitting during the incremental training stage. Continuing training at this time will cause the model to seriously overfit the new class, thereby losing many features learned in the previous stage. Therefore, the present invention sets a model performance threshold in the incremental training stage and stops training after reaching the performance threshold. Although this will slightly sacrifice the performance of the new class, it can greatly alleviate the model's catastrophic forgetting of the old class; the fine-tuning module only uses the cross entropy loss of the classification prediction as the total loss of the module:
[0062]
[0063] Based on key frame extraction and interpolation algorithms, the present invention achieves a balance between the size of the memory set and the number of replay samples, and achieves higher performance than other incremental behavior recognition algorithms while consuming the same storage space. At the same time, a fine-tuning module based on parameter freezing and early exit is designed, which can effectively alleviate the model's catastrophic forgetting of old classes, thereby improving the overall recognition performance.
Claims
1. A behavior recognition method based on new class feature space to alleviate the forgetting of old classes. include: Step 1: normalize the video frame images and perform frame image scaling processing; Initialize the memory set, extract key frames from the training sample video, and store them together with their corresponding behavior labels into the memory set; Step 2: Cut the video, and the cut length is the input length of the feature extraction network; Step 3: Use the behavior recognition model trained with behavior to perform behavior recognition on the video captured in step 2; The behavior recognition model includes: a feature extraction module and a behavior recognition module, wherein the feature extraction module is a deep residual convolutional neural network; Step 4: When adding a new behavior to be recognized, the original trained recognition model Φ K-1 Based on this, further training is performed to obtain a new recognition model Φ K ; The training sample data are the original training sample data and the new behavior sample data, and the training loss function Loss is: Loss=aLoss 1 +βLoss 2 +γLoss 3 Among them, α, β, and γ are weight parameters. Loss 1 Represents the recognition model Φ T The cross entropy loss between the predicted result and the true label, Y represents the number of categories, q y (x) is a binary symbol function. When the predicted label x is the same as the true label y, q y (x) is 1, otherwise it is 0; p y (x) represents the recognition model Φ T The confidence that the predicted label x is the true label y, Loss 2 It means that under the constraint of the importance matrix Im, the identification model Φ K-1 With the identification model Φ K The sum of the two norms of the corresponding neuron weight differences; L represents the recognition model Φ K The number of neural network layers, T represents Φ K The time length of the input frame sequence, C represents Φ T The number of channels per frame of the input frame sequence; is the importance matrix, representing the model Φ K The importance of neurons at the lth layer of the neural network with time dimension t and channel dimension c for recognizing old actions; Indicates that the input frame sequence is in Φ K The neuron weights at layer l, with time dimension t and channel dimension c, are similarly Indicates that the input frame sequence is in Φ K-1 The neuron weights at layer l, with time dimension t and channel dimension c; Loss 3 Represents Φ K Inter-frame orthogonality loss of the input frame sequence; Indicates that The splicing is performed on the time scale, and I is the identity matrix with only diagonal values of 1; Step 5: After completing the training in step 4, use the trained model Φ K Calculate the average of the fully connected feature matrix of all videos in the new behavior l final Representative model Φ K The last layer of The closest videos are extracted, and the key frames of these videos are stored in the memory set along with their corresponding behavior labels. Step 6: Identify the model Φ T Make fine adjustments; According to the importance matrix Im K Determine the first m most important neurons, freeze the parameters of these m most important neurons, and train the parameters of the remaining neurons. The input data is the memory set, the output is the behavior label, and the loss function is: Step 7: Use the recognition model Φ trained in step 6 K Perform action recognition on new videos.
2. A behavior recognition method based on a new class feature space to alleviate the forgetting of old classes as claimed in claim 1, It is characterized in that In step 4 and step 6, if the data lengths of the data in the memory set and the input data of the recognition model are different, upsampling or downsampling the data in the memory set to make the two data of the same length; Suppose the frame sequence of video k in the memory set is The upsampling method is; Simple repeated interpolation algorithm: Average interpolation algorithm: Inter-frame prediction model based on convolutional neural network Interpolation algorithm: Represents the frame sequence of video k after the inter-frame prediction model After that, the i-th predicted frame is output, where i≤8.
3. A behavior recognition method based on a new class feature space to alleviate the forgetting of old classes as claimed in claim 1, It is characterized in that Importance Matrix Im K It is represented as model Φ K-1 The sum of the gradient norms of the cross entropy loss for all samples. Specifically, for The calculation method is as follows: in Represented by the model Φ K-1 The calculated loss value loss is the reverse gradient propagation value at the lth layer of the neural network, the time dimension is t, and the channel dimension is c; Represents the memory set, which contains all memory set samples from category 1 to category K-1.