A video domain knowledge memorization and migration method based on an efficient video memory network
Patent Information
- Application Number
- CN202311750182.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-12-19
AI Technical Summary
第一类方法丢失了源视频域数据的部分时序信息,而且显式地存储帧图像存在安全隐患;第二类方法破坏了源视频域数据的时序结构
[0031]本发明所提出的基于高效视频记忆网络的视频域知识记忆与迁移方法能在6Mb存储空间限制下将100个源视频域数据隐式地存储到神经网络中,存储效率是现有方法的20倍。同时,蒸馏损失能有效地将源视频域知识迁移至目标视频域,识别准确率相比最先进的方法提升了3.98个百分点。
Smart Images

Figure CN117952162B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video image processing, and in particular to a video domain knowledge memory and transfer method based on an efficient video memory network for behavior recognition tasks. Background Technology
[0002] With the development of information technology, more and more video data is being uploaded to the internet, and a large number of video-related applications have emerged in people's daily lives. Due to the complexity and variability of real-world scenarios, deep learning models are required to transfer knowledge from the source video domain to the current model while learning knowledge from the target video domain. However, due to privacy concerns and storage limitations, only a portion of the source video domain training data is often retained when learning knowledge from the target video domain. In this scenario, deep learning models can only transfer source video domain knowledge to the target video domain to a limited extent. To improve the efficiency of deep learning models in memorizing source video domain knowledge under limited storage conditions and to transfer more source video domain knowledge to the target video domain, this invention, compared to existing methods, adopts a technical approach based on efficient video memory networks to memorize and transfer video domain knowledge.
[0003] Currently, traditional behavior recognition methods select 5 videos as a subset for each class based on the mean feature value of each class in the current network when storing source video domain data. Then, 8 frames are sampled from each video and stored, with each frame being a 3-channel color image at a resolution of 224x224. This storage method occupies 5x8x3x224x224 = 6Mb of space. However, this method can only store videos from 5 source video domains, failing to accurately reflect the distribution of the source video domains. Existing behavior recognition methods optimize storage construction, mainly falling into two categories: the first increases the number of video samples from 5 to 10, while reducing the number of frames sampled per video from 8 to 4; the second increases the number of video samples from 5 to 40, while performing a weighted summation of the 8 sampled frames for each video, storing only one image per video. The first method loses some temporal information from the source video domain data, and explicitly storing frame images poses security risks; the second method disrupts the temporal structure of the source video domain data. Summary of the Invention
[0004] To achieve efficient memorization and transfer of knowledge from the source video domain, this invention implicitly stores the knowledge in a neural network using an efficient video memory network. To facilitate knowledge transfer from the source video domain to the target video domain, a distillation loss mechanism is designed to ensure that the source and target video domain models represent the source video domain knowledge as closely as possible.
[0005] Existing research indicates that increasing the amount of source video domain data can improve the generalization ability of the target video domain model. To enhance the model's storage of source video domain knowledge within limited storage space, this invention proposes an efficient video memory network. Therefore, the technical solution of this invention is: a video domain knowledge memory and transfer method based on an efficient video memory network, the method comprising:
[0006] Step 1: Perform standard model training on the source video domain dataset to obtain the source video domain model S;
[0007] Step 2: Calculate the feature mean p of the source video domain model S for the i-th class of source video domain data to be stored. i Then, for the q-th frame image x of the j-th data of the source video domain to be stored in the i-th class... i,j,q Calculate the frame index t:
[0008] t = N × j + q
[0009] Where N is the number of source video domains to be stored in the i-th category;
[0010] Using the characteristic mean p i The spatial information of the frame is represented by the frame index t, which represents the temporal information of the frame, through the spatiotemporal encoder E. i Encode both to obtain spatiotemporal features O i,j,q Specifically, it is expressed as:
[0011] O i,j,q =E i (p i ,t)=FC(Γ(t))·p i,q
[0012] Where FC is a fully connected layer, p i,q Representative feature mean p i In the q-th dimension of the data, Γ(t) is the positional encoding (PE):
[0013] Γ(t)=(sin(b 0 πt),cos(b 0 πt),…,sin(b l-1 πt),cos(b l-1 πt))
[0014] Where b and l are hyperparameters for positional encoding;
[0015] Step 3: Obtaining the spatiotemporal features O i,j,q Then, from the spatiotemporal characteristics O i,j,q The frame image is reconstructed from the image and then introduced into the reconstruction decoder D. i Specifically:
[0016] First, the spatiotemporal feature O is processed through a fully connected layer. i,j,q Perform a linear mapping to obtain the mapping feature o i,j,q Then, an upsampling module consisting of convolutional layers, pixel recombination layers, and activation layers, by stacking four upsampling modules, maps the features o. i,j,q Convert to reconstructed frame image After inputting frame images of each video class from the source video domain, the knowledge of the source video domain is implicitly stored in an efficient video memory network G for each class, consisting of a spatiotemporal encoder E and a reconstruction decoder D. The reconstructed video set can only be obtained using the corresponding feature mean p and frame index set T.
[0017]
[0018] To reduce the difference between the reconstructed video and the original video and improve the accuracy of knowledge memory in the source video domain, the efficient video memory network G is supervised at both the pixel level and the feature level; where G(·,·) represents the efficient video memory network, E(·,·) represents the spatiotemporal encoder, and D(·) represents the reconstruction decoder.
[0019] Step 4: To reduce pixel-level differences, pixel consistency loss is used. To minimize the reconstructed video set Differences from the original video set x:
[0020]
[0021] Where B is the reconstructed video set The number of frames, α is the weight balance coefficient, ‖·‖1 is the absolute value loss, and SSIM(·,·) is the structural similarity loss;
[0022] Step 5: In addition to pixel-level supervision, the spatiotemporal features O are also supervised using the trained source video domain model S; the original video set x is fed into the source video domain model S to obtain the original video set features S(x), and then feature consistency loss is applied. Minimize the difference between the original video set features S(x) and the spatiotemporal features O:
[0023]
[0024] in, Represents the Frobenius norm;
[0025] The total loss of the efficient video memory network is:
[0026]
[0027] After training the efficient video memory network, loss was distilled. Transferring source video domain knowledge to the target video domain model F:
[0028]
[0029] in, Indicates distillation loss, and These represent the reconstructed video sets. The features output after being input into the target video domain model F and the source video domain model S.
[0030] Beneficial effects:
[0031] The video domain knowledge memorization and transfer method based on an efficient video memory network proposed in this invention can implicitly store 100 source video domain data into a neural network within a 6Mb storage space limit, achieving a storage efficiency 20 times higher than existing methods. Simultaneously, distillation loss effectively transfers source video domain knowledge to the target video domain, improving recognition accuracy by 3.98 percentage points compared to state-of-the-art methods. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of a video domain knowledge memory and transfer method based on an efficient video memory network.
[0033] Figure 2 This is an example of comparing frame images of the original video and the reconstructed video. The original video image and the reconstructed video image are corresponding vertically, with the original video image on top and the corresponding reconstructed video image on the bottom.
[0034] Figure 3 The images show the feature space before and after distillation. The left image shows the feature space before distillation, and the right image shows the feature space after distillation. Detailed Implementation
[0035] This invention aims to implicitly store more video data in an efficient video memory network with a size consistent with the storage space budget, under the supervision of the source video domain model, and transfer source video domain knowledge to the target video domain model through distillation loss. First, a spatiotemporal encoder encodes the class feature mean and frame index. Then, the encoded spatiotemporal features are fed into a reconstruction decoder to obtain reconstructed frame images, and the efficient video memory network is supervised at both the pixel and feature levels. Finally, the source video domain knowledge is stored through the efficient video memory network, and the distillation loss between the representations of source video domain knowledge in the source and target video domain models is calculated to transfer the knowledge to the target video domain.
[0036] This invention is implemented on a GPU server experimental platform and mainly includes the following steps: training a source video domain model on the source video domain dataset, using the source video domain model and the source video domain dataset to train an efficient video memory network for memorization, and transferring source video domain knowledge on the target video domain through distillation loss.
[0037] Step 1: Perform standard model training on the source video domain dataset to obtain the source video domain model;
[0038] Step 2: Calculate the feature mean of each class in the source video domain dataset using the source video domain model;
[0039] Step 3: Perform video sampling and frame sampling on each class of the source video domain dataset to obtain the original video set and frame index set for each class;
[0040] Step 4: Encode the feature mean and frame index set of each class using a spatiotemporal encoder to obtain spatiotemporal features;
[0041] Step 5: Decode the spatiotemporal features of each class using the reconstruction decoder for each class to obtain the reconstructed video set for each class;
[0042] Step 6: Use The loss is optimized at the feature level and pixel level for various types of efficient video memory networks consisting of a spatiotemporal encoder and a reconstruction decoder, each with a size of 6.0 Mb.
[0043] Step 7: During training in the target video domain, the original video set is reconstructed using feature mean, frame index set, and efficient video memory network. Under the storage space limit of 6.0Mb per class, 100 videos with 8 frames sampled per class can be reconstructed.
[0044] Step 8: Apply distillation loss to the target video domain By transferring knowledge from the source video domain, the recognition accuracy is improved by 3.98 percentage points compared to the state-of-the-art existing methods.
Claims
1. A method for video domain knowledge memorization and transfer based on an efficient video memory network, the method comprising: Step 1: Perform standard model training on the source video domain dataset to obtain the source video domain model S; Step 2: Calculate the feature mean p of the source video domain model S for the i-th class of source video domain data to be stored. i Then, for the q-th frame image x of the j-th data of the source video domain to be stored in the i-th class... i,j,q Calculate the frame index t: t = N × j + q Where N is the number of source video domains to be stored in the i-th category; Using the characteristic mean p i The spatial information of the frame is represented by the frame index t, which represents the temporal information of the frame, through the spatiotemporal encoder E. i Encode both to obtain spatiotemporal features O i,j,q Specifically, it is expressed as: About i,j,q =E i (p i ,t)=FC(Γ(t))·p i,q Where FC is a fully connected layer, p i,q Representative feature mean p i In the q-th dimension of the data, Γ(t) is the position code: Γ(t)=(sin(b 0 πt),cos(b 0 πt),…,sin(b l-1 πt),cos(b l-1 πt)) Where b and l are hyperparameters for positional encoding; Step 3: Obtaining the spatiotemporal features O i,j,q Then, from the spatiotemporal characteristics O i,j,q The frame image is reconstructed from the image and introduced into the reconstruction decoder D. i Specifically: First, the spatiotemporal feature O is processed through a fully connected layer. i,j,q Perform a linear mapping to obtain the mapping feature o i,j,q Then, an upsampling module consisting of convolutional layers, pixel recombination layers, and activation layers, by stacking four upsampling modules, maps the features o. i,j,q Convert to reconstructed frame image After inputting frame images of each video class from the source video domain, the knowledge of the source video domain is implicitly stored in an efficient video memory network G for each class, consisting of a spatiotemporal encoder E and a reconstruction decoder D. The reconstructed video set can only be obtained using the corresponding feature mean p and frame index set T. Where G(·,·) represents an efficient video memory network, E(·,·) represents a spatiotemporal encoder, and D(·) represents a reconstruction decoder; Step 4: To reduce pixel-level differences, pixel consistency loss is used. To minimize the reconstructed video set Differences from the original video set x: Where B is the reconstructed video set The number of frames, α is the weight balance coefficient, ‖·‖1 is the absolute value loss, and SSIM(·,·) is the structural similarity loss; Step 5: In addition to pixel-level supervision, the spatiotemporal features O are also supervised using the trained source video domain model S; the original video set x is fed into the source video domain model S to obtain the original video set features S(x), and then feature consistency loss is applied. Minimize the difference between the original video set features S(x) and the spatiotemporal features O: in, Represents the Frobenius norm; The total loss of the efficient video memory network is: After training the efficient video memory network, loss was distilled. Transferring source video domain knowledge to the target video domain model F: in, Indicates distillation loss, and These represent the reconstructed video sets. The features output after being input into the target video domain model F and the source video domain model S.
Citation Information
Patent Citations
Video implicit representation method based on decoupled space and time sequence information
CN115147275A
Method and system for communicating compressed video data
WO2009047692A2