A video feature extraction method based on multi-dimensional information interaction
By combining a prefix convolutional network and a spatiotemporally separable encoder, a video feature extraction method is developed that solves the problem of incompatibility between temporal and spatial information interaction in existing technologies, and achieves efficient extraction of video features and modeling of long-term sequence spans.
Patent Information
- Application Number
- CN202311058507.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2043-08-22
AI Technical Summary
Existing 2D image convolution algorithms + Transformer methods lose spatial dimensionality information when processing videos, while 3D video feature extraction methods can only focus on temporal information between adjacent frames, making it difficult to effectively model information interaction over long time series.
A video feature extraction method based on multi-dimensional information interaction is adopted, which combines a prefix convolutional network, a spatiotemporally separable encoder and a video classifier. Feature extraction is performed through a spatial self-attention module, a temporal self-attention module and a forward propagation module, and a loss function is constructed for training.
It enables effective interaction between temporal and spatial information, reduces computational overhead, and improves the modeling capability for long-term time series.
Smart Images

Figure CN117274855B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of video processing technology, specifically relating to a video feature extraction method based on multi-dimensional information interaction. Background Technology
[0002] With the widespread adoption of mobile internet and the rapid growth of communication bandwidth, video has become a crucial medium for modern information transmission. However, compared to images, video possesses an additional temporal dimension and carries a greater amount of information, making information extraction from video more complex than image processing. Existing 2D image convolution algorithms combined with Transformers can address the interaction problem of long-term sequences, but directly using convolutional features results in the loss of spatial dimension information. Existing 3D video feature extraction methods, such as C3D and VidSwin, can model spatial dimension information, but they can only focus on temporal information between adjacent frames. For information spanning longer temporal distances, long steps are required for interaction, which is detrimental to temporal modeling. Summary of the Invention
[0003] To address the incompatibility between the ability to model long-term time series and spatial dimensions, this invention provides a video feature extraction method based on multi-dimensional information interaction.
[0004] A video feature extraction method based on multi-dimensional information interaction includes the following steps:
[0005] Step 1: Dataset acquisition, using existing video classification datasets.
[0006] Step 2: Construct a video feature extraction network based on multi-dimensional information interaction;
[0007] The video feature extraction network based on multi-dimensional information interaction includes a prefix convolutional network, a spatiotemporally separable encoder, and a video classifier.
[0008] Step 3. Construct the loss function;
[0009] Construct a video classification loss;
[0010] L = CE(CLS,Y)
[0011] Where Y is the correct video classification label in the dataset.
[0012] Step 4. Train the constructed video feature extraction network based on multi-dimensional information interaction using the acquired dataset.
[0013] Furthermore, the specific method for step 1 is as follows:
[0014] Step 1.1: Download the Kinetics-600 dataset.
[0015] The Kinetics-600 dataset is an extension of the Kinetics-400 dataset, containing 370k training videos and 28.3k validation videos from 600 human motion categories. This dataset includes annotations for human motion.
[0016] Step 1.2: Download the Something-Something V2 (SSv2) dataset.
[0017] The Something-Something V2 (SSv2) dataset contains 168.9K training videos and 24.7K validation videos across 174 categories. This dataset is used to test temporal modeling capabilities.
[0018] Furthermore, the aforementioned prefix convolutional network pre-samples the video frames to the normal resolution using a convolutional network;
[0019] Let the input tensor be of size . Video Where T represents the number of frames, H represents the height of each frame, W represents the width of each frame, and C represents the number of channels in each frame. The prefix convolutional network is a ResNet50 network with the last two layers removed. It first modifies the input raw video frames to a uniform resolution, and finally outputs a 7*7*512 feature map for each frame. The overall process is represented as follows:
[0020]
[0021] Z∈[T,H,W,C]
[0022] Where Z represents the encoded video feature. H, W, and C represent the height, width, and number of channels of the video feature, respectively.
[0023] Furthermore, the spatiotemporally separable encoder is a Transformer Encoder, which consists of three spatiotemporally separable encoder layers stacked together.
[0024] The spatiotemporally separable encoder layer comprises three important components: Spatial Self Attention (SSA), Temporal Self Attention (TSA), and Forward Propagation (FFN). These three modules are all wrapped by Residual Links and Layer Normalization (LN) modules, meaning that each of these three modules is followed by a Residual Links and Layer Normalization (LN) module.
[0025] Residual Linking and Layer Normalization Module (LN): Assuming the input feature tensor is X, the LN formula is expressed as follows.
[0026] LN(X) = LayerNorm(f(X) + X)
[0027] Where LayerNorm represents the layer normalization algorithm, and f(·) represents the function enclosed by LN.
[0028] The Spatial Self-Attention (SSA) module first passes the feature tensor X through three independent convolution kernels to obtain the corresponding query vector Q, key vector K, and value vector V. This process can be represented by the following formula.
[0029] X_Q = ConvQ(X)
[0030] X_K = ConvK(X)
[0031] X_V = ConvV(X)
[0032] ConvQ, ConvK, and ConvV are three convolutional layers with a kernel size of 3, a stride of 1, 512 input channels, and 512 output channels.
[0033] Then, X_Q, X_K, and X_V are stretched to make the dimensions [T, H×W, C], where H×W forms the spatial dimension. A self-attention operation is then performed at the resolution level.
[0034] X s =Attention(X_Q,X_K,X_V)
[0035] Attention(·) represents attention calculation. The feature dimensions are then restored to [T,H,W,C].
[0036] Temporal Self-Attention Module (TSA), similar to SSA, assumes that the input feature tensor is X, and its process can be represented by the following formula.
[0037] X_Q = ConvQ(X)
[0038] X_K = ConvK(X)
[0039] X_V = ConvV(X)
[0040] The structure of the three convolutional layers ConvQ, ConvK, and ConvV is the same as that of the convolutional layers in the spatial self-attention encoder, but the parameters are independent, and the same notation is used here.
[0041] Then, the order of the first three dimensions of X_Q, X_K, and X_V is swapped, making their dimensions [H, W, T, C], where T forms the time dimension. Self-attention operations are then performed in the spatial dimension.
[0042] X t =Attention(X_Q,X_K,X_V)
[0043] Attention(·) represents attention calculation. The feature dimensions are then restored to [T, H, W, C]. The overall process is shown in the figure and denoted as SSA. Transpose represents the tensor transpose operation. Mul represents matrix multiplication.
[0044] The Forward Propagation Module (FFN) is an MLP layer with both input and output dimensions of C.
[0045] The output of the previous spatiotemporally separable encoder layer is used as the input of the next spatiotemporally separable encoder layer, and so on. The features output by the last spatiotemporally separable encoder layer are used for downstream classification tasks.
[0046] Furthermore, the video classifier is specifically implemented as follows: global pooling is performed on the features output by the last spatiotemporally separable encoder layer to obtain the overall video features, and an MLP with a softmax activation function is used for multi-classification.
[0047] CLS = Softmax(MLP(X))
[0048] Where CLS is the probability distribution for multi-class classification, and X represents the input feature tensor.
[0049] Furthermore, the specific method for step 4 is as follows;
[0050] For both datasets, the same training parameters were used. The learning rate was set to 0.01, and training was performed for 30 epochs. Every 10 epochs, the learning rate was reduced to 0.1 times the previous value. The spatiotemporally separable encoder consisted of three stacked spatiotemporally separable encoder layers.
[0051] The beneficial effects of this invention are as follows:
[0052] This invention combines temporal and spatial information interaction, overcoming the limitation that the two cannot coexist. By using a prefix convolutional network and a temporally and spatially separable attention mechanism, it significantly reduces computational overhead. Attached Figure Description
[0053] Figure 1 This is a schematic diagram of the spatial self-attention module structure;
[0054] Figure 2This is a schematic diagram of the temporal self-attention module structure;
[0055] Figure 3 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0056] The technical solution of the present invention will be further defined below with reference to the accompanying drawings and embodiments.
[0057] like Figure 3 As shown, this invention proposes a video feature extraction method based on multi-dimensional information interaction, the basic steps of which are as follows:
[0058] Step 1: Dataset acquisition, download two video classification datasets;
[0059] Step 1.1: Download the Kinetics-600 dataset.
[0060] The Kinetics-600 dataset is an extension of the Kinetics-400 dataset, containing 370k training videos and 28.3k validation videos from 600 human motion categories. This dataset includes annotations for human motion.
[0061] Step 1.2: Download the Something-Something V2 (SSv2) dataset.
[0062] The Something-Something V2 (SSv2) dataset contains 168.9K training videos and 24.7K validation videos across 174 categories. This dataset is used to test temporal modeling capabilities.
[0063] Step 2: Construct a video feature extraction network based on multi-dimensional information interaction;
[0064] The video feature extraction network based on multi-dimensional information interaction includes a prefix convolutional network, a spatiotemporally separable encoder, and a video classifier.
[0065] The aforementioned prefix convolutional network pre-samples video frames to the standard resolution using a convolutional network;
[0066] Step 2.1: Construct a prefix convolutional network;
[0067] Using video information directly at its original resolution incurs significant computational overhead, and the original video resolution is not standardized, making it difficult to utilize directly. Therefore, a convolutional network is used beforehand to downsample the video frames to a standardized resolution.
[0068] Let the input tensor be of size . Video Where T represents the number of frames, H represents the height of each frame, W represents the width of each frame, and C represents the number of channels in each frame. The prefix convolutional network is a ResNet50 network with the last two layers removed. It first modifies the input raw video frames to a uniform resolution, and finally outputs a 7*7*512 feature map for each frame. The overall process is represented as follows:
[0069]
[0070] Z∈[T,H,W,C]
[0071] Where Z represents the encoded video feature. H, W, and C represent the height, width, and number of channels of the video feature, respectively, and in this embodiment, their values are 7, 7, and 512.
[0072] Step 2.2: Spatiotemporally separable encoder;
[0073] A spatiotemporally separable encoder is a Transformer Encoder consisting of three spatiotemporally separable encoder layers stacked together.
[0074] The spatiotemporally separable encoder layer comprises three important components: Spatial Self Attention (SSA), Temporal Self Attention (TSA), and Forward Propagation (FFN). These three modules are all wrapped by Residual Links and Layer Normalization (LN) modules, meaning that each of these three modules is followed by a Residual Links and Layer Normalization (LN) module.
[0075] Residual Linking and Layer Normalization Module (LN): Assuming the input feature tensor is X, the LN formula is expressed as follows.
[0076] LN(X) = LayerNorm(f(X) + X)
[0077] Where LayerNorm represents the layer normalization algorithm, and f(·) represents the function enclosed by LN.
[0078] like Figure 1 As shown, the Spatial Self-Attention (SSA) module first passes the feature tensor X through three independent convolution kernels to obtain the corresponding query vector Q, key vector K, and value vector V. This process can be represented by the following formula.
[0079] X_Q = ConvQ(X)
[0080] X_K = ConvK(X)
[0081] X_V = ConvV(X)
[0082] ConvQ, ConvK, and ConvV are three convolutional layers with a kernel size of 3, a stride of 1, 512 input channels, and 512 output channels.
[0083] Then, X_Q, X_K, and X_V are stretched to make the dimensions [T, H×W, C], where H×W forms the spatial dimension. A self-attention operation is then performed at the resolution level.
[0084] X s =Attention(X_Q,X_K,X_V)
[0085] Attention(·) represents attention calculation. The feature dimensions are then restored to [T, H, W, C]. The overall process is as follows: Figure 1 As shown in the diagram, Flatten represents the tensor stretching operation, and squeeze represents the tensor dimension addition operation. Mul represents matrix multiplication.
[0086] like Figure 2 As shown, the Temporal Self-Attention Module (TSA), similar to the SSA, assumes that the input feature tensor is X, and its process can be represented by the following formula.
[0087] X_Q = ConvQ(X)
[0088] X_K = ConvK(X)
[0089] X_V = ConvV(X)
[0090] The structure of the three convolutional layers ConvQ, ConvK, and ConvV is the same as that of the convolutional layers in the spatial self-attention encoder, but the parameters are independent, and the same notation is used here.
[0091] Then, the order of the first three dimensions of X_Q, X_K, and X_V is swapped, making their dimensions [H, W, T, C], where T forms the time dimension. Self-attention operations are then performed in the spatial dimension.
[0092] X t =Attention(X_Q,X_K,X_V)
[0093] Attention(·) represents attention calculation. The feature dimensions are then restored to [T, H, W, C]. The overall process is shown in the figure and denoted as SSA. Transpose represents the tensor transpose operation. Mul represents matrix multiplication.
[0094] The Forward Propagation Module (FFN) is an MLP layer with both input and output dimensions of C.
[0095] The output of the previous spatiotemporally separable encoder layer is used as the input of the next spatiotemporally separable encoder layer, and so on. The features output by the last spatiotemporally separable encoder layer are used for downstream classification tasks.
[0096] Step 2.3: Construct a video classifier;
[0097] The video classifier is implemented as follows: global pooling is performed on the features output by the last spatiotemporally separable encoder layer to obtain the overall video features, and an MLP with a softmax activation function is used for multi-class classification.
[0098] CLS = Softmax(MLP(X))
[0099] Where CLS is the probability distribution for multi-class classification, and X represents the input feature tensor.
[0100] Step 3. Construct the loss function;
[0101] Construct a video classification loss;
[0102] L = CE(CLS,Y)
[0103] Where Y is the correct video classification label in the dataset.
[0104] Step 4. Train the constructed video feature extraction network based on multi-dimensional information interaction using the acquired dataset;
[0105] For both datasets, we used the same training parameters. We set the learning rate to 0.01, trained for 30 epochs, and reduced the learning rate to 0.1 times the previous value every 10 epochs. The spatiotemporally separable encoder consists of three stacked spatiotemporally separable encoder layers.
[0106] The above description, in conjunction with specific / preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. Those skilled in the art can make various substitutions or modifications to these described embodiments without departing from the inventive concept, and all such substitutions or modifications should be considered within the scope of protection of the present invention.
[0107] The parts of this invention not described in detail are well-known to those skilled in the art.
Claims
1. A video feature extraction method based on multi-dimensional information interaction, characterized in that, The steps include the following: Step 1: Dataset acquisition, using existing video classification datasets; Step 2: Construct a video feature extraction network based on multi-dimensional information interaction; The video feature extraction network based on multi-dimensional information interaction includes a prefix convolutional network, a spatiotemporally separable encoder, and a video classifier; the prefix convolutional network is a ResNet50 network with the last two layers removed. A spatiotemporally separable encoder is a Transformer Encoder consisting of three spatiotemporally separable encoder layers stacked together. The spatiotemporally separable encoder layer contains three important components: spatial self-attention module, temporal self-attention module, and forward propagation module FFN. These three modules are all wrapped by residual linking and layer normalization module LN, that is, each of these three modules is connected to residual linking and layer normalization module LN. Residual Linking and Layer Normalization Module (LN): Assuming the input feature tensor is X, the LN formula is expressed as follows; in Presentation layer normalization algorithm, This represents the function enclosed by LN; The Spatial Self-Attention Module (SSA) first applies the feature tensor After passing through three independent convolution kernels, the corresponding query vector Q, key vector K, and value vector V are obtained. The process can be represented by the following formula. It consists of three convolutional layers with a kernel size of 3, a stride of 1, 512 input channels, and 512 output channels. Then This makes the dimension become , This forms the spatial dimension; then, self-attention operations are performed at the resolution level; where T represents the number of frames, H represents the height of each frame, W represents the width of each frame, and C represents the number of channels in each frame. This represents attention computation; subsequently, the feature dimensions are reduced to... ; The Temporal Self-Attention Module (TSA), similar to the SSA, assumes that the input feature tensor is X, and its process can be represented by the following formula; in The structure of the three convolutional layers is the same as that of the convolutional layers in the spatial self-attention encoder, but the parameters are independent, and the same notation is used here; Then exchange The order of the first three dimensions makes their dimensions become , It forms the time dimension; it performs self-attention operations in the spatial dimension; This represents attention computation; subsequently, the feature dimensions are reduced to... ; The forward propagation module FFN is an MLP layer with both input and output dimensions of C; The output of the previous spatiotemporally separable encoder layer is used as the input of the next spatiotemporally separable encoder layer, and so on. The features output by the last spatiotemporally separable encoder layer are used for downstream classification tasks. Step 3. Construct the loss function; Construct a video classification loss; Where Y is the correct video classification label in the dataset; CLS is the probability distribution of multi-class classification; Step 4. Train the constructed video feature extraction network based on multi-dimensional information interaction using the acquired dataset.
2. The video feature extraction method based on multi-dimensional information interaction according to claim 1, characterized in that, The specific method for step 1 is as follows: Step 1.1: Download the Kinetics-600 dataset; The Kinetics-600 dataset is an extension of the Kinetics-400 dataset, containing 370k training videos and 28.3k validation videos from 600 human motion categories; this dataset is labeled with human motion. Step 1.2: Download the Something-Something V2 dataset The Something - Something V2 dataset contains 168.9 K training videos and 24.7 K validation videos across 174 categories; this dataset is used to test temporal modeling capabilities.
3. The video feature extraction method based on multi-dimensional information interaction according to claim 1, characterized in that, The aforementioned prefix convolutional network pre-samples video frames to the standard resolution using a convolutional network; Let the input tensor be of size . Video Where T represents the number of frames, H represents the height of each frame, W represents the width of each frame, and C represents the number of channels in each frame; the prefix convolutional network is a ResNet50 network with the last two layers removed; it first modifies the input original video frames to a uniform resolution, and finally outputs a 7*7*512 feature map for each frame; the overall process is represented as: Where Z represents the encoded video features; These represent the height, width, and number of channels of the video feature, respectively.
4. The video feature extraction method based on multi-dimensional information interaction according to claim 1, characterized in that, The video classifier is implemented as follows: global pooling is performed on the features output by the last spatiotemporally separable encoder layer to obtain the overall video features, and an MLP with a softmax activation function is used for multi-classification. Where CLS is the probability distribution for multi-class classification. This represents the input feature tensor.
5. The video feature extraction method based on multi-dimensional information interaction according to claim 4, characterized in that, The specific method for step 4 is as follows; For both datasets, the same training parameters were used; the learning rate was set to 0.01, and the training was conducted for 30 epochs. Every 10 epochs, the learning rate was reduced to 0.1 times the previous value; a total of 3 spatiotemporally separable encoder layers were stacked.
Citation Information
Patent Citations
Action video recognition method combining hybrid convolution residual network and attention
CN112149504A
Systems And Methods For Improved Video Understanding
US20230017072A1