Video action recognition method based on calculation of frame difference information of different frame intervals
Through the FISNet model, the problem of insufficient space-time information capture in the existing methods is solved, and the accuracy of video action recognition is achieved, especially the significant improvement on the HMDB51 dataset.
Patent Information
- Application Number
- CN202510309907.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing video action recognition method based on deep learning is insufficient to capture space-time information, and feature extraction is easily interfered with by redundant information, resulting in low recognition accuracy.
The FISNet model based on calculating frame difference information of different frame intervals is adopted, combining 3D convolutional layer, maximum pooling layer and multiple inter-frame differential information extraction modules, video features are enhanced through the temporal feature extraction module and the residual layer, and action recognition is performed using inter-frame differential information and spatial attention mechanism.
The accuracy of video action recognition is improved, especially the TOP1 classification accuracy on the HMDB51 dataset reaches 74.0%, which is better than the existing methods.
Smart Images

Figure CN120259938A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video action recognition method, and more particularly to a video action recognition method based on calculating frame difference information with different frame intervals. Background Art
[0002] Video image action recognition is an important research direction in the field of computer vision, which aims to enable a computer to automatically understand human actions in video content. This technology involves extracting action features from a video sequence and using machine learning or deep learning models to analyze these features to identify the ongoing actions or activities.
[0003] Action recognition is very important in many practical applications, such as video surveillance, human-computer interaction, sports analysis, healthcare, and the entertainment industry. In the field of security monitoring, action recognition can help detect abnormal behaviors or potential security threats. In sports events, it can be used to analyze the performance and techniques of athletes. In healthcare, action recognition can monitor the rehabilitation process and daily activities of patients to provide better care services, etc.
[0004] With the development of deep learning technology, deep learning has achieved great success in fields such as image classification, object detection, and speech recognition. Deep learning can automatically learn high-level feature representations from data, avoiding the complexity and limitations of manually designed features. The accuracy and efficiency of deep learning in human action recognition have been significantly improved. Usually based on models such as convolutional neural network (CNN), recurrent neural network (RNN), and graph neural network (GNN), it can learn complex spatio-temporal features from video data, and the addition of the attention mechanism further improves the performance of action recognition. There are two main methods for CNN: one is the method based on the two-stream network, which divides the video into RGB frames and optical flow frames, extracts spatial and temporal features respectively using a convolutional neural network (CNN), and then fuses the two features for classification; the other is the method based on the three-dimensional convolutional neural network (3D CNN), which can directly perform three-dimensional convolution on video blocks, learn spatial and temporal features simultaneously, and then use a fully connected layer or a pooling layer for classification. Both of these methods have their own advantages and disadvantages. For example, the two-stream network can better capture the details of actions, but it requires additional calculation of optical flow; 3D CNN can better model the overall action, but it requires more parameters and computing resources. Summary of the Invention
[0005] Aiming at the problems of insufficient capture of spatio-temporal information and easy interference of feature extraction by redundant information in the image action recognition method based on deep learning, the present invention proposes a video action recognition method based on calculating frame difference information with different frame intervals for video image processing.
[0006] The technical solution of the present invention is as follows:
[0007] 1. A video action recognition method based on calculating frame difference information of different frame intervals
[0008] First, preprocess the image frames of the original video to obtain the preprocessed video; then, input the preprocessed video into the FISNet model based on calculating inter-frame difference information, and the model outputs recognition features; then, input the recognition features into the action classification module, and the module outputs the action classification result.
[0009] The FISNet model based on calculating inter-frame difference information includes a 3D convolutional layer, a max pooling layer, and multiple inter-frame difference information extraction modules. The multiple inter-frame difference information extraction modules are cascaded in sequence. The input of the FISNet model is used as the input of the 3D convolutional layer. After passing through the max pooling layer, the 3D convolutional layer is connected to the first inter-frame difference information extraction module. The output of the last inter-frame difference information extraction module is used as the output of the FISNet model based on calculating inter-frame difference information.
[0010] Each of the inter-frame difference information extraction modules includes a connected residual layer and a temporal feature extraction module.
[0011] The temporal feature extraction module contains 3 dilated convolutional layers, which are respectively denoted as the first - third dilated convolutional layers. The dilation rates of the first - third dilated convolutional layers are different. The input of the temporal feature extraction module is used as the input of the first - third dilated convolutional layers respectively. After subtracting the outputs of the first - third dilated convolutional layers frame by frame along the temporal dimension respectively, the corresponding inter-frame difference feature maps are obtained respectively. After being processed by sigmoid, the 3 inter-frame difference feature maps are multiplied by the input of the temporal feature extraction module respectively to obtain 3 intermediate feature maps. The 3 intermediate feature maps are weighted and summed according to the weights corresponding to the 3 intermediate feature maps to obtain the final video feature map and used as the output of the temporal feature extraction module, where the weight corresponding to each intermediate feature map is a learnable parameter.
[0012] The dilation rates of the first - third dilated convolutional layers are (1, 1, 1), (2, 1, 1), and (3, 1, 1) respectively.
[0013] The residual layer includes multiple sequentially connected residual modules.
[0014] Each of the residual modules includes multiple convolutional blocks. The input of the residual module is used as the input of the first convolutional block and the input of the fourth convolutional block. After passing through the second convolutional block, the first convolutional block is connected to the third convolutional block. The output of the third convolutional block is added to the output of the fourth convolutional block, and the result is used as the output of the residual module.
[0015] The action classification module includes a global average pooling layer, a Dropout layer, and a classifier that are connected in sequence. The input of the action classification module serves as the input of the global average pooling layer, and the output of the classifier serves as the output of the action classification module.
[0016] II. A computer device
[0017] The device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the video action recognition method based on calculating frame difference information with different frame intervals are implemented.
[0018] III. A computer-readable storage medium
[0019] The medium stores a computer program, and when the computer program is executed by a processor, the steps of the video action recognition method based on calculating frame difference information with different frame intervals are implemented.
[0020] The beneficial effects of the present invention are as follows:
[0021] The present invention proposes a time feature extraction module, which is a plug-and-play module. The time feature extraction module aims to obtain weight scores based on the differential features at a certain frame interval, more effectively locate the motion difference information between two frames, and thereby strengthen the position features of the space where the human action is located, so as to realize the secondary enhancement of the space features based on the time features of the video. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a network framework diagram of the method of the present invention.
[0023] Figure 2 It is a schematic structural diagram of the time feature extraction module used in the present invention.
[0024] Figure 3 It is a schematic structural diagram of the residual module in the network of the present invention.
[0025] Figure 4 It is the classification accuracy rate of the present invention and Resnet on the HMDB51 dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The present invention will be further described below with reference to the drawings and embodiments.
[0027] The embodiments of the present invention and their implementation processes and situations are as follows:
[0028] The dataset of the present invention uses the HMDB51 dataset. HMDB51 contains 51 types of actions, with a total of 6,849 videos. Each action contains at least 51 videos, with a resolution of 320*240, from YouTube, Google videos, etc., and the total size is 2G.
[0029] The method proposed by the present invention includes the following steps:
[0030] First, preprocess the original video for image frames to obtain the preprocessed video. The preprocessing of image frames specifically normalizes the size of the read original input video image into a 3-channel RGB image of 64×112×112. 64×112×112 is used as the input size of the neural network. Then, standardize the 3-channel RGB image, mapping the 3-channel RGB image from an integer between 0 and 255 to a floating-point number between 0 and 1. Then, input the preprocessed video into the FISNet model based on calculating inter-frame difference information. The model outputs recognition features containing spatial information, motion information, and motion detail information, with a size of 32×112×112. These recognition features have achieved good results in the action recognition and classification task. Then, input the recognition features containing spatial information, motion information, and motion detail information into the action classification module, and the module outputs the action classification result.
[0031] As Figure 1 shown, the FISNet model based on calculating inter-frame difference information includes a 3D convolutional layer, a max pooling layer, and multiple inter-frame difference information extraction modules. The multiple inter-frame difference information extraction modules are cascaded in sequence. The input of the FISNet model is used as the input of the 3D convolutional layer. After passing through the max pooling layer, the 3D convolutional layer is connected to the first inter-frame difference information extraction module. The output of the last inter-frame difference information extraction module is used as the output of the FISNet model based on calculating inter-frame difference information.
[0032] Each inter-frame difference information extraction module includes a connected residual layer and a temporal feature extraction module (FIS module). As Figure 2As shown in the figure, the time feature extraction module includes three dilated convolutional layers, which are respectively denoted as the first - third dilated convolutional layers. The dilation rates of the first - third dilated convolutional layers are different, which are (1, 1, 1), (2, 1, 1), and (3, 1, 1) respectively. The input of the time feature extraction module is respectively used as the input of the first - third dilated convolutional layers, and the first - third dilated convolutional layers output video feature maps with different frame intervals; then, after subtracting each frame along the time dimension from the outputs of the first - third dilated convolutional layers respectively, the corresponding frame - difference feature maps are obtained. Specifically, for the output of the first dilated convolutional layer, subtract the second frame from the first frame in its output, subtract the third frame from the second frame, and so on until the subtraction is completed through traversal, to obtain the corresponding frame - difference feature map; after the three frame - difference feature maps are processed by sigmoid and then multiplied by the input of the time feature extraction module (i.e., the original video feature) respectively, three intermediate feature maps are obtained. According to the weights corresponding to the three intermediate feature maps, the three intermediate feature maps are weighted and summed to obtain the final video feature map, which is used as the output of the time feature extraction module, where the weight corresponding to each intermediate feature map is a learnable parameter.
[0033] The residual layer includes multiple sequentially connected residual modules. As Figure 3 shown, each residual module includes a convolutional block and a spatial attention layer. The input of the residual module is used as the input of the first convolutional block and the input of the fourth convolutional block. The first convolutional block is connected to the third convolutional block after passing through the second convolutional block. The first and second convolutional blocks include a convolutional layer, an activation layer, and a batch normalization layer connected in sequence. The third convolutional block includes a connected convolutional layer and a batch normalization layer. The fourth convolutional block is a convolutional layer. The result after adding the output of the third convolutional block and the output of the fourth convolutional block is used as the output of the residual module.
[0034] The numbers of residual modules included in the four residual layers are 3, 4, 6, and 3 respectively. Among them, the first residual module in each residual layer performs a downsampling operation, and its purpose is to reduce the dimension of the features. Each residual module includes a residual mapping and an identity mapping. Among them, the input of each residual module is denoted as the input feature tensor. After the input feature tensor is extracted by the residual mapping, the first spatio - motion feature tensor is obtained. At the same time, after the input feature tensor passes through the identity mapping, the second spatio - motion feature tensor is obtained. The output feature tensor of the residual module is obtained by adding the first spatio - motion feature tensor and the second spatio - motion feature tensor. The specific formula is as follows:
[0035] H(x) = F(x) + G(x)
[0036] where, H() is the output function of the residual module; F() is the residual mapping function; G() is the identity mapping function; x is the input feature tensor of the residual module.
[0037] In the identity mapping, it is determined whether the number of channels of the output feature tensor (i.e., the first feature tensor) of the residual mapping is the same as that of the input feature tensor. If they are the same, the input feature tensor is directly used as the output feature tensor of the identity mapping. If they are different, the feature tensor after pointwise convolution of the input feature tensor is used as the output feature tensor of the identity mapping, that is, the second feature tensor. It can be set by the following formula:
[0038]
[0039] where G() is the identity mapping function; is a convolution function with a convolution kernel size of 1×1×1 and an output channel of C.
[0040] The action classification module includes a global average pooling layer, a Dropout layer, and a classifier connected in sequence. The input of the action classification module is used as the input of the global average pooling layer, and the output of the classifier is used as the output of the action classification module. The classifier is a fully connected layer. In this embodiment, the mapping in the classifier is the probability values for 51 categories, and the one with the largest probability value is taken as the action category of the video image, which is processed by the following formula:
[0041] O(x) = Max(Linear(Dropout(GlobalAvgPool(x))))
[0042] where O() is the network output function; Linear() is the fully connected function; Dropout() is the regularization function that randomly discards neurons with a given probability; Max() is the maximum value function, GlobalAvgPool() is the global average pooling function, and x is the input feature tensor.
[0043] The HMDB51 dataset is divided into a training set and a test set. The FISNet model based on calculating inter-frame difference information and the action classification module are trained using the training set until the training is completed. Then, the accuracy of the trained FISNet model based on calculating inter-frame difference information and the action classification module is tested using the test set, and the results are as Figure 4 shown. It can be seen from the figure that the classification accuracy of the FISNet model based on calculating inter-frame difference information proposed by the present invention has been improved to a certain extent compared with the original Resnet. Table 1 is a comparison table of the TOP1 classification accuracy of the method proposed by the present invention and other existing methods. It can also be seen from Table 1 that the method proposed by the present invention introduces the time feature extraction module FIS and is also superior to the existing methods in terms of TOP1 classification accuracy.
[0044] Table 1 is a comparison table of the TOP1 classification accuracy of the method proposed by the present invention and other existing methods
[0045] Method TOP1 Resnet (baseline network) 70.12 DMC-Net 71.8 Prob-Distill 72.0 TVNet+IDT 72.6 R[2+1D]D-TwoStream 72.7 FISnet (the present invention) 74.01
[0046] Finally, it should be noted that the above embodiments and descriptions are only used to illustrate the technical solutions of the present invention rather than to limit them. Those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced. Without departing from the spirit and scope of the disclosure of the technical solutions of the present invention, they should all be covered by the protection scope of the claims of the present invention.
Claims
1. A video action recognition method based on calculating frame difference information with different frame intervals, characterized in that The method comprises the following steps: First, preprocess the image frames of the original video to obtain the preprocessed video; then, input the preprocessed video into the FISNet model based on calculating inter-frame difference information, and the model outputs recognition features; Next, input the recognition features into the action classification module, and the module outputs the action classification result.
2. The video action recognition method based on calculating frame difference information with different frame intervals according to claim 1, characterized in that, The FISNet model based on calculating inter-frame difference information includes a 3D convolutional layer, a max pooling layer, and multiple inter-frame difference information extraction modules. The multiple inter-frame difference information extraction modules are cascaded in sequence. The input of the FISNet model serves as the input of the 3D convolutional layer. After passing through the max pooling layer, the 3D convolutional layer is connected to the first inter-frame difference information extraction module, and the output of the last inter-frame difference information extraction module serves as the output of the FISNet model based on calculating inter-frame difference information.
3. A video action recognition method based on calculating frame difference information with different frame intervals according to claim 2, characterized in that, Each of the inter-frame difference information extraction modules includes a connected residual layer and a temporal feature extraction module.
4. A video action recognition method based on calculating frame difference information with different frame intervals according to claim 3, characterized in that, The temporal feature extraction module includes 3 dilated convolutional layers, which are respectively denoted as the first - third dilated convolutional layers. The dilation rates of the first - third dilated convolutional layers are different. The input of the temporal feature extraction module serves as the input of the first - third dilated convolutional layers respectively. After performing frame-by-frame subtraction along the temporal dimension on the outputs of the first - third dilated convolutional layers respectively, the corresponding inter-frame difference feature maps are obtained respectively; After being processed by sigmoid, the 3 inter-frame difference feature maps are respectively multiplied by the input of the temporal feature extraction module to obtain 3 intermediate feature maps. The 3 intermediate feature maps are weighted and summed according to the weights corresponding to the 3 intermediate feature maps to obtain the final video feature map, which serves as the output of the temporal feature extraction module. Among them, the weight corresponding to each intermediate feature map is a learnable parameter.
5. A video action recognition method based on calculating frame difference information with different frame intervals according to claim 4, characterized in that, The dilation rates of the first - third dilated convolutional layers are (1, 1, 1), (2, 1, 1), and (3, 1, 1) respectively.
6. A video action recognition method based on calculating frame difference information with different frame intervals according to claim 3, characterized in that The residual layer includes multiple sequentially connected residual modules.
7. A video action recognition method based on calculating frame difference information with different frame intervals according to claim 6, characterized in that Each of the residual modules includes multiple convolutional blocks. The input of the residual module serves as the input of the first convolutional block and the input of the fourth convolutional block. The first convolutional block is connected to the third convolutional block after passing through the second convolutional block. The output of the third convolutional block and the output of the fourth convolutional block are added together, and the result serves as the output of the residual module.
8. A video action recognition method based on calculating frame difference information with different frame intervals according to claim 7, characterized in that The action classification module includes a globally average pooling layer, a Dropout layer, and a classifier connected in sequence. The input of the action classification module serves as the input of the globally average pooling layer, and the output of the classifier serves as the output of the action classification module.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for video action recognition based on calculating frame difference information of different frame intervals according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for video action recognition based on calculating frame difference information of different frame intervals according to any one of claims 1 to 8.