A video action recognition method based on a two-stream network

The optical flow and space-time characteristics are extracted through the TVL1 algorithm and the residual network, combined with the attention mechanism and gating weight, the problem of insufficient fusion of motion information and space-time information in the dual-stream network is solved, and the accuracy of video action recognition is improved.

CN116189292BActive Publication Date: 2025-07-18CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310014498.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-05
Publication Date
2025-07-18
Estimated Expiration
2043-01-05

AI Technical Summary

Technical Problem

The existing dual-stream network fails to effectively integrate motion information and space-time information in video action recognition, resulting in inaccurate judgment of action type.

Method used

The optical flow data is extracted using the TVL1 algorithm, the residual network is used to extract spatiotemporal and motion characteristics, and the feature fusion is performed through the attention mechanism and gating weights, and the shared features are obtained, and finally action recognition is performed.

Benefits of technology

It realizes accurate judgment of action types in various video action scenarios, improves the information fusion and sharing of RGB and optical flow characteristics, and improves the accuracy of action recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189292B_ABST
    Figure CN116189292B_ABST
Patent Text Reader

Abstract

The present invention discloses a video action recognition method based on a two-stream network, belonging to the technical field of video image processing. It includes extracting a number of image frames from a video dataset to be recognized, and using the TVL1 algorithm to extract optical flow data from the image frames; using a residual network to extract spatio-temporal features from the image frames, and at the same time using a residual network to extract motion features from the optical flow data; adopting a residual fusion of the spatio-temporal features and the motion features based on an attention mechanism to obtain attention fusion features; calculating gating weights according to the attention fusion features, spatio-temporal features and motion features, and performing global feature extraction on the attention fusion features to obtain shared features; weighting the shared features and then fusing them with the spatio-temporal features and motion features respectively to obtain new spatio-temporal features and motion features, thereby completing the action recognition of the video dataset. The present invention improves the feature extraction of RGB and optical flow in the traditional two-stream method, and realizes the effective fusion and sharing of information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video image processing, and particularly relates to a video action recognition method based on a two-stream network. Background Art

[0002] In traditional video action recognition methods, it is necessary to manually mark optical flow features. However, in real life, videos are highly variable, and the overhead of manual marking is too large in all aspects. With the rise of deep learning methods, deep learning methods, driven by data, gradually exceed traditional methods in terms of recognition accuracy and speed, and traditional methods are gradually replaced by deep learning.

[0003] Currently, the mainstream deep learning methods in the field of video action recognition can be roughly divided into three categories: one is the two-stream network, which uses the input optical flow to compensate for motion information. Another is the 3D convolutional network, which adds one dimension on the basis of 2D convolution to model video temporal information. The last one is the network that uses a uniquely designed module to simulate and model the motion and spatio-temporal information of the video, and embeds the module into a deep network to achieve the purpose of feature extraction.

[0004] Although the two-stream network adds the input of optical flow to model the motion information of the video, the two-stream network is still essentially two independent networks. Currently, various fusion methods of the two-stream network still do not well fuse the spatio-temporal information and motion information modeled by the two networks. Since 3D convolution adds one dimension on the basis of 2D, it can better model video features, but it also brings a large parameter burden and cannot be well applied in real life. The uniquely designed module can replace the optical flow to model the motion and spatio-temporal information, making the network more lightweight. However, compared with the optical flow input, its ability to model motion information is insufficient.

[0005] In video action recognition, since actions in the video are continuous and each frame input of the video is related, the result of the fusion of motion information and spatio-temporal information has a great impact on the judgment of the action type in the video.

[0006] Therefore, how to fuse the motion information and spatio-temporal information in the video during modeling and accurately judge the action type that appears in the video is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] In view of this, the present invention provides a video action recognition method based on a two-stream network, which can accurately judge the action type that appears in the video when a video is given.

[0008] To achieve the above object, the present invention adopts the following technical solutions:

[0009] A video action recognition method based on a two-stream network, comprising the following steps:

[0010] Extract a plurality of consecutive image frames from the video dataset to be recognized, and extract optical flow data from the image frames by using the TVL1 algorithm;

[0011] Use a residual network to extract spatio-temporal features from the image frames, and at the same time use a residual network to extract motion features from the optical flow data;

[0012] Use an attention mechanism to perform residual fusion on the spatio-temporal features and motion features to obtain attention fusion features;

[0013] Calculate gating weights according to the attention fusion features, spatio-temporal features and motion features, and perform global feature extraction on the attention fusion features to obtain shared features;

[0014] Fuse the shared features weighted by the gating weights with the spatio-temporal features to obtain new spatio-temporal features, and fuse the shared features weighted by the gating weights with the motion features to obtain new motion features;

[0015] Obtain the action recognition result of the video dataset to be recognized according to the new spatio-temporal features and new motion features.

[0016] Preferably, extract a plurality of consecutive image frames from the video dataset to be recognized, and extract optical flow data from the image frames by using the TVL1 algorithm, specifically including:

[0017] Obtain the dataset to be recognized, and extract b consecutive image frames from the dataset to be recognized;

[0018] Use the TVL1 algorithm to perform optical flow extraction on adjacent image frames to obtain b - 1 frames of optical flow data.

[0019] Preferably, use a residual network to extract spatio-temporal features from the image frames, and at the same time use a residual network to extract motion features from the optical flow data, specifically including:

[0020] Input the extracted plurality of consecutive image frames into a residual network ResNet-50 with a first convolutional channel of 3 to obtain spatio-temporal features;

[0021] Input the optical flow data into a residual network ResNet-50 with a first convolutional channel of 2 to obtain motion features.

[0022] Preferably, use an attention mechanism to perform residual fusion on the spatio-temporal features and motion features to obtain attention fusion features, specifically including:

[0023] Block the spatio-temporal features and motion features respectively using 3D convolution to obtain a spatio-temporal feature cube sequence and a motion feature cube sequence;

[0024] Linearly map the spatio-temporal feature cube sequence and the motion feature cube sequence respectively to obtain a spatio-temporal marker sequence and a motion marker sequence;

[0025] Linearly map the spatio-temporal marker sequence as the attention query, linearly map the motion marker sequence once as the attention key, and linearly map the motion marker sequence twice as the attention value;

[0026] Perform a scaled dot product operation on the attention query and the attention key to obtain an attention distribution;

[0027] Multiply the attention distribution and the attention value dimension-wise, and perform a residual connection on the multiplication result and the attention query;

[0028] Perform dimension reconstruction and upsampling on the residual connection result to make the dimension of the residual connection result consistent with the spatio-temporal features and motion features, obtaining an attention fusion feature.

[0029] Preferably, calculate gating weights based on the attention fusion feature, spatio-temporal features, and motion features, and perform global feature extraction on the attention fusion feature to obtain a shared feature, specifically including:

[0030] Perform a 2D convolutional mapping on the attention fusion feature, spatio-temporal features, and motion features respectively and then add them together, activate the addition result through a Sigmoid activation function, and split the activation result into a spatio-temporal gating weight and a motion gating weight by channel averaging;

[0031] Perform global feature extraction on the attention fusion feature through a layer of ConvLSTM network to obtain a shared feature.

[0032] Preferably, fuse the shared feature weighted by the gating weight with the spatio-temporal features to obtain new spatio-temporal features, and fuse the shared feature weighted by the gating weight with the motion features to obtain new motion features, specifically including:

[0033] Multiply the shared feature by the spatio-temporal gating weight to obtain a spatio-temporally weighted shared feature;

[0034] Convert the dimension and number of channels of the spatio-temporally weighted shared feature to make the dimension of the spatio-temporally weighted shared feature consistent with the dimension of the spatio-temporal features, and at the same time make the number of channels of the spatio-temporally weighted shared feature consistent with the number of channels of the spatio-temporal features;

[0035] The spatio-temporal weighted shared feature after converting the dimension and the number of channels is added to the spatio-temporal feature to obtain a new spatio-temporal feature;

[0036] The shared feature is multiplied by the motion gating weight to obtain a motion-weighted shared feature;

[0037] Convert the dimension and the number of channels of the motion-weighted shared feature so that the dimension of the motion-weighted shared feature is consistent with the dimension of the motion feature, and at the same time make the number of channels of the motion-weighted shared feature and the number of channels of the motion feature consistent;

[0038] The motion-weighted shared feature after converting the dimension and the number of channels is added to the motion feature to obtain a new motion feature.

[0039] Preferably, the action recognition result of the video dataset to be recognized is obtained according to the new spatio-temporal feature and the new motion feature, specifically including:

[0040] The new spatio-temporal feature and the new motion feature are respectively input into the fully connected layer after global pooling and stretching to obtain two classification prediction scores;

[0041] And the two classification prediction scores are averaged and fused to obtain the final recognition result.

[0042] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a video action recognition method based on a two-stream network, which has the following beneficial effects:

[0043] The present invention provides a method for fusing and modeling motion information and spatio-temporal information in a video, which can realize better judgment of the type of action in various different video action scenarios.

[0044] The present invention improves the feature extraction of RGB and optical flow in the traditional two-stream method, and realizes the effective fusion and sharing of their information. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0046] Figure 1 It is a schematic flow chart of the steps of the video action recognition method based on a two-stream network provided by the present invention;

[0047] Figure 2 It is a schematic network structure diagram of the video action recognition method based on a two-stream network provided by the present invention;

[0048] Figure 3 Schematic diagram of the attention mechanism residual fusion process provided by the present invention;

[0049] Figure 4 Schematic diagram of the gated weight acquisition process provided by the present invention. Detailed implementation manners

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0051] As Figure 1 shown, an embodiment of the present invention discloses a video action recognition method based on a two-stream network, including the following steps:

[0052] Extract a plurality of consecutive image frames from the video dataset to be recognized, and extract optical flow data from the image frames by using the TVL1 algorithm;

[0053] Extract spatio-temporal features from the image frames by using a residual network, and at the same time extract motion features from the optical flow data by using a residual network;

[0054] Use the attention mechanism to perform residual fusion on the spatio-temporal features and motion features to obtain attention fusion features;

[0055] Calculate the gated weights according to the attention fusion features, spatio-temporal features and motion features, and perform global feature extraction on the attention fusion features to obtain shared features;

[0056] Fuse the shared features weighted by the gated weights with the spatio-temporal features to obtain new spatio-temporal features, and fuse the shared features weighted by the gated weights with the motion features to obtain new motion features;

[0057] Obtain the action recognition result of the video dataset to be recognized according to the new spatio-temporal features and new motion features.

[0058] The steps of the present invention will be further described below in conjunction with more specific implementation manners.

[0059] S1. Obtain the video dataset to be recognized; extract 9 consecutive image frames from the video dataset, and adjust all the image frames to a unified preset size; at the same time, obtain optical flow data for all 9 frames of images through the TVL1 algorithm. The TVL1 algorithm extracts one frame of optical flow for every two frames of images, so 8 frames of optical flow data are obtained;

[0060] S2. The first 8 frames of the extracted images are directly input into the Residual Network ResNet50 for spatio-temporal feature extraction, and the 8-frame optical flow extracted is input into the Residual Network ResNet50 for motion feature extraction, so as to obtain the spatio-temporal feature and motion feature of the fourth layer with ResNet50 as the backbone network. In the embodiment of the present invention, the fourth layer refers to the fourth stage part of the residual network, and the ResNet50 fourth stage structure includes residual blocks of bottleneck block. In this embodiment, only the first 8 frames of 9 consecutive image frames are input when extracting spatio-temporal features, and all 9 consecutive images can also be input in other embodiments.

[0061] In this embodiment, both spatio-temporal feature extraction and motion feature extraction use ResNet50. The difference is that when extracting spatio-temporal features, the first convolution channel of ResNet50 is selected as 3, and when extracting motion features, the first convolution channel of ResNet50 is selected as 3.

[0062] S3. Use the attention mechanism to perform residual fusion on the spatio-temporal feature and the motion feature to obtain the attention fusion feature:

[0063] In this step, the obtained spatio-temporal feature is used as the attention query, and the motion feature is used as the attention key value to participate in the attention calculation to obtain the output, and the attention query is residually connected with the output to complete the first feature fusion, so as to obtain the preliminary fusion feature;

[0064] The specific process is as follows:

[0065] S31. The spatio-temporal feature and the motion feature are each divided into multiple cubic blocks through a three-dimensional convolution, and the obtained cubic block sequences are linearly projected respectively to obtain the corresponding spatio-temporal token sequence and motion token sequence;

[0066] In this step, the block division can be performed through the following formula to obtain the spatio-temporal feature token sequence and the motion feature token sequence:

[0067] Ts = Conv3d(x)

[0068] In the above formula, the kernel size of the three-dimensional convolution is (1, 16, 16), the stride is (1, 4, 4), and the padding is (1, 7, 7); Ts represents the tensor of the token sequence; x represents the input of consecutive image frames in the form of (N, C, T, H, W), where N represents the sample number batch_size, C is the number of channels, T represents the time series length, H represents the height of the picture, and W represents the width of the picture;

[0069] S32. Linearly map the spatio-temporal token sequence as the attention query, linearly map the motion token sequence once as the attention key, and linearly map the motion token sequence twice as the attention value;

[0070] S33. Perform a scaled dot-product operation on the attention query and the attention key to obtain the attention distribution

[0071] S34. Multiply the attention distribution and the motion token sequence dimension-wise, perform a residual connection on the multiplication result and the attention query, and finally reconstruct the token sequence to the original feature size to obtain the final attention fusion feature;

[0072] In this process, the attention fusion feature can be expressed by the following formula:

[0073]

[0074] In the above formula, Attention(Q, K, V) represents the attention calculation result; Q represents the attention query generated by linearly mapping the spatio-temporal token sequence; K, V represent the attention key and value generated by linearly mapping the motion token sequence respectively; D represents the channel dimension, that is, the embedding dimension of the sequence.

[0075] S4. Calculate the gating weights based on the attention fusion feature, spatio-temporal feature, and motion feature, and perform global feature extraction on the attention fusion feature to obtain the shared feature;

[0076] This step mainly includes two aspects: gating weight calculation and shared feature extraction:

[0077] Among them, the gating weight calculation process is as follows:

[0078] Perform a 2D convolutional mapping on the attention fusion feature, spatio-temporal feature, and motion feature respectively, add the results, activate the addition result through the Sigmoid activation function, and split the activation result into the spatio-temporal gating weight and the motion gating weight on average by channel;

[0079] The specific calculation process of the gating weight can be expressed by the following formula:

[0080] GW1, GW2 = split(σ(Conv2d(R m ) + Conv2d(F m ) + Conv2d(M1)))

[0081] In the formula, GW1 and GW2 represent the generated spatio-temporal gating weight and motion gating weight respectively; split represents the split operation by channel dimension; σ represents the Sigmoid activation function; Conv2d represents the scale-invariant convolutional operation; R mRepresents the intermediate spatio-temporal feature; F m Represents the intermediate motion feature; M1 represents the attention fusion feature.

[0082] Shared feature extraction, the process is as follows:

[0083] The attention fusion feature is passed through a layer of ConvLSTM network for further global feature extraction to obtain the shared feature M2;

[0084] S5. After gating weight weighting the shared feature, it is fused with the spatio-temporal feature to obtain a new spatio-temporal feature, and after gating weight weighting the shared feature, it is fused with the motion feature to obtain a new motion feature:

[0085] In this step, the two gating weights obtained in step S4 are respectively multiplied by the shared feature to obtain two weighted shared features, and they are pooled to change their dimensions to 7x7 and the number of channels is changed to be consistent with the spatio-temporal feature and the motion feature through a two-dimensional convolution with a size of 1;

[0086] The above processing results are respectively added to the spatio-temporal feature and the motion feature to obtain a new spatio-temporal feature and a new motion feature;

[0087] The specific steps of gating weight fusion can be expressed by the following formula:

[0088] R n = R + Conv2d(AvgPool(GW1 * M2))

[0089] F n = F + Conv2d(AvgPool(GW2 * M2))

[0090] In the formula, R n Represents the new spatio-temporal feature; F n Represents the new motion feature; GW1 and GW2 respectively represent the generated spatio-temporal gating weight and motion gating weight; Conv2d represents a scale-invariant convolution operation; AvgPool is global pooling; R represents the spatio-temporal feature obtained in step S2; F represents the motion feature obtained in step S2; * represents a generalized multiplication operation; M2 represents the shared feature.

[0091] S6. The new spatio-temporal feature and the new motion feature are each passed through global pooling and stretching and then input into a fully connected layer for classification to obtain two classification prediction scores, and the two classification prediction scores are averaged and fused to obtain the final recognition result of the video dataset to be recognized;

[0092] The final recognition result of the video dataset to be recognized can be expressed by the following formula:

[0093] Class = AvgFusion{FCr (R n ),FC f (F n )}

[0094] In the above formula, R n represents the new spatio-temporal feature; F n represents the new motion feature; FC r , FC f are the fully connected layers of the spatio-temporal network and the motion network respectively; AvgFusion represents the average fusion operation. Class represents the finally predicted category.

[0095] During training, this example uses the cross-entropy between the classification prediction score and the true action category label as the loss function, and at the same time takes different weights for the two streams to obtain the overall loss:

[0096] Furthermore, the overall loss function in the embodiment of the present invention is:

[0097] L = L r + λL f

[0098] In the above formula, L r represents the cross-entropy loss between the classification prediction score of the spatio-temporal network and the true action category label score; L f represents the cross-entropy loss between the classification prediction score of the motion network and the true action category label; λ represents the weight coefficient. In this embodiment, λ is a hyperparameter set manually, and λ is set to 2.

[0099] In this embodiment, the classification prediction score represents the result after passing through the softmax of the output after classification by the fully connected layer; the true action category label represents the digital encoding representation of the action category in the code;

[0100] Explanation: Video action recognition is a classification problem, and the loss function is the cross-entropy loss function. Here, since both the spatio-temporal network and the motion network make predictions, there are two cross-entropy losses in the present invention. Therefore, a weighted sum is used as the total loss function, and its role is to train the spatio-temporal network and the motion network simultaneously;

[0101] By continuously reducing the loss between the prediction score and the true label, the best mapping relationship between the video frame and the action category can be obtained.

[0102] Furthermore, L r and L f can be calculated through the following formula:

[0103]

[0104] In the above formula, y represents the true label, Let \(y\) denote the predicted label, \(i\) denote the serial number of the current label, and \(n\) denote the total number of labels in the sample.

[0105] The following specifically elaborates on the video action recognition method:

[0106] After batch loading the dataset videos, 9 image frames are extracted from each video, and the sizes of all images are uniformly adjusted to \(224\times224\times3\). Among the extracted images, the first 8 frames are directly input into the spatio-temporal feature extraction network, and the full 9 frames are used to extract 8 frames of optical flow through the TVL1 algorithm and input into the motion feature extraction network, obtaining the intermediate spatio-temporal features and intermediate motion features of the fourth layer, as well as the spatio-temporal features and motion features of the last layer. After operations such as convolution and pooling in the spatio-temporal feature network and motion feature network, the sizes of the intermediate spatio-temporal features and intermediate motion features are \(14\times14\times1024\), and the sizes of the spatio-temporal features and motion features are \(7\times7\times2048\).

[0107] Specifically: After batch loading the dataset videos, 9 image frames are extracted from each video, and the sizes of all images are uniformly adjusted to \(224\times224\times3\). Among the images extracted from each video, the first 8 frames are input into the original ResNet-50 network to obtain the intermediate spatio-temporal features of \(14\times14\times1024\) and the spatio-temporal features of \(7\times7\times2048\). After extracting 8 frames of optical flow from the full 9 frames of images through the TVL1 optical flow extraction algorithm to obtain a tensor of size \(224\times224\times2\), based on the modified ResNet-50 network, the intermediate motion features of \(14\times14\times1024\) and the motion features of \(7\times7\times2048\) are obtained.

[0108] Since there are complex correlations between the extracted intermediate spatio-temporal features and intermediate motion features, it is difficult to explicitly fuse the relevant information of the two using ordinary channel concatenation or summation operations. To effectively filter and fuse the relevant information, a feature fusion method based on the attention mechanism is adopted in the shared module.

[0109] Given the intermediate spatio-temporal feature \(R\) m , \(F\) m \(\in\mathbb{R}\) C*H*W , where \(c\) represents the number of channels, and \(H\times W\) (height * width) represents the size of the intermediate feature.

[0110] Referring to Figure 2 and Figure 3 as shown, Figure 2 Attention Fusion in

[0111] Referring to Figure 3 as shown, the specific process of Attention Fusion is: the intermediate spatio-temporal feature \(R\)m and the intermediate motion feature F m respectively pass through a Conv3d(1024, 256, (1, 7, 7), (1, 7, 7)) (1024 refers to the number of channels of the input convolutional layer, 256 refers to the number of channels of the output convolutional layer, the first (1, 7, 7) refers to the size of the convolutional kernel in the three dimensions of time, height, and width, and the second (1, 7, 7) refers to the three-dimensional stride; setting the convolutional kernel size and stride to be the same as the block size to achieve the purpose of block division) for block embedding to obtain a spatio-temporal feature marker sequence and a motion feature marker sequence with a dimension of 8 * 8 * 8 * 256, merge the first three dimensions of the above two sequences to obtain a spatio-temporal feature marker sequence and a motion feature marker sequence with a dimension of 512 * 256, perform multi-head attention calculation based on this, and perform a residual connection between the attention query and the calculation output to complete the first feature fusion, obtaining a preliminary fusion feature with a dimension of 512 * 256.

[0112] The calculation process of the two-stream feature fusion based on the attention mechanism is as follows:

[0113] Perform block embedding on the two types of intermediate features respectively through the following formula:

[0114] p(x) = Conv3d(x),

[0115] In the above formula, p(x) represents the block embedding operation; x represents the intermediate feature.

[0116] Perform a linear mapping on the intermediate feature after block embedding through the following formula:

[0117] Q = R m W q , K = F m W k , V = F m W v

[0118] In the above formula, Q, K, V represent the attention query, key, and value; W q , W k , W v represents the three linear matrices corresponding to Q, K, V, with a dimension of 1024 * 256; R m , F m represents the intermediate spatio-temporal feature and the intermediate motion feature.

[0119] Perform residual fusion based on multi-head attention calculation through the following formula:

[0120]

[0121] In the above formula, M1 represents the attention fusion feature; Q represents the attention query generated by linearly mapping the spatio-temporal token sequence; K and V represent the attention key and value generated by linearly mapping the motion token sequence respectively; D represents the channel dimension, that is, the embedding dimension of the sequence, which is 256. Softmax(QK^T / sqrt(D)) in this formula is the attention distribution.

[0122] After obtaining the preliminary attention fusion feature, the shared module, on the one hand, further extracts the global feature through a ConvLSTM layer and outputs the shared feature while keeping the size unchanged. On the other hand, it calculates the gating weights to weight the shared feature and then adds it to the spatio-temporal feature and the motion feature for the second feature fusion to obtain the new spatio-temporal feature and motion feature.

[0123] Given the spatio-temporal feature and the motion feature R, F ∈ R c*H*W , where c represents the number of channels, and H*W (height * width) represents the size of the feature.

[0124] Refer to Figure 4 As shown, Gated Fusion represents the two-stream feature fusion based on gating weighting, where M2 refers to the shared feature, R n , F n refer to the new spatio-temporal feature and motion feature.

[0125] Refer to Figure 4 As shown, Gated Fusion is: R m , F m and M1 each perform a convolution transformation and then add them together (the dimension of all convolution kernels is 3*3*1024*2, and the channel output dimension of the convolution kernel is 2 to reduce the number of parameters), and after activation through the Sigmoid function, two gating weights are obtained by splitting by channel, with the dimension of 14*14*1. Multiply the two gating weights by the shared feature (the dimension is 14*14*1024, and due to the mismatch of the channel dimension, PyTorch will automatically perform broadcasting to repeat the gating weights 1024 times in the channel dimension and then multiply) to obtain two weighted shared features, and perform pooling on them to change the spatial dimension to 7*7 and transform the number of channels to be the same as that of the spatio-temporal feature and the motion feature through a two-dimensional convolution with a size of 1, and the final dimension is 7*7*2048.

[0126] Finally, the new spatio-temporal feature and the motion feature are respectively subjected to global pooling, stretching, and then input into the fully connected layer for classification, and the two classification prediction scores are averaged and fused to obtain the recognition result of the video dataset to be recognized.

[0127] During training, this example uses the cross-entropy between the prediction score and the true label as the loss function, and at the same time, different weights are used to weight the two streams to obtain the overall loss.

[0128] By continuously reducing the loss between the predicted score and the true label, the best mapping relationship between video frames and action categories can be obtained.

[0129] The picture contains rich information, and multi-modal picture input can mine the deep semantic information of the picture. After introducing the optical flow modal information, the original two-stream network has achieved good results in the video action recognition task. However, the spatio-temporal feature extraction network and the motion feature extraction network of this method are relatively independent, and do not fuse and mine the common semantic information of multiple modalities, resulting in a lack of information interaction between the two feature extraction networks and not fully exploiting the potential of feature mining and fusion of the two networks. In this embodiment, attention fusion and gating fusion are used to process the intermediate output features of the two feature extraction networks. Attention fusion will notice the common features of the output features of the two networks, and gating fusion will generate weights for the common features. The generated weights will weight the subsequent shared module and the network output, which can not only fuse the common semantic information of multiple models, but also transmit the learned common features to the end of the network through weights to weight the network output feature information.

[0130] Facing the extracted common feature information, if simple multi-layer convolution operations are used, the previously fused common feature information will be lost. Since the input is a video sequence and an action requires multiple frames to be jointly judged, the previous common feature information cannot affect the later input, which will lead to weak long-term dependence in the network. In this embodiment, ConvLSTM is used as the shared module. ConvLSTM takes into account the input in the form of features and has good long-term dependence ability, which enhances the long-term dependence of the common feature information while also enhancing the network's extraction of temporal information.

[0131] Adopting the technical solution of the present invention not only solves the problem that the original two-stream network is relatively independent and does not fully mine the common features of multi-modal feature information, but also enhances the network's extraction of temporal information. It improves the feature extraction of RGB and optical flow in the traditional two-stream method and realizes the fusion and sharing of their information.

[0132] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0133] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video action recognition method based on a two-stream network, characterized in that, The method includes the following steps: Extract a number of consecutive image frames from the video dataset to be recognized, and extract optical flow data from the image frames using the TVL1 algorithm; Extract spatio-temporal features from the image frames using a residual network, and at the same time extract motion features from the optical flow data using a residual network; Use an attention mechanism to perform residual fusion on the spatio-temporal features and motion features to obtain attention fusion features; Calculate gating weights based on the attention fusion features, spatio-temporal features, and motion features, and perform global feature extraction on the attention fusion features to obtain shared features; Fuse the shared features weighted by the gating weights with the spatio-temporal features to obtain new spatio-temporal features, and fuse the shared features weighted by the gating weights with the motion features to obtain new motion features; Obtain the action recognition result of the video dataset to be recognized based on the new spatio-temporal features and new motion features.

2. The video action recognition method based on a two-stream network according to claim 1, wherein Extract a number of consecutive image frames from the video dataset to be recognized, and extract optical flow data from the image frames using the TVL1 algorithm, specifically including: Obtain the dataset to be recognized, and extract b consecutive image frames from the dataset to be recognized; Use the TVL1 algorithm to perform optical flow extraction on adjacent image frames to obtain b-1 frames of optical flow data.

3. The video action recognition method based on a two-stream network according to claim 1, wherein Extract spatio-temporal features from the image frames using a residual network, and at the same time extract motion features from the optical flow data using a residual network, specifically including: Input the extracted number of consecutive image frames into a residual network ResNet-50 with a first convolutional channel of 3 to obtain spatio-temporal features; Input the optical flow data into a residual network ResNet-50 with a first convolutional channel of 2 to obtain motion features.

4. The video action recognition method based on a two-stream network according to claim 1, characterized in that Use an attention mechanism to perform residual fusion on the spatio-temporal features and motion features to obtain attention fusion features, specifically including: Block the spatio-temporal features and motion features respectively using 3D convolution to obtain a sequence of spatio-temporal feature cubes and a sequence of motion feature cubes; Perform linear mapping on the sequence of spatio-temporal feature cubes and the sequence of motion feature cubes respectively to obtain a spatio-temporal marker sequence and a motion marker sequence; Use the linearly mapped spatio-temporal marker sequence as the attention query, use the motion marker sequence linearly mapped once as the attention key, and use the motion marker sequence linearly mapped twice as the attention value; Perform a scaled dot product operation on the attention query and the attention key to obtain an attention distribution; Multiply the attention distribution and the attention value dimension by dimension, and perform a residual connection on the multiplication result and the attention query; Perform dimension reconstruction and upsampling on the residual connection result to make the dimension of the residual connection result consistent with the spatio-temporal features and motion features, and obtain attention fusion features.

5. The video action recognition method based on a two-stream network according to claim 1, characterized in that Calculate gating weights based on the attention fusion features, spatio-temporal feature maps, and motion features, and perform global feature extraction on the attention fusion features to obtain shared features, specifically including: Perform a 2D convolutional mapping on each of the attention fusion feature, spatio-temporal feature, and motion feature, then add them together, and activate the addition result through a Sigmoid activation function. Split the activation result into spatio-temporal gating weights and motion gating weights by channel averaging. Perform global feature extraction on the attention fusion feature through a layer of ConvLSTM network to obtain a shared feature.

6. The video action recognition method based on a two-stream network according to claim 5, wherein Fuse the shared feature weighted by gating weights with the spatio-temporal feature to obtain a new spatio-temporal feature, and fuse the shared feature weighted by gating weights with the motion feature to obtain a new motion feature. Specifically, it includes: Multiply the shared feature by the spatio-temporal gating weights to obtain a spatio-temporally weighted shared feature. Convert the dimensions and number of channels of the spatio-temporally weighted shared feature so that the dimensions of the spatio-temporally weighted shared feature are consistent with the dimensions of the spatio-temporal feature, and at the same time, the number of channels of the spatio-temporally weighted shared feature is consistent with the number of channels of the spatio-temporal feature. Add the spatio-temporally weighted shared feature with the dimensions and number of channels converted to the spatio-temporal feature to obtain a new spatio-temporal feature. Multiply the shared feature by the motion gating weights to obtain a motion-weighted shared feature. Convert the dimensions and number of channels of the motion-weighted shared feature so that the dimensions of the motion-weighted shared feature are consistent with the dimensions of the motion feature, and at the same time, the number of channels of the motion-weighted shared feature is consistent with the number of channels of the motion feature. Add the motion-weighted shared feature with the dimensions and number of channels converted to the motion feature to obtain a new motion feature.

7. The video action recognition method based on a two-stream network according to claim 1, wherein Obtain the action recognition result of the video dataset to be recognized according to the new spatio-temporal feature and the new motion feature. Specifically, it includes: Input the new spatio-temporal feature and the new motion feature into a fully connected layer after global pooling and stretching respectively to obtain two classification prediction scores. And average and fuse the two classification prediction scores to obtain the final recognition result.