Training, recognition method, device and medium of video action recognition model
By using the dual-stream coupled network model of residual attention network, hidden Markov network and fusion model in video action recognition, the problem of difficulty in extracting timing information in video action recognition is solved, and more efficient fusion and recognition effects are achieved.
Patent Information
- Application Number
- CN202210630186.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-06
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-06-06
AI Technical Summary
The prior art is difficult to effectively extract and characterize timing information in video action recognition, resulting in low recognition accuracy and high computational cost of 3D convolutional neural networks.
The dual-current coupled network model of residual attention network model, hidden Markov network model and fusion model are used to extract spatial features through residual attention network, hidden Markov network model timing features, and fusion of spatiotemporal features through fusion model.
It improves the ability to extract and characterize spatiotemporal information in video action recognition, and improves the performance, efficiency and accuracy of recognition.
Smart Images

Figure CN114863570B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer processing technologies, and in particular, to a method for training a video action recognition model, a video action recognition method, an apparatus, and a storage medium. Background Art
[0002] Video understanding is an important part of machine learning and computer vision in fields such as monitoring, security, patient monitoring, and sports video analysis. The recognition of video actions usually requires a sequence of images. It is necessary not only to extract the features of individual key frames, but also to understand the context of the entire video and capture the relationships between key frames in order to obtain a high-precision video action recognition model. Since in the video action recognition task, a complete action in a video is composed of a series of consecutive frames, it often causes redundant temporal information. The 2D convolutional neural network cannot perform the operation of extracting temporal information (the relationship between frames) in spatial feature extraction. Although the 3D convolutional neural network has achieved very good results in the action recognition task, its computational cost is very high. Therefore, a new technical solution for video action recognition is needed, which has better performance in video action recognition. Summary of the Invention
[0003] In view of this, one technical problem to be solved by the present invention is to provide a method for training a video action recognition model, a video action recognition method, an apparatus, and a storage medium.
[0004] According to a first aspect of the present disclosure, there is provided a method for training a video action recognition model, wherein the video action recognition model includes: a residual attention network model, a hidden Markov network model, and a fusion model; the method includes: using the residual attention network model and based on sample frame feature information corresponding to a video sample, obtaining sample frame spatial feature information corresponding to the video sample; generating sample frame fusion feature information according to the sample frame feature information and the sample frame spatial feature information, using the hidden Markov network model and based on the sample frame fusion feature information, obtaining sample frame temporal feature information corresponding to the video sample; using the fusion model to perform fusion and recognition processing on the sample frame spatial feature information and the sample frame temporal feature information, obtaining an action category corresponding to the video sample; adjusting the residual attention network model and the hidden Markov network model according to an overall loss function corresponding to the video action recognition model.
[0005] Optionally, constructing a first loss function corresponding to the residual attention network model; constructing a second loss function corresponding to the hidden Markov network model; generating the overall loss function based on the first loss function, the second loss function, and a corresponding balance coefficient.
[0006] Optionally, constructing the second loss function corresponding to the hidden Markov network model includes: determining the posterior probability information of the hidden Markov network model for processing the sample frame fusion feature information; generating an objective function based on the posterior probability information; constructing the second loss function according to the objective function; wherein, the second loss function is used to characterize the parameter value of the hidden Markov network model when the objective function is at the minimum value.
[0007] Optionally, the residual attention network model includes: an invariant branch sub-model and a variant branch sub-model; obtaining the sample frame spatial feature information corresponding to the video sample by using the residual attention network model and based on the sample frame feature information corresponding to the video sample includes: using the invariant branch sub-model and based on the sample frame feature information, obtaining the invariant branch feature information corresponding to the video sample; using the variant branch sub-model and based on the sample frame feature information, obtaining the variant branch feature information corresponding to the video sample; using a first activation function and based on the invariant branch feature information and the variant branch feature information, generating the sample frame spatial feature information.
[0008] Optionally, the invariant branch sub-model includes: a depthwise separable convolution (DW convolution) layer and a pointwise convolution (PW convolution) layer; obtaining the invariant branch feature information corresponding to the video sample by using the invariant branch sub-model and based on the sample frame feature information includes: using the DW convolution layer and based on the sample frame feature information, obtaining first feature information; inputting the first feature information into the PW convolution layer, and outputting the invariant branch feature information.
[0009] Optionally, the variant branch sub-model includes: a downsampling layer, an upsampling layer, and a variant branch convolution layer; obtaining the variant branch feature information corresponding to the video sample by using the variant branch sub-model and based on the sample frame feature information includes: using the downsampling layer and based on the sample frame feature information, obtaining downsampled feature information; using the upsampling layer and based on the downsampled feature information and the residual information corresponding to the sample frame feature information, obtaining upsampled feature information; using the variant branch convolution layer and based on the upsampled feature information, obtaining convolution feature information; using a second activation function and based on the convolution feature information, generating the variant branch feature information.
[0010] Optionally, obtaining the sample frame temporal feature information corresponding to the video sample by using the hidden Markov network model and based on the sample frame fusion feature information includes: performing temporal feature analysis on the sample frame fusion feature information by using the hidden Markov network model to generate a sample frame feature random vector sequence corresponding to the video sample; and obtaining the sample frame temporal feature information by using the hidden Markov network model and based on the sample frame feature random vector sequence.
[0011] Optionally, the fusion model includes an average pooling layer and a fully connected layer; using the fusion model to perform fusion and recognition processing on the sample frame spatial feature information and the sample frame temporal feature information to obtain the action category corresponding to the video sample includes: using the average pooling layer to perform fusion processing on the sample frame spatial feature information and the sample frame temporal feature information to generate fusion feature information; and obtaining the action category information by using the fully connected layer and based on the fusion feature information.
[0012] Optionally, the video action recognition model includes: a feature extraction model; the method further includes: performing feature extraction processing on each video frame of the video sample by using the feature extraction model to obtain the sample frame feature information; wherein, the feature extraction model includes: a convolutional neural network model.
[0013] According to a second aspect of the present disclosure, there is provided a video action recognition method, including: obtaining a trained video action recognition model; wherein, the video action recognition model is trained by the training method as described above, and the video action recognition model includes: a residual attention network model, a hidden Markov network model, and a fusion model; using the residual attention network model and based on the frame feature information corresponding to the video to be recognized, obtaining the frame spatial feature information corresponding to the video to be recognized; generating frame fusion feature information according to the frame feature information and the frame spatial feature information, and using the hidden Markov network model and based on the frame fusion feature information, obtaining the frame temporal feature information corresponding to the video to be recognized; and using the fusion model to perform fusion and recognition processing on the frame spatial feature information and the frame temporal feature information to obtain the action category corresponding to the video to be recognized.
[0014] Optionally, the residual attention network model includes: an invariant branch sub-model and a variant branch sub-model; the step of using the residual attention network model and based on the frame feature information corresponding to the video to be recognized to obtain the frame spatial feature information corresponding to the video to be recognized includes: using the invariant branch sub-model and based on the frame feature information, obtaining the invariant branch feature information corresponding to the video to be recognized; using the variant branch sub-model and based on the frame feature information, obtaining the variant branch feature information corresponding to the video to be recognized; using a first activation function and based on the invariant branch feature information and the variant branch feature information, generating the frame spatial feature information.
[0015] Optionally, the invariant branch sub-model includes: a depthwise separable convolution layer (DW convolution layer) and a pointwise convolution layer (PW convolution layer); the step of using the invariant branch sub-model and based on the frame feature information, obtaining the invariant branch feature information corresponding to the video to be recognized includes: using the DW convolution layer and based on the frame feature information, obtaining first feature information; inputting the first feature information into the PW convolution layer, and outputting the invariant branch feature information.
[0016] Optionally, the variant branch sub-model includes: a downsampling layer, an upsampling layer, and a variant branch convolution layer; the step of using the variant branch sub-model and based on the frame feature information, obtaining the variant branch feature information corresponding to the video to be recognized includes: using the downsampling layer and based on the frame feature information, obtaining downsampled feature information; using the upsampling layer and based on the downsampled feature information and the residual information corresponding to the frame feature information, obtaining upsampled feature information; using the variant branch convolution layer and based on the upsampled feature information, obtaining convolution feature information; using a second activation function and based on the convolution feature information, generating the variant branch feature information.
[0017] Optionally, the step of using the hidden Markov network model and based on the frame fusion feature information, obtaining the frame temporal feature information corresponding to the video to be recognized includes: using the hidden Markov network model to perform temporal feature analysis on the frame fusion feature information, generating a sequence of frame feature random vectors corresponding to the video to be recognized; using the hidden Markov network model and based on the sequence of frame feature random vectors, obtaining the frame temporal feature information.
[0018] Optionally, the fusion model includes an average pooling layer and a fully connected layer; the process of using the fusion model to fuse and recognize the frame spatial feature information and the frame temporal feature information to obtain the action category corresponding to the video to be recognized includes: using the average pooling layer to fuse the frame spatial feature information and the frame temporal feature information to generate fused feature information; using the fully connected layer and based on the fused feature information, obtaining the action category information.
[0019] Optionally, the video action recognition model includes: a feature extraction model; the method further includes: using the feature extraction model to perform feature extraction processing on each video frame of the video to be recognized to obtain the frame feature information; wherein, the feature extraction model includes: a convolutional neural network model.
[0020] According to a third aspect of the present disclosure, there is provided a training device for a video action recognition model, wherein the video action recognition model includes: a residual attention network model, a hidden Markov network model, and a fusion model; the device includes: a first spatial feature obtaining module, configured to use the residual attention network model and based on the sample frame feature information corresponding to a video sample, obtain the sample frame spatial feature information corresponding to the video sample; a first temporal feature obtaining module, configured to generate sample frame fused feature information according to the sample frame feature information and the sample frame spatial feature information, use the hidden Markov network model and based on the sample frame fused feature information, obtain the sample frame temporal feature information corresponding to the video sample; a first feature fusion module, configured to use the fusion model to fuse and recognize the sample frame spatial feature information and the sample frame temporal feature information to obtain the action category corresponding to the video sample; a model adjustment module, configured to adjust the residual attention network model and the hidden Markov network model according to the overall loss function corresponding to the video action recognition model.
[0021] According to a fourth aspect of the present disclosure, there is provided a video action recognition device, including: a model acquisition module configured to acquire a trained video action recognition model; wherein, the video action recognition model is trained by the training method as described above, and the video action recognition model includes: a residual attention network model, a hidden Markov network model, and a fusion model; a second spatial feature acquisition module configured to use the residual attention network model and based on the frame feature information corresponding to the video to be recognized, acquire the frame spatial feature information corresponding to the video to be recognized; a second temporal feature acquisition module configured to generate frame fusion feature information according to the frame feature information and the frame spatial feature information, and use the hidden Markov network model and based on the frame fusion feature information, acquire the frame temporal feature information corresponding to the video to be recognized; a second feature fusion module configured to use the fusion model to perform fusion and recognition processing on the frame spatial feature information and the frame temporal feature information, and acquire the action category corresponding to the video to be recognized.
[0022] According to a fifth aspect of the present disclosure, there is provided a video action recognition model training device, including: a memory; and a processor coupled to the memory, the processor being configured to execute the method as described above based on instructions stored in the memory.
[0023] According to a sixth aspect of the present disclosure, there is provided a video action recognition device, including: a memory; and a processor coupled to the memory, the processor being configured to execute the method as described above based on instructions stored in the memory.
[0024] According to a seventh aspect of the present disclosure, there is provided a computer-readable storage medium storing computer instructions, and the instructions are executed by a processor to perform the method as described above.
[0025] The training method, video action recognition method, device, and storage medium of the video action recognition model of the present disclosure train the video action recognition model, provide a two-stream coupling network model including a residual attention network model and a hidden Markov network model, and can achieve spatio-temporal feature fusion of video actions through the video action recognition model, which can improve the extraction and effective representation of spatio-temporal information in video action recognition, can improve the performance, efficiency, and accuracy of video recognition, and improve the user experience. Description of the Drawings
[0026] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0027] Figure 1 It is a schematic flowchart of an embodiment of a method for training a video action recognition model according to the present disclosure;
[0028] Figure 2 It is a schematic structural diagram of an embodiment of a video action recognition model;
[0029] Figure 3 It is a schematic flowchart of an embodiment of a video action recognition method according to the present disclosure;
[0030] Figure 4 It is a schematic module diagram of an embodiment of a training device for a video action recognition model according to the present disclosure;
[0031] Figure 5 It is a schematic module diagram of another embodiment of a training device for a video action recognition model according to the present disclosure;
[0032] Figure 6 It is a schematic module diagram of a model adjustment module in an embodiment of a training device for a video action recognition model according to the present disclosure;
[0033] Figure 7 It is a schematic module diagram of a first spatial feature acquisition module in an embodiment of a training device for a video action recognition model according to the present disclosure;
[0034] Figure 8 It is a schematic module diagram of an embodiment of a video action recognition device according to the present disclosure;
[0035] Figure 9 It is a schematic module diagram of another embodiment of a video action recognition device according to the present disclosure;
[0036] Figure 10 It is a schematic module diagram of a second spatial feature acquisition module in an embodiment of a video action recognition device according to the present disclosure;
[0037] Figure 11 It is a schematic module diagram of yet another embodiment of a training device for a video action recognition model according to the present disclosure;
[0038] Figure 12 It is a schematic module diagram of yet another embodiment of a video action recognition device according to the present disclosure. Specific Embodiments
[0039] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which exemplary embodiments of the present disclosure are illustrated. The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. It is obvious that the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the protection scope of the present disclosure. The technical solutions of the present disclosure will be described in various aspects below with reference to each figure and embodiment.
[0040] The "first", "second", etc. in the following text are only used for descriptive distinction and have no other special meanings.
[0041] Figure 1 FIG. is a schematic flowchart of an embodiment of a method for training a video action recognition model according to the present disclosure. The video action recognition model includes a residual attention network model, a hidden Markov network model, a fusion model, etc.; as Figure 1 shown:
[0042] Step 101: Use the residual attention network model and based on the sample frame feature information corresponding to the video sample, obtain the sample frame spatial feature information corresponding to the video sample.
[0043] In one embodiment, the video sample may be a video, including multiple video frames. Determine the action information in the video sample and perform annotation. The action information includes speaking, running, playing ball, walking, driving, arranging goods, etc. Based on the annotated video sample and the annotated action information, generate training samples to train the video action recognition model.
[0044] The sample frame feature information is the feature information corresponding to the video frames of the video sample. Multiple methods can be used to obtain the sample frame feature information. For example, the video action recognition model includes a feature extraction model, and the feature extraction model includes a convolutional neural network model, etc. Use the feature extraction model to perform feature extraction processing on each video frame of the video sample to obtain the sample frame feature information.
[0045] Step 102: Generate sample frame fusion feature information according to the sample frame feature information and the sample frame spatial feature information, and use the hidden Markov network model and based on the sample frame fusion feature information, obtain the sample frame temporal feature information corresponding to the video sample.
[0046] Existing multiple methods can be used to fuse the sample frame feature information and the sample frame spatial feature information to generate the sample frame fusion feature information.
[0047] Step 103: Use the fusion model to fuse and identify the spatial feature information and temporal feature information of the sample frames to obtain the action category corresponding to the video sample.
[0048] Step 104: Adjust the residual attention network model and the hidden Markov network model according to the overall loss function corresponding to the video action recognition model.
[0049] The video action recognition model of the present disclosure is a two-stream coupling network model, where one stream is a residual attention network model and the other stream is a hidden Markov network model. The video action recognition model can achieve the spatio-temporal feature fusion of video actions, and improve the extraction and effective representation of spatio-temporal information in video action recognition through two-stream coupling, thereby improving the performance of video recognition.
[0050] The residual attention network model in the video action recognition model can robustly remove background noise from features and obtain strong and reliable feature information. The residual attention network model is formed by combining residual blocks with an attention mechanism, which helps the two-stream coupling model improve the recognition accuracy. The residual attention network model obtains the benefit of reducing video background noise without significantly losing spatial information, making the spatial features richer.
[0051] The hidden Markov network model in the video action recognition model can model temporal features. The output of the residual attention network model is fused with the input of the hidden Markov network model, and the fusion strategy gradually combines the input frames according to the specified rules. In order to learn the complementary signals between the two streams, the features of the two streams can be fused at different levels.
[0052] In one embodiment, the residual attention network model includes an invariant branch sub-model and a variant branch sub-model. Use the invariant branch sub-model and based on the sample frame feature information, obtain the invariant branch feature information corresponding to the video sample. Use the variant branch sub-model and based on the sample frame feature information, obtain the variant branch feature information corresponding to the video sample. Use the first activation function and based on the invariant branch feature information and the variant branch feature information, generate the sample frame spatial feature information.
[0053] The invariant branch sub-model includes a DW convolutional layer and a PW convolutional layer. DW convolution: One convolutional kernel of Depthwise Convolution is responsible for one channel, and one channel is only convolved by one convolutional kernel. PW convolution: The operation of Pointwise Convolution is very similar to the conventional convolution operation, and the size of its convolutional kernel is 1×1×M, where M is the number of channels in the previous layer. Use the DW convolutional layer and based on the sample frame feature information, obtain the first feature information, and input the first feature information into the PW convolutional layer to output the invariant branch feature information.
[0054] The variant branch sub-model includes a downsampling layer, an upsampling layer, and a variant branch convolutional layer. The downsampling layer is used to obtain downsampled feature information based on the sample frame feature information. The upsampling layer is used to obtain upsampled feature information based on the downsampled feature information and the residual information corresponding to the sample frame feature information. The variant branch convolutional layer is used to obtain convolutional feature information based on the upsampled feature information. The second activation function is used to generate variant branch feature information based on the convolutional feature information.
[0055] In one embodiment, as Figure 2 shown, the video action recognition model includes three parts. The first part is a spatial residual attention network (residual attention network model) for processing spatial features; the second part is a temporal hidden Markov network (hidden Markov network model) for modeling temporal features; the third part is a fusion model for learning the fusion and classification processes between the two streams.
[0056] During the training of the model, first, the video samples are used to extract the information features (sample frame feature information) of the video samples by using an existing convolutional neural network backbone model, such as an existing multi-layer convolutional neural network model; the spatial residual attention network is used to fuse the sample frame feature information with each other to obtain more reliable spatial information features.
[0057] As Figure 2 shown, the residual attention network model includes an invariant branch sub-model and a variant branch sub-model. The invariant branch sub-model includes a DW convolutional layer and a PW convolutional layer, etc., and the variant branch sub-model includes a downsampling layer, an upsampling layer, a variant branch convolutional layer, etc.
[0058] Using the residual-based spatial residual attention network, the enhanced features of each sample frame feature information can be obtained. The calculation method of the residual is:
[0059] y = F(x) + x (1-1);
[0060] where x is the input feature of the first layer (sample frame feature information), and F(x) is the output of the non-linear superposition layer, which is the sample frame spatial feature information. Element-wise summation is used to process the outputs of the invariant branch M() and the variant branch V(), and the specific calculation is as follows:
[0061] F i (x) = M i (x) + V i (x) (1-2);
[0062] where i represents the i-th spatial position. As the input feature, x undergoes smooth convolution, Depthwise (DW) convolution + ReLU, and Pointwise (PW) convolution + ReLU operations to obtain Mi (x), reducing the number of convolution operation parameters through the operations of the DW layer and the PW layer, where M i (x) is the invariant branch feature information.
[0063] V i (x) is obtained after the attention mechanism is processed. V i (x) consists of downsampling, upsampling, convolution operations, and activation functions. The input features are downsampled (max pooling) multiple times to quickly increase the receptive field, and then a small number of residual units are extracted. After performing the downsampling operation to reach the lowest resolution, the global information is expanded through a symmetric top-down architecture. The processing of the spatial feature map mainly involves the downsampling and upsampling processes first to obtain the global features of the spatial features. Then, the global high-dimensional features extracted by upsampling are combined with the previous residual features to obtain rich features.
[0064] The spatial residual attention network processes the input feature x to create the selection and localization of important spatial features. This module captures the interaction characteristics by treating people / objects as nodes and their interactions as edges, and calculates the output value v as:
[0065] v = Conv(Us(Ds(x) + residual(x))) (1-3);
[0066] where Ds(x), Us(·), and Conv(·) represent downsampling (the original feature length and width are each reduced by half), upsampling (the original feature length and width are each doubled), and conventional convolution operations (variant branch convolution layer), respectively. residual(x) represents taking x as the residual term. The output value v is normalized using the sigmoid(·) activation function (the second activation function) to obtain the variant branch feature information. Finally, the sample frame spatial feature information is:
[0067] F i (x) = ReLU(M i (x) + V i (x)) (1-4);
[0068] Adding the ReLU activation function (the first activation function) to the variant branch obtains a non-linear feature vector, making it easier to optimize the residual mapping than the original mapping. Residual attention learning not only maintains the good characteristics of the original features but also enables it to bypass the upsampling and downsampling operations and forward the features to the top layer, while suppressing the feature selection ability of upsampling and downsampling.
[0069] In one embodiment, a Hidden Markov Network model is used to perform temporal feature analysis on the fused feature information of sample frames, generating a sequence of sample frame feature random vectors corresponding to the video sample. Using the Hidden Markov Network model and based on the sequence of sample frame feature random vectors, the sample frame temporal feature information is obtained.
[0070] Construct a first loss function corresponding to the Residual Attention Network model. Construct a second loss function corresponding to the Hidden Markov Network model. Based on the first loss function, the second loss function, and the corresponding balance coefficient, generate an overall loss function.
[0071] Multiple methods can be used to construct the second loss function. For example, determine the posterior probability information of the Hidden Markov Network model for processing the fused feature information of sample frames, generate an objective function based on the posterior probability information, and construct the second loss function according to the objective function; wherein, the second loss function is used to characterize the parameter values of the Hidden Markov Network model when the objective function is at its minimum value.
[0072] As Figure 2 shown, the temporal information task can be modeled using a Hidden Markov Network. First, divide the input video segment (fused feature information of sample frames) into multiple frames and extract features, perform temporal feature analysis (mainly temporal information) on each frame (T′∈T), thereby obtaining a sequence of feature vectors O={O t1 ,O t2 ,…O tn}, where O∈T′) is the input video segment, and tn represents the end time; each feature vector is a random vector that satisfies a certain probability distribution, and this sequence of random vectors changes over time.
[0073] Define a Hidden Markov model θ with a set of parameters (N, A, B, π), denoted as θ(N, A, B, π), where N is the number of states, A is the state transition probability matrix, B is the observation probability matrix, and π is the initial probability vector distribution state.
[0074] θ has a posterior probability where P(θ) is the prior probability of model θ, and P(O) is the probability of training sample O. As the number of training samples increases, the magnitude of P(O|θ) will increase without limit, while the magnitude of P(θ) is finite and constant. Therefore, θ has a relatively high prior probability as long as there are a sufficient number of training samples. The optimal model of the Hidden Markov model θ is:
[0075] θ m =arg max(P(θ|O))=arg max P(O|θ)(P(θ)) (1-5);
[0076] where θm is the optimal model; arg max P(*) means: the value of the variable when P(*) takes the maximum value. P(θ|O) given represents the probability of θ given O. The recognition process of temporal features can be described as:
[0077]
[0078] where, r is the recognition result, and L is the total number of pattern categories. P(θ r ) in formula (1-6) is the prior probability of pattern r.
[0079] Use the prior probability P(θ) expression to find the higher prior probability in the model space (N, A, B, π). According to information theory, the entropy of the θ probability distribution is as follows:
[0080]
[0081] According to the maximum entropy principle, the prior probability of θ is:
[0082] P(θ) = exp(-λH(θ)) / α (1-8);
[0083] where, λ and α are normalized to constants, λ is usually set to 1, exp is the exponential function with base e, and the prior probability distribution of θ is obtained.
[0084] For the training sample O with a fixed number of states N, its maximum a posteriori probability model is:
[0085]
[0086] From formula (1-5), formula (1-6) and formula (1-8), it can be further obtained that:
[0087] P(O|θ)P(θ) = log(P(O|θ)) - λH(θ) (1-10);
[0088] It is necessary to find the minimum value of the following objective function
[0089] f(θ) = -log(P(O|θ)) + λH(θ) (1-11);
[0090] The method of using this objective function for model training is: the first step is to seek the minimum value of the objective function E. During training, initialize the parameters randomly; the second step is the E-step, and the data expectation of the parameters is obtained through forward-backward calculation; the third step is the M-step, and the parameters are updated through E(X); the fourth step is to iterate the second step - the third step until the parameters are stable and a local optimal solution is obtained.
[0091] The posterior model θ* needs to satisfy:
[0092]
[0093] Using the multi - feature sequence O = {O t1 , O t2 , …O tn} observed in the video as the input of the model, the input is a multi - dimensional matrix, and the model θ is trained.
[0094] Using the shrinkage training method of maximum a posteriori probability as the optimization criterion, a temporal hidden Markov model with a set of undetermined parameters (N, A, B, π) is generated. Through the hidden Markov network model, the observed sequence state features are real video temporal features, which are more in line with the structure of temporal features and can extract more temporal information. Considering that there are obvious differences in the importance of the input (for example, when estimating O t2 , θ t3 is obviously more relevant than θ t2 ), the details contained are not easily overlooked.
[0095] Construct the first loss function corresponding to the residual attention network model, that is, use the cross - entropy loss function:
[0096]
[0097] where σ(·) and y are the softmax activation function and the true action label respectively, and N represents the number of samples.
[0098] Generate the overall loss function as:
[0099] L total = L SRA +τθ * (1 - 14);
[0100] where L SRA and θ* are the first loss function and the second loss function for the spatial residual attention network and the temporal hidden Markov network respectively, and τ is the balance coefficient.
[0101] In one embodiment, the fusion model includes an average pooling layer and a fully - connected layer. The average pooling layer is used to fuse the sample frame spatial feature information and the sample frame temporal feature information to generate fused feature information. The fully - connected layer is used and based on the fused feature information, the action category information is obtained.
[0102] As Figure 2 shown, the video sample passes through the convolutional neural network to obtain the feature x (sample frame feature information), and x is input into the residual attention network to obtain the feature ReLU(F i(x)) (Sample frame spatial feature information), (ReLU(F i (x)) + x) (Sample frame fusion feature information) is input into the temporal Hidden Markov Network to obtain the feature matrix x' (Sample frame temporal feature information).
[0103] Through the average pooling operation, the outputs of the spatial residual attention network and the temporal Hidden Markov Network are fused respectively. The fusion process is: AvgPool(ReLU(F i (x)) + AvgPool(x'), and the fused features include both rich spatial features and temporal features. The fused features are then refined through the fully connected layer, and finally the video action categories are output.
[0104] The video action recognition model of the present disclosure uses the residual attention network model coupled with the temporal Hidden Markov Network for video action recognition. The residual attention network model can not only correctly identify important regions, but also better eliminate the noise of irrelevant backgrounds, making the spatial features more abundant, and at the same time improving the recognition of representative features.
[0105] The residual attention network model also uses the combination of DW convolution and PW convolution to keep the deep features at a relatively large resolution to ensure that the features are not lost, while the number of parameters and the computational cost are relatively low. The Hidden Markov Network model performs temporal modeling, combines the enhanced features output by the spatial residual attention network, and learns the temporal information iteratively according to time. The output of the Hidden Markov Network and the output of the spatial residual attention network are fused, which can improve the recognition performance of the model.
[0106] Figure 3 It is a schematic flowchart of an embodiment of the video action recognition method according to the present disclosure, as Figure 3 shown:[[]]END]]
[0107] Step 301, obtain a trained video action recognition model; wherein, the video action recognition model is trained by the training method as above.
[0108] Step 302, use the residual attention network model and based on the frame feature information corresponding to the video to be recognized, obtain the frame spatial feature information corresponding to the video to be recognized.
[0109] In one embodiment, a feature extraction model is used to perform feature extraction processing on each video frame of the video to be recognized to obtain frame feature information.
[0110] Step 303, generate frame fusion feature information according to the frame feature information and the frame spatial feature information, and use the Hidden Markov Network model and based on the frame fusion feature information, obtain the frame temporal feature information corresponding to the video to be recognized.
[0111] Step 304: Use the fusion model to fuse and identify the frame spatial feature information and the frame temporal feature information, and obtain the action category corresponding to the video to be recognized.
[0112] In one embodiment, the residual attention network model includes an invariant branch sub-model and a variant branch sub-model. Use the invariant branch sub-model and based on the frame feature information, obtain the invariant branch feature information corresponding to the video to be recognized; use the variant branch sub-model and based on the frame feature information, obtain the variant branch feature information corresponding to the video to be recognized; use the first activation function and based on the invariant branch feature information and the variant branch feature information, generate the frame spatial feature information.
[0113] The invariant branch sub-model includes a DW convolutional layer and a PW convolutional layer. Use the DW convolutional layer and based on the frame feature information, obtain the first feature information; input the first feature information into the PW convolutional layer, and output the invariant branch feature information.
[0114] The variant branch sub-model includes a downsampling layer, an upsampling layer, and a variant branch convolutional layer. Use the downsampling layer and based on the frame feature information, obtain the downsampling feature information; use the upsampling layer and based on the downsampling feature information and the residual information corresponding to the frame feature information, obtain the upsampling feature information; use the variant branch convolutional layer and based on the upsampling feature information, obtain the convolutional feature information; use the second activation function and based on the convolutional feature information, generate the variant branch feature information.
[0115] In one embodiment, use the hidden Markov network model to perform temporal feature analysis on the frame fusion feature information, generate a sequence of frame feature random vectors corresponding to the video to be recognized; use the hidden Markov network model and based on the sequence of frame feature random vectors, obtain the frame temporal feature information.
[0116] The fusion model includes an average pooling layer and a fully connected layer. Use the average pooling layer to fuse the frame spatial feature information and the frame temporal feature information, and generate the fusion feature information; use the fully connected layer and based on the fusion feature information, obtain the action category information.
[0117] In one embodiment, as Figure 4 shown, the present disclosure provides a training device 40 for a video action recognition model, including: a first spatial feature obtaining module 41, a first temporal feature obtaining module 42, a first feature fusion module 43, and a model adjustment module 44.
[0118] The first spatial feature obtaining module 41 uses a residual attention network model and, based on the sample frame feature information corresponding to the video sample, obtains the sample frame spatial feature information corresponding to the video sample. The first temporal feature obtaining module 42 generates sample frame fusion feature information according to the sample frame feature information and the sample frame spatial feature information, and uses a hidden Markov network model and, based on the sample frame fusion feature information, obtains the sample frame temporal feature information corresponding to the video sample.
[0119] The first feature fusion module 43 uses a fusion model to perform fusion and recognition processing on the sample frame spatial feature information and the sample frame temporal feature information, and obtains the action category corresponding to the video sample. The model adjustment module 44 adjusts the residual attention network model and the hidden Markov network model according to the overall loss function corresponding to the video action recognition model.
[0120] As Figure 5 shown, the video action recognition model 40 further includes a feature extraction model 45. The feature extraction model 45 uses the feature extraction model to perform feature extraction processing on each video frame of the video sample, and obtains the sample frame feature information. The feature extraction model includes a convolutional neural network model and the like.
[0121] As Figure 6 shown, the model adjustment module 44 includes a first loss function construction unit 441, a second loss function construction unit 442, and an overall loss function construction unit 443. The first loss function construction unit 441 constructs a first loss function corresponding to the residual attention network model. The second loss function construction unit 442 constructs a second loss function corresponding to the hidden Markov network model. The overall loss function construction unit 443 generates an overall loss function based on the first loss function, the second loss function, and the corresponding balance coefficient.
[0122] The second loss function construction unit 442 determines the posterior probability information of the hidden Markov network model for processing the sample frame fusion feature information, generates an objective function based on the posterior probability information; constructs a second loss function according to the objective function; wherein, the second loss function is used to represent the parameter value when the parameter value of the hidden Markov network model makes the objective function the minimum value.
[0123] In one embodiment, the first temporal feature obtaining module 42 uses a hidden Markov network model to perform temporal feature analysis on the sample frame fusion feature information, and generates a sample frame feature random vector sequence corresponding to the video sample. The first temporal feature obtaining module 42 uses a hidden Markov network model and, based on the sample frame feature random vector sequence, obtains the sample frame temporal feature information.
[0124] The first feature fusion module 43 uses an average pooling layer to fuse the sample frame spatial feature information and the sample frame temporal feature information to generate fused feature information. The first feature fusion module 43 uses a fully connected layer and based on the fused feature information, obtains action category information.
[0125] In one embodiment, as Figure 7 shown, the first spatial feature obtaining module 41 includes: a first invariant feature obtaining unit 411, a first variant feature obtaining unit 412, and a first feature activation unit 413.
[0126] The first invariant feature obtaining unit 411 uses an invariant branch sub-model and based on the sample frame feature information, obtains invariant branch feature information corresponding to the video sample. The first variant feature obtaining unit 412 uses a variant branch sub-model and based on the sample frame feature information, obtains variant branch feature information corresponding to the video sample. The first feature activation unit 413 uses a first activation function and based on the invariant branch feature information and the variant branch feature information, generates sample frame spatial feature information.
[0127] For example, the first invariant feature obtaining unit 411 uses a DW convolutional layer and based on the sample frame feature information, obtains first feature information, inputs the first feature information into a PW convolutional layer, and outputs invariant branch feature information. The first variant feature obtaining unit 412 uses a downsampling layer and based on the sample frame feature information, obtains downsampled feature information; uses an upsampling layer and based on the downsampled feature information and residual information corresponding to the sample frame feature information, obtains upsampled feature information. The first variant feature obtaining unit 412 uses a variant branch convolutional layer and based on the upsampled feature information, obtains convolutional feature information; uses a second activation function and based on the convolutional feature information, generates variant branch feature information.
[0128] In one embodiment, as Figure 8 shown, the present disclosure provides a video action recognition device 80, including: a model acquisition module 81, a second spatial feature obtaining module 82, a second temporal feature obtaining module 83, and a second feature fusion module 84.
[0129] The model acquisition module 81 acquires a trained video action recognition model. The second spatial feature obtaining module 82 uses a residual attention network model and based on the frame feature information corresponding to the video to be recognized, obtains frame spatial feature information corresponding to the video to be recognized. The second temporal feature obtaining module 83 generates frame fusion feature information according to the frame feature information and the frame spatial feature information, and uses a hidden Markov network model and based on the frame fusion feature information, obtains frame temporal feature information corresponding to the video to be recognized. The second feature fusion module 83 uses a fusion model to perform fusion and recognition processing on the frame spatial feature information and the frame temporal feature information, and obtains the action category corresponding to the video to be recognized.
[0130] The second temporal feature obtaining module 83 uses a hidden Markov network model to perform temporal feature analysis on the frame fusion feature information, and generates a sequence of frame feature random vectors corresponding to the video to be recognized. The second temporal feature obtaining module 83 uses the hidden Markov network model and based on the sequence of frame feature random vectors, obtains the frame temporal feature information.
[0131] The second feature fusion module 85 uses an average pooling layer to perform fusion processing on the frame spatial feature information and the frame temporal feature information, and generates fusion feature information; uses a fully connected layer and based on the fusion feature information, obtains the action category information.
[0132] As Figure 9 shown, the video action recognition device further includes a frame feature generation model 85. The frame feature generation model 85 uses a feature extraction model to perform feature extraction processing on each video frame of the video to be recognized, and obtains frame feature information, wherein the feature extraction model includes: a convolutional neural network model.
[0133] In one embodiment, as Figure 10 shown, the second spatial feature obtaining module 82 includes a second invariant feature obtaining unit 821, a second variant feature obtaining unit 822, and a second feature activation unit 823. The second invariant feature obtaining unit 821 uses an invariant branch sub-model and based on the frame feature information, obtains invariant branch feature information corresponding to the video to be recognized. The second variant feature obtaining unit 822 uses a variant branch sub-model and based on the frame feature information, obtains variant branch feature information corresponding to the video to be recognized. The second feature activation unit 823 uses a first activation function and based on the invariant branch feature information and the variant branch feature information, generates frame spatial feature information.
[0134] For example, the second invariant feature obtaining unit 821 uses a DW convolutional layer and based on the frame feature information, obtains first feature information; inputs the first feature information into a PW convolutional layer, and outputs invariant branch feature information.
[0135] The second variant feature obtaining unit 822 uses a downsampling layer and based on the frame feature information, obtains downsampled feature information; uses an upsampling layer and based on the downsampled feature information and the residual information corresponding to the frame feature information, obtains upsampled feature information. The second variant feature obtaining unit 822 uses a variant branch convolutional layer and based on the upsampled feature information, obtains convolutional feature information; uses a second activation function and based on the convolutional feature information, generates variant branch feature information.
[0136] In one embodiment, Figure 11 is a schematic diagram of a module of an embodiment of a training device for a video action recognition model according to the present disclosure. As Figure 11As shown, the device may include a memory 111, a processor 112, a communication interface 113, and a bus 114. The memory 111 is used to store instructions. The processor 112 is coupled to the memory 111 and is configured to execute the training method of the above-mentioned video action recognition model based on the instructions stored in the memory 111.
[0137] The memory 111 may be a high-speed RAM memory, a non-volatile memory, etc. The memory 111 may also be a memory array. The memory 111 may also be partitioned, and the partitions may be combined into virtual volumes according to certain rules. The processor 112 may be a central processing unit CPU, or an application specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the training method of the video action recognition model of the present disclosure.
[0138] In one embodiment, Figure 12 is a schematic diagram of a module of an embodiment of a video action recognition device according to the present disclosure. As Figure 12 shown, the device may include a memory 121, a processor 122, a communication interface 123, and a bus 124. The memory 121 is used to store instructions. The processor 122 is coupled to the memory 121 and is configured to execute the video action recognition method described above based on the instructions stored in the memory 121.
[0139] The memory 121 may be a high-speed RAM memory, a non-volatile memory, etc. The memory 121 may also be a memory array. The memory 121 may also be partitioned, and the partitions may be combined into virtual volumes according to certain rules. The processor 122 may be a central processing unit CPU, or an application specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the video action recognition method of the present disclosure.
[0140] In one embodiment, the present disclosure provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the training method of the video action recognition model in any of the above embodiments, and / or the video action recognition method in any of the above embodiments.
[0141] The training method, video action recognition method, device, and storage medium of the video action recognition model provided by the above embodiments train the video action recognition model to provide a two-stream coupling network model including a residual attention network model and a hidden Markov network model. Through the video action recognition model, spatio-temporal feature fusion of video actions can be achieved, the extraction and effective representation of spatio-temporal information in video action recognition can be improved, the performance, efficiency, and accuracy of video recognition can be enhanced, and the user experience can be improved.
[0142] The methods and systems of the present disclosure may be implemented in many ways. For example, the methods and systems of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the methods is for illustrative purposes only. The steps of the methods of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers a recording medium storing a program for executing the methods according to the present disclosure.
[0143] The description of the present disclosure is given for purposes of illustration and description, and is not intended to be exhaustive or to limit the present disclosure to the disclosed form. Many modifications and variations are obvious to those of ordinary skill in the art. The embodiments are chosen and described in order to best explain the principles and practical applications of the present disclosure, and to enable those of ordinary skill in the art to understand the present disclosure and design various embodiments with various modifications suitable for specific purposes.
Claims
1. A training method for a video action recognition model, wherein, the video action recognition model includes: a residual attention network model, a hidden Markov network model, and a fusion model; the method includes: using the residual attention network model and based on the sample frame feature information corresponding to the video sample, obtaining the sample frame spatial feature information corresponding to the video sample; generating sample frame fusion feature information according to the sample frame feature information and the sample frame spatial feature information, using the hidden Markov network model and based on the sample frame fusion feature information, obtaining the sample frame temporal feature information corresponding to the video sample; using the fusion model to perform fusion and recognition processing on the sample frame spatial feature information and the sample frame temporal feature information, obtaining the action category corresponding to the video sample; adjusting the residual attention network model and the hidden Markov network model according to the overall loss function corresponding to the video action recognition model.
2. The method according to claim 1, further including: constructing a first loss function corresponding to the residual attention network model; constructing a second loss function corresponding to the hidden Markov network model; generating the overall loss function based on the first loss function, the second loss function, and the corresponding balance coefficient.
3. The method according to claim 2, the constructing the second loss function corresponding to the hidden Markov network model includes: determining the posterior probability information of the hidden Markov network model for processing the sample frame fusion feature information; generating an objective function based on the posterior probability information; constructing the second loss function according to the objective function; wherein, the second loss function is used to represent the parameter value when the parameter value of the hidden Markov network model makes the objective function the minimum value.
4. The method according to claim 1, wherein, the residual attention network model includes: an invariant branch sub-model and a variant branch sub-model; the using the residual attention network model and based on the sample frame feature information corresponding to the video sample, obtaining the sample frame spatial feature information corresponding to the video sample includes: using the invariant branch sub-model and based on the sample frame feature information, obtaining the invariant branch feature information corresponding to the video sample; using the variant branch sub-model and based on the sample frame feature information, obtaining the variant branch feature information corresponding to the video sample; using a first activation function and based on the invariant branch feature information and the variant branch feature information, generating the sample frame spatial feature information.
5. The method according to claim 4, wherein, the invariant branch sub-model includes: a DW convolutional layer and a PW convolutional layer; the using the invariant branch sub-model and based on the sample frame feature information, obtaining the invariant branch feature information corresponding to the video sample includes: using the DW convolutional layer and based on the sample frame feature information, obtaining first feature information; inputting the first feature information into the PW convolutional layer, and outputting the invariant branch feature information.
6. The method according to claim 4, wherein, the variant branch sub-model includes: a downsampling layer, an upsampling layer, and a variant branch convolutional layer; obtaining variant branch feature information corresponding to the video sample by using the variant branch sub-model and based on the sample frame feature information includes: using the downsampling layer and based on the sample frame feature information, obtaining downsampled feature information; using the upsampling layer and based on the downsampled feature information and residual information corresponding to the sample frame feature information, obtaining upsampled feature information; using the variant branch convolutional layer and based on the upsampled feature information, obtaining convolutional feature information; using a second activation function and based on the convolutional feature information, generating the variant branch feature information.
7. The method according to claim 1, the step of obtaining sample frame temporal feature information corresponding to the video sample by using the hidden Markov network model and based on the sample frame fusion feature information includes: performing temporal feature analysis on the sample frame fusion feature information by using the hidden Markov network model to generate a sequence of sample frame feature random vectors corresponding to the video sample; using the hidden Markov network model and based on the sequence of sample frame feature random vectors, obtaining the sample frame temporal feature information.
8. The method according to claim 1, wherein, the fusion model includes an average pooling layer and a fully connected layer; the step of performing fusion and recognition processing on the sample frame spatial feature information and the sample frame temporal feature information by using the fusion model to obtain an action category corresponding to the video sample includes: using the average pooling layer to perform fusion processing on the sample frame spatial feature information and the sample frame temporal feature information to generate fusion feature information; using the fully connected layer and based on the fusion feature information, obtaining the action category information.
9. The method according to any one of claims 1 to 8, wherein, the video action recognition model includes: a feature extraction model; the method further includes: performing feature extraction processing on each video frame of the video sample by using the feature extraction model to obtain the sample frame feature information; wherein, the feature extraction model includes: a convolutional neural network model.
10. A video action recognition method, including: obtaining a trained video action recognition model; wherein, the video action recognition model is trained by the training method according to any one of claims 1 to 9, and the video action recognition model includes: a residual attention network model, a hidden Markov network model, and a fusion model; using the residual attention network model and based on frame feature information corresponding to a video to be recognized, obtaining frame spatial feature information corresponding to the video to be recognized; generating frame fusion feature information according to the frame feature information and the frame spatial feature information, and using the hidden Markov network model and based on the frame fusion feature information, obtaining frame temporal feature information corresponding to the video to be recognized; Using the fusion model to perform fusion and recognition processing on the frame spatial feature information and the frame temporal feature information, and obtaining an action category corresponding to the video to be recognized.
11. The method according to claim 10, wherein, the residual attention network model includes: an invariant branch sub-model and a variant branch sub-model; and obtaining the frame spatial feature information corresponding to the video to be recognized by using the residual attention network model and based on the frame feature information corresponding to the video to be recognized includes: using the invariant branch sub-model and based on the frame feature information, obtaining invariant branch feature information corresponding to the video to be recognized; using the variant branch sub-model and based on the frame feature information, obtaining variant branch feature information corresponding to the video to be recognized; using a first activation function and based on the invariant branch feature information and the variant branch feature information, generating the frame spatial feature information.
12. The method according to claim 11, wherein, the invariant branch sub-model includes: a DW convolutional layer and a PW convolutional layer; and obtaining the invariant branch feature information corresponding to the video to be recognized by using the invariant branch sub-model and based on the frame feature information includes: using the DW convolutional layer and based on the frame feature information, obtaining first feature information; inputting the first feature information into the PW convolutional layer and outputting the invariant branch feature information.
13. The method according to claim 11, wherein, the variant branch sub-model includes: a downsampling layer, an upsampling layer, and a variant branch convolutional layer; and obtaining the variant branch feature information corresponding to the video to be recognized by using the variant branch sub-model and based on the frame feature information includes: using the downsampling layer and based on the frame feature information, obtaining downsampled feature information; using the upsampling layer and based on the downsampled feature information and the residual information corresponding to the frame feature information, obtaining upsampled feature information; using the variant branch convolutional layer and based on the upsampled feature information, obtaining convolutional feature information; using a second activation function and based on the convolutional feature information, generating the variant branch feature information.
14. The method according to claim 10, the step of obtaining the frame temporal feature information corresponding to the video to be recognized by using the hidden Markov network model and based on the frame fusion feature information includes: using the hidden Markov network model to perform temporal feature analysis on the frame fusion feature information, and generating a sequence of frame feature random vectors corresponding to the video to be recognized; using the hidden Markov network model and based on the sequence of frame feature random vectors, obtaining the frame temporal feature information.
15. The method according to claim 10, wherein, the fusion model includes an average pooling layer and a fully connected layer; and the step of using the fusion model to perform fusion and recognition processing on the frame spatial feature information and the frame temporal feature information, and obtaining an action category corresponding to the video to be recognized includes: using the average pooling layer to perform fusion processing on the frame spatial feature information and the frame temporal feature information, and generating fusion feature information; Use the fully-connected layer and obtain the action category information based on the fused feature information.
16. The method according to any one of claims 10 to 15, wherein, the video action recognition model includes: a feature extraction model; the method further includes: performing feature extraction processing on each video frame of the video to be recognized by using the feature extraction model to obtain the frame feature information; wherein, the feature extraction model includes: a convolutional neural network model.
17. A training device for a video action recognition model, wherein, the video action recognition model includes: a residual attention network model, a hidden Markov network model, and a fusion model; the device includes: a first spatial feature obtaining module, configured to use the residual attention network model and obtain sample frame spatial feature information corresponding to the video sample based on the sample frame feature information corresponding to the video sample; a first temporal feature obtaining module, configured to generate sample frame fused feature information according to the sample frame feature information and the sample frame spatial feature information, and use the hidden Markov network model and obtain sample frame temporal feature information corresponding to the video sample based on the sample frame fused feature information; a first feature fusion module, configured to perform fusion and recognition processing on the sample frame spatial feature information and the sample frame temporal feature information by using the fusion model to obtain the action category corresponding to the video sample; a model adjustment module, configured to adjust the residual attention network model and the hidden Markov network model according to the overall loss function corresponding to the video action recognition model.
18. A video action recognition device, including: a model acquisition module, configured to acquire a trained video action recognition model; wherein, the video action recognition model is trained by the training method according to any one of claims 1 to 9, and the video action recognition model includes: a residual attention network model, a hidden Markov network model, and a fusion model; a second spatial feature obtaining module, configured to use the residual attention network model and obtain frame spatial feature information corresponding to the video to be recognized based on the frame feature information corresponding to the video to be recognized; a second temporal feature obtaining module, configured to generate frame fused feature information according to the frame feature information and the frame spatial feature information, and use the hidden Markov network model and obtain frame temporal feature information corresponding to the video to be recognized based on the frame fused feature information; a second feature fusion module, configured to perform fusion and recognition processing on the frame spatial feature information and the frame temporal feature information by using the fusion model to obtain the action category corresponding to the video to be recognized.
19. A training device for a video action recognition model, including: a memory; and a processor coupled to the memory, the processor being configured to execute the method according to any one of claims 1 to 9 based on instructions stored in the memory.
20. A video action recognition device, including: a memory; and a processor coupled to the memory, the processor being configured to execute the method according to any one of claims 10 to 16 based on instructions stored in the memory.
21. A computer-readable storage medium that non-transitorily stores computer instructions, the instructions being executed by a processor to perform the method according to any one of claims 1 to 16.
Citation Information
Patent Citations
Temporal semantic fusion association determining sub-system based on multimodal emotion recognition system
CN108805087A
Behavior recognition method, terminal device and computer readable storage medium
CN110321761A