Multi-Resolution Attention Network for Video Action Recognition

MRANET addresses computational and scalability issues in VHAR by integrating 2D CNNs and multi-resolution analysis with attention mechanisms, enabling efficient and scalable video action recognition.

JP7713528B2Active Publication Date: 2025-07-25BEN GROUP INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023553165
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-11-16
Filing Date
2021-11-16
Publication Date
2025-07-25
Estimated Expiration
2041-11-16

AI Technical Summary

Technical Problem

Existing video-based human action recognition (VHAR) systems face challenges in requiring significant computational resources, relying on regression models or hand-crafted solutions that are costly and time-consuming, and struggle to learn long-term temporal relationships between frames.

Method used

A Multi-Resolution Attention Network (MRANET) that combines 2D convolutional neural networks with multi-resolution analysis and an attention mechanism, eliminating the need for bounding boxes or pose modeling, and using recursive attention weights based on kinematic derivatives to capture spatio-temporal context.

Benefits of technology

Enables efficient and scalable video action recognition with reduced computational requirements, effectively learning long-term temporal relationships without human intervention, suitable for industrial applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713528000045
    Figure 0007713528000045
  • Figure 0007713528000046
    Figure 0007713528000046
  • Figure 0007713528000047
    Figure 0007713528000047
Patent Text Reader

Abstract

The present invention classifies actions appearing in a video clip by receiving a video clip for analysis, applying a convolutional neural network mechanism (CNN) to frames in the clip to generate a 4D embedding tensor for each frame in the clip, applying a multi-resolution convolutional neural network mechanism (CNN) to each of the frames in the clip to generate a sequence of reduced-resolution blocks and calculating kinematic attention weights that estimate the amount of motion in the blocks, applying the attention weights to the embedding tensors for each frame in the clip to generate a weighted embedding tensor, i.e., context, that represents all frames in the clip at a resolution, combining the contexts across all resolutions to generate a multi-resolution context, performing 3D pooling to obtain a 1D feature vector, and classifying the primary action of the video clip based on the feature vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various embodiments generally relate to methods and systems for classifying actions within a video using a multi-resolution attention network.

Background Art

[0002] In recent years, end-to-end deep learning for video-based human action recognition (VHAR) from video clips has been attracting attention. Application examples have been confirmed in a wide range of fields such as security, gaming, and entertainment. However, human action recognition from video has serious problems. For example, constructing a video action recognition architecture involves capturing the extended spatio-temporal context between frames and requires a large amount of computational resources, which can limit the speed and usefulness of the industrial application of action recognition. Having a robust spatial object detection model or learning the interaction between objects in a scene in a pose model may require human operators to manually identify the objects in the image, which may create highly domain-specific data and can be time-consuming and costly to process.

Prior Art Documents

Non-Patent Documents

[0003]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The attention model is attractive because it can eliminate the need to use explicit regression models with high computational costs. Also, the attention mechanism can form the basis of an interpretable deep learning model by visualizing the image regions used by the network in both space and time during the HAR task. Current attention architectures for HAR may require significant computational resources for model training (e.g., up to 64 GPUs in some cases) because they rely on regression models or optical flow features. This is generally a problem faced by small enterprises and universities. Other attention models use hand-crafted solutions. That is, some of the parameters are pre-defined by experts (skeleton parts, human postures, or bounding boxes). Hand-crafted parameters are difficult to handle because they require human labor and domain expertise, which can reduce the scalability of solutions for new datasets. This is generally a problem faced in industrial applications. The spatial attention mechanism aims to automatically localize objects in a scene without the need for human intervention or expertise. However, prior art attention mechanisms may have difficulty learning long-term temporal relationships because they do not consider the temporal relationships between different frames.

[0005] Therefore, the present invention has been made in view of these considerations and the like.

Means for Solving the Problems

[0006] The present invention provides a novel end-to-end deep learning architecture for classifying (recognizing) human actions occurring within a video clip (VHAR). The present invention introduces an architecture called the Multi-Resolution Attention Network (MRANET) herein, which combines mechanisms provided by a two-dimensional (2D) convolutional neural network (2D-CNN), such as a stream network, keyframe learning, and multi-resolution analysis, in a unified framework.

[0007] To achieve high computational performance, MRANET constructs a multi-resolution (MR) decomposition of the scene using a two-dimensional (2D) convolutional neural network (2D-CNN). Different from prior art methods, this approach does not require bounding box or pose modeling to recognize objects and actions within a video. Video frames at multiple resolutions, i.e., the details of the image, commonly characterize distinct physical structures with different sizes (frequencies) and orientations in the MR space.

[0008] At the core of MRANET is an attention mechanism that calculates a vector of attention weights computed recursively. That is, the weight of the frame at time t is a function of the previous frame at time t - 1. In a particular embodiment, the recursive attention weights are calculated using first-order finite difference derivatives (velocity) and second-order finite difference derivatives (acceleration) for the sequence of frames in which an action occurs.

[0009] In one embodiment, MRANET receives a video clip for analysis, applies a convolutional neural network (CNN) mechanism to the frames within the clip to generate a 4D embedding tensor for each frame within the clip, applies a multi-resolution convolutional neural network (CNN) mechanism to each of the frames within the clip to generate a sequence of downsampled resolution blocks, calculates kinematic attention weights that estimate the amount of motion within the blocks, applies the attention weights to the embedding tensor for each frame within the clip to generate a weighted embedding tensor representing all the frames within the clip at the resolution, i.e., to generate a context, combines the contexts across all resolutions to generate a multi-resolution context, performs 3D pooling to obtain a 1D feature vector, and classifies the primary action of the video clip based on the feature vector, thereby classifying the actions that appear within the video clip.

[0010] Non-limiting and non-exhaustive embodiments of the present invention are described with reference to the following drawings. In the drawings, like reference numerals refer to like parts throughout the various drawings unless otherwise specified.

[0011] For a better understanding of the present invention, reference should be made to the following detailed description of the invention to be read in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

[0013] The drawings merely illustrate embodiments of the invention for the purpose of example. Those skilled in the art will readily understand from the following description that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles of the invention described herein.

[0014] The present invention will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which illustrate, by way of example, specific exemplary embodiments in which the invention may be practiced. However, the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art. In particular, the invention may be embodied as a method, process, system, business method, or device. Thus, the present invention may take the form of an embodiment that is entirely hardware, an embodiment that is entirely software, or an embodiment combining aspects of software and aspects of hardware. Accordingly, the following detailed description should not be taken in a limiting sense.

[0015] As used herein, the following terms have the meanings given below.

[0016] Refers to a video clip, clip, or segment of a video that includes a plurality of frames. As used herein, a video includes a primary action.

[0017] Subject - Refers to a person who performs an action captured within a video clip.

[0018] Human action or action - Refers to the movement within a video clip by a person. Although the present invention focuses on human actions, the present invention is not so limited and can also be applied to animals and inanimate objects such as automobiles and balls.

[0019] Posture or human posture - Refers to the body of the subject within a video frame. The posture may include the entire body or a partial body such as, for example, only the head.

[0020] VHAR - Refers to video human action recognition, which is a basic task in computer vision aimed at recognizing or classifying human actions based on actions performed within a video.

[0021] A machine learning model refers to an algorithm or a set of algorithms that takes structured and / or unstructured data inputs and generates predictions or results. The predictions are typically values or sets of values. A machine learning model may itself include one or more component models that interact to produce results. As used herein, a machine learning model refers to a neural network, including a convolutional neural network or another type of machine learning mechanism, that receives a video clip as input data and generates an estimate or prediction for a known validation data set. Typically, the model is trained through successive executions of the model. Typically, the model is executed continuously during the training phase and, after being successfully trained, is used operationally to evaluate new data and make predictions. It must be emphasized that this training phase can be executed thousands of times to obtain an acceptable model that can predict success metrics. Also, the model may discover thousands or tens of thousands of features. And many of these features can be completely different from the features provided as input data. Therefore, the model is not known in advance and cannot be calculated by mental effort alone.

[0022] Prediction - As used herein, it refers to a statistical estimate, i.e., an estimated probability, that an action within a video clip belongs to a specific action class or category of actions. Prediction may also refer to the estimated value or probability assigned to each class or category within a classification system that includes many individual classes. For example, Kinetics400, a data set from DeepMind, is a commonly used training data set that provides up to 650,000 video clips, each of which is classified into a set of 400 different human actions or action classes, each of which is called an action classification or set of action classifications.

[0023] Generalized operation The operations of some aspects of the present invention will be described below with reference to FIGS. 1-3.

[0024] Figure 1 is a generalized block diagram of a Multi-Resolution Attention Network (MRANET) system 100 for analyzing and classifying actions within a video clip. The MRANET server 120 computer-operates or executes an MRANET machine learning architecture 125, also referred to as MRANET 125. The MRANET server 120 accesses, for analysis, a data source 130 that provides a video clip, herein referred to as x c The video clip may be used during model training or operationally for analysis and classification. For example, YOUTUBE (registered trademark).COM, a website operated by GOOGLE, may be one of the data sources 130. Other data sources 130 may include television channels, movies, and video archives. Typically, the MRANET server 120 accesses video clips from the data source 130 over a network 140.

[0025] The user interacts with the MRANET server 120 to identify and provide training video clips for training the MRANET architecture 125. Typically, the user interacts with a user application 115 running on a user computer 110. The user application 115 may be a native application, a web application operating within a web browser such as MOZILLA's FIREFOX or GOOGLE's CHROME, or an app running within a mobile device such as a smartphone.

[0026] The user computer 110 may be any other computer that can execute a program to communicate on the network 140 in order to access a laptop computer, a desktop personal computer, a mobile device such as a smartphone, or the MRANET server 120. Generally, the user computer 110 may be a smartphone, a personal computer, a laptop computer, a tablet computer, or any other computer system equipped with a processor, non-transitory memory for storing program instructions and data, a display, and interactive devices such as a keyboard and a mouse.

[0027] MRANET 125 typically stores data and executes the MRANET method described below with reference to FIGS. 2 and 3A - 3B. The MRANET server 120 may be implemented by a single server computer, by a plurality of server computers operating in cooperation, or by a "cloud" service provided by a network service or a cloud service provider. Devices that can operate as the MRANET server 120 include, but are not limited to, personal computers, desktop computers, multiprocessor systems, microprocessor-based or programmable household appliances, network PCs, servers, network devices, and the like.

[0028] The network 140 enables the user computer 110 and the MRANET server 120 to exchange data and messages. The network 140 may include the Internet, in addition to a local area network (LAN), a wide area network (WAN), a direct connection, combinations thereof, and the like.

[0029] Multi - resolution attention network The teacher-aided machine learning model provides scores or probability estimates for each class in the classification set. The score (probability) indicates the likelihood that the video clip contains the actions represented by the class members. The class with the highest score can be selected when a single prediction is required. This class is considered to represent the action most likely to have occurred within the video clip, as performed by the subject. A validation dataset of video clips for which the primary class is known for each clip is used to train the model by continuously operating the model with different clips from the dataset, adjusting the model with each successive model execution to minimize the error.

[0030] MRANET is a deep end-to-end multi-resolution attention network architecture for video-based human action recognition (VHAR). Figure 3 is a diagram showing the overall architecture and processing steps executed by MRANET100. In the first learning step, MRANET100 performs frame-by-frame analysis of the video clip to encapsulate the spatial action representation. In a particular embodiment, a convolutional neural network (CNN) model or mechanism is used as the embedding model, which processes the video frames to extract features. In a particular embodiment, ResNet, a CNN implementation, i.e., a residual network, is used. ResNet has been confirmed to be effective for image recognition and classification. However, various commercially available CNN models, backbone architectures, or other processing systems that can sequentially extract image features used for image classification may also be used. In a particular embodiment, a ResNet model pre-trained on the ImageNet dataset is used as the embedding model (EM). Each of the T frames within the clip is submitted to CNN302 for feature extraction. Typically, CNN302 is a commercially available CNN model such as ResNet18. CNN302 is the video clip x cProcess each of the t frames inside sequentially or in parallel, and generate the embedded tensor e as the output of each frame t is generated.

[0031] As an example, the last convolutional layer generated by the ResNet CNN before average pooling can be used as the output embedded tensor e t and may then be used for further processing. Formally, EM represents the dynamics of the video clip in a feature volume or 4D embedded tensor (E), and E is defined as follows in Equation 1. E = [e1, …, e t , …, e T Equation 1 Here, E has the shape E ∈ R T×g·F·N×M where T is the number of frames in the clip, F is the number of channels or features in the embedded tensor, N×M is the cropped image dimensions (image dimensions), i.e., the spatial size, and g is the scale factor that increases the total number of channels of the ResNet model. Generally, the image dimensions are represented as an image of N×M, i.e., width N and height M. Thus, each of [e1, …, e t , …, e T is a 3D tensor that is a set of feature values with its dimensions specified as the spatial location of the width and height values in the (N×M) frames, and one value for each of the F channels.

[0032] The second step of the action representation is to generate a fine-to-coarse representation of the scene using a multi-resolution model (MRM: multi-resolution model) architecture, which is described in more detail with reference to Figure 4. The details of the images at multiple resolutions characterize distinct physical structures at different sizes or frequencies and orientations within the MR space. For example, at a coarse resolution (W 3 ) in this example, the low frequencies correspond to large object structures and provide the "context" of the image. Alternatively, at a finer level of the model's resolution layer (W0 , W 1 , W 2 ) learns from small object structures (details). The advantage of MRM is that it does not require either a bounding box or a human pose model to detect objects in a scene.

[0033] Figure 2 provides an example of an image and its feature representation in four consecutive low-resolution versions. Representation A shows the initial input image. Representation B shows the feature representation W of the image at the highest resolution, i.e., the highest-resolution feature representation. 0 is shown. Representation C shows the feature representation W of the image at 1 / 2 resolution. 1 is shown. Representation D shows the feature representation W of the initial image at 1 / 4 resolution. 2 is shown. And Representation E shows the feature representation W at 1 / 8 of the initial image representation. 3 is shown. These representations are essentially intermediate layers of the CNN model, and it can be understood that the extracted feature quantities shown in B - E usually do not correspond to real-world feature quantities.

[0034] In this specification, a spatio-temporal attention mechanism called multi-resolution attention (MRA) calculates a vector of kinematic attention weights using a kinematic model. The kinematic attention weights add a temporal regression calculation to the attention mechanism, enabling long sequential modeling. This means that the weights calculated for the image recorded at time t are calculated based on the weights and / or images recorded at time t - 1. MRA encapsulates each human action in a multi-resolution context. Finally, the contexts are stacked in the action recognition step and passed through a classifier for the final prediction. Note that the entire model is differentiable, and thus it is possible to train the model end-to-end using standard backpropagation. One area of novelty is the use of regression in the multi-resolution space of the attention weights.

[0035] Parameterization of Actions The parameterization of the action models, i.e., identifies, the actions performed by the subjects in the video clip. Returning to Figure 3, the model assumes that the raw input video clip is preprocessed to generate a sequence of T video frames referred to as [Number] Each clip is provided to the CNN 302 and the multi-resolution module (MRM).

[0036] Formally, the video clip is described as a 4D tensor x c as follows. [Number] where x c ∈ R T×3×W×H is the video clip encapsulating the dynamics within the scene, T is the number of frames, i.e., the number of 2D images within the clip, W refers to the frame width in pixels, i.e., another dimension, H refers to the frame height, and the value 3 refers to the 3-value color space such as RGB where each pixel has red, green, and blue values. Also, x c t ∈ R 3×W×H is the t-th frame in the video clip. Each frame contains the main action c, where c refers to the class of the frame, i.e., how the frame is classified by the classifier or labeled by the training set, and C is the number of classes. The right side of Equation 2 represents the average frame ( [Number] ). The batch size is omitted for notational simplicity. The result of the MRA 300 is the action classification, also known as the logit [Number] is an estimated or predicted action class - score that is referred to as.

[0037] Multiresolution model for spatial analysis Referring again to FIG. 3, the multiresolution model (MRM) 304 implements a ResNet model to obtain the fine - coarse MR representations {W c} of each frame of x j}, {j = 0, 1, 2, …, S - 1}, where S represents the number of reduced - resolution representations in the MR space, i.e., the dimension. Essentially, Equation 3 recursively calculates the MR decomposition on a frame - by - frame basis for each clip. Thus, W j is the clip representation in the MR space.

Number

Number

[0038] FIG. 4 shows the multiresolution representations called blocks generated by the MRM 304. Four separate models are shown, each typically implemented as a CNN model. Starting from the video frames of clip x c the first model 402 creates the full - resolution representation block W

Number

[0039] The following Table 1 shows a plurality of evaluated MRM architectures. The MR block [W0, W1, W2, W3] defined by Table 1 may be generated using a pre-activated ResNet18 model. However, there is a difference in that the Conv1 layer uses k = (3×3) instead of the standard kernel (7×7) used by the ResNet model.

[0040] In addition to using a ResNet CNN to calculate the reduced resolution block, other techniques such as averaging, interpolation, and subsampling may be used.

[0041] The output frame size (N×M) is reduced by 1 / 2 at each consecutive resolution W j . Thus, in the example of Table 1, when the frame size of the input data x c is V 0 = 112×112, the frame size of W 0 is 56×56, W 1 is 28×28, and so on.

[0042] The architecture of the model is inspired by the pre-activated ResNet18. However, there is one difference, where the initial Conv layer (pre-processed input) uses a kernel k=(3×3) instead of k=(7×7). The remaining part of the architecture structure is the same as the ResNet18 model, except for the number of channels and blocks. The number of channels and blocks may vary with respect to the original ResNet18 implementation and the target performance (faster computation due to fewer multiplication and addition operations) or accuracy. For example, a shallower model with fewer channels and thus a reduced amount of multiplication and addition operations may be constructed using the ResNet18 architecture.

[0043] Above, we have discussed the CNN network architecture centered around creating the MR block [W0,W1,W2,W3]. However, the same CNN network architecture used to create W0 may be used to generate the embedded output [e1,…e T . That is, similar or identical pre-activation and convolution steps may be used.

[0044] Temporal Modeling After MR processing, the 4D tensor W is passed through the attention model. As the first step of learning, the attention model calculates a vector of attention weights. These attention weights may also be called kinematic attention weights as they reflect the motion between frames in a clip. First, this mechanism performs a high-dimensional reduction of R 3D => R using dot product similarity followed by a 2D pooling operation. Next, the mechanism performs normalization (e.g., using the softmax function) to force the weights into the range [0,1]. Finally, the attention model performs a linear or weighted combination between the normalized weights and the model embedding E to calculate the context and make the final prediction.

[0045] Kinematic Attention Weights To calculate the attention weights applicable to the frame of the embedded model output E, various alternative methods may be used. Four alternative formulas for calculating the attention weights are presented below, namely, (1) forward velocity, (2) backward velocity, (3) backward acceleration, and (4) absolute position.

[0046] Given an action clip, regression calculations can be used to model the temporal dependence of the human posture by making the posture at time t+1 sensitive to the posture in the previous time frame t. To achieve this, an estimated value of velocity or acceleration can be used to calculate the kinematic attention weights using a finite difference derivative. An additional model calculates the positional attention weights that do not require velocity or acceleration. The kinematic attention weights enable the model to learn to focus on the posture at time t while tracking the posture in the previous frame.

[0047] Mathematically, the kinematic attention weights at time t may be estimated as follows from its first-order finite derivative, which can also be called the forward and backward velocities, and its second-order finite derivative, which can also be called the backward acceleration.

Number

Number

[0048] On the other hand, Equation 7 below tracks the posture based on the absolute position as follows.

Equation

[0049] One potential side effect of the first-order approximation is aliasing (high-frequency) addition, which is amplified by the stride-convolution operation and can lead to a degradation in accuracy. A well-known solution for anti-aliasing in any input signal is low-pass filtering before its downsampling. This operation can be performed either on the gradient operator or on the stride-convolution operation. In one embodiment, the low-pass filtering is performed on the gradient operator using a first-order approximation of the central-difference derivative. For a uniform grid and using Taylor expansion, the central derivative can be analytically calculated by summing the forward-backward derivatives (Equations 4 and 5) as given in Equation 8 below.

Equation

[0050] Equations 4, 5, and 8 use information at only two time points, but Equation 8 provides quadratic convergence. In practice, Equation 8 yields results with higher accuracy than forward or backward differences. It can also be observed that Equation 7 has non-time-dependent characteristics (i.e., it does not provide information regarding the order of the sequence). Therefore, when using Equation 7, the attention mechanism may have difficulty modeling sequences over long ranges. Thus, a reference frame may be added to impose a relative order between frames. Instead of using a specific frame, the attention weights may be centered using Equation 9 below.

Number

Number

Number

Number

Number

[0051] The decentralized attention weight models presented in Equations 4 to 7 can often yield acceptable results, but the re-arranged versions of these equations presented in Equations 9 to 12 have been shown to yield higher accuracy. As a result of the re-arrangement, the attention weights become smaller for short motion displacements from the average and larger for longer displacements. In other words, the model automatically learns to use a frame-by-frame strategy to pay attention to the most informative parts of the clip and assign weights to each frame that reflect the variability (amount) of the motion corresponding to the frame.

[0052] Therefore, referring again to FIG. 3, for use in the generation of the MR decomposition for j = 0, …, S−1, which is also called the kinematic tensor and is the tensor output from MRM304

Number

[0053] FIG. 5 illustrates the processing performed by MRA310, 312, 314, and 316 to generate the final context ctx or attention weights for each resolution.

[0054] In step 504, the kinematic tensors generated by MRM304 are stacked to create a block. Similarly, in step 502, the embedding output of CNN302 is stacked for later use, as will be described with respect to step 510 below.

[0055] Next, in step 506, 3D pooling is used to reduce the dimensions of the kinematic tensor using Equation 13 below.

Number

Number

Number

[0056] In step 508, the attention weight

Number

Number

Number

Number

Number

Number

Number

[0057] Other dimensionality reduction methods also exist and may be used to calculate the weights shown in Equation 14. For example, to remove the dimensionality of the filter and apply second-order statistics (mean pooling) to the (N×M) spatial locations, the dot product similarity (w^ t j ) > w^ t j may be used. As another solution, there is one that uses a fully connected layer to apply a series of linear transformations to reduce the dimensionality of the tensor (w^ j ) and uses the softmax function to normalize the weights, which is similar to the dot product solution.

[0058] Soft Attention and Residual Attention As given below in Equation 15, the attention vector

Number

Number

Number

[0059] The attention weight vector calculated above in Equation 14

Number

Number

[0060] The residual attention mechanism is configured by adding the embedding feature to Equation 15. Similar to the soft attention in Equation 15, the residual attention in Equation 16 first uses Equation 13 to reduce the dimension of the kinematic tensor using 3D pooling, and then uses Equation 14 to normalize the attention weights. Mathematically, this is given by

Number

Number

Number

Number

Number

[0061] The final attention, called Scaled Residual Attention (SRA), is scaled by 1 / T to make the context invariant to the clip. SRA is given by

Number

[0062] Equations 15 and 16 each compute a single 3D tensor of dimension F×N×M for each resolution j. These are alternative formulations of what is called the context ctx j Returning again to FIG. 3, ctx j is the output of the MRAs 310, 312, 314, 316.

[0063] Multi-Resolution Attention Returning to FIG. 3, at step 320, the contexts (ctx 0 , ctx 1 , …, ctx S ) are stacked with respect to resolution. Thus, since there are S resolutions each of which is a tensor of dimension FNM, the stacked context results in a block of dimension SFNM.

[0064] Next, at step 322, multi-resolution attention is computed that makes use of the fine-to-coarse context ctx j . The final multi-resolution attention (MRA) is computed as follows.

Number

Number

Number

[0065] MRA is similar to multi-head attention but has two main differences. First, instead of concatenating resolutions, the multiple resolutions are stacked and averaged to have smooth features. Second, the multiple resolution representations view the scene as different physical structures. This fine-to-coarse representation enables the attention model to automatically learn to first focus on the image details (small objects) in the highest resolution representation and then, in each progressively coarser (lower resolution) representation, focus on the larger structures remaining across various scales.

[0066] Unlike conventional attention weight modeling, the method 500 of implementing MRA 310, 312, 314, and 316 generates attention weights based on the feature representations of the images within the clip at various resolutions. Thus, features that may be apparent at a particular resolution but not at others are considered when generating the final context.

[0067] Next, in step 324, a 3D pooling operation is performed that averages the temporal and spatial dimensions, i.e., reduces N×M×T. This step can be performed using Equation 13. By reducing the temporal (T) and spatial (N×M) dimensions, a single 1×F feature vector is obtained, where the elements are normalized and weighted values or scores for each of the F features.

[0068] In certain embodiments, an operation of dropout 326 is performed on the 1×F feature vector. For example, dropout 326 may be performed if overfitting of the model is a concern due to relatively little training data for the number of features. Dropout 326 may be applied, for example, each time the model is run during training. Generally, dropout 326 eliminates features when there is not enough data to generate an estimate. One way to perform the dropout is described in Non-Patent Document 1.

[0069] The final step is called classification 328. That is, a single class from the set of classes is selected as the primary action for the input video x based on the feature vector. Since the number of classes in the classification set may not be equal to the number of features, a linear transformation that generates a classification vector with scores for each class in the classification set is performed in this step. Since this step is performed using a linear transformation, it can also be called linearization. Typically, c the class with the highest value or score, which may also be called the maximum, is the estimated or selected class.

Number

[0070] Action Recognition - Model Training Once the multi - resolution attention finishes the calculation, the MRA network learns to recognize human actions from the context of the actions. Since the logits are

Number

[0071] LLR initializes the learning rate (e.g., λ = 10 -2 ), and reduces it to one-tenth after several epochs. In another embodiment, generally called superconvergence, cyclical learning rate (CLR) update is used, which speeds up training and regularizes the model.

[0072] The above specification, examples, and data provide complete details of the manufacture and use of the compositions of the present invention. Since many embodiments of the present invention can be made without departing from the spirit and scope of the present invention, the present invention resides in the claims appended hereto.

[0073] [Table 1]

Claims

1. A computer-implemented method for classifying actions that appear within a video clip, comprising: receiving a video clip for analysis, the video clip including video frames in a time series; applying a convolutional neural network mechanism (CNN) to the frames within the clip to generate a 4D embedding tensor for each frame within the clip, wherein the four dimensions are time, feature amount, image width, and image height represented by a sequence of video frames within the clip; applying a multi-resolution convolutional neural network mechanism (CNN) to each of the frames within the clip to generate a sequence of reduced-resolution kinematic tensors, each kinematic tensor representing a frame at one of the reduced resolutions; calculating a kinematic attention weight for each reduced-resolution kinematic tensor to estimate the amount of motion within the corresponding video clip at the reduced resolution; applying the attention weight to the embedding tensor for each frame within the clip for each resolution to generate a weighted embedding tensor, called a context, representing all of the frames within the clip at the resolution; combining the contexts across all resolutions to generate a multi-resolution context; performing 3D pooling of the multi-resolution attention to obtain a 1D feature vector, wherein each value in the feature vector indicates the relative importance of the corresponding feature amount; classifying a primary action of the video clip based on the feature vector; and a computer-implemented method.

2. The method according to claim 1, wherein the step of classifying the video clip based on the feature vector includes calculating probabilities for each action class in an action classification set, and the action class probabilities specify the likelihood that the corresponding action occurred within the video clip.

3. The method according to claim 2, wherein the step of calculating the probability for each action class includes performing a linear transformation between the 1D feature vector and a 1D action class vector representing the action classification set to obtain the probability for each class in the action classification set.

4. The method according to claim 1, further comprising applying a dropout mechanism for excluding one or more feature quantities to the feature vector.

5. The method according to claim 1, wherein each successive reduced-resolution embedding tensor has a resolution that is 1 / 2 of the previous reduced-resolution embedding tensor.

6. The step of applying a multi-resolution attention mechanism to the reduced-resolution kinematic tensor includes calculating a tensor for each frame of each resolution representing the action at each spatial location in the corresponding video frame; and performing a 3D pooling operation for reducing the dimensions of the width, height, and feature quantity to obtain scalar attention weights for each frame at each resolution. The method according to claim 1.

7. The method according to claim 1, wherein the step of performing 3D pooling of multi-resolution attention includes averaging the kinematic tensor in the dimensions of the width, height, and feature quantity.

8. The step of generating a sequence of reduced-resolution kinematic tensors includes performing a convolutional neural network operation to generate a new convolutional layer; and reducing the resolution of the new convolutional layer using a technique selected from the group consisting of bilinear interpolation, averaging, weighting, subsampling, or application of a 2D pooling function. The method according to claim 1.

9. The step of calculating kinematic attention weights for estimating the amount of motion in the video includes generating a tensor representation of a video frame at time t using a method selected from the group consisting of a first-order finite derivative function, a second-order finite derivative function, and an absolute position based on time t; and centering the tensor representation around an average frame value. The method according to claim 1.

10. The step of combining the context over all resolutions includes stacking the context for each resolution; and calculating a single 3D tensor having feature quantity values for each 2D spatial location. The method according to claim 1, comprising

11. A server computer, comprising a processor; a communication interface in communication with the processor; data storage for storing video clips; a memory in communication with the processor for storing instructions, which when executed by the processor cause the server to receive a video clip for analysis, the video clip including video frames in a time series; apply a convolutional neural network mechanism (CNN) to the frames in the clip to generate a 4D embedding tensor for each frame in the clip, the four dimensions being time, feature amount, image width, and image height represented by a sequence of video frames in the clip; apply a multi-resolution convolutional neural network mechanism (CNN) to each of the frames in the clip to generate a sequence of reduced-resolution kinematic tensors, each kinematic tensor representing a frame at one of the reduced resolutions; calculate kinematic attention weights for estimating the amount of motion in the corresponding video clip at the reduced resolution for each reduced-resolution kinematic tensor; apply the attention weights to the embedding tensor for each frame in the clip for each resolution to generate a weighted embedding tensor, called a context, representing all the frames in the clip at the resolution; combine the contexts across all resolutions to generate a multi-resolution context; perform 3D pooling of the multi-resolution attention to obtain a 1D feature vector, each value in the feature vector indicating the relative importance of the corresponding feature amount; classify the primary action of the video clip based on the feature vector and a memory A server computer comprising

12. The step of classifying the video clip based on the feature vector includes the step of calculating a probability for each action class in the action classification set, and the action class probability specifies the likelihood that the corresponding action occurred within the video clip. The server computer according to claim 11.

13. The step of calculating a probability for each action class includes the step of performing a linear transformation between the 1D feature vector and a 1D action class vector representing the action classification set to obtain a probability for each class in the action classification set. The server computer according to claim 12.

14. The memory causes the server to apply a dropout mechanism that eliminates one or more features to the feature vector further. The server computer according to claim 11.

15. Each successive reduced-resolution embedding tensor has a resolution that is 1 / 2 of the previous reduced-resolution embedding tensor. The server computer according to claim 11.

16. The step of applying a multi-resolution attention mechanism to the reduced-resolution kinematic tensor includes the step of calculating a tensor for each frame of each resolution representing the motion at each spatial location within the corresponding video frame, and performing a 3D pooling operation that reduces the dimensions of the width, height, and number of features to obtain scalar attention weights for each frame at each resolution. The server computer according to claim 11.

17. The step of performing 3D pooling of the multi-resolution attention includes the step of averaging the kinematic tensor in the dimensions of the width, height, and number of features. The server computer according to claim 11.

18. The step of generating a sequence of reduced-resolution kinematic tensors includes the step of performing a convolutional neural network operation to generate a new convolutional layer, and reducing the resolution of the new convolutional layer using a technique selected from the group consisting of bilinear interpolation, averaging, weighting, subsampling, or application of a 2D pooling function. The server computer according to claim 11.

19. The step of calculating a kinematic attention weight that estimates the amount of motion within the video generating a tensor representation of a video frame at time t using a method selected from the group consisting of a first-order finite derivative, a second-order finite derivative, and an absolute position based on time t; centering the tensor representation about an average frame value; The server computer according to claim 11, comprising:

20. The step of combining the contexts across all resolutions comprises: stacking the contexts for each resolution; calculating a single 3D tensor having a value of a feature amount for each 2D spatial location; The server computer according to claim 11, comprising:

Citation Information

Patent Citations

  • Image processing device, and control method and program of the same

    JP2019144827A

  • Action Recognition in Video Using 3D Spatiotemporal Convolutional Neural Networks

    JP2020519995A

  • 4D convolutional neural networks for video recognition

    US10713493B1

  • Weakly-Supervised Action Localization by Sparse Temporal Pooling Network

    US20200272823A1