Human Behavior Recognition Method and System Based on Adaptive Spatiotemporal Convolution Network
By introducing time convolution residual blocks and multi-scale time convolution blocks into the adaptive spatiotemporal convolution network, combining spatial and temporal attention modules, the information problem that the existing technology cannot effectively aggregate non-European spatial graph structure data is solved, and more accurate human behavior recognition is achieved.
Patent Information
- Application Number
- CN202111628110.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-28
AI Technical Summary
When the existing 3D convolution algorithm and graph convolution neural network process graph structure data in non-European space, they cannot effectively aggregate all neighborhood information around the node, resulting in the inability to extract sufficient spatial features, which in turn affects the accuracy of behavior recognition.
A human behavior recognition method based on an adaptive spatiotemporal convolution network is adopted to construct multi-layer spatiotemporal convolution blocks, in which the fifth and eighth layers add residual blocks of time convolution. The other layers include two different spatial convolution blocks and multi-scale time convolution blocks. More feature information is extracted through the spatial and temporal attention modules, and parameter updates are participated in through the adjacency matrix of the topology structure.
Through the use of spatial and temporal attention modules, spatial features and time domain information can be made more fully utilized, the shortcomings of insufficient extraction features of ordinary spatial graph convolution models can be made up for, and the accuracy of behavior recognition can be improved.
Smart Images

Figure CN114463837B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of human behavior recognition in computer vision, and particularly relates to a human behavior recognition method and system based on an adaptive spatio-temporal convolutional network. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] For the behavior recognition task based on RGB videos, the most classic is the Convolution3D algorithm, that is, the 3D convolution algorithm. Based on the CNN (Convolutional Neural Network), this algorithm introduces the time dimension. Not only does it increase the dimension of the input data, but also the convolutional kernel, stride, padding, etc. in the convolution process add an additional time dimension. This algorithm extracts features from both spatial and temporal dimensions to capture the motion information encoded in multiple adjacent frames, and then classifies this motion information.
[0004] For the behavior recognition task based on skeleton datasets, from the initial graph convolutional network to the spatio-temporal graph convolutional network and the latest various new networks, they all rely on the basic GCN (Graph Convolutional Neural Network) module. Graph convolution generalizes convolution to non-Euclidean structures. However, the essence of convolution is still to aggregate the information of surrounding neighbor nodes. It's just that graph convolution faces data in non-Euclidean space forms. So the core of graph convolution is the multiplication between matrices. But with the continuous development of deep learning-related frameworks, many scholars have started to introduce temporal convolution in the graph convolution module to aggregate the motion information between different frames, or optimize the spatial graph convolution module to improve the accuracy of behavior recognition.
[0005] The problems existing in the above algorithms are as follows:
[0006] For the 3D convolution algorithm, it cannot effectively aggregate the information of graph structure data in non-Euclidean space, that is, it cannot obtain all the neighborhood information within the nodes, which will lead to the inability to extract sufficient spatial features during the convolution process, and thus cannot accurately identify the action categories.
[0007] For the graph convolutional neural network, the ordinary spatial graph convolution module can only focus on the local physical connections between joint points. And during the convolution process, the adjacency matrix does not participate in the parameter update in the backpropagation process. Without a convolutional kernel as a shared parameter, it simply aggregates the features of the graph structure data, so it cannot achieve a good recognition effect either. Summary of the Invention
[0008] To solve at least one of the technical problems existing in the above-mentioned background art, the present invention provides a human behavior recognition method based on an adaptive spatio-temporal convolutional network, which consists of ten basic spatio-temporal convolutional blocks, but only residual blocks with temporal convolution are added to the fifth and eighth layers. For each of the remaining spatio-temporal convolutional blocks, after extracting features by two different spatial convolutions, an information aggregation operation is performed, and then it is sent into a multi-scale temporal convolutional block for extracting temporal domain information, and then it passes through an activation function and is sent to the next basic spatio-temporal convolutional block.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] The first aspect of the present invention provides a human behavior recognition method based on an adaptive spatio-temporal convolutional network, including the following steps:
[0011] Obtain skeleton data;
[0012] According to the skeleton data and the adaptive spatio-temporal convolutional network, perform a classification operation, output a classification result, and obtain a human behavior recognition result according to the classification result; wherein, the construction process of the adaptive spatio-temporal convolutional network includes: constructing multiple spatio-temporal convolutional blocks, among which, residual blocks with temporal convolution are added to the fifth and eighth layers, and each of the remaining spatio-temporal convolutional blocks includes two different spatial convolutional blocks and a multi-scale temporal convolutional block, and motion information is extracted by the two different spatial convolutional blocks; according to the motion information and the multi-scale temporal convolutional block, the motion information is further extracted and aggregated to obtain temporal domain information.
[0013] The second aspect of the present invention provides a human behavior recognition system based on an adaptive spatio-temporal convolutional network, including:
[0014] A data acquisition module, configured to: obtain skeleton data;
[0015] A human behavior recognition module, configured to: according to the skeleton data and the adaptive spatio-temporal convolutional network, perform a classification operation, output a classification result, and obtain a human behavior recognition result according to the classification result; wherein, the construction process of the adaptive spatio-temporal convolutional network includes: constructing multiple spatio-temporal convolutional blocks, among which, residual blocks with temporal convolution are added to the fifth and eighth layers, and each of the remaining spatio-temporal convolutional blocks includes two different spatial convolutional blocks and a multi-scale temporal convolutional block, and motion information is extracted by the two different spatial convolutional blocks; according to the motion information and the multi-scale temporal convolutional block, the motion information is further extracted and aggregated to obtain temporal domain information.
[0016] The third aspect of the present invention provides a computer-readable storage medium.
[0017] A computer-readable storage medium stores a computer program thereon. When the program is executed by a processor, it implements the steps in the human behavior recognition method based on an adaptive spatio-temporal convolutional network as described above.
[0018] The fourth aspect of the present invention provides a computer device.
[0019] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the human behavior recognition method based on an adaptive spatio-temporal convolutional network as described above.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0021] In the multi-layer spatio-temporal convolutional block of the present invention, residual blocks with temporal convolution are only added to the fifth and eighth layers. For each of the remaining spatio-temporal convolutional blocks, after extracting features by two different spatial convolutions, an information aggregation operation is performed, and then it is sent into a multi-scale temporal convolutional block for extracting temporal domain information. After that, it passes through an activation function and is sent to the basic spatio-temporal convolutional block of the next layer. The spatial and temporal attention modules give different degrees of attention to the features of each joint, and the channel attention module helps the model enhance discriminative features according to the input samples. Using these two parts of spatial convolutional blocks to extract more feature information and perform feature fusion can make up for the shortcoming of insufficient feature extraction by ordinary spatial graph convolutional models. Moreover, during the process of network parameter update, the adjacency matrix of the topological structure participates in the update, which ensures the diversity of the extracted feature information. And the information extracted by different modules is different, which can achieve the full utilization of the information under spatial features, and thus a better recognition effect can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0023] Figure 1 is a flowchart of an adaptive spatio-temporal convolutional network for behavior recognition;
[0024] Figure 2 is an architecture diagram of an adaptive spatio-temporal convolutional network for behavior recognition;
[0025] Figure 3 is an architecture diagram of a spatial convolutional block of an adaptive spatio-temporal convolutional network for behavior recognition;
[0026] Figure 4 is an architecture diagram of a multi-scale temporal convolution of an adaptive spatio-temporal convolutional network for behavior recognition. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0028] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0029] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0030] The action recognition task refers to the recognition task of identifying the specific actions of people in a video through specific algorithms. Due to its great application value in virtual reality, intelligent monitoring, intelligent security, and athlete assistant training, etc., it has attracted extensive academic attention in recent years. The action recognition task generally has the following basic processes: preprocessing of data images, detection of the human body in motion, extraction of motion features, training and classification of features, and action recognition. The current action recognition tasks can be divided into action recognition tasks based on RGB videos and action recognition tasks based on skeleton datasets according to the dataset format. The method mentioned in this article is based on the skeleton dataset.
[0031] Embodiment 1
[0032] As Figures 1-4 shown, this embodiment provides a human action recognition method based on an adaptive spatio-temporal convolutional network, including the following steps:
[0033] Step 1: Obtain skeleton data;
[0034] In this embodiment, the dataset used is the NTU-RGBD60 / 120 dataset, which consists of many text files. Each file contains the number of frames of skeleton data, the number of people performing actions, the three-dimensional coordinates (xyz coordinates) of each joint point, etc.
[0035] Step 2: Preprocess and compose the skeleton data;
[0036] The preprocessing of the skeleton data includes:
[0037] Wrap the text data into a 5D matrix format of (N, C, T, V, W) so that it can be input into the adaptive spatio-temporal convolutional network, where N represents the amount of data fed into the network for each training, C represents the number of channels of node information, T represents the number of frames of each video, V represents the number of nodes in the skeleton graph, and W represents the number of people in motion in each frame.
[0038] The preprocessing part of the skeleton data is to extract specific information such as the coordinates of skeleton points, frame length, and number of joint points required for network training from the video, and finally encapsulate it into a format that can be input into the network using the Dataset and Dataloader modules provided by Pytorch, which is the five-dimensional vector (N, C, T, V, W), where the letters represent the batch size of each training, the number of channels, the number of frames, the number of nodes, and the number of people in motion in one frame respectively.
[0039] The composition part is mainly to construct the adjacency matrix A of the joint points according to the connection of the human body skeleton joint points. The size of this matrix is (3, V, V), where V represents the number of nodes, and the three dimensions represent the self-connection matrix of the joint points, the in-degree matrix of the joint points, and the out-degree matrix of the joint points respectively.
[0040] The composition part constructs the corresponding matrices based on the self-connection, out-degree, and in-degree of the joint points and stacks them into a three-dimensional vector format. This matrix is the adjacency matrix A of the nodes. The following formula is used to update the information of the neighbor nodes around node v using matrix A, a ij represents the connection strength between nodes i and j, X is the feature of the node, and W is the weight matrix for feature transformation.
[0041]
[0042] Step 3: According to the skeleton data and the adaptive spatio-temporal convolutional network, perform the classification operation, output the classification matrix, select the subscript of the largest number in each row of the classification matrix as the label of the action type, compare this label with the true label. If they are the same, then the hit count is incremented by one, and the higher the hit count, the better the recognition effect, and the human recognition result is obtained.
[0043] Among them, the construction process of the adaptive spatio-temporal convolutional network includes:
[0044] Among them, the data format of the classification matrix is (N, class), N is the amount of data for each training, and class is the number of action types. For example, N = 8 and class = 60.
[0045] Construct a multi-layer spatio-temporal convolutional block. Among them, residual blocks with temporal convolution are added to the fifth and eighth layers. Each of the remaining spatio-temporal convolutional blocks includes two different spatial convolutional blocks and a multi-scale temporal convolutional block. Motion information is extracted through the two different spatial convolutional blocks; based on the motion information and the multi-scale temporal convolutional block, the motion information is further extracted and aggregated to obtain temporal domain information.
[0046] In step 3, extracting motion information through two different spatial convolutional blocks includes:
[0047] Among them, the first convolutional block includes 3 different topological refinement graph convolutions, and their inputs are all (N*W, C, T, V). Each convolutional block learns the channel topology in a refined manner, while learning the shared topology and the correlation of specific channels. Finally, an accumulation operation is performed on the obtained results to obtain the output (N*W, C′, T′, V′).
[0048] The second spatial convolutional block includes a spatial attention module, a temporal attention module, a channel attention module, and a residual connection. The spatial attention module is used to give different degrees of attention to each joint point. The temporal attention module is used to give different degrees of attention to the same joint point in different frames. The channel attention module is used to enhance discriminative features according to the input samples and supplement the temporal domain information in the convolution process. Among them, these three attention modules respectively propose the input features to the output features, and then an accumulation operation is performed on the output features. At the same time, there is also a feature extracted by the residual connection. Finally, these two features are aggregated to obtain the output (N*W, C′, T′, V′).
[0049] Finally, the information extracted by the two spatial convolutional blocks is aggregated and sent to the next layer.
[0050] The formula is as follows:
[0051] unit_gcn i = Relu(f c ) + Softmax(f a ) (2)
[0052] In the first spatial convolutional block, the 3 different topological refinement graph convolutions include three channel-based refinement topological convolutional blocks. The topological convolutional block includes feature transformation, channel topology modeling, and a feature aggregation operation completed by an aggregation function. The adjacency matrix A is used as the shared topology of all channels, and the matrix A is updated through backpropagation. The second spatial convolutional block includes spatial, temporal, and channel attention blocks. The spatial and temporal attention modules give different degrees of attention to the features of each joint. The channel attention module helps the model enhance discriminative features according to the input samples. These two parts of the spatial convolutional blocks are used to extract more feature information and perform feature fusion.
[0053] In step 3, according to the motion information and the multi-scale temporal convolution block, the multi-scale temporal convolution block includes multiple convolution blocks, and after separately extracting and aggregating the motion information, temporal information is obtained;
[0054] To model actions with different durations, a multi-scale temporal convolution block is added to the model to process the temporal information from the spatial convolution block. It contains 4 temporal convolution blocks, and its input is (N*W, C′, T′, V′) from the spatial convolution block. We use fewer branches to improve the processing speed. The first two branches contain residual blocks of temporal convolution to reduce the training error. After training through layers of the network, global average pooling is performed to prevent overfitting, and finally, a fully connected layer is passed through for classification operation to obtain the output (N, class).
[0055] Among them, the multi-scale temporal convolution block is 4 encapsulated convolution blocks. The first two convolution blocks both include ordinary convolution, normalization, activation function, and a residual block of temporal convolution. The last two convolution blocks include operations such as ordinary convolution, normalization, activation function, and pooling.
[0056] These four convolution blocks respectively extract and aggregate the temporal information of the information from the previous layer and send it to the next basic spatio-temporal convolution block. The representation formula is as follows:
[0057]
[0058] Residual block of aggregated temporal convolution: Residual blocks of temporal convolution are added to the fifth layer and the eighth layer. The residual block consists of an ordinary Conv2d convolution and a normalization layer. The input directly comes from the data output by the previous spatio-temporal convolution block. The data output by the residual block of aggregated temporal convolution and the data output by the multi-scale temporal convolution block are subjected to an aggregation operation to obtain R i , which, via the activation function, is then sent to the next spatio-temporal convolution block. The formula is as follows:
[0059]
[0060] Among them, the process of performing the classification operation includes:
[0061] For the result data after the operations of all layers of spatio-temporal convolution blocks, the format of the data is (N*M, C, T, V,), where N, M, C, T, V respectively represent the meanings of the data's batch_size, N represents the number of people in motion in the video, M represents the number of channels, the number of frames, and the number of nodes. Global average pooling is performed on this data, and the average value of all pixel values in each channel map is calculated to obtain a new channel map to achieve the effect of reducing the dimensionality of the data. Then, through the dropout layer, some neurons in the network are inactivated, and the output is (the number of output channels, the number of classifications). Finally, classification is performed through the fully connected layer.
[0062] Example 2
[0063] This embodiment provides a human behavior recognition system based on an adaptive spatio-temporal convolutional network, including:
[0064] A data acquisition module, configured to: acquire skeleton data;
[0065] A human behavior recognition module, configured to: perform a classification operation according to the skeleton data and the adaptive spatio-temporal convolutional network, output a classification result, and obtain a human behavior recognition result according to the classification result; wherein, the construction process of the adaptive spatio-temporal convolutional network includes: constructing multiple layers of spatio-temporal convolutional blocks, wherein residual blocks with temporal convolution are added to the fifth and eighth layers, and each of the remaining spatio-temporal convolutional blocks includes two different spatial convolutional blocks and multi-scale temporal convolutional blocks, and motion information is extracted through the two different spatial convolutional blocks; according to the motion information and the multi-scale temporal convolutional blocks, the motion information is further extracted and aggregated to obtain temporal domain information.
[0066] Example 3
[0067] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the above-mentioned human behavior recognition method based on an adaptive spatio-temporal convolutional network.
[0068] Example 4
[0069] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the above-mentioned human behavior recognition method based on an adaptive spatio-temporal convolutional network.
[0070] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.
[0071] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0072] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0073] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0074] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-described method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0075] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A human behavior recognition method based on an adaptive spatio-temporal convolutional network, characterized in that, It includes the following steps: Obtain skeleton data; According to the skeleton data and the adaptive spatio-temporal convolutional network, perform a classification operation, output a classification result, and obtain a human behavior recognition result based on the classification result; wherein, the construction process of the adaptive spatio-temporal convolutional network includes: constructing multiple layers of spatio-temporal convolutional blocks, among which, residual blocks with temporal convolution are added to the fifth layer and the eighth layer. Each layer of spatio-temporal convolutional block includes two different spatial convolutional blocks and multi-scale temporal convolutional blocks. Extract motion information through the two different spatial convolutional blocks; according to the motion information and the multi-scale temporal convolutional blocks, re-extract and aggregate the motion information to obtain temporal domain information; The extracting motion information through the two different spatial convolutional blocks includes: The first convolutional block includes multiple different topology-refined graph convolutions. Each convolutional block learns the channel topology in a refined manner, simultaneously learns the shared topology and the correlation of specific channels, and finally performs an accumulation operation on the obtained results; The second spatial convolutional block includes a spatial attention module, a temporal attention module, and a channel attention module, and performs feature refinement operations through each attention module; Finally, aggregate the motion information extracted by the two spatial convolutional blocks.
2. The human behavior recognition method based on the adaptive spatio-temporal convolutional network according to claim 1, wherein The multiple different topology-refined graph convolutions include three channel-refined topology convolutional blocks. The topology convolutional block includes feature transformation, channel topology modeling, and feature aggregation operations completed by an aggregation function. Use the adjacency matrix as the shared topology of all channels and update the adjacency matrix through backpropagation.
3. The human behavior recognition method based on an adaptive spatio-temporal convolutional network according to claim 1, characterized in that The multi-scale temporal convolutional block includes multiple convolutional blocks. Each convolutional block re-extracts and aggregates the motion information to obtain temporal domain information respectively; wherein, the multi-scale temporal convolutional block is 4 encapsulated convolutional blocks. The first two convolutional blocks both include ordinary convolution, normalization, activation function, and a residual block with temporal convolution. The last two convolutional blocks include ordinary convolution, normalization, activation function, and pooling operation.
4. The human behavior recognition method based on an adaptive spatio-temporal convolutional network according to claim 1, wherein The residual block consists of an ordinary Conv2d convolution and a normalization layer.
5. The human behavior recognition method based on an adaptive spatio-temporal convolutional network according to claim 1, wherein The process of performing the classification operation includes: taking the average value of all pixel values in each channel map to obtain a new channel map, then passing through a dropout layer to inactivate some neurons in the network, obtaining the number of output channels and the number of classifications, and finally performing classification through a fully connected layer.
6. The human behavior recognition method based on an adaptive spatio-temporal convolutional network according to claim 1, characterized in that The skeleton data is preprocessed and composed into a graph before being input into the adaptive spatio-temporal convolutional network.
7. A human behavior recognition system based on an adaptive spatio-temporal convolutional network, characterized in that, It includes: A data acquisition module, configured to: obtain skeleton data; A human behavior recognition module, configured to: according to the skeleton data and the adaptive spatio-temporal convolutional network, perform a classification operation, output a classification result, and obtain a human behavior recognition result based on the classification result; wherein, the construction process of the adaptive spatio-temporal convolutional network includes: constructing multiple layers of spatio-temporal convolutional blocks, among which, residual blocks with temporal convolution are added to the fifth layer and the eighth layer. Each layer of spatio-temporal convolutional block includes two different spatial convolutional blocks and multi-scale temporal convolutional blocks. Extract motion information through the two different spatial convolutional blocks; according to the motion information and the multi-scale temporal convolutional blocks, re-extract and aggregate the motion information to obtain temporal domain information; The extracting motion information through the two different spatial convolutional blocks includes: The first convolutional block includes multiple different topological refinement graph convolutions. Each convolutional block learns the channel topology in a refined manner, while learning the shared topology and the correlation of specific channels, and finally performs an accumulation operation on the obtained results; The second spatial convolutional block includes a spatial attention module, a temporal attention module, and a channel attention module, and performs feature refinement operations through each attention module; Finally, the motion information extracted by the two spatial convolutional blocks is aggregated.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in the human behavior recognition method based on an adaptive spatio-temporal convolutional network according to any one of claims 1-6.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the human behavior recognition method based on an adaptive spatio-temporal convolutional network according to any one of claims 1-6.