A method for recognizing sports video actions based on an action granularity grouping structure
Through the grouping structure based on action granularity and the multi-scale spatio-temporal feature fusion method, the problems of high computing costs and difficult information fusion in sports video action recognition are solved, efficient recognition of athletes' movements is achieved, and the accuracy of action recognition is improved.
Patent Information
- Application Number
- CN202310507915.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-08
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-05-08
AI Technical Summary
Existing sports video action recognition technology is cost-effective in multi-level space-time modeling and is difficult to integrate multi-scale information, making it difficult to effectively capture the details and temporal relationships of athletes' movements.
A grouping structure based on action granularity is designed, multi-scale action information is extracted through four different spatial and temporal feature extraction modules, and feature fusion is used to use a hierarchical structure, including global time module, spatial motion module and local time module, and classified in convolutional neural network and fully connected layer.
It improves the accuracy of sports video action recognition, especially in event, set and element categories, and is suitable for action recognition in multi-level semantic categories.
Smart Images

Figure CN116524596B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and video action recognition, and particularly relates to a method for recognizing sports video actions based on an action granularity grouping structure. Background Art
[0002] Sports video action recognition technology refers to the computer's analysis and understanding of the actions of athletes in a video by inputting multiple frames of sports video images, and classifying the actions into specific action types. In practical applications, accurate sports action recognition helps correct athletes' action mistakes, helps coaches make correct decisions, and is applied to sports live broadcast scenarios.
[0003] Due to the successful application of deep learning methods in video action recognition tasks, the accuracy of video action recognition has been significantly improved in the past few years. However, in sports videos, the actions of athletes change very quickly, and the action classification granularity is fine. This requires the model to be able to judge the start and end times of each detailed action in terms of time; and in terms of semantics, to be able to distinguish subclasses of actions at a finer level on the premise of classifying coarse-grained actions. In such a complex multi-level spatio-temporal context, only considering static information modeling cannot achieve satisfactory results. The model needs to effectively model spatio-temporal dynamic information and accurately master time relationships. Therefore, designing an effective structure to capture spatio-temporal information has become a challenge for sports video action recognition. The present invention mainly aims at the problem of spatio-temporal information modeling in sports video action recognition tasks, and proposes a method for recognizing sports video actions based on an action granularity grouping structure.
[0004] One of the methods for sports video action recognition modeling is to use a modeling method based on 2DCNN. Initially, a two-stream network is constructed to extract temporal and spatial information separately. Simonyan et al. used handcrafted optical flow to construct a two-stream network (Simonyan K, Zisserman A. Two-stream convolutional networks for action recognition in videos[J]. Advances in neural information processing systems, 2014, 27). Wang et al. proposed a multi-path framework, which captures channel, motion, and temporal features by designing multiple paths respectively (Wang Z, She Q, Smolic A. Action-net: Multipath excitation for action recognition[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2021:13214-13223). However, with the increase in the number of streams, the computational cost of the multi-stream network grows exponentially. There are also researchers who design an end-to-end single-stream network to extract temporal and spatial information simultaneously. Wang et al. cascaded short-term and long-term temporal modeling modules to complete long-term and short-term temporal modeling in one path (Wang L, Tong Z, Ji B, et al. Tdn: Temporal difference networks for efficient action recognition[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021:1895-1904). But this approach cannot guarantee the independent output of short-term information and is not conducive to capturing motion change information. To solve this problem, some research uses channel grouping, and each group performs feature extraction separately. Hao et al. decomposed spatio-temporal features along the channel dimension into several groups, and each group extracted spatio-temporal features from different angles (Hao Y, Zhang H, Ngo C W, et al. Group Contextualization for Video Recognition[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022:928-938).However, this parallel structure cannot perform multi-scale information fusion, and multi-scale spatio-temporal information plays an important role in sports action recognition.
[0005] To better model temporal information, some researchers also use a modeling method based on 3D CNN. Carreira et al. proposed the I3D network, which extends 2D CNN to 3D CNN. The 2D CNN network is a network pre-trained using ImageNet, so I3D does not need to be trained from scratch (CARREIRA J, ZISSERMAN A. Quo Vadis, Action Recognition?A New Model and the Kinetics Dataset[C] / / 2017 IEEE Conference on Computer Vision and Pattern Recognition, 2017). This network inputs the RGB stream and the optical flow into 3D CNN respectively, and finally obtains the result of two-stream fusion. To reduce the complexity of the three-dimensional network, P3D decomposes the three-dimensional kernel into two independent operations: two-dimensional spatial convolution and one-dimensional temporal convolution (QIU Z, YAO T, MEI T. Learning Spatio-Temporal Representation with Pseudo-3D Residual Networks[C] / / 2017 IEEE International Conference on Computer Vision, 2017). Based on this, Xie et al. proposed the S3D network, which mixes two-dimensional and three-dimensional convolutions in a single network (XIE S, SUN C, HUANG J, et al. Rethinking Spatiotemporal Feature Learning: Speed-Accuracy Trade-offs in Video Classification[C] / / European Conference on Computer Vision, 2018). Some other researchers combine 2D CNN and 3D CNN. Action-net inserts a three-dimensional convolutional feature extractor into 2D CNN, which not only ensures the extraction of spatio-temporal information but also avoids the huge computational cost brought by three-dimensional convolution (WANG Z, SHE Q, SMOLIC A. Action-net: Multipath excitation for action recognition[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021:13214-13223).The CSN proposed by Tran et al. decomposes 3D convolution through channel-separable convolution and spatio-temporal interaction (TRAN D, WANG H, TORRESANI L, et al. Video Classification with Channel-Separated Convolutional Networks[C] / / IEEE International Conference on Computer Vision, 2019). Feichtenhofer proposed the X3D network, which gradually expands the 2D image classification architecture along multiple axes (such as time, frame rate, spatial resolution, width, and depth) (FEICHTENHOFER C. X3D: Expanding Architectures for Efficient Video Recognition[C] / / 2020 IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2020).
[0006] In recent years, with the application of Transformer to the visual field, modeling methods using Transformer have emerged. Transformer is very suitable for processing ordered data because its internal structure contains temporal information and can process the temporal information in videos without additional input of time information. Therefore, Transformer has high applicability in the processing of video content. Neimark et al. were the first to apply Transformer to the video action recognition task and proposed the Video Transformer Network (NEIMARK D, BAR O, ZOHAR M, et al. Video Transformer Network[J]. arXiv preprint arXiv:2102.00719, 2021. 17, 23). This network first uses 2DCNN to extract features from each frame, then learns the temporal relationship between frames through the Transformer encoder, and finally uses the MLP classification head to obtain the classification result. Bertasius et al. proposed a video action recognition model without convolution, TimeSformer (BERTASIUS G, WANG H, TORRESANI L. Is Space-Time Attention All You Need for Video Understanding?[J]. arXiv preprint arXiv:2102.05095, 2021). This model directly learns spatio-temporal features from the sequence of frame-level image patches, and applies the temporal attention mechanism and the spatial attention mechanism within each patch respectively. Arnab et al. proposed a pure Transformer model for video action recognition and proposed several methods for decomposing features along the spatial and temporal dimensions to improve efficiency and robustness (WU W, HE D, LIN T, et al. Mvfnet: Multi-view fusion network for efficient video recognition[C] / / Proceedings of the AAAI Conference on Artificial Intelligence. 2021, 35(4): 2943-2951).UniFormer integrates 3D convolution and spatio-temporal self-attention in the Transformer, which can simultaneously solve the problems of spatio-temporal information redundancy and spatio-temporal relationship dependence (LI K, WANG Y, GAO P, et al. Uniformer: Unified transformer for efficient spatio-temporal representation learning[J]. arXiv preprint arXiv:2201.04676, 2022). However, the Transformer requires a large amount of pre-training data and has a high computational cost.
[0007] To solve the problems of high computational cost and difficulty in multi-scale information fusion of existing methods, the present invention proposes a grouping structure based on action granularity and designs a lightweight multi-scale spatio-temporal modeling and information fusion mechanism. Summary of the Invention
[0008] To solve the above technical problems, the object of the present invention is to provide an effective spatio-temporal feature modeling method for sports video action recognition tasks. This method designs a grouping structure based on action granularity, which can extract action information of different granularities through four spatio-temporal feature extraction modules with different focuses, and uses a hierarchical structure to fuse multi-scale spatio-temporal features, which is applicable to both coarse-grained and fine-grained sports action recognition and improves the performance of sports video action recognition.
[0009] The technical solution adopted by the present invention is as follows:
[0010] A method for recognizing sports video actions based on action granularity grouping, comprising the following steps:
[0011] Step 1: Extract frames from the video data of the FineGym gymnastics dataset and store them as a number of images with a fixed width;
[0012] Use a video frame extraction tool to extract frames from the video data in the FineGym gymnastics dataset, unify the frame width to 256 pixels, and save the images; depending on the video length, the extraction results include dozens to hundreds of frames; store the video frames extracted from the same video in a folder and name them in chronological order;
[0013] Step 2: Use a random sampling algorithm to randomly sample the video frames extracted in Step 1 as the network input;
[0014] The video frames extracted from each video in Step 1 are evenly divided into T segments, and 1 frame is randomly sampled from each segment as the network input. The total input is T frames, and N videos are input simultaneously; Random sampling algorithm: First, calculate the average number of frames in each segment, denoted as a frames; When sampling the i-th frame, use a random function to generate a random integer r in the range [1, a]. i , and use the following formula to determine the position of the sampled frame:
[0015] t i =(i - 1)*a + r i (1)
[0016] where, t i represents that the order of the i-th frame sampled in all video frames is the t i -th frame, a represents the average number of frames in each segment, and r i represents the random number generated in the range [1, a];
[0017] Step 3: Preprocess the video frames extracted in Step 2;
[0018] Apply random scaling and corner cropping to the video frames extracted in Step 2 for data augmentation, and adjust the height and width of each frame to 224 pixels;
[0019] Random scaling randomly adjusts the size of the image, which can be achieved through the following steps: First, randomly select a scaling ratio range, such as [0.8, 1.2], indicating that the scaling ratio can be randomly selected between 0.8 and 1.2; Then, scale the original image according to the randomly selected scaling ratio to obtain the scaled image; Finally, crop the scaled image according to the size of the original image to obtain an image of the specified size.
[0020] Corner cropping is to crop a small square area from the corner of the image. It can be achieved through the following steps: First, randomly select a cropping size range, such as [0.8, 1.0], indicating that the cropping size can be randomly selected between 0.8 times and 1.0 times of the original image; Then, randomly select a scaled image and crop it according to the cropping size; Specifically, randomly select a point from the four corner points of the upper left, lower left, upper right, and lower right of the scaled image as the starting point of the cropping area; Then, determine the size of the cropping area according to the selected cropping size to obtain the cropped image; Finally, scale the cropped image to the specified size to obtain the final image.
[0021] Step 4: Input the video frames processed in Step 3 into a convolutional neural network and use convolutional blocks for feature extraction;
[0022] Input the preprocessed video frame sequence in step 3 into a multi-layer convolutional neural network, which consists of a convolutional layer, a batch normalization layer, a ReLU layer, and a max pooling layer, aiming to extract the shallow features of video information. Among them, the convolutional layer contains 64 convolutional kernels with a size of 7×7, a stride of 2, and a padding of 3; the batch normalization layer performs batch normalization on the output of the convolutional layer, making the mean and variance of each feature map close to 0 and 1; the ReLU activation function layer performs the ReLU activation function operation on the normalized feature map; the max pooling layer has a size of 3×3 and a stride of 2, performs max pooling operation on the feature map, and the number of output feature map channels is 64.
[0023] Step 5: Input the feature map obtained in step 4 into a continuous 4-stage action granularity grouping module to obtain high-level spatio-temporal features that fuse multi-scale spatio-temporal information. The specific content of this module is as follows: First, use a convolutional layer to adjust the number of channels, then divide the channels of the feature map into four groups on average, and each group uses four spatio-temporal feature extraction modules with different focuses. Use residual connections to construct a hierarchical grouping structure, then fuse the features of the four groups containing different granularity action information, and finally use a convolutional layer again to adjust the number of channels to the same as the input, and add the fused features to the input features to obtain high-level spatio-temporal features;
[0024] Specifically, input the feature map obtained in step 4 into a continuous 4-stage action granularity grouping module, and each stage contains 3, 4, 6, and 3 action granularity grouping modules respectively. This module first includes a 1×1 convolutional layer for adjusting the number of channels. Then divide the channels of the input feature map into 4 groups on average, insert 1 spatio-temporal feature extraction module corresponding to different action granularities into each group, and aggregate them into a hierarchical grouping structure using residual connections to capture multi-scale spatio-temporal features and effectively fuse them. The action granularity grouping modules form a hierarchical structure with gradually finer action granularity from top to bottom. The output of each group from top to bottom is used to identify multi-level sports action types from coarse-grained actions to fine-grained actions.
[0025] First, set the input feature after adjusting the number of channels as X, and its shape is set as [N×T×C×H×W], where N represents the batch size, T represents the number of video frames, C represents the number of channels, H represents the height of the frame image, and W represents the width of the frame image. Then divide X into four groups along the channel dimension, namely X1, X2, X3, and X4. The shape of each group is [N×T×C / 4×H×W]. Each group represents 1 action granularity level. For these 4 groups, the first group retains the original information and does not perform additional processing; the remaining 3 groups perform multi-scale spatio-temporal feature extraction, and the output of the second group is connected with the input of the third group by residual connection. The above process can be expressed as:
[0026]
[0027]
[0028]
[0029]
[0030] Among them represents the outputs of the 1st to 4th groups. GTM represents the Global Time Module, SMM represents the Spatial Motion Module, and LTM represents the Local Time Module. The 1st group is used for event-class sports action recognition, with the coarsest granularity. The 2nd and 3rd groups are used for set-class sports action recognition, with the coarser granularity. The 4th group is used for element-class sports action recognition, with the fine granularity.
[0031] The output of the 1st group is consistent with its input, representing the coarsest action granularity. It is used to recognize actions with high requirements for static information and is also the coarsest-grained action in sports, such as "high-low bar" and "vault" in gymnastics.
[0032] The main component of the 2nd group is the Global Time Module (GTM). This module focuses on extracting global time information, identifying the start, end, and duration of actions. First, the input X2 is pooled along the spatial dimension, and then a 1D temporal convolution with a convolution kernel of 3 is used to capture global time features. Finally, a sigmoid activation function and a residual connection are added.
[0033] The core of the 3rd group is the Spatial Motion Module (SMM), which is used to average time information and focus on spatial changes in actions. After first fusing with the coarse-grained features extracted by the 2nd group through a residual connection, the information is then averaged in the temporal dimension. The core operation is a 2D convolution of 3×3. The 2nd and 3rd groups are used to recognize set-level sports actions with a large time span, such as "mounting the horse" and "dismounting the horse" in the vault.
[0034] The key of the 4th group is the Local Time Module (LTM). The Local Time Module is used to capture local context information of surrounding positions for fine-grained action modeling and can recognize fine-grained element-level sports actions, such as "circle" and "twist" in the horizontal bar. Its key part is a 3D convolutional layer with a convolution kernel of 3×1×1.
[0035] After that, a simple concatenation strategy is used to aggregate the outputs of multiple groups to adapt to the recognition of multi-level action categories:
[0036]
[0037] where X o ∈R N×T×C×H×W , is a set of spatio-temporal features at different levels. [,] represents the concatenation operation.
[0038] Finally, use a 1×1 convolutional layer to adjust the number of channels to the original input size and add it to the original input.
[0039] Step 6: Input the high-level spatio-temporal features output in Step 5 into a fully connected layer for high-level spatio-temporal feature mapping, and use a weight function to output the classification result of sports video action recognition;
[0040] Input the multi-scale spatio-temporal features obtained in Step 5 into multiple consecutive fully connected layers, and finally map them to K neurons equal to the number of categories in the dataset. Then use the softmax function to map K real numbers to K probabilities in the range (0,1), while ensuring that all values sum to 1, as follows:
[0041]
[0042] where z j represents the output value of the j-th neuron, and K represents the number of neurons, which is also equal to the number of categories in the dataset. represents the sum of the outputs of K neurons. Finally, select the category with the largest softmax probability value as the prediction result. During training, use this result to compare with the label and update the parameters using cross-entropy loss; during testing, use this result as the prediction result.
[0043] Step 7: Use cross-entropy loss for training until convergence;
[0044] Use cross-entropy loss to train the category probabilities obtained in Step 6 until the network converges. p represents the distribution of true labels, and q represents the predicted label distribution of the trained model. The cross-entropy loss function can measure the similarity between p and q, and its calculation formula is as follows:
[0045]
[0046] where M represents the number of samples, and K represents the number of classification categories; y ij represents whether the i-th sample belongs to the j-th class, with only two values, 0 or 1; p ij represents the probability value that the i-th sample is predicted to belong to the j-th class, and the value range is [0,1].
[0047] Step 8: Verify the effect on the validation set of the FineGym dataset.
[0048] Perform accuracy testing using the FineGym test set. During the testing process, a center cropping and single-sampling evaluation mode is adopted. Center cropping means cropping the input image to only retain the central region of the image and keeping the width and height the same. For an input image with a size of 256×256, the central 224×224 region is cropped. The single-sampling evaluation mode means that during model evaluation, each sample is sampled only once, rather than sampling multiple times and taking the average.
[0049] Finally, compare the Top-1 accuracy. The Top-1 accuracy indicates that when the model makes a prediction, for each sample, only the class with the highest prediction probability is selected as the prediction result, and then the accuracy is obtained by dividing the number of correctly predicted samples by the total number of samples. Specifically, for a classification problem, assuming there are N samples, for each sample, the model outputs the prediction probabilities for each class, and then selects the class with the highest prediction probability as the prediction result. If the prediction result is consistent with the actual label, then the sample is considered a correctly predicted sample. Then, the Top-1 accuracy is the number of correctly predicted samples divided by the total number of samples.
[0050] The beneficial effects of the present invention are as follows: The present invention provides an effective spatio-temporal feature modeling method for sports video action recognition tasks. By designing a grouping structure based on action granularity, different granularity action information is extracted using four spatio-temporal feature extraction modules with different focuses, and the multi-scale spatio-temporal features are fused using the grouping structure, which is applicable to sports action recognition with multiple levels of categories and improves the performance of sports video action recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a detailed network structure schematic diagram of the sports video action recognition method based on action granularity grouping provided by the present invention;
[0052] Figure 2 It is a flowchart of the sports video action recognition method based on action granularity grouping provided by the present invention;
[0053] Figures 3(a), 3(b), and 3(c) are respectively schematic diagrams of three spatio-temporal feature extraction modules of action granularity provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0054] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. The present invention includes but is not limited to the following embodiments.
[0055] As Figure 1 shown, the present invention provides a sports video action recognition method based on action granularity grouping. As Figure 2 shown, the specific implementation process is as follows:
[0056] 1. Video frame extraction
[0057] Use the command 'ffmpeg -i "{} / {}" -threads 1 -vf scale=-1:256 -q:v 0 "{} / {} / %06d.jpg"' in the video processing tool Ffmpeg to achieve automatic frame extraction of videos in the FineGym dataset. Among them, "-i" means the ffmpeg tool processes the input, " "{} / {}" " fills in the source video path that needs to extract frames, "-threads 1" means the number of threads is set to 1, "-vf scale=-1:256" means the specified extraction frame width is 256, "-q:v 0" means using the default output quality, " "{} / {} / %06d.jpg" " means specifying the storage location of the extracted frame images and specifying the image naming format as 6-digit numbers, and saving the pictures in.jpg format. This tool runs through the cmd command, and multi-threading can be set to speed up the extraction process.
[0058] 2. Segment random frame sampling
[0059] Average segment the sampled video frames into 8 or 16 segments, corresponding to two strategies of high-efficiency segmentation and high-precision segmentation respectively. Then randomly sample one frame from each segment as the network input, and the total input is 8 frames or 16 frames in total. The sampling strategy adopts the random sampling strategy. The specific method is to first calculate the average number of frames in each segment, set it as a frames. When sampling the i-th frame, use a random function to generate a random integer r in the range of [1, a] i , and use the following formula to determine the position of the sampled frame:
[0060] t i =(i - 1)*a + r i (9)
[0061] where t i represents the i-th frame sampled, a represents the average number of frames in each segment, and r i represents the random number generated in the range of [1, a].
[0062] 3. Video frame preprocessing
[0063] Apply random scaling and corner cropping to the extracted video frames for data augmentation, and center crop the height and width of each frame to 224 pixels.
[0064] Random scaling is to randomly adjust the size of an image. The specific steps are as follows: First, randomly select a scaling ratio range, such as [0.8, 1.2], which means the scaling ratio can be randomly selected between 0.8 and 1.2. Then, scale the original image according to the randomly selected scaling ratio to obtain the scaled image. Finally, crop the scaled image according to the size of the original image to obtain an image of the specified size.
[0065] Corner cropping is to crop a small square area from the corner of an image. The steps are as follows: First, randomly select a cropping size range, such as [0.8, 1.0], which means the cropping size can be randomly selected between 0.8 times and 1.0 times of the original image. Then, randomly select a scaled image and crop it according to the cropping size. Specifically, a point can be randomly selected from the four corner points of the upper left, lower left, upper right, and lower right of the scaled image as the starting point of the cropping area. Then, determine the size of the cropping area according to the selected cropping size to obtain the cropped image. Finally, scale the cropped image to the specified size to obtain the final image.
[0066] 4. Input into the convolutional neural network
[0067] Input the preprocessed video frames into the convolutional neural network to extract shallow spatio-temporal features. The network consists of a convolutional layer, a batch normalization layer, a ReLU layer, and a max pooling layer. Among them, the convolutional layer contains 64 convolutional kernels with a size of 7×7, a stride of 2, and a padding of 3. The batch normalization layer normalizes the output of the convolutional layer so that the mean and variance of each feature map are close to 0 and 1. The ReLU activation function layer performs the ReLU activation function operation on the normalized feature map. The max pooling layer has a size of 3×3 and a stride of 2, and performs max pooling on the feature map, and the number of output feature map channels is 64. It is followed by 4 stages, each stage contains 3, 4, 6, 3 action granularity grouping modules respectively. Finally, it contains a fully connected layer, with a total of 50 layers of network.
[0068] 5. Input into the action granularity grouping module
[0069] The action granularity grouping module first uses a 1×1 convolutional layer to adjust the number of channels to 64. Subsequently, the feature map channels are evenly divided into 4 groups, and 1 spatio-temporal feature extraction module corresponding to different action granularities is inserted into each group, and aggregated into a hierarchical grouping structure using residual connections to capture multi-scale spatio-temporal features and effectively fuse them. The action granularity grouping module forms a hierarchical structure with gradually finer action granularity from top to bottom. The output of each group from top to bottom is used to identify multi-level sports action types from coarse-grained actions to fine-grained actions.
[0070] Specifically, let the input feature be X, and its shape be [N×T×C×H×W], where N represents the batch size, T represents the number of video frames, C represents the number of channels, H represents the height of the frame image, and W represents the width of the frame image. Then X is divided into four groups along the channel dimension, namely X1, X2, X3, and X4. The shape of each group is [N×T×C / 4×H×W]. Each group represents one action granularity level. Regarding these 4 groups, the first group retains the original information without additional processing. The remaining 3 groups perform multi-scale spatio-temporal feature extraction, where the output of the second group is connected to the input of the third group through a residual connection. The above process can be expressed as:
[0071]
[0072]
[0073]
[0074]
[0075] Among them represents the outputs of the first group to the fourth group. GTM represents the global time module, SMM represents the spatial motion module, and LTM represents the local time module. The first group is used for event-class sports action recognition with the coarsest granularity. The second and third groups are used for set-class sports action recognition with the second coarsest granularity. The fourth group is used for element-class sports action recognition with a fine granularity.
[0076] Among the 4 groups, the first group retains the original information without additional processing. The remaining 3 groups add specific feature extraction modules, and the module structure is shown in Figure 3. The second group adds the global time module (GTM) to extract global time information. First, the input is pooled along the spatial dimension, and then a one-dimensional temporal convolution with a convolution kernel of 3 is used to capture the global time feature. Finally, a sigmoid activation function and a residual connection are added. The specific operations are as follows:
[0077]
[0078]
[0079]
[0080] In formula (14), H represents the height of the frame image, and W represents the width of the frame image. X2 ∈ R N×T×C / 4×H×W , means summing the elements in the H dimension and the W dimension of X2 respectively. represents the result after spatial pooling of X2. In formula (15), Conv3 represents a one-dimensional temporal convolution with a convolution kernel of 3, δ represents the sigmoid activation function, Denotes the output feature of X2 after passing through the Global Time Module (GTM). In formula (16), · denotes element-wise multiplication, and + denotes residual connection. Denotes the output feature of the second group.
[0081] The third layer is the Spatial Motion Module (SMM), which is used to average the temporal information and focus on the impact of spatial motion in the action. First, it is fused with the coarse-grained features extracted from the upper layer through a residual connection, then the information is averaged in the temporal dimension, and then two-dimensional convolution with a 3×3 kernel is used for feature extraction. Finally, a sigmoid activation function and a residual connection are also added. The specific operations are as follows:
[0082]
[0083]
[0084]
[0085]
[0086] In formula (17), Denotes the output feature of the second group, and X3 denotes the input feature of the third group. Denotes the fused feature after adding to the second layer. In formula (18), Denotes the sum of the element values in the T dimension. Denotes The output after temporal pooling. In formula (19), Conv 3×3 Denotes two-dimensional convolution with a 3×3 kernel, and δ denotes the sigmoid activation function. Denotes the output feature of X3 after passing through the Spatial Motion Module (SMM). In formula (20), · denotes element-wise multiplication, and + denotes residual connection. Denotes the output feature of the third group.
[0087] The fourth layer is the Local Time Module (LTM), which is used to capture the local context information of surrounding positions and accurately model the fine-grained action features. First, it is fused with the multi-scale features extracted from the upper layer through a residual connection, and then three-dimensional convolution with a 3×1×1 convolution kernel is used to extract local spatio-temporal features, focusing on small action changes. The specific operations can be expressed by the following formula:
[0088]
[0089]
[0090] In formula (21), Conv 3×1×1It represents a 3D convolution with a convolution kernel of 3×1×1, and δ represents the sigmoid activation function. It represents the output feature of X4 after passing through the Local Time Module (LTM).
[0091] In formula (22), · represents element-wise multiplication, and + represents residual connection. It represents the output feature of the 4th group.
[0092] Finally, a simple concatenation strategy is used to aggregate the 4 groups of outputs and together to adapt to the recognition of multi-level action categories:
[0093]
[0094] where X o ∈R N×T×C×H×W , which is a set of spatio-temporal features at different levels. [,] represents the concatenation operation.
[0095] Finally, a 1×1 convolutional layer is used to adjust the number of channels to the original input size and add it to the original input.
[0096] 6. Use a fully connected layer and a softmax layer for class prediction
[0097] As the last linear layer of the network, the fully connected layer often corresponds to the number of classifications, that is, each real number represents the weight of the class it belongs to. Therefore, the multi-scale spatio-temporal features obtained first are input into multiple consecutive fully connected layers, and are mapped to K neurons equal to the number of classes in the dataset in the last fully connected layer. After the fully connected layer, softmax converts the digital output of the last linear layer of the neural network into probabilities by obtaining the exponent of each output and then normalizing each number by the sum of these exponents. Therefore, the entire output vector, that is, the sum of all probabilities should be 1. Specifically as follows:
[0098]
[0099] where z j represents the output value of the jth neuron, and K represents the number of neurons, which is also equal to the number of classes in the dataset. represents the sum of the outputs of K neurons. Finally, the class with the largest softmax probability value is selected as the prediction result of the class. During the training process, this result is compared with the label, and the parameters are updated using the cross-entropy loss; during the testing process, this result is used as the prediction result.
[0100] 7. Use cross-entropy loss to train the action categories
[0101] The cross-entropy loss can be used as a loss function in a neural network. Let p represent the distribution of true labels and q represent the predicted label distribution of the trained model. The cross-entropy loss function can measure the similarity between p and q, and its calculation formula is as follows:
[0102]
[0103] Among them, M represents the number of samples, and K represents the number of classification categories. y ij represents whether the i-th sample belongs to the j-th class, and there are only two values, 0 or 1. p ij represents the probability value that the i-th sample is predicted to be the j-th class, and the value range is [0, 1].
[0104] Experiments were conducted using the FineGym dataset. FineGym is a large-scale gymnastics movement dataset, and the movement categories and sub-movements are organized at three levels: event, set, and element. During the training process, the initial learning rate was set to 0.01, the learning rate decayed by 0.1 every 30 epochs, and the batch size was 64. A total of 100 epochs were trained.
[0105] 8. Validation
[0106] After convergence, the validation set was used to validate the model. The validation method adopted the mode of central cropping and sampling once. Central cropping means cropping the input image to only retain the central area of the image and keeping the width and height the same. For an input image with a size of 256×256, the central 224×224 area was cropped. The evaluation mode of sampling once means that during model evaluation, each sample was only sampled once, rather than sampling multiple times and taking the average value.
[0107] Finally, the Top-1 accuracy was compared. The Top-1 accuracy indicates that when the model makes predictions, for each sample, only the class with the highest predicted probability is selected as the prediction result, and then the accuracy is obtained by dividing the number of all correctly predicted samples by the total number of samples. Specifically, for a classification problem, assuming there are K samples, for each sample, the model will output the predicted probability of each class, and then select the class with the highest predicted probability as the prediction result. If the prediction result is consistent with the actual label, then the sample is considered a correctly predicted sample. Then, the Top-1 accuracy is the number of correctly predicted samples divided by the total number of samples.
[0108] The sports video action recognition method based on action granularity grouping proposed by the present invention has an improved recognition accuracy in terms of event, set, and element categories compared to the baseline network, reaching an advanced level.
[0109] In summary, the present invention discloses a method for recognizing sports video actions based on action granularity grouping. The present invention designs a hierarchical grouping structure based on action granularity, which can extract action information of different granularities by using four spatio-temporal feature extraction modules with different focuses, and fuse multi-scale spatio-temporal features by using the hierarchical grouping structure, which is applicable to the recognition of sports actions of multi-level semantic categories. The present invention provides an effective spatio-temporal feature modeling method to improve the performance of sports video action recognition.
[0110] Firstly, the convolutional layer is used to extract shallow spatio-temporal features. Secondly, the action granularity grouping module is used to extract effective multi-scale spatio-temporal features. The action granularity grouping module can extract global time features, spatial dynamic features and local spatio-temporal features, and hierarchically and residually fuse the multi-scale spatio-temporal features from the perspective of the action granularity from coarse to fine, so as to be applicable to the recognition of sports actions of multi-level semantic categories. Finally, the cross-entropy loss is used to train the network, effectively improving the accuracy of sports action recognition.
Claims
1. A method for recognizing sports video actions based on an action granularity grouping structure, characterized in that The steps are as follows: Step 1: Extract frames from the video data of the FineGym gymnastics dataset and store them as a number of images with a fixed width; Use a video frame extraction tool to extract frames from the video data in the FineGym gymnastics dataset, unify the frame width to 256 pixels, and save the images; depending on the video length, the extraction results contain dozens to hundreds of frames; store the video frames extracted from the same video in a folder and name them in chronological order; Step 2: Use a random sampling algorithm to randomly sample the video frames extracted in Step 1 as network inputs; The video frames extracted from each video in step 1 are evenly divided into segments, and 1 frame is randomly sampled from each segment as the network input. The total number of inputs is frames, and videos are input simultaneously; Random sampling algorithm: First, calculate the average number of frames per segment, denoted as frames; When sampling the th frame, use a random function to generate a random integer within the range of , and use the following formula to determine the position of the sampled frame: (1) , Among them, indicates that the order of the frame sampled in all video frames is the frame, represents the average number of frames per segment, represents a random number with a range of Step 3: Preprocess the video frames extracted in Step 2; Apply random scaling and corner cropping to the video frames extracted in Step 2 for data augmentation, and adjust the height and width of each frame to 224 pixels; Step 4: Input the video frames processed in Step 3 into a convolutional neural network and use convolutional blocks for feature extraction; Input the sequence of preprocessed video frames in Step 3 into a multi-layer convolutional neural network. The convolutional neural network mainly consists of a convolutional layer, a batch normalization layer, a ReLU layer, and a max pooling layer, which is used to extract the shallow features of the video frames to obtain feature maps; Step 5: Input the feature maps obtained in Step 4 into a continuous four-stage action granularity grouping module to obtain high-level spatio-temporal features that fuse multi-scale spatio-temporal information; the specific content of the continuous four-stage action granularity grouping module is as follows: first use a convolutional layer to adjust the number of channels of the feature map, and then evenly divide the channels of the feature map into four groups. Each group uses four spatio-temporal feature extraction modules with different focuses; use residual connections to construct a hierarchical grouping structure, then fuse the features of the four groups containing different granularity action information, and finally use a convolutional layer again to adjust the number of channels to the same as the input channel number, and add the fused features to the input features to obtain high-level spatio-temporal features; First, set the input feature after adjusting the number of channels to , and set its shape to , where N represents the batch size, T represents the number of video frames, C represents the number of channels, H represents the height of the frame image, and W represents the width of the frame image; then is divided into four groups along the channel dimension, which are respectively , , and ; the shape of each group is , and each group represents 1 action granularity level; the first group retains the original information and no additional processing is performed; the remaining 3 groups perform multi-scale spatio-temporal feature extraction; among them, the output of the second group is connected to the input of the third group through a residual connection; the above process is expressed as: (2) , (3) , (4) , (5) , Among them, , , representing the outputs from the 1st to the 4th groups; GTM represents the global time module, SMM represents the spatial motion module, and LTM represents the local time module; the 1st group is used for event-class sports action recognition with the coarsest granularity; the 2nd and 3rd groups are used for set-class sports action recognition with the coarser granularity; the 4th group is used for element-class sports action recognition with the finest granularity; Step 6: Input the high-level spatio-temporal features output in Step 5 into a fully connected layer for high-level spatio-temporal feature mapping, and use a weight function to output the sports video action recognition classification results; Input the multi-scale spatio-temporal features obtained in step 5 into multiple consecutive fully-connected layers, and finally map them to the number of neurons equal to the number of categories in the dataset ; then use the softmax function to map real numbers to class probabilities in the range of (0, 1), while ensuring that the sum of all values is 1, as follows: (7) , Among them, represents the output value of the -th neuron, represents the number of neurons, which is also equal to the number of dataset categories; represents summing the outputs of neurons; finally, select the category with the largest softmax probability value as the result of the predicted category; during the training process, use this result to compare with the label and update the parameters using cross-entropy loss; during the testing process, use this result as the prediction result; Step 7: Use cross-entropy loss for training until convergence; Use cross-entropy loss to train the class probabilities obtained in Step 6 until the network converges; p represents the distribution of true labels, and q is the predicted label distribution of the trained model. The cross-entropy loss function can measure the similarity between p and q, and its calculation formula is as follows: (8) , Among them, represents the number of samples, represents the number of categories for classification; represents whether the th sample belongs to the th category, with only two values, 0 or 1; represents the probability value that the th sample is predicted to belong to the th category, and the value range is [0, 1]; Step 8: Verify the effect on the validation set of the FineGym dataset; Use the FineGym test set for accuracy testing; adopt a center cropping and single-sampling evaluation mode during the testing process; center cropping means cropping the input image to only retain the central area of the image and keep the width and height the same; for an input image with a size of 256×256, crop the central 224×224 area; the single-sampling evaluation mode means that during model evaluation, each sample is only sampled once instead of sampling multiple times and taking the average; Finally, compare the Top-1 accuracy; the Top-1 accuracy represents the accuracy obtained by dividing the number of all correctly predicted samples by the total number of samples when the model, for each sample, selects only the class with the highest prediction probability as the prediction result during prediction; specifically, for a classification problem, assuming there are samples, for each sample, the model will output the prediction probabilities for each class and then select the class with the highest prediction probability as the prediction result; if the prediction result is consistent with the actual label, then the sample is considered a correctly predicted sample; then, the Top-1 accuracy is the number of correctly predicted samples divided by the total number of samples.
2. The method for recognizing sports video actions based on the action granularity grouping structure according to claim 1, wherein, In Step 3, Random scaling randomly adjusts the size of an image and is achieved through the following steps: First, randomly select a scaling ratio range; then, scale the original image according to the randomly selected scaling ratio to obtain a scaled image; finally, crop the scaled image according to the size of the original image to obtain an image of a specified size; Corner cropping crops a small square area from the corner of an image and is achieved through the following steps: First, randomly select a cropping size range; then, randomly select a scaled image and crop it according to the cropping size; specifically, randomly select a point from the four corner points of the top left, bottom left, top right, and bottom right of the scaled image as the starting point of the cropping area; Then, determine the size of the cropping area according to the selected cropping size to obtain a cropped image; finally, scale the cropped image to the specified size to obtain the final image.
3. The method for recognizing sports video actions based on an action granularity grouping structure according to claim 1, characterized in that In step 4, In the convolutional neural network, the convolutional layer contains 64 convolutional kernels with a size of 7×7, a stride of 2, and a padding of 3; the batch normalization layer then performs batch normalization on the output of the convolutional layer to make the mean and variance of each feature map close to 0 and 1; the ReLU activation function layer performs the ReLU activation function operation on the normalized feature map; the max pooling layer has a size of 3×3 and a stride of 2, performs max pooling on the feature map, and the number of output feature map channels is 64.
4. The sports video action recognition method based on an action granularity grouping structure according to claim 1, wherein In step 5, Input the feature map obtained in step 4 into the continuous 4-stage action granularity grouping module, with each stage containing 3, 4, 6, and 3 action granularity grouping modules respectively; the action granularity grouping module first includes a 1×1 convolutional layer for adjusting the number of channels; then, the channels of the input feature map are evenly divided into 4 groups, and 1 spatio-temporal feature extraction module corresponding to different action granularities is inserted into each group, and aggregated into a hierarchical grouping structure using residual connections to capture multi-scale spatio-temporal features and effectively fuse them; the action granularity grouping module forms a hierarchical structure with gradually finer action granularity from top to bottom; the output of each group from top to bottom is used to identify multi-level sports action types from coarse-grained actions to fine-grained actions; Output of Group 1 It is consistent with its input and represents the coarsest action granularity; The main component of Group 2 is the global time module, which focuses on extracting global time information and identifying the start, end, and duration of actions; first, the input is pooled along the spatial dimension, and then a 1D temporal convolution with a kernel size of 3 is used to capture global temporal features; finally, a sigmoid activation function and a residual connection are added. The core of the third group is the spatial motion module for averaging temporal information and focusing on spatial changes in the action; first, it is fused with the coarse-grained features extracted by the second group through residual connections, and then the information is averaged in the temporal dimension; the core operation is a 3×3 two-dimensional convolution; the second and third groups are used to identify set-level sports actions with a relatively large time span; The key of the fourth group is the local time module for capturing local context information of surrounding positions for fine-grained action modeling and can identify fine-grained element-level sports actions; its key part is a three-dimensional convolutional layer with a convolutional kernel of 3×1×1; After that, a simple concatenation strategy is used to aggregate the outputs of multiple groups to adapt to the recognition of multi-level action categories: (6) , Among them, is a set of spatio-temporal features at different levels; represents a splicing operation; Finally, use a 1×1 convolutional layer to adjust the number of channels to the original input size and add it to the original input.
Citation Information
Patent Citations
CCA and 2PKNN based automatic image annotation method
CN105808752A
Full-reference video quality evaluation method based on two stages of adaptive sampling and multi-scale time sequence
CN115239647A