A motion quality evaluation method based on time-aware feature learning

By using a time-aware attention mechanism and regression methods to evaluate motion quality, this approach addresses the problem of insufficient video feature extraction in existing technologies, achieving a more efficient motion quality evaluation.

CN113920584BActive Publication Date: 2025-11-28SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111207579.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-15
Publication Date
2025-11-28
Estimated Expiration
2041-10-15

AI Technical Summary

Technical Problem

Existing technologies suffer from poor performance and insufficient generalization ability in motion quality assessment due to the large video dimensionality and inconsistent features. Furthermore, 3D convolutional networks are prone to losing motion information changes during video feature extraction.

Method used

A time-aware attention mechanism is introduced. By dividing the video into segments and extracting segment features using the I3D network, and combining the subtitle generation module, time-aware module, and average aggregation module, motion change information is captured, and a regression method is used for evaluation.

Benefits of technology

More discernible video feature representations were extracted, improving the accuracy of action quality assessment and the stability of the model. The model performance was further enhanced by introducing adversarial loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113920584B_ABST
    Figure CN113920584B_ABST
Patent Text Reader

Abstract

The application discloses a motion quality evaluation method based on time-aware feature learning. The method learns the segment features in the video by using a 3D convolution network, and learns the relationship between the segments by a time-aware module. The relationship can capture the transformation information of the motion to improve the accuracy of the motion quality evaluation. Then, the segment relationship is used to aggregate the features of the whole video, and the video features can be directly used for score prediction of the motion. In addition, two auxiliary tasks of subtitle generation and motion recognition are introduced to enable the 3D convolution network to learn more rich feature representation. Finally, in order to ensure that the time-aware module can more accurately capture the transformation information of the motion, an adversarial loss is introduced to stabilize the whole model. The application can extract discriminative feature representation on the motion quality evaluation dataset, and effectively improve the Spearman correlation coefficient in the motion quality evaluation problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision, and more particularly to a method for evaluating action quality using a time-aware attention mechanism. Background Technology

[0002] Motion quality assessment is a technique widely used in various fields such as gesture detection, motion event analysis, and operational skill analysis, attracting considerable attention in the field of computer vision. It aims to evaluate the quality of people's actions in videos, assigning high scores to well-performing actions and low scores to poorly performed ones. Early work attempted to analyze people's actions using hand-designed features, but due to the large dimensionality of videos and the inconsistency of features for different events, this resulted in poor performance and insufficient generalization ability.

[0003] Due to the powerful processing capabilities of deep neural networks for high-dimensional problems, using deep learning for motion quality assessment has become a trend in recent years. Many studies have attempted to use 3D convolutional networks to extract features from video segments and then simply aggregate them into a video feature representation. This may lose the transformation of motion information, making it difficult for the computer to accurately capture the different impacts of various segments on the final assessment, thus affecting the assessment results. Attention mechanisms, which have performed exceptionally well in computer vision, have been proposed and have significantly improved the performance of existing methods. Therefore, it is natural to introduce attention mechanisms into motion quality assessment models to distinguish the impact of different segments on the assessment, thereby extracting more effective video feature representations. Summary of the Invention

[0004] Objective: In this paper, instead of simply aggregating segment features to obtain video features, we introduce a time-aware attention mechanism to capture action transformation information, thereby obtaining a more discriminative video feature representation. Then, a regression method is used to evaluate the quality of the action. This invention provides an action quality evaluation method based on time-aware feature learning.

[0005] Technical solution: A method for evaluating action quality based on time-aware feature learning, characterized by the following steps:

[0006] Step 1: Input the video Divided into A fragment And downsample and data augmentation are performed on each frame;

[0007] Step 2: Use the I3D network to extract features from each segment. and fragment-level features Input the data into the subtitle generation module to calculate the subtitle generation loss;

[0008] Step three: put the segment-level features into the time-aware module and the average aggregation module to get the video-level features respectively and ;

[0009] Step four: put the video-level features into the score prediction module to calculate the score prediction loss and the adversarial loss and respectively

[0010] Step five: put the video-level features into the action recognition module to calculate the action recognition loss ;

[0011] Step six: minimize the loss and update the model parameters

[0012] Step seven: the score calculated by the score prediction module after the video-level features is the predicted score of the video .

[0013] Further, in step one, the downsampling is center cropping of the picture, and the data augmentation is random horizontal rotation of the picture.

[0014] Further, in step two, the subtitle generation module is a gated recurrent unit (GRU), which takes the video features as the initial input to obtain an initial output vector and an initial state vector, then updates the output vector and the state vector by recycling the maximum length of the generated subtitles for a number of times, and outputs a word each time, finally composing the entire subtitle output; the subtitle generation loss is a negative log-likelihood loss function.

[0015] Further, in step three, the time-aware module consists of two fully connected layers, a ReLU activation function, and a Softmax normalization exponential function; the first fully connected layer compresses the features, the ReLU activation function increases the nonlinearity of the features, the second fully connected layer compresses the feature dimension to 1, and the Softmax normalizes the features to obtain weights; finally, the weights are point multiplied with the segment features and summed to obtain discriminative video features; the average aggregation module directly averages the input segment features to obtain the video features.

[0016] Further, in step four, the adversarial loss is a hinge loss composed of the difference between the predicted score and the true value and the difference between the predicted score and the true value.

[0017] ​​​Further, the score prediction module is a network structure composed of two fully connected layers; wherein the first fully connected layer performs dimension reduction processing on the features, and the second fully connected layer compresses the feature dimension to 1 to predict the score.

[0018] Further, in step five, the action recognition module is a classification model composed of multiple fully connected layers, and each sub-action has a fully connected layer for classification, wherein the position is classified into three categories, the twisting is classified into eight categories, the rotation type is classified into four categories, the rotation number is classified into ten categories, and the arm strength is classified into two categories, and the action recognition loss is a cross-entropy loss.

[0019] Further, in step six, the Adam algorithm is used to minimize the loss, and the loss includes the subtitle generation loss, the action recognition loss, the score prediction loss, and the adversarial loss.

[0020] Beneficial effects: the present application provides a deep learning method for action quality assessment, which can extract more recognizable video feature representations compared to the prior art. In order to improve the stability and effectiveness of the model, an adversarial loss is introduced to improve the performance of the model. The following examples show that the present application can effectively learn high-level features with action information discriminability in action quality assessment. In addition, the method proposed in the present application has good effect on public data sets. BRIEF DESCRIPTION OF DRAWINGS

[0021] Figure 1 The method flowchart of the present application;

[0022] Figure 2 The algorithm framework of the present application;

[0023] Figure 3 Comparison of the present application with other methods. DETAILED DESCRIPTION

[0024] The present application will be further described in detail below in combination with the drawings and specific embodiments:

[0025] The present embodiment provides a method for learning time-aware features and for MTL-AQA dataset action quality assessment, which divides the video into multiple segments, then obtains segment features through a 3D convolutional neural network, and finally captures action information through a time-aware module to obtain video features, which are regressed to obtain good evaluation effect.

[0026] The flow of the method is shown in Figure 1 The algorithm framework is shown in Figure 2

[0027] Step one: divide the input video into ​segments and down-sampling and data augmentation are applied to each frame;

[0028] Step two: I3D network is used to extract the features of each segment and the segment-level features are fed into the caption generation module to calculate the caption generation loss;

[0029] Step three: segment-level features are fed into the temporal-aware module and the average aggregation module to obtain video-level features and ;

[0030] Step four: video-level features and are fed into the score prediction module to calculate the score prediction loss and the adversarial loss;

[0031] Step five: video-level features are fed into the action recognition module to calculate the action recognition loss;

[0032] Step six: minimize the loss and update the model parameters;

[0033] Step seven: the scores calculated by the score prediction module after the video-level features are calculated are the predicted scores of the video .

[0034] In this example, the MTL-AQA dataset is the largest dataset dedicated to action quality assessment. It contains 1412 video samples. All samples come from 16 events in the diving competition of the competition, and each video contains 103 frames, which have different perspectives and camera angles. The dataset contains samples of male and female athletes, individual and synchronized diving, 3m and 10m diving, judges' final action quality scores, task difficulty levels, live commentary of the event, and fine-grained action labels (e.g., position, rotation method, rotation number).

[0035] During training, the 3D convolutional network loads the pre-trained parameters on the Kinects dataset. The weight of the score prediction loss is 1, the weight of the action recognition loss is 1, the weight of the caption generation loss is 0.01, and the weight of the adversarial loss is 0.05. The number of training times is 100, and the learning rate of the Adam optimizer is 10 -4, training batch size is 3. And data augmentation is performed by center crop, temporal augmentation and random horizontal flip. For 3D convolutional network, the size of input frame is cropped to 224x224, the extracted feature dimension is 1024, and the structure of fully connected layer of regression score is 1024x512x1. In addition, for the dimension input to the auxiliary task, only the corresponding output feature size is changed.

[0036] Some recent deep learning based methods and our method are compared on MTL-AQA dataset. The results are shown in Figure 3 Our method is better than most of the methods. And for ResNet34-(2+1)D-WD, its method is slightly higher than our method, but its deep neural network is deeper than our method, and the training time and prediction time are longer. Therefore, the algorithm proposed in the present application has great advantages in practical application.

Claims

1. A method for evaluating action quality based on time-aware feature learning, characterized in that, It includes the following seven steps: Step 1: Input the video Divided into A fragment And downsample and data augmentation are performed on each frame; Step 2: Use the I3D network to extract features from each segment. and fragment-level features The input subtitle generation module calculates the subtitle generation loss; Step 3: Extract fragment-level features Video-level features are obtained by inputting them into the time-aware module and the average aggregation module, respectively. and ; Step 4: Extract video-level features and Input the score prediction module separately to calculate the score prediction loss and the adversarial loss; Step 5: Extract video-level features The input action recognition module calculates the action recognition loss; Step 6: Minimize the loss and update the model parameters; Step 7: Utilize the score prediction module to analyze video-level features The score after calculation is the score for the video. The predicted score; In step one, downsampling involves centering the image, and data augmentation involves randomly rotating the image horizontally. In step two, the subtitle generation module uses a gated recurrent network (GRU), which takes video features as initial input to obtain an initial output vector and an initial state vector, then iterates through the maximum length of the subtitle generation to update its output vector and state vector. Each iteration outputs one word, which is then combined to form the entire subtitle output. The subtitle generation loss is the negative log-likelihood loss function of the generated subtitle and the actual subtitle. In step three, the time-aware module consists of two fully connected layers, a ReLU activation function, and a Softmax normalization exponential function. The first fully connected layer compresses the features, the ReLU activation function increases the non-linearity of the features, the second fully connected layer compresses the feature dimension to 1, and Softmax normalizes the features to obtain weights. Finally, the weights are multiplied by the segment features and summed to obtain discriminative video features. The average aggregation module directly averages and sums the input segment features to obtain the video features. In step four, the adversarial loss is caused by The difference between the predicted score and the true value, and the sum of... The page loss is composed of the difference between the predicted score and the true value; The score prediction module is a network structure consisting of two fully connected layers; the first fully connected layer performs dimensionality reduction on the features, and the second fully connected layer compresses the feature dimension to 1 to predict the score. In step five, the action recognition module is a classification model composed of fully connected layers. For each sub-action, there is a fully connected layer for classification. Position is divided into 3 categories, twisting into 8 categories, rotation type into 4 categories, number of rotations into 10 categories, and arm strength into 2 categories. The action recognition loss is cross-entropy loss. In step six, the Adam algorithm is used to minimize the loss, which includes caption generation loss, action recognition loss, score prediction loss, and adversarial loss.

Citation Information

Patent Citations

  • No-reference video quality evaluation method based on three-dimensional spatial-temporal feature decomposition

    CN112085102A

  • Multi-mode diving event intelligent evaluation method based on mark distribution learning

    CN113255489A