A method for evaluating the quality of athletes' behaviors based on deep learning

By adopting deep learning methods and pipeline self-attention mechanisms in the evaluation of athlete behavior quality, the problems of low evaluation accuracy, slow running speed and poor interpretability in the prior art are solved, and efficient and accurate evaluation of athlete behavior quality is achieved, and training efficiency and quality are improved.

CN113989920BActive Publication Date: 2025-06-03FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111193385.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-13
Publication Date
2025-06-03
Estimated Expiration
2041-10-13

AI Technical Summary

Technical Problem

The prior art has problems in the evaluation of athlete's behavior quality, low evaluation accuracy, slow running speed and poor interpretability, especially when dealing with long-term complex action sequences, it is difficult to achieve efficient analysis.

Method used

The athlete's behavior quality assessment method based on deep learning is adopted, combining human body tracking unit, human body posture estimation unit, action sequence feature extraction and enhancement unit, score prediction unit and display unit, feature extraction and enhancement through the I3D network and pipeline self-attention mechanism to achieve high-precision behavior quality assessment.

Benefits of technology

It realizes high-precision assessment of athletes' behavior quality, and can efficiently evaluate them at all behavioral levels and at various action stages, saving manpower and material resources in athletes' training process, and improving training efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989920B_ABST
    Figure CN113989920B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for evaluating the behavior quality of athletes based on deep learning. The evaluation system based on this includes a human body tracking unit, a human body pose estimation unit, an action sequence feature extraction and enhancement unit, a score prediction unit, and a display unit; the video is input into the human body tracking unit, target detection is performed on each frame of the video, and the detection frame of each frame is obtained as the tracking result, and the tracking result is visualized on the display unit; the human body pose estimation unit obtains the tracking result, estimates the pose of the athlete in each frame, and obtains the key point information as the pose estimation result, and the pose estimation result is visualized on the display unit; the action sequence feature extraction and enhancement unit takes the video, the tracking result, and the pose estimation result as inputs, and obtains video features after feature extraction, feature enhancement, and feature fusion; the score prediction unit takes the video features as inputs and performs full action process quality evaluation and stage behavior quality evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of athlete behavior quality assessment, and in particular to an athlete behavior quality assessment method based on deep learning. Background Art

[0002] In recent years, with the continuous development of real-time sports broadcast technology, more and more regular sports events have recorded the whole process of athletes' competitions and saved the final scores of athletes. How to make better use of these data to bring more effective help to the improvement of athletes' skills in subsequent training has become a problem worthy of research. At present, although there are some algorithms to evaluate the postures of athletes through visual information, they are only limited to the level of posture and short-term action sequence perception, and cannot efficiently analyze long-term complex action sequences. Therefore, there is an urgent need for an intelligent system in sports training that can automatically score and evaluate athletes' action sequences to save the labor and material costs in the training stage of athletes and improve the training efficiency.

[0003] In the prior art, for the evaluation of athletes' postures and behavior sequences, the existing models mainly focus on two technical points: motion perception technology and motion evaluation technology.

[0004] Motion perception technology is often considered first. Motion perception refers to locating the position of an athlete through raw video and image information, checking the athlete's posture, and semantic segmentation. There are already many such algorithms at present, such as the YOLO detector for object detection technology, the Alphapose algorithm for pose estimation tasks, and the Mask R-CNN for semantic segmentation tasks. These algorithms have achieved excellent performance in various public datasets. Although they can be directly used by motion evaluation systems, these algorithms are all based on deep learning technology, and the running process has high requirements for equipment, which limits their usage scenarios.

[0005] After completing motion perception, motion evaluation technology is needed to comprehensively evaluate the motion sequence to obtain the final prediction result and detect and feedback low-score behaviors. At present, although some work has designed behavior quality assessment models, these models uniformly use the entire video information as input, ignoring the differences between athletes and background information in the video. This undifferentiated feature extraction and feature enhancement will slow down the running efficiency of the evaluation model on the one hand, and on the other hand, it will lead to the mixing of video information, thus affecting the final behavior quality assessment result.

[0006] Therefore, the disadvantages of the prior art are mainly reflected in the following three aspects:

[0007] 1. Low evaluation accuracy: Currently, video-based behavior quality evaluation technologies only take the original video as input, extract features from the video through 3D convolution kernels, and finally complete score prediction through a regressor. This processing method does not consider the differences between the foreground and the background during feature extraction. For example, the movement area of athletes should receive more attention rather than the background advertisements and audiences. This unified processing method will cause important information to be submerged in complex background information, ultimately leading to a deterioration in the evaluation performance of the model.

[0008] 2. Slow running speed: The training and inference processes of 3D convolutional networks consume a large amount of memory and require the device to have extremely high computing power. These problems will lead to an excessive delay in the operation of the video analysis system, unable to provide timely behavior quality feedback, and ultimately reducing the training efficiency. Therefore, a low number of parameters, low computational volume, and high computational efficiency have become essential features of an excellent behavior quality evaluation system.

[0009] 3. Poor interpretability: The information processing process of the behavior quality evaluation system based on LSTM (Long Short-Term Memory) is divided into two major steps: frame-by-frame feature extraction and feature joint analysis. First, feature extraction is performed on each frame in the video through a 2D convolutional neural network, and then LSTM is used to aggregate and analyze the feature sequence, finally completing the behavior quality prediction. This method can only analyze the entire video and cannot be accurate to each movement stage, so it is impossible to conduct segmented evaluation of the video and provide improvement suggestions. Summary of the Invention

[0010] The object of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a method for evaluating the behavior quality of athletes based on deep learning. The evaluation system based on this method includes a human body tracking unit, a human body pose estimation unit, an action sequence feature extraction and enhancement unit, a score prediction unit, and a display unit. The human body tracking unit tracks the athletes in the original competition video to obtain continuous detection frames. The human body pose estimation unit detects the key points of the athletes' bodies. The action sequence feature extraction and enhancement unit uses the I3D convolutional neural network and the TubeSelf-attention Mechanism to complete feature extraction and feature enhancement respectively, and obtains video features. The score prediction unit takes the video features as input and predicts the evaluation result of the behavior quality. In the whole process, the tracking result and the pose estimation result are respectively extracted from the original competition video, and feature extraction is carried out through the I3D neural network, and the efficient and effective enhancement of the features is completed by using the tube attention mechanism, and finally the high-precision behavior quality evaluation is realized. The quality evaluation result provides local and global analysis for the behavior quality of the athletes. The athletes can carry out targeted training, which can save the human and material costs in the training stage of the athletes, improve the training efficiency, and is more instructive.

[0011] The object of the present invention can be achieved by the following technical solutions:

[0012] A method for evaluating the behavior quality of athletes based on deep learning. The evaluation system based on this method includes a human body tracking unit, a human body pose estimation unit, an action sequence feature extraction and enhancement unit, a score prediction unit, and a display unit. Specifically, the processes in each unit are as follows:

[0013] The video is input into the human body tracking unit. The human body tracking unit performs target detection on each frame of the video to obtain the detection frame of each frame as the tracking result, and visualizes the tracking result on the display unit.

[0014] The human body pose estimation unit obtains the tracking result, estimates the pose of the athletes in each frame, and obtains the key point information as the pose estimation result, and visualizes the pose estimation result on the display unit.

[0015] The action sequence feature extraction and enhancement unit takes the video, the tracking result, and the pose estimation result as input, and obtains video features after feature extraction, feature enhancement, and feature fusion.

[0016] The score prediction unit takes the video features as input and performs full-action process quality evaluation and stage behavior quality evaluation.

[0017] According to the quality evaluation of the athletes' movement behaviors in the video, the deficiencies of the athletes' movement behaviors are obtained, and the athletes are subjected to special training based on the deficiencies of the movement behaviors.

[0018] Further, the human body tracking unit uses the YOLO detector and the Siammask framework for object detection: the YOLO detector is used for object detection in the initial frame of the video to obtain the detection boxes of the initial frame, and the Siammask framework is used as a single-object tracker for object detection in the consecutive frames after the initial frame to obtain the detection boxes of each frame after the initial frame.

[0019] Further, the human body pose estimation unit uses the Alphapose framework to estimate the poses of the athletes in each frame: after obtaining the detection boxes of each frame in the video, the Alphapose framework is used to estimate the poses of the athletes in each frame, the Alphapose framework generates the key point information of each frame, and the Kalman filter algorithm is used to process the key point information of each frame to obtain the pose estimation result.

[0020] Further, the action sequence feature extraction and enhancement unit extracts features from the video to obtain the first feature, combines the first feature with the tracking result and uses the pipeline self-attention mechanism for feature enhancement to obtain the enhanced first feature, the action sequence feature extraction and enhancement unit extracts features from the pose estimation result to obtain the second feature, and the second feature and the enhanced first feature are fused through a fully connected layer to obtain the video feature.

[0021] Further, the action sequence feature extraction and enhancement unit uses the I3D neural network to extract features from the video to obtain the first feature, and uses the graph convolutional neural network to extract features from the pose estimation result to obtain the second feature.

[0022] Further, when performing the overall action process quality assessment, all video features are subjected to temporal global average pooling and then sent to the fully connected layer to complete the quality assessment. When performing the stage behavior quality assessment, the video features of a segment of video are sent to the fully connected layer to complete the quality assessment.

[0023] Further, the score prediction unit uses the I3D neural network. When performing the overall action process quality assessment, all video features are subjected to temporal global average pooling and then sent to the fully connected layer of the I3D neural network to complete the quality assessment. When performing the stage behavior quality assessment, the video features of a segment of video are sent to the fully connected layer of the I3D neural network to complete the quality assessment.

[0024] Further, the specific process of enhancing the first feature using the spatio-temporal self-attention mechanism in combination with the tracking result is as follows: Quantize and align the first feature with the detection box to generate a feature map mask, fuse the masks according to the ratio of the number of video frames to the number of the first features to generate a spatio-temporal pipeline, perform a sparse enhancement operation on the first feature using the spatio-temporal self-attention mechanism inside the spatio-temporal pipeline, and fuse the first feature with the first feature after the sparse enhancement through a residual connection to obtain the enhanced first feature.

[0025] Further, after obtaining the tracking result and the first feature, determine the ratio N:1 of the number of detection boxes to the number of temporal dimensions of the first feature according to the ratio of the number of video frames to the number of the first features, where N>1. Determine the mask corresponding to each detection box, and obtain the ratio of the detection box covering the feature network of the first feature. If the ratio is greater than the preset threshold, the first feature is selected; otherwise, the first feature is excluded. After calculating the masks of N detection boxes, complete the fusion of the masks through a bitwise AND operation to generate a spatio-temporal pipeline.

[0026] Further, the spatio-temporal self-attention mechanism inside the spatio-temporal pipeline is expressed as:

[0027]

[0028] where p represents the output position to be calculated, the pair (c, t, i, j) traverses all the positions of the first feature in the spatio-temporal pipeline, the output feature y has the same size as the input feature x, the f function is a distance metric function, the g function is a feature mapping function, and the response value is normalized by the normalization factor C(x)=∑ c ∑ t Ω c,t for normalization.

[0029] Compared with the prior art, the present invention combines a human tracking unit, a human pose estimation unit, an action sequence feature extraction and enhancement unit, a score prediction unit, and a display unit to realize multi-channel perception of the position and pose information of athletes. The video feature extraction is completed through the I3D network, and a spatio-temporal self-attention mechanism is proposed to efficiently and effectively enhance the action features of athletes by means of detection box information. The fusion of video features and second features is completed through a feature fusion technology, and two behavior quality evaluation modes, namely global and local, are designed in the score prediction unit. This system can efficiently complete the evaluation of the behavior quality of athletes at the full behavior level and each action stage level, save a large amount of manpower and material resources in the athlete training process, improve the training efficiency and quality, and provide a strong basic guarantee for the continuous improvement of the competitive level of athletes and the development of the sports cause. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is the system principle block diagram of the present invention;

[0031] Figure 2 is the principle block diagram of the human body tracking unit;

[0032] Figure 3 is the principle block diagram of the human body pose estimation unit;

[0033] Figure 4 is the principle block diagram of the action sequence feature extraction and enhancement unit;

[0034] Figure 5 is the principle block diagram of the score prediction unit;

[0035] Figure 6 is the schematic diagram of the pipeline self-attention mechanism. Specific implementation manners

[0036] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0037] Embodiment 1:

[0038] A method for evaluating the behavior quality of athletes based on deep learning, as Figure 1 shown, the evaluation system on which it is based includes a human body tracking unit, a human body pose estimation unit, an action sequence feature extraction and enhancement unit, a score prediction unit, and a display unit.

[0039] The human body tracking unit solves the problems of missed detection and false detection in traditional object detection schemes. Generally, the tracking algorithm first uses a frame-by-frame detection algorithm to obtain the bounding box of each frame (i.e., the detection box mentioned in this application, which can also be called the tracking box, etc.), and then completes the tracking task through the Kalman Filter and the Person Re-identification (ReID) algorithm. However, this method is only applicable to general monitoring environments, requiring that the human body cannot undergo huge deformations in a short time, and is not suitable for sports scenarios where the human body posture is often highly distorted and in a high-speed movement state. Therefore, the present invention introduces a Single Object Tracker (SOT) into the task of evaluating the behavior quality of athletes. Different from the frame-by-frame detection strategy in ordinary tracking algorithms, the single object tracker skips the frame-by-frame detection stage and can output stable tracking results in a coherent time series on the premise of a given initial frame bounding box, providing position information for subsequent feature enhancement. In this application, the initial frame (or called the first frame) uses the YOLO detector for object detection, and then irrelevant objects are filtered according to the size relationship of the detection boxes, providing initial frame information for the single object tracker.

[0040] The human body pose estimation unit can jointly analyze the human body pose in both spatial and temporal dimensions, and finally obtain high-precision human key point information. Traditional pose estimation algorithms often only focus on single-frame scenarios, and the pose estimation algorithms in videos only apply the algorithms for frame-by-frame detection and then splicing. This simple transfer method cannot properly handle the body distortion and self-occlusion problems of athletes in videos. The present invention adds a tracking mechanism based on a Kalman filter on the basis of single-frame key point detection to handle the misdetection and missed detection of key points. In addition, the neural network of the human body pose estimation unit in the present invention is 3-bit quantized to save computing resources and improve the inference efficiency of the neural network.

[0041] The action sequence feature extraction and enhancement unit takes the video, tracking result, and pose estimation result as inputs and outputs video features. First, the I3D neural network is used to extract features from the motion video segment. Then, the detection box and the first feature (which can also be called the video feature map) are quantized and aligned to generate a feature map mask (Mask), and the masks are fused according to the ratio of the number of video frames to the number of feature maps to generate a spatio-temporal tube. Inside the spatio-temporal tube, a tube self-attention mechanism is adopted to complete the sparse enhancement operation of the first feature, and the enhanced first feature and the original first feature are fused through a residual link to obtain the enhanced first feature. The graph convolutional network (GCN) is used to extract features from the pose estimation result to obtain the second feature. The second feature (i.e., the feature of the pose estimation result) and the enhanced first feature will be fused through the fully connected layer of the I3D neural network to generate video features for subsequent behavior quality evaluation.

[0042] The score prediction unit (for behavior quality evaluation) takes the video features output by the action sequence feature extraction and enhancement unit as inputs, makes predictions in the I3D neural network, outputs the aggregated video features, and completes the final score prediction. The score prediction unit is divided into two modes: full action process quality evaluation and stage behavior quality evaluation. In the full action process quality evaluation mode, it is necessary to first perform temporal global average pooling on the video features extracted from all video segments, and then send them to the fully connected layer to complete the prediction; in the stage behavior quality evaluation mode, the video features of each video segment are directly sent to the fully connected layer to complete the prediction, so the advantages and disadvantages of the athlete's actions in each stage can be observed.

[0043] The display module can visualize the results of each stage and the final result in the above execution process, and provide action improvement suggestions for the stage scores.

[0044] Specifically, the process within each unit is as follows:

[0045] (1) Video input to the human body tracking unit. The human body tracking unit performs object detection on each frame of the video, obtains the detection boxes of each frame as the tracking results, and visualizes the tracking results on the display unit;

[0046] In this embodiment, the schematic diagram of the human body tracking unit is as shown in Figure 2 shown. The human body tracking unit uses the YOLO detector and the Siammask framework for object detection: in the initial frame of the video, the YOLO detector is used for object detection to obtain the detection box of the initial frame, and the Siammask framework is used as a single-object tracker to perform object detection on the consecutive frames after the initial frame to obtain the detection boxes of each frame after the initial frame.

[0047] (1) Given the first-frame tracking box: Traditional tracking methods consist of two modules: single-frame object detection and fusion of consecutive-frame detection results. This method makes the tracking results highly dependent on the detection results (i.e., detecting the target in a single frame. In this application, the target to be detected is an athlete), and the detection effect in sports videos is not optimistic. The high-speed movement and severe deformation of athletes will lead to missed detections and false detections by the detector. At the same time, the audience in the background will also interfere with the object detection of athletes. Therefore, the present invention adopts a strategy based on a single-object tracker. Usually, athletes are in a static state during the preparation stage of sports action execution and are relatively easy to be recognized by the object detector. Therefore, the present invention first uses the YOLO detector to complete object detection in the initial frame, and then filters out irrelevant targets according to the size relationship of the detection boxes to provide initial frame information for the single-object tracker.

[0048] (2) Single-object tracking: Currently, many mature frameworks have been successively proposed in the field of single-object tracking. Considering the particularity of athlete target tracking in this application, the Siammask framework is adopted as the single-object tracker. Siammask is a simple method that can perform visual object tracking and semi-supervised object segmentation in real time. Siammask adopts a fully connected siamese network structure during the training process and enhances the loss function using a binary segmentation task; during testing, it can generate object segmentation masks and rotation bounding boxes at 55 FPS. The adoption of the single-object tracker strategy makes up for the detection difficulties in traditional single-frame detection methods and can obtain more accurate tracking boxes and segmentation information, providing important reference information for subsequent feature enhancement and behavior quality assessment.

[0049] (3) Visualization of tracking results: This part is a component in the display module. Siammask can generate both the bounding box and the segmentation mask of the tracking target. Therefore, the present invention visualizes the tracking results in the display module to provide reference for coaches and athletes.

[0050] (2) The human body pose estimation unit obtains the tracking result, estimates the poses of the athletes in each frame, obtains the key point information as the pose estimation result, and visualizes the pose estimation result on the display unit;

[0051] In this embodiment, the schematic diagram of the human body pose estimation unit is as Figure 3 shown. The human body pose estimation unit uses the Alphapose framework to estimate the poses of the athletes in each frame: after obtaining the detection boxes of each frame in the video, the Alphapose framework is used to estimate the poses of the athletes in each frame. The Alphapose framework generates the key point information of each frame, and the Kalman filter algorithm is used to process the key point information of each frame to obtain the pose estimation result.

[0052] (1) Single-frame pose estimation: The present invention uses the Alphapose framework to complete the pose estimation of the athletes in a single frame. The design purpose of the Alphapose framework is to handle two problems, the positioning error problem and the redundant detection problem. The positioning error problem is caused by the difference between the box given by the detector and the real box, that is, although the intersection over union (IoU>0.5) of the two satisfies the screening requirements, the detection box may only contain partial human body information, resulting in final false detection and missed detection; the redundant detection problem is caused by the repeated detection boxes generated by NMS. To handle these problems, A1phapose uses the regional multi-person pose estimation framework to complete the pose estimation under the condition of uncertain constraint box position.

[0053] (2) Kalman filter smoothing: After completing the single-frame pose estimation, it is necessary to smooth the key point information in time series. The present invention uses the Kalman filter algorithm to process the key point information generated by the Alphapose framework. The essence of the Kalman filter is a set of mathematical equations, which recursively estimate the state of the process, that is, minimize the mean of the root error. The Kalman filter can support the estimation of past, present and future states when the system accuracy is unknown. Its purpose is to estimate the state column vector x of the system, and generally estimate the state of the system through a difference equation containing random quantities:

[0054] x k = Ax k-1 + Bu k-1 + w k-1

[0055] where x k-1 is the state at the current moment, and x k is the state at the next moment. A is the transition matrix of size n×n, and B is the control matrix of size n×1. w k-1It is the noise in the state transition process. Since the observation of the system is not perfect and there will be some measurement noise, the observation equation is as follows:

[0056] z k = Hz k + v k

[0057] where H is an observation matrix of size m×n, which converts an n×1 state into an m×1 observation value and adds the deviation v in the observation process. k . Assume that the state transition process noise w and the measurement noise v both follow a normal distribution:

[0058]

[0059] where Q is called the process noise covariance matrix and R is called the measurement noise covariance matrix.

[0060] (3) Visualization of pose estimation results: This part is a component in the display module. The positions and confidence information of the key points are marked in the original video with different colors and different transparencies to provide reference for coaches and athletes.

[0061] (3) The action sequence feature extraction and enhancement unit takes the video, the tracking result, and the pose estimation result as inputs, and obtains video features after feature extraction, feature enhancement, and feature fusion.

[0062] In this embodiment, the schematic diagram of the action sequence feature extraction and enhancement unit is as shown in Figure 4 . The action sequence feature extraction and enhancement unit extracts features from the video to obtain the first feature, uses the pipeline self-attention mechanism to enhance the first feature by combining the first feature with the tracking result, obtains the enhanced first feature, extracts features from the pose estimation result to obtain the second feature, and fuses the second feature with the enhanced first feature through a fully connected layer to obtain video features.

[0063] (4) The score prediction unit takes the video features as inputs and performs overall action process quality assessment and stage behavior quality assessment.

[0064] (1) Overall action process quality assessment: The present invention can complete the overall assessment of the motion video at the global level and give the final action score. The present invention decouples the extraction and enhancement of video features from the score prediction stage, and transforms the features (i.e., the video features of the entire video) into 1024-d by adding a temporal average pooling layer between the two parts to achieve the overall action process quality assessment.

[0065] (2)Stagewise Behavior Quality Assessment: The present invention can evaluate the behavior quality of each video clip at the local level. After feature extraction and feature enhancement, each video clip becomes a 1024-d feature vector (i.e., the video feature of each video clip), and based on this feature vector, stagewise behavior quality assessment is realized.

[0066] (3)Behavior Quality Prediction Module: The present invention uses a multi-layer fully connected layer to complete the mapping from the feature vector to the behavior quality. The network layers adopted in this embodiment are: {FC(1024 - 512), ReLU}, {FC(512 - 128), ReLU}, {FC(128 - 1)}.

[0067] In this embodiment, the schematic diagram of the score prediction unit is as Figure 5 shown. The score prediction unit uses the I3D neural network. When performing full-action process quality assessment, all video features are subjected to temporal global average pooling and then sent to the fully connected layer of the I3D neural network to complete quality assessment. When performing stagewise behavior quality assessment, the video features of a video segment are sent to the fully connected layer of the I3D neural network to complete quality assessment.

[0068] In the action sequence feature extraction and enhancement unit and the score prediction unit, the action sequence feature extraction and enhancement unit uses the I3D neural network to extract features from the video to obtain the first feature, and uses the graph convolutional neural network to extract features from the pose estimation result to obtain the second feature.

[0069] When the score prediction unit performs full-action process quality assessment, all video features are subjected to temporal global average pooling and then sent to the fully connected layer to complete quality assessment. When performing stagewise behavior quality assessment, the video features of a video segment are sent to the fully connected layer to complete quality assessment.

[0070] (1) The I3D neural network is full name Two-Stream Inflated 3DConvNet. This network expands both the filters and pooling kernels in the 2D network to 3D, so the parameter initialization of the video network can be completed through a pre-trained model on an image dataset. I3D is extended from the Inception network, and its basic component is the Inception module. From the overall architecture, I3D consists of a convolutional layer, Inception modules, and a pooling layer. In ordinary video recognition tasks, the I3D network is regarded as a whole, and after the video is converted into a feature vector, it is used for classification tasks, regression tasks, etc. This invention divides the I3D network into two stages, and completes feature enhancement through the tracking result and pose estimation between the two stages.

[0071] In the first stage of the I3D network, feature extraction is performed. Assume that the input video contains a total of L frames. First, the Siammask tracker is used to track the athletes in the video to obtain the detection box information. In the video feature extraction stage, the video will be divided into N segments, and each segment contains M consecutive images. In this embodiment, N = 10 and M = 16. The video segments will be fed into the first stage of the I3D network to complete the feature extraction process and obtain features for subsequent feature enhancement.

[0072] (2) Feature enhancement: The current methods cannot complete effective and efficient feature enhancement. The limited receptive field of the convolutional operation makes it impossible to model long-term dependencies, and the characteristic of the RNN that it needs to store hidden states makes it impossible to perform efficient parallel computing. The present invention proposes a method of fusing the pipeline mechanism and the self-attention mechanism to complete the efficient enhancement of behavior features.

[0073] The pipeline self-attention module takes the detection box and the first feature as inputs for feature enhancement and finally generates the enhanced first feature. The pipeline self-attention mechanism does not change the size of the feature map, and this characteristic enables the pipeline self-attention mechanism to be embedded between any two layers of the network and can be stacked multiple layers.

[0074] (3) Feature fusion: The present invention fuses the first feature enhanced by the pipeline self-attention mechanism and the second feature (i.e., the pose feature) obtained by the graph convolutional network through feature connection to generate the fused feature X'. The feature X' will be fed into the second stage of the I3D to complete the subsequent feature extraction and finally generate H, where H represents the video feature and characterizes the behavior quality of the athlete.

[0075] Figure 6 is a schematic diagram of the pipeline self-attention mechanism, showing the quantization of the detection box and the mask generation process:

[0076] The specific process of using the pipeline self-attention mechanism to enhance the first feature by combining the first feature with the tracking result is as follows: Quantize (or discretize) and align the first feature with the detection box to generate a feature map mask, fuse the masks according to the ratio of the number of frames of the video and the number of the first features to generate a spatio-temporal pipeline, and use the pipeline self-attention mechanism inside the spatio-temporal pipeline to complete the sparse enhancement operation of the first feature. Fuse the first feature with the first feature after the sparse enhancement through a residual connection to obtain the enhanced first feature.

[0077] Specifically, after obtaining the tracking result and the first feature, determine the ratio N∶1 of the number of detection frames to the number of temporal dimension of the first feature according to the ratio of the number of frames of the video to the number of the first features, where N>1. Determine the mask corresponding to each detection frame, and obtain the ratio of the detection frame covering the feature network of the first feature. If the ratio is greater than a preset threshold, the first feature is selected; otherwise, the first feature is excluded. After completing the mask calculation of N detection frames, perform a bitwise AND operation to complete the fusion of the masks and generate a spatio-temporal pipeline.

[0078] The quantization of the bounding box and the mask generation process are shown in Figure 6 : After obtaining the tracking result and the first feature (which can also be understood as the feature map), it is necessary to screen the selected features in the feature map. Since there are two temporal pooling layers in the first stage of the I3D network, the ratio of the number of detection frames to the number of temporal dimensions of the feature map is not 1∶1, and since the detection frames generated by Siammask are skewed, it is impossible to directly complete the feature map screening. To address this problem, the present invention proposes a method for generating a feature map mask based on the discretization and aggregation of tracking frames for constructing a spatio-temporal pipeline. In this example, the ratio of the number of detection frames to the number of temporal dimensions of the feature map is 4∶1. It is necessary to first determine the mask corresponding to each detection frame. Then, according to the ratio of the feature grid covered by the detection frame and a preset threshold τ, determine whether the first feature at this position can be selected. If this ratio is greater than the threshold τ, the first feature at this position is selected; if it is less than, it is excluded. In this embodiment, the threshold τ = 0.5. After completing the mask calculation of the four detection frames, it is necessary to perform a bitwise AND operation to complete the aggregation of the masks to obtain the total mask. This mask contains all the positions of the selected first features. For more concise and clear representation, this embodiment converts it into a subscript set and then participates in subsequent operations.

[0079] Specifically, after completing the positioning of the selected first feature, that is, constructing the spatio-temporal pipeline, at this time, feature enhancement can be completed by introducing the self-attention mechanism. Similar to the Non-local module, the self-attention mechanism inside the spatio-temporal pipeline is expressed as:

[0080]

[0081] where p represents the output position to be calculated, (c, t, i, j) traverses all the positions of the first features in the spatio-temporal pipeline, the output feature y has the same size as the input feature x, the f function is a distance metric function, the g function is a feature mapping function, and the response value will be normalized by the normalization factor C(x)=∑ c ∑tΩ c,t for normalization.

[0082] To reduce the computational complexity, this embodiment uses a dot product operation as the similarity metric function;

[0083] To enhance the feature extraction ability of the subsequent I3D network, the present invention adds a residual connection to the pipeline self-attention module:

[0084] x′ p = W z y p + x p

[0085] where x′ p and x p have exactly the same dimensions, and W z is the parameter of the connection. Therefore, the pipeline self-attention mechanism can be embedded at any position in the network. To achieve a balance between high performance and high computational efficiency, in this embodiment, the pipeline self-attention module is placed after the Mixed_4e layer. Thus, T = 4, H = W = 14.

[0086] Compared with the Non-local module, when the pipeline self-attention module completes feature enhancement, it does not consider all features, but uses the self-attention mechanism to complete feature enhancement based on the spatio-temporal pipeline. This strategy greatly reduces the computational amount.

[0087] (5) Based on the quality assessment of the athlete's motion behavior in the video, the deficiencies of the athlete's motion behavior are obtained, and the athlete is subject to special training based on the deficiencies of the motion behavior.

[0088] This embodiment combines a human tracking unit, a human pose estimation unit, an action sequence feature extraction and enhancement unit, a score prediction unit, and a display unit to achieve multi-channel perception of the athlete's position and pose information, complete video feature extraction through the I3D network, and propose a pipeline self-attention mechanism to efficiently and effectively enhance the athlete's action features with the help of detection box information. The video features and the second features (i.e., the pose features obtained based on the pose estimation results, and the pose estimation results can be simply understood as the key points of the athlete's body) are fused through the feature fusion technology, and both global and local behavior quality assessment modes are designed in the score prediction unit. This system can efficiently complete the athlete behavior quality assessment at the full behavior level and each action stage level, save a large amount of manpower and material resources in the athlete training process, improve the training efficiency and quality, and provide a strong basic guarantee for the continuous improvement of the athlete's competitive level and the development of the sports cause.

[0089] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field according to the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A method for evaluating the quality of athletes' behaviors based on deep learning, characterized in that, the evaluation system it is based on includes a human body tracking unit, a human body pose estimation unit, an action sequence feature extraction and enhancement unit, a score prediction unit and a display unit. Specifically, the processes within each unit are as follows: The video is input into the human body tracking unit. The human body tracking unit performs object detection on each frame of the video to obtain the detection boxes of each frame as the tracking results, and visualizes the tracking results on the display unit; The human body pose estimation unit obtains the tracking results and estimates the poses of the athletes in each frame to obtain the key point information as the pose estimation results, and visualizes the pose estimation results on the display unit; The action sequence feature extraction and enhancement unit takes the video, the tracking results and the pose estimation results as inputs, and obtains video features after feature extraction, feature enhancement and feature fusion; The score prediction unit takes the video features as inputs and performs full action process quality evaluation and stage behavior quality evaluation; The human body tracking unit uses the YOLO detector and the Siammask framework for object detection: in the initial frame of the video, the YOLO detector is used for object detection to obtain the detection box of the initial frame, and the Siammask framework is used as a single object tracker to perform object detection on the consecutive frames after the initial frame to obtain the detection boxes of each frame after the initial frame; The human body pose estimation unit uses the Alphapose framework to estimate the poses of the athletes in each frame: after obtaining the detection boxes of each frame in the video, the Alphapose framework is used to estimate the poses of the athletes in each frame. The Alphapose framework generates the key point information of each frame, and the Kalman filter algorithm is used to process the key point information of each frame to obtain the pose estimation results; The action sequence feature extraction and enhancement unit extracts features from the video to obtain the first feature, combines the first feature with the tracking results and uses the pipeline self-attention mechanism for feature enhancement to obtain the enhanced first feature. The action sequence feature extraction and enhancement unit extracts features from the pose estimation results to obtain the second feature, and fuses the second feature with the enhanced first feature through a fully connected layer to obtain video features.

2. The method for evaluating the quality of athletes' behaviors based on deep learning according to claim 1, characterized in that, the action sequence feature extraction and enhancement unit uses the I3D neural network to extract features from the video to obtain the first feature, and uses the graph convolutional neural network to extract features from the pose estimation results to obtain the second feature.

3. The method for evaluating the quality of athletes' behaviors based on deep learning according to claim 2, characterized in that, When performing full action process quality evaluation, all video features are subjected to temporal global average pooling and then sent to a fully connected layer to complete the quality evaluation. When performing stage behavior quality evaluation, the video features of a segment of video are sent to a fully connected layer to complete the quality evaluation.

4. The method for evaluating the quality of athletes' behaviors based on deep learning according to claim 3, characterized in that, When the score prediction unit uses the I3D neural network to perform the quality assessment of the full action process, after performing temporal global average pooling on all video features, it sends them into the fully connected layer of the I3D neural network to complete the quality assessment. When performing the quality assessment of stage behaviors, it sends the video features of a segment of video into the fully connected layer of the I3D neural network to complete the quality assessment.

5. A method for evaluating the behavior quality of athletes based on deep learning according to claim 1, wherein, The specific method of using the pipeline self-attention mechanism to enhance the features by combining the first feature with the tracking result is as follows: Quantify and align the first feature with the detection box to generate a feature map mask, fuse the masks according to the ratio of the number of frames of the video to the number of the first features to generate a spatio-temporal pipeline, and use the pipeline self-attention mechanism inside the spatio-temporal pipeline to complete the sparse enhancement operation of the first feature. Fuse the first feature with the first feature after completing the sparse enhancement through residual connection to obtain the enhanced first feature.

6. A method for evaluating the behavior quality of athletes based on deep learning according to claim 5, wherein, After obtaining the tracking result and the first feature, determine the ratio N:1 of the number of detection boxes to the number of temporal dimensions of the first feature according to the ratio of the number of frames of the video to the number of the first features, N>1. Determine the mask corresponding to each detection box, and obtain the ratio of the detection box covering the feature network of the first feature. If the ratio is greater than the preset threshold, the first feature is selected; otherwise, the first feature is excluded. After completing the mask calculation of N detection boxes, complete the fusion of the masks through bitwise AND operation to generate a spatio-temporal pipeline.

7. A method for evaluating the behavior quality of athletes based on deep learning according to claim 5, wherein, The pipeline self-attention mechanism inside the spatio-temporal pipeline is expressed as: Among them, p represents the output position to be calculated. The (c, t, i, j) traverses all the first feature positions in the spatio-temporal pipeline. The output feature y and the input feature x have the same size. The f function is a distance metric function, and the g function is a feature mapping function. The response value will be normalized by the normalization factor C(x) = ∑ c ∑ t Ω c,t for normalization.

Citation Information

Patent Citations

  • Human body behavior recognition method and system based on multi-target tracking

    CN110399808A

  • Method and system for estimating and tracking human body posture in video

    CN113255429A