Video-based golf swing evaluation method and system

By employing a video-based golf swing evaluation method, utilizing a keyframe extraction network and a swing motion comparison and analysis module, the high cost of sensor-based evaluation is addressed, enabling efficient and low-cost swing motion analysis and evaluation.

CN116958859BActive Publication Date: 2025-11-28DONGHUA UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310671046.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-07
Publication Date
2025-11-28
Estimated Expiration
2043-06-07

AI Technical Summary

Technical Problem

Existing golf swing evaluation systems are expensive to obtain swing action analysis data through sensors, and have a high barrier to entry.

Method used

A video-based golf swing evaluation method is adopted, which utilizes a keyframe extraction network and a swing action comparison and analysis module. Image features are extracted through MobileNetV2 and multi-scale temporal MLPFormer network, and human pose estimation is performed in combination with the VideoPose3D algorithm to provide action comparison and standard score.

Benefits of technology

It reduces the cost of swing evaluation, improves the accuracy of critical event detection, can quickly and accurately locate motion problems, provides quantitative data analysis, and simplifies equipment requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116958859B_ABST
    Figure CN116958859B_ABST
Patent Text Reader

Abstract

One of the technical solutions of the present application is to provide a golf swing evaluation method based on video, which is characterized by including a key frame extraction network and a swing action comparison and analysis module. Another technical solution of the present application is to provide a golf swing evaluation system realized based on the above golf swing evaluation method, which is characterized by being divided into a presentation layer, a communication layer, a service layer and a data layer, and including a user management function module, a video management function module and an AI swing action comparison and analysis function module. The present application only needs a smart phone, compares and analyzes the swing action according to the swing key events, can more quickly and accurately locate the problem of the action, and can view the difference between the swing action and the professional player from multiple angles through 3D reconstruction of the skeleton model. At the same time, quantitative data analysis can be provided, such as the angle of some joints, and the distance between the skeletal points can be used as a standard score.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a video-based golf swing evaluation method and a golf swing evaluation system based on the method. BACKGROUND

[0002] Nomatics MySwing [1] discloses a golf swing evaluation system which provides 17 wireless sensor wearing nodes and club sensors. Its attached analysis software can analyze the motion trajectory, rotation angle, muscle force sequence and other data of the key parts of the golfer's body. At the same time, it is equipped with a pressure-sensitive foot pad to record the force of the foot kicking the ground. The product advantage lies in more accurate and comprehensive capture of actions, and more specific and professional quantitative data analysis. The disadvantage is that the cost of a set of sensor equipment is expensive, and because of the variety of professional sensor hardware devices, the player needs to have a good understanding of the system when using it, and the use threshold is high.

[0003] Liao et al. [2] provides an analysis tool to help golf beginners compare their swing actions with those of experts. The proposed application uses neural network-based encoder-extracted latent features to synchronize videos with different swing phase timings and detect key frames where inconsistent motions occur. They visualize the synchronized image frames and 3D poses to help users identify differences and key factors essential to improving swing skills.

[0004] As shown in Figure 1 Liao et al. [2] proposed a scheme to create a prototype application that can visualize the distance in the latent space of two input videos, detect difference frames using adaptive thresholds, and compare the detected frames with 3D human poses. They determine the action difference frames according to whether the distance of the two videos in the latent space exceeds the set threshold, which is likely to locate to some unimportant action moments that may be more related to personal habits and do not affect the final ball hitting effect, and is also likely to miss some unconfirmed action moments due to model accuracy problems.

[0005] McNally et al. [3] proposed SwingNet network to detect swing key events. The whole network structure of SwingNet is divided into two parts, first, a mature convolutional neural network is used to extract the features of each frame of the input video sequence. Since the focus of this network is on mobile deployment, it selects the lightweight convolutional neural network MobileNetV2 mentioned in the previous text for feature extraction. Then, considering that it is difficult to identify swing key events through a single frame, for example, the single frame images of the swing from top to bottom and from bottom to top to the same position are very similar, it tries to use the time information between video frames to help detection. SwingNet here uses a long short-term memory (LSTM) network to process the average-pooled output of MobileNetV2, and finally obtains class probabilities through a fully connected (FC) layer and a softmax activation function, where the weights of the fully connected layer are shared across frames. It is worth mentioning that SwingNet adds a background class in addition to the 8 key event classes, which is essentially a 9-class classification for each frame.

[0006] SwingNet correctly detects 8 golf swing key events with an average accuracy of 76.1%. Although it proposes a good swing key event detection baseline network, there is still much room for improvement in its performance. First, it is easy to notice a phenomenon in swing video data: the frames adjacent to the key event frames are very similar, but SwingNet does not pay special attention to these frames when performing classification, but directly uses cross-entropy as the loss function to identify adjacent frames as "other event" frames, which will make the model get completely different labels when facing similar pictures. This has a great impact on the model's learning, especially in slow-motion videos. In addition, the ability of LSTM to utilize temporal information is limited, so an updated and effective network needs to be applied to extract more temporal information between frames. At the same time, enhancing the image feature extraction ability of its backbone network can also greatly improve the performance. Improving the accuracy of key event detection is of great significance to subsequent swing analysis tasks.

[0007] With the development of computer vision and deep learning, many works provide some basis and inspiration for our method research, such as the target detection algorithm CenterNet[4] based on which the object center point is predicted to detect the target; the human body skeleton point coordinate estimation algorithm VideoPose3D[5] which can predict the 3D coordinates of 17 key skeleton points of the human body; some feature enhancement algorithms based on attention mechanism such as CBAM(Convolutional Block Attention Module)[6] which can provide attention enhancement from space and channel parts; and MetaFormer[7] proposed by Yu et al. which abstracts transformer into a general architecture and can be extended to more forms of information fusion. These works lay a certain foundation for the design and development of the method and evaluation system.

[0008] Reference:

[0009] [1] Nuoiten. Professional full-body motion capture golf swing evaluation and training system [EB / OL]. 2019. https: / / www.myswing.com.cn /

[0010] [2] Liao C C, Hwang D H, Koike H. How Can I Swing Like Pro?: Golf Swing Analysis Tool for Self Training[J]. 2021.

[0011] [3] McNally W, Vats K, Pinto T, et al. Golfdb: A video database for golf swing sequencing[C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition Workshops. 2019: 0-0.

[0012] [4] Zhou X, Wang D, P. Objects as points[J]. arXiv preprint arXiv:1904.07850, 2019.

[0013] [5] Pavllo D, Feichtenhofer C, Grangier D, et al. 3d human pose estimation in video with temporal convolutions and semi-supervised training [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019: 7753-7762.

[0014] [6] Woo S, Park J, Lee J Y, et al. Cbam: Convolutional block attention module [C] / / Proceedings of the European conference on computer vision (ECCV). 2018: 3-19.

[0015] [7] Yu W, Luo M, Zhou P, et al. Metaformer is actually what you need for vision [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022: 10819-10829. SUMMARY

[0016] The technical problem to be solved by the present application is that the cost of obtaining swing action analysis data through sensors is high.

[0017] In order to solve the above technical problems, one technical scheme of the present application provides a golf swing evaluation method and system based on video, characterized in that it comprises a key frame extraction network and a swing action comparison and analysis module, wherein:

[0018] The input of the key frame extraction network is a video frame sequence I t ∈ 3×H×W , wherein I tis the t-th frame image, T is the length of the sequence, H and W represent the height and width of each frame image respectively; the key frame extraction network processes each frame image in the video frame sequence I using a MobileNetV2 network to extract image features thereof; subsequently, the key frame extraction network performs global average pooling on the extracted image features, so that the information of each frame is represented by a vector, i.e. f t ∈R 1280 The key frame extraction network processes the image feature sequence f using a multi-scale time sequence MLPFormer to output embedded features Fc that fuse time sequence information; finally, the key frame extraction network uses a fully connected layer for classification for each frame to predict the event class e t For the video frame sequence I, the following is obtained e t ∈R C wherein C represents the number of event classes, wherein for each class, the frame with the highest predicted probability score is denoted as the corresponding key event;

[0019] The swing action comparison and analysis module further comprises an action comparison unit and a standard degree scoring unit:

[0020] The action comparison unit compares the swing key event action pictures, 3D skeleton models and joint angles of key body parts of ordinary players and professional players in the most intuitive form using the key frame pictures obtained by the key frame extraction network and the 2D and 3D human skeleton point coordinates obtained by the VideoPose3D[5] human pose estimation method;

[0021] The standard degree scoring unit provides a standard degree score with specific data display based on the difference presentation provided by the action comparison unit, and respectively calculates the distances of the skeleton points of each key event action of professional players and ordinary players.

[0022] Preferably, three CBAM[6] attention modules are added in the MobileNetV2 network: a CBAM module is added after the initial Conv2d, a CBAM module is added after the intermediate Bottleneck, and a CBAM module is added after the final Conv2d operation.

[0023] Preferably, the processing of the image feature sequence f by the multi-scale time sequence MLPFormer comprises the following steps:

[0024] The image feature sequence f obtains an embedded token sequence of the MLPFormer through an embedding layer;

[0025] The token sequence is input into a stacked B temporal MLPFormer block, each MLPFormer block first performs temporal MLP processing on the token sequence, and then inputs a feedforward layer and layer normalization added with a residual connection;

[0026] The outputs of each stage obtained by the B temporal MLPFormer blocks are linearly projected and upsampled, and the high semantic features of the last stage are added to the outputs of the previous stages to form semantic-rich features;

[0027] The outputs of different stages are spliced to obtain a video representation containing multi-scale information, and the video representation is classified using a fully connected layer.

[0028] Preferably, when training the key frame extraction network, the t-th frame label value of event class c is represented as , wherein t' c The key frame of event c is represented, and σ is an adjustable sequence length adaptive standard deviation. If the distance between the current frame and the event frame exceeds the threshold r, the label value is still 0, wherein the value of r is set as τ·δ, and δ is the error tolerance , wherein n is the number of frames between the key event "preparation action" and "hit", s is the sampling frequency, , which means the nearest integer near x, and τ is a multiplication coefficient.

[0029] Preferably, when training the key frame extraction network, the loss function is as follows:

[0030]

[0031] In the formula: , which represents the predicted probability of class c at frame t; m is a hard threshold, when the probability of negative samples is less than m, the negative samples are completely discarded; α and β represent adjustable exponential hyperparameters; N represents the number of key event frames contained in the input sequence.

[0032] Preferably, the action comparison unit uses the frame number of the key event obtained by the key frame extraction network to select the coordinate point at the corresponding time in the human body skeletal point coordinates of the entire swing video; then, according to the coordinate point, the skeletal models of professional players and ordinary players are generated in a canvas to form an intuitive comparison; in addition, the action comparison unit calculates the body joint angles of each key event using the coordinate point.

[0033] Preferably, the standard degree scoring unit simultaneously uses the difference in the length of the distance from the hip joint to the cervical vertebra of the ordinary player and the professional player as a scaling ratio to approximately normalize the skeletal points of the whole body to the same height; then the Euclidean distance of the 3D human skeletal point coordinates of the professional player and the ordinary player under each key event is calculated to obtain the standard degree score of each key event sub-action.

[0034] Another technical solution of the present application is to provide a golf swing evaluation system based on the above-mentioned golf swing evaluation method, characterized by being divided into a presentation layer, a communication layer, a service layer and a data layer, comprising a user management function module, a video management function module and an AI swing action comparative analysis function module, wherein:

[0035] In the presentation layer, a mobile phone interface of a WeChat style is designed by relying on the development platform of the WeChat applet, and the necessary video information for swing analysis is obtained by acquiring the camera and microphone of the user's smartphone;

[0036] In the communication layer, the network communication application programming interface of the WeChat applet is used to initiate a network request from the applet end to the server end in the form of hypertext transfer protocol to constitute the data communication between the front end and the back end.

[0037] In the service layer, information management, file uploading and method action analysis are realized, wherein the method action analysis is based on the above-mentioned golf swing evaluation method to generate human 3D skeletal key point coordinates corresponding to the swing event, and a series of subsequent quantitative comparative analysis and display are carried out accordingly;

[0038] In the data layer, the management and storage of data are realized through the WeChat cloud development function of the WeChat applet development platform;

[0039] The user management function module is responsible for the management of the account information and personal information of the system users, and controls the login and permissions of the users;

[0040] The video management function module realizes the function of the user using the smartphone to shoot and upload the video to the system for storage and display, and provides the function of the teacher user scoring and evaluating the swing video action, and the student user can also view the score of his own action;

[0041] The AI swing action comparative analysis function module is based on the above-mentioned golf swing evaluation device, divides the complete swing action into multiple key sub-actions according to the extracted key event frames, generates the human skeletal point coordinates at a specific moment through the 3D human pose estimation algorithm, carries out quantitative calculation and analysis based on the coordinates, and displays the swing sub-action at the current moment in 3D multi-angle view for the user to rotate and view.

[0042] Preferably, in the user management function module:

[0043] For student users, after logging into the system, the first step is to complete personal information, and then upload and view videos;

[0044] For teacher users, after logging into the system, they can choose to enter new students or view existing students. After entering new students, they can shoot the first video for the student, score it, and give comments. If they choose to view existing students, they can add new swing videos or view all existing videos for each student. In the video viewing page, teachers can modify the action score and comments of a video.

[0045] Preferably, the timing of the AI swing action comparison and analysis function module is:

[0046] Whether it is a teacher user or a student user, when viewing a certain swing video of a student, they can request AI analysis. The system listens to the user's request and jumps to the swing video analysis page. During the loading of the swing video analysis page, the WeChat applet downloads the current video file from the cloud database and sends a POST request to the server, passing the temporary address of the video to the background. After receiving the video file, the server judges whether the request method and file format are correct. If correct, the video is temporarily stored in the server for subsequent algorithm processing. Otherwise, an error prompt is returned. The server loads the algorithm model and its parameters based on the above golf swing evaluation device when it starts. After receiving the request, the video file temporarily stored in the local is preprocessed.

[0047] The video is divided into a frame of pictures and combined into a sequence of pictures of the same length using the OpenCV tool, and the pictures are scaled to the size required by the model input, converted to the Tensor form required by the algorithm model, and normalized to form a data sample. Then, through the Dataset and DataLoader classes in Pytorch, the video frame sequence is sequentially sent to the algorithm model for calculation, and the algorithm model outputs the sequence number of the frame where the key event occurs. Then, according to this sequence number, eight pictures are obtained from the original video. In addition, the algorithm model also uses the VideoPose3D algorithm to estimate the pose of the video. After obtaining the 2D human body skeleton point coordinates and 3D human body skeleton point coordinates of the current video, the background visualizes the 2D coordinates on the picture, and the 3D coordinates are temporarily stored on the server. The 3D coordinates of the "preparation action" event and the 3D coordinates of the default professional player in the expert library on the server are converted into base64 encoding and returned to the WeChat applet side in JSON format, and the rest of the pictures are temporarily stored. After receiving the picture code, the WeChat applet side decodes and renders it, displays it on the page, and displays the corresponding picture of the default professional player saved in the system in advance. At this point, the swing analysis page is loaded, and then the user can select different professional players and different swing key events in the WeChat applet. After confirming the selection, the applet side sends another POST request to obtain the video frame picture and the newly generated 3D comparison picture of the remaining events.

[0048] If the user issues a request to generate an angle, the WeChat applet side sends the currently selected key event and professional player information to the server. After receiving the request, the background loads the previously saved 3D human body skeleton point coordinates, and extracts the coordinate information of the student and the selected professional player corresponding to the event frame according to the information given by the front end. The specific body posture angle is returned to the WeChat applet side according to the coordinates.

[0049] If the user issues a request to generate a distance, the WeChat applet side sends the currently selected event and professional player information to the server. The backend calculates the distance between the student and the professional player based on the corresponding coordinate information, calculates the average distance as the final score of the overall swing using the Euclidean distance calculation method.

[0050] If the user issues a request to generate a distance, the WeChat applet side sends the currently selected event and professional player information to the server. The backend calculates the distance between the student and the professional player based on the corresponding coordinate information, calculates the average distance as the final score of the overall swing using the Euclidean distance calculation method.

[0051] The application is based on video acquisition of swing action analysis data, improves the swing key event detection precision, and compares the action deficiency according to the complete definition of 8 key events to prevent missed detection or detection of non-deterministic action difference. The required equipment of the application is only a smart phone, the swing action is analyzed according to the swing key event, the problem of the action can be positioned more quickly and accurately, the skeleton model is reconstructed in 3D, the difference between the swing action and the professional player can be viewed from multiple angles. At the same time, quantitative data analysis can be provided, for example, the angle of part of the joint, and the distance between the skeletal points of the professional player can be used as a standard degree score. BRIEF DESCRIPTION OF DRAWINGS

[0052] Figure 1 An overview of the method and visual analysis results proposed by Liao et al. [2] is illustrated;

[0053] Figure 2 The overall process of the improved golf swing key frame detection algorithm is illustrated

[0054] Figure 3 The architecture of MSTM is illustrated;

[0055] Figure 4 The overall architecture of the golf swing evaluation system is illustrated;

[0056] Figure 5 The system overall function module structure diagram is illustrated;

[0057] Figure 6 The activity diagram of the student using the video management function module is illustrated;

[0058] Figure 7 The activity diagram of the teacher using the video management function module is illustrated;

[0059] Figure 8 The AI swing action comparison and analysis function module flow chart is illustrated;

[0060] Figure 9 The AI swing action comparison and analysis function module timing diagram is illustrated. DETAILED DESCRIPTION

[0061] The application will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the application and not to limit the scope of the application. In addition, it should be understood that those skilled in the art can make various modifications or changes to the application after reading the content taught by the application, and these equivalent forms also fall within the scope defined by the appended claims of the application.

[0062] The application discloses a golf swing evaluation method based on video, which includes a key frame extraction network and a swing action comparison and analysis module.

[0063] The key frame extraction network includes three steps: (1) extracting frame-level image features from the video; (2) fusing the temporal information between the video frames; (3) classifying the event type in each frame. The overall process is as shown in Figure 2

[0064] Specifically, the input of the entire key frame extraction network is a video frame sequence I t ∈R 3×H×W , where I t is the t-th frame image, T is the length of the sequence, H and W represent the height and width of each frame image, respectively. Each frame image in the sequence is processed by the MobileNetV2 network to extract its image features. It should be noted that the model uses a simple and effective attention module CBAM, which is flexibly inserted into the convolutional neural network, and the specific details will be described later. Then, the extracted image features are globally averaged pooled, so that the information of each frame is represented by a vector, i.e. f t ∈R 1280 In the inter-frame information modeling part, the present application proposes a multi-scale time sequence MLPFormer to process the image feature sequence f, and outputs the embedded features Fc that fuse the temporal information. Finally, for each frame t, a fully connected layer is used for classification to predict the event class e t . For the entire sequence, we get e t ∈R C , where C represents the number of event classes. For each class, the frame with the highest predicted probability score is represented as the corresponding key event.

[0065] In order to promote mobile applications, SwingNet adopts the MobileNetV2 backbone, which is a lightweight convolutional neural network with an inverted residual structure. The present application integrates the CBAM module into the architecture of MobileNetV2 to improve its representation ability. CBAM is a simple but effective attention module that can be flexibly inserted into the CNN framework. The intermediate feature map can be adaptively refined through two concatenated sub-modules that affect the channel and spatial dimensions, respectively. The present application adds three CBAM modules in MobileNetV2 to enhance the feature representation under different states, which are added after the initial Conv2d, the middle Bottleneck, and the final Conv2d operation, respectively.

[0066] ​Due to the continuity of human motion in the golf swing process, efficient temporal modeling in videos is crucial. Yu et al. [7] proposed the concept of MetaFormer, a general architecture. It illustrates that the multi-head self-attention part of the Transformer can be replaced with any structure, and the components of this part are collectively called token mixer. Inspired by this, this paper uses MLP to process the inter-frame temporal information as a token mixer to process the video frame sequence, and this module is called MLPFormer. To obtain semantics at different temporal resolutions, the model uses a hierarchical structure with multiple stages. In each stage, the model first down-samples the intermediate features to pass through the temporal MLPFormer, and then up-samples the features to the same dimension as the first stage. Finally, the features of different stages are spliced together to obtain multi-scale temporal information, as shown in Figure 3

[0067] Specifically, the features f output by the network pooling layer of the feature extraction part first pass through the embedding layer to obtain the embedding tokens of the MLPFormer, which can be represented by the formula:

[0068] X = InputEmb(f) (1)

[0069] Then the token sequence X e R T×D with length T and channel dimension D is sent to a stack of B temporal MLPFormer blocks. Each MLPFormer block contains two parts: the main component of the first part is the temporal MLP, which can be represented by the formula:

[0070] X' = TemporalMLP(Norm(X T ) T + X (2)

[0071] where Norm(·) refers to layer normalization. By transforming the dimensions of the video frame sequence, the MLP can integrate temporal information between tokens. The second part is the same as the traditional Transformer block, which is a feedforward layer with residual connection and layer normalization. The so-called feedforward layer is to process the channel dimension of the features with MLP, which can be represented by the formula:

[0072] X'' = ChannelMLP(Norm(X')) + X' (3)

[0073] ​Meanwhile, the number of tokens is reduced and the feature dimension is increased by downsampling between stages (implemented by Conv1D) to make the model obtain time information at different scales. Each downsampling halves T and multiplies D by a gamma factor, where the value of gamma is set to 1.5. Then, referring to the method of Dai et al., the outputs of each stage X i are linearly projected and upsampled, and the high semantic features of the last stage are added to the outputs of the previous stages to form a semantic-rich feature F i . These operations can be expressed in formulas as follows:

[0074]

[0075] where W i is the learnable parameter in the linear layer, and Up(·) means that the output of each stage is first linearly projected and then upsampled to unify the output dimension to 64x512. When upsampling, the features are interpolated to the same dimension as the output of the first stage, but if the input dimension is already the same as the first stage, an identity mapping is performed. Finally, the outputs of different stages are spliced to obtain a video representation F c containing multi-scale information:

[0076] F c = Concat(F1, F2, …, F S ) (5)

[0077] After that, F c is sent to a fully connected layer for frame-level classification.

[0078] Considering that the adjacent frames of the key event frame have similar human poses, the method proposed by the present application applies a Gaussian kernel function to generate smooth label values. The inspiration for this idea comes from the method of generating a heat map in CenterNet [4] by Zhou et al. to predict object center points. In the heat map, the value of the center point is the largest, and it decreases along the radius outward according to the value of the Gaussian kernel function. Reducing this two-dimensional heat map generation method to one dimension can be well applied to time series. Specifically, unlike the traditional label, the key frame is 1 and the background frame is 0, the new label value applies a Gaussian function to form a one-dimensional heat map Y∈[0,1] T×1 . Then, the label value of the t-th frame of the event class c can be expressed as , where t' c represents the key frame of the event c, and σ is the adjustable sequence length adaptive standard deviation. If the distance between the current frame and the event frame exceeds the threshold r, the label value is still 0. In this way, frames that have a great similarity with the key event frame can be treated fairly.

[0079] Because different videos have different sequence lengths (especially slow motion videos), in order to set a proper r, the error tolerance defined in SwingNet is introduced here where n is the number of frames between the key event "preparation" and "hit", s is the sampling frequency, means the nearest integer around x. In this way, the value of r can be set as τ·δ, where τ is a multiplication factor. Finally, the heat maps of all key events are concatenated in parallel, i.e. Y ∈ [0,1] T×C , forming the true value of label smoothing.

[0080] For each key event class, there is only one correct key frame in the whole video sequence. This leads to a huge imbalance between positive and negative samples, which should be noted when training the model. Due to the scarcity of positive samples, it is necessary to emphasize the contribution of positive samples and weaken the influence of most negative samples. Therefore, the improved focal loss is adopted:

[0081]

[0082] The function makes the model perform binary classification on each c-class key event for each frame of the entire input sequence. Let denote the predicted probability of c-class at frame t, and the loss is calculated according to whether the label value Y tc is equal to 1. When the label value is 1, i.e. the frame is the actual key frame, the focal loss is used as before, focusing on difficult-to-classify frames using the term. When the label value is not 1, a hard threshold m is set, and when the probability of negative samples is less than m, the negative sample is completely discarded to filter out the gradient provided by too simple negative samples. And the function uses two calibration terms, first is (p m ) α term, where, This term aims to punish the wrong frames adjacent to the key frame, because the predicted probability value of the wrong frame should be small, if it is too large, it means that the prediction is wrong, at this time the term will further amplify this error, through the loss function to make the model pay attention to this sample. However, since the neighboring frames of the key frame are very similar to the correct frames, the information learned from the neighboring frames cannot be completely denied by the model, so (1-Y tc ) β is used to compensate for its punishment.

[0083] At the same time, inspired by the method of asymmetric loss (Asymmetric Loss, ASL), the coefficients of the penalty terms for positive and negative samples are decoupled, i.e. set α+ and α-, to reduce the contribution of easy negative samples and make the model pay more attention to positive samples.

[0084] Swing action comparison and analysis module:

[0085] People learn a new skill from imitation, so it is the most direct and effective way to contrast the action of ordinary players with the "standard template" of professional players' swings. Therefore, the swing action comparison and analysis module of the present application is aimed at swing key events, with the contrast of professional players' actions as the core, and is divided into two parts: action comparison and standard degree scoring.

[0086] Action comparison is to use the 8 key frame pictures obtained by the key event detection algorithm and the 2D and 3D human body skeleton point coordinates obtained by the VideoPose3D human body pose estimation algorithm, and then compare the swing key event action pictures, 3D skeleton models and joint angles of key body parts of ordinary players and professional players. In the most intuitive form, the gap is obvious at a glance.

[0087] Specifically, first, the frame numbers of the 8 key events obtained by the key frame recognition algorithm are used to select the coordinate points at the corresponding time in the human body skeleton point coordinates of the entire swing video. Then, according to the coordinate points, the skeleton models of professional players and ordinary players are generated in one canvas to form an intuitive comparison, and the position difference of body parts such as arms or lumbar vertebrae can be easily found. In addition, the coordinate points are used to calculate the joint angles of the body at each key event. For example, the angle between the arm and the body can reflect the incomplete swing of the arm; keeping the curvature angle of the spine unchanged is one of the most important factors to achieve stable hitting; during the backswing, keeping the angle of the right knee bent and placing the center of gravity of the body on the right knee will generate the maximum torsion for the backswing. According to the scalar product formula

[0088] , the following can be obtained:

[0089]

[0090] wherein, and represents the space vector obtained from the 3D human body skeleton point coordinates, and is the angle between the two vectors. At this time, 3 specific joint coordinate points are selected, and the relevant joint angles can be obtained through the inverse trigonometric function.

[0091] Standard degree scoring is to add a standard degree scoring with specific data display based on the difference presentation provided by the action comparison, and to calculate the distance of the skeleton points of professional players and ordinary players for each key event.

[0092] First, since the body size and height of ordinary players and professional players may be different, the swing data of multiple professional players are collected as an expert library for selection, and the difference in the length of the distance from the hip joint to the cervical vertebrae is used as a scaling ratio to approximately standardize the skeletal points of the whole body to the same height. Then, the Euclidean distance of the 3D human skeletal point coordinates of the professional players and the ordinary players under each key event is calculated to obtain the standard degree score of each key event sub-action. The distance can well reflect the standard degree of the swing of ordinary players, and the smaller the distance, the closer to the action of professional players.

[0093] The application also provides a golf swing evaluation system realized based on the above method, which takes the WeChat applet as a basic development framework, cooperates with a server background in which a deep learning algorithm model is deployed, and provides convenient and simple user experience. The overall architecture of the system is as shown in Figure 4

[0094] According to Figure 4 It can be seen that the entire system is divided into four levels: presentation layer, communication layer, service layer and data layer.

[0095] In the presentation layer, the system relies on the development platform of the WeChat applet to design a WeChat-like mobile interface, which can be conveniently obtained within WeChat and brings users the original APP experience. And by obtaining the camera and microphone of the user's smartphone and other devices, the necessary video information for swing analysis is obtained.

[0096] In the communication layer, through the network communication application programming interface (API) of the WeChat applet, a network request is initiated from the applet end to the server end in the form of Hyper Text Transfer Protocol (HTTP) to constitute the data communication between the front and back ends.

[0097] In the service layer, the system realizes the backend service through the Flask Web micro-framework of Python. There are three business services in total: information management, file upload and algorithm action analysis. Among them, the information management part mainly provides editing and management of user accounts, personal information and video information; the file upload part provides the function of receiving the video uploaded by the mobile phone from the applet end; the algorithm action analysis part is the core business of the system, that is, through the key event frame recognition algorithm and human pose estimation algorithm, the 3D skeletal key point coordinates corresponding to the swing event are generated, and a series of quantitative comparison and analysis and display are carried out based on this.

[0098] ​At the data layer, the management and storage of data are realized through the WeChat cloud development function of the WeChat applet development platform. This function is a development service jointly launched by the WeChat team and Tencent Cloud. Its advantage is that it does not need to specially build a server and can easily interact with the database provided by the applet itself through the API. This function can add, delete, modify and query ordinary user data, and also store image and video file data.

[0099] According to the analysis of the functional requirements of the system, the golf full swing evaluation system is divided into three functional modules: user management, video management and AI swing action comparison and analysis. The overall functional module structure of the system is shown in Figure 5 .

[0100] Among them, the user management module is mainly responsible for the management of the account information and personal information of the system users, and controls the login and permissions of the users; the video management module realizes the function of users using mobile phones to shoot and upload videos to the system for storage and display, and provides the function of teachers scoring and evaluating swing video actions, and students can also check their action scores; the AI swing action comparison and analysis module can divide the complete swing action into multiple key sub-actions according to the extracted key event frames, generate human body skeletal point coordinates at a specific time through 3D human body posture estimation algorithm, and then perform quantitative calculation and analysis of joint angles and skeletal point distances through these coordinates. Finally, the current swing sub-action is displayed in 3D multi-angle view for users to rotate and view.

[0101] The user management function is an indispensable part of every system, which sets up information barriers between different users, controls the identity of each user when using the system, and the access permission of data, etc. In this module, users can register, log in, and add or modify personal information, etc. When registering, each account is unique, and different users of different identities have different interfaces after logging in.

[0102] In order to enable users to shoot and upload, and view all previous swing videos and analysis records at any time, the system designs a swing video management function module. This module has different use permissions for users of different identities.

[0103] For student users, after logging into the system, the first thing to do is to complete personal information, and then upload and view the video. This is to make the student's personal video correspond to the teacher's list, and also to facilitate retrieval. If you do not complete the information directly by clicking on the record video or view the video, you will be prompted to add personal information first. After completing personal information, click on Add Video to jump to the video shooting upload interface, click on Record to call the phone camera to shoot the swing video. Of course, you can also choose to upload the video from the phone album. After the video is uploaded successfully, click on Submit to store the video in the database according to the student's personal information. Students can also view all swing videos uploaded by themselves or teachers, and see the scores and comments given by teachers. The activity diagram of the student using the video management function module is shown in Figure 6 .

[0104] For teacher users, after logging into the system, you can choose to enter new students or view existing students. Click on Enter New Student, then enter the student's ID, class, name and gender, and click on Confirm to add. Enter the video upload interface. At this time, the teacher can shoot the first video for the student, and can also score and give comments. If you choose to view existing students, click on the student class, and the system will display all existing student information in the database according to the class. Click on a class to list all students in the class. For each student, you can add a new swing video or view all existing videos. In the video viewing page, the teacher can modify the action score and comments of a video, especially the record score and comments of the student's personal upload. The activity diagram of the teacher using the video management function module is shown in Figure 7 .

[0105] The AI swing action comparison and analysis function module is the core function module of the system, and will not be distinguished according to user identity. Teachers can use the analysis results as a reference for scoring and comments, or as a basis for subsequent classroom teaching and training plan changes. Students can use the analysis results to obtain feedback on their swing training at any time, helping them quickly find weak links in their actions for targeted training.

[0106] The function enters from the swing video viewing interface, at this time the server will algorithmically process the video being viewed. First, the 8 key event frames in the current video and the selected professional player demonstration video are located according to the key event detection algorithm proposed in the application, and then the human body pose estimation algorithm is used to generate the human body skeleton point coordinates of each key event. The 2D coordinates are visualized on the picture corresponding to each key event, presented at the top of the interface, with the professional player on the left and the student on the right, forming a sharp contrast. The 3D coordinates are used to generate a multi-view 3D model display of the current key event action. The actions of the professional player and the student are generated on one canvas and presented in red and blue respectively, and four buttons are set for user view rotation operations. In addition, the angle between the arm and the body, the spine bending angle and the right knee bending angle are calculated according to the 3D coordinates. Finally, the distances of the 17 skeleton points of the professional player and the student at the 8 key events are calculated.

[0107] In addition, when the page is loaded, the system will default to selecting a male professional player and a preparation action key event as the initial display of the page. Then, as the user selects a professional player in the "expert library" and switches the key events, all the above-mentioned displays, quantitative calculation analyses, etc. will be switched. The flow chart of the whole AI swing action comparison and analysis function module is as shown in Figure 8

[0108] The core function module of the system is AI swing action comparison and analysis. The two function modules of user management and video management can be realized by simple logical judgment and database operation, and will not be described here.

[0109] ​Whether it is a teacher user or a student user, when viewing a certain swing video of a student, the AI analysis button can be seen in the "action standard degree score" column. The system jumps to the swing video analysis page by listening to the user click event. During the loading of this page, the WeChat applet downloads the current video file from the cloud database and sends a POST request to the server, passing the temporary address of the video to the background. After the server receives the video file, it judges whether the request method and file format are correct. If correct, the video is temporarily stored in the server for subsequent algorithm processing; otherwise, an error prompt is returned. The server loads the algorithm model and its parameters when it starts, and does not need to load in real time when the video is transmitted. After receiving the request, the server temporarily stores the video file in the local data preprocessing: using OpenCV tools to divide the video into a frame of pictures and combine them into multiple groups of the same length of picture sequence. At the same time, the picture is scaled to the size required by the model input, and then converted into the Tensor form required by the deep learning model and normalized to form a data sample. Then through the Dataset and DataLoader classes in Pytorch, the video frame sequence is sequentially sent into the model for calculation. The model outputs the sequence number of the 8 key events, and then obtains the eight pictures from the original video according to the sequence number. In addition, the VideoPose3D algorithm is also used to estimate the posture of the video. After obtaining the 2D and 3D human body skeleton point coordinates of the current video, the background visualizes the 2D coordinates on the 8 pictures, and the 3D coordinates are temporarily stored in the server. Then, using the matplotlib package, the 3D coordinates of the "ready to act" event and the 3D coordinates of the default professional player in the expert library on the server are made into a 3D comparison image on the same canvas, and the corresponding picture of the event is selected from the 8 original video frame pictures generated before, and together converted into base64 encoding, returned to the WeChat applet in the form of JSON, and the rest of the pictures are temporarily stored. After the applet receives the picture code, it is decoded and rendered, displayed on the page, and the corresponding picture of the default professional player saved in the system in advance is also displayed. At this point, the swing analysis page is loaded. Then the user can select different professional players and different swing key events in the applet, and after determining the selection, the applet will send another POST request to obtain the video frame picture and the newly generated 3D comparison picture of the rest of the events.

[0110] After that, if the user clicks the "generate angle" button, the applet side will send the currently selected key event and professional player information to the server. After receiving the request, the background loads the previously saved 3D human skeleton point coordinates, and extracts the coordinate information of the student and the selected professional player corresponding to the event frame according to the information given by the front end. According to the coordinates, the specific body posture angle is returned to the applet side. If the user clicks the "generate distance" button, the same information is sent to the server, and the backend calculates the distance according to the corresponding coordinate information. Using the calculation method of Euclidean distance, the distance between the student and the professional player's skeleton point coordinates under the 8 key event frames is calculated, and the average is taken as the final score of the overall swing. This distance can well reflect the standard degree of the student's swing. The smaller the distance, the closer the student's action to the professional player's action, and the higher the score obtained.

[0111] Finally, if the user clicks the rotation button while viewing the 3D swing action diagram, the applet records the number of clicks of the up, down, left, and right buttons to obtain the current user's viewing angle of the 3D diagram, and then requests the server to display the picture under the user's required viewing angle. The timing diagram of the AI swing action comparison and analysis function module is shown in Figure 9 .

Claims

1. A video-based golf swing evaluation method, characterized by, The swing action comparison and analysis module comprises a key frame extraction network and a swing action comparison and analysis module, wherein: An input of the key frame extraction network is a video frame sequence I t ∈R 3×H×W , where I t is the t-th frame image, T is the length of the sequence, H and W represent the height and width of each frame image respectively; the key frame extraction network processes each frame image in the video frame sequence I using a MobileNetV2 network to extract image features thereof; subsequently, the key frame extraction network performs global average pooling on the extracted image features, so that the information of each frame is represented by a vector, i.e. f t ∈R 1280 ; the key frame extraction network processes the image feature sequence f using a multi-scale time sequence MLPFormer to output embedded features Fc that fuse time sequence information; finally, the key frame extraction network uses a fully connected layer for classification for each frame to predict an event class e t , for the video frame sequence I, obtain e t ∈R C , where C represents the number of event classes, wherein for each class, the frame with the highest predicted probability score is represented as the corresponding key event; The swing action comparison and analysis module further comprises an action comparison unit and a standard degree scoring unit: The action comparison unit compares the swing key event action pictures, 3D skeleton models and joint angles of the body key parts of ordinary players and professional players in the most intuitive form by using the key frame pictures obtained by the key frame extraction network and the 2D and 3D human body skeleton point coordinates obtained by the VideoPose3D human body pose estimation algorithm. The standard degree scoring unit provides standard degree scores with specific data display on the basis of the difference presentation provided by the action comparison unit, and calculates the distances of the skeleton points of each key event action of professional players and ordinary players. The processing of the image feature sequence f by the multi-scale time sequence MLPFormer comprises the following steps: The image feature sequence f obtains an embedded token sequence of the MLPFormer through an embedding layer; The token sequence is input into a stacked B time sequence MLPFormer block, each MLPFormer block performs time sequence MLP processing on the token sequence, and then inputs a feedforward layer and layer normalization with residual connection added; The outputs of each stage obtained by the B time sequence MLPFormer blocks are linearly projected and upsampled, and the high semantic features of the last stage are added to the outputs of the previous stages to form semantic-rich features; The outputs of different stages are spliced to obtain a video representation containing multi-scale information, and the video representation is classified by using a fully connected layer; When training the key frame extraction network, the t-th frame label value of event class c is represented as wherein t c ' represents the key frame of event c, σ is an adjustable sequence length adaptive standard deviation, and the label value is still 0 if the distance between the current frame and the event frame exceeds the threshold value r, wherein the value of r is set as τ·δ, and δ is an error tolerance wherein n is the number of frames between the key event "preparation action" and "ball hitting", s is the sampling frequency, which refers to the nearest integer near x, and τ is a multiplication coefficient.

2. The video-based golf swing evaluation method of claim 1, wherein, Three CBAM modules are added in the MobileNetV2 network: a CBAM module is added after the initial Conv2d, a CBAM module is added after the intermediate Bottleneck, and a CBAM module is added after the final Conv2d operation.

3. The video-based golf swing evaluation method of claim 1, wherein, When the key frame extraction network is trained, the loss function is as follows: In the formula: The prediction probability of class c at frame t is represented as c; m is a hard threshold, when the probability of negative samples is less than m, the negative samples are completely discarded; Alpha and beta represent adjustable exponential hyperparameters; N represents the number of key event frames contained in the input sequence.

4. The video-based golf swing evaluation method of claim 1, wherein, The action comparison unit selects the coordinate points of the corresponding time from the human body skeleton point coordinates of the entire swing video by using the frame number of the key event obtained by the key frame extraction network, and then generates the skeleton models of professional players and ordinary players in a canvas according to the coordinate points to form an intuitive comparison. In addition, the action comparison unit calculates the body joint angles of each key event by using the coordinate points.

5. The video-based golf swing evaluation method of claim 1, wherein, The standard degree scoring unit simultaneously uses the difference degree of the distance length from the hip joint to the cervical vertebra of ordinary players and professional players as a scaling ratio to approximately standardize the skeleton points of the whole body to the same height; then, the Euclidean distance of the 3D human body skeleton point coordinates of professional players and ordinary players at each key event is calculated to obtain the sub-action standard degree score of each key event.

6. A golf swing evaluation system implemented based on the golf swing evaluation method of claim 1, characterized by, The system is divided into a display layer, a communication layer, a service layer and a data layer, and comprises a user management function module, a video management function module and an AI swing action comparison and analysis function module, wherein: At the presentation layer, a WeChat-like mobile interface is designed by relying on the development platform of the WeChat applet, and the necessary video information for swing analysis is obtained by acquiring the camera and microphone of the user's smartphone; At the communication layer, the network communication application programming interface of the WeChat applet is used to initiate a network request from the applet end to the server end in the form of hypertext transfer protocol, so as to constitute data communication between the front end and the back end; At the service layer, information management, file uploading, and algorithm action analysis are implemented, wherein the algorithm action analysis is based on the golf swing evaluation method of claim 1 to generate human 3D skeleton key point coordinates corresponding to the swing event, and a series of subsequent quantitative comparison and analysis and display are performed accordingly; At the data layer, the management and storage of data are realized through the WeChat cloud development function of the WeChat applet development platform; The user management function module is responsible for the management of user account information and personal information, and controls the login and permissions of users; The video management function module realizes the functions of users using smartphones to shoot videos and upload them to the system for storage and display, as well as the functions of teacher users scoring and evaluating swing video actions, and student users being able to view the scores of their own actions; The AI swing action comparison and analysis function module is based on the golf swing evaluation method of claim 1, divides the complete swing action into multiple key sub-actions according to the extracted key event frames, generates human skeleton point coordinates through a 3D human pose estimation algorithm, and performs quantitative calculation and analysis based on the coordinates, and displays the current swing sub-action in 3D multi-angle view for users to rotate and view.

7. A golf swing evaluation system as claimed in claim 6, wherein, In the user management function module: For student users, after logging into the system, they first need to complete their personal information, and then can upload and view videos; For teacher users, after logging into the system, they can choose to enter new students or view existing students, and after entering a new student, they can shoot the first video for the student and score and give comments on it; If they choose to view existing students, they can add new swing videos for each student or view all existing videos, and on the video viewing page, they can modify the action score and comments of a certain video.

8. A golf swing evaluation system as claimed in claim 7, wherein, The timing of the AI swing action comparison and analysis function module is as follows: Both teacher users and student users can request AI analysis when viewing a certain swing video of a student, and the system jumps to the swing video analysis page by listening to the user's request; during the loading of the swing video analysis page, the WeChat applet end downloads the current video file from the cloud database and sends a POST request to the server end, and the temporary address of the video is transmitted to the background; after receiving the video file, the server end judges whether the request method and file format are correct: if correct, the video is temporarily stored in the server for subsequent algorithm processing; Otherwise, an error prompt is returned; the server loads the algorithm model and its parameters of the golf swing evaluation method at startup, and after receiving the request, the video file temporarily stored in the local is preprocessed: The video is divided into a frame of pictures and combined into a sequence of pictures of the same length using the OpenCV tool, and the pictures are scaled to the size required by the model input, converted to the Tensor form required by the algorithm model, and normalized to form a data sample. Then, through the Dataset and DataLoader classes in Pytorch, the video frame sequence is sent into the algorithm model for calculation, and the algorithm model outputs the sequence number of the frame where the key event occurs. Then, according to this sequence number, eight pictures are obtained from the original video. In addition, the algorithm model also uses the VideoPose3D algorithm to estimate the pose of the video. After obtaining the 2D human body skeleton point coordinates and 3D human body skeleton point coordinates of the current video, the 2D coordinates are visualized on the picture, and the 3D coordinates are temporarily stored on the server. The 3D coordinates of the "ready to move" event and the 3D coordinates of the default professional player in the expert library on the server are converted into a 3D comparison image using the matplotlib package, and the corresponding picture of the event is selected from the original video frame picture generated earlier. Together, they are converted into base64 encoding and returned to the WeChat applet side in JSON format. The rest of the pictures are temporarily stored. After receiving the picture encoding, the WeChat applet side decodes and renders it, and displays it on the page. At the same time, the corresponding picture of the default professional player saved in the system in advance is displayed. At this point, the swing analysis page is loaded, and then the user can select different professional players and different swing key events in the WeChat applet. After confirming the selection, the applet side will send another POST request to obtain the remaining event video frame picture and the newly generated 3D comparison picture; If the user issues a request to generate an angle, the WeChat applet side sends the current selected key event and professional player information to the server side. After receiving the request, the background loads the previously saved 3D human body skeleton point coordinates, and extracts the coordinate information of the student and the selected professional player corresponding to the event frame according to the information given by the front end. The specific body posture angle is returned to the WeChat applet side according to the coordinates; If the user issues a request to generate a distance, the WeChat applet side sends the current user-selected event and professional player information to the server side. The backend calculates the distance between the student and the professional player's skeleton point coordinates at the key event frame using the Euclidean distance calculation method, and calculates the average as the final score of the overall swing; If the user issues a request to generate a distance, the WeChat applet side sends the current user-selected event and professional player information to the server side. The backend calculates the distance between the student and the professional player's skeleton point coordinates at the key event frame using the Euclidean distance calculation method, and calculates the average as the final score of the overall swing; If the user issues a request to generate a distance, the WeChat applet side sends the current user-selected event and professional player information to the server side. The backend calculates the distance between the student and the professional player's skeleton point coordinates at the key event frame using the Euclidean distance calculation method, and calculates the average as the final score of the overall swing;