Method for extracting tennis catching and playing skills from sports video
By adopting physics-based simulation correction in tennis skills learning, and combining deep learning and hybrid control strategies, the problem of low exercise quality caused by cumulative errors is solved, and a more realistic and accurate tennis skills learning is achieved.
Patent Information
- Application Number
- CN202510146881.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, when learning tennis players' catch and hitting movements, cumulative errors are prone to problems such as low sports quality.
The physics-based simulation correction is used to estimate motion, combined with deep learning and hybrid control strategies, to predict and correct motion embed errors, thereby improving motion quality.
Through the combination of physics simulation and deep learning, tennis skills can be learned more realistically and accurately, and adapted to different game scenarios and opponents.
Smart Images

Figure FT_1 
Figure FT_2 
Figure QLYQS_1
Abstract
Description
Technical Field
[0001] The present invention relates to a method for extracting tennis receiving and hitting skills from sports videos, belonging to the technical field of the combination of computer vision and artificial intelligence. Background Art
[0002] Tennis is a global sport, deeply loved by global enthusiasts for its excellent competitiveness, entertainment value and fitness value. With the rapid development of artificial intelligence technology, combined with broadcast video and physical simulation technology, the learning of tennis skills has ushered in a new way, opening up new possibilities for the popularization and promotion of this sport.
[0003] The multi-modal variational auto-encoder (MVAE: Multi-Modal Variational Auto-Encoder) is a deep learning model framework based on the variational auto-encoder (VAE: Multi-Modal Variational Auto-Encoder), which is good at processing multi-modal data. The VAE is a generative model, consisting of an encoder and a decoder. The encoder maps the input data to a latent space, which usually has a lower dimension and serves as a compressed representation of the data. The decoder then reconstructs the input data from the latent space and learns the distribution of the data by minimizing the reconstruction error. The MVAE aims to process data of multiple modalities simultaneously, and these data of different modalities usually have different feature representations and statistical characteristics. By learning the joint distribution of multi-modal data, the MVAE can effectively fuse information of different modalities and extract more comprehensive and representative feature representations.
[0004] The character motion controller can achieve more realistic, personalized, and interactive motion simulations by combining physical simulation, animation technology, and artificial intelligence algorithms. (1) The character motion controller can utilize physical simulation technology to establish physical models of athletes and tennis balls, accurately simulating the mechanical principles and physical laws during the motion process. For example, by considering factors such as gravity, friction, and air resistance, it can achieve a more realistic tennis trajectory and the athlete's ball-catching action. (2) Animation technology can provide more vivid and lifelike visual effects for the character motion controller. By using techniques such as keyframe animation, motion capture, and physical simulation animation, it can achieve various complex actions of the athlete, such as hitting the ball, running, and jumping. (3) Artificial intelligence algorithms can provide more intelligent motion control strategies for the character motion controller. For example, by using machine learning algorithms, it can automatically adjust the controller's parameters based on the athlete's motion data and the game scenario, simulate different difficulty levels of game scenarios and training tasks to meet the needs of tennis learners at different levels, and achieve more personalized motion simulations. The character motion controller can provide personalized training programs according to the learner's physical characteristics, skill level, and motion style. Generally speaking, when an athlete is undergoing tennis training, the character motion controller can visually present the athlete's actions and the game scenario, providing a more intuitive analysis and evaluation tool for learners and coaches. By analyzing the athlete's motion data and game videos, problems and deficiencies can be discovered, and corresponding improvement measures can be formulated.
[0005] In summary, based on the multimodal variational autoencoder, combined with physical simulation and character control, a method for extracting tennis receiving and hitting skills from motion videos is proposed. This method can effectively learn the complex tennis skills in the video. The provided method is of great significance for facilitating the daily training of tennis players. Summary of the Invention
[0006] The present invention proposes a method for extracting tennis receiving and hitting skills from motion videos. It uses motion videos to learn the ball-catching and hitting actions of tennis players, and uses a motion controller to control the character to complete the tasks of tennis ball-catching and hitting. To solve the motion quality problem caused by the cumulative error in learning motion embeddings, the patent uses physics-based simulation to correct the estimated motion and uses a hybrid control strategy to correct the errors in the predicted motion embeddings.
[0007] Controlling the imitation of the reference motion of the character is an important step in learning tennis skills. The process of simulating and controlling the character to imitate the reference motion is represented as a Markov decision process, which is defined by a tuple consisting of state S, action A, transition dynamics T, reward function r, and discount factor Υ as follows:
[0008] MDP = (S, A, T, r, Υ)
[0009] Initialize the state s of the simulated character 0 to be the same as the initial state of the reference motion; according to the policy π(a t |s t ) at each stage s t perform iterative sampling on a t ; according to T(s t+1 |s t , a t ) transition to the next stage s t+1 and output a corresponding reward function r t , as shown in the process Figure 1 .
[0010] Generally speaking, when facing videos of a large number of tennis match instances, the patent can learn complex tennis receiving and hitting skills, and use simple rewards and annotated hitting types to realistically concatenate multiple hitting actions to form a long - term tennis motion cycle.
[0011] The general process of the present invention is as shown in Figure 1 , and the main steps are as follows:
[0012] Step 1: Estimate the global foot trajectory of the tennis player's hitting and playing postures, and construct a tennis motion video dataset;
[0013] Step 2: Train an initial imitation policy to simulate the tennis player's motion, control the behavior of the simulated character, and generate an initial motion dataset;
[0014] Step 3: Fit the conditional VAE to the initial motion dataset to obtain a low - dimensional motion embedding and generate athlete - like motions;
[0015] Step 4: Combine the body motion output of the motion embedding and the predicted correction of the character's wrist motion to train a motion planning strategy to guide the tennis receiving and hitting motions of the athlete in the imitation video.
[0016] Furthermore, the specific steps of step 1 are as follows:
[0017] Step 1.1: Use the Yolov4 network model to track the motions of the players on both sides of the tennis video sports field, obtain the point boundaries and the player bounding boxes, use the vision transformer with ViT as the core to extract the key points of the tennis player's postures, and use the inverse dynamics method to estimate the postures of the tennis player hitting the tennis ball; the tennis player bounding box is used to define the athlete's motion area and is input into the pose estimation module;
[0018] Step 1.2: Since the inverse dynamics method outputs the position and orientation of the tennis player's heel in the camera coordinates, it is necessary to convert them into the court coordinates in the global coordinate system (the origin is located at the center of the tennis court); detect the boundary lines and their intersections of the tennis court, and solve the camera matrix using the multi-view perspective transformation algorithm; calculate the position of the player's heel (the center of the two ankle key points), and convert the position into the tennis court coordinates using the inverse projection transformation; calculate the global heel coordinates using the camera transformation to minimize the projection error between the 2D key points of the player's heel and the 3D joint positions, and correct the heel trajectory to obtain the motion dataset;
[0019] Step 1.3: In the modeling stage of tennis movement, mark the frames containing tennis players hitting the ball in different styles, and mark the identities of the tennis players in the frame sequence;
[0020] Furthermore, the specific steps of Step 2 are as follows:
[0021] Step 2.1: Use the following features to represent the character state
[0022]
[0023] where p t is the heel coordinate of the tennis player, represents the linear velocity of the heel coordinate of the tennis player, q t is the joint rotation, represents the angular velocity of joint rotation, represents the heel position of the target, represents the angular velocity of the target joint rotation.
[0024] Step 2.2: Calculate the body torque of the tennis player using a proportional derivative (SD) controller at each non-heel joint position. Action a t specifies the target joint angle u t of the SD controller, and the joint torque T t is calculated by the formula: where k p and k d represent the stiffness and damping of the joint, and represent the rotation and angular velocity of the non-heel joint nr;
[0025] Step 2.3: Include joint rotation r t o , velocity r i υ , position r t p , key point r t k and penalty term r t e : rt = ω o r t o + ω υ r t v + ω p r t p + ω k r t k + ω e r t e , where ω 0 , ω v , ω p , ω k , ω e are the weights of the corresponding terms respectively:
[0026] Step 2.4: The training is carried out in two stages. In the first stage, the motion strategy is trained using the human motion capture and animation database so that the simulated athlete learns to imitate general actions; the penalty term r t e is used to weaken the influence of the difference between M kin frames; the corrected tennis motion data M corr is output.
[0027] Furthermore, the specific steps of step 3 are as follows:
[0028] Step 3.1: Use MVAE to predict the motion phase of the cyclic playing posture of the tennis player from receiving the ball to continuous hitting at a fixed position on the court, and use the cyclic phase variable θ to represent the motion phase of each frame; when θ = π, the athlete hits the ball; when θ = 0 or θ = 2π, the athlete resumes.
[0029] Furthermore, the specific steps of step 4 are as follows:
[0030] Step 4.1: In order to overcome the inaccuracy of the simulated motion data, the present invention adopts a hybrid control method, in which the whole body motion is controlled by the reference trajectory generated by MVAE, while the wrist motion is controlled by the motion strategy. Each action of the simulated athlete state includes generating the target posture of the next frame by the latent code of MVAE, and the joint correction of the swinging arm.
[0031] Step 4.2: Use the reward function to control the strategy of the simulated motion so that the simulated athlete can hit the incoming tennis ball and bounce at the desired position (the bouncing position of the ball) and the target rotation direction on the court;
[0032] Before contacting the tennis ball, apply the racket position reward r t r , to minimize the center of the racket head when the person hits the ball Distance to the ball position Between:
[0033]
[0034] After contacting the tennis ball, apply the ball position reward r t b To minimize the estimated bounce position of the ball And the target bounce position Between, while ensuring that the rotation direction of the ball is consistent with the target rotation direction:
[0035]
[0036] Wherein, where s b And Are binary variables, representing the rotation directions of the simulated ball and the target ball respectively. When the value is 1, it represents topspin (the ball rotates forward), and when the value is 0, it represents backspin (the ball rotates backward).
[0037] The beneficial effects of the present invention are as follows: Physical simulation provides physical modeling of tennis motion, including the trajectory of the tennis ball, the actions of the athlete, and the mechanics of hitting and receiving the ball. Deep learning extracts the action strategies and skills of the athlete from the motion video. Combining physical simulation and deep learning makes the learned tennis skills more realistic, accurate, and adaptable to different game scenarios and opponents. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Is the flowchart of the present invention.
[0039] Figure 2 Is the Markov decision process diagram. DETAILED DESCRIPTION OF THE INVENTION
[0040] The present invention provides a method for extracting tennis receiving and hitting skills from a motion video. The following further describes the present invention in detail with reference to the accompanying drawings and specific implementation methods.
[0041] The training steps of the model are divided into three stages: First, using human motion capture and animation data, the tennis ball and the athlete's actions are trained in two stages. Second, training is carried out by embedding motion so that the model can learn human tennis motion. The motion strategy first uses the racket position reward r t r To train the model, use the learning rate (1e -4 ), the action distribution variance ∑π(0.25), and the simulation frequency (120Hz) to implement the simulation; use the reward function r t b To train, and use the learning rate (2e -5) The action distribution variance ∑π(0.04) and the simulation frequency (360 Hz) are used to ensure more accurate simulation of racket contact; the third stage is carried out using the learning rate (1e -5 ) and the action distribution variance ∑π(0.0025).
[0042] In summary, the present invention discloses a method for extracting tennis receiving and hitting skills from sports videos, which can successfully simulate the pulling competition of athletes in tennis match videos.
[0043] The above describes the basic principles, main features and advantages of the present invention. Those of ordinary skill in the art should understand that the above embodiments do not limit the protection scope of the present invention in any form. Any technical solutions obtained by means of equivalent replacement and the like fall within the protection scope of the present invention.
[0044] Parts not involved in the present invention are the same as or can be implemented using the prior art.
Claims
1. The present invention provides a method for extracting tennis hitting and receiving skills from sports videos, and the technical solutions adopted are mainly as follows: Step 1: Estimate the global foot trajectory of tennis players’ hitting and playing postures and construct a tennis motion video dataset; Step 2: Train an initial imitation policy to simulate the tennis player's movements, control the behavior of the simulated character, and generate an initial motion dataset; Step 3: Fit the conditional variational autoencoder VAE to the initial motion dataset to obtain motion embedding and generate athlete-like motion; Step 4: Combine the predictions of the motion-embedded body motion and the character’s wrist motion to train a motion planning strategy to guide the simulated video player’s tennis hitting and receiving motion.
2. According to the method for extracting tennis hitting skills from sports videos of claim 1, the specific steps of step 1 are as follows: Step 1.1: Use the Yolov4 network model to track the movement of players on both sides of the tennis video playing field, obtain point boundaries and player bounding boxes, use the visual transformer with ViT as the core to extract the key points of the tennis player's posture, and use the inverse dynamics method to estimate the posture of the tennis player hitting the tennis ball; the tennis player's bounding box is used to define the player's movement area and input into the posture estimation module; Step 1.2: Since the inverse dynamics method outputs the position and direction of the tennis player's heel in the camera coordinates, it needs to be converted to the court coordinates in the global coordinate system (the origin is located at the center of the tennis court); detect the tennis court boundary lines and their intersections, and use the multi-view perspective transformation algorithm to solve the camera matrix; calculate the position of the tennis player's heel (the center of the two ankle key points), and use the inverse projection transformation to convert the position to the tennis court coordinates; use the camera transformation to calculate the global heel coordinates to minimize the projection error between the player's heel 2D key points and the 3D joint position, correct the heel trajectory, and obtain the motion data set M kin ; Step 1.3: In the modeling phase of tennis motion, the frames containing tennis players hitting the tennis ball in different styles are labeled, and the identities of the tennis players in the frame sequences are labeled.
3. According to the method for extracting tennis hitting and receiving skills from sports videos of claim 1, the specific steps of step 2 are as follows: Step 2.1: Use the following features to represent the character state in, p t is the heel coordinates of the tennis player, represents the linear velocity of the tennis player's heel coordinate, q t is the joint rotation, represents the angular velocity of the joint, represents the target's heel position, Indicates the target's joint rotation angular velocity. Step 2.2: Calculate the tennis player's body torque using a proportional derivative (SD) controller at each non-heel joint position, action a t Specify the target joint angle u of the SD controller t , joint torque T t The calculation formula is: Among them, k p and k d represents the stiffness and damping of the joint, and represents the rotation and angular velocity of the non-heel joint nr; Step 2.3: Including joint rotation speed Location Key Points and penalties Among them, ω0, ω v ,ω p ,ω k ,ω e are the weights of the corresponding items respectively; Step 2.4: The training is carried out in two stages. In the first stage, the motion strategy is trained using human motion capture and animation database so that the simulated athlete can learn to imitate general movements; the penalty term is used Weakened M kin The influence of the difference between frames; output the corrected tennis motion data M corr .
4. According to the method for extracting tennis hitting and receiving skills from sports videos as claimed in claim 1, the specific steps of step 3 are as follows: Step 3.1: Use a multimodal variational autoencoder (MVAE) to predict the motion phase of a tennis player’s cyclical playing posture from receiving the ball to continuous hitting the ball at a fixed position on the court, and use the cyclic phase variable θ to represent the motion phase of each frame; when θ = π, the player hits the ball; when θ = 0 or θ = 2π, the player recovers.
5. According to the method for extracting tennis hitting and receiving skills from sports videos as claimed in claim 1, the specific steps of step 4 are as follows: Step 4.1: In order to overcome the inaccuracy of simulated motion data, the present invention adopts a hybrid control method, in which the whole body motion is controlled by the reference trajectory generated by MVAE, while the wrist motion is controlled by the motion strategy, and each action of the simulated athlete state includes the implicit code of MVAE to generate the target posture of the next frame, and the joint correction of the arm swing. Step 4.2: Use the reward function to control the policy of the simulated movement so that the simulated player can hit the incoming tennis ball so that it bounces at the desired position on the court (the bounce position of the ball) and in the target rotation direction; Before touching the tennis ball, set the racket position reward To minimize the center of the racket head when the person hits the ball Ball position Distance between: After touching the tennis ball, set the ball position reward To minimize the estimated ball bounce position Rebound position with target The distance between the two balls, while ensuring that the ball's rotation direction is consistent with the target's rotation direction: in, where s b and is a binary variable, representing the rotation direction of the simulation ball and the target ball respectively. When the value is 1, it means topspin (the ball rotates forward), and when the value is 0, it means backspin (the ball rotates backward).