A method for intelligently identifying the spin of a table tennis serve
Patent Information
- Application Number
- CN202610862801.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-11
AI Technical Summary
[0003]乒乓球智能训练系统在图像清晰度与落点精度方面取得一定成果,然而,普遍存在以下问题:(1)旋转类型识别能力弱
Smart Images

Figure CN122737592A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of table tennis training assistance technology, specifically relating to an intelligent method for recognizing the spin of a table tennis serve. Background Technology
[0002] Research on table tennis serve spin techniques is a global research hotspot, playing a significant role in improving both the overall ability and training level of table tennis players. Currently, intelligent table tennis training systems are gradually transitioning from technology demonstration to practical feedback, exhibiting the following development trends: (1) Multimodal perception fusion. The system is evolving from single video data to a fusion of video + IMU + biomechanical data, improving the accuracy of spin recognition and action prediction. (2) Lightweight intelligent prediction models. To adapt to on-site training deployment, deep learning models need to be compressed into edge-deployable versions. (3) Visualized feedback interaction system. A teaching interaction interface with real-time feedback, action scoring, and problem prompts enhances the user experience. (4) Knowledge-driven intelligent recommendation. A knowledge graph and data-driven mechanism are established by combining expert experience and training logic to achieve automatic diagnosis and optimization suggestions.
[0003] The intelligent training system for table tennis has achieved certain results in terms of image clarity and landing point accuracy. However, the following problems are common: (1) Weak spin type recognition ability. The recognition accuracy drops sharply under complex conditions such as strong light, occlusion, and non-frontal serves. (2) Intensity prediction cannot be quantified. It is difficult to accurately estimate the ball's angular velocity and spin components. (3) Lack of hand shape and spin mapping mechanism. No prediction model has been established from the serve hand shape to the spin state. (4) Poor real-time feedback capability. Most systems require offline processing and lack low-latency training feedback function. (5) Expensive system price and high deployment threshold. Some high-end systems have strict requirements for venues and high equipment maintenance costs, making them difficult to popularize. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides an intelligent method for recognizing the spin of a table tennis serve, thereby resolving the issues in the prior art. The technical solution adopted by this invention is as follows: A method for intelligent recognition of spin in table tennis serves includes the following steps: Step 1, Data Acquisition and Preprocessing: Collect multimodal data of the athlete's serving motion, including hand movements, racket angle, and ball motion data; process the raw multimodal data to generate a multimodal time-series dataset; Step 2, Feature Process and Analysis: Based on computer vision and sports biomechanics, key hand points, racket posture and action phase features are extracted from multimodal time series datasets to construct four types of sequence features: key hand point sequence, hand action parameter sequence, racket Euler angle time series, and serve action phase label sequence. Step 3, Multimodal Model Construction and Training: Construct a rotation prediction model based on four types of sequence features. After processing by the rotation prediction model, output the rotation type identification result and the rotation intensity classification result. Step 4, Model Post-processing: The rotation prediction model is subjected to lightweight processing; Step 5, Assisted Training and Guidance: During the athlete's training process, assisted training and guidance are provided based on the real-time output of the rotation prediction model.
[0005] Furthermore, in step 1, the multimodal data is unified with a global timestamp through a linear calibration model, expressed as the following formula: in, For the calibrated global unified timestamp; 1 represents the original timestamp output by each acquisition device; 'a' represents the time scaling factor; and 'b' represents the time offset.
[0006] Furthermore, in step 1, the processing of the original multimodal data includes the calculation of the true value of the sphere's rotation, which includes: Three marker points are pre-set on the ping-pong ball, and the three-dimensional coordinates of the marker points are constructed in the world coordinate system: in, Let be the depth scaling factor for the i-th camera; , Let be the pixel coordinates of the marker point on the image plane of the i-th camera; Let be the intrinsic parameter matrix of the i-th camera; Let be the rotation matrix for the i-th camera; Let be the translation vector of the i-th camera; , , This refers to the three-dimensional coordinates of the marker point in the world coordinate system. Then, the true value of the three-dimensional angular velocity is obtained by calculating the inter-frame displacement of three marked points on the surface of the sphere: in, Let be the vector of the marker point in the k-th frame relative to the center of the sphere. This is the orthogonal rotation matrix of the sphere between the two frames; Rotation matrix traces; The rotation angle of the sphere between the two frames; The frame rate of the high-speed camera; , is the three-dimensional angular velocity vector of the sphere; Rotation matrix The element in the i-th row and j-th column.
[0007] Furthermore, in step 2, the MediaPipeHands model is used to extract the coordinates of 21 key points on the hand, which are then normalized to a local coordinate system with the hand as the origin, resulting in a normalized sequence of hand key points, including: Calculate the coordinates of key points on the hand and the length of the palm: in, , where is the original three-dimensional coordinate of the j-th hand key point; These are the normalized coordinates of the j-th hand key point; , where is the three-dimensional coordinate of key point 0 on the hand; The length of a hand; , which represents the three-dimensional coordinates of the key point at the tip of the middle finger; For frame t of the serve video, the original 3D coordinates of 21 key points of the hand are extracted using the MediaPipeHands model: in, Number the key points of the hand in the MediaPipeHands model; These are the key points of the hand at frame t. The original three-dimensional coordinates; Frame-by-frame normalization calculation: in, , representing the three-dimensional coordinates of the wrist root in frame t; Let J be the normalized 3D coordinates of the j-th keypoint in frame t. Finally, the normalized 3D coordinates of all frames are arranged in chronological order to obtain the final normalized sequence of hand key points.
[0008] Furthermore, in step 2, based on the normalized sequence of hand key points, the hand flexion-extension angle and hand torsion angle are calculated to obtain the sequence of hand movement parameters: Hand flexion and extension angle sequence : in, Let t be the wrist flexion / extension angle in frame t. Let be the forearm vector of frame t; Let t be the hand vector of the t-th frame; Hand twisting angle sequence : in, Let t be the wrist twist angle in frame t. Let t be the projection vector of the hand vector on the forearm normal plane in frame t; Let t be the reference vertical vector for the t-th frame; Finally, the hand motion parameter sequence is obtained by calculating frame by frame. .
[0009] Furthermore, in step 2, the racket rotation matrix is solved using the three-dimensional coordinates of the marked points on the racket frame 4, converted into the corresponding Euler angles, and the racket Euler angle sequence is obtained, including: For each frame of the image, the three-dimensional coordinates of the four marker points on the racket in the world coordinate system are acquired, as follows: in, , where is the three-dimensional coordinate of the i-th racket marker point in frame t; Calculate the racket plane normal vector and the handle direction vector, and orthogonalize them to obtain the three axes of the racket's local coordinate system: in, The unit vector normal to the racket face; The unit vector along the axis of the racket handle; The local coordinate system of the racket is orthogonal to the unit axis; Arrange the racket's local coordinate system to obtain the racket rotation matrix for frame t: Expand to standard format: in, Let be the rotation matrix of the racket from the racket's local coordinate system to the world coordinate system in frame t; are the unit direction vectors of the X-axis, Y-axis, and Z-axis of the racket's local coordinate system at frame t, respectively. The element in the i-th row and j-th column of the rotation matrix; Extract Euler angles in the order of yaw, pitch, and roll: in, The pitch angle is the angle by which the racket rotates around the Y-axis; The yaw angle is the angle by which the racket rotates around the Z-axis. The roll angle is the angle by which the racket rotates around the X-axis. Finally, the racket Euler angle time series was calculated frame by frame: in, This represents the total number of frames in the video of the serve.
[0010] Furthermore, in step 2, the serving motion is divided into four stages based on changes in hand speed: preparation, swing, contact, and follow-through, and a sequence of labels for each stage of the serving motion is constructed. Calculate the linear velocity of key points in the hand: in, Let be the linear velocity of the key point of the hand at time t; Let be the three-dimensional coordinates of the key points of the hand at time t; The time interval between two adjacent frames. , The frame rate of the high-speed camera; Calculate the acceleration at key points of the hand: in, Let be the acceleration of the key points of the hand at time t; The differential symbol; Preparation phase judgment rules: And the duration of this state is greater than the stabilization time threshold. ; Rules for determining the swing phase: Furthermore, the speed is continuously increasing, but has not yet reached its peak speed; Contact Phase Determination Rules: And this moment is a local maximum point of the velocity curve; Follow-up phase determination rules: And the motion has not completely stopped; in, Speed threshold for the preparation phase; The speed threshold for the following phase; The swing acceleration threshold; This is the deceleration threshold; This is the minimum stable duration for the preparation phase; Each frame is judged individually according to the judgment rules for each stage, and a corresponding preparation stage label, swing stage label, contact stage label, and follow-up stage label are assigned. The labels of all frames are arranged in chronological order to obtain the service action stage label sequence.
[0011] Furthermore, step 3 includes: Step 3.1: The normalized hand keypoint sequence and hand motion parameter sequence are concatenated and input into a 3D convolutional network. Simultaneously, the spatiotemporal features of joint coordinates and mechanical parameters are extracted and represented as a spatiotemporal feature vector of hand motion. in, Let be the spatiotemporal feature vector of the hand movement in frame t; For the time-sensing field window size; For feature concatenation operations; This is a 3D convolution operation. These are the trainable weight parameters for the 3D convolution kernel; Represents the sequence of hand flexion and extension angles In the middle, with frame t as the center, before and after each frame... A continuous subsequence of frames; Represents the sequence of hand twisting angles In the middle, with frame t as the center, before and after each frame... A continuous subsequence of frames; Input the racket Euler angle sequence into a 1D temporal convolutional layer to extract the racket posture temporal feature vector: in, Let be the temporal feature vector of the racket posture in frame t. This is a 1D convolution operation. These are the trainable weight parameters for the 1D convolutional kernel; In the time series of Euler angles of the racket, with frame t as the center, before and after each frame... A continuous subsequence of frame roll angles; In the time series of Euler angles of the racket, with frame t as the center, before and after each frame... A continuous subsequence of the pitch angles of a frame; In the time series of Euler angles of the racket, with frame t as the center, before and after each frame... A continuous subsequence of yaw angles of a frame; Finally, attention weights are generated based on the label sequence during the serve phase, and the two feature sets are weighted and then concatenated: Where A is the attention weight of the spatiotemporal feature vector of the hand movement in frame t, and B is the attention weight of the temporal feature vector of the racket posture in frame t. Let be the global feature vector after fusion in frame t; Step 3.2: Input the fused global feature vector into the fully connected layer for dimensionality transformation to obtain a high-dimensional feature vector for classification. in, Input the feature vector into the classifier; , These are the trainable weights and bias parameters for the fully connected layer; Rotation type recognition: Input the classifier into the feature vector Input a 5-class softmax layer and output the probability distribution of the 5 basic rotations: in, , which are the probability vectors for the five types of rotations. These represent topspin, backspin, left-handed spin, right-handed spin, and compound spin, respectively. Trainable parameters for the type classification layer; Pick The recognition result corresponding to the maximum value is used as the final rotation type recognition result; Rotation intensity identification: Will The intensity probability distribution is output by calculating the cosine similarity between the features and the five intensity prototypes: in, , is the probability vector for intensity level 5. The probabilities are extremely weak, weak, medium, strong, and extremely strong, respectively. Let be the prototype feature vector of the i-th intensity level; Pick The level corresponding to the maximum value is used as the final rotation intensity classification result.
[0012] Furthermore, step 5 includes: The four types of sequence features obtained from step 2 of the trainee are compared frame by frame with the standard sequences of professional athletes to calculate the phased action similarity score: Where S is the overall similarity score of the serving motion; Let i be the standard value of the i-th feature in the k-th stage of a professional athlete; , where represents the measured value of the i-th feature of the athlete in the k-th stage; k is the stage number of the serving action, k=1 is the preparation stage, k=2 is the swing stage, k=3 is the contact stage, and k=4 is the follow-up stage; The number of key features in the k-th stage; Athletes to be trained receive supplementary training based on their overall similarity score to their serving motion.
[0013] The present invention has the following beneficial effects: (1) This invention combines multimodal vision technology, deep learning algorithms, and sports biomechanics to propose a new model for table tennis spin recognition and trajectory prediction based on multi-source data fusion. It is expected to break through the limitations of traditional table tennis kinematics research that relies on single videos or laboratory conditions, and promote the cross-integration of sports science and artificial intelligence. The research results will provide a theoretical basis for the dynamic capture and intelligent analysis of complex and fast targets in sports.
[0014] (2) This invention can reduce the traditional large amount of trial and error training through intelligent training platform, reduce the risk of sports injury to athletes, and achieve resource conservation; promote green and low-carbon digital sports training methods, and reduce the energy consumption and wear and tear of excessive reliance on physical equipment. Attached Figure Description
[0015] Figure 1 For flowcharts; Figure 2 This is a schematic diagram of the preliminary experiment. Detailed Implementation
[0016] The following will be described in conjunction with embodiments of the present invention. Figures 1-2 The technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art.
[0017] This invention constructs a complete technology chain from data acquisition and motion analysis to intelligent feedback through multimodal visual fusion and sports biomechanical analysis, promoting the transformation of table tennis training from experience-driven to data-driven.
[0018] like Figure 1 This invention proposes a method for intelligent recognition of spin in table tennis serves, comprising the following steps: Step 1, Data Acquisition and Preprocessing: Collect multimodal data of the athlete's serving action, including hand movements, racket angle and ball motion data; after time calibration, noise reduction and spin ground truth calculation, the raw multimodal data is used to generate a multimodal time series dataset; Multiple high-speed cameras and other multi-view high-speed video equipment were used to capture the serving motions of professional athletes. Combined with motion capture equipment and IMU sensors, the focus was on capturing key elements such as hand shape changes, hand movements, racket angle, and contact point. Video capture was also performed on standard movements and variations of different types of spin serves. Biomechanical parameters were simultaneously acquired using motion capture equipment. The raw data underwent noise reduction, synchronization, and formatting. During this stage, professional athletes were invited to annotate the movements, establishing a mapping relationship between professional terminology and computer features, and constructing a hand shape and movement classification system to form a standardized descriptive language.
[0019] In Step 1, the subject warms up thoroughly. Reflective markers are affixed to key body parts (shoulder, elbow, and wrist joint centers) and designated locations on the racket, and the subject wears an IMU sensor. Afterward, the subject completes the required static calibration posture. The subject is instructed to sequentially perform topspin, backspin, left side spin, right side spin, and compound spin. Each spin type of serve is repeated 10-15 times, with the landing point controlled as close to the designated area as possible. All data acquisition devices (high-speed camera, motion capture device, IMU, force table) are triggered synchronously to ensure the consistency of data streams in time.
[0020] Specifically, during synchronous data acquisition and time calibration across multiple devices, a linear calibration model is used to unify the global timestamp, eliminating clock discrepancies between different devices, as expressed in the following formula: in, For the calibrated global unified timestamp; 1 represents the original timestamp output by each acquisition device; 2 represents the time scaling factor, calculated from the time difference of the rising edge of the synchronization trigger signal, with an ideal value of 1; 3 represents the time offset, which is the average difference between the timestamps of 10 consecutive synchronization triggers between the master device (high-speed camera) and the slave device (motion capture device, IMU sensor).
[0021] The motion capture equipment and IMU sensors employ existing technology of Kalman filtering for noise reduction to eliminate high-frequency electromagnetic noise and motion jitter.
[0022] The calculation of the truth value during rotation includes: Three non-collinear markers are pre-set on the ping-pong ball. The three-dimensional coordinates of the reflective markers in the world coordinate system are reconstructed based on a pinhole camera model. in, Let be the depth scaling factor for the i-th camera; , Let be the pixel coordinates of the marker point on the image plane of the i-th camera; Let be the intrinsic parameter matrix of the i-th camera, which includes focal length, principal point coordinates, and distortion coefficients; Let be the rotation matrix of the i-th camera, describing the rotation relationship from the camera coordinate system to the world coordinate system; Let be the translation vector of the i-th camera, describing the translation relationship from the camera coordinate system to the world coordinate system; , , This represents the three-dimensional coordinates of the marker point in the world coordinate system.
[0023] Then, the true value of the three-dimensional angular velocity is obtained by calculating the inter-frame displacement of three non-collinear marker points on the surface of the sphere: in, Let be the vector of the marker point in the k-th frame relative to the center of the sphere. This is the orthogonal rotation matrix of the sphere between the two frames; Rotation matrix traces; The rotation angle of the sphere between the two frames; The frame rate of the high-speed camera; , is the three-dimensional angular velocity vector of the sphere; Rotation matrix The element in the i-th row and j-th column.
[0024] All data were uniformly resampled to 240Hz, and professional table tennis players labeled the spin type (topspin, backspin, left side spin, right side spin, and compound spin) and the phase of the action.
[0025] The multimodal time-series dataset constructed in this invention includes synchronously acquired high-speed video, IMU data, motion capture, force table, and true values of sphere rotation, among other time-series data.
[0026] Step 2, Feature Process and Analysis: Based on computer vision and sports biomechanics, key hand points, racket posture and action phase features are extracted from multimodal time series datasets to construct key hand point sequences, hand action parameter sequences, racket Euler angle time series, and serve action phase label sequences. Step 2, when implemented in practice, includes: This study extracts key information about hand key points and posture, as well as racket posture and trajectory. The mature MediaPipeHands model will be used to extract the coordinates of 21 key hand points. For occlusion situations, multimodal information can be fused or kinematic constraints can be used for optimization to ensure the accuracy and robustness of key point localization (target accuracy error controlled within <2mm). A hand skeletal key point tracking algorithm will be developed to accurately locate finger joint positions and movement trajectories. A hand motion analysis model will be constructed to quantify key motion parameters such as pronation and supination. Using deep learning target detection and posture estimation algorithms, the real-time position, orientation, and angle of the racket will be extracted from the video. Four non-collinear and non-coplanar markers will be set on the racket frame to extract the rotation matrix. The extracted key point coordinates, posture parameters, and rotation matrix will be used to construct time-series data. The serve technique will be decomposed into four stages: preparation, swing, contact, and follow-through, and key features and temporal relationships of each stage will be extracted.
[0027] Specifically, step 2 includes the following sub-steps: Step 2.1: From the multimodal time-series dataset, the MediaPipeHands model is used to extract the coordinates of 21 key points of the hand, which are then normalized to a local coordinate system with the hand as the origin, resulting in a normalized sequence of hand key points; specifically: First, calculate the coordinates of the key points of the hand and the length of the palm: in, , where is the original three-dimensional coordinate of the j-th hand key point; These are the normalized coordinates of the j-th hand key point; , where is the three-dimensional coordinate of key point 0 on the hand; The length of the palm is defined as the Euclidean distance from the hand to the tip of the middle finger. Here, represents the 3D coordinates of the fingertip key point (key point number 12) of the middle finger. It should be noted that the 21 key points of the hand in the MediaPipeHands model are existing technology, which defines the human hand as 21 3D key points, numbered in a fixed order from 0 to 20, where 1 to 4 are the thumb; 5 to 8 are the index finger; 9 to 12 are the middle finger; 13 to 16 are the ring finger; 17 to 20 are the little finger; and key point 0 is the center of the wrist base.
[0028] Next, for frame t of the serve video, the original 3D coordinates of 21 key points of the hand are extracted using the MediaPipeHands model: in, Number the key points of the hand in the MediaPipeHands model; These are the key points of the hand at frame t. The original three-dimensional coordinates; Frame-by-frame normalization calculation: in, , represents the three-dimensional coordinates of the wrist root in frame t, i.e., the origin of the local hand coordinate system; The length of a hand; Let J be the normalized 3D coordinates of the j-th keypoint in frame t. Finally, the normalized 3D coordinates of all frames are arranged in chronological order to obtain the final normalized hand keypoint sequence. , T represents the total number of frames in the serving action video, 21 represents the spatial dimension of the 21 key points of the hand, and 3 represents the coordinate dimension of the X, Y, and Z three-dimensional spatial coordinates.
[0029] Step 2.2, based on the normalized hand keypoint sequence Calculate the hand flexion and extension angles and the hand twisting angles to quantify the characteristics of the serve's power. Hand flexion and extension angle sequence The amplitude of the wrist's forward and backward swing directly determines the strength of the serve's topspin and backspin. in, Let t be the wrist flexion / extension angle in frame t. Let t be the forearm vector of the t-th frame, pointing from the elbow coordinates to the wrist coordinates; Let t be the hand vector in the t-th frame, pointing from the wrist coordinates to the midpoint of the metacarpal bone; Hand twisting angle sequence The range of wrist rotation (inward or outward) directly determines the strength of the left or right spin in a serve. in, Let t be the wrist twist angle in frame t. Let t be the projection vector of the hand vector on the forearm normal plane in frame t; Let t be the reference vertical vector for the t-th frame, which is obtained by the cross product of the vertical upward vector and the forearm vector; Finally, the hand motion parameter sequence is obtained by calculating frame by frame. .
[0030] Step 2.3: Solve the racket rotation matrix using the 3D coordinates of the four non-collinear and non-coplanar marked points on the racket frame, convert it to the corresponding Euler angles, and obtain the racket Euler angle sequence, including: For each frame of the image, the three-dimensional coordinates of the four marker points on the racket in the world coordinate system are acquired, as follows: in, , where is the three-dimensional coordinate of the i-th racket marker point in frame t; Calculate the racket plane normal vector and the handle direction vector, and orthogonalize them to obtain the three axes of the racket's local coordinate system: in, The unit vector normal to the racket face; The unit vector along the axis of the racket handle; The local coordinate system of the racket is orthogonal to the unit axis; Arrange the racket's local coordinate system to obtain the racket rotation matrix for frame t: Expand to standard format: in, Let be the rotation matrix of the racket from the racket's local coordinate system to the world coordinate system in frame t; are the unit direction vectors of the X-axis, Y-axis, and Z-axis of the racket's local coordinate system at frame t, respectively. Let be the element in the i-th row and j-th column of the rotation matrix, representing the projection of the j-th axis of the racket's local coordinate system onto the i-th axis of the world coordinate system, for example... This represents the projection component of the Y-axis of the racket's local coordinate system onto the X-axis of the world coordinate system.
[0031] Then, extract the Euler angles in the order of yaw, pitch, and roll: in, The pitch angle is the angle by which the racket rotates around the Y-axis; Yaw angle is the angle by which the racket rotates around the Z-axis; The roll angle is the angle by which the racket rotates around the X-axis.
[0032] Finally, the time series of Euler angles of the racket was calculated frame by frame: in, This represents the total number of frames in the video of the serve.
[0033] Step 2.4: Based on the changes in hand speed, the serving action is divided into four stages: preparation, swing, contact, and follow-through, and a sequence of labeling steps for the serving action is obtained. First, calculate the linear velocity of the key points of the hand: in, Let be the linear velocity of the key point of the hand at time t; Let be the three-dimensional coordinates of the key points of the hand at time t; The time interval between two adjacent frames. , The frame rate of the high-speed camera; And calculate the acceleration of key points of the hand: in, Let be the acceleration of the key points of the hand at time t; This is the differential symbol.
[0034] Preparation phase judgment rules: And the duration of this state is greater than the stabilization time threshold. ; Rules for determining the swing phase: Furthermore, the speed is continuously increasing, but has not yet reached its peak speed; Contact Phase Determination Rules: And this moment is a local maximum point of the velocity curve; Follow-up phase determination rules: And the motion has not completely stopped; in, Speed threshold for the preparation phase; The speed threshold for the following phase; The swing acceleration threshold; This is the deceleration threshold; This is the minimum stable duration for the preparation phase; Based on the above judgment rules and threshold conditions, each frame is judged individually and assigned corresponding labels for the preparation phase, swing phase, contact phase, and follow-up phase. The labels of all frames are arranged in chronological order to obtain the serving motion phase label sequence. The thresholds for each phase are determined statistically from multiple sets of professional serving motions.
[0035] Step 3, Multimodal Model Construction and Training: Construct a spin prediction model. Input the normalized hand key point sequence, hand motion parameter sequence, racket Euler angle time sequence, and serve action stage label sequence obtained in Step 2. After processing by the spin prediction model, output the spin type recognition result and spin intensity classification result. Step 3 includes the following sub-steps: Step 3.1: Construct a dual-path hybrid architecture consisting of a 3D convolutional network and a 1D temporal convolutional layer, extracting features from the four types of input data and fusing them using stage-based attention weighted fusion. The normalized hand keypoint sequence and the hand motion parameter sequence are concatenated and then input into a 3D convolutional network. At the same time, the spatiotemporal features of joint coordinates and mechanical parameters are extracted and represented as a spatiotemporal feature vector of hand motion: in, Let be the spatiotemporal feature vector of the hand movement in frame t; The size of the time receptive field window is set to 3, meaning that the context information of the preceding and following 3 frames is used for each frame. For feature concatenation operations; This is a 3D convolution operation. These are the trainable weight parameters for the 3D convolution kernel; Represents the sequence of hand flexion and extension angles In the middle, with frame t as the center, before and after each frame... A continuous subsequence of frames; Represents the sequence of hand twisting angles In the middle, with frame t as the center, before and after each frame... A continuous subsequence of frames.
[0036] Input the racket Euler angle sequence into a 1D temporal convolutional layer to extract the racket posture temporal feature vector: in, Let be the temporal feature vector of the racket posture in frame t. This is a 1D convolution operation. These are the trainable weight parameters for the 1D convolutional kernel; In the time series of Euler angles of the racket, with frame t as the center, before and after each frame... A continuous subsequence of frame roll angles; In the time series of Euler angles of the racket, with frame t as the center, before and after each frame... A continuous subsequence of the pitch angles of a frame; In the time series of Euler angles of the racket, with frame t as the center, before and after each frame... A continuous subsequence of yaw angles of a frame.
[0037] Finally, attention weights are generated based on the label sequence during the serve phase, and the two feature sets are weighted and then concatenated: Where A is the attention weight of the spatiotemporal feature vector of the hand movement in frame t, and B is the attention weight of the temporal feature vector of the racket posture in frame t. Let be the global feature vector after fusion in frame t.
[0038] Step 3.2: Input the fused global feature vector into the fully connected layer for dimensionality transformation to obtain a high-dimensional feature vector for classification. in, Input the feature vector into the classifier; , These are the trainable weights and bias parameters for the fully connected layer; (1) Rotation type recognition: Input the classifier into the feature vector Input a 5-class softmax layer and output the probability distribution of the 5 basic rotations: in, , which are the probability vectors for the five types of rotations. These represent topspin, backspin, left-handed spin, right-handed spin, and compound spin, respectively. These are the trainable parameters for the type classification layer.
[0039] Pick Maximum value, final rotation type recognition result for: (2) Rotation intensity identification: Will The input prototype learning fine-classification module calculates the cosine similarity between features and the five intensity prototypes, and outputs the intensity probability distribution: in, , is the probability vector for intensity level 5. The probabilities are extremely weak, weak, medium, strong, and extremely strong, respectively. The prototype feature vector of the i-th intensity level is obtained from the average features of similar samples in the training set.
[0040] Pick The level corresponding to the maximum value is used as the final rotation intensity classification result. : In addition, the global feature vector sequence is coarsely classified by rotation type and then finely classified by rotation intensity, with training supervised by a weighted cross-entropy loss function. in, The total loss value for model training; Cross-entropy loss for rotation type classification, including 5 types: upspin, downspin, left spin, right spin, and compound rotation; The cross-entropy loss is used to classify spin intensity, with 5 levels: extremely weak, weak, medium, strong, and extremely strong. S is the intensity loss weight coefficient, determined by statistical analysis of multiple sets of professional athletes' serve data, and is set to 0.6.
[0041] Step 4, Model Post-processing: The rotation prediction model is subjected to lightweight processing; The rotation prediction model is compressed using a combination of channel pruning and INT8 linear quantization, while retaining the core feature extraction layer. in, The compression ratio of model parameters; This represents the total number of parameters in the original model. The effective number of parameters in the pruned and quantized model, with a target compression ratio of ≥70%.
[0042] Finally, the system was tested using the multimodal time-series dataset constructed in steps 1-2 to verify the core performance metrics: in, This represents the total end-to-end latency of the system. This represents the total time spent on feature extraction in step 2. The inference time per frame for the rotation prediction model; This includes the time required for outputting results and rendering the interface. Typically, the required accuracy is ≥92% for rotation type recognition, ≥87% for five-level rotation intensity classification, and ≤100ms for end-to-end latency.
[0043] Step 5, Assisted Training and Guidance: During the athlete's training process, assisted training and guidance are provided based on the real-time output of the rotation prediction model.
[0044] The step 2 feature sequence of the athlete to be trained is compared frame by frame with the standard sequence of professional athletes to calculate the phased movement similarity score: Where S is the overall similarity score of the serving action, with a maximum score of 100 points; Let i be the standard value of the i-th feature in the k-th stage of a professional athlete; , where represents the measured value of the i-th feature of the athlete in the k-th stage; k is the stage number of the serving action, k=1 is the preparation stage, k=2 is the swing stage, k=3 is the contact stage, and k=4 is the follow-up stage; The number of key features in the k-th stage; During the preparation and swing phases, as well as These refer to the hand flexion-extension angle and the hand torsion angle. During the contact phase, The racket's roll angle, pitch angle, and yaw angle; During the follow-up phase, Trajectory features of key points in the hand.
[0045] Furthermore, step 5 can generate personalized training based on stage defects, statistically analyze the similarity scores and rotation recognition accuracy of each action stage, locate weak points, and automatically adjust the training scheme: in, Let be the training intensity adjustment coefficient for the m-th type of spin serve; The average motion similarity score of the athlete's m-type spin serve; The target action similarity score is set for this stage; Adjustments to the training program include: when The focus of training is on wrist movement and racket angle during the contact phase of this type of serve; when Increase training in compound rotation combinations; when Introduce competitive serve and receive training.
[0046] The above embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Any modifications, alterations, alterations, or substitutions made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for intelligent recognition of spin in table tennis serves, characterized in that, Includes the following steps: Step 1, Data Acquisition and Preprocessing: Collect multimodal data of the athlete's serving motion, including hand movements, racket angle, and ball motion data; process the raw multimodal data to generate a multimodal time-series dataset; Step 2, Feature Process and Analysis: Based on computer vision and sports biomechanics, key hand points, racket posture and action phase features are extracted from multimodal time series datasets to construct four types of sequence features: key hand point sequence, hand action parameter sequence, racket Euler angle time series, and serve action phase label sequence. Step 3, Multimodal Model Construction and Training: Construct a rotation prediction model based on four types of sequence features. After processing by the rotation prediction model, output the rotation type identification result and the rotation intensity classification result. Step 4, Model Post-processing: The rotation prediction model is subjected to lightweight processing; Step 5, Assisted Training and Guidance: During the athlete's training process, assisted training and guidance are provided based on the real-time output of the rotation prediction model.
2. The intelligent recognition method for table tennis serve spin according to claim 1, characterized in that, In step 1, the multimodal data is unified with a global timestamp through a linear calibration model, expressed as the following formula: in, For the calibrated global unified timestamp; 1 represents the original timestamp output by each acquisition device; 'a' represents the time scaling factor; and 'b' represents the time offset.
3. The intelligent recognition method for table tennis serve spin according to claim 1, characterized in that, In step 1, the processing of the original multimodal data includes the calculation of the true value of the sphere's rotation, including: Three marker points are pre-set on the ping-pong ball, and the three-dimensional coordinates of the marker points are constructed in the world coordinate system: in, Let be the depth scaling factor for the i-th camera; , Let be the pixel coordinates of the marker point on the image plane of the i-th camera; Let be the intrinsic parameter matrix of the i-th camera; Let be the rotation matrix for the i-th camera; Let be the translation vector of the i-th camera; , , This refers to the three-dimensional coordinates of the marker point in the world coordinate system. Then, the true value of the three-dimensional angular velocity is obtained by calculating the inter-frame displacement of three marked points on the surface of the sphere: in, Let be the vector of the marker point in the k-th frame relative to the center of the sphere. This is the orthogonal rotation matrix of the sphere between the two frames; Rotation matrix traces; The rotation angle of the sphere between the two frames; The frame rate of the high-speed camera; , is the three-dimensional angular velocity vector of the sphere; Rotation matrix The element in the i-th row and j-th column.
4. The intelligent recognition method for table tennis serve spin according to claim 1, characterized in that, In step 2, the MediaPipeHands model is used to extract the coordinates of 21 key points on the hand, which are then normalized to a local coordinate system with the hand as the origin, resulting in a normalized sequence of hand key points, including: Calculate the coordinates of key points on the hand and the length of the palm: in, , where is the original three-dimensional coordinate of the j-th hand key point; These are the normalized coordinates of the j-th hand key point; , where is the three-dimensional coordinate of key point 0 on the hand; The length of a hand; , which represents the three-dimensional coordinates of the key point at the tip of the middle finger; For frame t of the serve video, the original 3D coordinates of 21 key points of the hand are extracted using the MediaPipeHands model: in, Number the key points of the hand in the MediaPipeHands model; These are the key points of the hand at frame t. The original three-dimensional coordinates; Frame-by-frame normalization calculation: in, , representing the three-dimensional coordinates of the wrist root in frame t; Let J be the normalized 3D coordinates of the j-th keypoint in frame t. Finally, the normalized 3D coordinates of all frames are arranged in chronological order to obtain the final normalized sequence of hand key points.
5. The intelligent recognition method for table tennis serve spin according to claim 4, characterized in that, In step 2, based on the normalized sequence of hand key points, the hand flexion-extension angle and hand torsion angle are calculated to obtain the sequence of hand movement parameters: Hand flexion and extension angle sequence : in, Let t be the wrist flexion / extension angle in frame t. Let be the forearm vector of frame t; Let t be the hand vector of the t-th frame; Hand twisting angle sequence : in, Let t be the wrist twist angle in frame t. Let t be the projection vector of the hand vector on the forearm normal plane in frame t; Let t be the reference vertical vector for the t-th frame; Finally, the hand motion parameter sequence is obtained by calculating frame by frame. .
6. The intelligent recognition method for table tennis serve spin according to claim 1, characterized in that, In step 2, the racket rotation matrix is solved using the three-dimensional coordinates of the marked points on the racket frame 4, converted into the corresponding Euler angles, and the racket Euler angle sequence is obtained, including: For each frame of the image, the three-dimensional coordinates of the four marker points on the racket in the world coordinate system are acquired, as follows: in, , where is the three-dimensional coordinate of the i-th racket marker point in frame t; Calculate the racket plane normal vector and the handle direction vector, and orthogonalize them to obtain the three axes of the racket's local coordinate system: in, The unit vector normal to the racket face; The unit vector along the axis of the racket handle; The local coordinate system of the racket is orthogonal to the unit axis; Arrange the racket's local coordinate system to obtain the racket rotation matrix for frame t: Expand to standard format: in, Let be the rotation matrix of the racket from the racket's local coordinate system to the world coordinate system in frame t; are the unit direction vectors of the X-axis, Y-axis, and Z-axis of the racket's local coordinate system at frame t, respectively. The element in the i-th row and j-th column of the rotation matrix; Extract Euler angles in the order of yaw, pitch, and roll: in, The pitch angle is the angle by which the racket rotates around the Y-axis; Yaw angle is the angle by which the racket rotates around the Z-axis; The roll angle is the angle by which the racket rotates around the X-axis. Finally, the racket Euler angle time series was calculated frame by frame: in, This represents the total number of frames in the video of the serve.
7. The intelligent recognition method for table tennis serve spin according to claim 4, characterized in that, In step 2, the serving motion is divided into four stages based on changes in hand speed: preparation, swing, contact, and follow-through, and a sequence of labels for each stage of the serving motion is constructed. Calculate the linear velocity of key points in the hand: in, Let be the linear velocity of the key point of the hand at time t; Let be the three-dimensional coordinates of the key points of the hand at time t; The time interval between two adjacent frames. , The frame rate of the high-speed camera; Calculate the acceleration at key points of the hand: in, Let be the acceleration of the key points of the hand at time t; The differential symbol; Preparation phase judgment rules: And the duration of this state is greater than the stabilization time threshold. ; Rules for determining the swing phase: Furthermore, the speed is continuously increasing, but has not yet reached its peak speed; Contact Phase Determination Rules: And this moment is a local maximum point of the velocity curve; Follow-up phase determination rules: And the motion has not completely stopped; in, Speed threshold for the preparation phase; The speed threshold for the following phase; The swing acceleration threshold; This is the deceleration threshold; This is the minimum stable duration for the preparation phase; Each frame is judged individually according to the judgment rules for each stage, and a corresponding preparation stage label, swing stage label, contact stage label, and follow-up stage label are assigned. The labels of all frames are arranged in chronological order to obtain the service action stage label sequence.
8. The intelligent recognition method for table tennis serve spin according to claim 1, characterized in that, Step 3 includes: Step 3.1: The normalized hand keypoint sequence and hand motion parameter sequence are concatenated and input into a 3D convolutional network. Simultaneously, the spatiotemporal features of joint coordinates and mechanical parameters are extracted and represented as a spatiotemporal feature vector of hand motion. in, Let be the spatiotemporal feature vector of the hand movement in frame t; For the time-sensing field window size; For feature concatenation operations; This is a 3D convolution operation. These are the trainable weight parameters for the 3D convolution kernel; Represents the sequence of hand flexion and extension angles In the middle, with frame t as the center, before and after each frame... A continuous subsequence of frames; Represents the sequence of hand twisting angles In the middle, with frame t as the center, before and after each frame... A continuous subsequence of frames; Input the racket Euler angle sequence into a 1D temporal convolutional layer to extract the racket posture temporal feature vector: in, Let be the temporal feature vector of the racket posture in frame t. This is a 1D convolution operation. These are the trainable weight parameters for the 1D convolutional kernel; In the time series of Euler angles of the racket, with frame t as the center, before and after each frame... A continuous subsequence of frame roll angles; In the time series of Euler angles of the racket, with frame t as the center, before and after each frame... A continuous subsequence of the pitch angles of a frame; In the time series of Euler angles of the racket, with frame t as the center, before and after each frame... A continuous subsequence of yaw angles of a frame; Finally, attention weights are generated based on the label sequence during the serve phase, and the two feature sets are weighted and then concatenated: Where A is the attention weight of the spatiotemporal feature vector of the hand movement in frame t, and B is the attention weight of the temporal feature vector of the racket posture in frame t. Let be the global feature vector after fusion in frame t; Step 3.2: Input the fused global feature vector into the fully connected layer for dimensionality transformation to obtain a high-dimensional feature vector for classification. in, Input the feature vector into the classifier; , These are the trainable weights and bias parameters for the fully connected layer; Rotation type recognition: Input the classifier into the feature vector Input a 5-class softmax layer and output the probability distribution of the 5 basic rotations: in, , which are the probability vectors for the five types of rotations. These represent topspin, backspin, left-handed spin, right-handed spin, and compound spin, respectively. Trainable parameters for the type classification layer; Pick The recognition result corresponding to the maximum value is used as the final rotation type recognition result; Rotation intensity identification: Will The intensity probability distribution is output by calculating the cosine similarity between the features and the five intensity prototypes: in, , is the probability vector for intensity level 5. The probabilities are extremely weak, weak, medium, strong, and extremely strong, respectively. Let be the prototype feature vector of the i-th intensity level; Pick The level corresponding to the maximum value is used as the final rotation intensity classification result.
9. The intelligent recognition method for table tennis serve spin according to claim 1, characterized in that, Step 5 includes: The four types of sequence features obtained from step 2 of the trainee are compared frame by frame with the standard sequences of professional athletes to calculate the phased action similarity score: Where S is the overall similarity score of the serving motion; Let i be the standard value of the i-th feature in the k-th stage of a professional athlete; , where represents the measured value of the i-th feature of the athlete in the k-th stage; k is the stage number of the serving action, k=1 is the preparation stage, k=2 is the swing stage, k=3 is the contact stage, and k=4 is the follow-up stage; The number of key features in the k-th stage; Athletes to be trained receive supplementary training based on their overall similarity score to their serving motion.