Table tennis technical identification method with track data fused with multi-source information

By fusing multi-source data and processing data in a time and space, a fused feature vector is generated, which solves the problems of posture dependence and recognition error in the existing technology. This enables high-precision and rapid recognition of table tennis techniques and meets the real-time analysis needs of professional competition training.

CN121353733APending Publication Date: 2026-01-16NINGBO UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511315845.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing technologies rely excessively on single posture features when recognizing technical movements in table tennis match videos, leading to recognition bias and errors. This fails to meet the requirements for high-precision analysis, and the trajectory data is not effectively integrated, making it difficult to balance recognition efficiency and accuracy, thus failing to meet the immediate needs of professional competition training.

Method used

By preprocessing multi-source data, the data on the trajectory of the ping-pong ball, the trajectory of the racket, and the posture of the athlete are spatiotemporally synchronized. A channel attention mechanism and a spatiotemporal dual-stream fusion structure are adopted to generate a fusion feature vector. Combined with coarse classification and fine-grained verification, accurate recognition of technical movements is achieved.

Benefits of technology

It significantly improves the accuracy and robustness of technical motion recognition, solves the limitations of posture dependence and confusion of similar movements, and meets the real-time analysis needs of professional competition training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353733A_ABST
    Figure CN121353733A_ABST
Patent Text Reader

Abstract

The invention provides a table tennis technical identification method with track data fused with multi-source information. The table tennis technical identification method comprises the steps that S1, any table tennis competition video is obtained, and table tennis track data, racket movement track data and athlete posture data are extracted; s2, by taking the ball hitting moment of the table tennis ball as a time anchor point, aligning time axes of the table tennis ball track data, the racket movement track data and the athlete posture data through a dynamic time warping algorithm, and performing space normalization in a unified three-dimensional table coordinate system; s3, weighting the table tennis track data, the racket motion track data and the athlete posture data after space normalization by adopting a channel attention mechanism to obtain weighted features, and inputting the weighted features into a space-time double-flow fusion structure to generate a fusion feature vector; and S4, performing basic classification and similar action distinguishing based on the fusion feature vector, and outputting a final technical action type. The method has the beneficial effect that the table tennis technical actions can be quickly and accurately recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of table tennis data analysis, and more specifically, to a method for table tennis technique recognition by fusing trajectory data with multi-source information. Background Technology

[0002] Currently, the technology of recognizing movements from table tennis match videos is widely used in the field of competition and training analysis. Its core purpose is to extract key data from stored match videos (such as MP4 and AVI formats) to determine the type of technical movements used by athletes (such as forehand fast attack and backhand flick), providing a basis for coaches to assess athletes' technical levels and develop targeted training plans. The core approach of existing technologies heavily relies on "athlete body posture recognition," that is, extracting athlete limb features (such as joint angles and trunk posture) through computer vision technology, and then combining them with preset rules or simple classification models to complete the technical movement determination. This is currently the mainstream technological direction in the industry.

[0003] Existing technology 1: Technical recognition scheme based on single pose feature This solution extracts athlete posture data using computer vision technologies (such as MediaPipe and OpenPose), focusing on changes in the angles of the shoulder, elbow, and wrist joints, as well as the torso position. It constructs a "posture-technique" mapping rule to identify technical movements: for example, if the athlete's elbow angle is less than 90°, the shoulder joint is on the left side of the body (relative to the table), and the torso is biased towards the backhand area, it is determined as a "backhand flick"; if the elbow angle is greater than 120°, the shoulder joint is on the right side of the body, and the torso is leaning forward, it is determined as a "forehand fast attack". This solution does not require complex calculations and can run independently on edge devices (such as mobile devices). The time for single-frame posture extraction is ≤30ms, and the time for basic technical movement (such as forehand attack and backhand attack) recognition is ≤100ms.

[0004] Existing technology 2: A technical recognition scheme based on "attitude + simple trajectory supplementation" This solution, based on single posture recognition, supplements the extraction of coarse trajectory features of the ping-pong ball (e.g., only identifying the landing area as the forehand / backhand zone, without calculating ball speed and trajectory curvature). The two types of data are then concatenated and input into a lightweight CNN classification model. For example, "athlete's backhand posture + ping-pong ball landing in the backhand zone" is initially identified as a backhand technique, and further identified as a "backhand topspin" based on "wrist joint rotation angle > 30°" in the posture. If the wrist joint rotation angle is < 10°, it is identified as a "backhand flick". This solution requires uploading video frames to the cloud to complete model inference and relies on open-source object detection models (such as YOLOv5) to achieve ping-pong ball area recognition. The time for single technique action recognition is ≤ 250ms.

[0005] The existing technology has the following drawbacks: Over-reliance on pose features highlights the limitations of a single data source. Existing technologies rely entirely on body posture data, ignoring the motion of the table tennis ball and racket, leading to significant errors in technique recognition. For example, the difference in athlete posture between "backhand fast drive" and "backhand fast rip" is minimal (elbow joint angle is only ±5°, wrist joint rotation angle is only ±8°), making it impossible to distinguish them based solely on posture features. The confusion rate between the two types of movements is ≥40%. Furthermore, when an athlete uses "backhand side attack" in the forehand area, the posture features show that the torso is biased towards the forehand area, which can easily be misjudged as "forehand fast attack," resulting in an error rate of ≥35%, which cannot meet the requirements for high-precision analysis. The trajectory data is only used as a supplement and has not been effectively integrated. While existing technology 2 incorporates table tennis trajectory data, it only serves as a "regional" auxiliary judgment (such as the forehand / backhand zone) and fails to explore the spatiotemporal relationship between trajectory and posture. For example, the core difference between "forehand fast attack" and "forehand fast drive" lies in "ball speed" (fast attack ball speed is usually >38m / s, fast drive <35m / s) and "swing timing" (the ball is on the rise during a fast attack, and at its peak during a fast drive). However, this solution does not extract key trajectory features such as ball speed and hitting sequence, relying solely on posture, resulting in an accuracy rate of ≤70% for both types of actions. Furthermore, the trajectory and posture data are not spatiotemporally aligned (e.g., there is a 2-3 frame deviation between the moment the table tennis ball lands and the moment the posture is captured), further amplifying the recognition error. Pose recognition fails in complex scenarios, exhibiting poor robustness. The core bottleneck of existing technologies lies in the fact that "pose extraction is easily interfered with": In doubles matches, athletes are easily obstructed by their partners (such as when hitting a backhand shot, the elbow joint is obstructed by the forehand partner), resulting in incomplete pose feature extraction, with an elbow joint angle error of ≥20°, and the technical recognition accuracy drops sharply to ≤50%; in high-speed offensive and defensive transition scenarios (such as long rallies), the athlete's limb movements are blurred, and the joint point detection miss rate is ≥15% (such as the wrist joint cannot be recognized), directly causing the technical recognition to be interrupted; in addition, in low-light environments (such as uneven lighting on the court), the joint positioning deviation of the pose extraction algorithm is ≥10px (pixels), further reducing the recognition accuracy; It is difficult to balance recognition efficiency and accuracy, and the adaptability is insufficient. While existing technology one is highly efficient (it can run at the edge), its accuracy is extremely low (high confusion rate for similar actions), failing to meet the needs of professional competition training. Technology two introduces simple trajectory data to improve accuracy, but requires uploading all videos to the cloud for inference. The technical recognition time for a single match (2 hours) is ≥40 minutes, which cannot meet the coach's immediate need to "quickly view technical statistics during training breaks." At the same time, neither of these solutions performs lightweight processing for multi-source data. The data size of a single frame during cloud inference is ≥15KB, which can easily lead to bandwidth congestion when multiple users use it simultaneously, increasing the recognition latency to ≥600ms. Summary of the Invention

[0006] The technical problem to be solved by this invention is how to achieve rapid and accurate recognition of table tennis techniques. In order to overcome the defects of the above-mentioned existing technologies (or related technologies), this invention provides a table tennis technique recognition method that integrates trajectory data with multi-source information.

[0007] This invention provides a method for recognizing table tennis techniques by fusing trajectory data from multiple sources, comprising the following steps: Step S1, Multi-source data preprocessing: Obtain any table tennis match video and extract table tennis trajectory data, racket movement trajectory data and athlete posture data from the table tennis match video; Step S2, Multi-source data spatiotemporal synchronization alignment: Using the moment of hitting the ping-pong ball as the time anchor point, the time axes of the ping-pong ball trajectory data, the racket motion trajectory data, and the athlete posture data are aligned through a dynamic time warping algorithm, and spatial normalization is performed in a unified three-dimensional table coordinate system; Step S3, Multi-source data deep feature fusion: The spatially normalized table tennis ball trajectory data, racket motion trajectory data and athlete posture data are weighted using a channel attention mechanism to obtain weighted features, which are then input into the spatiotemporal dual-stream fusion structure to generate a fusion feature vector; Step S4, Technical Action Recognition: Based on the fused feature vector, perform basic classification and similar action differentiation, and output the final technical action type.

[0008] The table tennis technique recognition method based on trajectory data fusion and multi-source information of this invention has the following advantages compared with existing technologies: In this invention, video data is extracted in step S1, the time axis of the three types of data is aligned in step S2, the data is weighted and calculated in step S3 to generate a fusion feature vector, and the action is classified and distinguished in step S4 to obtain the final technical action type. This invention breaks through the limitation of existing technologies that rely too much on single posture data. By fusing three types of multi-source information, namely table tennis ball trajectory data, racket motion trajectory data and athlete posture data, and through spatiotemporal synchronization and deep feature fusion, the overall accuracy and robustness of technical action recognition are significantly improved.

[0009] In one possible implementation, the table tennis ball trajectory data extracted in step S1 includes the physical coordinates of the table tennis ball, the coordinates of the landing point, the ball speed, and the trajectory curvature; the racket motion trajectory data includes the physical coordinates of the racket, the swing direction, and the racket face angle; and the athlete posture data includes the athlete's key joint coordinates, forehand and backhand labels, and joint angles.

[0010] Compared with existing technologies, the above technical solution can limit the specific data dimensions extracted in step S1, clarify the core features of various types of data, and ensure the completeness and efficiency of information extraction. Furthermore, the extracted physical coordinates, landing point coordinates, ball speed, trajectory curvature of the ping-pong ball, physical coordinates, swing direction, racket face angle of the racket, as well as the coordinates of the athlete's key joints, forehand and backhand labels, and joint angles are all the most critical and direct parameters for distinguishing different technical movements. This lays a solid data foundation for subsequent accurate fusion and recognition, and avoids computational redundancy caused by useless information.

[0011] In one possible implementation, in step S2, the condition for determining the time of hitting the ping-pong ball is to determine whether the physical distance between the ping-pong ball and the racket at the current moment is less than or equal to 2cm and whether the sudden change in ball speed is greater than or equal to 5m / s: if yes, then the current moment is determined to be the time of hitting the ping-pong ball; if no, then the current moment is determined not to be the time of hitting the ping-pong ball.

[0012] Compared with existing technologies, the above-mentioned technical solution clarifies the criteria for determining the critical time anchor point of the ball's impact. By using the dual physical conditions of "whether the physical distance is less than or equal to 2cm and whether the sudden change in ball speed is greater than or equal to 5m / s", the accuracy of the ball's impact frame positioning is greatly improved compared with relying solely on image recognition or distance judgment. This provides a reliable time reference for the time synchronization of subsequent multi-source data and fundamentally reduces the recognition error caused by timing misalignment.

[0013] In one possible implementation, in step S2, a dynamic time warping algorithm is used to align the time axes of the racket trajectory data and the athlete posture data to the hitting frame, and the time difference between the ping-pong ball trajectory data, the racket trajectory data, and the athlete posture data within the range of three frames before and after the hitting frame is adjusted to be less than or equal to 8.3ms.

[0014] Compared with existing technologies, the above-mentioned technical solution limits the specific algorithm and accuracy requirements for time synchronization. The dynamic time warping algorithm can flexibly align time-series data of different lengths, controlling the time difference of the three types of data at the critical moment of hitting the ball within 8.3ms. This ensures that at the core moment of the technical action, the posture, racket and ball motion are highly correlated and matched, thereby greatly improving the effectiveness of feature vector fusion.

[0015] In one possible implementation, in step S3, a channel attention mechanism is used to assign a feature weight of 0.3 to the ping-pong ball trajectory data, a feature weight of 0.4 to the racket motion trajectory data, and a feature weight of 0.3 to the athlete posture data, and then a weighted calculation is performed.

[0016] Compared with existing technologies, the above-mentioned technical solution clarifies the weight allocation strategy in the channel attention mechanism, overturning the unbalanced situation in existing technologies where posture features dominate. By assigning the highest weight to racket trajectory data and balancing the contributions of table tennis trajectory data and athlete posture data, it is more helpful to distinguish technical movements with similar postures but different racket usage and ball trajectories, so that attention is focused on the most distinguishable information, fundamentally solving the problem of high confusion rate of similar movements.

[0017] In one possible implementation, the spatiotemporal dual-stream fusion structure includes a spatial flow branch and a temporal flow branch. In step S3, the fusion feature vector is statically spatially extracted through the spatial flow branch to obtain a 256-dimensional spatial feature vector, and the fusion feature vector is dynamically spatially extracted through the temporal flow branch to obtain a 256-dimensional temporal feature vector. Subsequently, the spatial feature vector and the temporal feature vector are concatenated to form a 512-dimensional feature vector, and then the dimensionality is reduced to 256-dimensional feature vector using a 1x1 convolution kernel as the fusion feature vector.

[0018] Compared with existing technologies, the above-mentioned technical solution defines the specific composition of the spatiotemporal dual-stream fusion structure, while capturing information on both the static spatial configuration and dynamic temporal evolution of technical actions. The spatial flow branch extracts the "spatial relative relationship between the body, racket, and ball at a certain moment", and the temporal flow branch extracts dynamic processes such as "ball speed change, swing trajectory, and posture continuity". Finally, the fusion and dimensionality reduction are performed. This design enables the generated 256-dimensional fusion feature vector to contain rich static and dynamic information, with extremely strong representation capabilities. Moreover, the dimensionality reduction maintains lightweightness, which is conducive to efficient classification in the future.

[0019] In one possible implementation, in step S4, the fused feature vector is input into a fully connected neural network to output preliminary probabilities corresponding to 12 basic technical actions. The basic technical action with the highest preliminary probability is taken as the coarse classification result. Subsequently, a contrastive loss mechanism and a fine-grained feature verification mechanism are used to enhance the distinguishability between the basic technical action in the coarse classification result and similar basic technical actions to obtain a fine-grained feature verification result. The final technical action type is obtained by combining the coarse classification result and the fine-grained feature verification result.

[0020] Compared with existing technologies, the above-mentioned technical solution defines the specific process of technical action recognition. The two-level recognition mechanism of "coarse classification + fine verification" greatly improves the accuracy and reliability of the final judgment. First, a preliminary probability classification is performed through a fully connected neural network to ensure the basic recognition rate. Then, for easily confused categories, contrast loss and threshold verification are specifically used for secondary identification. This mechanism is particularly suitable for solving the problem of distinguishing between many "similar in form but not in spirit" actions in table tennis, and is the key to achieving high accuracy.

[0021] In one possible implementation, in step S7, for the basic technical action in the coarse classification result and any similar basic technical action, a cosine distance is calculated based on the fusion feature vector of the two basic technical actions, and the cosine distance between basic technical actions of the same class is minimized, while the cosine distance between basic technical actions of different classes is maximized.

[0022] Compared with existing technologies, the above technical solution limits the specific implementation of the contrast loss mechanism. It actively and explicitly increases the distance between different types of samples and reduces the distance between similar samples at the feature space level. It learns the subtle but essential differences between similar actions, enhances the sensitivity and discrimination ability to subtle differences, and makes actions such as "backhand quick tear" and "backhand quick pull" no longer difficult due to similar features.

[0023] In one possible implementation, in step S7, the fine-grained feature verification mechanism includes threshold verification for distinguishing between forehand fast attack and fast drive, and threshold verification for distinguishing between backhand fast rip and fast drive. In the threshold verification for distinguishing between forehand fast attack and forehand fast drive, a swing direction > 60° and ball speed > 38m / s indicates a forehand fast attack, while a swing direction < 45° and ball speed < 35m / s indicates a forehand fast drive. In the threshold verification for distinguishing between backhand fast rip and backhand fast drive, a racket face angle > 65° and ball speed > 38m / s indicates a backhand fast rip, while a racket face angle < 55° and ball speed < 35m / s indicates a backhand fast drive.

[0024] Compared with existing technologies, the above-mentioned technical solution clarifies the specific physical rule thresholds used in fine-grained feature verification, providing a clear and interpretable physical basis for the final judgment of similar basic technical actions. These thresholds are derived from prior knowledge of professional table tennis techniques, forming a reliable decision boundary. By combining data-driven model classification results with knowledge-driven rule verification, the correctness of the recognition results is guaranteed in two ways. Attached Figure Description

[0025] Figure 1 This is a flowchart of the steps of the present invention. Detailed Implementation

[0026] First, those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention. Those skilled in the art can make adjustments as needed to adapt to specific application scenarios.

[0027] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0028] See Figure 1 This invention discloses a method for recognizing table tennis techniques by fusing trajectory data from multiple sources, comprising: Step S1, Multi-source data preprocessing: Obtain any table tennis match video and extract table tennis trajectory data, racket motion trajectory data and athlete posture data from the table tennis match video; Step S2, Spatiotemporal synchronization and alignment of multi-source data: Using the moment of hitting the ping-pong ball as the time anchor point, the time axes of the ping-pong ball trajectory data, racket motion trajectory data and athlete posture data are aligned through a dynamic time warping algorithm, and spatial normalization is performed in a unified three-dimensional table coordinate system. Step S3, deep feature fusion of multi-source data: The channel attention mechanism is used to weight the spatially normalized table tennis ball trajectory data, racket motion trajectory data and athlete posture data to obtain weighted features, and input them into the spatiotemporal dual-stream fusion structure to generate fusion feature vector; Step S4, Technical Action Recognition: Based on the fused feature vector, perform basic classification and similar action differentiation, and output the final technical action type.

[0029] In this embodiment of the invention, in step S1, the table tennis match video is analyzed to extract table tennis trajectory data, racket motion trajectory data, and athlete posture data, respectively, to construct a data source system of "trajectory-driven and posture-assisted". An improved YOLOv8s model with SKNet attention mechanism is used to detect table tennis balls in the video frames of the table tennis match video to solve the problem of missing small targets under low light. Candidate trajectories are generated by combining the inter-frame difference method. The trajectory continuity is optimized by the tree pruning trajectory restoration algorithm of Yu Yang's team (2024) to reduce the breakpoint rate caused by motion blur. Finally, the physical coordinates of the table tennis ball (based on the three-dimensional coordinate system of the table, with the center point of the table as the origin, the x-axis along the length of the table, the y-axis along the width, and the z-axis perpendicular to the table surface), the coordinates of the landing point, the ball speed, and the trajectory curvature are output. The formula for calculating the ball speed is as follows: in, Indicates ball speed. Indicates the first video frames Axis coordinates Indicates the first video frames Axis coordinates Indicates the first video frames Axis coordinates Indicates the first video frames Axis coordinates Indicates the first The video frame and the The time difference between video frames; The formula for calculating the trajectory curvature is as follows: in, Represents the curvature of the trajectory. This represents the second derivative of the x-coordinate with respect to time. express The first derivative of the coordinate with respect to time, This represents the first derivative of the x-coordinate with respect to time. express The second derivative of the coordinate with respect to time.

[0030] In this embodiment of the invention, in step S1, the racket is detected based on the YOLOv8s model with optimized 2DCNN branch (Jiao Yuming 2024), and the coordinates of the racket center point and the racket face angle are located (calculated by the angle between the major axis and the x-axis of the racket bounding box); the optical flow features are extracted using the GoogLeNet network improved by Zhang Ao (2023) to supplement the racket motion trend and generate the racket motion trajectory; the physical coordinates of the racket, the swing direction (the angle between the trajectory tangent direction and the x-axis), and the racket face angle (the angle with the table surface) are output to distinguish similar technical actions (such as backhand fast drive / fast tear). 23 key skeletal points of the athlete were extracted using MediaPipe, focusing on the three-dimensional coordinates of the shoulder, elbow, and wrist joints; core posture features were calculated: forehand and backhand determination (if the y-coordinate of the shoulder joint is less than the y-coordinate of the center point of the table and the elbow angle is less than 100°, it is determined to be a backhand posture, otherwise it is a forehand posture).

[0031] In this embodiment of the invention, step S1 uses the "ping-pong ball striking moment" as the core anchor point to solve the problem of spatiotemporal misalignment between posture and trajectory data in the prior art, ensuring the correlation of the three types of data. The striking moment is accurately determined through the ping-pong ball trajectory data: when the physical distance between the ping-pong ball and the racket is ≤2cm (contact threshold) and the ball speed changes abruptly (Δv≥5m / s), the frame is marked as the "striking frame" (timestamp t0). The Dynamic Time Warping (DTW) algorithm is used to align the time axis of the racket motion trajectory data and the athlete posture data to the striking frame, ensuring that the time difference of the three types of data within t0±3 frames is ≤1 frame (≤8.3ms at 120fps), avoiding "posture and ball / racket motion mismatch" caused by timing deviation.

[0032] In this embodiment of the invention, in step S2, a unified three-dimensional coordinate system for the table tennis table is established (the origin is the center point of the table tennis table, the x-axis range is [-1.525m, 1.525m], the y-axis range is [-0.7625m, 0.7625m], and the z-axis range is [0, 2m]). The table tennis ball coordinates, racket coordinates, and skeleton point coordinates extracted in step S1 are converted into physical coordinates using the pixel-physical distance mapping formula. The calibration coefficient is calculated from the pixel length of the edge of the table tennis table of known size. For example, if the table tennis table length is 2.74m and the corresponding pixel length is 1000px, then the calibration coefficient = 2.74 / 1000 = 0.00274m / px. This achieves spatial normalization and eliminates data deviations caused by differences in pixel dimensions.

[0033] In this embodiment of the invention, step S3 introduces a CBAM attention mechanism to assign weights based on the technical differentiation value of the three types of data, thus solving the imbalance problem of "excessive posture weight" in the prior art. The weighting formula used for the feature weighting calculation of the hitting frame t0±3 frames is as follows: in, Indicates weight, Indicates the current frame. The frame representing the hit is designed to ensure that the feature weight of the hit moment (the core instant of the technical action) is the highest, while minimizing interference from non-key frames. The weights are allocated according to the "value of technical differentiation": table tennis trajectory data (coordinates of the landing point, ball speed) has a weight of 0.3, racket motion trajectory data (swing direction, racket face angle) has a weight of 0.4, and athlete posture data (forehand and backhand labels, elbow joint angle) has a weight of 0.3, to avoid over-reliance on posture; Output weighted single-frame features: 8-dimensional single-frame features including table tennis ball features (placement point x / y, ball speed, curvature), racket features (swing direction, racket face angle), and posture features (forehand and backhand labels, elbow joint angle).

[0034] In this embodiment of the invention, in step S3, a dual-branch fusion structure of "spatial flow branch + temporal flow branch" is constructed to capture both static features and dynamic changes. The weighted single-frame features of frames t0±3 are input into the improved GoogLeNet network (refer to Zhang Ao 2023) through the spatial flow branch to extract static spatial features (such as the spatial combination features of "backhand posture + racket swing direction 55° + table tennis ball landing backhand area"), and output a 256-dimensional spatial feature vector. The table tennis ball trajectory sequence, racket trajectory sequence, and posture sequence of frames t0±5 are input into the LSTM network through the temporal flow branch to capture dynamic temporal features (such as the dynamic changes of technical actions such as "ball speed increases from 35m / s to 41m / s + swing direction changes from 40° to 55°"), and output a 256-dimensional temporal feature vector. Finally, the spatial feature vector and the temporal feature vector are concatenated into a 512-dimensional vector, and the dimensionality is reduced to 256 dimensions through a 1×1 convolution kernel to obtain the final "fusion feature vector", which retains the key information of each data source and eliminates information redundancy.

[0035] In this embodiment of the invention, in step S4, a two-level recognition model of "basic classification + similarity differentiation" is constructed based on the fused feature vector. This model focuses on solving the problem of confusion between similar actions in existing technologies. The 256-dimensional fused feature vector is input into a fully connected neural network, which outputs the preliminary probabilities of 12 basic technical actions (forehand fast attack, forehand fast drive, backhand fast flick, backhand fast drive, forehand topspin, backhand topspin, etc.). The category with the higher preliminary probability is used as the coarse classification result. The classification model is trained using the cross-entropy loss function, and the training dataset consists of 100,000+ labeled samples (including table tennis balls, rackets, posture data, and corresponding technical labels) to ensure that the preliminary recognition accuracy of basic technical actions is ≥90%.

[0036] In this embodiment of the invention, in step S4, for easily confused similar actions in the coarse classification results (such as forehand fast attack / drive, backhand fast tear / drive), a dual mechanism of "contrast loss + fine-grained feature verification" is introduced to enhance the discriminative power. The cosine distance is calculated for the fused feature vectors of similar actions, as shown in the following formula: in, Represents the cosine distance. This represents the fused feature vector of one of the similar actions. This represents a fusion feature vector representing another similar action. By minimizing the cosine distance between similar basic techniques and maximizing the cosine distance between dissimilar basic techniques, it learns subtle differences between similar actions. Based on the key features of the racket and the table tennis ball, a preset discrimination threshold is established to directly solve the confusion problem caused by similar postures. For example: Forehand fast attack vs. forehand fast drive: Forehand fast attack is defined as a swing direction with an angle greater than 60° to the x-axis and a ball speed greater than 38m / s; forehand fast drive is defined as a swing direction less than 45° and a ball speed less than 35m / s. Backhand flick vs. backhand drive: Backhand flick is defined as racket face angle > 65° and ball speed > 38m / s; backhand drive is defined as racket face angle < 55° and ball speed < 35m / s. Finally, by combining the coarse classification results (weight 0.6) and the fine-grained feature verification results (weight 0.4), the final technical action type is determined (e.g., "coarse classification as backhand fast rip + ball speed 41m / s + racket face angle 68° → final identification as backhand fast rip").

[0037] In this embodiment of the invention, to address the shortcomings of existing technologies such as "posture being easily affected by occlusion and motion blur", a supplementary data completion and error correction mechanism can be used to ensure recognition stability in complex scenarios. When key joints of the posture are occluded (such as the elbow joint being occluded by a partner), the correlation of skeletal nodes is modeled based on the ST-GCN network, and the angle and position of the occluded joint are predicted using the coordinates of the unoccluded shoulder and wrist joints. The prediction error is ≤8°, avoiding posture failure caused by the absence of a single joint. When the overall posture is blurred (such as limb blurring caused by high-speed movement), the posture features are inferred by the movement trends of the racket trajectory and the ping-pong ball trajectory. For example, if the racket swing direction is in the backhand area and the ball speed is >38m / s, the athlete's wrist joint rotation angle is inferred to be >30°, supplementing the key parameters of the blurred posture.

[0038] In this embodiment of the invention, when a breakpoint in the ping-pong ball trajectory is generated (no ping-pong ball is detected for two consecutive frames), the trajectory continuity constraint of the tree pruning algorithm is used, and the ball coordinates at the breakpoint are predicted by Kalman filtering in combination with the ball speed and the landing point position. When there is racket occlusion (such as the racket being blocked by an arm), the racket coordinates and swing direction of the occluded frame are completed based on the racket trajectory trend of the previous 5 frames (through linear fitting), with a completion error of ≤3°, to ensure that key distinguishing data are not missing.

[0039] Example 1 This embodiment takes "analyzing a stored table tennis singles match video (MP4 format, 120fps, resolution 1920×1080, duration 2 hours) and identifying the athlete's 'backhand fast flick' technique" as an example to explain the implementation process in detail; Step 1: Multi-source data preprocessing 1.1 Ping Pong Ball Trajectory Extraction (Core Anchor Point) An improved YOLOv8s model incorporating the SKNet attention mechanism was used to detect the pixel coordinates (1100px, 580px) of the ping-pong ball at 20 minutes and 15 seconds (frame number 145800, timestamp t=1215s). The coordinates were then converted to physical coordinates using pixel-to-physical mapping in the table's 3D coordinate system (calibration coefficient = 2.74m / 1920px ≈ 0.001427m / px). The landing area was determined to be the backhand zone (y < 0); the ball speed was calculated using a tree pruning trajectory reconstruction algorithm: the coordinate difference between adjacent frames (145799-145800). , , ball speed The average ball speed calculated over 10 consecutive frames was 41 m / s; the trajectory curvature was... ; 1.2 Racket Trajectory Extraction (Key Differentiation) Based on the YOLOv8s model with optimized 2DCNN branch, the physical coordinates of the racket center point (1.55m, -0.08m, 0.26m) are detected, and the angle between the racket face and the table surface is 68°. By improving the GoogLeNet network to extract optical flow features, the angle between the swing direction and the x-axis is found to be 55° (forward along the length of the table). 1.3 Attitude Extraction (Basic Support) MediaPipe extracts skeletal points: shoulder joint (1.4m, 0.1m, 1.65m), elbow joint (1.5m, -0.05m, 1.45m), wrist joint (1.54m, -0.07m, 1.35m); calculates elbow joint angle: Since the elbow joint y-coordinate (-0.05m) is less than the table center point y-coordinate (0), it is determined to be a backhand posture.

[0040] Step 2: Spatiotemporal synchronization and alignment of multi-source data 2.1 Time Synchronization When frame number 145800 is detected, the physical distance between the ping-pong ball and the racket is = The ball speed changes abruptly from 38m / s to 41m / s (Δv=3m / s), and this frame is marked as the hitting frame (t0=1215s). The DTW algorithm was used to align the racket trajectory (frames 145797-145803) with the attitude data (frames 145798-145802), with a time difference of ≤8.3ms (1 frame). 2.2 Spatial Synchronization All three types of data were converted into physical coordinates in the three-dimensional coordinate system of the table tennis table, with no spatial deviation (e.g., the table tennis ball x=1.57m and the racket x=1.55m, which is consistent with the spatial correlation at the time of hitting the ball).

[0041] Step 3: Deep Feature Fusion of Multi-Source Data 3.1 Channel Attention Weighting Time weight: t0 frame weight passed Calculations show that after Softmax normalization, the proportion is 42%. Spatial weights: Table tennis ball characteristics (landing point, ball speed) 0.3, racket characteristics (swing direction, racket face angle) 0.4, posture characteristics (backhand label, elbow angle) 0.3; Output 8-dimensional single-frame features: [backhand landing point (1.57, -0.05), ball speed 41m / s, curvature 0.004, swing direction 55°, racket face 68°, backhand posture, elbow angle 88°]; 3.2 Spatiotemporal Dual-Stream Fusion Spatial Flow: Input the weighted features of t0±3 frames into the improved GoogLeNet to extract the spatial features of "backhand posture + 55° swing + backhand landing", and output a 256-dimensional vector; Time Flow: Input the trajectory / attitude sequence of t0 ± 5 frames into LSTM to capture the dynamic features of "ball speed increases from 38m / s to 41m / s + swing direction changes from 50° to 55°", and output a 256-dimensional vector; After concatenation, the dimensionality is reduced to 256 dimensions through 1×1 convolution to fuse the feature vector.

[0042] Step 4: Technical Action Recognition 4.1 Basic Classification The fusion of feature vectors into a fully connected neural network yields the probabilities of 12 basic technical actions: backhand quick tear (83%), backhand quick drive (14%), and others (3%). The coarse classification result is "backhand quick tear". 4.2 Differentiating Similar Actions Contrast loss calculation: cosine distance between the eigenvectors of backhand fast tear and fast pass =0.38 (Distance between different classes > 0.3, discrimination meets the standard); Fine-grained verification: ball speed 41m / s > 38m / s (fast tear threshold), swing direction 55° > 45° (fast drive threshold), racket face angle 68° > 55° (fast drive threshold); Final result: Combining the coarse classification probability (0.83×0.6=0.498) and the verification result (0.9×0.4=0.36), it is determined to be "reverse hand fast tear".

[0043] Step 5: Robustness Optimization for Complex Scenes (Additional Example of Occlusion Scene) If frame number 145801 shows "elbow joint is obscured by athlete's torso", based on the ST-GCN network, the elbow joint angle is predicted to be 86° (error 2°) using the coordinates of the shoulder joint (1.4m, 0.1m, 1.65m) and wrist joint (1.54m, -0.07m, 1.35m). At the same time, by using the racket swing direction of 55° and the ball speed of 41m / s, the "wrist joint rotation angle > 30°" is deduced to supplement the missing posture information without affecting the final recognition result.

[0044] In the description of this invention, the terms "one embodiment," "some embodiments," "in this embodiment," "specific example," or "some examples," etc., refer to specific features, mechanisms, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, mechanisms, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0045] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for table tennis technology recognition by trajectory data fusion of multi-source information, characterized in that, The method comprises the following steps: Step S1, multi-source data preprocessing: acquiring any table tennis match video and extracting table tennis trajectory data, racket movement trajectory data and athlete posture data from the table tennis match video; Step S2, multi-source data space-time synchronization alignment: taking the table tennis hitting moment as the time anchor point, aligning the time axes of the table tennis trajectory data, the racket movement trajectory data and the athlete posture data through a dynamic time warping algorithm, and performing space normalization in a unified three-dimensional table coordinate system; Step S3, multi-source data deep feature fusion: adopting a channel attention mechanism to weight the space-normalized table tennis trajectory data, racket movement trajectory data and athlete posture data to obtain weighted features, and inputting the weighted features into a space-time dual-flow fusion structure to generate a fusion feature vector; Step S4, technical action recognition: performing basic classification and similar action differentiation based on the fusion feature vector, and outputting a final technical action type.

2. The table tennis technique recognition method according to claim 1, characterized in that, The table tennis trajectory data extracted in the step S1 includes physical coordinates, landing point coordinates, ball speed and trajectory curvature of the table tennis, the racket movement trajectory data includes physical coordinates, swing direction and face angle of the racket, and the athlete posture data includes key joint coordinates, right-left hand labels and joint angles of the athlete.

3. The table tennis technique recognition method according to claim 1, characterized in that, In the step S2, the determination condition of the table tennis hitting moment is whether the physical distance between the table tennis and the racket at the current moment is less than or equal to 2 cm and whether the sudden change of the ball speed is greater than or equal to 5 m / s: if yes, the current moment is determined as the table tennis hitting moment; if no, the current moment is determined as not the table tennis hitting moment.

4. The table tennis technique recognition method according to claim 1, characterized in that, In the step S2, the dynamic time warping algorithm is adopted to align the time axes of the racket movement trajectory data and the athlete posture data to the hitting frame, and the time difference of the table tennis trajectory data, the racket movement trajectory data and the athlete posture data within three frames before and after the hitting frame is adjusted to be less than or equal to 8.3 ms.

5. The table tennis technique recognition method according to claim 1, characterized in that, In the step S3, the channel attention mechanism is adopted to give the table tennis trajectory data a feature weight of 0.3, the racket movement trajectory data a feature weight of 0.4, and the athlete posture data a feature weight of 0.3 for weighted calculation.

6. The table tennis technique recognition method according to claim 1, characterized in that, The space-time dual-flow fusion structure includes a space flow branch and a time flow branch. In the step S3, the fusion feature vector is subjected to static space feature extraction through the space flow branch to obtain a 256-dimensional space feature vector, and is subjected to dynamic space feature extraction through the time flow branch to obtain a 256-dimensional time feature vector, and then the space feature vector and the time feature vector are spliced into a 512-dimensional feature vector, which is reduced to a 256-dimensional feature vector through a 1x1 convolution kernel as the fusion feature vector.

7. The table tennis technique recognition method according to claim 1, characterized in that, In the step S4, the fusion feature vector is input into a fully connected neural network to output a preliminary probability corresponding to 12 basic technical actions, and the basic technical action with the highest preliminary probability is taken as a coarse classification result. Then, a contrast loss mechanism and a fine-grained feature verification mechanism are used to strengthen the discrimination between the basic technical action in the coarse classification result and similar basic technical actions, to obtain a fine-grained feature verification result. Finally, the final technical action type is obtained by combining the coarse classification result and the fine-grained feature verification result.

8. The table tennis technique recognition method according to claim 7, characterized in that, In the step S7, for the basic technical action in the coarse classification result and any similar basic technical action, the cosine distance is calculated according to the fusion feature vectors of the two basic technical actions, and the cosine distance between similar basic technical actions is minimized, and the cosine distance between dissimilar basic technical actions is maximized.

9. The table tennis technique recognition method according to claim 7, characterized in that, In the step S7, the fine-grained feature verification mechanism includes a differentiation threshold verification of forehand fast attack and fast serve, and a differentiation threshold verification of backhand fast tear and fast serve. In the differentiation threshold verification of forehand fast attack and fast serve, a swing direction > 60° and a ball speed > 38 m / s are taken as forehand fast attack, and a swing direction < 45° and a ball speed < 35 m / s are taken as forehand fast serve. In the differentiation threshold verification of backhand fast tear and fast serve, a racket face angle > 65° and a ball speed > 38 m / s are taken as backhand fast tear, and a racket face angle < 55° and a ball speed < 35 m / s are taken as backhand fast serve.