Cloud-side-end video analysis method for table tennis
Through the cloud-edge-end video analysis method, the inter-frame difference method and deep learning model are used to identify the keyframes and hitting landing points in the table tennis game video, and a tactical action map is constructed, which solves the problem of inefficient analysis of table tennis game videos, and realizes the identification and tactical intention analysis of multi-round tactical combinations.
Patent Information
- Application Number
- CN202510394400.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-29
AI Technical Summary
现有技术在乒乓球比赛视频分析中存在整体效率低下,无法识别多回合战术组合,无法帮助用户进行战术意图分析的问题。
The cloud-edge-end video analysis method is used to detect keyframe images through the inter-frame differential method of the mobile terminal, and the hitting landing point is identified using the YOLOv5s model, and tactical tags are constructed through the LSTM model to construct tactical action maps to store them in the cloud server.
It improves the efficiency of table tennis game video analysis, can identify multi-round tactical combinations, and guides users to perform tactical intention analysis through tactical action maps.
Smart Images

Figure CN120388317A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video analysis, and more particularly, to a cloud-edge-end video analysis method for table tennis sports. Background Art
[0002] Visual analysis is a product of the field of data visualization, dedicated to promoting analytical reasoning through technologies such as interactive visual interfaces and computer vision to obtain data insights. With the rapid development of computer science and technology, visual analysis technology, as an effective data analysis means, is increasingly applied to the analysis of sports competition data. It can not only provide guidance and assistance for the training and competition of coaches and athletes, but also improve the communication effect of various media. In the field of professional competition data analysis, visual analysis technology can more directly and effectively display the spatio-temporal characteristics of the competition and the technical and tactical characteristics of athletes, attracting the attention of many professional sports data analysts.
[0003] For table tennis match videos, ball detection and stroke action recognition are the two most crucial elements in sports data visual analysis. Precise ball recognition can not only better visualize the match video, such as showing the trajectory of the ball, but also assist downstream data mining tasks, such as stroke action classification. Stroke action recognition is the key to rally segmentation and tactical analysis, and can also provide statistical data for multi-shot rallies. Based on multi-rally stroke action data, the player's ball path and tactics can be accurately and efficiently analyzed, which can not only provide reliable support for the athletes of the national table tennis team to study opponents, refine tactics, and prepare for competitions efficiently, but also provide an economical and efficient "professional data analyst" for amateur enthusiasts to improve their ball paths and study tactics.
[0004] Computer vision methods have gradually been applied to the data extraction of sports video visual analysis. However, there are also many challenges in using deep neural network models to achieve real-time, end-to-end ball detection and stroke action recognition in live table tennis match videos: firstly, since the table tennis ball in table tennis match videos is not only tiny but also in a high-speed motion state, often presenting a blurred elongated tail shadow, which poses new challenges compared to traditional detection and recognition tasks; secondly, the stroke actions of athletes in table tennis match videos are often very concealed in amplitude, and high-speed motion is likely to cause image blurring. For an end-to-end model with multiple tasks, it is very important to extract generalizable visual features.
[0005] Currently, the common analysis method for table tennis motion videos is to upload the complete game video to the cloud server and use a deep learning model (such as 3D-CNN) to analyze the hitting actions and landing point distributions offline to generate statistical reports. In this method, non-critical frame images (such as pauses and ball picking) account for more than 70%, resulting in low overall analysis efficiency, long single analysis time, and only supporting single-hit classification. It cannot identify multi-round tactical combinations and cannot help users analyze tactical intentions. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to improve the overall analysis efficiency, identify multi-round tactical combinations, and help users analyze tactical intentions. To overcome the defects of the above-mentioned prior art (or related technologies), the present invention provides a cloud-edge-end video analysis method for table tennis motion.
[0007] The present invention provides a cloud-edge-end video analysis method for table tennis motion, including the following steps: Step S1, obtain at least one table tennis motion video from the mobile terminal, and detect and obtain a set of key frame images containing hitting events in the table tennis motion video by using the inter-frame difference method to form a key frame image set; Step S2, input the key frame image set into a pre-configured YOLOv5s model to generate feature maps of different scales, and perform object detection on each of the feature maps to obtain corresponding hitting landing points; Step S3, construct a hitting sequence feature vector according to each of the hitting landing points, and input the hitting sequence feature vector into a pre-configured LSTM model to output corresponding tactical labels; Step S4, extract tactical actions as nodes from the tactical labels, extract the transition probabilities between the tactical actions as edge weights, and construct a tactical action graph and store it in the cloud server.
[0008] Compared with the prior art, the cloud-edge-end video analysis method for table tennis motion of the present invention has the following advantages: In the present invention, video recognition and extraction of key frame images are performed through step S1, generation of feature maps and intelligent recognition of hitting landing points are performed through step S2, construction of a hitting sequence feature vector and recognition of tactical labels are performed through step S3, and construction of a tactical action graph is performed through step S4. In the whole process, only the key frame images in the table tennis motion video are analyzed, reducing the interference of unnecessary images, improving the overall analysis efficiency, and being able to identify multi-round tactical combinations composed of hitting landing points and tactical labels, and then displaying them to the user through the tactical action graph to guide the user to analyze tactical intentions.
[0009] In a possible implementation, in step S1, after detecting the key frame image containing the hitting event in the table tennis motion video by using the inter-frame difference method, the video images within the range of 10 to 30 frames before and after the key frame image are retained and included in the key frame image set.
[0010] In a possible implementation, in step S1, a plurality of video images sorted in chronological order are detected in the table tennis motion video by using the inter-frame difference method. For each video image, any pixel point in the video image is extracted as the current pixel point, and the pixel point at the same position as the current pixel point in the previous video image of the video image is extracted as the historical pixel point. It is judged whether the pixel value difference between the current pixel point and the historical pixel point satisfies a preset trigger condition: If so, the video image is used as the key frame image; If not, exit.
[0011] In a possible implementation, the trigger condition in step S1 is the following expression: Wherein, represents the pixel value corresponding to the current pixel point; represents the current time; represents the pixel value corresponding to the historical pixel point; represents the previous time; represents the horizontal coordinate corresponding to the current pixel point or the historical pixel point; represents the vertical coordinate corresponding to the current pixel point or the historical pixel point; represents the adaptive threshold.
[0012] In a possible implementation, in step S1, it further includes: For each key frame image, the table tennis table area and the non-table tennis table area in the key frame image are identified by using a color filtering algorithm and a shape detection algorithm. The table tennis table area is encoded by using H.265, and the non-table tennis table area is encoded by using H.264.
[0013] In a possible implementation, step S2 includes: Step S21: For each key-frame image in the key-frame image set, input the key-frame image into the YOLOv5s model to perform feature extraction operations of multi-layer convolution and pooling to obtain feature maps of different scales; Step S22: For each feature map, perform object detection on the feature map to obtain the table tennis ball position; Step S23: Frame the table tennis ball position with a bounding box, obtain the upper left coordinate and the lower right coordinate of the bounding box, and obtain the hitting landing point according to the upper left coordinate and the lower right coordinate.
[0014] In a possible implementation manner, the hitting sequence feature vector in step S3 is the following expression: F = [x1, y1, t1; x2, y2, t2,... x i , y i , t i where, F represents the hitting sequence feature vector; x i represents the horizontal coordinate of the hitting landing point in the image plane at the i-th hit; y i represents the vertical coordinate of the hitting landing point in the image plane at the i-th hit; t i represents the time point corresponding to the hitting event at the i-th hit.
[0015] In a possible implementation manner, in step S4, it further includes: Obtain at least one historical table tennis motion video, for any two of the tactical actions, count the number of times of transferring from one tactical action to another tactical action in the historical table tennis motion video, and obtain the transfer weight between the two tactical actions and include it in the tactical action map.
[0016] In a possible implementation manner, in step S8, the transfer weight is obtained through the following calculation formula: where, represents the transfer weight from tactical action to tactical action ; represents the number of times of transferring from tactical action to tactical action ; Indicates the number of transfers from tactical action to tactical action . Brief Description of the Drawings
[0017] Figure 1 is the flowchart of the steps of the present invention. Detailed Embodiment
[0018] First of all, those skilled in the art should understand that these embodiments are only used to explain the technical principles of the embodiments of the present invention, and are not intended to limit the protection scope of the embodiments of the present invention. Those skilled in the art can adjust them as needed to adapt to specific application scenarios.
[0019] The present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0020] Refer to Figure 1 , an embodiment of the present invention discloses a cloud-edge-end video analysis method for table tennis sports, including: Step S1, obtain at least one table tennis sports video from the mobile terminal, and detect a key frame image set composed of multiple key frame images containing hitting events in the table tennis sports video through the inter-frame difference method; Step S2, input the key frame image set into a pre-configured YOLOv5s model to generate feature maps of different scales, and perform object detection on each feature map to obtain corresponding hitting landing points; Step S3, construct a hitting sequence feature vector according to each hitting landing point, and input the hitting sequence feature vector into a pre-configured LSTM model to output a corresponding tactical label; Step S4, extract the tactical action as a node from the tactical label, extract the transition probability between the tactical actions as the edge weight, and construct a tactical action graph and store it in the cloud server.
[0021] In this embodiment, the lightweight mobile terminal is used as the edge side, and the table tennis sports video transmitted from the edge side is obtained. The hitting event in the table tennis sports video is detected through the inter-frame difference method, and only 10-30 frames (adjustable parameters) before and after the hitting event are retained to reduce the amount of transmitted data.
[0022] A smart phone or a customized sports camera equipped with a Qualcomm Snapdragon 888 processor is used, and a six-axis gyroscope sensor is built in for video anti-shake processing to ensure the video clarity in high-speed motion scenarios. The mobile terminal device transmits data through Wi-Fi6 or 5G network, supports the dynamic bandwidth allocation algorithm, and preferentially uploads the key frame image set to the edge side after detecting the hitting event.
[0023] By performing operations on each pixel point (x, y) in the key-frame image, the pixel value difference between two adjacent key-frame images is calculated, and then it is determined whether the triggering condition for the hitting event is satisfied. The triggering condition in step S1 is the following expression: Wherein, represents the pixel value corresponding to the current pixel point; represents the current moment; represents the pixel value corresponding to the historical pixel point; represents the previous moment; represents the horizontal coordinate corresponding to the current pixel point or the historical pixel point; represents the vertical coordinate corresponding to the current pixel point or the historical pixel point; represents the adaptive threshold, which takes a value of 0.5 in this embodiment.
[0024] After the hitting event is recognized, hierarchical coding and transmission are performed. Since the color of the table tennis table area is relatively fixed (usually dark green or blue) and the shape is a regular rectangle, the color filtering algorithm can be used to first screen out the pixel area within a specific color range, and then the table tennis table area is determined by the shape detection algorithm (such as using the Hough transform to detect lines to determine the rectangle boundary). The table tennis table area (ROI) is encoded using H.265 (QP = 30), and other areas are downgraded to H.264 (QP = 40), reducing the bit rate by 55%.
[0025] The threshold setting of the inter-frame difference method in step S1 adopts a dynamic adjustment strategy. The key-frame image is divided into nine-grid areas according to the illumination conditions of the table tennis table area, and the local threshold is calculated for each sub-area. The formula is as follows: Wherein, represents the local threshold of the k-th sub-area, represents the global reference threshold, which is preset to 0.5, represents the brightness compensation coefficient, and its value range is 0.2 - 0.8, represents the average brightness value of the k-th sub-area.
[0026] In step S1, the exponential weighted moving average algorithm (EWMA) can be introduced to eliminate instantaneous interference. The update formula for the historical frame difference value is: Wherein, represents the smoothed difference value of the t-th hitting event, represents the original difference value of the t-th hitting event, represents the smoothing coefficient.
[0027] In this embodiment, the edge computing node is used as the edge side. The key frame image set after intelligent video slicing is input into the YOLOv5s model. Feature extraction operations such as multi-layer convolution and pooling are performed on the input key frame images. The loss function is used to calculate the difference between the prediction result and the image annotation. The model parameters of the YOLOv5s model are continuously adjusted through the backpropagation algorithm to generate feature maps of different scales, and then object detection is performed on these feature maps to obtain the position of the table tennis ball.
[0028] Deploy an edge server cluster based on NVIDIA Jetson AGX Xavier, configure a dual-GPU parallel computing architecture, with 16GB video memory for each node, support joint inference acceleration of the YOLOv5s model and the LSTM model, and use TensorRT for model quantization optimization, with the inference speed increased to 43FPS (1080p input).
[0029] After the YOLOv5s model detects the position of the table tennis ball, obtain the upper left coordinate (x1, y1) and the lower right coordinate (x2, y2) of the bounding box of the table tennis ball position. Then the coordinates of the hitting point are x = (x1 + x2) / 2, y = (y1 + y2) / 2, and the hitting point (x, y) can be output.
[0030] When training the YOLOv5s model, a large amount of image data containing hitting type annotations (forehand / backhand) is input into the model. The model learns the image features of forehand and backhand hits, and learns the feature patterns such as the position relationship, action postures, etc. among the table tennis ball, racket, and athlete. In real-time detection, the model matches the currently extracted key frame image with the image data learned during training, detects the position of the table tennis ball in the key frame image based on the features learned by the model, outputs the bounding box of the table tennis ball position, and at the same time, by calculating the probabilities of different hitting types, outputs the hitting type label (forehand / backhand) with the highest probability.
[0031] In this embodiment, the following improvements can be made to the YOLOv5s model: 1. Data augmentation strategy Generate adversarial network (GAN) to synthesize high-speed motion blurred images, simulate the hitting scenario with a ball speed of 20m / s, enhance the reflection spots of the racket, and randomly adjust the saturation (±15%) and brightness (±10%) in the HSV color space; 2. Improvement of the loss function Adopt the combined loss function of CIoU Loss + Focal Loss to solve the problem of detecting small targets (the diameter of the table tennis ball is about 4cm); 3. Multi-scale feature fusion Based on the original three detection heads (20×20, 40×40, 80×80), a new 160×160 fine-grained detection head is added, and the recall rate for small-sized table tennis balls (image occupancy <0.1%) is increased to 92.3%.
[0032] In step S3, a batting sequence feature vector is constructed and input into a lightweight LSTM model to output tactical labels (such as "fast attack and oppression", "short pendulum control", etc.). The batting sequence feature vector is expressed as follows: F=[x1,y1,t1;x2,y2,t2,...x i ,y i ,t i Among them, F represents the batting sequence feature vector; x i represents the horizontal coordinate of the batting landing point in the image plane at the i-th batting; y i represents the vertical coordinate of the batting landing point in the image plane at the i-th batting; t i represents the time point corresponding to the batting event at the i-th batting. By recording the time of each batting event, the sequence and interval of batting actions in the time dimension can be reflected. Combining with the batting landing point coordinates helps analyze the rhythm and tactical coherence of batting. Set appropriate number of layers and hidden units in the LSTM model, then connect the fully connected layer, and the output layer uses the softmax activation function to convert the output into the probability distribution of each tactical label category, and select the category with the highest probability as the predicted tactical label.
[0033] In this embodiment, the server is used as the cloud. The tactical labels output by the lightweight LSTM model represent the specific batting sequence and tactical execution situation, which are important data sources for constructing the tactical knowledge graph. Based on the graph neural network (GNN), the tactical transfer relationship is modeled. The nodes are tactical actions (such as "flick", "chop long", etc.), and the edge weights are transfer probabilities (such as the probability of "short pendulum → chop long" = 65%). Systematically sorting out the associations and conversion rules between tactical actions in table tennis matches can help athletes and coaches better understand the dynamic changes of match tactics, so as to formulate more reasonable tactical strategies.
[0034] An Alibaba Cloud ECS elastic computing instance is used to build a distributed graph database (Neo4j cluster) to store the tactical action graph constructed from more than 500,000 professional match data, and dynamic scaling is achieved through Kubernetes, with the response time controlled within 200ms.
[0035] In step S4, at least one historical table tennis video is obtained. For any two tactical actions, the number of transitions from one tactical action to another in the historical table tennis video is counted. Based on the number of transitions, the transition weights of the two tactical actions are obtained and included in the tactical action graph. The transition weights are calculated using the following formula: in, Indicates tactical action To tactical action The transfer weight of Indicates tactical action Shift to tactical action the number of times; Indicates tactical action Shift to tactical action the number of times; A weight matrix describing the transfer relationship between tactical actions is calculated through formulas. Combined with the tactical knowledge graph constructed by the graph neural network (GNN), the association and conversion rules between tactical actions in table tennis games can be analyzed.
[0036] Through historical data statistics, we use data analysis tools and statistical methods to conduct in-depth analysis. Taking the landing point of the ball as an example, we count the opponent's hitting frequency in different positions and calculate indicators such as the forehand utilization rate (number of forehand hits ÷ total number of hits × 100%). If the forehand utilization rate is greater than 70%, it means that the opponent has a clear preference for the forehand position in choosing the landing point of the ball.
[0037] Based on the opponent's tactical preferences derived from the analysis, combined with the team's own technical characteristics and advantages, targeted strategic suggestions are generated. If the opponent's forehand usage rate is found to be > 70%, it means that the opponent is more accustomed to and proficient in the forehand position. In this case, the suggestion of "increasing the proportion of backhand landing points" can be generated. By hitting the ball more towards the backhand position, the opponent's batting rhythm is disrupted, making it difficult for the opponent to play to the advantage of the forehand position.
[0038] In this embodiment of the present invention, after a hitting event is detected (the inter-frame difference value exceeds a threshold), a video clip (720p@30fps) 2 seconds after the hitting event is intercepted. The data size of the table area (ROI) after encoding is reduced from 25MB of original data to 9MB after compression. The YOLOv5s model detects that the landing point of the hit is the opponent's forehand short ball and the hit type is a backhand twist and pull. The LSTM model inputs the feature vectors of three consecutive hitting sequences and outputs a probability of "serve and attack" of 85%. Then, the tactical knowledge graph is queried and the historical optimal response strategy for "backhand twist and pull short ball" is matched as "forehand fast diagonal". The opponent's diagonal defensive point loss rate (62%) is calculated and the strategy is sent to the mobile terminal for display.
[0039] In the description of the present invention, the descriptions referring to terms such as "one embodiment", "some embodiments", "in this embodiment", "specific examples", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0040] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A cloud-edge-end video analysis method for table tennis sports, characterized in that, Including the following steps: Step S1: Obtain at least one table tennis motion video from a mobile device, and detect a set of key frame images containing hitting events within the table tennis motion video through the inter-frame difference method, which constitutes the key frame image set; Step S2: Input the key frame image set into a pre-configured YOLOv5s model to generate feature maps of different scales, and perform object detection on each of the feature maps to obtain corresponding hitting landing points; Step S3: Construct a hitting sequence feature vector based on each of the hitting landing points, and input the hitting sequence feature vector into a pre-configured LSTM model to output corresponding tactical labels; Step S4: Extract tactical actions as nodes from the tactical labels, extract the transition probabilities between the tactical actions as edge weights, and construct a tactical action graph and store it in a cloud server.
2. The cloud-edge-end video analysis method according to claim 1, wherein In step S1, after detecting the key frame images containing hitting events within the table tennis motion video through the inter-frame difference method, retain the video images within the range of 10 to 30 frames before and after the key frame images and include them in the key frame image set.
3. The cloud-edge-end video analysis method according to claim 1, wherein In step S1, detect multiple video images sorted by time in the table tennis motion video through the inter-frame difference method. For each video image, extract any pixel point within the video image as the current pixel point, extract the pixel point at the same position as the current pixel point in the previous video image of the video image as the historical pixel point, and determine whether the pixel value difference between the current pixel point and the historical pixel point satisfies a preset trigger condition: If so, use the video image as the key frame image; If not, exit.
4. The cloud-edge-terminal video analysis method according to claim 3, characterized in that The trigger condition in step S1 is the following expression: Where, Indicates the pixel value corresponding to the current pixel point; Indicates the current moment; represent the pixel value corresponding to the historical pixel point; Indicates the previous moment; represents the horizontal coordinate corresponding to the current pixel point or the historical pixel point; represents the vertical direction coordinate corresponding to the current pixel point or the historical pixel point; Indicates an adaptive threshold.
5. The cloud-edge-end video analysis method according to claim 1, wherein, In step S1, it further includes: For each key frame image, use a color filtering algorithm and a shape detection algorithm to identify the table tennis table area and non-table tennis table area in the key frame image, encode the table tennis table area using H.265, and encode the non-table tennis table area using H.
264.
6. The cloud-edge-terminal video analysis method according to claim 1, wherein Step S2 includes: Step S21: For each key frame image in the key frame image set, input the key frame image into the YOLOv5s model to perform feature extraction operations of multi-layer convolution and pooling to obtain feature maps of different scales; Step S22: For each feature map, perform object detection on the feature map to obtain the position of the table tennis ball; Step S23: Frame the position of the table tennis ball with a bounding box, obtain the upper left coordinate and the lower right coordinate of the bounding box, and obtain the hitting landing point according to the upper left coordinate and the lower right coordinate.
7. The cloud-edge-terminal video analysis method according to claim 1, characterized in that, The hitting sequence feature vector in step S3 is the following expression: F = [x1, y1, t1; x2, y2, t2,... x i , y i , t i Where, F represents the hitting sequence feature vector; x i represents the horizontal coordinate in the image plane of the hitting point when hitting the ball for the i-th time; y i represents the vertical direction coordinate of the hitting point in the image plane when hitting the ball for the i-th time; t i It represents the time point corresponding to the hitting event at the i-th hit.
8. The cloud-edge-end video analysis method according to claim 1, wherein, In step S4, it further includes: Obtain at least one historical table tennis movement video. For any two of the said tactical actions, count the number of times of transferring from one of the said tactical actions to another in the historical table tennis movement video, and obtain the transfer weight between the two said tactical actions according to the number of times, and include it in the tactical action map.
9. The cloud-edge-terminal video analysis method according to claim 8, wherein In the step S8, the transfer weight is obtained through the following calculation formula: Wherein, Indicates tactical action To tactical action The transfer weight of Indicates the number of times of transitioning from a tactical action to a tactical action; Indicates the number of times of transfer from a tactical action to a tactical action.