Table Tennis Intelligent Live Broadcasting Technology Method Based on Visual Video Analysis Technology
Through the intelligent table tennis guide method based on visual video analysis technology, combined with multi-scale confidence generation and sparse feature fusion technology, the problem of fast lens switching speed and frequent movements in table tennis event guide is solved, fast and accurate action recognition is achieved, and the director efficiency is improved.
Patent Information
- Application Number
- CN202210752392.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-06-28
AI Technical Summary
The existing technology cannot effectively solve the problem of fast camera switching speed and frequent movements in table tennis event directors, and the timing action positioning technology is slow, the accuracy is insufficient, and the real-time requirements cannot be met.
The intelligent ping-pong guide method based on visual video analysis technology is adopted, and through the steps of data set construction, preprocessing, feature extraction, timing action positioning and action recognition, combined with multi-scale confidence generation and sparse feature fusion technology, the rapid and accurate recognition of ping-pong movements is achieved.
It realizes intelligent director of table tennis, with fast detection speed and high action recognition rate, which can greatly reduce human resources, improve the speed of directors, and provide reference for intelligent directors in other competition fields.
Smart Images

Figure CN115115987B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a table tennis intelligent live broadcast technology method based on visual video analysis technology. Background Art
[0002] With the wide rise of the live broadcast field and people's high attention to sports events, the real-time requirement for sports event broadcasting is increasing. In addition, due to the gradually wide range of event content and the increasing number of events held over the years, the requirements for human resources and live broadcast level of the live broadcaster are increasing day by day. Especially for sports event live broadcasters, in table tennis events, for example, the camera switching speed is fast and the actions are frequent, which increases the difficulty for human live broadcasters. Thanks to the development of technologies in the field of deep learning, especially the rise of technologies such as video action recognition and temporal action localization in the field of video understanding, it has laid the foundation for the emergence of intelligent event live broadcast based on artificial intelligence technology. However, firstly, the current temporal action localization technology is slow and difficult to meet the real-time requirement, and the accuracy of short-time action localization is insufficient. Secondly, although the action detection scheme has been gradually improved, there is currently no complete technical method for sports event live broadcast.
[0003] The temporal action localization technology is a technology that inputs a complete video segment and identifies the position information of actions therein through deep learning technology. Currently, there are mainly three popular mainstream technologies. One is to only judge the start frame and end frame of the action, and the accuracy of this technology is generally not high. One is to generate the probability of the action after judging the start and end frames of the action, represented by algorithms such as BMN and BSN++. This technology has high accuracy but is slow. One is the technology that directly generates candidate boxes based on the Transformer framework, represented by TACNet. This technology can ensure both high speed and accuracy, but has high requirements for the data volume of training data, which is not conducive to application.
[0004] The action recognition technology is a technology that, given an action segment, classifies the action through deep learning methods to confirm the category of the action. Currently, the action recognition technology is mainly divided into the two-stream method (constructing image optical flow information and image information) such as TSN, 3D convolution (CNN technology that also calculates in time) such as I3D, and the Transformer framework such as VIVit. With the development of deep learning, the accuracy and applicability rate of the Transformer framework have gradually increased, and it has become the mainstream development direction of action recognition.
[0005] The problems that cannot be solved by the prior art are divided into two aspects. On the one hand, the solution centered on table tennis detection cannot adapt to the lens switching in different environments. On the other hand, it only optimizes the action recognition of complete videos, but cannot process real-time video streams, and the processing of complete videos is slow in the video segmentation processing task. In addition, there is currently no deep learning intelligent solution for table tennis game live broadcast technology. Summary of the Invention
[0006] The purpose of the present invention is to provide a table tennis intelligent live broadcast technology method based on visual video analysis technology to solve the problems raised in the above background technology.
[0007] To achieve the above purpose, the present invention provides the following technical methods:
[0008] The table tennis intelligent live broadcast technology method based on visual video analysis technology includes the following steps:
[0009] S101. Dataset construction: Collect the feature information of standard multi-camera high-definition live broadcast images in international and domestic table tennis competitions up to the current year, annotate the in-round racket-swinging actions of the athletes facing the camera based on the feature information, and then collect and form a dataset, and perform enhanced training on the annotated multi-camera lens dataset;
[0010] S102. Preprocessing: Convert the video into a 20 - 40s video stream and send it into the feature extraction network;
[0011] S103. Model construction: Include feature extraction network, temporal action localization and action recognition;
[0012] S104. Testing: Adopt the technical means of multi-scale confidence generation and sparse feature fusion of temporal object localization to solve the two problems of inaccurate short-time action localization and difficult extraction of high-fine-grained features;
[0013] S105. Video logic processing, live broadcast video result.
[0014] Preferably, the specific processing method of S102 is as follows:
[0015] Use the main camera and the side camera for action localization and recognition. Send the videos of the main camera and the side camera into the network every 20 - 40s. The main camera is at 30 degrees obliquely above the back of the athlete, and the perspective includes the athlete and the table tennis table. The side camera is at 10 degrees obliquely above the front side, and the perspective includes the two athletes. The images are uploaded and processed simultaneously. There are also other cameras including the competition seat, the studio, and the close-up camera of the athlete, which are only used for lens switching and not for recognition.
[0016] Preferably, the specific steps of S103 are as follows:
[0017] S1301. Feature extraction using a feature extraction network: The feature extraction network uses a TSN two-stream network for extraction;
[0018] S1302. Using a sparse multi-level boundary generator to locate actions in the video stream:
[0019] Benchmark feature extraction module: Input the output features of the TSN, and output the encoded action features after multiple layers of convolution;
[0020] Start and end generation module: Input the encoded features, and the output is the start and end position confidences, which is the same as the method in BMN;
[0021] Multi-scale confidence map feature generation module: Perform multi-scale convolution on the start and end position features of the action, and splice and fuse them in the form of head and tail Spans into confidence map features;
[0022] Sparse global feature fusion module: This module inputs the confidence map features, fuses the action information at different positions through the information under the sparse field of view, and differentiates the intermediate information of the action;
[0023] S1303. Output the segment confidence map. All the generated action start and end times and the segment confidence map pass through S-NMS, and finally generate the confidence sequence of the action;
[0024] S1304. Send all actions with confidences exceeding the preset threshold score into the action recognition network, which is a simple convolutional classifier, to generate action categories. The generable categories include: background, picking up the ball, serving, chopping, forehand drive, backhand slice, etc., more than 10 actions.
[0025] Preferably, the specific operation method of S105 is as follows:
[0026] When a specific action is detected and meets the requirement of switching the camera, the output camera will automatically switch channels. The priority order of the camera switching methods is: serving, celebration action - close-up of the athlete > rest, pause action - studio > picking up the ball action - side view > other actions - main camera position. The output camera will be sent for manual review, and the recording during the live broadcast is controlled within one minute.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] The present invention realizes the intelligent live broadcast of table tennis for the first time using deep learning technology. This method is optimized for the real-time performance of intelligent live broadcast, has a fast detection speed, is optimized for table tennis movements, has a higher recognition rate for table tennis actions, can greatly reduce the manpower when applied to the live broadcast of table tennis matches, and improve the speed of live broadcast. And this method can provide a reference paradigm for the intelligent live broadcast in other competition fields. Brief Description of the Drawings
[0029] Figure 1 This is a schematic flow diagram of the present invention; Detailed Embodiments
[0030] Next, the technical method will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0031] Embodiment
[0032] Please refer to Figure 1 , the present invention provides a technical method: a table tennis intelligent live broadcast technology method based on visual video analysis technology, including the following steps:
[0033] S101. Dataset construction: Collect the feature information of standard multi-camera high-definition live broadcast images in international and domestic table tennis competitions up to the current year, annotate the in-round racket-swinging actions of the athletes facing the camera based on the feature information, and then collect and form a dataset, and perform enhanced training on the annotated multi-camera lens dataset;
[0034] S102. Preprocessing: Convert the video into a 20 - 40s video stream and send it into the feature extraction network;
[0035] S103. Model construction: Include feature extraction network, temporal action localization and action recognition;
[0036] S104. Testing: Use the multi-scale confidence generation and sparse feature fusion techniques of temporal object localization to solve the two problems of inaccurate short-term action localization and difficult extraction of high-fine-grained features;
[0037] S105. Video logic processing, live broadcast video result.
[0038] Specifically, the specific processing method of S102 is as follows:
[0039] Use the main camera and the side camera for action localization and recognition. Send the videos of the main camera and the side camera into the network every 20 - 40s. The main camera is at 30 degrees obliquely above the back of the athlete, and the viewing angle includes the athlete and the table tennis table. The side camera is at 10 degrees obliquely above the front side, and the viewing angle includes the two athletes. The images are uploaded and processed simultaneously. There are also other cameras including the competition seat, the studio, and the close-up camera of the athlete, which are only used for camera switching and not for recognition.
[0040] Specifically, the specific steps of S103 are as follows:
[0041] S1301. Feature extraction using a feature extraction network: The feature extraction network uses a TSN two-stream network for extraction;
[0042] S1302. Using a sparse multi-level boundary generator to locate actions in the video stream:
[0043] Benchmark feature extraction module: Input the output features of the TSN, and output the encoded action features after multiple layers of convolution;
[0044] Start and end generation module: Input the encoded features, and the output is the confidence of the start and end positions, which is the same as the method in BMN;
[0045] Multi-scale confidence map feature generation module: Perform multi-scale convolution on the start and end position features of the action, and splice and fuse them in the form of head and tail spans into confidence map features;
[0046] Sparse global feature fusion module: This module inputs the confidence map features, fuses the action information at different positions through the information under the sparse field of view, and differentiates the intermediate information of the action;
[0047] S1303. Output the segment confidence map. All the generated start and end times of actions and the segment confidence map pass through S-NMS, and finally generate the confidence sequence of the action;
[0048] S1304. Send all actions with confidence scores exceeding the preset threshold into the action recognition network, which is a simple convolutional classifier, to generate action categories. The generable categories include: background, ball picking, serving, slicing, forehand drive, backhand chop, etc., more than 10 actions.
[0049] Specifically, the specific operation method of S105 is as follows:
[0050] When a specific action is detected and meets the requirement of switching the camera, the output camera will automatically switch channels. The priority order of the camera switching methods is: serving, celebration action - close-up of the athlete > rest, pause action - studio > ball picking action - side view > other actions - main camera position. The output camera will be sent for manual review, and the recording during the live broadcast is controlled within one minute.
[0051] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A table tennis intelligent live broadcast technology method based on visual video analysis technology, characterized in that, it includes the following steps: S101. Dataset construction: Collect the feature information of standard multi-camera high-definition live broadcast images in international and domestic table tennis competitions up to the current year. Based on the in-round racket-swinging actions of the athletes facing the camera in the feature information, annotations are made, and then collected to form a dataset, and the annotated multi-camera lens dataset is enhanced for training; S102. Preprocessing: Convert the video into a video stream of 20 - 40s and send it into the feature extraction network; S103. Model construction: It includes a feature extraction network, temporal action localization, and action recognition. The specific steps are as follows: S1301. Feature extraction using the feature extraction network: This feature extraction network uses the TSN two-stream network for extraction; S1302. Use a sparse multi-level boundary generator to locate actions in the video stream: Benchmark feature extraction module: Input the output features of TSN, and output encoded action features after multiple layers of convolution; Start and end generation module: Input the encoded features, and the output is the start and end position confidences, which is the same as the method in BMN; Multi-scale confidence map feature generation module: Perform multi-scale convolution on the start and end position features of the action, and splice and fuse them in the form of head and tail Spans into confidence map features; Sparse global feature fusion module: This module inputs the confidence map features, fuses action information at different positions through information under a sparse field of view, and differentiates the intermediate information of the action; S1303. Output the segment confidence map. All generated action start and end times and segment confidence maps go through S-NMS, and finally generate the confidence sequence of the action; S1304. Send all actions with confidences exceeding the preset threshold score into the action recognition network, which is a simple convolutional classifier, to generate action categories. The categories that can be generated include: background, picking up the ball, serving, chopping, forehand drive, backhand cut, and several other actions; S104. Testing: Use the multi-scale confidence generation and sparse feature fusion technical means of temporal object localization to solve the two problems of inaccurate short-time action localization and difficult extraction of high-fine-grained features; S105. Video logic processing, live broadcast video result.
2. The table tennis intelligent live broadcast technology method based on visual video analysis technology according to claim 1, characterized in that, the specific processing method of S102 is as follows: Use the main camera and the side camera for action localization and recognition. Send the videos of the main camera and the side camera into the network every 20 - 40s. The main camera is at a 30-degree oblique upward position behind the athlete, and the viewing angle includes the athlete and the table. The side camera is at a 10-degree oblique upward position on the front side, and the viewing angle includes the two athletes. The images are uploaded and processed simultaneously. There are also other cameras including the competition seat, the studio, and the close-up camera of the athlete, which are only used for camera switching and not for recognition.
3. The table tennis intelligent live broadcast technology method based on visual video analysis technology according to claim 1, characterized in that, the specific operation method of S105 is as follows: When a specific action is detected and the need to switch the camera is met, the output camera will automatically switch channels. The priority order of camera switching methods is as follows: serve, celebration action - close-up of the athlete > rest, pause action - studio > ball picking action - side view > other actions - main camera position. The output camera will be sent for manual review, and the recording during the live broadcast is controlled within one minute.