A Cloud-based Intelligent Table Tennis Training System Based on YOLOv5

Through the cloud table tennis smart training system based on YOLOv5, the Jetson nano and monocular camera combined with the improved YOLO v5 model and multi-threaded control, low-cost and efficient table tennis landing point detection and scoring are achieved, solving the problems of high cost and poor real-time performance of the existing system, and adapting to the intelligent training needs of different user levels.

CN116870444BActive Publication Date: 2025-07-25DALIAN NATIONALITIES UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310736473.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-20
Publication Date
2025-07-25
Estimated Expiration
2043-06-20

AI Technical Summary

Technical Problem

The existing table tennis training system is costly and has poor real-time performance, making it difficult to achieve full automation and accurately judge the table tennis landing point, especially in multi-camera systems, there are problems such as high hardware costs and complex synchronization.

Method used

The intelligent training system of cloud table tennis on-the-clock training system based on YOLOv5 is adopted, and the Jetson nano and monocular camera are used to combine the four-frame differential slope method and the polygon area determination method to perform table tennis landing detection and scoring, and real-time calculation and training strategy adjustment are carried out through the cloud server.

Benefits of technology

It realizes low-cost and efficient table tennis landing point detection and scoring, reduces hardware overhead, improves detection accuracy and real-time training, and adapts to the intelligent training needs of different user levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116870444B_ABST
    Figure CN116870444B_ABST
Patent Text Reader

Abstract

The present invention discloses a cloud-based intelligent table tennis training system based on YOLOv5, which relates to the field of artificial intelligence technology and is targeted at enterprise-level users. The enterprise-level users include a host computer module, a ball feeder module, a client module, a cloud-based table tennis landing point detection and scoring module, a function module for displaying detection results, an audio prompt interaction module, and a control module. The present invention uses an improved YOLO v5 model, combines the coordinate detection of table tennis balls, training scoring with multi-threaded timing signal interaction, and performs detection based on Jetson nano and a single industrial camera, saving a large amount of costs and being able to successfully complete closed-loop intelligent training. And aiming at the problem of low precision in small object detection, an exponential mean brightness equalization algorithm capable of automatically adjusting brightness and an improved YOLO v5 anchor generation mechanism are proposed, which not only improve the accuracy of model detection, but also can improve the accuracy of table tennis landing point detection, giving users a better experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a cloud-based intelligent table tennis training system based on YOLOv5. Background Art

[0002] Table tennis is widely pursued worldwide due to its simple equipment, strong fun, and diverse trajectory changes.

[0003] Traditional table tennis training is mainly manual training. However, manual training has high costs, consumes a large amount of manpower, and due to the characteristics of table tennis such as small size, fast flight speed, and complex motion models, it is difficult for the human eye to judge the landing point of the table tennis ball and the area to which the landing point belongs. Therefore, the difficulty of manual training is also very high, and it is difficult to review.

[0004] Today, with the rapid development of technology, artificial intelligence technology has gradually penetrated into production and life. Using the means of the computer vision field to replace the human eye for the capture and positioning of table tennis has become a hot topic in modern technology. Many companies at home and abroad have conducted research on this, such as the table tennis hawkeye system in China, which uses two or three hawkeye cameras to shoot the hitting video to perform three-dimensional reconstruction of the table tennis trajectory and analyze information such as the table tennis trajectory, landing point, and ball speed. However, multiple cameras also bring many other problems, such as increased costs, complex calibration of the cameras, synchronization of the cameras, etc. And a system using multiple cameras must have a very powerful processor to process more frames simultaneously, which will greatly reduce the real-time performance of the system. In addition, the existing semi-automated training system can only support two-person doubles, and there is no fully automated intelligent table tennis training system with a ball feeder. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a low-cost, high-efficiency, and multi-functional intelligent table tennis training system. By connecting the upper computer, the ball feeder, and the server through Jetson nano, and combining the detection of table tennis balls, the calculation of landing points, and the calculation of scores, the training strategy is updated in real time, which can save a large amount of costs while ensuring accuracy, and the landing point detection and intelligent training of table tennis balls can be completed through a monocular camera.

[0006] To this end, the present invention provides the following technical solutions:

[0007] The present invention provides a cloud-based intelligent table tennis training system based on YOLOv5, including:

[0008] Client: Responsible for signal interaction with the upper computer, the ball feeder, and the server, and transmitting the recorded images to the server in real time, and dynamically adjusting the training scoring strategy according to the results returned by the server;

[0009] Server: The cloud server acts as the server, receives images sent by the client, uses the table tennis detection model based on the improved YOLO v5 to perform table tennis detection on the cached images, detects the landing point of the table tennis ball, calculates the score, and returns the calculated data to the client.

[0010] Furthermore, the client is developed based on the Jetson Nano development board.

[0011] Furthermore, the recorded images are transmitted to the server in real time, including:

[0012] The recorded images are compressed and transmitted to the server in real time using the UDP protocol.

[0013] Furthermore, the calculation of the landing point of the table tennis ball and the area to which it belongs includes:

[0014] The four-frame differential slope method and polygonal area determination method are used to calculate the landing point of the table tennis ball and the area to which it belongs.

[0015] Furthermore, a table tennis detection model based on improved YOLO v5 is established, including:

[0016] Depthwise separable convolution is used in the YOLO v5 network structure;

[0017] Modify the anchor generation mechanism in the YOLO v5 network structure. The modifications include:

[0018] Count the frequency of the area and aspect ratio of the real annotation box;

[0019] The aspect ratio selects the interval value with the highest frequency, and the ping-pong ball area selects the middle and upper quartiles and is selected based on the input ratio. The two values are used as the main basis of the anchor mechanism;

[0020] Finally, four combinations of anchor sizes with aspect ratios of 1:1.2 and 1:1.6 and areas of 256 and 512 were determined.

[0021] Furthermore, the training scoring strategy is dynamically adjusted according to the results returned by the server, including:

[0022] determining a user-selected training mode;

[0023] In training mode, the score corresponding to each area on the trainer's side is dynamically adjusted according to the difficulty of the serving and returning areas, and the serving area and the area where the returning ball lands in each round are calculated to calculate the total score;

[0024] Analyze the data and total score sent back by the server, and use audio to provide the user with training round status, training round optimization suggestions, and other prompt information during training.

[0025] Furthermore, the training modes include: basic mode, defensive mode, offensive mode, and star mode;

[0026] After the training, the training situation is given and scored, including:

[0027] In the basic mode, the user independently adjusts the serving speed, spin direction, serving angle, and serving interval according to their own training situation, which is convenient for the user to conduct targeted training. After each training session, the scoring situation is uniformly given;

[0028] In the defensive mode, the table on the trainer's side is divided into three areas: I, II, and III. In the first round of serving, the ball machine serves the balls to areas I, II, and III in sequence. After the serving is completed, the scoring situations of the three areas are calculated and sorted respectively. In the next round of targeted training, the balls are first served to the area with the lowest score and m1 more balls are served than in the previous round, m2 more balls are served to the second area, m1 > m2, and so on;

[0029] In the offensive mode, the ball machine serves the balls randomly in the ratio of straight ball: spin ball: spin ball + corner ball = 3:2:1. At the same time, the sound feedback system tells the trainer the nine-square grid area for the return ball in this round. The trainer has three scoring methods: getting full marks if the return ball lands in the designated area in the round; getting 1 point if the return ball does not land in the designated area but lands on the table; getting no points if the return ball does not land on the table;

[0030] In the star mode, the system records the hitting methods and characteristics of domestic and foreign star players. The user independently selects different star players, and the ball machine simulates the serving of star players. At the same time, the sound feedback system tells the trainer the return ball area, and the score is calculated according to the scoring method of the offensive mode.

[0031] Furthermore, on the client side, the method of multi-thread control is used to conduct data interaction with the host computer, the ball machine, and the cloud server respectively, and multi-threads are used to monitor each serial port and Bluetooth to obtain the mode signal of the host computer, send the serving signal to the ball machine, and upload the locally recorded video to the cloud server.

[0032] Furthermore, it also includes: detecting the positions of the four corner points of the table through the Harris corner detection algorithm to establish an overall model of the table.

[0033] Furthermore, it also includes: interference area cropping and noise elimination; including:

[0034] Assuming that the contour area interval of the table tennis ball measured in multiple experiments is [T1, T2], then non-table tennis objects with contour areas outside the interval [T1, T2] are discarded;

[0035] Let A iIt is the set of contour coordinates of all objects extracted after denoising in the i-th frame image, and the relevant mathematical model is as follows:

[0036]

[0037]

[0038] Among them, is the contour area of the n-th object extracted in the i-th frame image, is the set of contour coordinates of the n-th object.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] The intelligent table tennis training system in the present invention is composed of a low-frame-rate camera, a Jetson nano, a ball feeder, and a rented cloud server. First, on the basis of traditional ball feeder training, a Jetson nano and a camera are added for video recording and video transmission, and the landing point of the table tennis ball is detected. It can send corresponding mode signals according to the needs of users. After the signals are decomposed by the Jetson nano, the ball feeder is controlled to serve. It can intelligently customize the serving strategy according to the training situation of users and display the landing point situation in real time. The main application advantages are that it can save the costs of sensors, cameras, controllers, etc., train targeted according to the user's level, and is convenient to train without a coach. The main process can be simplified as follows: after the user issues a training mode instruction, the ball feeder conducts a round of serving in the corresponding mode, and the results are displayed to the user after each round of serving. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.

[0042] Figure 1 is the overall structure diagram of the cloud-based intelligent table tennis training system according to the embodiment of the present invention;

[0043] Figure 2 is the method flow chart of cloud-based landing point detection and scoring according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0046] The present invention provides a cloud-based intelligent table tennis training system based on YOLOv5. By using a multi-thread control method on the client side, data interaction is carried out with the host computer, the ball feeder, and the cloud server respectively. On the client side, multi-thread technology is used to obtain the mode signal of the host computer, send the serving signal to the ball feeder, and upload the locally recorded video to the cloud server, etc. The development of the cloud-based intelligent table tennis training system based on YOLO v5 is completed by using the Bluetooth, serial port data interaction method and the PyTorch framework. This system is a system based on the Linux platform. By modifying the startup initialization file of the Linux system, the project files can be started automatically when the computer boots up, and users can directly operate the host computer to conduct table tennis training. This system mainly uses the client to control the ball feeder and detect the landing point of the table tennis ball in the cloud video frame. The system can display the result data information after training in the form of audio and text, and give optimized training opinions.

[0047] As Figure 1 shown, the embodiments of the present invention provide a cloud-based intelligent table tennis training system based on YOLOv5, which is designed for enterprise-level user terminals. Enterprise-level users include a host computer, a ball feeder, a client, and a cloud server. This system can be used for table tennis training in remote ethnic areas. The system includes:

[0048] Client: The Jetson Nano development board serves as the client and signal center, responsible for signal interaction with the host computer, ball dispenser, and server. It compresses the recorded images and uses the UDP protocol to transmit them to the server in real time. When the round ends but the calculation result from the server has not been returned, the ball dispenser makes relaxed and random ball serves, which are not counted towards the training score. The client dynamically adjusts the training scoring strategy based on the results sent back by the server.

[0049] Jetson Nano is an embedded development board based on a GPU processor. Its ARM side is transplanted and loaded with the Ubuntu 18.04 LTS system, supporting popular AI frameworks and algorithms such as TensorFlow, Caffe / Caffer2, PyTorch, and Keras. It also includes the operating environment of the GPU such as OpenGL and Opencv, as well as the corresponding call interfaces.

[0050] Specifically, the client includes:

[0051] Training module: Basic training mode, defensive training mode, offensive training mode, star training mode. The ball dispenser needs to be turned on first, and after its self-check is completed, the training mode can be selected.

[0052] Scoring module: Dynamically adjusts the scores corresponding to each area on the trainer's side according to the difficulty of the ball serving and return areas, calculates the ball serving area and the area where the ball return landing point is located in each round, and calculates the total score.

[0053] Audio interaction module: Jetson Nano analyzes the data sent back by the cloud server and the score calculated by the scoring module, and uses audio to give the user the training situation of each round, suggestions for optimized training in each round, and other prompt information during training.

[0054] Cloud server: The cloud server serves as the server, asynchronously receives and decompresses the images transmitted by the client using the UDP protocol, uses the YOLO v5 detection algorithm to detect table tennis balls in the cached images, uses the four-frame differential slope method and the polygon area determination method to calculate the landing point of the table tennis ball and the area it belongs to, and returns the calculated data to the client.

[0055] YOLOv5 is a single-stage object detector that can detect multiple objects in images or videos in real time. Its main features are high speed and high accuracy, and it supports object detection on images of different sizes and resolutions. In terms of network architecture, it uses a backbone network based on FPN (Feature Pyramid Network), in which the CSP (Cross Stage Partial) module is used to improve the effect of feature maps. The SPP (Spatial Pyramid Pooling) module and PAN (Path Aggregation Network) module are used to better process objects of different sizes. In addition, the Mosaic data augmentation method is introduced, which combines multiple different images into a large image to expand the dataset and improve the generalization ability of the model. Moreover, it uses Focal Loss similar to RetinaNet to alleviate the problem of sample imbalance and further improve the accuracy of the model. This system uses the lightweight network ResNet as the backbone network. For different feature extraction networks, comparative experiments are conducted on its own dataset multiple times. It can be found that ResNet not only has high accuracy but also is much faster than other feature extraction networks in terms of detection speed.

[0056] The detection part of YOLO v5 can be mainly divided into four modules:

[0057] (1) backbone. That is, the feature extraction network, which is used to extract features and includes the CBL module, Focus module, CSP module, and SPP module. Through a group of conv+relu+pooling layers, the feature maps of the image are extracted to shrink the feature maps. To improve the detection speed, the backbone selected by this system is ResNet. ResNet is constructed using a deep residual network, which can ensure that deeper networks can be trained without gradient disappearance and is a relatively lightweight network, which can minimize the computational amount and memory occupancy while ensuring high detection accuracy.

[0058] (2) Focus structure. The Focus layer is a special downsampling method in YOLO v5. Its main purpose is to reduce the number of layers, reduce parameters, reduce FLOPS, reduce CUDA memory, increase the forward and backward speeds, and at the same time minimally affect the mAP. Its operation is similar to nearest neighbor downsampling, which converts the information on the w-h plane to the channel dimension and then extracts different features through 3*3 convolution. Using this method can reduce the information loss caused by downsampling.

[0059] (3) Head network. The head obtains features from different levels of the backbone, processes them, and finally obtains prediction results from 3 feature maps respectively.

[0060] (4) Neck module. The neck interpolates into the feature map in the way of upsampling through Upsample, making the scale of the feature map larger, so as to fuse the feature maps from the backbone and perform upward fusion of features, and the feature map keeps getting larger; then continue to perform downsampling. One is to obtain feature maps of different scales, and the other is to better fuse the graphic features of the shallow layer and the semantic features of the deep layer, rather than just simple concat. The role of the neck layer is to combine the graphic features of the shallow layer and the semantic features of the deep layer to obtain more complete features.

[0061] (5) CSP structure. CSPNet mainly splits the feature map into two parts, one part performs convolution operations, and the other part is concated with the result of the convolution operation of the previous part. In classification problems, using CSPNet can reduce the computational amount, but the accuracy improvement is very small; in object detection problems, using CSPNet as the Backbone brings a relatively large improvement, which can effectively enhance the learning ability of CNN and also reduce the computational amount. YOLO v5 designs two CSP structures. The CSP1_X structure is applied to the Backbone network, and the CSP2_X structure is applied to the Neck network.

[0062] The detection dataset constructed by this system simulates the situation under a fixed angle of the camera (the camera uses an undistorted industrial camera with a frame rate of 60fps and a resolution of 640*480, and the ball dispenser is a self-developed ball dispenser. The camera is 166 cm above the ground and 75 cm away from the edge of the table). The camera on the Jetson nano records the video and transmits it to the cloud in real time. The detection model on the cloud server detects the ball hit by the athlete and analyzes the landing point. The detection mainly obtains the coordinate information of the table tennis ball in the video frame. According to the serving area and the area where the returned ball lands, calculate the score of this round of table tennis training, and then send the training situation and the training score from the cloud to the client. The voice outputs and the host computer displays the training situation of this round to prompt the user. This system combines the client and the server side, which not only greatly reduces the hardware cost, but also can achieve near-real-time detection and reduce the time for users to wait for the results.

[0063] Such as Figure 2As shown in the figure, the working method of the system in the above embodiment is as follows: First, an improved YOLO v5 detection algorithm is used to establish a table tennis detection model; then an adaptive brightness equalization algorithm is proposed and a brightness equalization module is added; then the Hough corner detection is used to establish a table tennis table corner area model, and the accuracy of table tennis landing point detection is improved by cropping the interference area, and finally real-time detection is achieved. Specifically, it includes the following aspects:

[0064] 1. Establish a table tennis detection model:

[0065] Modify the backbone network of the YOLO v5 network so that the network is dedicated to realizing the detection of table tennis. Use this network for training multiple times, select the best model weights, and save the detection model and weights. Record videos from the client and upload them to the cloud server in real time using the UDP protocol. Put the saved model on the cloud server for prediction. The results such as the landing rate and score obtained from the prediction are returned to the client through the UDP protocol, and the results are broadcast in audio form on the client. The upper computer presents the table tennis training situation in real time (including the given landing point area, the actual landing point area, the table tennis landing rate, and the number of serves). Table tennis training equipment is deployed on corresponding edge devices with limited computing resources according to its different functions. It is crucial to effectively control the model parameters, computational volume, memory access volume, etc. of object detection at the edge with limited computing resources. In order to improve the speed and accuracy of the model and reduce the number of model parameters, this design focuses on the research of model lightweighting. Lightweight networks mainly include knowledge distillation, weight quantization, pruning operation compression models, and directly training lightweight network models. The present invention directly designs a lightweight network based on YOLO v5 through depthwise separable convolution for the detection of table tennis. In order to further improve the speed of the model, modify the anchor generation mechanism in the YOLO v5 network structure. The main modification steps are as follows: 1. Statistically count the frequency of the area and aspect ratio of the real annotation box; 2. Select the interval value with the highest frequency of occurrence for the aspect ratio, and select the median and upper quartile for the table tennis area and scale according to the input ratio. These two values are used as the main basis for the anchor mechanism; 3. Finally, determine the anchor sizes of four combinations with aspect ratios of 1:1.2 and 1:1.6 and areas of 256 and 512.

[0066] In the specific implementation, the upper computer first sends a training mode signal to the client. After the client records and uploads video frames to the server in real time, the server performs table tennis detection, and the results show that the detection can be accurately completed.

[0067] 2. Calculate the table tennis landing point:

[0068] By observing multiple sets of table tennis trajectories, it is found that the direction of the table tennis ball changes suddenly after the landing point. The signs of the slopes formed by the centroid coordinates of the table tennis ball in the two frames before the landing point are opposite to those in the two frames after the landing point. The intersection coordinate of the two lines where the slopes are located is the landing point coordinate of the table tennis ball. The slope models formed by the centroid coordinates of the table tennis ball in the two frames before and after the landing point are shown in the following formula:

[0069]

[0070] Among them, respectively represent the centroid coordinates of the table tennis ball in the (i - 3)th, (i - 2)th, (i - 1)th, and ith frames of the image. k i-2 and k i respectively represent the slopes of the lines formed by the centroid coordinates of the table tennis ball in the (i - 3)th and (i - 2)th frames of the image, and in the (i - 1)th and ith frames of the image. The slope k i-2 and k i The corresponding linear equations are shown in the following formula.

[0071]

[0072] Among them, l i-2 represents the line passing through the centroid coordinate and with a slope of k i-2 , and l i represents the line passing through the centroid coordinate and with a slope of k i .

[0073] In the dataset, the moving direction of the valid table tennis ball is the positive direction of the x-axis. When k i-2 < 0 and k i > 0, the moving direction of the table tennis ball changes suddenly, and the intersection of the lines l i-2 and l i can be represented as the landing point, as shown in the following formula.

[0074]

[0075] (O x , O y ) is the landing point coordinate of the table tennis ball in this round.

[0076] 3. Scoring algorithm:

[0077] The table tennis intelligent training system has a total of four training modes, namely the basic mode, the defensive mode, the offensive mode, and the star mode, and the training difficulty increases in turn. During the training, the camera automatically records the side of the table tennis table of the ball feeder and transmits it to the Jetson nano for real-time table tennis detection and landing point calculation. After the training, the training situation is given and scored.

[0078] In the basic mode, users can independently adjust variables such as serving speed, spin direction, serving angle, and serving interval according to their own training conditions, which is convenient for users to conduct targeted training. After each training session, the scoring situation is uniformly given.

[0079] In the defensive mode, the table on the trainer's side is divided into three areas: I, II, and III. When serving the ball for the first round, the ball serving machine serves the ball to areas I, II, and III in sequence. After the serving is completed, the scoring situations of the three areas are calculated and sorted respectively. For the next round of targeted training, the ball is first served to the area with the lowest score and m1 more balls are served than in the previous round, m2 more balls are served to the second area (m1 > m2), and so on.

[0080] In the offensive mode, the ball serving machine serves the ball in a random order according to the ratio of straight ball: spin ball: spin ball + corner ball = 3:2:1. At the same time, the sound feedback system tells the trainer the nine-square grid area for the return ball in this round. The trainer has three scoring methods: 1) If the return ball lands in the designated area, the round gets a full score; 2) If the return ball does not land in the designated area but lands on the table, 1 point is scored; 3) If the return ball does not land on the table, no score is obtained.

[0081] In the star mode, the system enters the hitting methods and characteristics of domestic and foreign star players. Users independently select different star players, and the ball serving machine simulates the serving of star players. At the same time, the sound feedback system tells the trainer the return ball area, and the score is calculated according to the scoring method of the offensive mode.

[0082] 4. Harris corner detection algorithm:

[0083] In the embodiment of the present invention, the positions of the four corner points of the table are detected by the Harris corner detection algorithm. For each pixel (x, y), within the (blockSize x blockSize) neighborhood, the covariance matrix M(x, y) of the gradient map is calculated, and then the result map is obtained through the corner response function in the second step above. The corner points in the image can be the local maximum values of the result map. The main formulas for Harris corner detection are as follows:

[0084] E(u, v) = ∑ x,y w(x, y)[I(x + u, y + v) - I(x, y)] 2 ;

[0085] det(M) = λ1λ2;

[0086] trace(M) = λ1 + λ2;

[0087] R = det(M) - k(traceM) 2 ;

[0088] (u, v) represents the point to be determined currently, and its local window is K(u,v) , where \(u\) and \(v\) are the offsets in the horizontal and vertical directions respectively; \(I(x, y)\) represents the brightness value of the image \(I\) at the point \((x, y)\); \(w(x, y)\) is the local window weighting function, and corner points are detected by judging \(R>0\).

[0089] 5. Interference area cropping and noise elimination:

[0090] To avoid false detection, the serving machine and the ball basket areas in the video frame are masked, eliminating the possibility of false detection at the ball basket. Assuming that the contour area interval of the table tennis ball measured in multiple experiments is \([T1, T2]\), non-table tennis objects with contour areas outside the interval \([T1, T2]\) are discarded. Let \(A\) i be the set of contour coordinates of all objects extracted after noise reduction in the \(i\)-th frame image, and the related mathematical model is as follows:

[0091]

[0092]

[0093] Among them, is the contour area of the \(n\)-th object extracted in the \(i\)-th frame image, is the set of contour coordinates of the \(n\)-th object.

[0094] 6. Combination of table tennis detection and landing point detection and multi-thread control:

[0095] The table tennis intelligent training system on the cloud uses the method of multi-thread control on the client side to interact with the upper computer, the serving machine and the cloud server respectively, and uses multi-threads to monitor each serial port and Bluetooth, so as to obtain the mode signal of the upper computer, send the serving signal to the serving machine and upload the locally recorded video to the cloud server, etc. The multi-thread control in the embodiments of the present invention includes the following steps:

[0096] S1. Create multiple threads: The processing unit creates multiple threads, and each thread can execute different tasks simultaneously. The tasks in the present invention mainly include the task of monitoring the serial port ttyUSB0, the task of monitoring the Bluetooth serial port ttyUSB1, the task of UDP receiving and sending video frame data, and the audio control task.

[0097] S2. Allocate tasks: The processing unit allocates different tasks to different threads so as to execute multiple tasks simultaneously.

[0098] S3. Thread synchronization: The processing unit can use a synchronization mechanism to ensure the correctness of data exchange and operations between different threads, and avoid data competition and deadlocks.

[0099] S4. Thread Pool Management: The processing unit can maintain a thread pool and dynamically adjust the number of threads to make full use of the computing resources of the processing unit.

[0100] S5. Exception Handling: The processing unit can capture and handle exceptions that occur during the execution of threads, avoiding the collapse of the entire system due to exceptions in individual threads.

[0101] S6. Thread Destruction: The processing unit can destroy threads after tasks are completed to release resources.

[0102] Compared with traditional single-thread control, the multi-thread control in the embodiments of the present invention can improve the computing efficiency and response speed of the system, better allocate the computing tasks of the client, enable the mechanical components of the cloud-based table tennis intelligent training system to work in coordination, enhance the stability and reliability of the system, and better meet the requirements of large-scale data processing and complex mechanical control algorithms of the system.

[0103] The cloud-based table tennis intelligent training system relies on the server to detect the landing points of table tennis balls and calculate scores, and returns the ball landing point data and scoring situation of one round to the client. The client conducts targeted training according to the landing rate of different regions, so that the ball feeder focuses on serving balls to the regions with a low landing rate. In order to reduce hardware costs, research is carried out on table tennis detection and landing point detection under the condition of a single camera. Calculate the queue of table tennis coordinates detected by YOLO v5 in the cloud to determine whether there is a landing point. If there is, record the landing point coordinates and the corresponding round, and return the values to the client, which are sent to the host computer for real-time display. If there is no landing point, set the landing point coordinates of this round to (-1, -1), and the host computer remains unchanged. After all the balls in this round are served, the server calculates the scoring situation of this round and returns it to the client. The client plays the scoring situation in the form of audio and sends the total score to the host computer for display.

[0104] In several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0105] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0106] In addition, each functional unit in various embodiments of the present invention may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0107] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0108] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of various embodiments of the present invention.

Claims

1. A cloud-based intelligent table tennis training system based on YOLOv5, characterized in that, include: Client: Responsible for signal interaction with the host computer, ball machine and server, and real-time transmission of recorded images to the server, and dynamic adjustment of training scoring strategy based on the results sent back by the server; Server: The cloud server acts as the server, receives the image sent by the client, uses the table tennis detection model based on the improved YOLO v5 to perform table tennis detection on the cached image, detects the landing point of the table tennis ball, calculates the score, and returns the calculated data to the client; Among them, a table tennis detection model based on improved YOLO v5 is established, including: Depthwise separable convolution is used in the YOLO v5 network structure; Modify the anchor generation mechanism in the YOLO v5 network structure. The modifications include: Count the frequency of the area and aspect ratio of the real annotation box; The aspect ratio selects the interval value with the highest frequency, and the ping-pong ball area selects the middle and upper quartiles and is selected based on the input ratio. The two values are used as the main basis of the anchor mechanism; Finally, four combinations of anchor sizes with aspect ratios of 1:1.2 and 1:1.6 and areas of 256 and 512 were determined.

2. The cloud-based intelligent table tennis training system based on YOLOv5 according to claim 1, wherein, The client is developed based on the Jetson Nano development board.

3. The cloud-based intelligent table tennis training system based on YOLOv5 according to claim 1, wherein, Transmit recorded images to the server in real time, including: The recorded images are compressed and transmitted to the server in real time using the UDP protocol.

4. The cloud-based intelligent table tennis training system based on YOLOv5 according to claim 1, wherein Calculate the point where the table tennis ball lands and the area it belongs to, including: The four-frame differential slope method and polygonal area determination method are used to calculate the landing point of the table tennis ball and the area to which it belongs.

5. The cloud-based intelligent table tennis training system based on YOLOv5 according to claim 1, wherein, Dynamically adjust the training scoring strategy based on the results returned by the server, including: determining a user-selected training mode; In the training mode, the score corresponding to each area on the trainee's side is dynamically adjusted according to the difficulty of the serving and returning areas, and the serving area and the area where the returning ball lands in each round are calculated to calculate the total score; The data returned by the server and the total score are analyzed, and audio is used to provide the user with training round status, training round optimization suggestions, and other prompt information during training.

6. The cloud-based intelligent table tennis training system based on YOLOv5 according to claim 5, characterized in that, The training modes include: basic mode, defense mode, offense mode and star mode; After the training, the training status will be given and scored, including: In the basic mode, users can adjust the serve speed, spin direction, serve angle, and serve interval according to their own training conditions, which is convenient for users to train in a targeted manner. After each training session, the score will be given uniformly. In the defensive mode, the table on the trainee's side is divided into three areas: I, II, and III. In the first round of serving, the serving machine serves balls to the three areas in order. After serving, the scores of the three areas are calculated and ranked. In the next round, targeted training is carried out. First, serve to the area with the lowest score and serve m1 more balls than in the previous round. Then serve m2 more balls to the second area, m1>m2, and so on. In the offensive mode, the ball feeder serves balls randomly in the ratio of straight ball: spin ball: spin ball + corner ball = 3:2:

1. Meanwhile, the sound feedback system tells the trainer the nine-square grid area for the return ball in this round. The trainer has three ways to score: getting a full score if the return ball lands in the designated area; getting 1 point if the return ball does not land in the designated area but hits the table; getting no score if the return ball does not hit the table. In the star mode, the system records the hitting styles and characteristics of domestic and foreign star players. The user can independently select different star players. The ball feeder simulates the star players' serves. Meanwhile, the sound feedback system tells the trainer the return ball area, and the score is calculated according to the scoring method of the offensive mode.

7. The intelligent table tennis training system based on YOLOv5 on the cloud according to claim 1, characterized in that, On the client side, the method of multi-thread control is used to interact with the host computer, the ball feeder, and the cloud server respectively. Multi-threads are used to monitor each serial port and Bluetooth to obtain the mode signal from the host computer, send the serving signal to the ball feeder, and upload the locally recorded video to the cloud server.

8. The intelligent table tennis training system based on YOLOv5 on the cloud according to claim 1, characterized in that, It also includes: Detect the positions of the four corner points of the table tennis table through the Harris corner detection algorithm to establish an overall model of the table tennis table.

9. The cloud-based intelligent table tennis training system based on YOLOv5 according to claim 1, characterized in that, It also includes: Interference area cropping and noise elimination; including: Assume that the contour area interval of the table tennis ball measured in multiple experiments is [T1, T2]. Then, non-table tennis ball objects with contour areas outside the interval [T1, T2] are discarded. Let A i be the set of contour coordinates of all objects extracted after noise reduction in the i-th frame image. The relevant mathematical model is as follows: Among them, is the contour area of the nth object extracted from the ith frame image, is the set of contour coordinates of the nth object.

Citation Information

Patent Citations

  • Ping-pong ball drop point identification and scoring method and system based on target detection and tracking

    CN110458100A

  • Table tennis ball drop point detection method and system based on improved color gamut recognition technology

    CN114387354A

  • Auxiliary training system and method for ping-pong sports

    CN114618142A