Auxiliary penalty method based on artificial intelligence and live-action eagle eye system

Through an assisted penalty method based on artificial intelligence, a single camera and a lightweight multi-task joint learning network model is used to solve the problem of high cost and insufficient real-time performance of the Hawkeye system, high-precision tennis penalty is achieved, and the popularization of the Hawkeye system in the field of national fitness is promoted.

CN120339906AActive Publication Date: 2025-07-18ORANGE LION SPORTS (ZHEJIANG) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510396300.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-18
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing Hawkeye system is expensive and lacks real-time performance, making it difficult to popularize in the field of national fitness, and the detection accuracy of ordinary camera solutions is difficult to ensure.

Method used

Using an auxiliary penalty method based on artificial intelligence, a single camera is used to detect tennis positions and landing points with lightweight multi-task joint learning network model, and a computer vision algorithm is used to identify tennis contours and field boundaries to generate penalty results.

Benefits of technology

It reduces system costs, improves detection accuracy and real-time performance, and realizes the popularization of Hawkeye system in the field of national fitness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339906A_ABST
    Figure CN120339906A_ABST
Patent Text Reader

Abstract

The invention provides an auxiliary penalty method based on artificial intelligence and a real scene eagle eye system, the method comprises the following steps: obtaining continuous image frames collected by a camera device in real time, the image frames comprising a tennis court and a tennis ball located in the tennis court; detecting a tennis ball position and a falling point position of the tennis ball in a tennis ball ground frame from the continuous image frames; taking a falling point position of a tennis ball as a center in a tennis ball ground frame, respectively taking N pixel points leftwards and rightwards, and respectively taking M pixel points upwards and downwards to form a region of interest; identifying a tennis ball contour and a field boundary line in the region of interest; a penalty result is generated according to the position relation between the tennis ball contour and the field boundary line. According to the eagle eye system, the precision can be guaranteed while the system cost is reduced, the real-time performance is improved, and the eagle eye system is of great significance in popularization in the national fitness field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sports equipment, and in particular, to an assisted penalty method and a live hawkeye system based on artificial intelligence. Background Art

[0002] "Hawkeye" is also known as an instant replay system, which is used to assist referees in making accurate penalties. Professional hawkeye systems such as those used in ATP, WTA, and the four Grand Slams usually use 8-10 high-speed cameras and complete the judgment and display of results within 10 seconds. Although the accuracy is extremely high, due to the high cost of high-speed cameras themselves, a set of systems can cost hundreds of thousands to millions. In addition, simultaneously calculating 8-10 channels of high-frame and high-resolution data also results in high latency. Therefore, it is costly and lacks real-time performance. Lightweight solutions based on video analysis on the market use ordinary cameras and open-source object detection algorithms to detect tennis balls. Although the cost is reduced, it is difficult to guarantee the accuracy of detecting the tennis ball trajectory at high resolutions. Therefore, for referees, players, and audiences, the hawkeye results are difficult to be persuasive.

[0003] Therefore, how to reduce the system cost while ensuring the detection accuracy and improving the real-time performance is of great significance for popularizing the hawkeye system in the field of national fitness. Summary of the Invention

[0004] In view of the above problems, the present invention is proposed to provide an assisted penalty method and a live hawkeye system based on artificial intelligence that solve the above technical problems or at least partially solve the above technical problems.

[0005] In one aspect of the present invention, an assisted penalty method based on artificial intelligence is provided. The method includes:

[0006] Obtain continuous image frames collected in real time by a camera device, where the image frames include a tennis court and a tennis ball located within the tennis court;

[0007] Detect the position of the tennis ball and the landing position of the tennis ball in the tennis ball bounce frame from the continuous image frames;

[0008] Taking N pixel points to the left and right respectively, and M pixel points up and down respectively with the landing position of the tennis ball as the center in the tennis ball bounce frame to form a region of interest;

[0009] Identify the tennis ball contour and the court boundary line in the region of interest;

[0010] Generate a penalty result according to the positional relationship between the tennis ball contour and the court boundary line.

[0011] Further, detecting the position of the tennis ball and the landing position of the tennis ball in the tennis ball bounce frame from the continuous image frames includes:

[0012] Use a preset lightweight multi-task joint learning network model to detect the position of the tennis ball in consecutive image frames and the landing position of the tennis ball in the tennis bounce frame;

[0013] The multi-task joint learning network model includes an encoder, a feature enhancement network, and two task heads, namely a tennis detection task head network and a landing point recognition task head network; among them, the encoder is used to extract tennis features of different scales from the input image frames, the feature enhancement module is used to fuse the tennis features of different scales to obtain the tennis fusion features of semantic information and edge texture information, the tennis detection task head network is used to identify the position of the tennis ball in the image frame according to the tennis fusion features, and the landing point recognition task head network is used to identify the landing position of the tennis ball in the image frame according to the tennis fusion features and determine the tennis bounce frame.

[0014] Furthermore, the encoder includes a first downsampling network layer and a second downsampling network layer connected in sequence;

[0015] The first downsampling network layer consists of a 3×3 standard convolutional layer and a 3×3 max pooling layer, and is used to extract the primary feature map of the image to achieve preliminary dimensionality reduction;

[0016] The second downsampling network layer includes a downsampling module and at least three feature transformation modules. The downsampling module is used to extract tennis features of different scales from the primary feature map, including a first branch composed of 3×3 depthwise separable convolutions arranged in parallel and a second branch composed of a 1×1 standard convolution, a 3×3 depthwise separable convolution, and a 1×1 standard convolution connected in sequence. The feature output channels of the first branch and the second branch are mixed; each feature transformation module is used to perform feature transformation on the tennis features of different scales obtained by the downsampling module, including a third branch with an identity mapping function arranged in parallel and a fourth branch composed of a standard 1×1 convolution, a 3×3 depthwise separable convolution, and a 1×1 standard convolution connected in sequence. After mixing the feature output channels of the downsampling module, they are evenly divided into two parts, which are used as the inputs of the third branch and the fourth branch respectively, and the feature output channels of the third branch and the fourth branch are mixed.

[0017] Furthermore, the feature enhancement network includes a first upsampling network layer and a second upsampling network layer connected in sequence;

[0018] The first upsampling network layer includes a 3×3 RepVGG convolutional layer and a first upsampling module. The 3×3 RepVGG convolutional layer is used to perform enhanced local feature extraction on the output features of the second downsampling network layer. The first upsampling module is used to upsample the output features of the RepVGG convolutional layer using bilinear interpolation. The output channels of the upsampled features are concatenated with the feature input channels of the second downsampling network layer to obtain 72 feature output channels.

[0019] The second upsampling network layer includes a 3×3 standard convolutional layer and a second upsampling module. The 3×3 standard convolutional layer is used to perform feature extraction on the output features of the 72 feature output channels of the first upsampling network layer and convert them into 48 feature output channels for output. The first upsampling module is used to upsample the currently input features using bilinear interpolation. The output channels of the upsampled features are concatenated with the feature output channels of the first downsampling network layer to obtain 72 feature output channels.

[0020] Further, the tennis detection task head network includes a third upsampling module, a 3×3 RepVGG convolutional layer, and a 3×3 standard convolutional layer connected in sequence. The third upsampling module is used to upsample the output features of the feature enhancement network using bilinear interpolation. The 3×3 RepVGG convolutional layer is used to perform enhanced local feature extraction on the output features of the third upsampling module. The 3×3 standard convolutional layer is used to transform the 72-channel input feature channels into 3-channel feature output channels.

[0021] Further, the landing point recognition task head network includes three groups of 3×3 RepVGG convolutional layers, a max pooling layer, three groups of 3×3 RepVGG convolutional layers, a channel attention layer, and a fully connected layer arranged in series. By establishing a channel attention mechanism in the landing point recognition task head network, the channel weights of each feature input channel of the landing point recognition task head network are dynamically adjusted according to the channel attention mechanism to enhance the attention of the landing point recognition task head network to the motion features of the tennis in the image, and the landing position of the tennis in the image frame is output.

[0022] Further, after obtaining a series of consecutive image frames collected in real time by the camera device, the method further includes:

[0023] Merge a preset number of consecutive image frames in the channel dimension to achieve pre-fusion of the image frames.

[0024] Further, the method further includes:

[0025] The positions of the tennis balls detected from multiple sets of consecutive image frames and the landing positions of the tennis balls in the tennis ball bounce frames are plotted on a preset real - scene field map to form a dynamic video or static image corresponding to the tennis ball landing event, marked with the pre - bounce movement trajectory, landing position, and post - bounce movement trajectory of the tennis ball.

[0026] Further, the identifying the tennis ball contour and the field boundary line in the region of interest includes:

[0027] Obtain at least a third image frame whose acquisition time is earlier than the tennis ball bounce frame as a background frame, and select the region of interest of the background frame in the same way with the coordinate of the landing position of the tennis ball in the tennis ball bounce frame as the center;

[0028] Convert the regions of interest of the tennis ball bounce frame and the background frame into grayscale images through color conversion operations;

[0029] Use the frame difference method to calculate the absolute difference between the regions of interest of the tennis ball bounce frame and the background frame to obtain a local difference image, and binarize the local difference image after taking the absolute difference;

[0030] Erode the binarized image, and perform contour detection on the eroded image to obtain the tennis ball contour;

[0031] Shrink the width and height of the region of interest of the tennis ball bounce frame in a preset ratio in the tennis ball bounce frame to obtain a rectangular sideline region;

[0032] Convert the sideline region into a grayscale image;

[0033] Binarize the grayscale image and perform dilation, and perform contour detection on the dilated image to obtain the field boundary line contour.

[0034] Another aspect of the present invention provides an artificial - intelligence - based real - scene hawkeye system, and the system includes:

[0035] At least one camera device for photographing a tennis court and tennis balls and human bodies located in the tennis court to collect live images in real time;

[0036] A server, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the steps of the method according to any one of claims 1 - 9.

[0037] The auxiliary penalty judgment method and live eagle eye system based on artificial intelligence provided by the embodiments of the present invention can realize the auxiliary penalty judgment of the eagle eye system only by using a single camera in combination with computer vision algorithms. Compared with the scheme of multiple high-speed cameras, the installation and deployment cost of the eagle eye system is greatly reduced; compared with the schemes of other computer vision algorithms, the present invention simultaneously learns tennis detection and landing point recognition. By combining tennis detection and landing point recognition, the time and position accuracy of landing point recognition are effectively improved, the tennis trajectory is more reliable, and the accuracy of eagle eye recognition is effectively guaranteed. The present invention can ensure the accuracy while reducing the system cost and improving the real-time performance, which is of great significance for realizing the popularization of the eagle eye system in the field of national fitness.

[0038] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the description. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. In the drawings:

[0040] Figure 1 is a flowchart of the auxiliary penalty judgment method based on artificial intelligence according to the embodiment of the present invention;

[0041] Figure 2 is a schematic diagram of the output of the tennis detection task head in the embodiment of the present invention;

[0042] Figure 3 is a schematic diagram of the output of the eagle eye system algorithm in the embodiment of the present invention;

[0043] Figure 4 is a schematic diagram of 3D virtual scene annotation in the embodiment of the present invention;

[0044] Figure 5 is a schematic diagram of the display of the penalty result in the 3D virtual scene in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0046] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with their meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined.

[0047] An embodiment of the present invention provides an auxiliary penalty determination method based on artificial intelligence, as Figure 1 shown, the auxiliary penalty determination method based on artificial intelligence proposed by the present invention includes the following steps:

[0048] S1. Obtain continuous image frames collected in real time by a camera device, where the image frames include a tennis court and a tennis ball located within the tennis court; specifically, the continuous image frames can be optionally three consecutive image frames. The tennis court, the tennis ball, and the human body within the tennis court can be photographed by a single camera device deployed on the tennis court to collect the on-site images in real time and obtain multiple sets of continuous image frames. In this embodiment, the system input is three consecutive image frames, which can effectively utilize the temporal information, better capture the object displacement between consecutive frames, and maintain the coherence of the event development (such as the appearance → movement → disappearance of an object), and can reduce the misjudgment of a single frame.

[0049] S2. Detect the position of the tennis ball and the landing position of the tennis ball in the frame where the tennis ball bounces on the ground from the continuous image frames. Specifically, detect the position of the tennis ball and whether the tennis ball bounces on the ground and the landing position from the continuous image frames. The frame where the tennis ball bounces on the ground can be selected as the image frame in the middle position of the continuous images. In a specific example, if the input is three consecutive image frames, the middle frame of the three consecutive input images is the frame where the tennis ball bounces on the ground.

[0050] S3. Take N pixel points to the left and right respectively, and M pixel points up and down respectively with the landing position of the tennis ball as the center in the frame where the tennis ball bounces on the ground to form a region of interest.

[0051] S4. Identify the contour of the tennis ball and the boundary line of the court in the region of interest.

[0052] S5. Generate a penalty determination result according to the positional relationship between the contour of the tennis ball and the boundary line of the court.

[0053] Specifically, after learning the positional relationship between the tennis ball's contour and the court boundary line, that is, whether the landing point is on the line or not. If it is on the line, it represents a line call, which is judged as IN according to the tennis rules. If it is not on the line, the shot needs to be bound. By binding the landing point to the shot (the shot event data can be obtained from the existing system), the effective landing area is determined based on the shot type (serve, return) and the game mode (singles, doubles), and then the true judgment result of IN / OUT for the landing point is calculated.

[0054] The assisted penalty judgment method based on artificial intelligence provided by the embodiments of the present invention can realize the assisted penalty judgment of the Hawk-Eye system only by using a single camera in combination with computer vision algorithms. Compared with the scheme of multiple high-speed cameras, the installation and deployment cost of the Hawk-Eye system is greatly reduced; compared with the schemes of other computer vision algorithms, the present invention simultaneously learns tennis detection and landing point recognition. By combining tennis detection and landing point recognition, the time and position accuracy of landing point recognition are effectively improved, the tennis trajectory is more reliable, and the accuracy of Hawk-Eye recognition is effectively guaranteed. The present invention can ensure the accuracy while reducing the system cost and improving the real-time performance, which is of great significance for realizing the popularization of the Hawk-Eye system in the field of national fitness.

[0055] In the embodiments of the present invention, after obtaining the continuous image frames continuously collected by the imaging device, the method further includes: merging a preset number of continuous image frames in the channel dimension to achieve pre-fusion of the image frames.

[0056] Specifically, different ways can be used to utilize the timing information. Among them, pre-fusion merges multiple frames of data at the input stage, such as stacking three frames of images in the channel dimension. In addition, intermediate fusion can also be used. Intermediate fusion is performed in the middle layer of the network, such as after extracting features at different stages, and then fusing the timing information. Post-fusion can also be used. Post-fusion is performed after the network output, such as integrating the segmentation results of multiple frames. Although pre-fusion will increase the number of input channels, the processing method is simple and the calculation efficiency is relatively high compared to the other two, especially after being optimized by the TensorRT inference engine. Intermediate fusion can capture features at different levels, but it may increase the network complexity and calculation amount. Post-fusion can optimize the final result, but it needs to process the output of multiple frames, which affects the real-time performance. Different fusion methods can be designed according to actual needs in the specific implementation and process.

[0057] In order to achieve a balance between accuracy and speed, since it is real-time detection, the inference speed is crucial. Therefore, the fusion method cannot be too complex. At the same time, the tennis ball is a small target moving quickly, and the timing information needs to effectively capture the movement trajectory to avoid missing detections or false detections. In a specific example, the embodiments of the present invention adopt the pre-fusion method, which is simple to implement. The input directly merges multiple frames, and the network can automatically learn spatio-temporal features.

[0058] The core algorithm of the real - scene Hawk - Eye system lies in accurately judging the tennis bounce frame and the bounce position, that is, it has relatively high requirements for the time accuracy and position accuracy of the landing point. To improve the time and position accuracy, a multi - task joint learning scheme is used in the Hawk - Eye system to simultaneously learn the landing - point recognition task and the tennis detection task. The tennis detection task is to locate the position of the tennis in the image. By introducing supervised learning for tennis detection, the model can focus more on the movement of the tennis in the full - frame image, thereby improving the accuracy of landing - point recognition.

[0059] In the embodiment of the present invention, the specific implementation method for detecting the position of the tennis and the landing point of the tennis in the tennis bounce frame from continuous image frames is as follows: A preset lightweight multi - task joint learning network model is used to detect the position of the tennis and the landing point of the tennis in the tennis bounce frame from continuous image frames; the multi - task joint learning network model includes an encoder, a feature enhancement network, and two task heads, and the two task heads are respectively a tennis detection task - head network and a landing - point recognition task - head network; among them, the encoder is used to extract tennis features of different scales from the input image frames, the feature enhancement module is used to fuse tennis features of different scales to obtain a tennis fusion feature with semantic information and edge texture information, the tennis detection task - head network is used to identify the position of the tennis in the image frame according to the tennis fusion feature, and the landing - point recognition task - head network is used to identify the landing point of the tennis in the image frame according to the tennis fusion feature and determine the tennis bounce frame. In this embodiment, a multi - task network model is used to predict the bounce frame. The input of this model is three consecutive image frames, and the output has two paths. One path is the prediction of the tennis position, and the other path is whether there is a bounce in the three consecutive image frames. If there is a bounce, the middle frame of the three frames is the bounce frame. Specifically, the landing point is where the bounce occurs. The landing - point recognition task - head is used to find the tennis bounce frame. This task - head will output whether there is a bounce in these three frames. If there is a bounce, the middle frame is the bounce frame. Due to the assistance of the tennis position feature, it is possible to more accurately identify whether a bounce has occurred.

[0060] In this embodiment, the Hawk - Eye system algorithm is designed using a convolutional neural network. In order to accurately and efficiently deploy and apply the Hawk - Eye system, a lightweight multi - task joint learning network model is designed. The multi - task joint learning network model consists of an encoder, a feature enhancement module, and two task heads. Among them, the encoder is used to extract tennis features, the feature enhancement module is used to fuse tennis features of different scales, and the task heads further process the features to obtain task - oriented outputs, respectively outputting the position of the tennis and the result of tennis landing - point recognition.

[0061] Specifically, the encoder includes a first down - sampling network layer and a second down - sampling network layer connected in sequence;

[0062] The first downsampling network layer consists of a 3×3 standard convolutional layer and a 3×3 max pooling layer, which is used to extract the primary feature map of the image to achieve preliminary dimensionality reduction;

[0063] The second downsampling network layer includes a downsampling module and at least three feature transformation modules. The downsampling module is used to extract tennis features of different scales from the primary feature map, including a first branch composed of 3×3 depthwise separable convolutions arranged in parallel and a second branch composed of a 1×1 standard convolution, a 3×3 depthwise separable convolution, and a 1×1 standard convolution connected in sequence. The feature output channels of the first branch and the second branch are mixed; each feature transformation module is used to perform feature transformation on the tennis features of different scales obtained by the downsampling module, including a third branch with an identity mapping function arranged in parallel and a fourth branch composed of a standard 1×1 convolution, a 3×3 depthwise separable convolution, and a 1×1 standard convolution connected in sequence. After mixing the feature output channels of the downsampling module, they are evenly divided into two parts, which are used as the inputs of the third branch and the fourth branch respectively, and the feature output channels of the third branch and the fourth branch are mixed.

[0064] In a specific example, the encoder consists of 2 stages. Considering that the tennis ball occupies a relatively small proportion in the image, if 16-fold or 32-fold downsampling is used as in many common object detection models in the prior art, too much target information will be lost. Therefore, only 8-fold downsampling is performed in the model design.

[0065] 1. The initial layer, that is, the first downsampling network layer (Stage 1)

[0066] · Input resolution: 288×512

[0067] · Layer type: Standard convolutional layer + max pooling layer

[0068] · Design details:

[0069] o Standard convolutional layer: 3×3 convolution, stride = 2, output channels = 24.

[0070] o Max pooling layer: 3×3 pooling, stride = 2, used to quickly downsample to a resolution of 72×128.

[0071] · Function: Quickly extract low-level features such as edges and textures, and perform preliminary dimensionality reduction.

[0072] 2. Stage 2, that is, the second downsampling network layer

[0073] · Input resolution: 72×128 → Output resolution: 36×64 (downsampled by the first downsampling Block)

[0074] · Structure: 1 downsampling Block (i.e., downsampling module) + 3 ordinary Blocks (i.e., feature transformation modules). To increase the network depth and enhance the feature extraction ability, at least 3 ordinary Blocks are required.

[0075] · Design of the downsampling Block:

[0076] o Branch 1: 3×3 depthwise separable convolution (DWConv), stride = 2, output channels = 24.

[0077] o Branch 2: 1×1 convolution + 3×3 depthwise separable convolution (stride = 2) + 1×1 convolution, output channels = 24.

[0078] o Channel merging: Merge the outputs of Branch 1 and Branch 2 (a total of 48 channels), and then mix the channels. Assume the output of Branch 1 is 1234 and the output of Branch 2 is abcd. Mixing means recombining these eight results into 1a2b3c4d.

[0079] · Design of the ordinary Block:

[0080] o Channel splitting: Input channels = 48 → Split into 24 + 24.

[0081] o Branch 1: Identity mapping (24 channels). The features obtained by depthwise separable convolution correspond to a low-dimensional space with fewer features. Connecting an identity mapping afterwards can retain most of the features.

[0082] o Branch 2: 1×1 convolution + 3×3 depthwise separable convolution + 1×1 convolution.

[0083] o Channel merging: The merged output is 48 channels → Channel mixing.

[0084] Specifically, the feature enhancement network includes a first upsampling network layer and a second upsampling network layer connected in sequence;

[0085] The first upsampling network layer includes a 3×3 RepVGG convolutional layer and a first upsampling module. The 3×3 RepVGG convolutional layer is used to enhance local feature extraction for the output features of the second downsampling network layer. The first upsampling module is used to upsample the output features of the RepVGG convolutional layer using bilinear interpolation. The output channels of the upsampled features are concatenated with the feature input channels of the second downsampling network layer to obtain 72 feature output channels;

[0086] The second upsampling network layer includes a 3×3 standard convolutional layer and a second upsampling module. The 3×3 standard convolutional layer is used to extract features from the output features of the 72 feature output channels of the first upsampling network layer and convert them into 48 feature output channels for output. The first upsampling module is used to upsample the currently input features using bilinear interpolation. The feature output channels after upsampling are concatenated with the feature output channels of the first downsampling network layer to obtain 72 feature output channels.

[0087] In a specific example, the feature enhancement module consists of 2 upsampling structures, namely the first upsampling network layer and the second upsampling network layer. The features downsampled by 8 times have good semantic information, but the detailed information such as texture and color is severely lost. Therefore, it is necessary to combine the features downsampled by 8 times with the underlying high-resolution features of the network to further improve the recognition effect. Considering the speed during deployment, a convolutional layer is usually used to reduce the number of channels before upsampling to reduce the computational complexity, and then the upsampling operation is performed. To avoid a large number of parameters in the upsampling process, the present invention uses the upsample module for upsampling.

[0088] The first upsampling network layer: Using the features downsampled by 8 times, since the number of channels is small at this time (48 channels), there is no need to reduce the dimensionality of the channels at this time. A RepVGG convolution is used to enhance the features, and it can be reparameterized during deployment to improve the speed.

[0089] · Upsampling Block design:

[0090] o RepVGG convolution: 3×3 convolution, stride = 2, output channels = 48.

[0091] o Upsampling operation: Bilinear interpolation + 3*3 standard convolution.

[0092] o Channel concatenation operation with the input of stage2, i.e., the second downsampling network layer (48 + 24 = 72 channels). Here, instead of elementwise addition, channel concatenation is used, which can better preserve the feature details.

[0093] The second upsampling network layer:

[0094] · Upsampling Block design:

[0095] o 3*3 standard convolution: Output channels 72 → 48.

[0096] o Upsampling operation: Bilinear interpolation + 3*3 standard convolution.

[0097] o Channel concatenation operation with the output of stage1, i.e., the first downsampling network layer (48 + 24 = 72 channels).

[0098] In this embodiment, by concatenating the features after upsampling by the first upsampling network layer with the input features of the second downsampling network layer, and concatenating the features after upsampling by the second upsampling network layer with the output features of the first downsampling network layer, the features downsampled by 8 times can be fused with the detailed features such as edge textures at the bottom layer of the network, improving the effect of tennis detection.

[0099] Specifically, the tennis detection task head network includes a third upsampling module, a 3×3 RepVGG convolutional layer, and a 3×3 standard convolutional layer connected in sequence. The third upsampling module is used to upsample the output features of the feature enhancement network using the bilinear interpolation method. The 3×3 RepVGG convolutional layer is used to enhance the local feature extraction of the output features of the third upsampling module. The 3×3 standard convolutional layer is used to transform the input feature channels of 72 channels into the feature output channels of 3 channels.

[0100] In this application, by stacking multiple 3×3 convolutional layers, the nonlinearity and expressive ability of the model can be enhanced without increasing the number of parameters, significantly improving the performance and efficiency of the model.

[0101] In this embodiment, the tennis detection task head uses the features output by the feature enhancement module to further process and generate predictions for the positions of tennis balls. Tennis detection adopts the semantic segmentation idea. Compared with the object detection method of only detecting a center point and predicting the distances from the upper, lower, left, and right borders, the supervision of semantic segmentation is a mask of a tennis ball, so there can be more supervision information. Therefore, this task head will continue to perform upsampling and finally output the positions of tennis balls.

[0102] · Design of the tennis detection task head:

[0103] o Upsampling operation: Bilinear interpolation + 3*3 convolution.

[0104] o RepVGG convolution: 3×3 convolution, stride = 2, output channels = 72.

[0105] o 3×3 convolutional layer: Input channels 72, output channels 3, and the final output is batch*3*288*512.

[0106] Figure 2 This is an output example of the tennis detection task head, and the pink dots mark the trajectories of tennis movements.

[0107] Specifically, the landing point recognition task head network includes three groups of 3×3 RepVGG convolutional layers, a max pooling layer, three groups of 3×3 RepVGG convolutional layers, a channel attention layer, and a fully connected layer connected in sequence. By establishing a channel attention mechanism in the landing point recognition task head network, the channel weights of each feature input channel of the landing point recognition task head network are dynamically adjusted according to the channel attention mechanism, so as to enhance the attention of the landing point recognition task head network to the motion features of the tennis ball in the image, and output the landing point position of the tennis ball in the image frame.

[0108] In this embodiment, the landing point recognition task head integrates spatial and temporal information to identify whether the tennis ball bounces. Adding an attention mechanism makes the model more focused on the motion of the tennis ball in the image.

[0109] · Design of the landing point recognition task head:

[0110] o Three groups of RepVGG convolutions + max pooling layer + three groups of RepVGG convolutions.

[0111] o Se channel attention mechanism.

[0112] o The fully connected layer outputs the result of the tennis ball bouncing.

[0113] The auxiliary penalty judgment method based on artificial intelligence provided by the embodiment of the present invention further includes: plotting the positions of the tennis balls detected from multiple consecutive image frames and the landing point positions of the tennis balls in the tennis ball bounce frames on a preset real scene map to form a dynamic video or static image corresponding to the tennis ball landing event and marked with the pre-bounce motion trajectory, landing point position, and post-bounce motion trajectory of the tennis ball.

[0114] Specifically, when the player triggers a Hawk-Eye challenge, find the 2-3 second video corresponding to the landing point frame in the Hawk-Eye system service host, and use video editing technology to add the tennis ball motion trajectory, camera movement effect, etc. to this video segment. Through this step, the judgment result of the Hawk-Eye is generated into a visual picture and displayed to the referee, player, and audience. The entire rendering process runs on webgl and can be rendered and played in real time on mobile phones, large screens, and live streams, greatly improving the timeliness. It can be completed within a few seconds from triggering the Hawk-Eye challenge to visual viewing.

[0115] Further, after the special effect live-action video is produced, it needs to be provided for external playback. In the traditional technical solution, the video is uploaded to the cloud, and then the playback end obtains the video content from the cloud. However, this process requires network transmission, which affects the real-time performance of the Hawk-Eye system. To this end, the present invention effectively avoids the impact on real-time performance by adding a layer of reverse proxy technology. The specific steps are as follows: Deploy a cloud server in the cloud, and the cloud server proxies the Hawk-Eye host in the venue. The playback end can find the video in the Hawk-Eye host by accessing the address in the cloud server and play it, thus eliminating the network transmission process.

[0116] Figure 3 It is an output example of the Hawk-Eye system algorithm. The red dots identify the pre-bounce movement trajectory of the tennis ball, the green dots identify the post-bounce movement trajectory of the tennis ball, and the blue dots identify the bounce point. It can be seen that the Hawk-Eye system accurately identifies the time point and bounce position of the tennis ball bounce.

[0117] In the embodiment of the present invention, the implementation steps of identifying the tennis ball contour and the court boundary line in the region of interest specifically include:

[0118] (1) Tennis precise contour detection step: Obtain at least the third frame image frame whose acquisition time is earlier than the tennis ball bounce frame as the background frame, and select the region of interest of the background frame in the same way with the landing position coordinates of the tennis ball in the tennis ball bounce frame as the center; Convert the regions of interest of the tennis ball bounce frame and the background frame into grayscale images through color conversion operations; Use the frame difference method to calculate the absolute difference between the regions of interest of the tennis ball bounce frame and the background frame to obtain a local difference image, and binarize the local difference image after the absolute difference; Erode the binarized image and perform contour detection on the eroded image to obtain the tennis ball contour. In this embodiment, at least the third previous frame is taken as the background frame because the background frame cannot contain moving objects. The purpose of this step is to find the region of interest through the frame difference method. Since the current frame contains the target, the background frame should be as clean as possible, only pure background without tennis balls. Specifically, the region of interest intercepted from the current bounce frame is the foreground, and the third frame is taken forward and the region of interest is intercepted as the background. The two are first converted into grayscale images, and then the pixels are directly subtracted to obtain a difference image. After obtaining the difference image, those with pixel values greater than 5 are directly set to 255, and those less than or equal to 5 are set to 0, that is, a binarized image is obtained.

[0119] Specifically, on the detected tennis ball bouncing frame, with the tennis ball bouncing position as the center, N pixels are taken to the left and right, and M pixels are taken to the top and bottom to form a region of interest, where the value of N can be the image width divided by the first preset statistical value and rounded, and the value of M can be the image height divided by the second preset statistical value and rounded. In this embodiment, the first preset statistical value can be 10, and the second preset statistical value can be 4. The specific value of the preset statistical value can be obtained by converting the ratio of the original image to the target image, and the present invention does not make specific limitations on this. The purpose of taking the region of interest is to only focus on the moving area of the tennis ball when using the frame difference method, and other parts will cause interference;

[0120] Take three frames forward from the tennis ball bouncing frame as the background frame. Use the above-mentioned ROI cropping method to crop a ROI for the background frame. Since the specific bouncing position (x, y) in the bouncing frame is known, the same position (x, y) in the background frame can be cropped.

[0121] The region of interest of the ground frame and the background frame is converted into a single-channel grayscale image through a color conversion;

[0122] Add Gaussian blur to smooth out noise in the image;

[0123] The frame difference method calculates the absolute difference between two regions of interest and binarizes the differenced image;

[0124] Erosion is performed on the binary image, because only the bottom part of the tennis ball touches the ground.

[0125] Perform contour detection on the eroded image to obtain a refined tennis ball contour.

[0126] (2) The steps for detecting the boundary line of the tennis court are as follows: the width and height of the area of interest in the tennis ball bouncing frame are reduced by a preset ratio to obtain a rectangular boundary area; the boundary area is converted into a grayscale image; the grayscale image is binarized and expanded, and the contour of the expanded image is detected to obtain the contour of the court boundary line.

[0127] Specifically, taking the tennis ball bouncing frame as the center, the width and height of the region of interest obtained in the previous step are divided by a third preset statistical value and rounded up, and the third preset statistical value can be 10, so as to obtain a square sideline region;

[0128] Convert the RGB three-color channels of the edge area into a single-channel grayscale image;

[0129] Grayscale Figure 2 Value and expansion;

[0130] Perform contour detection on the expanded image to obtain a refined edge contour.

[0131] The auxiliary penalty judgment method based on artificial intelligence provided by the embodiment of the present invention further includes the implementation of virtual scene construction. Specifically, the camera pose is restored from the 2D-3D corresponding points according to the real-scene video image, and the camera is correctly placed in the three-dimensional scene of the modeled tennis court, achieving the seamless connection effect between the three-dimensional virtual scene and the real-scene image. The specific steps are as follows:

[0132] (1) Camera internal parameter matrix (K):

[0133] ·f x f y : The focal lengths in the x and y directions (in pixel units), calculated from the physical focal length of the camera and the sensor size.

[0134] ·c x c y : The coordinates of the optical center in the image (in pixel units), calculated from the resolution of the image.

[0135]

[0136] (2) Distortion parameters (distCoeffs):

[0137] The distortion parameters include radial distortion (k1, k2, k3) and tangential distortion (p1, p2), which are calibrated by the checkerboard calibration method and are in the form of:

[0138] distCoeffs = [k1, k2, p1, p2, k3]

[0139] (3) Prepare 19 groups of 2D-3D corresponding coordinate points, obtained through image annotation and 3D scene annotation. The schematic diagram of 3D virtual scene annotation is as Figure 4 shown.

[0140] (4) Use the solvePnP open-source algorithm to solve the external parameter rotation matrix (R) and translation vector (T) of the camera.

[0141] (5) Vertical field of view angle (fov) of the camera frustum:

[0142]

[0143] Among them, h is the height of the camera sensor (in pixel units), obtained from the supplier's specification sheet.

[0144] Finally, input the camera rotation matrix R, translation vector T, and vertical field of view angle fov into the camera of the 3D engine, so as to restore the real camera pose in the virtual scene.

[0145] The embodiments of the present invention can also achieve prediction of the landing angle. Since the landing point is an ellipse and has a directionality, based on the bound hitting position and the landing position, a movement route is predicted, and the angle between the route and the Y-axis is calculated to obtain the slope of the landing ellipse. Specifically, a line is connected between the hitting point and the landing point, and the angle with the Y-axis is calculated, and this angle represents the landing orientation of the ellipse.

[0146] Furthermore, when making a penalty decision, the obtained 2D coordinates of the Hawkeye landing point can be converted into 3D coordinates, and the landing point is drawn in the three-dimensional virtual scene. At the same time, the camera is panned, and finally a top-down view of the landing point is formed, which can clearly observe the relationship between the landing point and the sideline, and IN / OUT marks are made. The penalty result display in the three-dimensional virtual scene is as Figure 5 shown. Specifically, when drawing the landing ellipse, the orientation of the ellipse can be determined according to the angle obtained in the landing angle prediction process.

[0147] The assisted penalty decision method based on artificial intelligence provided by the embodiments of the present invention can use only a single camera and combine computer vision algorithms, greatly reducing the installation and deployment cost of the Hawkeye system. It has a lower cost compared to the solution of multiple high-speed cameras. Compared with other computer vision algorithm solutions, it simultaneously learns tennis detection and landing point recognition, and the tennis trajectory is more reliable, effectively ensuring the accuracy of Hawkeye recognition. At the same time, using webgl real-time rendering technology, a display effect combining reality and virtuality is produced, which can be played on mobile phones, large screens, live streams, etc., and the Hawkeye result can be played within 5 seconds, effectively improving the timeliness of the stadium.

[0148] The present invention proposes a Hawkeye recognition solution based on artificial intelligence. The core of the Hawkeye system lies in the proposed lightweight landing point recognition model. By combining tennis detection and landing point recognition, the time and position accuracy of landing point recognition are effectively improved, and based on the high-precision landing point, a virtual-reality combined Hawkeye penalty decision result picture is output and rendered and played in real time on each terminal, also effectively improving the timeliness of the stadium.

[0149] For the method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.

[0150] Another embodiment of the present invention also provides a real-scene Hawkeye system based on artificial intelligence, and the system includes:

[0151] At least one camera device for photographing a tennis court, tennis balls, and human bodies located within the tennis court to collect on-site images in real time;

[0152] A server, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for assisted penalty judgment based on artificial intelligence in any of the above embodiments.

[0153] For the system embodiment, since its implementation of the process for assisted penalty judgment based on artificial intelligence is basically similar to that of the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment, and it has corresponding technical effects.

[0154] In addition, those skilled in the art can understand that although some of the embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, any of the claimed embodiments can be used in any combination.

[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. However, such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An artificial intelligence-based auxiliary penalty judgment method, characterized in that, The method includes: Obtaining continuous image frames collected in real time by a camera device, where the image frames include a tennis court and a tennis ball located within the tennis court; Detecting the position of the tennis ball and the landing position of the tennis ball in the tennis ball bounce frame from the continuous image frames; Taking N pixel points to the left and right respectively, and M pixel points up and down respectively, centered on the landing position of the tennis ball in the tennis ball bounce frame, to form a region of interest; Identifying the tennis ball contour and the court boundary line in the region of interest; Generating a penalty result based on the positional relationship between the tennis ball contour and the court boundary line.

2. The method according to claim 1, wherein Detecting the position of the tennis ball and the landing position of the tennis ball in the tennis ball bounce frame from the continuous image frames includes: Using a preset lightweight multi-task joint learning network model to detect the position of the tennis ball and the landing position of the tennis ball in the tennis ball bounce frame from the continuous image frames; The multi-task joint learning network model includes an encoder, a feature enhancement network, and two task heads, namely a tennis ball detection task head network and a landing point recognition task head network; among them, the encoder is used to extract tennis ball features of different scales from the input image frames, the feature enhancement module is used to fuse the tennis ball features of different scales to obtain a tennis ball fusion feature with semantic information and edge texture information, the tennis ball detection task head network is used to identify the position of the tennis ball in the image frame according to the tennis ball fusion feature, and the landing point recognition task head network is used to identify the landing position of the tennis ball in the image frame according to the tennis ball fusion feature and determine the tennis ball bounce frame.

3. The method according to claim 2, wherein The encoder includes a first downsampling network layer and a second downsampling network layer connected in sequence; The first downsampling network layer consists of a 3×3 standard convolutional layer and a 3×3 max pooling layer, and is used to extract a primary feature map from the image to achieve preliminary dimensionality reduction; The second downsampling network layer includes a downsampling module and at least three feature transformation modules. The downsampling module is used to extract tennis ball features of different scales from the primary feature map, including a first branch composed of 3×3 depthwise separable convolutions arranged in parallel and a second branch composed of a 1×1 standard convolution, a 3×3 depthwise separable convolution, and a 1×1 standard convolution connected in sequence. The feature output channels of the first branch and the second branch are mixed; Each feature transformation module is used to perform feature transformation on the tennis ball features of different scales obtained by the downsampling module, including a third branch with an identity mapping function arranged in parallel and a fourth branch composed of a standard 1×1 convolution, a 3×3 depthwise separable convolution, and a 1×1 standard convolution connected in sequence. After mixing the feature output channels of the downsampling module, they are evenly divided into two parts, which are used as the inputs of the third branch and the fourth branch respectively, and the feature output channels of the third branch and the fourth branch are mixed.

4. The method according to claim 3, wherein The feature enhancement network includes a first upsampling network layer and a second upsampling network layer connected in sequence; The first upsampling network layer includes a 3×3 RepVGG convolutional layer and a first upsampling module. The 3×3 RepVGG convolutional layer is used to perform enhanced local feature extraction on the output features of the second downsampling network layer. The first upsampling module is used to upsample the output features of the RepVGG convolutional layer using bilinear interpolation. The output channels of the upsampled features are concatenated with the feature input channels of the second downsampling network layer to obtain 72 feature output channels. The second upsampling network layer includes a 3×3 standard convolutional layer and a second upsampling module. The 3×3 standard convolutional layer is used to perform feature extraction on the output features of the 72 feature output channels of the first upsampling network layer and convert them into 48 feature output channels for output. The first upsampling module is used to upsample the currently input features using bilinear interpolation. The output channels of the upsampled features are concatenated with the feature output channels of the first downsampling network layer to obtain 72 feature output channels.

5. The method according to claim 2, wherein The tennis detection task head network includes a third upsampling module, a 3×3 RepVGG convolutional layer, and a 3×3 standard convolutional layer connected in sequence. The third upsampling module is used to upsample the output features of the feature enhancement network using bilinear interpolation. The 3×3 RepVGG convolutional layer is used to perform enhanced local feature extraction on the output features of the third upsampling module. The 3×3 standard convolutional layer is used to transform the 72-channel input feature channels into 3-channel feature output channels.

6. The method according to claim 2, wherein The landing point recognition task head network includes three groups of 3×3 RepVGG convolutional layers, a max pooling layer, three groups of 3×3 RepVGG convolutional layers, a channel attention layer, and a fully connected layer arranged in series. By establishing a channel attention mechanism in the landing point recognition task head network, the channel weights of each feature input channel of the landing point recognition task head network are dynamically adjusted according to the channel attention mechanism to enhance the attention of the landing point recognition task head network to the motion features of the tennis in the image, and the landing point position of the tennis in the image frame is output.

7. The method according to any one of claims 1-6, characterized in that, After obtaining consecutive image frames continuously collected by the imaging device, the method further includes: Merging a preset number of consecutive image frames in the channel dimension to achieve pre-fusion of the image frames.

8. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Plotting the positions of the tennis detected from multiple groups of consecutive image frames and the landing point positions of the tennis in the tennis bounce frames on a preset real scene map to form a dynamic video or static image corresponding to the tennis landing event and marked with the pre-bounce motion trajectory, landing point position, and post-bounce motion trajectory of the tennis.

9. The method according to any one of claims 1-6, characterized in that The identifying the tennis contour and the court boundary line in the region of interest includes: Obtaining at least a third image frame whose acquisition time is earlier than the tennis bounce frame as a background frame, and selecting the region of interest of the background frame in the same way with the coordinate of the landing point position of the tennis in the tennis bounce frame as the center; Converting the regions of interest of the tennis bounce frame and the background frame into grayscale images through color conversion operations; Calculating the absolute difference between the regions of interest of the tennis bounce frame and the background frame using the frame difference method to obtain a local difference image, and binarizing the local difference image after the absolute difference. Erode the binarized image, perform contour detection on the eroded image to obtain the tennis ball contour; In the tennis ball bounce frame, reduce the width and height of the region of interest of the tennis ball bounce frame by a preset ratio to obtain a rectangular sideline region; Convert the sideline region to a grayscale image; Binarize the grayscale image and perform dilation, then perform contour detection on the dilated image to obtain the court boundary line contour.

10. An artificial intelligence-based real-scene eagle-eye system, characterized in that, The system includes: At least one camera device for photographing the tennis court, the tennis ball and the human body located in the tennis court to collect the live scene images in real time; A server including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the method according to any one of claims 1-9 are implemented.

Citation Information

Patent Citations

  • System and method utilizing single camera to accurately determine ball drop point

    CN108596942A

  • Intelligent auxiliary judgment method and system for badminton sports

    CN114005072A

  • Method for rebroadcasting table tennis match data acquisition and visualization technology

    CN118200750A

  • Tennis match video analysis system, device and method

    CN118570704A

  • Live-action auxiliary penalty method and system

    CN119649267A