Computer-implemented system for the automatic analysis of a volleyball match
A computer-implemented system with cameras and neural networks automates volleyball match analysis, addressing the limitations of costly and time-consuming human scouting by providing detailed metadata and statistics, enhancing scouting efficiency and reducing costs.
Patent Information
- Application Number
- PCT/IB2025/051886
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-22
- Filing Date
- 2025-02-21
- Publication Date
- 2025-08-28
AI Technical Summary
Current scouting methods in volleyball are costly, time-consuming, and limited in providing in-depth information about player positions and movements during matches, relying heavily on subjective human analysis that is not affordable by most clubs and lacks real-time capabilities.
A computer-implemented system using multiple cameras and neural networks to automatically analyze volleyball matches, generating metadata on player positions, actions, and ball dynamics, enabling efficient, rapid, and cost-effective scouting through a processing platform that includes analysis modules for action detection, synchronization, and player identification.
Provides comprehensive, real-time analysis of volleyball matches, offering detailed metadata and statistics that enhance scouting efficiency, reducing human effort and costs, and allowing for in-depth insights into team dynamics and player performance.
Smart Images

Figure IB2025051886_28082025_PF_FP_ABST
Abstract
Description
[0001] COMPUTER-IMPLEMENTED SYSTEM FOR THE AUTOMATIC ANALYSIS OF A VOLLEYBALL MATCH
[0002] Technical Field
[0003] The present invention relates to a computer-implemented system for the automatic analysis of a volleyball match, particularly employable for scouting in volleyball.
[0004] Background Art
[0005] In the context of volleyball, it is well known and very important to scout the matches played by one’s own team or opposing teams in order to analyze each individual phase of the match, specifically serve, receive, high shot, attack, block, defense and counterattack, and in order to evaluate the effectiveness of each fundamental according to a standardized rating scale.
[0006] It is also known that such scouting activity is commonly carried out by specialized professionals (scouts) who are responsible for studying each match (particularly of opposing teams) by extracting useful data and information and editing footage displaying the various actions so that it can be later used effectively by coaches and players.
[0007] However, such scouting activity is a tool that can only be used by teams with greater financial means.
[0008] In addition to this, currently manually-performed scouting does not contain much of the information useful to any coach of a volleyball team (e.g., the position of defensive players in the moments before an attack) because the scout cannot extract all this information except by processing and tagging the same video several times and adding additional information each time. Gathering such additional information would then require multiple scouts for each team or, in any case, the expenditure of several hours of processing for each individual match.
[0009] Therefore, the use of one or more specialized and dedicated professionals to analyze matches in real time or videos of previous matches, then for several hours, involves high costs or, in any case, costs that are not affordable by volleyball clubs outside the professional circuit.
[0010] In addition, even considering the professional circuit, current tools do not allow for tracking in- depth information such as, e.g., the position of players on the court at the time of each individual action, the movement of an individual player, and position adjustment just before an attack / defense.
[0011] Moreover, such analysis has inherent limitations related to the specific technical level of the scouting professional, his or her subjective preferences and, most importantly, the amount of time each individual professional can devote to analyzing matches.
[0012] Therefore, there is an increasing need to find alternative tools to acquire and process useful information related to the dynamics of the match during a volleyball match, in real time or from videos of previous matches, effectively, quickly and with limited costs, in order also to avoid the performance of tedious repetitive and wearisome human activities.
[0013] Description of the Invention
[0014] The main aim of the present invention is to devise a computer-implemented system for the automatic analysis of a volleyball match, particularly employable for scouting in volleyball, which enables the automatic acquisition and processing of useful information related to match dynamics during a volleyball match in an efficient, rapid, normalized and cost-effective manner.
[0015] The aforementioned objects are achieved by the present computer-implemented system for the automatic analysis of a volleyball match, particularly employable for scouting in volleyball, according to the characteristics described in claim 1.
[0016] Brief Description of the Drawings
[0017] Other characteristics and advantages of the present invention will become more apparent from the description of a preferred, but not exclusive, embodiment of a computer-implemented system for the automatic analysis of a volleyball match, particularly employable for scouting in volleyball, illustrated by way of an indicative, yet non-limited example, in the attached tables of drawings in which:
[0018] Figure l is a general diagram of the computer-implemented system according to the invention;
[0019] Figure 2 is a general diagram showing the main functional blocks of a computer-implemented system analysis module according to the invention;
[0020] Figure 3 is a block diagram detailing the operation of an action detection block of the analysis module in Figure 2;
[0021] Figure 4 shows a possible neural network input of the action detection block on an individual frame;
[0022] Figures 5, 6 and 7 show possible outputs of the neural network, again on an individual frame;
[0023] Figure 8 is a diagram showing a possible architecture of the neural network used in the action detection block;
[0024] Figure 9 shows a possible architecture of a hit classification head of the action detection block neural network;
[0025] Figure 10 shows a possible architecture of a hit localization head of the action detection block neural network;
[0026] Figure 11 shows the determination of the angles on a frame using a court line detection block of the analysis module;
[0027] Figure 12 shows the operation of a synchronization block of the analysis module;
[0028] Figure 13 shows the operation of a clip detection block of the analysis module;
[0029] Figure 14 shows the architecture of a head detection neural network of a player detection block of the analysis module; Figure 15 shows the architecture of a pose estimation neural network of the player detection block of the analysis module;
[0030] Figure 16 shows a general diagram of the player detection block of the analysis module;
[0031] Figure 17 shows an example of the output of the head detection neural network and of the pose detection neural network;
[0032] Figure 18 shows a possible architecture of an OCR network used in the player detection block for reading players’ jersey numbers;
[0033] Figure 19 schematically shows the calculation of the 3D position of the ball using a post-processing block of the analysis module;
[0034] Figure 20 is a general diagram showing the main functional blocks of an analysis module according to a possible alternative embodiment of the computer-implemented system according to the invention.
[0035] Embodiments of the Invention
[0036] With particular reference to the general diagram shown in Figure 1, reference numeral 1 globally denotes a computer-implemented system for the automatic analysis of a volleyball match, particularly employable for scouting in volleyball.
[0037] Specifically, the system 1 comprises at least one camera 2 positioned in the proximity of a volleyball court and oriented to capture at least one video of the players within the court.
[0038] In addition, the system 1 comprises at least one processing platform 3, comprising at least one remote and / or local processing unit, operationally connected to the aforementioned at least one camera 2 and configured to process one or more videos V captured to obtain data related to the game dynamics during a volleyball match.
[0039] According to a preferred embodiment, the system 1 comprises two cameras 2 placed in the proximity of respective opposite short sides of the court, substantially central to the short sides, and oriented to capture two respective videos V of the players within their respective halves of the court.
[0040] Preferably, the two cameras 2 are placed on the short sides of the court, central, at a predefined distance and height (varying according to the available space, indicatively at a distance of between 5m and 25m). The cameras 2 should have adequate resolution (not less than 1280x720), adequate focal length and adequate frame rate (not less than 25 fps).
[0041] Different embodiments cannot however be ruled out wherein the system 1 comprises a different number and a different placement of the cameras 2.
[0042] For example, the system 1 may use an individual camera 2 arranged on one side of the court.
[0043] In addition, the system 1 according to the invention can also work with an individual camera 2 placed in the proximity of only one of the short sides of the court. In this case, the system 1 will provide no cues for the team in the uncam eraed court and partial cues for the team in the cameraed court. Since there is a change of court after each set, in this case no team will have a complete overview of the match.
[0044] As for the processing platform 3, this can be implemented through a desktop application and / or a mobile application, operationally connected to at least one cloud or local computing unit. Specifically, as schematically shown in Figure 1, the processing platform 3 comprises: at least one analysis module 4, configured to process one or more videos V captured by means of the aforementioned at least one camera 2 and to generate metadata related to the game dynamics during the volleyball match; at least one storage 5 configured to store the captured videos V and generated metadata; at least one scouting analysis module 6 configured to provide on-demand the analysis data processed from the generated metadata to an operator O.
[0045] In addition, the system 1 comprises a man-machine interface 7 configured to manage the uploading of the videos V directly from the at least one camera 2 and / or indirectly from an operator O and configured to interface an operator O (the same operator uploading videos or another operator) with the scouting analysis module 6.
[0046] According to a preferred embodiment, the analysis module 4 produces the metadata listed below, which are stored in the storage 5 along with the match videos V.
[0047] Specifically, for each ball touch, the analysis module 4 calculates the following information: instant in the match when that hit occurred; number of the player who hit the ball; type / fundamental of ball touch: set, forearm pass, block, spike, serve; specialization of the fundamental: high / medium / fast / tight high shot, spin / float serve, attack spike / lob; logical action: attack, defense, receive, block, serve, high shot.
[0048] For each ball touch on the ground, the analysis module 4 calculates the following information: instant in the match when that hit occurred; court position where the ball hit the ground.
[0049] For each instant of play, the analysis module 4 calculates the following information: the position on the court of all players, identified by their jersey number; which of these players is on the jump; the position of the ball, both in 2D and 3D.
[0050] The generated metadata, along with the match videos, can then be downloaded by an operator who through the scouting analysis module 6 can perform the analyses described below.
[0051] One possible analysis using the scouting analysis module 6 involves selecting a set of metadata, from one match or multiple matches, from the same league (if you want to analyze one team over multiple matches) or from different leagues (if you want to analyze one player over multiple leagues).
[0052] Using the scouting analysis module 6, it is possible to perform a set of metadata filtering, for example, by: which set(s), player’s jersey number, type of fundamental, type of logical hit, position of hit performance, position of hit arrival (ground or other player), outcome of hit (in play / end point), outcome of point (win / lose) or a combination of all the previous selections.
[0053] In addition, through the scouting analysis module 6, aggregate statistics processing and relevant display can be carried out.
[0054] For example, it is possible to make quantitative aggregate statistics regarding how often a hit occurs. In this case, one or more elements must be specified on which to aggregate the filtered metadata, specifically: if you filter by a player and aggregate by fundamental, you can count how many hits that player made for each fundamental; if you filter by player and fundamental and aggregate by point outcome, you can count how many times the point ended in favor or against the team when that player performed that fundamental in the action.
[0055] The display of processed aggregate statistics is preferably tabular in nature.
[0056] In addition, through the scouting analysis module 6 it is possible to make temporal statistics, that is, when in the set and match a hit happens.
[0057] Such time statistics are used to analyze how points are distributed within a set, for example, whether more errors were made at the end of a set than at the beginning, or to check the sequence of points won / lost consecutively.
[0058] Preferably, the display of time statistics is of the donut type.
[0059] Using the scouting analysis module 6, it is possible to perform geographic statistics, that is, where a hit happens in the court.
[0060] These geographic statistics show where certain hits were taken and their direction and point of receive / hit to the ground.
[0061] Preferably, geographic statistics are represented by arrows on the court seen from above colored according to the quality of the hit, to the player who made it, or to other criteria.
[0062] Alternatively, in case you are interested in analyzing only the start or arrival location of the hit, geographic statistics can be represented by heatmaps.
[0063] In addition, through the scouting analysis module 6 it is possible to perform ball statistics, i.e., velocity, time at net, maximum height and 2D and 3D trajectory of the ball.
[0064] Such ball statistics are useful for understanding the game dynamics within the team, e.g., high shot-spike or power of attacks.
[0065] Through the analysis module 6, different displays can be managed depending on what specifically you want to display. For example, you can display a histogram of ball velocities or maximum ball height or time between high shot and spike, or an overlay of 2D trajectories on the video or a plan view of the ball trajectories from above.
[0066] Finally, through the scouting analysis module 6 it is possible to perform position statistics, that is, where players are in the instants before a hit.
[0067] Such position statistics are useful in understanding the game dynamics and the responsiveness of other allied or opposing players around the high shot and attack.
[0068] Preferably, the position statistics are displayed by plan view with heatmaps and arrows indicating the movement of the players in the last few seconds.
[0069] Figure 2 shows the main functional blocks of the analysis module 4.
[0070] Specifically, for each camera 2 used, the analysis module 4 comprises an action detection block 41 configured to receive as input the entire video V captured by the respective camera 2 and to generate as output the following data: for each hit detected in the video V: time instant, pixel coordinates of the ball touch, pixel coordinates of the head of the player who hit the ball, type of fundamental; for each frame of the video V: pixel coordinates of the ball.
[0071] In addition, again for each camera 2 used, the analysis module 4 comprises a court line detection block 42 configured to receive as input a manual selection of the four court comers made by operator O or, if not manually indicated by the operator O, N frames from the video V of that camera 2.
[0072] The court line detection block 42 is configured to generate as output the pixel coordinates of the intersections of the play lines of the midcourt near the relevant camera 2.
[0073] In case at least two cameras 2 are used, the analysis module 4 comprises a synchronization block 43 configured to receive as input a manual time alignment of the videos V acquired from each camera 2, performed by an operator O or, if not manually indicated by an operator O, the outputs of each action detection block 41 related to each camera 2.
[0074] The synchronization block 43 is configured to generate as output the time difference in seconds between the two videos V.
[0075] In addition, the analysis module 4 comprises a clip detection block 44 configured to receive as input the output of the action detection block 41 and, if any, of the synchronization block 43.
[0076] The clip detection block 44 is configured to generate as output the following data for each clip in the match: clip start instant, consisting of the instant of the serve that starts the point; clip end instant, consisting of the instant of the last hit that closes the point. For each camera 2 used and for each generated clip, the analysis module 4 comprises a player detection block 45 configured to receive as input the following data: the entire video V captured by each of the cameras 2, the clips start and end instants and the court lines near the camera 2. The player detection block 45 is configured to generate as output the following data: for each player, their jersey number, if visible; for each frame, the metric position of the player on the court and the pixel coordinates of the head of the player.
[0077] In addition, the analysis module 4 comprises an action-player matching block 46 configured to receive as input the following data: the output of the action detection block 41; the output of the player detection block 45; the clip start and end instants generated by the clip detection block 44. The action-player matching block 46 is configured to generate as output which player performed which action.
[0078] Finally, for each clip, the analysis module 4 comprises a post-processing block 47 configured to receive as input the following data: the clip start and end instants; for each camera 2: the output of the action detection block 41; the output of the player detection block 45; the output of the action-player matching block 46.
[0079] The post-processing block 47 is configured to generate as output an improvement in the actions recognized within each clip.
[0080] Therefore, the analysis module 4 works preferentially on two videos V that record the match from the two opposite short sides of the court. The two videos V undergo sometimes independent processing (such as action detection or player detection) and sometimes joint processing (such as synchronization, clip splitting and post processing).
[0081] The purpose of the analysis module 4 is to provide the metadata that substantially answers the question, “at each instant of the match where are the players and what are they doing?” to then allow an operator O to view the analysis starting from this information.
[0082] Although the analysis module 4 works preferentially on two videos V (recorded from opposite short sides of the court), it can also work in the following two configurations: only one camera 2 on the short side: in this case the synchronization block 43 and the postprocessing block 47 will not be exploited and in general all results and analysis will be provided only for the cameraed midcourt near the camera; only one camera 2 on the long side, in this case the synchronization block 43 and the postprocessing block 47 will not be exploited, and in general it will not be possible to identify the players automatically, but it will still be possible to provide the splitting into clips, the performance of actions within a point and the court position of all players of both teams. Figure 3 details the operation of the action detection block 41 of the analysis module 4. Specifically, first of all, the action detection block 41 comprises a step 100 of resizing the original full video V to a predefined fixed size, and a step 101 of reducing the frame rate of the video V to a predefined frame rate.
[0083] Next, the action detection block 41 comprises a step 102 of splitting the video V (resized and with reduced frame rate) into frame blocks.
[0084] Next, the action detection block 41 comprises a step 103 of analyzing the video V by means of a neural network NN for the extraction of the actions.
[0085] The neural network NN is configured to receive as input the video V, specifically the frame blocks, and to produce as output the following types of output: hit classification output 104 to determine the fundamental of the action performed, selected from serve, block, spike, set, forearm pass and ball to the ground; the hit classification output comprises for each frame and for each fundamental the probability that such a fundamental occurred in that frame; hit location output 105: a value for each pixel and for each frame, indicating the probability of the presence of any fundamental at that position in the image; player location output 106 to determine the position of the head of the player who performed the action, comprising a vector for each pixel and for each frame indicating the displacement that is necessary to follow from the considered pixel to obtain the position of the player’s head, if an action had occurred in that pixel; ball location output 107 to determine the position of the ball during the performance of the point, the network predicts a value for each pixel and for each frame, which indicates the probability of the ball’s presence at that position in the image.
[0086] As an example, Figure 4 shows a possible neural network input of the action detection block 41 on an individual frame, while Figures 5, 6 and 7 show possible neural network outputs, again on that individual frame. Specifically, Figure 5 shows hit classification, Figure 6 shows hit and ball localization while Figure 7 shows player localization.
[0087] Since the hit classification output 104 and the hit location output 105 are of the probabilistic type (i.e., continuous between 0 and 1), the action detection block 41 comprises a post-processing step 108 of the outputs to determine the actual presence and position of the various actions.
[0088] Specifically, such post-processing step 108 comprises the application of thresholds to discretize the hit classification output 104 and the hit location output 105 (preferably, 0.5 for hit classification, 0.2 for hit localization). The outcome will then be information that is no longer continuous but binary: in each frame (and pixel) either there is one action or there is not.
[0089] Next, because active (i.e., above-threshold) contiguous frames (and pixels) likely refer to the same action, the post-processing step 108 comprises clustering these active contiguous frames (and pixels) into classification-related components 109 and location-related components 110 in order to avoid predicting the same action several times.
[0090] Next, the action detection block 41 comprises an association step 111 of the location-related components 110 with the classification-related components 109 on the basis of temporal proximity, to thus obtain complete predicted actions comprising the position and touch type.
[0091] Next, the action detection block 41 comprises an action-player association step 112, starting from the complete predicted actions obtained in the association step 111 of connected components and from the player location output 106 generated by the neural network NN.
[0092] Specifically, the action-player association step 112 comprises the following steps: considering the pixel position in the hit location output 105; at that same position, reading the components (vx, vy) of the predicted vector related to the position of the player’s head; applying such a vector at the point where it was extracted to determine a new position corresponding to the position of the head that performed the action.
[0093] Finally, the action detection block 41 comprises a determination step 113 of the ball position at each frame, starting from the ball location output 107 by choosing the pixel coordinates with the highest probability and above a predefined threshold (preferably equal to 0.5).
[0094] A possible architecture of the neural network NN used in the action detection block 41 is schematically shown in Figure 8.
[0095] Specifically, the neural network NN implemented for the action detection block 41 consists of two main parts: a central part B (known as the backbone) configured to extract features FT from the frame blocks FR as input and characterized by spatiotemporal operators (3D convolutions); and an end part T configured to produce the hit classification outputs 104, the hit location outputs 105, the player location outputs 106 and the ball location outputs 107.
[0096] Specifically, according to a preferred embodiment, the end part T comprises a small final neural network (known as head) for each output, specifically: a hit classification head Hl, a hit location head H2, a player location head H3 and a ball location head H4.
[0097] According to a possible embodiment, schematically shown in Figure 9, the hit classification head Hl consists of a global average pooling block G to aggregate the spatial dimensions, followed by a fully-connected layer L to obtain a hit classification output 104 that has, for each frame, a probability vector of the cardinality of the correct number of classes.
[0098] Again according to one possible embodiment, schematically shown in Figure 10, the hit location head H2 consists of: a global average pooling block G to aggregate spatial dimensions; a first block BL1 comprising a first step of upsampling followed by a first convolutional layer; a second block BL2 comprising a second step of upsampling followed by a second convolutional layer.
[0099] In particular, the two blocks BL1 and BL2 are configured to switch from the reduced spatial size of the features after the backbone B to the original size of frames. In particular, the second convolutional layer of the second block BL2 will output one channel only, which indicates the probability that an action occurred in each pixel of each frame.
[0100] According to one possible embodiment, the player location head H3 is implemented by means of an architecture similar to that described above for the hit location head H2, but it has two output channels that do not represent probabilities but the vx and vy components of vectors that originate in each pixel and that indicate the position of the player who performed an action, if at that pixel the network really predicted an action.
[0101] Finally, according to one possible embodiment, the ball localization head H4 is implemented by means of an architecture similar to that described above for the hit localization head H2, with only one output channel indicating the probability that there is or is not a ball in each pixel of that frame. As for the training of the neural network NN of the action detection block 41, the dataset used comprises clips taken from videos of matches containing the play actions: from the beginning of the point with setting the ball into play to the end of the point, i.e., when the ball falls to the ground. In this regard, it should be pointed out that a point of play generally consists of an ordered sequence of actions performed by the two teams during the point, these actions are the fundamentals, and are those predicted by the network described above: serve, block, forearm pass, set, spike.
[0102] According to a preferred embodiment, for training the neural network NN 5000 clips taken from the data described above were manually annotated by domain experts.
[0103] On average, each clip contained four actions on each side of the court, so the network was collectively trained on about 20,000 actions.
[0104] The ball was randomly annotated on about 20% of the frames of the clips away from the hits (as ball position and hit position coincide during a fundamental). During training, the neural network NN is penalized for wrong positions only on frames where the ball was annotated.
[0105] Also according to this preferred embodiment, the neural network NN was trained by dividing the dataset described above into training set (80%) and test set (20%). After that in the training phase, random horizontal flip was applied as data augmentation method, and the network was trained on the training set using Adam as optimizer and le-4 as the fixed learning rate value. During training, the blocks are sampled so that the percentage of blocks containing an action and of blocks that do not contain an action are balanced. The court line detection block 42 of the analysis module 4 is described below.
[0106] In particular, the court line detection block 42 is important to be able to correlate the position of players on the image to the position of players in the plan (top-down view), in a metric space.
[0107] According to one possible embodiment, the operator O can correct or manually enter the position of the court lines in the loading phase of the video V.
[0108] The court line detection block 42 allows automatically and starting from a video V or, better, from some frames of a video V, to extract the lines.
[0109] In addition, given the lines, the court line detection block 42 is configured to correlate the players’ position on the court and on the image plane.
[0110] Preferably, the court line detection block 42 comprises a neural network configured to predict some known points in the court.
[0111] In addition, the court line detection block 42 comprises a step of filtering out such known points over multiple frames.
[0112] According to a preferred embodiment, the neural network of the court line detection block 42 is based on the architecture known in literature as Unet and generates four output maps, one per court corner. Each generated map has the size of the input image and represents the probability that there is that court corner in that pixel. Preferably, the neural network of the court line detection block 42 is trained by minimizing the mean square error between prediction and ground truth values on the training dataset.
[0113] Similarly to the neural network NN of the action detection block 41, the neural network of the court line detection block 42 is trained starting from a manually annotated dataset of about 5,000 images from different matches, gyms and leagues. Preferably, the dataset is divided into training set (80%) and test set (20%), and the images are resized to a fixed size of 176x320.
[0114] According to a preferred embodiment, the court line detection block 42 is configured to perform the following steps: selecting a predefined number of frames per video (preferably 20 frames per video); via the neural network, determining the angles A on all the selected frames (Figure 11); calculating the average of the angles A determined between the various frames, excluding predictions too different from the average, preferably by means of an outlier rejection procedure; having obtained the average angles, homography to switch from the image plane to the court metric plane.
[0115] In this regard, we point out that homography means a known problem in computer vision where, starting from a correspondence of at least three 2D (found from the network) and 3D points (known measurements of the court), we can relate these two worlds. The result of homography in fact, allows us to switch from the image plane to the court metric plane and thus know where the players are at each instant.
[0116] Court line detection is carried out by the block 42 for each camera V independently.
[0117] The synchronization block 43 of the analysis module 4 is described below.
[0118] During the match, each camera 2 records a video V independently of the others, which can then start at different times. The synchronization block 43 is important in order to fully exploit the benefit of having several cameras 2 to reinforce the results of each.
[0119] Possibly, the operator O can also enter a manual synchronization at the time of loading the videos V or can correct the automatic one obtained by means of the synchronization block 43.
[0120] The synchronization block 43 receives as input the different videos V from the cameras 2, in addition to the output from the action detection block 41, which produced a localization and classification of the actions independently on each camera.
[0121] Synchronization by means of the synchronization block 43 is based on the intuition that certain hits follow each other in the match of volleyball, that is, after a serve we expect a forearm pass or set, and if there is a block we expect a spike first.
[0122] Specifically, as schematically shown in Figure 12, the synchronization block 43 is configured to perform the following steps: determining the first hits (serve and subsequent receive of the first set) from each side of the court to have a first (inaccurate) alignment between the two videos V; around that first alignment, determining several possible alignments with increasing time deltas and, for each alignment, evaluation of the time compatibility of the following combinations of fundamentals: spike-block, serve-forearm pass; of all evaluated alignments, selecting the alignment with the greatest compatibility.
[0123] The clip detection block 44 is described in detail below.
[0124] The clip detection block 44 is configured to receive as input the outputs from the action detection block 41 and from the synchronization block 43 and to determine, starting from these outputs, the start and end times of each point played during the match.
[0125] Specifically, as schematically shown in Figure 13, the clip detection block 44 is configured to cycle through the actions predicted by the neural network NN of the action detection block 41 in order of occurrence, by performing the following steps: keeping a state that indicates whether you are during an action or not, initially the state will be on “no action”. the state switches from “no action” to “action” if the current state is “no action” and if any action occurs but other than ball touching the ground (a serve, or a receive because the other team has served); the state switches from “action” to “no action” if the current state is “action” and if:
[0126] - both in the case of one camera 2 and two cameras 2 being used: a ball action touching the ground occurs in the midcourt near the camera; in this case the action is over; in case of using just one camera 2: no action occurs for a longer period of time than a certain threshold of seconds, in which case it means that the other team failed to send the ball back into the court near the camera and, therefore, the point is over; in case of using two cameras 2: a ball action touching the ground occurs in the opponent’s midcourt; in this case it means that the action is over.
[0127] The player detection block 45 is described below.
[0128] The player detection block 45 receives as input the outputs from the clip detection block 44 and from the court line detection block 42.
[0129] The player detection block 45 is performed independently for each camera 2 and for each of the clips.
[0130] Specifically, the player detection block 45 is configured to: find all players on the side of the court framed by the clip, discarding all other people in the scene (other players, referee, coaches, fans); identify players by reading their jersey number; determine, for each frame, the court position of the players; follow the movements of the players within the clip.
[0131] Given the complexity of finding all the players especially because of the frequent occlusions between players that happen during a clip, the player detection block 45 comprises two neural networks: a head detection neural network HD configured to identify the head of the players; a pose estimation neural network PE configured to identify all joints of the players (hands, feet, pelvis, head, shoulders, etc...).
[0132] In fact, while the position of the feet is necessary in order to correctly position a player on the court, often the head is the only visible element of a player.
[0133] In contrast to the action detection block 41, the two neural networks HD and PE work independently on the individual frames of the clips and for each frame information returns regarding all the people in the image.
[0134] Figure 14 shows the architecture of the head detection neural network HD of the player detection block 45.
[0135] According to one possible embodiment, the head detection neural network HD comprises a backbone B, preferably implemented by means of a RESNET50, configured to receive the images I as input from the video V and to extract the features FT from these images I. In addition, the head detection neural network HD comprises a head localization head H configured to regress a map M indicating the probability that there is a head of any person in each pixel.
[0136] In this way, with an individual forward pass, it is possible to recognize the heads of all the people within a frame of the video.
[0137] The head detection neural network HD is trained using an internal dataset of about 5000 images, manually annotated. Preferably, the dataset is divided into training set (80%) and test set (20%), and the images resized to a predefined fixed size. Adam with a learning rate of le-3 is preferably employed as the optimizer.
[0138] Figure 15 shows the architecture of the pose estimation neural network PE of the player detection block 45.
[0139] According to one possible embodiment, the pose estimation neural network PE comprises a backbone B, preferably implemented using a RESNET50, configured to receive the images I as input from the video V and to extract the features FT from those images I.
[0140] In addition, the pose estimation neural network PE comprises a Feature Pyramid Network FPN configured to process the extracted features FT to generate scaling-invariant fused features FF, so as to be able to correctly recognize the player joints even at different scaling values.
[0141] In addition, the pose estimation neural network PE comprises a final head H configured to regress the coordinates of the bounding box and the coordinates of each joint (keypoints) for each person in the initial image. Preferably, the pose estimation neural network PE is also trained using a dataset of about 10,000 images, manually annotated. The dataset is divided into training set (80%) and test set (20%), and the images resized to a predefined fixed size. Adam with a learning rate of le-3 is also preferably employed as the optimizer.
[0142] The two neural networks HD and PE produce results that are independent, that is, for a player, in a given frame, there could be only the head, both the head and the pose or only the pose.
[0143] A general diagram of the player detection block 45 is shown in Figure 16.
[0144] As can be seen in that figure, the outputs produced by the two neural networks HD and PE are first subjected to a filtering phase which allows retaining only the information of the players in the near part of the court and then the information obtained is merged with each other and over time (i.e., in two consecutive frames which heads and which poses refer to the same player).
[0145] Specifically, the player detection block 45 comprises a first filter Fl and a second filter F2 configured to filter the heads and poses predicted by the head detection neural network HD and the pose estimation neural network PE, respectively.
[0146] Specifically, the first filter Fl is configured to exclude heads with lower confidence than a predefined threshold (preferably equal to 0.5), while the second filter F2 is configured to discard poses that have lower confidence among all pose keypoints below a predefined threshold (preferably equal to 0.3). On the poses, the second filter F2 then applies a “non maxima suppression” (nms) to filter out poses that refer to the same player.
[0147] Figure 17 shows an example of the output of the head detection neural network HD (dots visible at the players’ heads) and of the pose detection neural network PD on a specific player (top right, skeleton in which hands and feet are highlighted in particular). In the graphical representation, results are shown only for players in the court near the camera but the two neural networks HD and PD can produce results for all people in the scene, which will then need to be filtered by position as detailed below.
[0148] As illustrated in Figure 16, downstream of the first and of the second filters Fl and F2, the player detection block 45 comprises a first tracklet creation block T.
[0149] The tracklet creation block T is configured to associate a head in a given frame i with a head in the frame i+1 if their spatial distance is below a certain threshold and if there are no other heads in frame i+1 with a lower distance than the threshold (thus compatible). In the latter case, ambiguity would occur and the association could be wrong, introducing an irreversible error.
[0150] In the case where there is not even one head close to that in frame i+1, then the tracklet creation block T is configured to also consider the average heads extracted from the poses (only the poses after filtering on the average head confidence) as possible candidates for association. This association is made with the same criterion as stated above before: distance below a certain threshold and no ambiguity with other candidate heads.
[0151] Next, the player detection block 45 comprises an association block P-T of all the poses extracted from the pose estimation network PE to the tracklets constructed from the block T by the heads. Specifically, the association block P-T is configured to associate poses with the tracklets if the spatial distance between the head of the tracklet and the average head of the pose (thus looking at all keypoints on the head: ears, eyes,...), is below a predetermined threshold and if there are no association ambiguities.
[0152] It is important to note that no tracklets are modified at this stage, only poses are added to the already defined tracklets. In fact, heads are more reliable than poses, and yet, the poses might add some extra detection that lengthen the tracklets, allowing players to be tracked for more frames. Pose, on the other hand, is essential to be able to correctly position the player on the court, which is also useful to filter out people who are not players on the court under consideration (such as referee, audience, coach or people on the bench).
[0153] Next, the player detection block 45 comprises a position-based filtering block F3.
[0154] Specifically, since only the tracklets corresponding to the six players playing in the midcourt near the camera 2 are of interest, the filtering block F3 allows all out-of-court tracklets to be eliminated. This is done by computing, for each tracklet, the plane positions of each point that has an associated pose. The pose is necessary because the in-court position is computed using the position of the feet, via a homography between the image plane and the playing court plane. The homography is calculated as described above for the court line detection block 42, that is, by a correspondence between some intersections of the court lines in the image and the actual dimensions of the court in the metric space.
[0155] In order to achieve the correct plane position, it is necessary for the players’ feet to be in contact with the court, so it is first necessary to exclude all poses that are considered jumping.
[0156] For this purpose, the filtering block F3 is configured to trigger a jump condition when both of the following conditions are met: one pose has his hands above his head; the feet of a pose move vertically upwards.
[0157] If both conditions are verified, the filtering block F3 is configured to consider poses in the next 25 frames of that same tracklet in a jumping condition and, therefore, those poses are not used for the calculation of the plane position.
[0158] In addition, the filtering block F3 is configured to exclude out-of-court players after calculating the plane position for each non-jump pose by applying the following steps: for each pose of a tracklet, determining a position state selected from outside or inside, where: if a pose is beyond the midcourt it is considered to be in the outside state; if a pose is inside the playing rectangle it is considered to be in the inside state; if a pose is before the net but outside the rectangle the position is determined by means of a motion analysis; determining the global position of the tracklet considered.
[0159] Specifically, the motion analysis comprises the following steps: if the distance of the feet from the previous frame is less than a predefined threshold, then the tracklet is considered stationary (it is probably a coach and so it is considered outside the court); if the distance of the feet from the previous frame is greater than a predefined threshold, then the tracklet is considered to be moving (it is probably a player and so it is considered inside the court).
[0160] In addition, according to a preferred embodiment, the determination of the global tracklet position is made by calculating the number of individual poses in the court and the number of individual poses out of the court, and comparing them as follows: the global tracklet position is considered in-court if the total number of individual poses incourt is greater than a fixed threshold and the total number of individual poses out-of-the- court multiplied by a safety factor (e.g., equal to 2 or 3); the global tracklet position is considered out-of-court if the total number of individual poses out-of-court is greater than a fixed threshold and the total number of individual poses in-court multiplied by a safety factor (e.g., 2 or 3); the global tracklet position is considered ambiguous if neither of the previous two conditions are met.
[0161] Importantly, after this filtering it is possible to obtain results previously thought to be unassociable due to ambiguity.
[0162] For this reason, the tracklet creation block B, the association block P-T and the filtering block F3 are applied in sequence and iteratively until convergence is reached, that is, when no further tracklets can be discarded or merged.
[0163] Next, the player detection block 45 comprises a jersey number reading block R.
[0164] In particular, at this phase tracklets of unidentified players with position certainly on the court, or ambiguous, are available.
[0165] In order to assign an identity to each of the tracklets, the jersey number reading block R is used, comprising an OCR network that is configured to operate on players’ back images to recognize the jersey number.
[0166] Since the poses are associated with each frame of a tracklet, it is possible to extract a bounding box of the back on which to perform an OCR reading to extract a number.
[0167] Reading using the OCR network is done on all poses associated with a tracklet on the court (no reading is done on those with ambiguous poses because it could be a player from the other team), so that multiple readings can be available and any OCR errors can be filtered.
[0168] Also, in the case where the operator O has manually specified the players in the match, a reading by the OCR network is considered correct only if it is part of the pool of known numbers defined by the operator O.
[0169] In addition, for each pose, the reading using the OCR network is actually performed only if the player is placed with his back to the camera 2, so he is not facing and not at an angle. This check is done using the position of the left and right shoulder and based on their distance from each other, specifically: if the two shoulders are too close together, it means the player is at an angle; if the right shoulder is further to the left than the left shoulder, it means the player is placed frontally.
[0170] Finally, the reading block R is configured to assign a number to each tracklet according to the following criterion: the most frequent number among the various tracklet poses with at least two readings is assigned, otherwise the tracklet remains momentarily unidentified.
[0171] A possible architecture of the OCR network used in the jersey number reading block R is shown in Figure 18. According to this possible embodiment, the OCR network is configured to receive as input a crop of the original image, resized to a fixed 64x64 size, which contains the player’s back with the number.
[0172] In addition, the OCR network comprises a backbone B, preferably implemented using RESNET50, to extract the features FT.
[0173] Finally, the OCR network comprises a classification head H for up to two digits, configured to produce as output a probability vector for each digit. Each vector is eleven in length because the ten digits plus the empty character, i.e., no number, are considered.
[0174] The OCR network is trained on an internal dataset of about 5000 images, manually annotated, divided into training set (80%) and test set (20%). During training, the crops are sampled with different shifts so that the number to be read is not always in the center of the image. Preferably, it is used as Adam optimizer with a learning rate value of le-3 until error convergence.
[0175] Once the OCR network has been trained, for each frame of each tracklet where a pose is present, numbers are read and only those predicted with greater confidence than a threshold (0.3) for each digit are taken into account. The number associated with the tracklet will be the number with the most occurrences, considering a minimum of two that is, a tracklet cannot be identified by an individual read, nor by low confidence readings. This is important to avoid making identification errors.
[0176] Unidentified tracklets through the OCR network can still be identified through exclusion logics presented below.
[0177] As shown in Figure 16, the player detection block 45 comprises a final identification block E configured to identify tracklets with no number associated but that belong to the same player.
[0178] The identification block E is executed iteratively until convergence, that is, until at least two tracklets can be joined.
[0179] The identification block E comprises heuristics based on exclusion, geometric and tracking principles.
[0180] Specifically, the identification block E is configured to iteratively perform the following steps in the following order: union of the tracklets with same number previously assigned by the OCR network, if not overlapping in time; principles of exclusion; tracking; association of two tracklets; jersey number assignment.
[0181] With reference to the union of tracklets, if the tracklets are overlapped over time, it means that the reading of one of them is wrong, and in this case only the longer tracklet is considered, while the shorter tracklet is considered as an unidentified tracklet.
[0182] With reference to the principles of exclusion, the following steps are carried out: elimination of all unidentified tracklets that coexist at least once with six other identified tracklets; identification of a tracklet if it sees five other identified tracklets during its lifetime and at these times it is always the only unidentified tracklet that is “certainly in the court” with the considered id (the only one it never “sees”) as the only compatible id; identification of a tracklet if it sees five other identified tracklets during its lifetime, even if there are other unidentified tracklets at the same time that are compatible with the sixth and last identity, in which case the longest tracklet is identified; identification of a tracklet that is “certainly in court” if in a given frame there are exactly 6 tracklets in court (comprising the one being considered) and the one being considered is the only one that has some identity as compatible.
[0183] With reference to the tracking, for each tracklet in court, a linear prediction of the plane position is made, both “in the future” and “in the past” by a versor estimated with the last K or first K points of the tracklet.
[0184] Association occurs if and only if: the distance between the linear prediction and the position of the other tracklet is below a predefined threshold (e.g., 30 px for tracking on image or 30 cm for tracking on plane); the two tracklets considered have the same and only compatible identity in common; there is no ambiguity (must be the only possible association for both tracklets).
[0185] The tracking step is applied first on plane, then considering the metric position of the players and only the “certainly in court” tracklets; and then on the image, therefore considering the position of the players’ heads in pixel coordinates and all tracklets, even those with ambiguous positions. With reference to the step of associating two tracklets, such an association is made if the tracklets are compatible on two levels: temporal compatibility: they must never exist simultaneously in any frame; spatial compatibility: the distance run in meters between the end of one tracklet and the beginning of the other must be below a predefined threshold.
[0186] With reference to the jersey number assignment step, this assignment is made based on the starting position of the players, based on the players on the court in the previous step and based on the pool of players available to the team.
[0187] Eventually, the identification block E may comprise a final step of retrieving the OCR network readings on the remaining unidentified tracklets only on compatible identities, accepting lower reading confidence than before. The action-player matching block 46 of the analysis module 4 (shown in Figure 2) is described below.
[0188] Specifically, the action-player matching block 46 is configured to determine which player performed which action.
[0189] Specifically, the action-player matching block 46 is configured to search for an association based on the spatial proximity (in the image plane) between the coordinate of the head predicted by the neural network NN of the action detection block 41 and the head of the various tracklets extracted from the player detection block 45, only considering the frame in which the action occurs.
[0190] Finally, the post-processing block 47 of the analysis module 4 (shown in Figure 2) is described below.
[0191] In the post-processing block 47, the actions of the various players are compared between the various cameras 2 with various lenses.
[0192] In particular, the post-processing block 47 is configured to detect incompatible situations.
[0193] For example, two actions may occur in two different parts of the court at the same time, and this may happen if an action detection block 41 detected an action that was actually the job of the other camera 2. In this case, the post-processing block 47 is able to detect the error by observing the actions preceding and following the incorrect one.
[0194] In addition, the post-processing block 47 is configured to determine the 2D trajectory of the ball in plane, then its direction and velocity, again in 2D, from the full dynamics of the action on both sides of the court.
[0195] Moreover, having also observed the ball position in all the frames, the post-processing block 47 is configured to make a rough estimate of the 3D trajectory of the ball, i.e., estimate its maximum height point.
[0196] Specifically, the post-processing block 47 is configured to determine an estimate of the 3D ball trajectory by performing the following steps: calculation of the 3D ball position, which in turn comprises the following steps (Figure 19): approximation of the x- and y-coordinate of the ball at the moment of the impact with the player’s hands by homography relating the metric plane of the court to the image plane (described above with reference to the court line detection block 42), taking the position of the player’s feet as an approximation and transforming the coordinate into meters by homography; starting from that coordinate, projection of the ball position on the metric plane of the playing court obtaining the z-coordinate; starting from the 3D position of the ball at the beginning and at the end of a hit and the elapsed time, calculation of the trajectory by performing the following steps: calculation of the vx and vy components of the velocity as the ratio of the distance run on the respective axis to the elapsed time; calculation of the vz component of the velocity from the formula describing parabolic motion on the z axis having as a priori known values the initial and final positions, the time elapsed and the gravitational constant as the value of acceleration.
[0197] Conveniently, by means of the scouting analysis module 6 and through the man-machine interface 7 it is also possible to edit the video clips where the hits and actions analyzed took place, eliminating dead time and keeping only the actions of interest with ball in play.
[0198] Specifically, the man-machine interface 7 is configured to create editing of the clips extracted from the clip detection block 44 depending on the scouting results provided by the scouting analysis module 6.
[0199] For example, it is possible to create an individual video containing all the clips of the serves performed by a certain player. This is possible by selecting a sequence of frames of a few seconds around the exact frame in which the serve is performed. The exact moment at which each action is performed is provided by the action detection block 41, while the identity of the player who performed the action is provided by the player detection block 45 and by the action-player matching block 46. Following this procedure, it is possible to create an editing of many small clips containing the actions defined by the scouting analyses done by the scouting analysis module 6 at the request of the operator O.
[0200] In general, clip editing can refer to any result returned by scouting such as, for example, all high shots at a certain position, all spike points of a team in a set, all actions of a player, or in the case of no filtering, all points played in the match. In all cases, editing has the advantage of compressing the actions by removing all moments when the ball is not in play and thus speeding up the coach’s visual analysis of the match.
[0201] Referring to a possible alternative embodiment of the system 1, schematically shown in Figure 20, the analysis module 4’ comprises a post-processing block 47’ configured to receive as input the data from an action detection block 41’ and from a synchronization block 46’.
[0202] The post-processing block 47’ is also configured to receive as input the output of a court line detection block 42’ (geometry and camera calibration) in order to take advantage of the 3D ball triangulation (block 45’).
[0203] The post-processing block 47’ is configured to produce neat, filtered actions as output that are sent as input to the next clip detection block 48’, which exploits these neat, filtered actions from both videos to split the game actions.
[0204] Next, the output of the clip detection block 48’ constitutes one of the inputs to an action-player matching block 50’, which is the last block of the analysis module 4’. Thus, the flow within the analysis module 4’ is reorganized so that the end of an action (clip detection) is detected by exploiting the filtered action list (post-processing) of the videos V of both cameras 2.
[0205] Also according to this alternative embodiment, the court line detection block 42’ comprises a neural network that does not determine some known points in the court (line intersections), but is configured to determine court lines. This configuration of the court line detection block 42’ allows the model to be able to make greater use of the information on the image (there are more line pixels than intersection pixels) and be more robust to occlusions and inaccuracies.
[0206] Thus, the output maps of the court line detection block 42’ are no longer four (as for the four corners) but five for the following lines: the two side lines, the midcourt line, the end line (court near the camera), and the three-meter line (also court near the camera).
[0207] Starting from predictions of the court line detection block 42’, a robust fitting of lines is made and then the intersections are derived. We then always arrive at the correspondence between known points in the court, but instead of predicting them directly, we derive them from line prediction, making the whole process very robust.
[0208] Furthermore, as an alternative to homography, the court line detection block 42’ comprises a calibration of each camera 2, estimating the intrinsic (focal) and extrinsic parameters of the camera (rotation and shift). Such a calibration allows, in addition to all that homography allows, also to relate 3D elements out of the plane of the court and, thus, allows, for example, to triangulate the position of the ball through similar information also processed from the video V of the other camera 2.
[0209] In addition, multiple frames are always selected, but instead of performing independent predictions, the court line detection block 42’ is configured to first perform a median to join the frames and to subsequently make an individual prediction. The pixel-to-pixel median operation ensures that occasionally occluded elements by players are well displayed in the final result. In fact, the resulting image tends to have no players in the scene, which allows the network to work under the best working conditions.
[0210] The dataset and training procedures remain unchanged from the first embodiment described above. Still referring to the possible further embodiment illustrated in Figure 20, the post-processing block 47’ is configured to triangulate the 3D ball trajectory starting from the 2D detections on each video stream V and taking advantage of the relevant calibration of the two cameras 2. Specifically, the 3D ball is found only if a 2D detection is present in both videos for that frame, and then the detections are interpolated and filtered.
[0211] Still with reference to this additional embodiment, the clip detection block 48’ is configured to split the match into individual action clips, jointly considering the actions occurred and found on both sides of the court.
[0212] In particular, the poses of the players at the phase of splitting the video into clips are also considered. When events are detected, there is an option to create a game action or to discard such events as false positives.
[0213] Specifically, the following criteria are evaluated and weighed to determine whether the ball is in play or not:
[0214] - number of players on the court;
[0215] - uniform position of players on the court;
[0216] - players looking in the right direction;
[0217] - uniqueness of the ball in play.
[0218] Still referring to the possible further embodiment, the analysis module 4’ comprises a player detection block 49’ that receives as input the outputs of the court line detection block 42’, the clip detection block 48’, the head detection block 43’ and the pose estimation block 44’.
[0219] According to this possible embodiment, the player detection block 49’ is configured to perform a forced split of the tracklets with the same id that overlap due to incorrect OCR readings. This allows the OCR to be more robust and never have situations in the court that would definitely be wrong.
[0220] In addition, the analysis module 4’ comprises an action-player matching block 50’ configured to search not only in the same frame, but in a neighborhood of K (K=+-5) frames and among all players closest to the ball by a threshold (50px), considering the one closest overall.
[0221] It has in practice been ascertained that the described invention achieves the intended objects.
[0222] In particular, the fact is emphasized that the system according to the invention enables the automatic acquisition and processing useful information related to the game dynamics during a volleyball match effectively, quickly, normalized and with limited costs.
Claims
CLAIMS1) Computer-implemented system (1) for the automatic analysis of a volleyball match, particularly employable for scouting in volleyball, characterized by the fact that it comprises: at least one camera (2) positioned in the proximity of a volleyball court and oriented to capture at least one video (V) of the players within the court; at least one processing platform (3), comprising at least one remote and / or local processing unit, operationally connected to said at least one camera (2) and configured to process one or more videos (V) captured to obtain data related to the game dynamics during a volleyball match; wherein said processing platform (3) comprises: at least one analysis module (4), configured to process one or more videos (V) captured by means of said at least one camera (2) and to generate metadata related to the game dynamics during the volleyball match; at least one scouting analysis module (6) configured to provide on-demand the analysis data processed from said generated metadata to an operator (O).2) System (1) according to claim 1, characterized by the fact that it comprises two cameras (2) arranged in the proximity of respective opposite short sides of the court, substantially central to the short sides, and oriented to capture two respective videos (V) of the players within the respective halves of the court.3) System (1) according to one or more of the preceding claims, characterized by the fact that it comprises a human-machine interface (7) configured to manage the uploading of said videos (V) directly from said at least one camera (2) and / or indirectly from an operator (O) and configured to interface an operator (O) with said scouting analysis module (6).4) System (1) according to one or more of the preceding claims, characterized by the fact that, for each camera (2) used, said analysis module (4) comprises an action detection block (41) configured to receive as input the entire video (V) captured by said camera (2) and to generate as output the following data: for each ball hit detected in said video (V): time instant, pixel coordinates of the ball touch, pixel coordinates of the head of the player who hit the ball, type of fundamental and relevant specialization; for each frame of said video (V): pixel coordinates of the ball.5) System (1) according to one or more of the preceding claims, characterized by the fact that, for each camera (2) used, said analysis module (4) comprises a court line detection block (42), configured to: receive as input a manual selection of the court lines made by an operator (O) or N framesfrom said video (V) of said camera (2); generate as output the pixel coordinates of the intersections of the playing lines of the half court near said camera (2).6) System (1) according to one or more of the preceding claims, characterized by the fact that, in case at least two cameras (2) are used, said analysis module (4) comprises a synchronization block (43) configured to: receive as input a manual alignment of the videos V acquired from each of said cameras (2), performed by an operator (O) or the outputs of each action detection block (41) related to each of said cameras (2); generate as output the time difference in seconds between said videos (V).7) System (1) according to one or more of the preceding claims, characterized by the fact that said analysis module (4) comprises a clip detection block (44) configured to: receive as input the output of said action detection block (41) and, if any, of said synchronization block (43); generate as output the following data, for each clip in the match: clip start instant, consisting of the instant of the serve that starts the point; clip end instant, consisting of the instant of the last hit that closes the point.8) System (1) according to one or more of the preceding claims, characterized by the fact that, for each camera (2) used and for each generated clip, said analysis module (4) comprises a player detection block (45) configured to: receive as input the following data: the entire video (V) captured by each of said cameras (2), the clip start and end instants and the lines of the court near each of said cameras (2); generate as output the following data: for each player, their jersey number, if visible; for each frame, the metric position of the player on the court and the pixel coordinates of the head of the player.9) System (1) according to one or more of the preceding claims, characterized by the fact that said analysis module (4) comprises an action-player matching block (46) configured to: receive as input the following data: the output of said action detection block (41); the output of said player detection block (45); the clip start and end instants generated by said clip detection block (44); generate as output which player performed which action.10) System (1) according to one or more of the preceding claims, characterized by the fact that for each of said clips, said analysis module (4) comprises a post-processing block (47) configured to: receive as input the following data: the clip start and end instants; for each camera (2), theoutput of said action detection block (41), the output of said player detection block (45), the output of said action-player matching block (46); generate as output an improvement in the actions recognized within each clip.11) System (1) according to one or more of the preceding claims, characterized by the fact that said action detection block (41) is configured to perform at least the following steps: an analyzing step (103) of said video (V) by means of a neural network (NN) for the extraction of said actions, in which said neural network (NN) is configured to receive as input said video (V) and to produce as output the following types of output: hit classification output (104) to determine the fundamental of the action performed; hit location output (105); player location output (106) to determine the position of the head of the player who performed the action; ball location output (107) to determine the position of the ball during the performance of the point; a post-processing step ( 108) of the outputs of said neural network (NN) to determine the actual presence and location of the various actions, comprising: application of thresholds to discretize said hit classification output (104) and said hit location output (105); clustering into classification-related components (109) and location-related components (110); an association step (111) of said classification-related components (109) with said location- related components (110) on the basis of temporal proximity, to obtain complete predicted actions comprising location and touch type; an action-player association step (112), starting from said complete predicted actions and from said player location output (106) generated by the neural network (NN); a determination step (113) of the location of the ball at each frame, starting from said ball location output (107).12) System (1) according to one or more of the preceding claims, characterized by the fact that said neural network (NN) of the action detection block (41) consists of: a central part (B) configured to extract features (FT) from frame blocks (FR) as input and characterized by spatiotemporal operators; and an end part (T) comprising a hit classification head (Hl), a hit location head (H2), a player location head (H3) and a ball location head (H4).13) System (1) according to one or more of the preceding claims, characterized by the fact that said court line detection block (42) is configured to perform the following steps: selection of a predefined number of frames per video; using a neural network, determination of the angles (A) on all the selected frames; execution of the average of said angles (A) determined between the various frames, excluding predictions too different from the average;having obtained the average angles, homography to switch from the image plane to the court metric plane.14) System (1) according to one or more of the preceding claims, characterized by the fact that said synchronization block (43) is configured to perform the following steps: determining the first hits from each side of the court to have a first alignment between said two videos (V); around said first alignment, determination of several possible alignments with increasing time delta and, for each alignment, evaluation of the time compatibility of the following combinations of fundamentals: spike-block, serve-forearm pass; of all the evaluated alignments, select the alignment with the greatest compatibility.15) System (1) according to one or more of the preceding claims, characterized by the fact that said clip detection block (44) is configured to cycle through the actions predicted by said neural network (NN) of action detection block (41) in order of occurrence, by performing the following steps: keep a state that indicates whether you are during an action or not, initially the state will be on “no action”. the state switches from “no action” to “action” if the current state is “no action” and if any action occurs but other than touching the ball on the ground; the state switches from “action” to “no action” if the current state is “action” and if: both in the case of one camera (2) and two cameras (2) being used: a ball touch action occurs on the ground in the half of the court near the camera; in this case the action is over; in the case of using just one camera (2): no action occurs for a period of time longer than a certain threshold of seconds; in which case it means that the other team failed to send the ball back into the court near the camera and, therefore, the point is over; in the case of using two cameras (2): a ball touch action occurs in the opponent’s half of the court; in which case it means that the action is over.16) System (1) according to one or more of the preceding claims, characterized by the fact that said player detection block (45) comprises: a head detection (HD) neural network configured to identify the head of the players; a pose estimation (PE) neural network configured to identify all joints of the players; a tracklet creation block (T) starting from the heads identified by said head detection (HD) neural network; an association block (P-T) of all the poses extracted from the pose estimation (PE) network with said created tracklets; a reading block (R) of the player jersey number.17) System (1) according to one or more of the preceding claims, characterized by the fact that said action-player matching block (46) is configured to search for an association based on the proximity between the coordinate of the head predicted by said neural network (NN) of the action detection block (41) and the head of the various tracklets extracted from said player detection block (45), considering only the frame in which the action occurs.18) System (1) according to one or more of the preceding claims, characterized by the fact that it comprises an analysis module (4’) comprising a post-processing block (47’) configured to receive as input the data coming from an action detection block (41’), from a synchronization block (46’) and from a court line detection block (42’), wherein said post-processing block (47’) is configured to produce neat, filtered actions as output that are sent as input to the next clip detection block 48’, which exploits these neat, filtered actions from both videos to split the play actions and wherein the output of said clip detection block (48’) constitutes one of the inputs of an action-player matching block (50’).
Citation Information
Patent Citations
Ball game video analysis device and ball game video analysis method
US20210150220A1
Performance interactive system
US20230196770A1