System and method for implementing high-dimensional spatial data feature sets in sports prediction

By using a generative adversarial network (GAN) model and a robust feature set, the problem of prediction accuracy of traditional models under abnormal conditions is solved, and more accurate action probability prediction is achieved in sports events.

CN121586903APending Publication Date: 2026-02-27STAT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480045664.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-23
Filing Date
2024-08-26
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict the probability of actions in abnormal or marginal situations during sporting events, especially when the goalkeeper is poorly positioned or other players interfere, where traditional models exhibit low prediction accuracy.

Method used

By using a generative adversarial network (GAN) model and incorporating a robust feature set, including details such as the position of opposing players, the distance and angle between the defender and the goal, the initial prediction probabilities are modified to generate more accurate updated predictions.

Benefits of technology

It improves prediction accuracy in abnormal or marginal situations, especially in sports events such as football and tennis, enabling more accurate prediction of the probability of a shot or hit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121586903A_ABST
    Figure CN121586903A_ABST
Patent Text Reader

Abstract

A method of generating a probability of a first action of a sporting event by implementing a set of features, the method comprising: obtaining an initial data set related to the first action of the sporting event, the initial data set at least comprises the position of a first player on the field surface and the position of a target area on the field surface; generating an initial prediction score probability based on the initial data set through a machine learning model; generating a feature set related to the sports event; and modifying, by the machine learning model, the initial predicted score probability to an updated score probability using the feature set.
Need to check novelty before this filing date? Find Prior Art

Description

Cross Reference to Related Applications

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 578,741, filed August 25, 2023, the entire contents of which are incorporated herein by reference for all purposes. TECHNICAL FIELD

[0002] Various aspects of the present disclosure generally relate to machine learning for sports applications. In particular, various aspects relate to systems and methods for implementing a high-dimensional spatial data feature set in sports prediction. BACKGROUND

[0003] As sports have become increasingly popular, there is an increasing desire for accurate granular predictions of events during sporting events. For example, predicting the probability of a goal, a win, a point, or predicting a shot location for a sport can be of particular interest to media members, broadcasters (whether primary broadcast or second screen experience), and fans, sports betting, and fantasy / gaming applications. Moreover, in many cases, such predictions / probabilities can be influenced by detailed in-game statistics.

[0004] Unless otherwise indicated herein, the description herein of materials described in this section is not an admission that any or all of the described material is or was part of the prior art, or was common general knowledge, and the description herein of prior art or common general knowledge is not an admission of any description in prior art documents or common general knowledge. SUMMARY

[0005] In some aspects, the technology described herein relates to a method of generating a probability of a first action of a sporting event by implementing a feature set, the method comprising: obtaining an initial data set related to the first action of the sporting event, the initial data set comprising at least a position of a first player on a playing surface and a position of a target area on the playing surface; generating, by a machine learning model, an initial predicted scoring probability based on the initial data set; generating a feature set related to the sporting event, the feature set derived from any of: a position of a second player on the playing surface, and at least one of a distance or an angle between any two of the first player, the second player, and the target area on the playing surface; and modifying, by the machine learning model, the initial predicted scoring probability to an updated scoring probability using the feature set.

[0006] In some aspects, the technology described herein relates to a method wherein the machine learning model comprises a generative adversarial network (GAN) model, and wherein the GAN model comprises at least one of a set of monotonic constraints or a weight of each of the monotonic constraints.

[0007] In some aspects, the technology described herein relates to a method wherein the feature set comprises a previous action type specifying a previous action performed prior to the first action occurring.

[0008] In some aspects, the technology described herein relates to a method, wherein any of the initial data set and the feature set includes a historical save probability of the second player.

[0009] In some aspects, the technology described herein relates to a method, wherein the first player and the second player belong to opposing teams in the sporting event.

[0010] In some aspects, the technology described herein relates to a method, wherein the feature set includes: a virtual line pointing from the location of the first player and the location of the goal area; and a distance between the second player and a region of the virtual line.

[0011] In some aspects, the technology described herein relates to a method, wherein the feature set includes a save probability profile of the second player that modifies, for each of a set of shot locations at different angles from the virtual line between the location of the first player and the location of the goal area, a save probability.

[0012] In some aspects, the technology described herein relates to a method of generating a predicted save probability for a first action of a sporting event using a feature set, the method comprising: obtaining an initial data set related to the first action of the sporting event, the initial data set including at least a location of an attacking player on a field and a location of a goal on the field; generating, by a machine learning model, an initial predicted save probability based on the initial data set; generating a feature set related to the sporting event, the feature set derived from any of: a location of a defending player on the field, a distance or an angle between at least one of the defending player, the attacking player, or the goal, or an existence of one or more additional players in the vicinity of the defending player, the attacking player, or the goal; and modifying, by the machine learning model, the initial predicted save probability to an updated predicted save probability using the feature set.

[0013] In some aspects, the technology described herein relates to a method, wherein the feature set includes: a virtual line pointing from the location of the attacking player and the location of the goal; and a distance between the location of the defending player and the virtual line between the location of the attacking player and the location of the goal.

[0014] In some aspects, the technology described herein relates to a method, wherein the feature set includes a goalkeeper save probability profile that modifies, for each of a set of shot locations at different angles from the virtual line between the location of the attacking player and the location of the goal, a goalkeeper save probability.

[0015] In some aspects, the technology described herein relates to a method, wherein the feature set includes a clarity value that specifies a number of interfering players positioned in the vicinity of the virtual line between the location of the attacking player and the location of the goal.

[0016] In some aspects, the technology described herein relates to a method, wherein the set of features includes a previous action type that specifies a previous action performed prior to the first action occurring.

[0017] In some aspects, the technology described herein relates to a method, wherein the first action includes a shot on goal, the sporting event includes a soccer match, and the defending player is a goalkeeper.

[0018] In some aspects, the technology described herein relates to a method, wherein the set of features includes a shot archetype, the shot archetype specifying one or more of: a shot type, whether the shot is under pressure, a distance range of the shot, and / or whether the shot is from a central portion of the field or a side portion of the field.

[0019] In some aspects, the technology described herein relates to a method, wherein the machine learning model includes a generative adversarial network (GAN) model, and wherein the GAN model includes at least one of a set of monotonic constraints or a weight of each of the monotonic constraints.

[0020] In some aspects, the technology described herein relates to a method, wherein any of the initial data set or the set of features includes a historical save probability of the defending player.

[0021] In some aspects, the technology described herein relates to a method of generating a predicted score win probability for a shot in a tennis match using a set of features, the method comprising: obtaining an initial data set related to a first action of a sporting event, the initial data set including at least a location of a first player on a court surface and a location of a target area on the court; generating, by a machine learning model, an initial predicted score win probability based on the initial data set; generating a set of features related to the sporting event, the set of features derived from any of: a location of a second player on the court, a distance and / or an angle between the second player, the first player, and the target area, or a presence of any additional players in the vicinity of the second player, the first player, and the target area; and modifying, by the machine learning model, the initial predicted score win probability to an updated predicted score win probability using the set of features.

[0022] In some aspects, the technology described herein relates to a method, wherein the machine learning model includes a generative adversarial network (GAN) model, and wherein the GAN model includes at least one of a set of monotonic constraints or a weight of each of the monotonic constraints.

[0023] In some aspects, the technology described herein relates to a method, wherein the set of features includes: a virtual line between the location of the first player and the location of the target area; and a distance between the location of the second player and the virtual line between the location of the first player and the target area.

[0024] In some aspects, the technology described herein relates to a method, where the feature set includes shot trajectory data, 3D body pose information of the first player and the second player, current shot count, speed and acceleration information of the first player and the second player, court surface type, and / or shot type. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various example aspects and together with the description, explain the principles of the disclosed aspects.

[0026] FIG. 1A is a block diagram of an example tracking and analysis environment in accordance with one or more embodiments.

[0027] FIG. 1B is a block diagram of a prediction environment in accordance with one or more embodiments.

[0028] FIG. 2 is a flowchart of an example method for generating probabilities of actions of a sporting event by implementing a feature set in accordance with one or more embodiments.

[0029] FIG. 3 depicts an example view of a goalkeeper and attacking players in a soccer game in accordance with one or more embodiments.

[0030] FIG. 4A to FIG. 4C shows a representation of goalkeeper save probabilities based on shot location in accordance with one or more embodiments.

[0031] FIG. 5A to FIG. 5B shows graphical representations of feature curves, calibration points, precision-recall curves, and log loss of goalkeeper operations for shots based on various shot locations in accordance with one or more embodiments.

[0032] FIG. 6A to FIG. 6B shows an example representation of indicators arranged by relative position in accordance with one or more embodiments.

[0033] FIG. 7A to FIG. 7B shows various representations of calibrating a model across multiple model versions in accordance with one or more embodiments.

[0034] FIG. 8 is a graphical representation of predicted goal save probabilities versus actual conversion rates for various teams in a game season in accordance with one or more embodiments.

[0035] FIG. 9 is a graphical illustration of a comparison indicator in accordance with one or more embodiments.

[0036] FIG. 10A to FIG. 10B shows an example view of a shot and a graph specific to that shot in accordance with example embodiments.

[0037] FIG. 11A to FIG. 11B An example view of a shot and a shot-specific graph are shown in accordance with one or more embodiments.

[0038] FIG. 12 An example shot graph is shown in accordance with one or more embodiments.

[0039] FIG. 13 is an example heat map output of a predicted shot in accordance with one or more embodiments.

[0040] FIG. 14 An example output of a prediction model when applied to an example tennis match is depicted in accordance with one or more embodiments.

[0041] FIG. 15 A flowchart for training a machine learning model is depicted in accordance with one or more embodiments.

[0042] FIG. 16 An example computing device is depicted in accordance with one or more embodiments.

[0043] It is noteworthy that certain aspects of the drawings can depict the general configuration of various embodiments. Descriptions and details of the known features and techniques can be omitted in order to avoid unnecessarily obscuring other features of the present disclosure. The elements in the figures are not necessarily drawn to scale. The dimensions of some features can be exaggerated relative to other elements in the figure in order to improve the understanding of the example embodiments. DETAILED DESCRIPTION

[0044] Various aspects of the present disclosure generally relate to machine learning for sports applications. In particular, various aspects relate to systems and methods for generating a predicted save probability for an action of a sports event using a set of features.

[0045] One or more embodiments disclosed herein can provide systems and methods for generating probabilities of actions of a sporting event by implementing a machine learning model. The systems described herein can first utilize an initial dataset of a sporting event to generate initial probabilities of actions of the sporting event (e.g., scoring probabilities). The initial dataset can include at least a location of a first player on a playing surface / field and a target area (e.g., a scoring area of a sport with designated scoring areas, a given location on a playing surface, a given portion of such a playing surface, a given location likely to result in a score, etc.) on the playing surface (e.g., field). Additionally, a feature set related to the sporting event can be generated. The feature set can include more nuanced information for the sporting event, such as a location of an opposing player, a first player’s line of sight clarity, a shot context, etc. The feature set can include detailed information about the location of the opposing player relative to the target area and the first player. The initial predicted probabilities of actions can be modified or enhanced by a machine learning model based on the inputted feature set. This data augmentation with the feature set can allow for edge case predictions to be considered, for example, when the opposing player location is poor, and allow for more accurate overall predictions from the machine learning model.

[0046] As used herein, a “machine learning model” generally encompasses instructions, data, and / or a model configured to receive input and apply one or more of weights, biases, classifications, or analyses to the input to generate an output. For example, the output can include a classification of the input, an analysis based on the input, a design, a process, a prediction, or a recommendation associated with the input, or any other suitable type of output. Machine learning models are typically trained using training data, e.g., samples of experience data and / or input data, that is fed into the model in order to establish, adjust, or modify one or more aspects of the model, e.g., weights, biases, criteria for forming classifications or clusters, etc. Various aspects of a machine learning model can operate on input linearly, in parallel, or via any suitable configuration of networks, e.g., neural networks.

[0047] Execution of a machine learning model can include deploying one or more machine learning techniques, such as linear regression, logistic regression, random forest, gradient boosting machine (GBM), deep learning, and / or deep neural networks. Supervised and / or unsupervised training can be employed. For example, supervised learning can include providing training data and labels corresponding to the training data, e.g., as ground truth. Unsupervised approaches can include clustering, classification, etc. K-means clustering or K-nearest neighbors can also be used, which can be supervised or unsupervised. A combination of K-nearest neighbors and unsupervised clustering techniques can also be used. Any suitable type of training can be used, e.g., random, gradient boosting, random seeding, recursive, round-robin, or batch-based, etc.

[0048] While several of the examples herein relate to certain types of machine learning, it should be understood that the techniques in accordance with the present disclosure can be adapted to any suitable type of machine learning. It should also be understood that the above examples are illustrative only. The techniques and methodologies of the present disclosure can be adapted to any suitable activity.

[0049] Machine learning techniques can be used to measure player performance in a sporting event (e.g., soccer) using detailed data related to the player / team. One key metric in soccer can be related to goal scoring performance (xG), and this metric can use a small set of features (e.g., shot location, previous action, big opportunity flag). However, models using such a small feature set can not adequately capture specific circumstances related to a shot, such as when the goalkeeper is poorly positioned, or whether other players are blocking the goalkeeper’s view of the shot.

[0050] The accuracy of the predicted metric can be improved with more detailed information, such as goalkeeper position, defensive pressure, shot clarity (i.e., whether the forward can see the goal or whether the player is facing away from the goal), shot context (i.e., header, left-footed volley, right-footed volley), etc. The systems and methods described herein can incorporate more detailed information in order to achieve more accurate predictions. This can include implementing input features with high dimensionality (e.g., greater than a threshold number of inputs, such as, for example, greater than thirty inputs).

[0051] Machine learning models that implement input features with high dimensionality (e.g., neural networks and tree-based methods such as decision forests and boosting trees) can have issues with leakage or overfitting. Traditional machine learning model systems can implement a carefully curated training dataset, a development dataset, and a test dataset, and through regularization, can avoid overfitting. However, many traditional systems incorporate imbalanced datasets (i.e., where the training data can be biased, such as 10% of shots resulting in a goal) and many examples (e.g., hundreds of thousands of training examples). These traditional systems can not have a high success rate in determining anomalies or exceptional scenarios in a sporting event.

[0052] Traditional systems can incorporate various models in an attempt to address exceptional scenarios in a sporting event. For example, these traditional techniques can include applying a simple model (e.g., a logistic regressor), a model with a small number of features, or a separate model for exceptional cases (i.e., a different model when the goalkeeper is far from the goal), and the model can generate noisy outputs given that the case is exceptional, as the scenario can not occur often.

[0053] The systems and methods described herein can more accurately generate one or more predictions by determining and incorporating robust features based on detailed sports data to improve upon the above limitations of traditional systems. In an example use case, the method can be used for soccer, tennis, hockey, volleyball, etc.

[0054] An example scenario in which the systems and methods described herein can be implemented is a soccer game when the goalkeeper is outside the penalty area and when a player of the opposing team is in possession of the ball on the half of the goalkeeper's team. Because the position of the goalkeeper is only a single variable, and this situation can not occur often in training data, traditional systems can not determine that this data point is valuable in determining the probability of an action (e.g., the probability of a player scoring a goal). In this scenario, traditional models have lower accuracy in predicting the opportunity to score a goal, resulting in lower reliability of these traditional models.

[0055] The systems and methods described herein can include a machine learning system that incorporates a feature dataset (i.e., additional data points), such as the position of the opposing player. For example, when implementing the systems and methods described herein with respect to soccer, the feature dataset can include a goalkeeper position feature, allowing the system to more accurately predict the likelihood of scoring a goal in an unusual scenario of a soccer game. In an example use case, the systems and methods described herein can improve upon the above limitations of traditional systems by creating a robust goalkeeper feature based on the arrangement of different positions of the goalkeeper. This type of sport-specific data augmentation can allow for more accurate predictions in edge case scenarios.

[0056] For example, a machine learning model can use an initial dataset related to an action of a sporting event (e.g., a goal in soccer) to generate an initial prediction save probability. The initial dataset can include information related to the position of the ball and / or the striker relative to the goal. However, using only the initial dataset to generate such metrics can result in inaccurate predictions, particularly in unusual situations (e.g., whether the goalkeeper is in a position to block the shot, whether the goalkeeper has a line of sight to the shot).

[0057] In response, the systems described herein can determine a feature set related to the sporting event. The feature set can include a robust set of details related to the action. The feature set can stem from any of the position of the defending player on the field, the distance and / or angle between any two of the defending player, the attacking player, and the goal, and the presence of any additional players in the vicinity of the goal. The machine learning model can use the feature set to modify the initial prediction save probability to an updated prediction save probability. The updated model implementing the feature dataset can have more accurate predictions, particularly in specific scenarios that do not occur often in the sporting event.

[0058] The systems and methods described herein can also relate to other sports (e.g., other team or individual sports). The systems and methods can be incorporated to assist in generating predictions of unusual actions in a sporting event. An example use case can include when a player in tennis makes a high-pressure smash on an open court but does not succeed, or when a goalie in hockey leaves the goal and the attacking team has an open shot on goal. Another example can include basketball, where a shot can be tipped or even intercepted. Rather than creating specific models for these anomalies, which can be impractical or resource-intensive, robust features that capture the nuances of a particular sport can be incorporated and utilized to determine accurate predictions. The arrangement methods described herein can address such issues.

[0059] One or more embodiments described herein can determine robust features that both capture the nuances of a sport and address unusual situations. Unusual situations, such as determining the probability of a save when the goalie is not in the goal, can be arranged with an initial prediction by implementing additional nuanced data in the feature set. The arrangement can be done exhaustively with operator feedback across all positions, or can be done with a machine learning model (e.g., a generative adversarial network (GAN)) approach. When the features are calculated, the features can then be fed back into the model to essentially “regularize” for these situations (i.e., the goalie leaving the goal).

[0060] In an example case, the models described herein can quantify the quality of each opportunity in a match by estimating the likelihood of a successful event (e.g., a shot on goal, a shot on score, etc.). A feature xG can model the likelihood of a successful event (e.g., a goal or a score) at the instant the ball is struck, while an expected shot on goal (xGoT) can model the likelihood of hitting the goal target (or being saved) by including information about the shot’s landing point and trajectory.

[0061] FIG. 1A is a block diagram illustrating a tracking and analysis environment 100 according to an example aspect of the present disclosure. The environment 100 includes a tracking system 102, a computing system 104, and a client device 108 connected via a network 105. In the depicted example, the tracking system 102 obtains various attributes of a match and transmits measurements across the network 105 to the computing system 104, where the measurements can be used in conjunction with one or more machine learning models. In one example, the tracking system 102 can be configured to obtain initial, simpler data (e.g., player positions, ball position) in addition to more nuanced information (e.g., ball / player distances and angles, save probabilities, field of view clarity, shot context, and / or previous actions). According to one example, the nuanced information can be a subset or detail related to the simpler data.

[0062] The tracking system 102 can be positioned in, adjacent to, or near the venue 106. Non-limiting examples of the venue 106 include a stadium, a court, a site, and a field. The venue 106 includes agents 112A-N (e.g., players, objects, officials, etc.). The tracking system 102 can be configured to record the motion and actions of the agents 112A-N on the playing field, which can include related objects (e.g., a ball, a referee, etc.). Although the environment 100 generally depicts the agents 112A-N as players, it should be understood that, according to certain embodiments, the agents 112A-N can correspond to players, objects, markers (e.g., playing field markers), officials, etc.

[0063] The tracking system 102 and / or the data storage 118 can be configured to measure and / or record various statistics related to the sporting event. These can include, but are not limited to: player and ball positions; player distance and angle relative to a target area; defensive player save or rebound probability; pressure exerted by the opposing team’s players (e.g., distance between the opposing team’s players and the ball); clarity (e.g., view of the ball by the player in possession relative to the target area); shot context (e.g., shot type, such as a free kick, a header, a volley, a first touch, or a one-on-one in soccer); and previous action (e.g., a corner kick, a throw-in, a direct pass, a fast break, or a bounce ball in soccer). Additional statistics recorded can include shot trajectory (e.g., shot impact location, speed, spin, or three-dimensional position and time); player position (e.g., three-dimensional position of all players); current shot count (e.g., shot count in a rally in tennis); player speed and acceleration; court surface (e.g., court surface in tennis); game state (e.g., score or set number in tennis); shot type (e.g., serve, return, forehand, backhand, lob, volley, or underarm shot in tennis); temperature and wind information; time of day; and / or player attributes (e.g., skill rating for a particular shot, such as serve, return, forehand, backhand, quality rating).

[0064] In some aspects, the tracking system 102 can be an optical-based system using, for example, cameras 103. While one camera is depicted, additional cameras are possible. For example, a system of six fixed, calibrated cameras can be used that project the three-dimensional positions of the players and the ball onto a two-dimensional overhead view of the court. Additional tracking can be based on a radio-based system that uses, for example, radio frequency identification (RFID) tags worn by the players or embedded in objects to be tracked.

[0065] In another example, a mix of fixed and non-fixed cameras can be used to capture the motion of all agents 112A-N on the playing field, as well as the motion of one or more related objects. Utilizing such a tracking system (e.g., tracking system 102) can produce a number of different camera perspectives of the court (e.g., high sideline perspective, penalty line perspective, player huddle perspective, scrum perspective, end zone perspective, etc.). In some aspects, the tracking system 102 can be used for a broadcast feed of a given match. In such aspects, each frame of the broadcast feed can be stored in a match file.

[0066] The tracking system 102 can be configured to communicate with the computing system 104 via the network 105. The computing system 104 can be configured to manage and analyze data captured by the tracking system 102. The computing system 104 can include a web client application server 114, a pre-processing agent 116 (e.g., a processor), a data store 118, a predictor 126, and a third-party application programming interface (API) 138. Related to FIG. 16 One example of the computing system 104 is depicted.

[0067] The pre-processing agent 116 can be configured to process data retrieved from the data store 118 or the tracking system 102 prior to input to the predictor 126. The pre-processing agent 116 and / or the predictor 126 can include one or more software modules. The one or more software modules can be a collection of code or instructions stored on a medium (e.g., a memory of the computing system 104) that represents a series of machine instructions (e.g., program code) that implement one or more algorithmic steps. Such machine instructions can be actual computer code that is parsed by a processor of the computing system 104 to implement the instructions, or alternatively, can be a higher level of instruction coding that is parsed to obtain the actual computer code. The one or more software modules can also include one or more hardware components. One or more aspects of an example algorithm can be performed by the hardware components (e.g., circuitry) themselves, rather than as a result of instructions.

[0068] The data store 118 can be configured to store different kinds of data (e.g., in one or more formats). In one example, the data store 118 can store raw tracking data received from the tracking system 102. The data store 118 can include historical match data, in-match data, outputs derived from any of the models described herein, and / or non-match data, such as player data, injury data, training data, etc. Historical data can include aggregated in-match statistics (e.g., a goalkeeper’s save percentage or a tennis player’s return percentage for various types of shots).

[0069] The predictor 126 includes one or more machine learning models 128A-N. The one or more machine learning models can include a generative adversarial network (GAN) and / or a tree-based method (e.g., a decision forest and a boosting tree). The one or more machine learning models 128A-N can be configured to receive one or more features and output a prediction related to a sporting event. The output can include a predicted goal scorer, a shot winner, a shot trajectory, and / or a shot location.

[0070] The client device 108 can communicate with the computing system 104 via the network 105. The client device 108 can be operated by a user. For example, the client device 108 can be a mobile device, a tablet, a desktop computer, or any computing system having the capabilities described herein. The user can include, but is not limited to, an individual, such as, for example, a subscriber, client, potential client, or customer of an entity associated with the computing system 104, such as an individual who has obtained, will obtain, or is likely to obtain a product, service, or consultation from the entity associated with the computing system 104.

[0071] The client device 108 can include one or more applications 109. The applications 109 can represent a web browser that allows access to a website or can be standalone applications. The client device 108 can access the applications 109 to access one or more functions of the computing system 104. The client device 108 can communicate over the network 105 to, for example, request a web page from the web client application server 114 of the computing system 104. For example, the client device 108 can be configured to execute the applications 109 to access content managed by the web client application server 114. Content displayed to the client device 108 can be transmitted from the web client application server 114 to the client device 108 and subsequently processed by the applications 109 to be displayed through a graphical user interface (GUI) of the client device 108.

[0072] The client device can include a display 110. Examples of the display 110 include, but are not limited to, a computer display, a light-emitting diode (LED) display, and the like. Output or visualizations generated by the applications 109 can be displayed on the display 110.

[0073] The functionality of the sub-components shown in computing system 104 can be implemented in hardware, software, or some combination thereof. For example, software components can be a set of codes or instructions stored on a medium such as a non-transitory computer-readable medium (e.g., a memory of computing system 104) that when executed by a processor of computing system 104, cause the processor to perform a series of machine-readable instructions (e.g., program code) to implement one or more method operations. Such machine instructions can be actual computer code that is resolved by the processor of computing system 104 to implement the instructions, or alternatively, can be higher level instruction code that is resolved to obtain the actual computer code. One or more software modules can also include one or more hardware components. Examples of components include processors, controllers, signal processors, neural network processors, and the like.

[0074] Network 105 can be any suitable type, including individual connections via the Internet, such as cellular or Wi-Fi networks. In some aspects, network 105 can connect terminals, services, and mobile devices using direct connections, such as radio frequency identification (RFID), near field communication (NFC), Bluetooth™, Bluetooth™ Low Energy (BLE), Wi-Fi™, ZigBee™, Ambient Backscatter Communication (ABC) protocols, USB, WAN, or LAN. As the information transmitted can be personal or confidential, security concerns can require one or more of these types of connections to be encrypted or otherwise secured. However, in some aspects, the information transmitted can not be as personal and, as such, network connections can be selected for convenience rather than security.

[0075] Network 105 can include any type of computer networking arrangement for exchanging data or information. For example, network 105 can be the Internet, a private data network, a virtual private network using a public network, and / or other suitable connections that enable components in computing environment 100 to send and receive information between components of environment 100.

[0076] FIG. 1B is a block diagram of a prediction environment 101 according to one or more embodiments. Prediction environment 101 can be configured to predict one or more statistics related to a sporting event. Prediction environment 101 can include an initial input generator 160, a feature set generator 164, and a prediction model 162. One or more of the modules of prediction environment 101 (e.g., initial input generator 160, feature set generator 164, and prediction model 162) can be located within computing system 104. FIG. 1A The modules of prediction environment 101 can depict the components and modules of computing system 104 described in FIG. 1A The modules of prediction environment 101 can depict the components and modules of computing system 104 described in

[0077] Each of the initial input generator 160, the feature set generator 164, and the prediction model 162 can include one or more software modules. The one or more software modules can be a set of code or instructions stored on a medium (e.g., a memory of the organization computing system 104) that represents a series of machine instructions (e.g., program code) that implement one or more algorithmic steps. Such machine instructions can be actual computer code that is parsed by a processor of the organization computing system 104 to implement the instructions, or alternatively, can be higher level instruction coding that is parsed to obtain the actual computer code. The one or more software modules can also include one or more hardware components. One or more aspects of the example algorithms can be performed by the hardware components (e.g., circuitry) themselves, rather than as a result of instructions.

[0078] The initial input generator 160 can include tools or a series of instructions to derive a dataset (or initial dataset) to be used by a machine learning system (e.g., the predictor 126) to generate initial sports predictions. In one example, the initial dataset can include simple input data such as the positions of the players and the ball on the court, previous actions (e.g., pass, dribble), and historical information of one or more players of the opposing team (e.g., a goalkeeper’s save percentage or an opposing tennis player’s return percentage). For example, an initial dataset related to tennis can include all player position data and ball position data and corresponding velocity and acceleration data for all players and the ball. The positions of the players and the ball can be actual format (e.g., global positioning coordinates) or grid coordinates relative to the court. The initial dataset can also include shot types and previous shot types (i.e., forehand, backhand, smash), etc.

[0079] Predictive model 162 may include a machine learning model configured to generate one or more sports predictions. Predictive model 162 may include a generative adversarial network (GAN) model with various monotonic constraints and weights to calibrate the model, as described herein. In other examples, predictive model 162 may be a transformer or any other type of machine learning model. Predictive model 162 may be configured to generate more accurate predictions in edge-case scenarios of sports events. Predictive model 162 may include human-in-the-loop mechanisms to tune the model. Predictive model 162 may be configured to update the values ​​of the corresponding weights such that the results conform to the expected behavior of human experts (this can be human or can be automatically set). Predictive model 162 may not minimize the loss function, but rather has a desired output prediction space, toward which predictive model 162 is guided to address all edge cases that tend to occur in sports. Predictive model 162 may receive initial data as input via initial input generator 160 and is configured to determine one or more predictions. Predictive model 162 may also refine the determined predictions by incorporating additional feature datasets provided by feature set generator 164, as described below. Predictive model 162 can be updated (e.g., fine-tuned, modified, specialized, etc.) to make various predictions. For example, predictive model 162 can be configured to determine whether a particular shot will result in a goal (e.g., in football, hockey, basketball, handball, etc.). Predictive model 162 can also be configured to determine the winner of a point (e.g., predicting which player will win the point in a tennis match), the winner of a shot (e.g., the chance of a particular shot winning the point), the shot position, and / or the trajectory of the next shot (e.g., the impact point, velocity, spin, and / or three-dimensional position and time). Predictive model 162 can be constrained by monotonic constraints (e.g., such as...). FIG. 4A (as shown) implements various software libraries (e.g., XGBoost).

[0080] Predictive model 162 can be retrained with additional data. In some scenarios, the additional data could be marginal / abnormal situations in a sporting event. Predictive model 162 can output adjusted weights that it has learned from the training data. Therefore, predictive model 162 can first be trained on the initial training data and learn the initial model weights. Then, predictive model 162 can be trained on additional training data to address marginal cases, and then retrained with the updated training data, and so on, until convergence. Convergence can be defined as the moment when the output is as expected (e.g., accurate).

[0081] In soccer, the input to the expected goal-line save prediction model 162 can include event-based data that includes information about the exact striker and goalkeeper positions, the proximity of the defending players to the shot, and qualifiers about the shot and assist. The output can include a continuous value between 0-1 xG for all runs and direct free-kick shots. Penalty kicks can be assigned a constant value (e.g., 0.7884) xG, and the output can exclude own goals. In an example case, the prediction model 162 can have been trained on data from 800,000 shots in 49 games (including 9 women’s games) between 2018 / 19 and 2021 / 22. The prediction model 162 can have been trained on 70% of the data, with 30% of the data used to test the model.

[0082] Features of the prediction model 162 can include the distance of the goalkeeper and attacking player (or ball) and the angle of attack relative to the goalkeeper and / or goal net. The goalkeeper save probability can be situational and determined separately for different scenarios. These scenarios can include indicators of the goalkeeper making a save given the distance of the ball to the goal or goalkeeper, the position of the goalkeeper, and / or the angle between the ball and the goal or goalkeeper.

[0083] In one example, there can be multiple prediction models 162. Separate prediction models 162 can be created with separate weights for different groups of athletes with different statistics (e.g., men’s and women’s sports; or based on skill level, such as amateur, semi-professional, or professional). Further, separate prediction models 162 can be generated for individual players and / or playing surfaces in a sport.

[0084] The feature set generator 164 can be configured to generate a feature set, as described herein. For example, the feature set generator 164 can include a tool or set of instructions that is capable of generating robust data related to an action (e.g., a shot on goal or a return of serve in tennis) and deriving features, as described herein. The features of the feature set can be converted to values that can be used to modify the predictions generated by the prediction model 162. In some cases, each feature can be weighted to further increase the accuracy of the predictions. The feature set generator 164 can extract data from the data store 118. FIG. 1A

[0085] The feature set generator 164 can generate features that include, but are not limited to, the distance and angle of the player relative to the goal or goalkeeper, the defensive player pressure (e.g., the proximity of the defensive player to the goalkeeper), the clarity (e.g., the line of sight between the goalkeeper and the ball). Other example model features can include the shot context (e.g., free kick, header, overhead kick, first touch, one-on-one) and / or the previous action type (e.g., corner kick, throw-in, direct pass, fast break, bounce ball). ​

[0086] FIG. 2 is an example method 200 for generating a probability of an action of a sporting event by implementing a set of features in accordance with one or more embodiments. The method 200 can be implemented by the environment 100 and the prediction environment 101 described above in FIG. 1A and FIG. 1B .

[0087] At step 202, the method can include obtaining an initial data set related to an action of a sporting event, the initial data set including at least a location of a first player on a field and a location of a target area on a field surface (e.g., field, court, etc.). The initial data set can be obtained from the initial input generator 160. The location of the first player can refer to a player preparing to take a shot (e.g., a player attempting to shoot a goal). The target area can refer to a goal of a soccer field or a service area of a tennis court. For example, the action can be a shot by an attacking player during a soccer game or a serve by a player in a tennis game.

[0088] In some cases, the initial data set can include a historical save probability of a defensive player. For example, the historical save probability can include a previous save percentage of a goalkeeper during a previous game or season. The initial data can also include a tackle or shot block percentage of a defensive player in a soccer game. The historical save probability can also include a return of serve percentage of a tennis player in a past game, set, or match. The historical save probability can also be a forehand or backhand completion rate and win rate, or it can be a player’s efficiency in edge-of-the-court shots, such as a lob or drop shot. This type of feature can capture a player’s tendencies at unique moments.

[0089] At step 204, the method can include generating an initial prediction score probability by a machine learning model (e.g., the prediction model 162) based on the initial data set from step 202. In another example, as part of and / or in addition to step 204, the method can include generating an alternative prediction, such as a predicted shot location or shot trajectory. The prediction can be a numerical output (e.g., see FIG. 3 indicating a 0.74 output, which represents a 74% likelihood of a goal based on the shot). The initial prediction can be based only on the initial data provided. The machine learning model can include a GAN model, where the GAN model includes any of a set of monotonic constraints and a weight for each of the monotonic constraints. An example of the assigned weights can be shown in FIG. 4A and discussed below.

[0090] At step 206, the method can include generating a feature set related to the sporting event, the feature set derived from a position of the second player on the field, and any of a distance and / or an angle between the first player, the second player, and any two of the target region on the field surface (e.g., court). The feature set can be generated by the feature set generator 164 of FIG. 1B In one example, the second player can be a goalkeeper in a soccer match. The second player can be an opposing player in a tennis match.

[0091] The generated feature set can include a virtual line between the position of the first player and the position of the target region. The virtual line can be used to calculate an angle of a shot on goal relative to the goalkeeper position. The feature set can be derived from a position of the defensive player on the field, a distance and / or an angle between the defensive player, the offensive player, and the goal, and any of a presence of any additional players in the vicinity of the defensive player, the offensive player, and the goal. The feature set can also include a ball strike trajectory (e.g., ball strike impact position, velocity, spin, or three-dimensional position and time); three-dimensional body pose information of the player (including skeletal position information and racquet position); a current number of strokes (e.g., number of strokes in a volley in tennis); a speed and acceleration of the player; a court surface (e.g., court surface in tennis); a game state (e.g., score or game number in tennis); a stroke type (e.g., serve, return, forehand, backhand, lob, volley, or underarm stroke in tennis); temperature and wind information; time of day; and / or player attributes (e.g., skill rating for a particular stroke, quality rating for a serve, return, forehand, backhand).

[0092] The feature set can also include a distance between the second player and the virtual line region. The feature set can include a save probability arrangement for the second player that modifies a save probability for each of a set of shot positions at different angles from the virtual line between the position of the first player and the position of the target region (e.g., as depicted in FIG. 4B The example graphical representation of such a feature can be seen in FIG. 4B to FIG. 4C and FIG. 6A to FIG. 6B For example, in FIG. 4CIn the image, the goalkeeper 404 is positioned away from the line of sight between the attacking player 406 and the goal 410. As described in more detail below, various factors, such as the distance between the goalkeeper's line of sight and the striker's, or the position of any interfering players, can affect the predicted save probability for a given shot. For example, in a football match, the goal can be crucial information, so a virtual line can be created to capture the distance of the player / ball from the center of the goal. The virtual line can include the relative distance and angle between the attacking and defending players, thus capturing the relative distance and angle of the goal. Instead of using the center of the goal, the virtual line can be configured relative to the near or far post, allowing the system to know the player's position relative to each respective post. Virtual lines can also be defined and implemented in tennis use cases. The virtual line can display the distance between the player and the baseline. This can be paired with metadata showing the bounce position of the shot relative to the baseline.

[0093] In some cases, the feature set may include a clarity value, which specifies the number of interfering players positioned near the virtual line between the attacking player's position and the goal's position. For example, "near the virtual line" could be considered as being within 1 foot, 3 feet, 5 feet, or 10 feet of the virtual line. FIG. 4C In this scenario, interfering player 408 can be positioned near attacking player 406 and the goal 410. Given that additional players may obstruct the goalkeeper's view and cause shots to deflect, extra players near the goalkeeper, forwards, and the net may affect the predicted goal save probability.

[0094] In some cases, the feature set may include a prior action type specifying a preceding action performed before the action occurs. Example prior action types in football may include free kicks, corner kicks, throw-ins, through balls, fast breaks, and bounces. Alternative examples may include serves, returns, forehands, backhands, overhead smashes, volleys, or between-the-legs shots in tennis.

[0095] In some cases, the feature set may include a shot prototype, which specifies the type of shot, whether the shot is marked, the range of the shot, and / or whether the shot comes from the center or the side of the field. Other example shot prototypes may include, but are not limited to, long shots, shots into an open goal (with foot or head), direct free kicks, close-range shots, headers, etc. When the sport is a tennis match, alternative shot types may include serves, returns, forehands, backhands, overhead smashes, volleys, or between-the-legs shots.

[0096] At step 208, the method can include modifying, by a machine learning model (e.g., the prediction model 162), the initial prediction score probability resulting from step 204 to an updated score probability by using a feature set. The updated save probability can incorporate the feature set information and provide a more accurate prediction. In particular, the prediction can be accurate for less occurring situations (e.g., when the opposing player is further away from the goal area). The prediction result can be a prediction output from 0 to 1, where 1 indicates a 100% chance of the prediction (e.g., a goal or winning the point) occurring, and 0 indicates a 0% chance of the prediction occurring. In alternative predictions, the system can be configured to output a continuous prediction (see FIG. 14 ) as well as a grid of the corresponding percentage chances of the ball landing in a particular area.

[0097] As described below FIG. 3 to FIG. 12 Aspects of the method 200 and the environment 100 as described above can be described as applied to a soccer game. FIG. 3 to FIG. 12 Particular aspects of the description can apply to sports events other than soccer.

[0098] FIG. 3 An example of a game scenario of a goalkeeper 302 and an attacking player 304 in a soccer game is depicted in accordance with example embodiments. The method 200 can have been applied to the game scenario 300 to determine an expected goal probability (xG) when the attacking player 304 attempts to shoot. The prediction model (e.g., the prediction model 162) can have first received initial data related to the position of the attacking player 304, the position of the object 301, and / or the goal 306. An initial prediction, not shown, can have occurred. The system can have also received more detailed information and determined a feature set to provide to the prediction model. With the updated information, a prediction 308 can have been generated indicating a 74% chance of a goal.

[0099] FIG. 4A A graphical depiction 400A of a monotonic constraint of the prediction model 162 is depicted in accordance with example embodiments. FIG. 4B and FIG. 4C A graphical representation of a predicted save probability output by the prediction model 162 for an example scenario is depicted.

[0100] Checking FIG. 4AFigures 402A-E depict graphical models of the monotonic constraints implemented by prediction model 162. These figures depict partial dependencies and various metrics related to shots to the goalkeeper. For example, Figure 402A may depict how much dependence (e.g., how much weight) should be given to the distance (e.g., virtual line) and angle between the attacking player and the goal in the model prediction. Weights that can be adjusted during the calibration of the prediction model described herein can be assigned to the virtual line and angle. Furthermore, figures 402B-E may show various metrics (e.g., angle of shot to goal, goalkeeper save probability, pressure, and / or clarity) mapped to the partial dependencies applied by the prediction model.

[0101] The angle of a shot to the goal can be the angle from the attacking player to the center of the goal. A larger angle between the ball and the goal reduces the probability of a goal because it makes it more difficult for the attacking player to hit the ball into the net. Goalkeeper save probability can be the probability that the goalkeeper will save any shot or a shot from a specific position on the field. Pressure can be an indicator specifying the number of players within a certain distance of the goalkeeper. Additionally, clarity specifies the line of sight between the goalkeeper and the ball, as interfering players between the goalkeeper and the ball can obstruct the goalkeeper's view.

[0102] FIG. 4B A first representation 400B depicting the goalkeeper's save probability with a predicted arrangement of shot positions, according to an example embodiment, is shown. FIG. 4B As shown, the positions of the goalkeeper 404, attacking player 406, and other players 408 can be derived relative to the goal 410. As mentioned above, the goalkeeper's position relative to the goal and the ball can modify the probability of the goalkeeper making a save based on the shooting position. The output of the goalkeeper's save probability arranged by the goalkeeper's position and the predicted shooting position can be shown in output 412. As shown in output 412, the probability of the goalkeeper making a save decreases when the shot is taken further away from the goalkeeper's position.

[0103] FIG. 4C A second representation 400C depicting the goalkeeper's save probability with shot position arrangement according to an example embodiment is shown. For example... FIG. 4C As shown, this representation shows the goalkeeper's save probability based on the position of attacking player 406. When goalkeeper 404 is outside the line of sight between attacking player 406 and goal 410, the save probability may be lower due to the position of goalkeeper 404. Line 413 visualizes the distance from attacking player 406 to goal 410. FIG. 4CAs shown, based on the position of the goalkeeper 404, certain areas on the field (e.g., area 414) can have a higher probability of a goal than other areas on the field (e.g., area 416) that are located a similar distance from the goal 410. This illustrates how a model (e.g., the prediction model 162) can identify that, despite the goalkeeper being further from the goal 410, the goalkeeper can still be in a favorable position to make a save (e.g., in area 416) based on the position of the attacking player. Further, it depicts that a player shooting from a greater distance (e.g., outside the penalty area) can have a significant chance of scoring, despite the goalkeeper being further from the goal and not being in line of sight between the player and the goalkeeper.

[0104] According to embodiments disclosed herein, the outputs (e.g., predictions, probabilities, likelihoods, etc.) discussed herein can be used to output recommended actions. For example, a machine learning model can be trained to output a recommended action based on input data that includes one or more of the outputs discussed herein. Such a machine learning model can be trained using training data that includes one or more of historical or simulated match data, historical or simulated actions, historical or simulated player data, etc. A trained version of the machine learning model can receive one or more of the outputs discussed herein and can output a recommended player action, formation, team action, etc. As one example, a goalkeeper save probability for a shot from a first direction can be provided to a machine learning model. Based on the save probability, the machine learning model can generate a recommended goalkeeper action to set up in a given area relative to a post to maximize the save probability. One or more recommended actions can be used to generate a visual depiction of the recommended action. The type of visual depiction can be dynamically determined based on the recommended action and can include one or more objects, players, etc. determined based on the recommended action (e.g., a plot depicting where a goalkeeper should stand based on the save probability and the initial penalty kick location).

[0105] Example performance of the prediction model 162 can be represented by an area under the curve of a receiver operating characteristic (AUC-ROC) and a log loss score. FIG. 5A A plot 500A of a receiver operating characteristic curve for one or more versions of the prediction model 162 can be depicted. As FIG. 5A shown, a first representation 502A can include a true positive rate and a false positive rate of a receiver operating characteristic curve. Further, a second representation 502B can include an observed shot conversion rate and an expected shot conversion rate based on calibration points. A third representation 502C can include a precision rate and a recall rate of a precision-recall curve.

[0106] For example, performance of a model can include: xG1: 0.80 AUC-ROC, 0.250 log loss xG2.1 : 0.84 AUC-ROC, 0.236 log loss xG2.2: 0.81 AUC-ROC, 0.248 log loss

[0107] The performance of the model can remain robust across various shot locations. FIG. 5B is a representation 500B of shot log loss based on various shot locations. As shown, the log loss at various locations on the field can indicate the likelihood of scoring a goal at the various locations on the field. For example, the size of the circular indicators can correspond to the likelihood of scoring a goal from the location corresponding to the circular indicator. FIG. 5B

[0108] Further, the goal scoring indicators as described herein can be tracked based on various shot archetypes. Example shot archetypes can include a header, a tap-in, a free kick, a long shot, a wide angle shot, etc. Table 1 provides an example table of conversion rates and biases for each shot archetype. Table 1

[0109] Further, the indicators of the prediction model 162 can be arranged using a relative striker and / or goalkeeper position. FIG. 6A to FIG. 6B Example representations 600A-B of xG indicators arranged by a relative striker and / or goalkeeper position are shown in accordance with example embodiments. FIG. 6A to FIG. 6B The representations shown in FIGS. 6A-B can be used to evaluate the optimality of both the player at the time of the shot and the goalkeeper positioned to attempt to save the shot.

[0110] For example, FIG. 6A A representation 600A of xG indicators (e.g., goal scoring probability) arranged by shot location is shown. As shown, FIG. 6A The xG indicators can be based on the position of the goalkeeper 604, the position of the striker 606, the goal 608, any other player 610, or the distance 612 between the striker 606 and the goal 608, as another example, FIG. 6B A representation of xG indicators arranged by goalkeeper position is shown. As shown, FIG. 6B The position of the goalkeeper 604 relative to the angle between the striker 606 and the goal 608 can predict a goal save probability based on the prediction of a shot near the line of sight between the striker 606 and the goal 608.

[0111] ​The prediction model 162 can be calibrated using past football season data. In some cases, the model (e.g., the prediction model 162) can be calibrated such that the total number of goals scored closely approximates the number of goals scored by any large sample of shots, such as the total xG across a match season or a team season. This can be referred to as a prediction bias or a higher than expected number of goals. Calibration can include synthetically adding more examples around edge case scenarios (e.g., training data for scenarios such as when a goalkeeper is poorly positioned). Calibration can include utilizing feedback from external sources. For example, an external source can not just look at one metric to calculate the total loss, but review many metrics (i.e., how the output prediction space looks on the pitch or field). Calibration can include data from all parts of the pitch / field and if any part spikes and overfits (e.g., different from the generated or expected outcome), additional calibration can be performed with more training data for those scenarios.

[0112] FIG. 7A to FIG. 7B Various representations 700A-B of calibrating a model (e.g., the prediction model 162) across multiple model versions (e.g., XG1, XG2.1, XG2.2 (men), and XG2.2 (women) representing updated models) are shown in accordance with example embodiments. FIG. 7A A first representation of a model for a match season (e.g., 700A) and a team season (e.g., 702A) is shown. As shown, FIG. 7A As the model is updated to newer versions (e.g., when the prediction model is trained on additional training data), the model can be recalibrated.

[0113] Prediction bias can be similar for men’s and women’s matches. FIG. 7B A representation 700B of a model that can be calibrated for women’s matches is shown. While noise within a single season can mean that predictions are sometimes high or low, prediction bias can be centered around zero (e.g., the median in a box plot in FIG. 7A to FIG. 7B For FIG. 7A and FIG. 7B the dashed line can represent the mean, the center of each box can be the median prediction, and the outline of the respective box represents the 25th to 75th percentile.

[0114] FIG. 8 is a graphical representation 800 of the predicted goal save probability versus the actual conversion rate for various teams in a match season. As shown, FIG. 8 The graphical representation 800 can show the observed shot conversion rate and the number of goals minus the mean of the predicted goal save probability. FIG. 8A trend line 802 can be provided for each team's conversion rate over each match season. An extreme prediction bias value 804 is examined, which includes a season in which a soccer team took 4,632 shots and observed 515 goals, with a predicted 422 xG, which represents a prediction that is 93 xG lower or 0.02 xG lower per shot on average.

[0115] Further, in this example, the xG bias can be correlated with the observed shot conversion rate, which explains 46% of the total variance in the bias. In this example, when a player's opportunity conversion rate is higher than expected, there can be a prediction that is lower, and vice versa. As an illustrative example, a first team can have one of the highest observed conversion rates in any match season, as shown by the value 804 in FIG. 8 The residual bias can be equal to only 43 xG or 0.01 xG per shot on average. While the residual bias is still high, this can be on the order of what can only be expected to see in one out of 10 match seasons.

[0116] The xG metrics described herein can be compared to similar metrics of another model that implements traditional methods to generate soccer tournament predictions. The present xG metrics can have similar aggregations as the different model. The performance of the xG generated using the model described herein, across all shots in a match and within a particular shot archetype, can be more accurate and less biased than other metrics. Table 2 shows example conversion rates and biases within a particular shot archetype: Table 2

[0117] For example, an xG metric (e.g., also referred to as "xG2.2") from the model described herein (e.g., prediction model 162) can have a log loss = 0.267 and a bias = 4.9 goals above expectation. In this example, other metrics can have a log loss = 0.273 and a bias = 13.7 goals above expectation. The results can show improved results for the model and methods described herein.

[0118] FIG. 9 is a graphical illustration of the xG metrics as described herein according to example embodiments versus one or more of the other metrics described above. As shown in FIG. 9 Figure 900 can show the performance improvement of the xG derived herein relative to the xG of another model for a particular tournament.

[0119] In some embodiments, the explainable prediction explanation can include a Shapley Additive explanations (SHAP) plot to provide a human-interpretable explanation of any single shot prediction. FIG. 10A to FIG. 10B shows an example view of a shot (e.g., FIG. 10A ) and a plot specific to that shot (e.g.,FIG. 10B As shown in FIG. 1000, the positions of the goalkeeper 1002, the forward 1004, etc. can be tracked and determined relative to the field. FIG. 10A As shown in FIG. 1000, the positions of the goalkeeper 1002, the forward 1004, etc. can be tracked and determined relative to the field.

[0120] Further, FIG. 10B A graph (e.g., 1000B) in FIG. 1000 can show a base value for the xG metric 1006. The base value can be modified based on individual features that increase (e.g., 1008A) or decrease (1008B) the modeled probability of scoring, which in this case can be 0.36 xG (e.g., final value 1010).

[0121] For example, the relatively low goalkeeper save probability (about 43%) from the shot from that location, the short distance (about 9 meters), and the unobstructed view of the goal (clarity = 1, pressure = 1) all contribute to increasing the xG. The fact that the shot is both an overhead kick and the first touch by the forward can have a negative impact on the xG.

[0122] FIG. 11A to FIG. 11B An example view of the shot (e.g., image 1100A) and a graph (e.g., 1100B) specific to the shot are shown in FIG. 1100 in accordance with example embodiments. FIG. 11A As shown in FIG. 1100, the positions of the goalkeeper 1102, the forward 1104, etc. can be tracked and determined relative to the field, where the goalkeeper 1102 is severely out of position from the forward 1104 and the goal 1111. FIG. 11B As shown in FIG. 1100, the positions of the goalkeeper 1102, the forward 1104, etc. can be tracked and determined relative to the field, where the goalkeeper 1102 is severely out of position from the forward 1104 and the goal 1111. FIG. 11A

[0123] Further, FIG. 11B A graph (e.g., 1100B) in FIG. 1100 can show a base value for the xG metric 1106. The base value can be modified based on individual features that increase (e.g., 1108A) or decrease (1108B) the modeled probability of scoring, which can be 0.54 xG (e.g., final value 1110). In this example, there can be very low goalkeeper save probability (<0.01%) and short distance (about 8 meters), which have the largest positive impact on the xG. The fact that the shot is the first touch and relatively blocked by a defending player (clarity = 2) can provide a small negative impact.

[0124] Other example outputs can include xG shot graphs, number of goals above xG, shot value added (i.e., xGoT - xG), goalkeeper positioning optimality, etc. FIG. 12 An example xG shot graph 1200A is shown. As shown in FIG. 1200, the shot plot in the shot graph 1200 can show the number of goals out of the total number of shots, as well as the xG for each shot. FIG. 12 An example xG shot graph 1200A is shown. As shown in FIG. 1200, the shot plot in the shot graph 1200 can show the number of goals out of the total number of shots, as well as the xG for each shot.

[0125] Table 3 below shows an example table mapping the number of goals at or above xG for a soccer league: ​ Table 3

[0126] As described below FIG. 13 and FIG. 14 Aspects of the method 200 and environment 100 as described above can be shown when applied to tennis. When applying the prediction model 162 to tennis, input features that can be utilized can include: a shot trajectory (e.g., shot impact location, velocity, spin, or three-dimensional position and time); a player’s location (e.g., three-dimensional location of all players); a current shot count (e.g., shot count in a volley in tennis); a player’s velocity and acceleration; a court surface (e.g., court surface in tennis); a game state (e.g., score or game number in tennis); a shot type (e.g., serve, return, forehand, backhand, lob, volley, or under-the-leg shot in tennis); temperature and wind information; time of day; and / or player attributes (e.g., skill rating for a particular shot, such as quality rating for serve, return, forehand, backhand). The prediction model 162 can predict outputs such as a next point winner, a next shot winner, a shot location, and / or a next shot trajectory.

[0127] FIG. 13 is an example heat map output 1300 from a prediction model (e.g., prediction model 162) according to one or more embodiments. The prediction output can be a predicted location of a shot. The output 1300 can include a location of a player 1302 hitting a tennis ball, a location of where a previous shot was made 1304, and a corresponding path of the tennis ball prior 1306. The prediction model can also include location information on a target zone 1308, and project a heat map prediction 1310 within the target zone 1308 that depicts the likelihood of a shot landing. The prediction model can incorporate any of the aforementioned feature sets, including a location of an opposing team player 1312, to help predict a shot location. Additional feature set information can include that the previous shot was not a serve, the previous shot had topspin, and the previous shot did not have backspin. XLN can refer to a shot landing within the court boundary. Conversely, x_out would be 7.6% in this example. The model can predict additional outputs such as a location of each zone or a location of winning the point.

[0128] FIG. 14An exemplary output 1400 of a prediction model 162, according to one or more embodiments, when applied to an exemplary tennis match can be depicted. For example, during a tennis match, a first player 1402 may hit a tennis ball, aiming at a target area 1404 on the court. In addition to initial data, the prediction model 162 may incorporate various feature data to generate one or more predictions. The data may include the position of a second player 1406 from the opposing team. For example, the model may output a constant prediction 1408 of the probability that each player will win a particular point. This prediction 1408 may be updated with each hit in a rally. The model may also predict a specific probability that the shot will land in one or more grid segments 1410 of the target area 1404.

[0129] Overview of Neural Network Training and Computation Systems

[0130] FIG. 15 A flowchart illustrating the training of a machine learning model based on one aspect is provided. FIG. 15 As shown in flowchart 1510, training data 1512 may include one or more of stage inputs 1514 and known results 1518 associated with the machine learning model to be trained. Stage inputs 1514 may originate from any applicable source, including the components or sets shown in the figures provided herein. Known results 1518 may be included for machine learning models generated based on supervised or semi-supervised training. Supervised machine learning models may be trained without using known results 1518. Known results 1518 may include known or expected outputs for future inputs that are similar to or in the same category as stage inputs 1514 that do not have corresponding known outputs.

[0131] Training data 1512 and training algorithm 1520 can be provided to training component 1530, which can apply training data 1512 to training algorithm 1520 to generate a trained machine learning model 1550. According to one embodiment, a comparison result 1516 can be provided to training component 1530, which compares the previous output of the corresponding machine learning model to apply the previous result to retrain the machine learning model. Training component 1530 can use comparison result 1516 to update the corresponding machine learning model. Training algorithm 1520 can utilize machine learning networks and / or models, including but not limited to, deep learning networks such as deep neural networks (DNN), convolutional neural networks (CNN), fully convolutional networks (FCN), and recurrent neural networks (RCN), probabilistic models (such as Bayesian networks and graphical models), and / or discriminative models (such as decision forests and maximum margin methods). The output of flowchart 1510 can be the trained machine learning model 1550.

[0132] The machine learning models disclosed herein can be trained by adjusting one or more weights, layers, and / or biases during a training phase. During the training phase, historical or simulated data can be provided as input to the model. The model can adjust one or more of its weights, layers, and / or biases based on such historical or simulated information. The adjusted weights, layers, and / or biases can be configured based on the training to a production version of the machine learning model (e.g., a trained model). Once trained, the machine learning model can output an output of the machine learning model in accordance with the subject matter disclosed herein. In accordance with an embodiment, one or more machine learning models disclosed herein can be continuously or periodically updated based on feedback associated with the use or implementation of the machine learning model output.

[0133] It should be appreciated that aspects in the application are exemplary only and that other aspects can include various combinations of features from other aspects, as well as additional or fewer features.

[0134] Generally, any processes or operations discussed in the application that are understood to be computer-implemented, such as the processes shown in the flowcharts disclosed herein, can be performed by one or more processors of a computer system, such as any of the systems or devices in the example environments disclosed herein, as described above. Processes or process steps performed by one or more processors can also be referred to as operations. The one or more processors can be configured to perform the processes by accessing instructions (e.g., software or computer-readable code) that, when executed by the one or more processors, cause the one or more processors to perform the processes. The instructions can be stored in a memory of the computer system. The processors can be central processing units (CPUs), graphics processing units (GPUs), or any suitable type of processing units.

[0135] A computer system, such as a system or device that implements processes or operations in the examples described above, can include one or more computing devices, such as one or more of the systems or devices disclosed herein. The one or more processors of the computer system can be included in a single computing device or distributed among multiple computing devices. The memory of the computer system can include a respective memory of each of the multiple computing devices.

[0136] FIG. 16is a simplified functional block diagram of a computer 1600 that can be configured as an apparatus for performing the methods disclosed herein according to the exemplary aspects of the application. For example, the computer 1600 can be configured as a system according to the exemplary aspects of the application. In various aspects, any of the systems herein can be a computer 1600 that includes, for example, a data communication interface 1620 for packetized data communication. The computer 1600 can also include a central processing unit ("CPU") 1602 in the form of one or more processors for executing program instructions for the methods disclosed herein. While the computer 1600 can receive programming and data via network communication, the computer 1600 can include an internal communication bus 1608 and a storage unit 1606 (such as a ROM, HDD, SDD, etc.) that can store data on a computer readable medium 1622.

[0137] While the instructions 1624 can be temporarily or permanently stored within other modules of the computer 1600 (e.g., the processing unit 1602 and / or the computer readable medium 1622), the computer 1600 can also have a memory 1604 (such as a RAM) that stores the instructions 1624 for performing the techniques presented herein, such as the methods described above in relation to FIG. 2 The computer 1600 can also include input and output ports 1612 and / or a display 1610 to connect with input and output devices, such as a keyboard, mouse, touchscreen, monitor, display, etc. Various system functions can be implemented in a distributed manner across multiple similar platforms to distribute processing loads. Alternatively, the system can be implemented by appropriate programming of one computer hardware platform.

[0138] The programmatic aspect of this technology can be considered a "product" or "manufactured item," typically carried on a machine-readable medium or in the form of embodied executable code and / or associated data. "Storage" type media includes computers, processors, etc., or their associated modules, any or all of tangible storage such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. Sometimes, all or part of the software can communicate via the Internet or various other telecommunications networks. For example, such communication enables the loading of software from one computer or processor to another, such as from a management server or host of a mobile communication network to a server's computer platform and / or from a server to a mobile device. Therefore, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as physical interfaces between local devices, used via wired and fiber optic networks, and via various air links. Physical elements carrying such waves, such as wired or wireless links, fiber optic links, etc., can also be considered as media carrying software. As used herein, unless limited to non-transitory, tangible "storage" media, the term "readable medium" for a computer or machine refers to any medium that participates in providing instructions to a processor for execution.

[0139] While the disclosed methods, apparatus, and systems are described with reference to illustrative data transmission, it should be understood that the disclosed aspects are applicable to any environment, such as desktop or laptop computers, car entertainment systems, home entertainment systems, etc. Furthermore, the disclosed aspects are applicable to any type of Internet protocol.

[0140] It should be understood that in the foregoing description of exemplary aspects of the invention, various features of the invention are sometimes combined in a single aspect, drawing, or description therein to simplify the invention and aid in understanding one or more of the various inventive aspects. However, this approach of the invention should not be interpreted as reflecting an intention that the claimed embodiments require more features than expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects exist in fewer features than all features in a single foregoingly disclosed aspect. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, wherein each claim exists independently as a separate aspect of the invention.

[0141] Furthermore, while some aspects described herein include features included in other aspects but not others, combinations of features from different aspects should be within the scope of this invention and form different aspects, as those skilled in the art should understand. For example, any of the claimed aspects may be used in any combination in the following claims.

[0142] Thus, although certain aspects have been described, it is to be understood that other, additional or further aspects can be made within the scope of the present application, and that aspects described can be combined in other manners to constitute further aspects within the scope of the present application. For example, functionality described from the perspective of one functional block can be added to or removed from other functional blocks. Operations can be interchanged among functional blocks. Operations can be added or removed from the described methods.

[0143] The disclosed subject matter is to be considered illustrative, and not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other implementations falling within the true spirit and scope of the present application. Therefore, the scope of the present application is to be determined not with reference to the foregoing description, but rather should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. The disclosures of each patent, patent application, and publication cited herein are hereby incorporated herein by reference in their entirety. Although various embodiments of the present application have been described, further embodiments can be made without departing from the scope of the present application. Accordingly, no limitation is placed on the scope of the present application by the scope of the claims and / or the embodiments set forth.

Claims

1. A method of generating a probability of a first action of a sporting event by implementing a set of features, the method comprising: obtaining an initial data set relating to the first action of a sporting event, the initial data set comprising at least a location of a first player on a field surface and a location of a target area on the field surface; generating, by a machine learning model, an initial predicted score probability based on the initial data set; generating a set of features relating to the sporting event, the set of features derived from any of: a location of a second player on the field surface, and at least one of a distance or an angle between any two of the first player, the second player, and the target area on the field surface; and modifying, by the machine learning model, the initial predicted score probability to an updated score probability using the set of features.

2. The method of claim 1, wherein, the machine learning model comprises a generative adversarial network (GAN) model, and wherein the GAN model comprises at least one of a set of monotonic constraints or a weight of each of the monotonic constraints.

3. The method of claim 1, wherein, the set of features comprises a previous action type specifying a previous action performed prior to the first action occurring.

4. The method of claim 1, wherein, any of the initial data set and the set of features comprises a historical save probability of the second player.

5. The method of claim 1, wherein, the first player and the second player belong to opposing teams in the sporting event.

6. The method of claim 1, wherein, the set of features comprises: a virtual line between the location of the first player and the location of the target area; and a distance between the location of the second player and a region of the virtual line.

7. The method of claim 6, wherein, the set of features comprises a save probability arrangement of a second player that modifies a save probability for each of a set of shot locations at different angles from the virtual line between the location of the first player and the location of the target area.

8. A method of generating a predicted save probability of a first action of a sporting event using a set of features, the method comprising: obtaining an initial data set relating to the first action of a sporting event, the initial data set comprising at least a location of an attacking player on a field and a location of a goal on the field; generating, by a machine learning model, an initial predicted save probability based on the initial data set; generating a set of features relating to the sporting event, the set of features derived from any of: a location of a defending player on the field, at least one of a distance or an angle between any two of the defending player, the attacking player, and the goal, or a presence of one or more additional players in the vicinity of the defending player, the attacking player, and the goal; and modifying, by the machine learning model, the initial predicted save probability to an updated predicted save probability using the set of features.

9. The method of claim 8, wherein, the set of features comprises: a virtual line between the location of the attacking player and the location of the goal; and a distance between the location of the defending player and the virtual line between the location of the attacking player and the location of the goal.

10. The method of claim 9, wherein, The feature set includes a goalkeeper save probability ranking that modifies a goalkeeper save probability for each of a set of shot locations at different angles from the virtual line between the location of the attacking player and the location of the goal.

11. The method of claim 9, wherein, The feature set includes a clarity value that specifies a number of interfering players positioned near the virtual line between the location of the attacking player and the location of the goal.

12. The method of claim 8, wherein, The feature set includes a previous action type that specifies a previous action performed prior to the first action occurring.

13. The method of claim 8, wherein, The first action includes a shot, the sporting event includes a soccer match, and the defending player is a goalkeeper.

14. The method of claim 13, wherein, The feature set includes a shot archetype that specifies one or more of a shot type, whether the shot is under pressure, a distance range of the shot, and / or whether the shot is from a central portion of the field or a side portion of the field.

15. The method of claim 8, wherein, The machine learning model includes a generative adversarial network (GAN) model, and wherein the GAN model includes at least one of a set of monotonic constraints or a weight of each of the monotonic constraints.

16. The method of claim 8, wherein, Any of the initial data set or the feature set includes a historical save probability of the defending player.

17. A method of generating a predicted score win probability for a shot in a tennis match using a feature set, the method comprising: obtaining an initial data set related to an action of a sporting event, the initial data set including at least a location of a first player on a court surface and a location of a target area on the court; generating, by a machine learning model, an initial predicted score win probability based on the initial data set; generating a feature set related to the sporting event, the feature set derived from any of a location of a second player on the court, a distance and / or an angle between any two of the second player, the first player, and the target area, or a presence of any additional players near the second player, the first player, and the target area; and modifying, by the machine learning model, the initial predicted score win probability to an updated predicted score win probability using the feature set.

18. The method of claim 17, wherein, The machine learning model includes a generative adversarial network (GAN) model, and wherein the GAN model includes at least one of a set of monotonic constraints or a weight of each of the monotonic constraints.

19. The method of claim 17, wherein, The feature set includes: a virtual line between the location of the first player and the location of the target area; and a distance between the location of the second player and the virtual line between the location of the first player and the target area.

20. The method of claim 17, wherein, The feature set includes shot trajectory data, 3D body pose information of the first player and the second player, a current shot count, velocity and acceleration information of the first player and the second player, a court surface type, and / or a shot type.