System and method for combining top-down and bottom-up team and player predictions for athletic sports

By combining top-down team prediction and bottom-up player prediction, and using transformer neural networks, we address the problem of insufficient prediction accuracy in existing technologies and achieve real-time, accurate and consistent player and team action prediction.

CN120677484APending Publication Date: 2025-09-19STAT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380086224.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-30
Filing Date
2023-12-29
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing sports event prediction systems are unable to accurately predict the specific actions of players and teams, especially when considering multiple factors and complex interactions, resulting in insufficient prediction accuracy.

Method used

A transformer-based neural network is used to combine top-down team predictions and bottom-up player predictions. It receives multiple feature vectors and generates predictions using embedding layers, transformer encoder layers, and fully connected layers, updating them in real time and maintaining the consistency of the predictions.

Benefits of technology

It improves the accuracy and consistency of player and team action predictions, can update and reflect game dynamics in real time, adapt to complex game environments and player interactions, and enhances the market alignment of predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120677484A_ABST
    Figure CN120677484A_ABST
Patent Text Reader

Abstract

The present invention relates to a method of generating predictions for teams and players of each team associated with a sporting event, the method comprising: receiving one or more top-down predictions for the sporting event; providing the top-down predictions as one or more top-down feature vectors to a computing system; receiving, by the computing system, a second set of feature vectors including data for one or more players associated with one or more respective teams and data for one or more teams associated with a sporting event; inputting the one or more feature vectors from top to bottom and the second group of feature vectors into a neural network based on a converter; and generating one or more predictions for the sporting event using the transformer-based neural network based on the one or more top-down feature vectors and the second set of feature vectors.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 478,056, filed December 30, 2022, the contents of which are incorporated herein by reference for all purposes. Technical Field

[0002] The present invention generally relates to machine learning techniques for generating predictions for players and teams for sporting events. Background Art

[0003] With the growing popularity of sports, there's a growing desire for accurate, granular predictions of what will happen during a sporting event. For example, predicting the number of passes or shots a specific football player (e.g., Lionel Messi) will take in a given match is likely to be of particular interest to members of the media, broadcasters (whether on the primary feed or in a second-screen experience), sports betting, and fantasy / gamification applications. Existing solutions are unable to accurately make such predictions, thus creating a need for new solutions.

[0004] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art or suggestions of prior art by inclusion in this section. Summary of the Invention

[0005] In some aspects, the technology described herein relates to a method for generating predictions for teams and players on each team associated with a sporting event, the method comprising: receiving one or more top-down predictions for the sporting event; providing the top-down predictions as one or more top-down feature vectors to a computing system; receiving, by the computing system, a second set of feature vectors, the second set of feature vectors comprising data for one or more players associated with one or more respective teams and data for one or more teams associated with the sporting event; inputting the one or more top-down feature vectors and the second set of feature vectors into a transformer-based neural network; and generating, using the transformer-based neural network, one or more predictions for the sporting event based on the one or more top-down feature vectors and the second set of feature vectors.

[0006] In some aspects, the technology described herein relates to a method in which one or more top-down forecasts are based on neural network forecasts or market information.

[0007] In some aspects, the technology described herein relates to a method wherein a team-level prediction is a team's prediction of one or more of a plurality of goals, shots, shots on target, assists, passes, fouls, yellow cards, or red cards for a match.

[0008] In some aspects, the technology described herein relates to a method that also includes causing one or more predictions for the sporting event to be displayed on a display device.

[0009] In some aspects, the technology described herein relates to a method wherein the transformer-based neural network further comprises: a set of embedding layers; a transformer encoder layer; and a fully connected layer.

[0010] In some aspects, the technology described herein relates to a method in which data for one or more players includes movements of one or more actors on a playing field received from a tracking device. In some aspects, the technology described herein relates to a method that also includes: receiving updated data for one or more players or teams from a tracking device; providing the updated data to one or more of the first neural network or the transformer-based neural network; and generating one or more updated predictions for the sporting event based on the updated data. In some aspects, the technology described herein relates to a method that also includes: using a trigger processing step to access a data platform at set time intervals to determine when a sporting event occurs; and using a feature creator processing step to create a second set of feature vectors by querying data from the data platform.

[0011] In some aspects, the technology described herein relates to a system for generating predictions for teams and players on each team associated with a sporting event, the method comprising: receiving one or more top-down predictions for the sporting event; providing the top-down predictions as one or more top-down feature vectors to a computing system; receiving, by the computing system, a second set of feature vectors, the second set of feature vectors comprising data for one or more players associated with one or more respective teams and data for one or more teams associated with the sporting event; inputting the one or more top-down feature vectors and the second set of feature vectors into a transformer-based neural network; and generating, using the transformer-based neural network, one or more predictions for the sporting event based on the one or more top-down feature vectors and the second set of feature vectors.

[0012] In some aspects, the technology described herein relates to a system in which one or more top-down forecasts are based on neural network forecasts or market information.

[0013] In some aspects, the technology described herein relates to a system wherein a team-level prediction is a team's prediction of one or more of a plurality of goals, shots, shots on target, assists, passes, fouls, yellow cards, or red cards for a match.

[0014] In some aspects, the technology described herein relates to a system that also includes causing one or more predictions for the sporting event to be displayed on a display device.

[0015] In some aspects, the technology described herein relates to a system wherein the transformer-based neural network further comprises: a set of embedding layers; a transformer encoder layer; and a fully connected layer.

[0016] In some aspects, the technology described herein relates to a system in which data for one or more players includes movements of one or more actors on a playing field received from a tracking device. In some aspects, the technology described herein relates to a system that also includes: receiving updated data for one or more players or teams from a tracking device; providing the updated data to one or more of the first neural network or the transformer-based neural network; and generating one or more updated predictions for the sporting event based on the updated data.

[0017] In some aspects, the technology described herein relates to a system that further includes: using a trigger processing step to access a data platform at set time intervals to determine when a sporting event occurs; and using a feature creator processing step to create a second set of feature vectors by querying data from the data platform.

[0018] In some aspects, the technology described herein relates to a non-transitory computer-readable medium for generating predictions for teams and players on each team associated with a sporting event, the method comprising: receiving one or more top-down predictions for the sporting event; providing the top-down predictions as one or more top-down feature vectors to a computing system; receiving, by the computing system, a second set of feature vectors, the second set of feature vectors comprising data for one or more players associated with one or more respective teams and data for one or more teams associated with the sporting event; inputting the one or more top-down feature vectors and the second set of feature vectors into a transformer-based neural network; and generating, using the transformer-based neural network, one or more predictions for the sporting event based on the one or more top-down feature vectors and the second set of feature vectors.

[0019] In some aspects, the technology described herein relates to a non-transitory computer-readable medium, wherein one or more top-down forecasts are based on neural network forecasts or market information.

[0020] In some aspects, the technology described herein relates to a non-transitory computer-readable medium, wherein a team-level prediction is a team's prediction of one or more of a plurality of goals, shots, shots on target, assists, passes, fouls, yellow cards, or red cards for a game.

[0021] In some aspects, the technology described herein relates to a non-transitory computer-readable medium, further comprising causing one or more predictions for a sporting event to be displayed on a display device.

[0022] Additional objects and advantages of the disclosed aspects will be set forth in part in the following description and will become apparent from the description, or may be learned by practicing the disclosed aspects. The objects and advantages of the disclosed aspects will be realized and obtained by means of the elements and combinations particularly pointed out in the appended claims.

[0023] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed aspects, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various exemplary aspects and, together with the description, serve to explain the principles of the disclosed aspects.

[0025] Figure 1 is a block diagram of an exemplary tracking and analysis environment according to an example embodiment.

[0026] Figure 2A is an example correlation between input features and predicted distributions according to an example embodiment.

[0027] Figure 2B is an illustration of a set of neural networks that derive a predictive distribution for a set of input features in accordance with an example embodiment.

[0028] Figure 3A is a transformer neural network for player and team predictions according to an example embodiment.

[0029] Figure 3B is a method for Figure 3A An exemplary model of the input tensor of a Transformer neural network.

[0030] Figure 3C According to an example embodiment Figure 3A Example model of the linear embedding layer, axial transfer encoder layer, and fully connected layer of the Transformer neural network.

[0031] Figure 3D According to an example embodiment Figure 3A An example model of the axial attention layer of a Transformer neural network.

[0032] Figure 4 is a flow chart of an example top-down prediction model according to an example embodiment.

[0033] Figure 5 is a flow of an example top-down prediction model during an example sporting game, according to an example embodiment.

[0034] Figure 6 is an exemplary graphical representation of predicted statistical output from a transformer neural network during a sporting event, according to an example embodiment.

[0035] Figure 7 is an example flow for trigger processing steps according to an example embodiment.

[0036] Figure 8 is an example flow for feature creation process steps according to an example embodiment.

[0037] Figure 9 is an example flow of attribute predictor processing steps according to an example embodiment.

[0038] Figure 10 is an example flow for team attribute predictor processing steps according to an example embodiment.

[0039] Figure 11 is an example flow for player attribute predictor processing steps according to an example embodiment.

[0040] Figure 12 is an example flow for combining attribute predictor processing steps according to an example embodiment.

[0041] Figure 13 is an example process for generating predictions for teams and players associated with a sporting event, according to an example embodiment.

[0042] Figure 14 Depicted is a flow diagram for training a machine learning model according to one aspect.

[0043] Figure 15 An example of a computing device according to one aspect is depicted.

[0044] It is worth noting that, for simplicity and clarity of illustration, certain aspects of the drawings depict the general configurations of various embodiments. Descriptions and details of well-known features and techniques may be omitted to avoid unnecessarily obscuring other features. Elements in the figures are not necessarily drawn to scale; the dimensions of some features may be exaggerated relative to other features to improve understanding of the example embodiments. DETAILED DESCRIPTION

[0045] The techniques described herein use machine learning to make predictions related to sports. For example, certain aspects include a combined top-down and bottom-up prediction system for generating both team prediction outputs and player prediction outputs for a given sporting event. Thus, the system can include a first, or top-down, machine learning model and a second, or bottom-up, machine learning model. In one example, the top-down machine learning model applies a first technique to make team-level predictions based on team data, and the bottom-up machine learning model applies a second technique to make player-level predictions and / or team-level predictions based on player data and / or team data, as discussed herein.

[0046] As discussed above, it may be difficult to achieve an acceptable threshold level of prediction accuracy for sporting events that occur during a sporting event. This difficulty may be due in part to the large number of variables involved. For example, various factors may affect what will happen in a given sporting event. For example, predicting the performance of a particular player may be challenging because the performance of that particular player may be affected based on the player's specific role, the performance of one or more other players, the opponent's team's strategy and structure, the current state of the game (e.g., if a team is leading 2-0, the team's offensive needs are reduced, with more players entering the defensive capacity), etc.

[0047] Furthermore, it can be difficult to account for certain unexpected events when making such predictions. For example, if a team is trailing 2-0 in a tiebreaker, one team might sacrifice defensive players in exchange for more offensive players. In this scenario, the prediction needs to account for offensive players taking more defensive positions. Therefore, when predicting a player's performance during a specific match, historical expectations of that player's performance may be insufficient. Other information, such as, but not limited to, the other players on both teams currently in the match, the current game environment, etc., can be used to improve the accuracy of the prediction, as discussed herein.

[0048] The systems and methods described herein describe a system that is configured to make one or more predictions for one or more players and / or for one or more teams within a sporting event. Furthermore, the predictions may relate to several actions for each player (e.g., shots taken, passes made, goals completed, fouls committed, etc.). Furthermore, these predictions may be real-time in that they are updated as the game progresses (e.g., in real time or near real time).

[0049] The actions that occur during a game (e.g., a match, a contest, a round, etc.) may be the product of a complex system of interactions between players and teams. The task of predicting the number of actions that will be performed by a single player over the duration of a game has several considerations. An exemplary prediction may be to predict the total number of successful passes made by a particular player at a particular point in time during the game. The prediction may be based on several factors, including but not limited to: the number of passes already completed by the player; the offensive strength of the player's teammates, and the strategy the player is executing; the defensive strength of the opposing team, and their strategy; the state of the game (i.e., the current score, time remaining, players sent off, etc.); the context of the game (i.e., the stage of competition in which the player's game is located, such as whether the game is a knockout final, or a season-ending match between two mid-table teams); and / or the likelihood that a player will be substituted, and at what time. It should be understood that the prediction may be based on one or more factors other than those exemplified above.

[0050] The systems and methods described herein may include models that capture interactions between players while incorporating factors such as team information, match state, and / or match environment. Thus, the system may be able to accept input features for a set of players and collectively make predictions for each of the players based on the impact of each of the other players. The system may further receive multi-resolution inputs for: players; teams; match state and environment; and / or features indicating a priori strengths and playing styles for each player and team in a principled manner. Predictions may be made for all players who could potentially play in a match (e.g., for a matchday lineup that includes the starting lineup and any corresponding substitutes).

[0051] In addition, the systems and methods described herein may be consistent in their predictions for different actions. In some cases, this consistency may be straightforward. For example, the number of shots on target for a particular player will not be greater than the total number of shots taken. In other cases, the consistency output generated by the system may be relatively more complex. For example, the total number of assists for all players on a team must not be greater than the total number of goals predicted for the team. In addition, the methods and systems disclosed herein can output predictions by associating the actions of multiple players and / or multiple teams. For example, a dominant team may tend to make more passes and shots, but fewer fouls and yellow cards. Therefore, player-based predictions can take into account the dominant trends based on the team. The system described herein can jointly predict the total number of games for all actions in a consistent manner by learning the correlations and patterns between actions (e.g., player actions, team actions, etc.).

[0052] Furthermore, the systems and methods described herein can be configured to update predictions in real time as events unfold during a sporting event. Thus, the predictions output by the system can be temporally consistent, such that the predicted total should never be less than the actual running total, be locally smooth (e.g., except when significant events such as goals occur), and converge to the actual total at the end of the game (e.g., as Figure 6 The system can also be configured to capture dynamics in the game, such as changes in momentum, the effects of goals, and / or player exits.

[0053] Thus, the system described herein can receive as input a set of features that are not represented in an implicit order, where the features can be represented as player vectors, team vectors, and / or game day vectors. The system can also be configured to predict multiple individual actions that are consistent with each other. The predictions can also be continuous (e.g., updated throughout the game) and consistent in time (e.g., updated as the game time passes). Furthermore, the system can be configured to capture momentum and dynamics in the game.

[0054] The system may include a transformer network as described herein, which may be configured to solve the constraints described herein. The transformer network may include layers that alternate between applying attention temporally to a sequence of action events and applying attention spatially across the set of players and teams at each event time step. The transformer neural network may include learned embeddings that are the output of the final transformer layer. These can be used to predict the total number of each action for each player using a single linear layer. Because each predicted action for a particular player is an affine transformation of the same embedding, prediction consistency for each action is maintained.

[0055] This Transformer Neural Network architecture may be resource and / or time efficient. For example, the Transformer Neural Network may not replicate any features from the input. The Transformer Neural Network may make multiple predictions per time step. For example, the Transformer Neural Network may predict a total of 672 predictions per time step for all players in the match-day lineup and for each of the 16 actions for each of the two teams, with approximately 1,000 predictions per game-day.

[0056] By using a transformer network architecture, the system can capture relationships between and within players. However, in many cases, bottom-up approaches may not align with team totals predictions, which are often quite "effective" because they can leverage market information (such as for sports betting). By utilizing team predictions using a top-down approach while ensuring that team totals are aligned with the market, the system can improve accuracy by assigning output predictions to each player using a graph neural network (GNN) (or "bottom-up" representation) model. Using only a transformer network architecture, the system can capture relationships between and within players. However, in many cases, bottom-up approaches alone may not align with team totals predictions, which are often quite "effective" because they can leverage market information (such as for sports betting markets). By utilizing team predictions using a top-down approach, the system can ensure that predictions for each player are aligned with market predictions. In other words, the system can use predictions generated by a top-down model to normalize predictions generated by a bottom-up model.

[0057] The transformer neural network described herein can be configured to make predictions for one or more players or teams simultaneously. By performing predictions simultaneously, the predictions may be more accurate than if the predictions were made non-simultaneously.

[0058] As used herein, a "machine learning model" generally comprises instructions, data, and / or a model configured to receive an input and apply one or more of weights, biases, classifications, or analyses to the input to generate an output. For example, the output may include a classification of the input, an analysis based on the input, a design, process, prediction, or recommendation associated with the input, or any other suitable type of output. Machine learning models are typically trained using training data, e.g., empirical data and / or samples of input data, which is fed into the model to establish, adjust, or modify one or more aspects of the model, e.g., weights, biases, criteria for forming classifications or clusters, etc. Various aspects of a machine learning model may operate on the input linearly, in parallel, via a network (e.g., a neural network), or via any suitable configuration.

[0059] Execution of the machine learning model may include deploying one or more machine learning techniques, such as linear regression, logistic regression, random forest, gradient boosting machine (GBM), deep learning, and / or deep neural networks. Supervised and / or unsupervised training may be employed. For example, supervised learning may include providing training data and labels corresponding to the training data, for example, as ground truth. Unsupervised methods may include clustering, classification, etc. K-means clustering or K-nearest neighbors may also be used, which may be supervised or unsupervised. A combination of K-nearest neighbors and unsupervised clustering techniques may also be used. Any suitable type of training may be used, for example, random, gradient boosting, random seeding, recursive, round-robin, or batch-based, etc.

[0060] While several of the examples herein relate to certain types of machine learning, it should be understood that the techniques according to the present invention may be adapted for any suitable type of machine learning. It should also be understood that the above examples are merely illustrative. The techniques and methods of the present invention may be adapted for any suitable activity.

[0061] Furthermore, while various aspects are discussed with respect to a given sport, these aspects are merely illustrative examples. The disclosed technology is in no way limited to any particular sport. For example, the present aspects can be implemented for use with other sports or activities, such as soccer, rugby (i.e., American football), basketball, baseball, hockey, cricket, rugby, tennis, and the like.

[0062] Figure 1 1 is a block diagram illustrating a computing environment 100 according to example aspects of the disclosed subject matter. Environment 100 includes a tracking system 102, a computing system 104, and a client device 108 connected via a network 105. In the depicted example, tracking system 102 obtains various measurements of a game and transmits the measurements across network 105 to computing system 104, where the measurements can be used in conjunction with one or more machine learning models. In one example, one or more machine learning models described herein can be configured to receive as input an embedding of player, team, and / or game information and determine one or more predictions for the players and / or teams.

[0063] Tracking system 102 can be in communication with venue 106 and / or can be located within, adjacent to, or near venue 106. Non-limiting examples of venue 106 include a stadium, a field, a venue, and a court. Venue 106 includes actors 112A-N (players, officials, coaches, objects, markers, etc.). Tracking system 102 can be configured to record the movements and actions of actors 112A-N on the playing field, as well as one or more associated other objects (e.g., a ball, a referee, etc.). Although environment 100 generally depicts actors 112A-N as players, it should be understood that, according to certain embodiments, actors 112A-N can correspond to players, objects, markers, etc.

[0064] In some aspects, tracking system 102 can be an optical-based system using, for example, a camera. While one camera is depicted, additional cameras are possible. For example, a system consisting of six fixed, calibrated cameras can be used that projects the three-dimensional positions of the players and ball onto a two-dimensional overhead view of the court.

[0065] In another example, a mix of fixed and non-fixed cameras can be used to capture the movement of all actors 112A-N on the playing field, as well as the movement of one or more related objects. Utilizing such a tracking system (e.g., tracking system 102), many different camera views of the field can be generated (e.g., high sideline view, free throw line view, huddle view, scrimmage view, end zone view, etc.). In some aspects, tracking system 102 corresponds to or utilizes a broadcast feed of a given game. In such aspects, each frame of the broadcast feed can be stored in a game file.

[0066] The tracking system 102 can be configured to communicate with the computing system 104 via the network 105. The computing system 104 can be configured to manage and analyze the data captured by the tracking system 102. The computing system 104 can include a network client application server 114, a pre-processing agent 116 (e.g., a processor and / or pre-processor), a data storage device 118, and a third-party application programming interface (API) 138. Figure 15 One example of computing system 104 is depicted.

[0067] The pre-processing agent 116 can be configured to process data retrieved from the data storage device 118 or the tracking system 102 before inputting it into the predictor 126. The pre-processing agent 116, the predictor and / or the predictive model analysis engine 122 can be composed of one or more software modules. One or more software modules can be a collection of codes or instructions stored on a medium (e.g., a memory of the organizational computing system 104), which represents a series of machine instructions (e.g., program code) that implement one or more algorithmic steps. Such machine instructions can be actual computer code that the processor of the organizational computing system 104 parses to implement the instructions, or alternatively, can be a higher-level instruction encoding that is parsed to obtain the actual computer code. One or more software modules can also include one or more hardware components. One or more aspects of the example algorithm can be performed by the hardware component (e.g., circuit) itself, rather than as a result of the instructions.

[0068] The data storage device 118 can be configured to store different types of data. In one example, the data storage device 118 can store raw tracking data received from the tracking system 102. The data storage device 118 can include historical game data, real-time data, features, and / or predictions. The historical game data can include historical team and player data for one or more sporting events. The real-time data can include data received from the tracking system 102, for example, in real time or near real time. The game data can include broadcast data or content related to a game (e.g., a match, a tournament, a round, etc.) and / or can include tracking data generated by the tracking system 102 or in response to data generated by the tracking system 102. The data storage device 118 can be configured to store features (e.g., feature vectors) generated for a particular sporting event, which include player, team, and game features.

[0069] According to aspects disclosed herein, the data storage device 118 can receive and / or store game files. A game file can include one or more game data types. Game data types can include, but are not limited to, position data (e.g., player position, object position, etc.), change data (e.g., position change, player change, object change, etc.), trend data (e.g., player trends, position trends, object trends, team trends, etc.), game data, etc. A game file can be a single game file or can be segmented (e.g., grouped by one or more data types, grouped by one or more players, grouped by one or more teams, etc.). The pre-processing agent 116 and / or the data storage device 118 can operate (e.g., using applicable code) to receive tracking data in a first format, store the game file in a second format, and / or output the game data in a third format (e.g., to the predictor 126). For example, the pre-processing agent 116 can receive an intended destination for the game data (or data generally stored in the data storage device 118) and format the data into a format acceptable to the intended destination.

[0070] The predictor 126 includes one or more machine learning models 128A-N. Examples include transformer neural networks, graph neural networks, recurrent neural networks, convolutional neural networks, and / or feedforward neural networks. The predictor 126 may include a system that implements a series of neural network instances (e.g., feedforward network (FFN) models) connected via transformer neural networks (e.g., graph neural network (GNN) models). FFNs can be unidirectional networks such that information in the model flows in a forward direction from input nodes or layers, through hidden nodes or layers, and to output nodes or layers (e.g., without any cycles or loops). FFNs can be trained using backpropagation techniques to iteratively train the model by calculating the necessary parameter adjustments to gradually minimize error. GNNs can capture the dependencies of a graph (e.g., its corresponding nodes and edges) via message passing between one or more nodes of the graph.

[0071] Each neural network instance can be configured to generate predictions for one or more actions for each player or team during a game. Furthermore, the neural network instance can forward feature vector information or generated predictions to a transformer neural network to transfer such information between neural network instances. This information transfer across neural network instances can allow player-to-player interactions to be factored into predictions, which can improve prediction accuracy, as described herein.

[0072] The machine learning models 128A-N may include top-down models to generate team predictions for a sporting event. The top-down models may be fed with features and / or other information for the sporting event, such as third-party predictions (e.g., provided by a third-party application programming interface (API) 138). The top-down models may implement a feed-forward neural network or another similar machine learning algorithm.

[0073] The machine learning models 128A-N may include bottom-up models for generating player predictions for a sporting event. The bottom-up models may be fed with team predictions and / or features for the sporting event in generating player predictions. For example, the bottom-up models may implement a transformer neural network or another similar machine learning algorithm, such as a graph neural network (GNN). The bottom-up models (e.g., transformers) may be configured to receive output from the top-down models. The bottom-up models may utilize the top-down model outputs as input.

[0074] Client device 108 can communicate with computing system 104 via network 105. Client device 108 can be operated by a user. For example, client device 108 can be a mobile device, a tablet computer, a desktop computer, or any computing system with the capabilities described herein. Users can include, but are not limited to, individuals such as, for example, subscribers, clients, potential clients, or customers of an entity associated with computing system 104, such as individuals who have obtained, will obtain, or may obtain products, services, or consulting from an entity associated with computing system 104.

[0075] The client device 108 may include one or more applications 109. The applications 109 may represent a web browser that allows access to websites or may be stand-alone applications. The client device 108 may access the applications 109 to access one or more functions of the computing system 104. The client device 108 may communicate over the network 105 to, for example, request a web page from a web client application server 114 of the computing system 104. For example, the client device 108 may be configured to execute the applications 109 to access content managed by the web client application server 114. The content displayed to the client device 108 may be transmitted from the web client application server 114 to the client device 108 and then processed by the applications 109 for display via a graphical user interface (GUI) of the client device 108.

[0076] The client device may include a display 110. Examples of the display 110 include, but are not limited to, a computer monitor, a light emitting diode (LED) display, etc. Output or visualization generated by the application 109 (e.g., a GUI) may be displayed on or using the display 110.

[0077] The functionality of the subcomponents shown in computing system 104 may be implemented in hardware, software, or some combination thereof. For example, a software component may be a collection of code or instructions stored on a medium, such as a non-transitory computer-readable medium (e.g., a memory of computing system 104), that represents a series of machine instructions (e.g., program code) that implement one or more method operations. Such machine instructions may be actual computer code that a processor of computing system 104 parses to implement the instructions, or alternatively, may be a higher-level instruction encoding that is parsed to obtain the actual computer code. One or more software modules may also include one or more hardware components. Examples of components include processors, controllers, signal processors, neural network processors, and the like.

[0078] The network 105 can be of any suitable type, including a separate connection via the Internet, such as a cellular or Wi-Fi network. In some aspects, the network 105 can use a direct connection to connect terminals, services, and mobile devices, such as radio frequency identification (RFID), near field communication (NFC), Bluetooth, or other similar communication methods. TM , Bluetooth Low Energy TM (BLE), Wi-Fi TM 、ZigBee TM , Ambient Backscatter Communication (ABC) protocol, USB, WAN, or LAN. Because the information transmitted may be personal or confidential, security concerns may require that one or more of these types of connections be encrypted or otherwise protected. However, in some aspects, the information transmitted may not be so personal, and therefore, a network connection may be chosen for convenience rather than security.

[0079] The network 105 may include any type of computer networking arrangement for exchanging data or information. For example, the network 105 may be the Internet, a private data network, a virtual private network using a public network, and / or other suitable connections that enable components in the computing environment 100 to send and receive information between components of the environment 100.

[0080] Transformer Neural Networks for Player and Team Prediction in Soccer

[0081] In one example, before a sporting event occurs, a predictor (e.g., predictor 126) can use team-level features and regularize the total number of predictions for upcoming events based on the team-level features. A top-down model can be implemented to incorporate a neural network to generate initial predictions based on match context information. Figure 4 and Figure 5 Such top-down models are further discussed. As used herein, a top-down model may generally refer to a model that observes or predicts a team's performance as a whole, and may also be used to assign such observations or predictions to one or more players on a team. Furthermore, in some cases, a top-down model may retrieve and / or utilize third-party (or "market") forecast information to compare with the forecasts generated as described herein to improve the accuracy of forecast generation. As used herein, a bottom-up model may generally refer to a model that analyzes each action or attribute associated with an event or player and makes observations or predictions based on the actions or attributes associated with a given player or team.

[0082] Conventional systems can utilize machine learning models to make player and team predictions by leveraging market data. The problem with using market data alone is that it may not be possible to combine player (and intra-player) and style-specific data to make accurate predictions based on factors other than the influence of a given player. Market data might refer to statistical odds of a particular team winning, points scored, game totals, etc. Furthermore, market data may not take into account cross-relationships between multiple groups of players or between multiple teams.

[0083] The present embodiment generally provides a transformer machine learning model and an associated data feed that receives data from multiple machine learning instances across sporting events to generate real-time outputs. The transformer machine learning model can be implemented using a transformer network hierarchy that includes a GNN that connects multiple models (e.g., feed-forward neural networks) that generate predictions for each player and team for a specific sporting event (e.g., a football game).

[0084] Furthermore, predictions made by conventional systems may not adequately capture correlations between teammates, opposing teams and players, and other contextual features (e.g., as shown below in Figure 2B ). Conventional systems including isolated models can make substantially more model calls for sporting events than systems incorporating a hierarchy of machine learning models, as described herein.

[0085] Figure 2A is a graph 200 depicting an example correlation between input features 202 and a predicted distribution 204 , according to an example embodiment. Figure 2AThe input features can be configured as input to Figure 3A The predicted distribution 204 can be Figure 3A The output of the transformer-based neural network 302.

[0086] Input features 202 may include player, team, and match-specific features. For example, match features may include crowd size, field capacity, weather conditions, season timing, team head-to-head record, and the like. Team features may include recent form, number of days since the last match, ranking or position, coaching information, team composition, injury information, substitution information, and the like. Player features may include physiological characteristics, recent form, player statistics, player attributes (e.g., height, weight, age, etc.), player trends, player preferences, and the like. In addition, input features 202 may also include in-game data, such as current match statistics, player performance, substitutions, and / or time remaining in the match. For example, input features 202 may be input into the transformer neural network as a tuple of input tensors. For example, a tuple of three tensors may be provided, wherein the first tensor corresponds to all players in the match, the second tensor corresponds to the two teams in the match, and the third tensor corresponds to the match status. A given tensor may be composed by stacking single-dimensional vectors (e.g., one-dimensional tensors, such as a player tensor, a team tensor, a match tensor, etc.) to form a two-dimensional tensor. Prediction distribution 204 may correspond to a prediction for each player and team for a game. Each prediction distribution 204 (e.g., distributions 204a, 204b, 204c, 204e) may correspond to a prediction for a particular player or team. For example, there may be a prediction distribution for each player and team in a game. For example, the prediction may be an empirical estimate of a sports statistic, such as the number of shots taken, the number of goals scored, the number of passes passed, etc. The output may include a distribution estimate for a random variable related to the number of shots. For example, prediction distribution 204 may be output as a learned embedding by including a tuple of tensors.

[0087] Figure 2B is an illustration of a set of neural networks (eg, a feed-forward network (FFN) 206 ) that derive a predictive distribution 204 for a set of input features 202 , according to an example embodiment.

[0088] In some cases, a neural network may be assigned to each player and team. Figure 2B is an illustration of a neural network that derives a predictive distribution from input features for each player, team, and match. Figure 2BAs shown, each input feature 202 can be specific to a player, team, or match. Furthermore, a neural network (e.g., a feed-forward network (FFN) 206) can be assigned to each player, team, and match to predict the actions of each corresponding player, team, etc. (e.g., the number of shots taken during the match). However, an architecture with FFNs assigned to each individual player and team may not interact with each other. This may prevent interaction between models, which may not capture interactions between players that are part of the same team. Figures 3A to 3D The system described in can overcome these problems.

[0089] As follows Figures 3A to 3D As described in

[15] , a neural network layer can be connected to each of multiple separate models (e.g., instances) to capture the interactions between players and teams. For example, such a neural network can include a graph neural network (GNN) that is configured to predict one or more actions for each player, team, game, etc. Figure 2B Compared to the distributed neural network shown in Figure 3A The neural network can replace the distributed FFN model, making Figure 3A A neural network (e.g., a GNN neural network) collectively makes all predictions. As a specific example, a transformer neural network can be used to output predictions related to a sporting event. Although transformer neural networks are generally discussed herein, it should be understood that any applicable GNN or other neural network that can be interpreted using a graph can be used to perform the techniques discussed herein with reference to transformer neural networks. As discussed herein, a "graph" can refer to a network of nodes and their corresponding edges. The transformer neural network can be configured to pass information (e.g., feature vectors, predictions) between network nodes, such as those representing and / or assigned to each player, team, match, or similar state, to enhance the predictions made by each model to take into account the context of the game and player interactions. The GNN can be configured to include self-attention, which can allow inputs to interact with each other ("self") and determine who or which node they should focus on ("attention"). The GNN can be fully connected (e.g., there may be an edge between each pair of nodes), and the self-attention mechanism can be configured to pass messages between pairs of connected nodes so that they can update their state. Once the node state has been updated through the mechanism described herein, simultaneous predictions can be made for each node. For example, predicted market data determined from a separate system can be input into the transformer neural network described herein. The output of the transformer neural network can include an aggregation of interactions, including interactions related to the predicted market data.

[0090] In some cases, the model can access third-party (or market) data that provides predictions made by third-party sources. The model can also access other top-down data, such as league embedding data, referee embedding data, etc. The third-party data can be output from a separate neural network system. This market data can be input as an embedding into the transformer neural network described herein.

[0091] Furthermore, transformer neural networks can be used to model and predict player-to-player interactions for sporting events. As mentioned above, for each sporting event, the difference between players on the same team and players playing against that team can significantly alter a given player's predicted performance. For example, if a player is paired against a strong defensive team, the player's predicted number of passes may be lower than the overall average number of passes for that player. Transformer-based neural networks can use team-level features and regularize the overall prediction.

[0092] In some cases, the input to the machine learning hierarchy as described herein may include proposed modifications to one or more aspects of a sporting event. For example, a request may include a proposal to replace a first player with a second player in a lineup for a football game. The machine learning hierarchy as described herein may include one or more natural language processing (NLP) techniques to parse the query and update features to account for the proposed changes. For example, a network node in a graph neural network may be assigned to each player and team for a football game. Predictions may be made for each player and team, and the graph neural network may facilitate interactions between each feedforward neural network to better account for player interactions and other game context for the sporting event.

[0093] For example, the task can be to predict the total number of shots that each player and team will take before a sporting event begins or during a sporting event. A series of input features can be obtained and processed to derive a predictive distribution for each player and team that predicts the shots taken.

[0094] Figure 3A An exemplary transformer-based neural network 302 for player and team predictions according to an example embodiment is included.

[0095] like Figure 3AAs shown, input features 202 can be divided into bottom-up features 304 and top-down features 306. Bottom-up features 304 can include features (or feature vectors) for each player and team, while top-down features 306 can be specific features (or feature vectors) of the game. For example, bottom-up features 304 and top-down features 306 can be formatted as tensors, which can be defined as multidimensional data arrays. Tensors can be organized in tuples. Example tensors within a tuple may be related to player feature vectors, player strength features, team features, team strength features, real-time game features, and game background features (context features). These input features can be described in more detail below.

[0096] Input features 202 can be fed into each node in the transformer network 302 corresponding to each player, team, and match to allow interaction between nodes. Each node included in the GNN (e.g., a transformer-based neural network 302) can derive a predictive distribution 204 for the corresponding player or team. Nodes can obtain player interactions and other predictive data generated by other node instances to enhance the predictions made by each node.

[0097] As mentioned above, GNNs allow nodes to interact with each other. These interactions can be modeled in a manner similar to language modeling, which helps understand the meaning of words in a sentence. In this example, language modeling can use the meaning of the word itself and the context provided by other words in the sentence. For each word, a model can be used to understand other words in the sentence that should be considered.

[0098] The transformer-based neural network 302 can process the input features to derive the predicted distribution 204 for each player, team, and game. As an example, the predicted distribution 204 can include the predicted total number of passes made by each team for each team, and the total number of passes made by each team during the game.

[0099] In some cases, the transformer-based neural network 302 can make a variety of predictions (e.g., 16 types). For each player and team, example predictions that are output by the transformer-based neural network 302 can include predicted goals, assists, shots, shots on target, passes, fouls, yellow cards, red cards, and minutes (player only). In addition, the model call can be made multiple times at a given time (e.g., Features may include historical features for games, teams, players, referees, and features used in games for teams and players (e.g., event counts).

[0100] In some embodiments, the player output generated by the bottom-up model 300 can be normalized by the total team prediction output generated by the top-down model (e.g., by the top-down features 306). For example, if the total number of predicted passes for each player on a team is 300, but the team's predicted number of passes is only 280, the player's predicted number of passes can be normalized to match the team's total of 280.

[0101] The top-down features 306 may be output from a separate machine learning system (e.g., Figure 4 Model 404 as depicted in, or as Figure 5 ). In another example, the top-down features 306 can be extracted from market information (e.g., odds for sports) or game context information and / or can be determined directly by a single or non-machine learning system. The top-down features 306 can be input into the transformer-based neural network 302, as discussed further herein. The top-down features 306 can be received via a feed (e.g., an API connected to a separate system) or input via an interface accessible to a user. The top-down features 306 can align the overall statistics of the entire team with information that does not exist in history (e.g., such as information provided at least in part by the bottom-up features 304). The top-down features 306 can work as a wrapper or interaction layer on top of the transformer-based neural network 302.

[0102] Thus, top-down features 306 may include one or more of match predictions (e.g., goals, shots, shots on target, assists, passes, fouls, yellow cards and / or red cards, etc.) based on one or more machine learning models (e.g., neural networks that output team-level predictions), market information, and / or match context information. Such top-down features may be provided before a given sporting event (e.g., based on pre-match data) and / or may be provided during the match (e.g., based on data on actions, events, situations, states, etc. that occur during the sporting event), and may be updated throughout the given sporting event. For example, top-down features may be provided during the match based on actions or events that occur during the sporting event. Such in-match top-down features may be provided over a long period of time (e.g., upon expiration of a predetermined or dynamically determined time period) and / or may be provided upon the occurrence of a condition that triggers the top-down feature. The dynamically determined time period may be, for example, a time period that is adjusted based on the duration of the expiration time or a given in-match time. For example, in-match top-down features may be provided at a lower frequency during the middle of a given sporting event than at a higher frequency immediately before the end of the sporting event. The conditions that trigger the top-down features can be actions or events that occur during the game (e.g., score, timeout, free throw, change of possession, substitution, etc.). The top-down features 306 can be used to enhance, update, and / or replace the prior inputs to and / or the prior predictions made by the transformer-based neural network 302 in real time or near real time.

[0103] As used herein, market information may refer to odds for a sporting event or may be based on odds for a sporting event. Odds for such sports may depend on market liquidity and / or may tend to be effective for a major market (e.g., for a total number of teams). Odds for such sports may include probabilities for a given outcome (e.g., an action), where the outcome may be a binary outcome (e.g., win or lose), an occurrence (e.g., a score), an occurrence (e.g., a hole-in-one), etc. Market information may be generated (e.g., based on historical data), may be provided by an individual (e.g., via user input using an interface or via an API), and / or may be provided by one or more entities, automated systems, or individuals. According to one example, market information may be provided by multiple entities, which provide multiple input data points that may be used by the transformer-based neural network 302 to output one or more predictions. Alternatively, market information provided by multiple entities may be filtered or normalized (e.g., averaged, weighted and averaged, transformed, etc.), and such filtered or normalized market information may be provided as top-down features 306 to the transformer-based neural network 302.

[0104] As used herein, game context information may refer to information related to a sporting event. Such game context information may be generated or provided by an expert, an expert system, and / or a specialized system. For example, such game context information may include one or more of environmental conditions (e.g., weather conditions, environmental conditions, etc.), player information (e.g., injury information, mentality information, training information, travel information, health information, etc.), expert predictions (e.g., team-level predictions, player-level predictions, etc.), crowd information (e.g., crowd density, crowd size, crowd excitement, crowd demographics, etc.). Game context information may be provided by an individual (e.g., via user input using an interface or via an API) and / or may be provided by one or more entities, automated systems, or individuals.

[0105] The transformer-based neural network 302 can accept input features 202 at different resolutions and in a temporal sequence. The transformer-based neural network 302 can use a single transformer encoder backbone to learn an embedding that is configured to predict multiple actions for each agent (e.g., each player and / or team). The transformer-based neural network 302 backbone can be a series of axial transformer encoder layers, where each layer alternately applies attention along the time and agent dimensions.

[0106] An exemplary objective of the transformer-based neural network 302 may be to predict the total number of actions at the end of a sporting event for all players participating in the sporting event at time step t. Each objective may be a discrete random variable Ya,p, where a is the action being predicted and p is the player. These objectives may correspond to the prediction distribution 204 of the transformer-based neural network 302. According to an embodiment, the transformer-based neural network 302 may be configured to not predict the joint distribution of all objectives received as input. For example, the model may not predict: P(Y): Y = {Ya,p|a∈A,p∈P}, where A is a set of actions and P is a set of players. The model may not be able to predict this for 22 players and 16 actions, and the space of possible outcomes, which would be at least 2^(22×16), may not be tractable.

[0107] The transformer-based neural network 302 may have been trained to predict the marginal distribution of each target. The model may utilize the following function: where Y\Ya,p represents the set of all targets except Ya,p. This can allow the transformer-based neural network 302 to approximate the joint distribution and use this approximation to determine the marginal distribution of each target. This can ensure consistency between the distributions of each target.

[0108] Figure 3B is a method for Figure 3AAn exemplary model of an input tensor for a transformer-based neural network 302. Figure 3B The input tensor may correspond to Figure 2A and Figure 3A Input features 202.

[0109] The transformer-based neural network 302 is trained on a set of training samples X, where N = |X| is the number of training samples. The granularity of each training sample is a single game, and for each game i, there are T games in which action events occur. (i) time steps. There is also a P in the match day lineup (i) players, and two teams participate in the game i.

[0110] Each training sample can be a tuple (X (i) ,Y (i) ), where X (i) , is a tuple of input tensors, and Y (i) is a tuple of target tensors for game i.

[0111] The input tuple can contain the following tensors:

[0112] is a tensor of real-time player features, where D 球员 is the dimension of the player feature vector. The features of contain indicator features for player position and team, as well as a running total of moves that have been made by the player.

[0113] is a tensor of player strength characteristics, where D 球员实力 is the dimension of the player strength feature vector. The player strength feature can be configured to capture the player's prior strength and is primarily an aggregated statistical data of the player's actions in previous games, such as, for example, the average number of passes in the previous 5 games, the maximum number of fouls in the previous 10 games, etc. In addition, there are features for the time since the last game, the distance from the player's home court, and the distance from the previous game position.

[0114] is a tensor of real-time team features, where D 球队 is the dimension of the team feature vector. The features contain indicator features for the teams, as well as a running total of moves that have been made by the teams, similar to real-time player features.

[0115] is a tensor of team strength characteristics, where D 球队实力 is the dimension of the team strength feature vector. Team strength features are mainly aggregated statistics of the team’s actions in previous games, similar to player strength features.

[0116] is a tensor of real-time game status features, where the feature vector has dimension D 比赛 The Game Status feature contains properties of the current event, such as event type, game clock time, and event location.

[0117] is of dimension D 比赛背景 The event context features capture the context in which the game was played, such as indicator features for the competition (such as the league the game belongs to), the duration the game was played, etc.

[0118] The input tensor X mentioned above (i) Each of these can correspond to Figure 3A Input features 202.

[0119] according to Figure 3B The input tensor X (i) It can be arranged along the time and agent dimensions. The columns can represent agents (e.g., corresponding tensors for players, teams, and matches), and the rows can represent the time aspects of the match. Column 312 represents the initial input to the aforementioned tensors. Column 314 depicts how the input tensors are provided over time throughout the match.

[0120] Check the input of the transformer-based neural network 302 (e.g., the input tensor X (i) ), features are not copied. Pass tensor X (i) Each feature provided as input can be categorized by its granularity. For example, the granularity can be (1) player level, (2) player frame level, and / or (3) match level. Each granularity can be represented by a tensor within a tuple received by the transformer neural network. The transformer-based neural network 302 can rely on attention to learn which features at each granularity are relevant to predictions at other granularities (e.g., how important the current game state is for predicting a particular player).

[0121] Target tuple Y (i) Contains tensors for each of the modeled actions for each player and / or team. For a given action a, the target tensor contains the tensors for The remaining number of actions for each player in . For example, a can represent the number of shots, which is calculated for time steps ∈[1,T (i) ], and player p∈[1,p (i) ],Y a (i)can represent the number of shots remaining for player p at time step 5. The count of remaining actions can be used as the true value rather than the total count, as this facilitates making distributional assumptions, in particular using a Poisson distribution for the interval [t,T (i) ] to model the remaining actions in the target tuple Y. (i) can be used for training, and the tuple Y (i) May correspond to the predicted distribution 204 of the transformer-based neural network 302 .

[0122] Figure 3C According to an example embodiment Figure 3A 316 of an exemplary model of a transformer-based neural network 302 with a linear embedding layer 317, a transformer encoder layer 318, and a fully connected layer 319. The embedding layer 317 can map the input component tensors into a tensor with a common feature dimension. The transformer encoder layer 318 can perform attention along the time and agent dimensions. The fully connected layer 319 can map the output embedding from the final transformer layer into a tensor with the requested feature dimension for each target metric.

[0123] Examining the embedding layer 317, the embedding layer 317 may contain a linear block for each input tensor, and each block may map the input tensor into a tensor with a common feature dimension D. For example, for the player input tensor, the linear block may be represented by the following function:

[0124] The output of a linear layer may be a tuple of tensors, each of which has a common feature dimension, and the tensors can be concatenated along the time and agent dimensions to form a single tensor:

[0125] Next, the tensor can be passed to a transformer encoder layer 318 (e.g., a series of L axial transformer encoder layers, e.g., time linear embedding 321). The axial encoder layer accepts a tensor of dimension 3 and applies attention first along the time dimension and then along the agent dimension. During the training phase, an autoregressive attention mask is applied (i.e., an upper triangular mask with an offset of 1) so that elements at a time step in a layer can only attend to previous time steps. This may be followed by standard feedforward, layer normalization, and skip connections, such as Figure 3D Depicted.

[0126] Figure 3D Including according to example embodiments Figure 3A An exemplary model 320 of an axial attention layer of a transformer-based neural network 302. The model 320 may be located at Figure 3Cwithin the transformer encoder layer 318 .

[0127] The temporal attention step 322 may allow agent j in time step i to pay attention to the previous time step [1..i-1]. The agent attention step 324 may allow agent j in time step i to pay attention to all other agents in the current time step (e.g., to other players, teams, and game status). Next, in a residual connection and layer normalization step 326, the model may apply layer normalization to the sum of sub-layer inputs and outputs to update and stabilize the model. Next, the steps of feedforward processing 328 and further residual connection and layer normalization steps 330 may be applied.

[0128] The attention mechanism implemented by the Transformer Encoder layer 318 can have a graphical interpretation on a dense graph, where each element is a node and the attention mask is the inverse of the adjacency matrix defining the edges between nodes (thus, the absence of the attention mask implies a fully connected graph). In the case of axial attention used here, with attention masking on the time (row) dimension, the nodes in the graph can be arranged in a grid, and each node can be connected to all nodes in the same column, as well as to all previous nodes in the same row. In this case, attention can be message passing, where each node can accept the state of nodes depicted near it and then update its own state based on these messages. This attention scheme can mean that when making predictions for a particular player, the model may consider (i.e., attend to): nodes containing the player's previous states along the time series; as well as nodes containing the states of other players, teams, and the current game state at the current time step. Nodes may not need to be homogeneous, except for having the same feature dimension, and thus a node representing a player can receive messages from a node representing their team or from a node representing the player's strength. Thus, the model can learn interactions between agents and ensure consistent predictions for each agent along the time series.

[0129] Since the transformer encoder layer 318 includes two or more layers, the state of any node can indirectly be aware of the input states of all other nodes, except for the nodes in subsequent time steps. This can eliminate the need to copy any information in the input and thus support Figure 3B A valid input representation of a tensor described in .

[0130] The output of the transformer encoder layer 3181 is a tensor of the same shape as the input, and thus the layer is a function represented as follows:

[0131] The output embedding of the previous layer can be used as the input of the current layer. Later in EL (i) Each vector of the final output of [i, j,;] can be the embedding of agent j at time step i, which has been updated by the attention layer to contain information from all input features up to and including the current time step. This embedding can then be used to directly make predictions for the desired target.

[0132] The final layer of the transformer-based neural network 302 may be a fully connected layer 319. These layers may map the output embeddings of the final transformer layer of the transformer encoder layer 318 to the feature dimensions of each target metric. Perhaps only the player at each in-game event needs to be predicted, and the output embedding Z may be used. (i) =E L (i) The input to the linear layer is a slice of [1; 4; :]. For example, if the goal is the final goal count of a player and the model uses a Poisson distribution, you may want to estimate a single parameter λ for each player at each time step. The linear layer can be a function like:

[0133] Different distributional assumptions can be made, and the only difference in the models may be the output parameter dimensions of the corresponding linear layers. The transformer-based neural network 302 described herein can utilize assumptions including: Bernoulli, Poisson, log-Gaussian, and "model-free" discrete distributions.

[0134] Training of the transformer-based neural network 302 is further described herein. A corresponding loss function can be selected for each distribution hypothesis of the output target. For example, the loss function can be the Poisson negative log-likelihood for a Poisson distribution, the binary cross entropy for a Bernoulli distribution, etc. During training, a loss can be calculated based on the ground truth value of each target in the training set, and the loss values ​​can be summed, and an optimizer can be used to update the weights of the model based on the total loss. The learning rate may have been adjusted according to a cosine annealing schedule without the need for warm restarts.

[0135] As discussed above, the transformer-based neural network 302 may be configured to receive top-down features 306 as input. Figure 4 Depicted is an exemplary top-down prediction model 404. The output of the top-down prediction model 404 may be the top-down features 306 input that the transformer-based neural network 302 may receive.

[0136] For example, before a sporting event occurs, a predictor can use team-level features and regularize predictions of the total number of upcoming events. Top-down models can be implemented to incorporate neural networks to generate initial predictions based on match context information. Furthermore, in some cases, top-down models can retrieve and / or utilize third-party (or "market") forecast information to compare with the forecasts generated as described herein to improve the accuracy of forecast generation.

[0137] More specifically, Figure 4 Depicted is a process 400 of an example top-down prediction model 404 according to an example embodiment. Figure 4 As shown, a top-down prediction model 404 can obtain game context information 402. Game context information 402 can be derived from a source of pre-game information 408. Pre-game information 408 can include historical information related to each team identified as part of the upcoming game, such as historical team pass feature data 410, historical team home / away pass features 412, and historical team pass-and-miss features 414. Pre-game information 408 data can be obtained from one or more data sources, and the data can be combined to create a set of game context information 416. Game context information 416 can refer to compiled pre-game information 408, which can be configured to be received by a neural network. For example, the game context can be input as a single vector. It should be understood that the game context 402 features provided herein are merely examples, and one or more other game context 404 features can be used based on the game type, input data, pre-game information 408, etc.

[0138] The combined game context information 416 may be provided to the neural network model 404. For example, the neural network model 404 may utilize machine learning techniques such as linear regression, logistic regression, random forest, gradient boosting machine (GBM), deep learning, and / or deep neural networks. The neural network model 404 may process the feature vectors included in the game context information 416 for each team to derive one or more specified predictions 406 for each team. For example, the output may include predictions 406 including the number of passes 420 for the first team at the end of the game and the number of passes 422 for the second team at the end of the game. These predictions 406 may be organized into features and fed as top-down features 306 to the neural network model 404. Figure 3A In the transformer-based neural network 302.

[0139] In addition, the top-down prediction model can be extended to in-game predictions. For example, a prediction model as described herein (e.g., model 404) can update predictions during a game to take into account contextual information that occurs during the game.

[0140] Figure 5FIG. 5 is a flow chart 500 of an example top-down prediction model 504 during an example sports game, according to an example embodiment. Figure 5 As shown, game context information 502 can include both pre-game information 508 and in-game information 510. For example, pre-game information 508 can include historical information such as each team's historical passes per second and each team's pass percentage. For example, in-game information 510 can include a team's passes per second, goal difference, and red card difference during the game. Pre-game information 508 and in-game information 510 can be combined to create game context information 502. A top-down prediction model 504 can capture the game context and generate a prediction 506 for the remainder of the game. For example, prediction 506 can include remaining game passes per second or a team's pass percentage (e.g., predicting which portion of the total passes a team will execute in the game).

[0141] Prediction 506 can be combined with current game statistics to generate output 507 for the remainder of the game. Output 507 can include predictions based on the game, such as the remaining passes of the game (e.g., remaining game passes per second and remaining seconds). Prediction 506 can also include predictions based on teams, such as the remaining passes of team 1 (remaining game passes per second, remaining seconds and the passing ratio of team 1). Output 507 can also include, for example, the remaining passes 418 of team 2 (remaining game passes per second, remaining seconds and the passing ratio of team 1). For example, output 507 can be dynamically updated as the game proceeds, and top-down prediction model 504 can obtain new information, such as the remaining seconds in the game and the current statistics of the game. Therefore, prediction 506 can be a point-in-time prediction, which can be updated within the entire duration of a given game. Prediction 506 can be based at least in part on time-based input (e.g., elapsed time, remaining time, team possession time, etc.).

[0142] In some cases, in-game predictions can be performed by updating pre-game predictions using a decay function. A pre-game prediction can be generated, and then, to obtain a real-time in-game prediction, the number of minutes in the game can be considered as part of the decay function. For example, in-game prediction = pre-game prediction * (90-60) / 90 + actual event count at 60 minutes. As another example, in-game prediction = (100 passes) * (90-60) / 90 + (50 passes so far) = 85.5 passes.

[0143] Exceptions can be made in some cases, such as when a player is sent off due to a red card or substituted. In these cases, the "prediction" can be the actual number up to that moment. For example, for in-game team predictions, the model could predict the final pass count for each team during the game.

[0144] As an illustrative example, a system as described herein may dynamically update predictions (eg, the total number of passes in a game) as the game progresses. Figure 6 6 is an example graphical representation of the predicted number of passes and the actual number of passes during a sporting event. The graphical representation 600 shows the number of passes (y-axis 602) as a function of time (minutes played along the x-axis 604). First, the number of major events for each team can be tracked. For example, major events can include goals, as depicted by circles, or red cards, as depicted by squares, which are tracked with the corresponding time each major event occurs. Each major event can modify the prediction, such as a red card reducing the predicted number of passes (e.g., as shown in 606).

[0145] Graphical representation 600 may also include the current actual pass total 610 as the game progresses and the final actual pass total 608. As the game progresses, the model described herein may be updated so that the predicted pass count more closely matches the current actual pass count and the final pass total at the end of game time.

[0146] Cloud computing pipeline

[0147] The cloud computing pipeline can use a processing pipeline to generate pre-game and / or in-game predictions. The processing pipeline can include a series of processing steps, which are also called "Lambdas". Such processing steps or lambdas can refer to programming language features that enable the creation of functions. For example, a processing pipeline can include 6 processing steps, where the goal of the pipeline is to respond to a specific event, trigger a model, generate a prediction, and send the prediction to a defined client. For example, you can use Figures 7 to 11 The exemplary processing steps depicted in Figures 3A to 3D Transformer-based neural networks.

[0148] The first processing step may include a trigger processing step that can trigger the pipeline. For example, the trigger processing step may include checking whether any sporting events are scheduled to start in a specified time window (e.g., 11 to 12 hours before the start of the game) or in real time at a given or dynamic duration (e.g., every hour or every minute).

[0149] When relevant atomic actions related to the game are detected (e.g., a goal, a foul, a lineup announcement, etc.), such as when a game is about to start, when a game is live, when a game lineup is provided, or when the odds of a game change, the trigger processing step can trigger the model to generate a prediction for the event.

[0150] Figure 7 is an example process 700 for trigger processing step 702. Figure 7As shown, the trigger processing step 702 can periodically check a source (e.g., a data platform (Gold Standard Data Platform GSDP), Redshift Machine Learning Platform) at a first elapsed time (e.g., every hour 706) and periodically check a database (e.g., DynamoDB) at a second elapsed time (e.g., every minute 704). The trigger processing step 702 can obtain data from various sources (e.g., odds data source 708, data source 710, attribute activity match lookup 712) to determine whether a new event has been triggered.

[0151] The feature creator process step can be triggered by the trigger process step and can initiate the creation of a feature (or update a created feature) for an event. For example, if a game is in the upcoming window or a new lineup is available, the feature creator process step can query data from the data platform to create a feature and save the feature in the database.

[0152] Figure 8 is an example flow 800 for the feature creation process step 802. Figure 8 As shown in FIG, a feature creator process step 802 may be triggered by a trigger process step 702 and may generate features for a match, as described herein. The system may obtain historical data for each player and team to generate features for the match. The feature creator process step 802 may query data from a data platform (e.g., GSDP Redshift 806) and save the features in a database (e.g., feature storage 808). These features may then be automatically transferred and sent to Figure 3A Transformer-based neural network 302.

[0153] The attribute predictor process step can be triggered by both the trigger process step and the feature creator process step in response to the created (or updated) features. The attribute predictor can retrieve features from the feature store, send the features to a third-party API, receive predictions from the API, and send the predictions to the data platform, which can then be forwarded to another platform.

[0154] Figure 9 is an example flow 900 of the attribute predictor processing step 902. Figure 9 As shown, the trigger processing step 702 and the feature creator processing step 802 can trigger the attribute predictor processing step 902. The attribute predictor processing step 902 can obtain features from the feature storage device 808 and provide the features to the API 904. Predictions from one or more third-party sources can be obtained via the API 904 and can be fed into the data platform 906.

[0155] The team attribute predictor process step may be triggered by both the trigger process step and the feature creator process step in response to the created (or updated) features. The team attribute predictor process step may retrieve features (and in some cases, third party predictions) from the feature store, derive team predictions, and store the predictions in the prediction store. The team predictions may be made by a top-down model as described herein (e.g., as described in Figure 4 and Figure 5 as depicted in ).

[0156] Figure 10 1000 is an example flow of a team attribute predictor process step 1002. The team attribute predictor process step 1002 may be triggered by either the trigger process step 702 or the feature creator process step 802. In addition, the team attribute predictor process step 1002 may retrieve features from the feature store 808, derive team predictions, and provide the predictions to the prediction store 1004 and the data platform 1006.

[0157] The player attribute predictor processing step may be triggered by generating team predictions by a team attribute predictor processing step, e.g., 1002. The player attribute predictor may take features and team predictions and may generate player predictions using a bottom-up model as described herein.

[0158] Figure 11 1100 is an example flow chart of player attribute predictor processing step 1102. Figure 11 As shown, the trigger processing step 702 and the feature creator processing step 802 may trigger the team attribute predictor processing step 1002. The team attribute predictor processing step 1002 may trigger the player attribute predictor processing step 1102. The player attribute predictor processing step 1102 may obtain the team predictions generated by the team attribute predictor processing step 1002 (e.g., from a prediction storage device) and features previously generated for the sporting event (e.g., from 808). The player attribute predictor processing step 1102 may also generate player predictions using a bottom-up model as described herein (e.g., using Figure 3A The player predictions may be stored in a prediction storage device 1004.

[0159] The combined attribute predictor process step may be triggered to generate player predictions by a player attribute predictor process step (e.g., 1102). The combined attribute predictor may obtain team predictions and player predictions and may combine predictions for a sporting event. The combined attribute predictor process step may normalize the player predictions based on the corresponding team predictions.

[0160] Figure 12is an example flow 1200 of combining attribute predictors processing step 1202. Figure 12 As shown, trigger processing step 702 and feature creator processing step 802 may trigger team attribute predictor processing step 1002. Team attribute predictor processing step 1002 may trigger player attribute predictor processing step 1102. Player attribute predictor processing step 1102 may trigger combined attribute predictor λ 1202. Combined attribute predictor processing step 1202 may retrieve team predictions and player predictions from prediction storage 1004 and may normalize the player predictions based on the team predictions, as described herein. Combined attribute predictor processing step 1202 may interact with data platform 1206 to generate a combined prediction, and the combined prediction may be stored in the prediction storage.

[0161] Figure 13 is an example process 1300 for generating predictions for teams and players associated with a sporting event according to an example embodiment. For example, the process 700 may be performed by Figures 3A to 3D Transformer-based neural network 302 and Figure 4 Top-down prediction model 404 or Figure 5 The top-down prediction model 504 is used to perform.

[0162] At step 1302, one or more top-down predictions for a sporting event may be generated by a first neural network, the top-down predictions being team-level predictions. The one or more top-down predictions may be generated by incorporating at least one historical team information or market information as input into the first neural network. The team-level predictions may be predictions of one or more of a plurality of goals, shots, shots on target, assists, passes, fouls, yellow cards, or red cards for the match.

[0163] At step 1304 , the top-down predictions may be provided to a computing system as one or more top-down feature vectors.

[0164] At step 1306, the computing system may receive a second set of feature vectors, the second set of feature vectors including data for one or more players associated with one or more corresponding teams and data for one or more teams associated with the sporting event. The data for the one or more players may include movements of one or more actors on the playing field received from the tracking device.

[0165] The one or more top-down feature vectors and the second set of feature vectors may be input into a transformer-based neural network at step 1308. The transformer-based neural network may include a set of embedding layers, transformer encoder layers, and fully connected layers.

[0166] At step 1310, one or more predictions for the sporting event may be generated based on the one or more top-down feature vectors and the second set of feature vectors.The one or more predictions for the sporting event may be displayed on a display device.

[0167] Additionally, updated data for one or more players or teams may be received from the tracking device. The updated data may be provided to one or more of the first neural network or the transformer-based neural network. One or more updated predictions for the sporting event may be generated based on the updated data.

[0168] A trigger processing step can be used to access the data platform at set time intervals to determine when a sporting event occurs. A feature creator processing step can be used to create a second set of feature vectors by querying data from the data platform.

[0169] Overview of Neural Network Training and Computation Systems

[0170] Figure 14 Depicted is a flow chart for training a machine learning model according to one aspect. Figure 14 As shown in flowchart 1410 of , training data 1412 may include one or more of stage inputs 1414 and known results 1418 associated with the machine learning model to be trained. Stage inputs 1414 may come from any applicable source, including the components or sets shown in the figures provided herein. For machine learning models generated based on supervised or semi-supervised training, known results 1418 may be included. Supervised machine learning models may be trained without using known results 1418. Known results 1418 may include known or expected outputs for future inputs that are similar to or in the same category as stage inputs 1414 that do not have corresponding known outputs.

[0171] The training data 1412 and the training algorithm 1420 can be provided to a training component 1430, which can apply the training data 1412 to the training algorithm 1420 to generate a trained machine learning model 1450. According to one embodiment, a comparison result 1416 can be provided to the training component 1430, which compares the previous output of the corresponding machine learning model to apply the previous results to retrain the machine learning model. The training component 1430 can use the comparison result 1416 to update the corresponding machine learning model. The training algorithm 1420 can utilize machine learning networks and / or models, including but not limited to deep learning networks such as deep neural networks (DNNs), convolutional neural networks (CNNs), fully convolutional networks (FCNs), and recurrent neural networks (RCNs), probabilistic models (such as Bayesian networks and graphical models), and / or discriminative models (such as decision forests and maximum margin methods). The output of flowchart 1400 can be a trained machine learning model 1450.

[0172] The machine learning models disclosed herein can be trained by adjusting one or more weights, layers and / or biases during a training phase. During the training phase, historical or simulated data can be provided as input to the model. The model can adjust one or more of its weights, layers and / or biases based on such historical or simulated information. The adjusted weights, layers and / or biases can be configured with a production version of the machine learning model (e.g., a trained model) based on the training. Once trained, the machine learning model can output the output of the machine learning model according to the subject matter disclosed herein. According to one embodiment, one or more machine learning models disclosed herein can be continuously updated based on the use or implementation of the output of the machine learning model.

[0173] It will be understood that the aspects of the present invention are merely exemplary, and that other aspects may include various combinations of features from other aspects, as well as additional or fewer features.

[0174] Generally, any process or operation discussed in the present invention that is understood to be computer-implementable, such as the process shown in the flowchart disclosed herein, can be performed by one or more processors of a computer system, such as any one of the systems or devices in the exemplary environment disclosed herein, as described above. The process or process steps performed by one or more processors can also be referred to as operations. One or more processors can be configured to perform these processes by accessing instructions (e.g., software or computer-readable code) that, when executed by one or more processors, cause one or more processors to perform the process. Instructions can be stored in a memory of a computer system. The processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a processing unit of any suitable type.

[0175] A computer system, such as the system or device implementing the processes or operations in the examples above, may include one or more computing devices, such as one or more of the systems or devices disclosed herein. The one or more processors of the computer system may be included in a single computing device or distributed across multiple computing devices. The memory of the computer system may include a corresponding memory for each of the multiple computing devices.

[0176] Figure 15 is a simplified functional block diagram of a computer 1500 that can be configured as an apparatus for performing the methods disclosed herein according to exemplary aspects of the present invention. For example, according to exemplary aspects of the present invention, the computer 1500 can be configured as a system. In various aspects, any of the systems herein can be a computer 1500 that includes, for example, a data communication interface 1520 for packet data communication. The computer 1500 can also include a central processing unit ("CPU") 1502 in the form of one or more processors for executing program instructions. Although the computer 1500 can receive programming and data via network communications, the computer 1500 can include an internal communication bus 1508 and a storage unit 1506 (such as, ROM, HDD, SDD, etc.), which can store data on a computer readable medium 1522.

[0177] Although the instructions 1524 may be temporarily or permanently stored in other modules of the computer 1500 (e.g., the processor 1502 and / or the computer-readable medium 1522), the computer 1500 may also have stored therein instructions for performing the techniques presented herein, such as those related to Figure 13 Memory 1504 (such as RAM) for instructions 1524 for the described methods. Computer 1500 may also include input and output ports 1512 and / or a display 1510 for connecting to input and output devices such as a keyboard, mouse, touch screen, monitor, display, etc. Various system functions can be implemented in a distributed manner on multiple similar platforms to distribute the processing load. Alternatively, the system can be implemented by appropriately programming a computer hardware platform.

[0178] The programmatic aspects of the technology can be considered a "product" or "article of manufacture," which typically takes the form of executable code and / or associated data carried on or embodied in a type of machine-readable medium. "Storage"-type media include computers, processors, etc., or their associated modules, such as any or all of various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. Sometimes, all or part of the software can be communicated via the Internet or various other telecommunications networks. For example, such communication can enable software to be loaded from one computer or processor to another, such as from a management server or host of a mobile communication network to a server's computer platform and / or from a server to a mobile device. Therefore, another type of medium that can carry software elements includes optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and fiber optic route networks, and through various air links. Physical elements that carry such waves, such as wired or wireless links, fiber optic links, etc., can also be considered as media that carry software. As used herein, unless restricted to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.

[0179] Although the disclosed methods, apparatuses, and systems are described with exemplary reference to transmitting data, it should be understood that the disclosed aspects may be applicable to any environment, such as a desktop or laptop computer, a car entertainment system, a home entertainment system, etc. In addition, the disclosed aspects may be applicable to any type of Internet protocol.

[0180] It should be understood that in the above description of exemplary aspects of the invention, various features of the invention are sometimes grouped together in a single aspect, figure, or description thereof to simplify the invention and aid in understanding one or more of the various inventive aspects. However, this approach to the invention should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. On the contrary, as reflected in the following claims, inventive aspects exist in features that are less than all of the features of a single aforementioned disclosed aspect. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim existing independently as a separate aspect of the invention.

[0181] Furthermore, although some aspects described herein include some features included in other aspects and not other features, combinations of features from different aspects should be within the scope of the present invention and form different aspects, as will be understood by those skilled in the art. For example, in the following claims, any of the claimed aspects may be used in any combination.

[0182] Thus, while certain aspects have been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended that all such changes and modifications fall within the scope of the invention. For example, functionality may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Operations may be added or deleted from the methods described within the scope of the invention.

[0183] The subject matter disclosed above should be considered illustrative, not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other embodiments that fall within the true spirit and scope of the invention. Therefore, to the maximum extent permitted by law, the scope of the present invention should be determined by the broadest permissible interpretation of the following claims and their equivalents, and should not be restricted or limited by the foregoing detailed description. Although various embodiments of the present invention have been described, it will be apparent to those skilled in the art that more embodiments may be within the scope of the present invention. Therefore, the present invention is not limited except in accordance with the appended claims and their equivalents.

Claims

1. A method of generating predictions for teams and players on each team associated with a sporting event, the method comprising: receiving one or more top-down predictions related to the sporting event; providing the top-down predictions to a computing system as one or more top-down feature vectors; receiving, by the computing system, a second set of feature vectors comprising data for one or more players associated with one or more respective teams and data for one or more teams associated with a sporting event; inputting the one or more top-down feature vectors and the second set of feature vectors into a transformer-based neural network; as well as One or more predictions for the sporting event are generated using the transformer-based neural network based on the one or more top-down feature vectors and the second set of feature vectors.

2. The method of claim 1 , wherein the one or more top-down predictions are one or more of a neural network output, market information, or match context information. 3 . The method of claim 1 , wherein the top-down prediction comprises a first top-down prediction based on pre-game data and a second top-down prediction based on in-game data.

4. The method according to claim 1, further comprising: The one or more predictions for the sporting event are caused to be displayed on a display device.

5. The method of claim 1 , wherein the transformer-based neural network further comprises: A set of embedding layers; Transformer encoder layer; as well as Fully connected layer.

6. The method of claim 1, wherein the data for the one or more players comprises movements of one or more actors on a playing field received from a tracking device.

7. The method according to claim 1, further comprising: receiving updated data for the one or more players or teams from a tracking device; providing the update data to the transformer-based neural network; as well as One or more updated predictions for the sporting event are generated based on the updated data.

8. The method according to claim 1, further comprising: using a trigger processing step to access the data platform at set time intervals to determine when the sporting event occurs; as well as The second set of feature vectors is created by querying data from the data platform using a feature creator processing step.

9. A system for generating predictions for teams and players on each team associated with a sporting event, the system comprising: a non-transitory computer-readable medium configured to store processor-readable instructions; as well as a processor operably connected to the non-transitory computer-readable medium and configured to execute the instructions to perform operations comprising: receiving one or more top-down predictions related to the sporting event; providing the top-down predictions to a computing system as one or more top-down feature vectors; receiving, by the computing system, a second set of feature vectors comprising data for one or more players associated with one or more respective teams and data for one or more teams associated with a sporting event; inputting the one or more top-down feature vectors and the second set of feature vectors into a transformer-based neural network; and One or more predictions for the sporting event are generated using the transformer-based neural network based on the one or more top-down feature vectors and the second set of feature vectors.

10. The system of claim 9, wherein the one or more top-down predictions are one or more of neural network outputs, market information, or match context information.

11. The system of claim 9, wherein the top-down prediction comprises a first top-down prediction based on pre-game data and a second top-down prediction based on in-game data.

12. The system of claim 9, wherein the operations further comprise: The one or more predictions for the sporting event are caused to be displayed on a display device.

13. The system of claim 9, wherein the transformer-based neural network comprises: A set of embedding layers; Transformer encoder layer; as well as Fully connected layer.

14. The system of claim 9, wherein the data for the one or more players includes movements of one or more actors on the playing field received from a tracking device.

15. The system of claim 9, wherein the operations further comprise: receiving updated data for one or more players or teams from a tracking device; providing the update data to the transformer-based neural network; as well as One or more updated predictions for the sporting event are generated based on the updated data.

16. The system of claim 9, further comprising: using a trigger processing step to access the data platform at set time intervals to determine when the sporting event occurs; as well as A feature creator processing step is used to create the second set of feature vectors by querying data from the data platform.

17. A non-transitory computer-readable medium configured to store processor-readable instructions, wherein when executed by a processor, the instructions perform operations comprising: receiving one or more top-down predictions related to the sporting event; providing the top-down predictions to a computing system as one or more top-down feature vectors; receiving, by the computing system, a second set of feature vectors comprising data for one or more players associated with one or more respective teams and data for one or more teams associated with a sporting event; inputting the one or more top-down feature vectors and the second set of feature vectors into a transformer-based neural network; as well as One or more predictions for the sporting event are generated using the transformer-based neural network based on the one or more top-down feature vectors and the second set of feature vectors.

18. The non-transitory computer-readable medium of claim 17, wherein the one or more top-down predictions are one or more of neural network outputs, market information, or game context information.

19. The non-transitory computer-readable medium of claim 17, wherein the top-down predictions include a first top-down prediction based on pre-game data and a second top-down prediction based on in-game data.

20. The non-transitory computer-readable medium of claim 17, wherein the operations further comprise: The one or more predictions for the sporting event are caused to be displayed on a display device.