Systems and methods for athletic motion tracking using diffusion models
By combining geospatial and tagged event data with a Transformer-based diffusion model, we address the occlusion and error issues in broadcast tracking systems, generating high-fidelity player trajectories and enabling in-depth tactical analysis.
Patent Information
- Application Number
- CN202380082258.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-30
- Filing Date
- 2023-12-29
- Publication Date
- 2025-09-16
AI Technical Summary
Existing sports broadcast tracking systems cannot accurately predict player trajectories, especially in the presence of long-term occlusions and tracking errors, and cannot generate high-fidelity and complete multi-agent tracking data.
A Transformer-based diffusion model is adopted, combined with geospatial data and labeled event data of sports events, to generate high-fidelity player trajectories through event encoder and tracking decoder, and spatiotemporal axial attention and temporal convolution are used to handle occlusion and tracking errors, fusing multi-agent trajectories with semantic context.
It enables high-fidelity tracking of multiple actors in sports events, can interpolate occlusions and correct tracking errors, generate highly realistic player trajectories, and support in-depth tactical analysis and downstream data applications.
Smart Images

Figure CN120660093A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 477,932, filed December 30, 2022, the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0002] Various aspects of the present invention generally relate to machine learning for sports applications, and more particularly, various aspects relate to systems and methods for performing high-fidelity sports tracking based on Transformer-based models and diffusion models. Background Art
[0003] With the growing popularity of sports, there is a growing desire to make accurate, granular predictions about what will happen during a sporting event. For example, predicting the number of passes or shots a particular soccer player (e.g., Lionel Messi) will take in a given match (e.g., the World Cup Final) (both before and during the World Cup Final) may be of particular interest to members of the media, broadcasters (whether on the primary feed or on a second screen experience), and fantasy / gamified applications. Existing solutions are unable to accurately make such predictions. In particular, existing solutions may not be able to accurately make predictions about the trajectories of one or more players during a match. In particular, existing solutions may not be able to accurately predict trajectories due to, for example, long-term occlusions and tracking errors in broadcast footage. Therefore, new solutions are needed.
[0004] Unless otherwise indicated herein, the materials described in this section are not prior art to the claims in this application and are not admitted to be prior art or suggestions of prior art by inclusion in this section. Summary of the Invention
[0005] In some aspects, the technology herein relates to a method for tracking one or more individuals during a sporting event, the method comprising: receiving as input geospatial data for the sporting event; receiving as input tagged event data based on sports broadcast footage of the sporting event; performing multi-object tracking on one or more subjects of the received geospatial data to determine one or more vectors; inputting the tagged event data and the one or more vectors into a diffusion model; and determining one or more trajectory sequences for the one or more subjects using the diffusion model.
[0006] In some aspects, the technology described herein relates to a method wherein the tagged event data comprises a continuous stream of one or more major events throughout a sporting event, the major events comprising at least one of a pass, a shot, a steal, a foul, a turnover, a free throw, a goal, a score, or a substitution from the sporting event, and wherein the geospatial data comprises sports broadcast footage, in-field footage, global positioning system (GPS) data, near field communication (NFC), and / or radio frequency identification data (RFID).
[0007] In some aspects, the technology described herein relates to a method wherein one or more vectors include at least one of the following: two-dimensional coordinates of an actor on a field of a sporting event, the actor's location, the actor's team, an indicator indicating that the actor is a ball, or player visibility information.
[0008] In some aspects, the technology described herein relates to a method in which event data is represented as a two-dimensional space-time grid, with the grid representing a stack of events for each player.
[0009] In some aspects, the technology described herein relates to a method in which a diffusion model applies spatiotemporal axial attention to received event data and one or more vectors, where self-attention is applied to the temporal axis and the spatial axis, respectively.
[0010] In some aspects, the technology described herein relates to a method in which a diffusion model includes: an event encoder; and a track decoder, wherein the event encoder encodes labeled event data and the track decoder conditionally decodes a set of tracks.
[0011] In some aspects, the technology described herein relates to a method wherein an event encoder embeds event data, wherein embedding the event data further comprises: tokenizing the labeled event data using a linear projection; applying a sinusoidal position embedding to specify a temporal occurrence of the event data; processing the event data with a stacked encoder; and outputting the event embedding.
[0012] In some aspects, the techniques described herein relate to a method in which a tracking decoder uses attention to embed and fuse one or more vectors with an event embedding.
[0013] In some aspects, the techniques described herein relate to a method further comprising: a second tracking decoder; and a transposed temporal convolution configured to extend the track to its original temporal dimension.
[0014] In some aspects, the technology herein relates to a system for tracking one or more individuals during a sporting event, the system comprising: a non-transitory computer-readable medium configured to store processor-readable instructions; and a processor operably connected to the memory and configured to execute the instructions to perform operations, the operations comprising: receiving geospatial data for the sporting event as input; receiving tagged event data based on sports broadcast footage of the sporting event as input; performing multi-object tracking on one or more subjects of the received geospatial data to determine one or more vectors; inputting the tagged event data and the one or more vectors into a diffusion model; and determining one or more trajectory sequences for the one or more subjects using the diffusion model.
[0015] In some aspects, the technology described herein relates to a system wherein the tagged event data comprises a continuous stream of one or more major events throughout a sporting event, the major event comprising at least one of a pass, a shot, a steal, a foul, a turnover, a free throw, a goal, a score, or a substitution from the sporting event, and wherein the geospatial data comprises sports broadcast footage, in-field footage, global positioning system (GPS) data, near field communication (NFC), and / or radio frequency identification data (RFID).
[0016] In some aspects, the technology described herein relates to a system wherein one or more vectors include at least one of the following: two-dimensional coordinates of an actor on a field of a sporting event, the actor's location, the actor's team, an indicator indicating that the actor is a ball, or player visibility information.
[0017] In some aspects, the technology described herein relates to a system in which event data is represented as a two-dimensional space-time grid, with the grid representing a stack of events for each player.
[0018] In some aspects, the technology described herein relates to a system in which a diffusion model applies spatiotemporal axial attention to received event data and one or more vectors, wherein self-attention is applied to the temporal axis and the spatial axis, respectively.
[0019] In some aspects, the technology described herein relates to a system wherein a diffusion model includes: an event encoder; and a track decoder, wherein the event encoder encodes labeled event data and the track decoder conditionally decodes a set of tracks.
[0020] In some aspects, the technology described herein relates to a system wherein an event encoder embeds event data, wherein embedding the event data further comprises: tokenizing the labeled event data using linear projection; applying a sinusoidal position embedding to specify temporal occurrences of the event data; processing the event data with a stacked encoder; and outputting the event embedding.
[0021] In some aspects, the techniques described herein relate to a system in which a tracking decoder uses attention to embed and fuse one or more vectors with an event embedding.
[0022] In some aspects, the techniques described herein relate to a system further comprising: a second tracking decoder; and a transposed temporal convolution configured to extend the track to its original temporal dimension.
[0023] In some aspects, the technology herein relates to a non-transitory computer-readable medium configured to store processor-readable instructions for tracking one or more individuals during a sporting event, the instructions performing operations comprising: receiving as input geospatial data for the sporting event; receiving as input tagged event data based on sports broadcast footage of the sporting event; performing multi-object tracking on one or more subjects of the received geospatial data to determine one or more vectors; inputting the tagged event data and the one or more vectors into a diffusion model; and determining one or more trajectory sequences for the one or more subjects using the diffusion model.
[0024] In some aspects, the technology described herein relates to a non-transitory computer-readable medium, wherein the tagged event data comprises a continuous stream of one or more major events throughout a sporting event, the major events comprising at least one of a pass, a shot, a steal, a foul, a turnover, a free throw, a goal, a score, or a substitution from the sporting event, and wherein the geospatial data comprises sports broadcast footage, in-field footage, global positioning system (GPS) data, near field communication (NFC) data, and / or radio frequency identification data (RFID).
[0025] Additional objects and advantages of the disclosed aspects will be set forth in part in the following description and will become apparent from the description, or may be learned by practicing the disclosed aspects. The objects and advantages of the disclosed aspects will be realized and obtained by means of the elements and combinations particularly pointed out in the appended claims.
[0026] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed aspects, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate various exemplary embodiments and, together with the description, serve to explain the principles of the disclosed embodiments.
[0028] Figure 1 A block diagram of an exemplary tracking and analysis environment is depicted in accordance with one or more embodiments.
[0029] Figure 2Depicted is an exemplary block diagram of a system for generating a network of transformers for a player's trajectory in accordance with one or more embodiments.
[0030] Figure 3 Depicted are exemplary scenarios of a backdiffusion process performed by a system in accordance with one or more embodiments.
[0031] Figure 4 Depicted are example diagrams of player occlusion during an example game in accordance with one or more embodiments.
[0032] Figure 5A Depicted is an exemplary block diagram of a diffusion neural network in accordance with one or more embodiments.
[0033] Figure 5B Depicted is an example encoder for a diffusion neural network in accordance with one or more embodiments.
[0034] Figure 5C Depicted is another exemplary decoder for a diffusion neural network in accordance with one or more embodiments.
[0035] Figure 6A Depicted is a first exemplary visualization of predicted trajectories of a diffusion neural network along with input data to the network and ground truth data in accordance with one or more embodiments.
[0036] Figure 6B Depicted is a second exemplary visualization of a predicted trajectory of a diffusion neural network along with input data to the network and ground truth data in accordance with one or more embodiments.
[0037] Figure 6C Depicted is a third exemplary visualization of predicted trajectories of a diffusion neural network along with input data to the network and ground truth data in accordance with one or more embodiments.
[0038] Figure 7A and Figure 7B Depicted are example occlusions of players from an example basketball game in accordance with one or more embodiments.
[0039] Figure 7C Depicts Figure 7A and Figure 7B Legend for .
[0040] Figure 8 Depicted is an exemplary block diagram of a multi-agent trajectory system in accordance with one or more embodiments.
[0041] Figure 9 Depicted are exemplary compositing mask strategies in accordance with one or more embodiments.
[0042] Figure 10Depicted is an example visualization of a diffusion neural network's predicted trajectory for a basketball possession in accordance with one or more embodiments.
[0043] Figure 11 Depicted are exemplary results of using a diffusion neural network to track one or more players in accordance with one or more embodiments.
[0044] Figure 12 Depicted is an exemplary flow diagram of a method for tracking one or more individuals during a sporting event in accordance with one or more embodiments.
[0045] Figure 13 Depicted is a flow diagram for training a machine learning model according to one aspect.
[0046] Figure 14 An example of a computing device according to one aspect is depicted.
[0047] It is worth noting that, for simplicity and clarity of illustration, certain aspects of the drawings depict the general configurations of various embodiments. Descriptions and details of well-known features and techniques may be omitted to avoid unnecessarily obscuring other features. Elements in the figures are not necessarily drawn to scale; the dimensions of some features may be exaggerated relative to other features to improve understanding of the example embodiments. DETAILED DESCRIPTION
[0048] Various aspects of the present invention generally relate to machine learning for sports applications, and more particularly, various aspects relate to systems and methods for transformer networks for generating player trajectories.
[0049] According to embodiments disclosed herein, a guided diffusion model can receive broadcast tracking data and event data for a sporting event as input. The guided diffusion model can generate high-fidelity tracking data based on the received input data. The diffusion model can include an event encoder and a track decoder that can embed and fuse the received event and broadcast tracking data. The output embedding can be fed into a score-based diffusion model to generate trajectories for one or more players in the sporting event.
[0050] As used herein, a "machine learning model" generally comprises instructions, data, and / or a model configured to receive an input and apply one or more of weights, biases, classifications, or analyses to the input to generate an output. For example, the output may include a classification of the input, an analysis based on the input, a design, process, prediction, or recommendation associated with the input, or any other suitable type of output. Machine learning models are typically trained using training data, e.g., empirical data and / or samples of input data, which is fed into the model to establish, adjust, or modify one or more aspects of the model, e.g., weights, biases, criteria for forming classifications or clusters, etc. Various aspects of a machine learning model may operate on the input linearly, in parallel, via a network (e.g., a neural network), or via any suitable configuration.
[0051] Execution of the machine learning model may include deploying one or more machine learning techniques, such as linear regression, logistic regression, random forest, gradient boosting machine (GBM), deep learning, and / or deep neural networks. Supervised and / or unsupervised training may be employed. For example, supervised learning may include providing training data and labels corresponding to the training data, for example, as ground truth. Unsupervised methods may include clustering, classification, and the like. K-means clustering or K-nearest neighbors may also be used, which may be supervised or unsupervised. A combination of K-nearest neighbors and unsupervised clustering techniques may also be used. Any suitable type of training may be used, for example, random, gradient boosting, random seeding, recursive, round-robin, or batch-based, and the like.
[0052] While several of the examples herein relate to certain types of machine learning, it should be understood that the techniques according to the present invention may be adapted for any suitable type of machine learning. It should also be understood that the above examples are merely illustrative. The techniques and methods of the present invention may be adapted for any suitable activity.
[0053] While football and football-related aspects (e.g., predicted trajectories of one or more players during a game) are described in this aspect as illustrative examples, this aspect is not limited to such examples. For example, this aspect can be implemented for other sports or activities, such as football, basketball, baseball, hockey, cricket, rugby, tennis, etc.
[0054] The uniform, complete and scalable digitization of broadcast video of sports into tracking data (positions of players and balls over time) can be considered a landmark challenge in computer vision for sports. Conventional vision-centric systems can track players based in part on broadcast footage, however, the visual constraints of broadcast footage cause the tracking data to suffer from frequent long-term occlusions and / or tracking errors. The systems and methods described herein address this problem by using a diffusion-based multi-agent trajectory generation model that jointly interpolates occluded behaviors and denoises tracking errors. The multi-agent trajectory generation model can fuse large amounts of event data (e.g., a temporal semantic data stream of a football) with the original broadcast tracking data. The multi-agent trajectory generation model described herein can generate highly realistic behaviors and can allow the injection of domain-specific constraints via guidance.
[0055] As sports become more commercialized and popularized, tracking data (e.g., player and ball position over time) can facilitate deeper analysis of individual and collective performance. Traditionally, sports tracking data has been captured by polyphasic on-field tracking systems that continuously monitor all active players and their positions to generate complete tracking. However, the cost of installing and operating these systems has limited their adoption to only the most high-profile leagues.
[0056] In a sporting event, complete tracking data for all players while an object (e.g., a ball) is in play may be critical for tactical analysis of the players and teams. Furthermore, given that high leverage events (e.g., goal scoring opportunities) are rare in sporting events, having complete tracking data for each of these events may be necessary for downstream analysis. Conventional broadcast tracking systems may be vision-centric, and conventional tracking systems may rely primarily on visual perception. However, using such visual perception, it is not possible to generate complete tracking data from broadcast footage (e.g., using vision alone). A disadvantage of conventional vision-centric systems may be caused by the limited receptive field of the camera, where important context leading to high probability goal scoring opportunities may be missed. This example may represent the impact of occlusion (e.g., where a subset of the actors cannot be intuitively perceived) on a broadcast tracking system. Occlusion may have multiple sources, which are caused by the limited monocular receptive field of the broadcast camera, close-ups, replays, and changing native camera angles. During a sporting event such as football, most occlusions may be short-term (≤10 seconds, however, many occlusions last much longer, with some lasting for more than 60 seconds (e.g., as in Figure 4as shown). Additional challenges in conventional systems can be caused by the system having to be calibrated based on a moving camera, the homogeneous clothing of players within a team, and the small size and fast movement of the ball, which together result in frequent object detection errors and misidentification of actors in conventional systems. These occlusions and tracking errors reduce the coverage and quality of the original broadcast tracking and may need to be addressed to generate accurate and complete remote tracking data.
[0057] The systems and techniques described herein can utilize generative modeling to jointly impute and denoise broadcast tracking data. The systems and techniques described herein can be applicable to any sport. For example, the systems and techniques described herein can be applicable to sports with game players having an inherent structure, such as sports in which players abide by short-term and long-term policies within an environment based on opposing teams and / or sports in which rigid single-actor and multi-actor spatio-temporal behaviors recur within and across games. The systems and techniques described herein can be configured to learn such structures, and multi-actor trajectory generation methods can be used to impute the behaviors of occluded actors and correct erroneous tracking data.
[0058] Disadvantages of conventional techniques include challenges within multi-actor trajectory modeling. First, for example, most trajectory modeling tasks can be done on independently segmented short-term windows (about <l0 seconds) where trajectory start and / or end anchors can be observed. Since the duration of a soccer game is long (e.g., 90+ minutes) and the frequency of long-term occlusions is high (more than about 10 seconds), this framework is not suitable for broadcast tracking. Second, for example, the systems described herein can include trajectories that can jointly predict, impute, and denoise, rather than performing prediction (e.g., predicting future trajectories) or imputation (e.g., generating occluded behaviors).
[0059] The systems and techniques described herein may include learning the coarse and / or long-term latent structure of a given sport. The systems and techniques described herein may include a multimodal infrastructure. For example, a system or one or more components may be transformer-based, in which partially observed multi-agent trajectories (e.g., based on raw broadcast tracking data) are fused with a coarse stream of semantic information (e.g., event data). Event data may refer to a continuous stream of all events or all major events throughout a given sporting game (e.g., passes, shots, steals, fouls, turnovers, free throws, goals, scores, substitutions, etc.). Event data may provide the necessary signals for reconstructing portions of the game not covered by the raw broadcast tracking data. The systems described herein may include joint modeling of event and tracking data. This may represent a paradigm shift away from conventional systems that perform unimodal processing of sports trajectories. Given the long duration of subject occlusion during a game, a key benefit of the systems and techniques described herein may be the ability of the system to ingest large temporal contexts. For example, the system may adapt the use of spatiotemporal axial attention and temporal convolution to jointly process tracking data and event context (e.g., three minutes of tracking data with ten minutes of event context) at once.
[0060] The system described herein can, for example, incorporate a conditional diffusion model capable of generating multi-agent tracking data. The conditional diffusion model can be configured to learn a conditional probability density function over the multi-agent trajectories, which can enable the synthesis of highly realistic multi-agent tracking data.
[0061] The system described in this paper can be configured to fuse a large amount of tracking and semantic context. By leveraging a diffusion model, realistic trajectories can be generated that inject domain-specific constraints via guidance.
[0062] To improve upon existing methods, one or more techniques described herein utilize transformer-based neural networks to predict player positions. For example, using transformer-based neural networks, the current system can generate or simulate the remainder of a given game at the player trajectory level. For example, the system can generate player sequences based on the player's underlying representation. One benefit of this approach may be that it can be combined with a tracking system (e.g., computer vision-based player and player tracking) or used together as an auxiliary aid in the data collection process. This approach can also be used to highlight potentially erroneous data points for evaluation (e.g., evaluation based on a human operator or evaluation based on an automated system).
[0063] Figure 11 is a block diagram illustrating a computing environment 100 according to example aspects of the disclosed subject matter. Environment 100 includes a tracking system 102, a computing system 104, and a client device 108 connected via a network 105. In the depicted example, tracking system 102 obtains various measurements of a game and transmits the measurements across network 105 to computing system 104, where the measurements can be used in conjunction with one or more machine learning models. In one example, one or more machine learning models described herein can be configured to receive broadcast tracking data and event data as input and perform conditional guided diffusion to generate trajectories for one or more players in a sporting event.
[0064] Tracking system 102 can be in communication with venue 106 and / or can be located within, adjacent to, or near venue 106. Non-limiting examples of venue 106 include a stadium, a field, a venue, and a court. Venue 106 includes actors 112A-N (players). Tracking system 102 can be configured to record the movements and actions of actors 112A-N on the playing field, as well as one or more related other objects (e.g., a ball, a referee, etc.). Although environment 100 generally depicts actors 112A-N as players, it should be understood that, according to certain embodiments, actors 112A-N can correspond to players, officials, coaches, objects, markers, etc.
[0065] In some aspects, tracking system 102 can be an optical-based system, such as one using camera 103. While one camera is depicted, additional cameras are possible. For example, a system consisting of six fixed, calibrated cameras can be used that projects the three-dimensional positions of the players and ball onto a two-dimensional overhead view of the court.
[0066] In another example, a mix of fixed and non-fixed cameras can be used to capture the movement of all actors 112A-N on the playing field, as well as the movement of one or more related objects. Utilizing such a tracking system (e.g., tracking system 102), many different camera views of the field can be generated (e.g., high sideline view, free throw line view, huddle view, scrimmage view, end zone view, etc.). In some aspects, tracking system 102 corresponds to or utilizes a broadcast feed of a given game. In such aspects, each frame of the broadcast feed can be stored in a game file.
[0067] The tracking system 102 can be configured to communicate with the computing system 104 via the network 105. The computing system 104 can be configured to manage and analyze the data captured by the tracking system 102. The computing system 104 can include a network client application server 114, a pre-processing agent 116 (e.g., a processor and / or pre-processor), a data storage device 118, and a third-party application programming interface (API) 138. Figure 1 One example of computing system 104 is depicted.
[0068] The pre-processing agent 116 can be configured to process data retrieved from the data storage device 118 or the tracking system 102 before inputting it into the predictor 126. The pre-processing agent 116, the predictor and / or the predictive model analysis engine 122 can be composed of one or more software modules. One or more software modules can be a collection of codes or instructions stored on a medium (e.g., a memory of the organizational computing system 104), which represents a series of machine instructions (e.g., program code) that implement one or more algorithmic steps. Such machine instructions can be actual computer code that the processor of the organizational computing system 104 parses to implement the instructions, or alternatively, can be a higher-level instruction encoding that is parsed to obtain the actual computer code. One or more software modules can also include one or more hardware components. One or more aspects of the example algorithm can be performed by the hardware component (e.g., circuit) itself, rather than as a result of the instructions.
[0069] The data storage device 118 can be configured to store different types of data. In one example, the data storage device 118 can store raw tracking data received from the tracking system 102. The data storage device 118 can include historical game data, real-time data, features, and / or predictions. The historical game data can include historical team and player data for one or more sporting events. The real-time data can include data received from the tracking system 102, for example, in real time or near real time. The game data can include broadcast data or content related to a game (e.g., a match, a competition, a round, etc.) and / or can include tracking data generated by the tracking system 102 or in response to data generated by the tracking system 102. The data storage device 118 can be configured to store features (e.g., feature vectors) generated for a particular sporting event, which include player, team, and game features. The data storage device 118 can also be configured to store event data for a sporting event.
[0070] According to aspects disclosed herein, the data storage device 118 can receive and / or store competitive files. Competitive files can include one or more competitive data types. Competitive data types can include, but are not limited to, positional data (e.g., player positions, object positions, etc.), change data (e.g., position changes, player changes, object changes, etc.), trend data (e.g., player trends, position trends, object trends, team trends, etc.), match data, etc. The competitive file can be a single competitive file, or it can be segmented (e.g., grouped by one or more data types, grouped by one or more players, grouped by one or more teams, etc.). The pre-processing agent 116 and / or the data storage device 118 can operate (e.g., using applicable code) to receive tracking data in a first format, store the competitive file in a second format, and / or output the competitive data in a third format (e.g., to the predictor 126). For example, the pre-processing agent 116 can receive the intended destination of the competitive data (or data generally stored in the data storage device 118) and format the data into a format acceptable to the intended destination.
[0071] The predictor 126 includes one or more machine learning models 128AN. Examples include a Transformer neural network, which may include one or more encoders and / or decoders. The Transformer can be configured to generate tracking data based on broadcast tracking data and event data. The Transformer can also be configured to generate predictions for the trajectories of one or more players during a game. The Transformer neural network can be configured to fuse partially observed multi-agent trajectories (e.g., raw broadcast tracking data) with a coarse semantic information stream based on sports (e.g., event data). The Transformer neural network can also include a diffusion model capable of generating multi-agent tracking data. The Transformer-based neural network can be configured to generate or simulate the remainder of a given game at the player trajectory level. For example, instead of generating trajectories for a possession, the Transformer network can be configured to generate trajectories for multiple possessions and even for the remainder of the sporting event. In addition, the Transformer network can also be configured to generate event data for the game. In this way, the Transformer network can be used to generate commentary of the game via text / speech or 3D models of player behavior.
[0072] Client device 108 can communicate with computing system 104 via network 105. Client device 108 can be operated by a user. For example, client device 108 can be a mobile device, a tablet computer, a desktop computer, or any computing system with the capabilities described herein. Users can include, but are not limited to, individuals such as, for example, subscribers, clients, potential clients, or customers of an entity associated with computing system 104, such as individuals who have obtained, will obtain, or may obtain products, services, or consulting from an entity associated with computing system 104.
[0073] The client device 108 may include one or more applications 109. The applications 109 may represent a web browser that allows access to websites or may be stand-alone applications. The client device 108 may access the applications 109 to access one or more functions of the computing system 104. The client device 108 may communicate over the network 105 to, for example, request a web page from a web client application server 114 of the computing system 104. For example, the client device 108 may be configured to execute the applications 109 to access content managed by the web client application server 114. The content displayed to the client device 108 may be transmitted from the web client application server 114 to the client device 108 and then processed by the applications 109 for display via a graphical user interface (GUI) of the client device 108.
[0074] The client device may include a display 110. Examples of the display 110 include, but are not limited to, a computer monitor, a light emitting diode (LED) display, etc. Output or visualization generated by the application 109 (e.g., a GUI) may be displayed on or using the display 110.
[0075] The functionality of the subcomponents shown in computing system 104 may be implemented in hardware, software, or some combination thereof. For example, a software component may be a collection of code or instructions stored on a medium, such as a non-transitory computer-readable medium (e.g., a memory of computing system 104), that represents a series of machine instructions (e.g., program code) that implement one or more method operations. Such machine instructions may be actual computer code that a processor of computing system 104 parses to implement the instructions, or alternatively, may be a higher-level instruction encoding that is parsed to obtain the actual computer code. One or more software modules may also include one or more hardware components. Examples of components include processors, controllers, signal processors, neural network processors, and the like.
[0076] The network 105 can be of any suitable type, including a separate connection via the Internet, such as a cellular or Wi-Fi network. In some aspects, the network 105 can use a direct connection to connect terminals, services, and mobile devices, such as radio frequency identification (RFID), near field communication (NFC), Bluetooth, or other similar communication methods. TM , Bluetooth Low Energy TM (BLE), Wi-Fi TM 、ZigBee TM , Ambient Backscatter Communication (ABC) protocol, USB, WAN, or LAN. Because the information transmitted may be personal or confidential, security concerns may require that one or more of these types of connections be encrypted or otherwise protected. However, in some aspects, the information transmitted may not be so personal, and therefore, a network connection may be chosen for convenience rather than security.
[0077] The network 105 may include any type of computer networking arrangement for exchanging data or information. For example, the network 105 may be the Internet, a private data network, a virtual private network using a public network, and / or other suitable connections that enable components in the computing environment 100 to send and receive information between components of the environment 100.
[0078] Diffusion-based multi-agent trajectory generation model
[0079] The system described in this paper can use Transformer-based neural networks to fuse multi-agent trajectories with semantic data of sports, even streaming data. The system can implement a score-based diffusion framework as described below.
[0080] Figure 2 An exemplary block diagram depicts a system of a Transformer network (e.g., diffuser 210) for generating trajectories for a player (e.g., motion tracking information 212) in accordance with one or more embodiments. For example, the system can provide conditional guided diffusion (e.g., provided by diffuser 210) to generate one or more trajectories for a player (e.g., motion tracking information 212) based on a limited field of view (e.g., based on video input data 202 including occlusions).
[0081] For example, system 200 may include video input data 202 for a sports broadcast. For example, video input data 202 may have a limited receptive field. For example, occlusions may occur when a subset of players cannot be visually displayed on video input data 202. These occlusions may occur due to various sources, including the limited monocular receptive field of a broadcast camera, close-ups, replays, and changing camera angles. For example, video input data (e.g., broadcast data) may be a subset of geospatial data. Geospatial data may be any content, information, or feed that allows for tracking of one or more objects, as discussed further herein. For example, geospatial data may refer to broadcast footage, in-field footage, global positioning satellite (GPS) data, radio frequency identification data (RFID), near-field communication (NFC), triangulation data, and the like. Geospatial data and subsequently processed geospatial data (e.g., processed by video analysis system 204) may be received as input, for example, by diffuser 210 as described herein. Video input data may refer to broadcast footage or in-field computer vision system output, which may be or include, for example, raw video content. For example, an on-field computer vision system can record video footage of the entire playing field during the entire game.
[0082] The video input data 202 may be input into a video analysis system 204, for example. For example, the video analysis system 204 may perform one or more functions. First, the video analysis system 204 may implement one or more computer vision algorithms to determine broadcast tracking data 206. For example, the broadcast tracking data 206 may be output as a multi-agent trajectory for each of the players in the game. The one or more computer vision algorithms may be configured to (1) detect players in the sporting event; (2) group the detected players into one or more teams; (3) identify "logical identities" of the identified players so as to maintain the identities and track the players over time; (4) identify a ground plane for the sporting event; and / or (5) identify an assigned number for each player on the field. The one or more computer vision algorithms may also provide tracking of the identified players over time. For example, the broadcast tracking data 206 may be stored in a JavaScript Object Notation (JSON) file. The broadcast tracking data 206 may, for example, include two-dimensional tracking of one or more players in the game, the players' respective teams, and the players' respective identification numbers (e.g., the players' respective jersey numbers).
[0083] Broadcast tracking data 206 may be based on publicly available broadcast data and / or footage associated with the sporting event generated or broadcast at least in part using one or more cameras or camera systems. Figure 1The tracking system 102 generates broadcast tracking data 206. For example, the broadcast tracking data may include a tracking flow determined using a computer algorithm applied to a broadcast feed. The tracking flow may represent the movement of an agent (e.g., a player, other individual, object, etc.). The broadcast tracking flow may be represented as b∈R T×E×D b, where each observation contains the 2D coordinates of the agent, the agent type (i.e., outfield player, ball, goalkeeper, etc.), team affiliation, and indicators of whether the ball is in play and whether the agent is visible.
[0084] A second function of the video analysis system 204 may be to determine event data 208. Event data 208 may refer to a continuous stream of all major events throughout a given game (e.g., passes, shots, tackles, fouls, turnovers, free throws, goals, points, substitutions, etc.). Event data 208 may provide the necessary signals for reconstructing portions of a game not covered by the original broadcast tracking data. Event data 208 may be automatically detected by the computing system or input from a user reviewing the video input data 202, for example. For example, event data 208 may be input by a user viewing the video input data 202 (e.g., a broadcast feed). Event data 208 may be unified into a two-dimensional space-time grid. This may be done by stacking (with padding) the events for each player, forming an event stream s∈R L ×E×D s performed, where L is the maximum number of events performed by a single agent within a specified time range, and each event includes the timestamp of the event, 2D coordinates, agent type, and event category (e.g., pass). Event data 208 may be referred to herein as "labeled event data."
[0085] The determined broadcast tracking data 206 and event data 208 may be input to, for example, a diffuser 210. The diffuser 210 may be, for example, Figure 5A The diffusion neural network 500 is depicted and described in more detail below. The diffuser 210 can include a Transformer-based neural network. The diffuser 210 can include an encoder (e.g., for operations at least partially related to event data 208) and one or more tracking decoders (e.g., for fusing event encoder output with broadcast tracking data 206), as further discussed herein. The diffuser 210 can generate trajectories and output them as sports tracking information 212. These can be output as vectors for further analysis and / or presentation.
[0086] As discussed above, processed geospatial data can be received as input by diffuser 210 (e.g., as an alternative to or in addition to video data 202). For example, the geospatial data can be based on wearable technology worn by one or more actors on the field. For example, GPS, RFID, and / or NFC data can be received by system 200. GPS, RFID, and / or NFC data can correspond to location data tracked using GPS sensors, satellite tracking, proximity sensors, tags, and the like. Such location data can provide useful context to system 200 when sensor information (e.g., broadcast data, in-field sensor information, etc.) is noisy or missing. Geospatial data can also be based on an in-field computer vision system. In-field computer vision data can be used to denoise the input (or incorporated into event data 208). Event data 208 can be received and / or transformed into the reference frame being tracked. For example, event data can have a reference frame originating from (0, 0, 100, 100), whereby the field coordinates can be (0, 0, 106, 68). Therefore, any suitable scaling technique, such as translation, shifting, normalization, etc., may be used to transform the event data into a (0, 0, 100, 100) reference frame.
[0087] In addition, the diffuser 210 can be configured to receive labeled inputs, such as human labeled inputs (e.g., only event data 208). The system 200 can be configured to interpolate the positions of one or more objects based on the event data 208 (e.g., based on only one event data 208). Such labeled input can, for example, be received in the form of text and can be converted into tracking data based on analysis of the text and / or based on providing the text to a machine learning model that is trained to output tracking data based on labeled text input. In another example, the system can be configured to interpolate the event data 208 based on one or more inputs discussed herein. For example, frames regarding the time interval at which an event occurred.
[0088] Figure 3 An exemplary scenario 300 is depicted of a backdiffusion process performed by system 200 according to one or more embodiments. Figure 3 The sampling process and how the data is denoised are depicted. Each image (eg, image 302 , image 304 , image 306 , image 308 ) depicts the steps of iterative denoising performed by the diffuser 210 . Figure 3The application of a back-diffusion process using, for example, diffuser 210, which can be a conditional diffusion model capable of generating multi-agent tracking data, is depicted. Diffuser 210 can learn a conditional probability density function over the multi-agent trajectories, thereby enabling the synthesis of highly realistic multi-agent tracking data. The iterative application of this function is illustrated sequentially via images 302, 304, 306, and 308.
[0089] Figure 4 An exemplary graph 400 depicts player occlusions during an exemplary game according to one or more embodiments. Graph 400 depicts the number of occlusions that occurred during an exemplary sporting event broadcast. Graph 400 depicts that most occlusions of players were short-term (e.g., less than 10 seconds), however, many occlusions lasted for longer periods of time, including periods exceeding 60 seconds. Despite the presence of occlusions as depicted in graph 400, the systems and methods described herein determine tracking data (e.g., including tracks) for a given event based in part on the broadcast footage.
[0090] diffuser
[0091] A denoised diffusion model can be implemented by the system 200 described herein. Such a diffusion model can consider a family of distributions p(x,σ) where Gaussian noise with standard deviation σ is added to the distributions with standard deviation σ. 数据 The data distribution p 数据 (x). The Gaussian noise standard deviation can be maximized (i.e., σ max ), such a perturbed data distribution may be almost indistinguishable from pure Gaussian noise. Therefore, samples from this data distribution can be obtained by max ,...,σ N-2 ,σ N-1 On x0~N(0,σ 2 max I) Iteratively perform denoising so that xi ~ p(xi,σi) is generated. The fractional diffusion model can formulate this reverse diffusion process as an ordinary differential equation (ODE), where the derivative of the noise sample x is given by:
[0092] in, Given a score function, σ(t) is the noise level at diffusion step t, and is the time derivative of σ. The score function may be a vector field that gives the direction in which the probability density function grows fastest, from which the underlying probability density function can be inferred. The score function of the probability distribution can be obtained by training a conditional denoising model D parameterized by θ. θ (x,σ,c) is obtained to minimize the L2 reconstruction loss between the perturbed data samples and the original data samples,
[0093] Where q represents the distribution of σ during training, and y = x + n. Following this definition, the score is given by:
[0094] The training and preprocessing of the diffuser model used in this article can be implemented. Such models (e.g., deep models) can be learned most efficiently when their inputs and outputs are scaled to have unit variance. In addition, it is easier to predict the noise level n when the value of σ is low, and it is easier to predict the clean original signal x when the value of σ is high. Therefore, the diffuser described in this article (e.g., diffuser 210) can add preprocessing terms to scale the variance of the model input, and add jump connections to enable the model to adaptively predict the noise level or clean signal for different levels of σ, rather than directly returning the original output of the denoiser neural network. The denoiser can be written as: D θ (y;σ,c)=c skip (σ)y+c out (σ)F θ (c input (σ)y;c noise (σ),c) (4)
[0095] Make F θ is the output of the original neural network, c input The variance of the modulated perturbation trajectory, c noise Variance of modulation noise, c out The variance of the modulated output, and c skip Modulated skip connections. To normalize the loss within the range of σ, the reconstruction loss of each sample is calculated as λ(σ) = 1 / c 2 c can represent the raw input of the neural network and can be assumed to be modulated.
[0096] The diffuser described herein can apply constrained sampling. The diffusion model described herein (e.g., diffuser 210) can learn a conditional score function of the probability distribution of a set of multi-agent trajectories However, it is usually preferable to sample from the joint score function:
[0097] The second term represents the constrained gradient score of the manifold q on y. This constrained manifold can represent any loss function: L:R T×E×2 →R, which can be differentiated with respect to y. Scaling by the hyperparameter α, the constrained gradient score can be calculated as follows:
[0098] Using the ODF dynamics described in equation (1) above, sampling from the diffuser 210 may be performed using, for example, approximately 128 inference steps of a Henu sampler.
[0099] To prepare (e.g., train) and / or validate the diffuser 210, the diffuser 210 is provided with access to multiple streams of spatiotemporal data, such as video input data 202 (e.g., including broadcast tracking data 206 and / or event data 208), and can be provided with intra-field tracking data. Such streams can be represented as a spatiotemporal grid consisting of a temporal dimension T that specifies the length of the trajectory, a spatial dimension that represents the number of agents (e.g., of size E = 23), followed by a feature dimension. The perturbed intra-field trajectory can be written as y∈R T×E×2 , where each observation specifies the perturbed 2D position of the agent. Similarly, the broadcast tracking flow can be represented as b∈R T×E×Db , where each observation contains the 2D coordinates of the agent, the agent type (i.e., outfield player, ball, goalie), team affiliation, and / or an indicator of whether the ball is in play and whether the agent is visible. Invisible observations can have the 2D coordinates of the agent reset to zero. While event data can typically be represented as a 1D time stream, the data stream of event data 208 is represented as a 2D space-time grid. This can be done by stacking (with padding) the events for each agent, forming an event stream s∈R L×E×D s implementation, where L is the maximum number of events performed by a single agent within a specified time range, and each event includes the timestamp of the event, 2D coordinates, agent type, and event category (e.g., pass).
[0100] The diffuser 210 can apply spatiotemporal axial attention. The diffuser 210 can process modalities in a way that maintains their underlying spatiotemporal structure. While spatiotemporal data has a clear temporal total ordering (i.e., in chronological order), there is no such natural ordering of actors in space. In football, because there are two teams, each with 10 outfield players with no natural ordering, the actor index may have (10!) 2 To avoid combinatorial growth in complexity, the spatial dimensions of the space-time grid can be treated in a permutation-equivariant manner. That is, for example, for each permutation of the agent index, the following equation holds:
[0101] Among them, y p and c pcan represent the permutation of agent indices used to perturb the intra-field tracking and context vectors, respectively.
[0102] This property can be obtained using spatiotemporal axial attention, where self-attention is applied across the temporal and spatial axes respectively (e.g., as Figure 5A Using this scheme, the motion of individual agents can be learned via temporal attention, while the dynamics of clusters can be learned via spatial attention, without applying artificial ordering to the agents. Another benefit of axial attention may be its computational efficiency. Standard self-attention may have quadratic performance with respect to sequence length, and thus joint attention across spatial and temporal axes has O(T 2 ·E 2 ). In the case where the sequence length T dominates the number of actors E, the axial attention alone has O(T 2 )+O(E 2 )=O(T 2 ) complexity. This efficiency improvement in the diffuser 210 can allow processing of multi-agent trajectories that are much larger than conventional system lengths.
[0103] Figure 5A An exemplary block diagram of a diffusion neural network 500 is depicted in accordance with one or more embodiments. The diffusion neural network 500 may be Figure 2 An exemplary embodiment of a diffuser 210 is shown.
[0104] The diffusion neural network 500 may be configured to embed and fuse event data (e.g., event data 208) and broadcast trace context c (e.g., broadcast trace data 206). The diffusion neural network 500 may include: Figure 5B The event encoder 508 and the Figure 5C Depicted is a trace decoder 514. Additionally, the diffusion neural network may include a module 502 comprising an event encoder 508 and a trace decoder 514. The output of module 502 may be input to a score-based diffusion model, as described below.
[0105] The low temporal dimensionality of event data relative to tracking data may mean that encoding these modalities separately increases the amount of event data that can be encoded at once. The added context can improve the accuracy of denoised trajectories.
[0106] The diffusion neural network 500 may include an event encoder 508 . Figure 5B An exemplary encoder 508 of the diffusion neural network 500 is depicted in accordance with one or more embodiments. The event stream 504 may first be subjected to temporal convolution 506 as described above and then fed to the event encoder 508. The event encoder 508 may include an attention-based architecture that embeds the event stream s∈RL×E×DS . The event stream can first be tokenized using a linear projection before adding sinusoidal position embeddings to specify the temporal occurrence of events. Since the event stream consists of stacked events for each agent over a time range, the position embeddings can be based on the time step of each event, rather than a time index. The event tokens can then be processed by a stacked event encoder 508, which includes a time axis attention module 526 that establishes agent weights for the events, followed by a transformer encoder. The transformer encoder can include a self-attention module 528, a first residual connection and layer normalization module 530, a feedforward module 532, and a second residual connection and layer normalization module 534. The self-attention module 528 can perform tasks related to self-attention. Due to the low temporal dimensionality of the event stream relative to the tracking stream, the self-attention module 528 may not have a significant impact on the computational efficiency of the event encoder. The transformer encoder can include a first residual connection and layer normalization module 530, which includes residual connections and layer normalization. The Transformer encoder may further include a feedforward module 532 that passes the output through a series of linear projections followed by a nonlinear activation function. The Transformer encoder may include a second residual connection and layer normalization module 534 that includes residual connections and layer normalization. The stacked event encoder produces an event embedding Z with dimension d s ∈R L×E×d This may be the output of the event encoder 508. This output may be forwarded to the trace decoder 514.
[0107] The diffusion neural network 500 may include a tracking decoder 514 . Figure 5C An exemplary decoder 514 of the diffusion neural network 500 is depicted in accordance with one or more embodiments. The broadcast trace data 510 (e.g., the broadcast trace data 206) may first have a temporal convolution 518 applied thereto. The corresponding output and the event embeddings emitted from the event encoder 508 may be input to the trace decoder 514. The trace decoder 514 may utilize attention to embed and fuse the broadcast trace stream b∈R T×E×Db Embed Zs with events.
[0108] The combination of the high temporal dimensionality of broadcast tracking and the quadratic explosion of self-attention with respect to sequence length can traditionally limit the length of multi-agent trajectories that can be processed by transformers. This can be addressed by the system described in this paper, where a large amount of temporal context can be exploited to denoise long-term occlusions and tracking errors. For example, the tracking decoder 514 can compress the broadcast tracking stream b by applying temporal convolutions with kernel size and stride k, dimension d, and zero padding. conv ∈R(T / / k)×E×d By increasing the amount of trajectory context used, the output accuracy of trajectory generation can be improved.
[0109] After this compression, a sinusoidal position embedding can be added to specify the temporal ordering of the convolved trajectory tokens. The tokens can then be processed by a temporal attention operation 536, followed by a spatial attention 538. Next, the embedded tokens are temporally cross-attended with the event encoding (e.g., each agent can only attend to its own event tokens). This may occur in a temporal x attention operation 540. This allows the agents to be modeled jointly, avoiding egocentric representations as has been done before. The final modules within the tracking decoder are the Transformer standard normalization 544 and the feed-forward layer 542. Module 502 (e.g., via the tracking decoder 514) may return Z b ∈R (T / / k)×E×d , which represents the joint encoding of events and broadcast trace streams for each actor. This output may be, for example, the output of module 502.
[0110] The diffusion neural network 500 cannot directly use the embedding output by module 502 to deterministically predict the behavior of the player, but can generate trajectories through a score-based diffusion model. The score-based diffusion model can learn the conditional joint probability distribution over the set of multi-agent trajectories, thereby generating more realistic and controllable behaviors than the model trained to regress the positions of the agents independently. Specifically, the trained / activated version of the diffuser 210 can represent the denoising diffusion neural network F θ (y,σ,c). This network can form the parameterized part of the denoising function shown in equation (4) and is trained with the objective in equation (2). conv ∈R (T / / k)×E×d Before time compression, Figure 5A The perturbation trajectory 516y∈R T×E×2 The noise level c can be modulated first with, for example, noise (σ) is concatenated with the Fourier embedding of σ. As in the architecture of module 502, this can enable processing of larger windows of trajectory information. These compressed trajectory tokens can be positionally encoded using a sinusoidal position embedding. The second tracking decoder 520 can be used to fuse the context embedding Z b and the labeled perturbation trajectory y conv Finally, to extend the trajectories to their original time dimension, a transposed temporal convolution 522 may be applied. The output of this process may be a denoised trajectory 524.
[0111] At inference time (e.g., when using a trained diffuser 210 or diffusion neural network 500), the system can employ constrained trajectory sampling to allow for greater control over the generated behavior. The diffusion neural network 500 can leverage this control to enforce physical constraints specific to the sport. Events (e.g., event stream 504) can be labeled at 1-second intervals by any automated system or human, and therefore, pass events may frequently be out of phase with the tracking, which encourages the ball and the passing player to be close together when the pass occurs. This guidance term is intended to reduce the temporal misalignment between events and the tracking stream. Formally, L pass can be defined as follows:
[0112] in, includes the time step t and the agent index of each pass event, b is the spatial index of the ball, and y ^ =D θ (y;σ,c) is the denoised trace, and ∈ is a small constant added to maintain numerical stability.
[0113] Evaluation of diffuse neural networks
[0114] The diffusion neural network 500 can be evaluated to determine accuracy. In order to evaluate the coarse and granular authenticity of the generated trajectory, three indicators can be used. Coarse behavior can be evaluated using the average displacement error (ADE): the average distance (m) between the generated position and the true position of the behavior subject. This indicator is reported as the total average across all trajectories for all trajectory segments obscured by short-term occlusion (STO) (≤10 seconds) and long-term occlusion (LT0) (>10 seconds). In order to quantify the granular authenticity of the generated trajectory, two forms of trajectory generation failure can be defined. First, the speed failure rate can be defined as the average proportion of players who exhibit an instantaneous speed above a threshold of 12m / s (defined as the maximum human sprint speed) within a three-minute window. Second, pass failure may report the percentage of passes that occur outside a 2.5m radius of the passing player. Only passes that maintain this characteristic in on-field tracking can be considered for evaluation. These indicators can be averaged across games.
[0115] The evaluation can utilize a dataset containing on-field tracking, broadcast tracking, and event data from, for example, approximately 124 professional football matches. The on-field tracks may be obtained using a live multi-camera tracking system and therefore represent ground truth behavior. The broadcast tracks may be generated by a commercial computer vision tracking system collected from publicly accessible broadcast footage. These systems may consist of object detectors, small segment track association algorithms, and camera calibration models. Event data can be provided on a large scale through manual labeling. Referring to an experiment, the evaluation used 109 games for training and 15 games for evaluation. In the training matches, the system only had access to broadcast tracks for 9 games, and for the remaining games, broadcast tracks were synthesized from on-field tracks. A heuristic-based approach is used to simulate various forms of tracking errors (e.g., object detection errors, player misidentification, camera calibration errors) and occlusions (e.g., due to limited receptive field of view of the camera, close-ups, stop-motion).
[0116] The system setup described herein can differ significantly from conventional trajectory modeling settings in terms of occlusion length, available data streams, and generative objectives (i.e., denoising, interpolation, and prediction). Therefore, based on the experiments discussed above, three baselines were designed to evaluate the diffusion neural network 500. First, a linear interpolator was used, which interpolated occluded actions by linearly transitioning the player from its last visible position to the next. Second, a basic transformer (VanillaTransformer) was used, which independently denoised each agent's trajectory using the player's broadcast trace and event stream. Third, a spatiotemporal axial transformer was used. The spatiotemporal axial transformer is similar to a transformer encoder, where multi-head self-attention is implemented as a temporal attention module followed by a spatial attention module. This architecture is able to represent dependencies between agents. The Vanilla Transformer and the spatiotemporal axial transformer jointly process 30 seconds of broadcast traces and event context. These two transformer baselines make no assumptions about trajectory visibility, ingest extended trajectory context, and fuse event data with tracking data, thus representing significantly stronger baselines than any previous approach derived from conventional systems.
[0117] The implementation of the diffusion neural network 500 can include downsampling of the broadcast and intra-field tracking streams to 5 Hz. According to the experiment, the diffuser utilizes 180 seconds of trajectory and event context, where each segment also contains 500 seconds of event data before and after the main window. Each attention module of the diffusion neural network 500 uses a hidden dimension of 128 and a feedforward dimension of 512, as well as 8 attention heads. For module 502, the event encoder 508 and the tracking decoder 514 have 2 layers and 4 layers, respectively, and the second tracking decoder 520 has 8 layers. Each tracking decoder uses a temporal convolution with a stride and kernel size k=5 (e.g., temporal convolution 518). Each module 502 can be pre-trained with an L2 loss target to reconstruct the intra-field tracking and then fine-tuned on the target (using the above formula (2)).
[0118] Comparison of Diffusion Neural Networks:
[0119] The quantitative performance of the exemplary use of the diffusion neural network 500 can be compared with the three baselines discussed (linear interpolator, vanilla transformer, and spatiotemporal axial transformer models) as shown in Table 1 below. Table 1. Comparison of diffuser 210 to baseline.
[0120] The first two baselines (Linear Interpolator and Vanilla Transformer) may have poor performance in the ADE metric because both process the trajectory of each agent independently and may not use important inter-agent context to denoise and interpolate agent behaviors. In contrast, because the spatiotemporal axial Transformer models inter-agent dependencies, it can have much stronger performance in reconstructing agent positions. However, each baseline exhibits a high proportion of velocity failures. The linear interpolator only interpolates missing behaviors and is therefore unable to correct velocity failures present in the broadcast tracking data (e.g., caused by frequent camera calibration errors and misidentifications). While the Vanilla Transformer and the spatiotemporal axial Transformer may not be able to denoise unrealistic behaviors in the original broadcast tracking, they may be limited by their training regimes. These models can be trained to minimize the L2 reconstruction loss, and although they generally generate independently reasonable positions, these positions generally do not collectively represent realistic human motion. Example use cases of the diffusion neural network 500 have both a lower ADE metric and a lower rate of velocity failures. Example use cases of the diffusion neural network 500 can be, for example, Figures 6A to 6C Describe in.
[0121] Next, the key components of the diffusion neural network 500 can be ablated to determine their relative impact on the model performance. The quantitative results of the ablation experiment (ablation study) are shown in Table 2 below: Table 2 Ablation study of the diffusion neural network 500 (eg, diffuser 210)
[0122] Ablation studies include ablating diffusion. Examining ablations, the ablation study first ablates the use of diffusion by directly using the output of the pre-trained module 502. The model may have very strong performance in terms of the average displacement error (ADE) metric, thereby outperforming the base architecture in reconstructing the short-term target (STO) and the long-term target (LTO). This may be mainly because the event 2 tracking module 502 is trained to directly minimize the L2 reconstruction loss. However, as seen in the baseline architecture, training a multi-agent trajectory generation model via the reconstruction loss may result in unrealistic collective trajectories, as exemplified by the high rate of failure.
[0123] An ablation study included long trajectory context. In this ablation, 30-second trajectory segments were used instead of 180 seconds, examining the importance of extending the denoising timeframe. Shortening trajectory context can reduce the accuracy of reconstructed positions, as evidenced by the significantly weaker AED metrics. This ablation also produced failures at approximately twice the rate of the base architecture, further reinforcing the importance of utilizing a large amount of trajectory context.
[0124] Ablation studies include dilation of events. This involves matching the event time range to 180-second trajectory segments. Reducing this context window results in weaker performance across all ADE metrics and speed failure rates. This also demonstrates the advantage of incorporating as much context as possible when denoising motion trajectories.
[0125] The effect of "pass guidance" on the exemplary use of the diffusion neural network 500 was investigated. This was performed by quantitatively comparing the performance of the exemplary use of the diffusion neural network 500 without and with pass guidance sampling, as shown in Table 3. Table 3 Effect of pass guidance on diffuser 210 (ours)
[0126] While the exemplary use of the diffuser has shown strong performance relative to the AED and speed failure rate metrics, it can also have a high pass failure rate. Possible reasons for this include the large field area, noisy broadcast tracking, and synchronization issues between event and tracking streams. However, with the use of pass guidance, the pass failure rate is significantly reduced while maintaining strong performance in other metrics, demonstrating the granular control provided by the guided diffusion model.
[0127] Figure 6A A first exemplary visualization 600a is depicted of a predicted trajectory 608a of the diffusion neural network 500 along with input data to the network (eg, raw broadcast tracking data 602a and event data 604a) and ground truth data 606a in accordance with one or more embodiments.
[0128] Figure 6B A second exemplary visualization 600b depicts a predicted trajectory 608b of the diffusion neural network 500 along with input data to the network (eg, raw broadcast tracking data 602b and event data 604b) and ground truth data 606b in accordance with one or more embodiments.
[0129] Figure 6C A third exemplary visualization 600c depicts a predicted trajectory 608c of the diffusion neural network 500 along with input data to the network (eg, raw broadcast tracking data 602c and event data 604c) and ground truth data 606c in accordance with one or more embodiments.
[0130] Figures 6A to 6C The tail of the trajectory depicted in can visualize, for example, the previous two seconds of motion. The event data 604a, 604b, 604c show a sequence of three previous events that occurred in the scene. Figure 6A In the example presented, the diffuser neural network can generate highly realistic behaviors for 12 / 23 agents that were not visible in the original broadcast traces. Figure 6B Depicts the diffuser denoising camera calibration errors in the original broadcast tracking, which results in unrealistically fast motion of the four visible players. Figure 6C It is shown how the diffuser model can generate realistic and accurate behavior (e.g., smooth trajectories, rough position accuracy of players, reconstruction of rough events in the ball's trajectory) even in parts of the game where no original broadcast tracking is available.
[0131] Thus, the diffusion neural network 500 can generate complete game tracking without an on-field vision system, representing an important step in providing scalable and unified game analysis across a given sport. The diffusion neural network 500 can include several key technical components. The architecture of module 502 can be a multimodal base model. Second, the diffusion neural network 500 can integrate the base model with guided diffusion, demonstrating that this setup greatly improves the granularity and controllability of the generated behavior. The diffusion neural network 500 can be configured to apply the module 502 architecture to other game analysis tasks that require both trajectory and semantic perception.
[0132] Masked Autoencoders for Interpolating Fine-Grained Adversarial Multi-Agent Trajectories Corrupted by Long-Term Occlusions:
[0133] Due to ease of setup and cost, conventional adversarial multi-agent behavior is captured via video (e.g., broadcast video, as described in this paper). While valuable, the raw pixel information may be of little use for downstream analysis purposes beyond the ability of a human end user to view the video. Compressing the raw pixel information into tracking data (e.g., the spatial location of the agent) can provide a compact yet interpretable mid-level representation of the agent's behavior. A key issue that may not be solvable using conventional techniques is how to interpolate missing tracking segments caused by long-term occlusions. Conventional systems may not be able to interpolate these forms of long-term occlusions, especially when the start and end of the analysis stream are not available.
[0134] The systems and methods described herein can address these limitations with a Multi-Agent Masked Autoencoder (MAT-MAE) system that can be robust to various forms of multi-agent occlusions. These methods can include interpolating long-term occlusions by modeling the distant temporal interdependencies that may exist between the trajectory and coarse semantic data streams. The system disclosed herein can utilize basketball broadcast tracking data in an exemplary embodiment. The system can outperform several baselines. Based on experiments (such as those described herein), an exemplary use of the system can increase the proportion of visible frames in a basketball game, for example, from 59.48% to 86.23%.
[0135] In recent years, video and sensor data from sporting events has grown continuously, allowing for fine-grained behavioral analysis of sporting events. This can include behavioral analysis of single agents (e.g., motion capture), multiple agents (e.g., self-driving cars and human trajectory prediction), and multiple adversarial agents (e.g., sports and athletics). A problem that arises in fine-grained multi-agent behavioral analysis can be the problem of trajectory interpolation, where the movements of multiple agents are reconstructed based on partial observations, as discussed in this paper. This problem can be severe in tracking systems with limited visual perception that cannot track agents that are out of view (e.g., those based on broadcast tracking). In these cases, trajectory generation techniques can be used to interpolate the occluded appearance information.
[0136] Broadcast tracking in sports can form a valuable testbed in which to validate the methods discussed in this article. Unlike traditional on-field tracking systems that have a continuous, unobstructed view of the entire game, broadcast tracking systems can track players directly from the broadcast footage. While this works well when all players are in view, broadcast footage contains different sources of both partial and full occlusion, as discussed in this article.
[0137] The method described herein can classify occlusions into two categories: short-term and long-term occlusions. Short-term occlusions (STOs) can occur when a subset of players are out of view for a short period of time (e.g., ≤5 seconds). Long-term occlusions (LTOs) can refer to situations where no appearance information is observed for a long period of time (e.g., >5 seconds) due to, for example, broadcasts, on-screen graphics, or close-ups of players, coaches, or fans.
[0138] Figure 7A and Figure 7B Depicted are example occlusions of players from an example basketball game in accordance with one or more embodiments. Figure 7C Depicts Figure 7A and Figure 7B Legend 710 may depict which players are obscured and / or observed at the time of the close-up. Figure 7A An example of STO is depicted where, due to the camera angle, a subset of players are not observed for a short period of time, as shown in depiction 702. This scenario is depicted temporally in diagram 704. Figure 7B An example of LTO is depicted where the positions of all players are occluded (e.g., due to a close-up), as shown in depiction 706. This form of occlusion can be represented temporally in graph 708, where no appearance information is observed for a longer period of time.
[0139] Conventional systems that utilize multi-agent trajectories in sports can exploit the coarse spatial structure that exists within sports to roughly estimate the positions of occluded agents. Conventional systems have explored trajectory generation methods to estimate occluded appearance information. An exemplary conventional system can perform acceptable performance in interpolating STO by fusing bidirectional multi-agent context. However, this approach may not be well suited for interpolating LTO. In an exemplary conventional system, a non-autoregressive long-term interpolation method can be implemented. This method may not support partial occlusions (e.g., where a subset of players are occluded at a single frame), which frequently occurs in broadcast footage. Finally, conventional methods all model fixed-length trajectories where the start and end points of the trajectory are strictly visible, which may be a very realistic assumption for broadcast tracking data for sports.
[0140] Unlike conventional systems that use only trajectory information to interpolate behavior, semantic data streams from sports can also be leveraged to more accurately predict behavior, as described in this paper. Such data streams are event detection data that can specify ball-handling events of a game, such as passes, rebounds, and dribbles, with timestamps and player identities. Semantic information can provide a coarse reconstruction of granular multi-agent behavior during a game and can therefore be used to contextualize large portions of a game where trajectory information is completely obscured. Consequently, semantic information can significantly improve the ability to accurately interpolate LTO.
[0141] The multi-agent trajectory interpolation setting may be deeply similar to the mask modeling framework, where the reconstruction of partially masked inputs can be used as a self-supervised pre-training task. Additionally, when fusing sparse semantic information and heavily occluded trajectory data, attention-based models such as the Transformer may have the beneficial property of being able to model long-term temporal dependencies between agents.
[0142] like Figure 8 The depicted Multi-Agent Trajectory Masked Autoencoder (MAT-MAE) 806 can address the aforementioned issues. The MAT-MAE 806 model can be similar in some ways to the diffusion neural network 500 discussed herein. The MAT-MAE 806 can leverage semantic data streams to improve the fidelity of interpolated trajectories, particularly with LTO. The MAT-MAE 806 can be configured to perform interpolation for different classes of occlusions, including trajectories where the start and end points of a possession are heavily occluded. The MAT-MAE 806 can include a new application of the Masked Autoencoder (MAE) architecture in the domain of multi-agent behavior generation. The MAT-MAE 806 is evaluated using an exemplary single game of heavily occluded basketball broadcast trading data, as discussed below.
[0143] The MAT-MAE 806 can be configured to perform long-term trajectory interpolation for different classes of occlusions. The MAT-MAE 806 can utilize both a trajectory stream (e.g., based on the broadcast tracking data 206) and a semantic data stream (e.g., based on the event data 208) to generate a trajectory interpolation framework. The MAE framework can be adapted to interpolation settings with different unseen forms of LTO. The MAT-MAE 806 can explicitly represent permutation-invariant relational dynamics in a team-based multi-agent setting.
[0144] Figure 8An exemplary block diagram of a multi-agent trajectory system 800 is depicted in accordance with one or more embodiments. Entity 1 805 and entity 11 804 can each include visible and occluded trajectory information over time. An encoder (e.g., encoder 508) can operate on masked inputs, and a decoder (e.g., decoder 514) can reconstruct the masked trajectory information. A relational agent encoding approach represents the adversarial and cooperative dynamics of a team-based environment. MAT-MAE 806 can be configured to output one or more trajectories for players of a team (e.g., entity 1 808 and / or entity 11 810) for a sporting event.
[0145] Multi-agent trajectory interpolation can include the task of predicting the occluded position of an agent in a trajectory sequence. Using basketball as an example, all trajectory sets can include K=11 agents (e.g., agents 112A-N), which consist of five offensive players, five defensive players, and a single ball. In addition to partially occluded trajectory information, the system can also utilize a fully visible stream of player events with the ball. Each event includes a timestamp, the player who performed the event, and the event type (including, a shot made, a shot missed, an offensive rebound, a defensive rebound, a turnover, a pass, an inside pass, a block, an assist, and a dribble). Additionally, for certain events, supplementary coarse spatial information may be provided. For a shot, it can be specified whether the shot is a three-pointer, a mid-range shot, or a close-range shot. For an inside pass, the area of the court from which the pass originated (i.e., the baseline, the frontcourt, the backcourt) can be included.
[0146] The problem solved by MAT-MAE 806 can be formalized as discussed further herein. Each ball possession can be represented as a set of entity trajectories The trajectory of agent k consists of T observations At each time step t, the observation of entity k is represented as It includes the (x,y) position of the entity and its event category, i.e. (x,y,event-category). The occlusion can be represented by a mask m∈R T×K×3 Definition, where m t,k,d = 1 specifies that the dimension d of the observation of entity k at time step t is occluded. In this framework, m t,k,2 = 0, which means that the player event may be strictly visible. The interpolation task can be to use the mask input to generate the trajectory set The generated trajectory of entity k is represented as For the input trajectory, each prediction specifies the generated (x,y) position of entity k at time step t. The goal during training can be to make the L2 reconstruction loss of the mask trajectory segment, i.e. Minimize, where ο represents the Hadamard product.
[0147] Multi-agent trajectories can be fundamentally non-Euclidean because there is no natural spatial ordering of entities. Therefore, the best approach to modeling a set of trajectories may be permutation invariant. The system can utilize aspects of permutation invariant approaches for pedestrian trajectory prediction, where inter-agent (agent to other agents) and intra-agent (agent to itself) dependencies are modeled separately. However, in an adversarial team-based environment, inter-agent and intra-agent relationships may also depend on the agent's team affiliation. Using basketball as an example, an offensive player's interaction with another agent may depend on whether the agent is another offensive player, a defensive player, or the ball. Furthermore, the way offensive players attend to themselves may be different from the way defensive players attend to themselves. As a result, both inter-agent attention and intra-agent attention may incorporate player identity.
[0148] Thus, MAT-MAE 806 may include a permutation-invariant agent-based positional encoding method that models the relational dynamics of the adversarial environment of sports. To reflect the effect of team identity on inter-agent attention, a separate relative encoding for each possible ordered pair of team identities may be interpolated. A unique intra-agent relative encoding for each player identity may also be interpolated. More formally, when the player at index k src And the team belongs to a src The behavior subject and index is k dst And the team belongs to a dst When attention is paid between the actors,
[0149] The relational behavior subject code γA is interpolated as follows, intra-agent=1(k src ≡k dst )(9) γ A (a src ,a dst )=embedding((a src ,a dst ,intra-agent))(10)intra-agent: between behavioral subjects; embedding: embedding
[0150] Automatic encoding may be implemented by the MAT-MAE 806. The MAT-MAE 806 may include the ability to represent dependencies between long-term acting agents in LTO with sparse semantic information.
[0151] Within the MAT-MAE 806 framework, each token can be represented by an observation of agent k at time step t. To include the spatiotemporal inductive biases present in multi-agent trajectories, the system can use Shaw's relative position encoding. Different encoding methods can be used for both the agent and time dimensions of multi-agent trajectories. In terms of time, the system can utilize a learned position encoding with a maximum relative attention window rel max = ±40 frames, i.e., ±8 seconds, where the tracking data is sampled at 5 Hz. src and the destination time step t dst The relative time codes γT between these can be interpolated as follows, t diff =clip(t src -t dst ,-rel max ,rel max )(11) γ T (t src ,t dst )=embedding(t diff ).(12)
[0152] For the relative position coding of the behavior subject, the relational behavior subject coding method described in this paper is implemented. and The relative position encoding γ can be interpolated as the sum of the temporal and behavioral subject encodings as follows:
[0153] MAE can exploit symmetric autoencoders. Since there are different numbers of strictly visible semantic events in trajectory sequences, the encoder can handle both masked and non-masked tokens.
[0154] Conventional mask modeling methods can use imputation as a pre-training method. These methods can only study the impact of each strategy in isolation. The system described in this paper can extend these methods by using different sets of synthetic random masks during training to enable strong generalization to a set of unseen, different masks. For each batch during training, the system can randomly select from one of the following five strategies: random, timestep, block, start, and end. Figure 9 These masking strategies are depicted in . Furthermore, since the occlusion ratio in the evaluation competition is also unseen and variable across tracks, the system can implement a random mask ratio P that is sampled in the range [50%, 95%] for each batch.
[0155] Figure 9 An exemplary synthesis masking strategy 900 according to one or more embodiments is depicted. For example, the masking ratio for visualization may be P=60. The squares with slashes may represent visible tracks, and the unfilled squares represent occluded tracks. Masking strategy 902(a) may depict a random masking: where P% of the tokens are randomly masked. Masking strategy 904(b) may depict a time step masking: where P% of the time steps are randomly masked. Masking strategy 906(c) may depict a block masking: where P% of the blocks are randomly masked (blocks of size 3 are used for this visualization). Masking strategy 908(d) may depict a start masking strategy: where the first P% of the time steps are masked. Masking strategy 910(e) may depict an end masking strategy: where the last P% of the time steps are masked.
[0156] The implementation of the MAT-MAE 806 may now be described with reference to an experiment. It should therefore be understood that similar (though not identical) techniques, scopes, data, inputs, and outputs to those described with reference to the experiment may be used for implementation.
[0157] For training, a dataset of 100 games of National Basketball Association (NBA) on-court tracking data was implemented. The data was downsampled from 25Hz to 5Hz, and the games were divided into possessions. A possession begins when a team first establishes possession of the ball in active play and ends when the opposing team establishes possession of the ball or when play becomes inactive (e.g., due to an out-of-bounds). During training, basketball possessions were divided into eight-second segments.
[0158] The MAT-MAE 806 model was trained for 500 epochs using a batch size of 64, an optimizer with a learning rate of 1e-3, and default exponential decay hyperparameters of b1=.9 and b2=0.999. MAT-MAE 806 can be implemented as a symmetric autoencoder with r layers of Transformers, each with a hidden dimension of 64 and 4 attention heads. Evaluation can be performed on a cluster of GPUs.
[0159] During evaluation, training can be performed on fixed-length trajectories of eight seconds, and an autoregressive strategy can be implemented to interpolate trajectories of longer lengths. For example, an eight-second context sliding window may have been implemented, where the system autoregressively updates in four-second increments at a time.
[0160] MAT-MAE 806 is validated by performing two experiments, each constructing a single game of heavily occluded college basketball tracking data. In the first experiment, the system is evaluated on its ability to reconstruct clean on-court tracking data from synthetic masks representing occlusions from broadcast footage. This experiment facilitates a granular evaluation of each method's ability to interpolate various realistic forms of occlusion. In the second experiment, the system is evaluated on its ability to reconstruct noisy, naturally occluded broadcast tracking data. This experiment reflects the practical real-world application of trajectory interpolation methods, and we evaluate the reconstruction using macro-level performance analysis metrics from the sports domain.
[0161] During the evaluation of MAT-MAE 806, the interpolation method was applied to a single college basketball game. For this game, both full in-court tracking data and tracking data with natural occlusions could be processed. The game was downsampled to 5 Hz and divided into possessions. In this game, there were a total of 116 possessions. The frequency of occlusions for each category is shown in Table 5. Table 5 Frequency of occlusion categories in broadcast traces For each category, the occlusion contribution % is reported, which specifies the total % of the entire game frame that is occluded due to each occlusion category.
[0162] To evaluate the fine-grained reconstruction of the affected trajectories, we compute the L2 reconstruction loss, which is the average distance per possession between the ground truth and the generated positions in the occluded part of the trajectory. This metric is reported separately for each entity and occlusion class.
[0163] The following baselines may have been used in the experiments: (i) Linear: The baseline may have performed linear interpolation using the visible portion of the trajectory. (ii) Bidirectional LSTM: The baseline may have used forward and backward LSTM models to interpolate each player's trajectory separately. (iii) Graph Interpolator: The baseline may have included a random GNN-based method that uses bidirectional multi-agent context to interpolate occlusions.
[0164] These baselines may not be able to be interpolated natively without the track boundaries (e.g., the start and end points of the track). Therefore, the system implements a search-based approach that may have interpolated the first and last seconds of the occluded track. This search method can find similar actor tracks from the dataset of unoccluded tracks and copy the first / last seconds of the track. Similar tracks can be defined by significant starting events (e.g., a baseline serve) and / or ending events (e.g., a three-point shot) as well as the visible track information of the entity.
[0165] To investigate the relative contributions of the main components of the MAT-MAE 806 architecture, various ablation studies were conducted. The first ablation study focused on the relational agent encoding module. This model was compared to an approach that uses an absolute positional encoding of the agents, with random indexing of the agents within each team. The second and third ablation studies were conducted on the autoencoder architecture using a conventional system. The ablation studies explored the impact of using an asymmetric autoencoder with a shallow two-layer decoder and a shallow symmetric autoencoder with a two-layer Transformer. The ablation studies further investigated the impact of using a random masking strategy with a masking ratio of 60% and a synthetic masking strategy with a block size of 5 and a masking ratio of 60%.
[0166] Further quantitative analysis was performed. Table 6 below depicts the results for various baselines and ablations. Table 6 Quantitative results of on-field tracking interpolation All results are reported separately for each type of entity (which is represented as attack / defense / ball).
[0167] The system described in this paper has an average L2 reconstruction loss of 2.02 / 2.20 / 2.01 and 5.09 / 6.95 / 4.30 for STO and LTO, respectively, where the loss is reported separately for each class of entity (offense / defense / ball). This outperforms every baseline method for every type of entity and every category of occlusion.
[0168] Experiments were also performed on various ablations of the MAT-MAE 806 architecture. Overall, the base architecture showed the strongest performance. However, an ablation specifically trained with synthetic block masks achieved strong performance for STO interpolation, with results of 1.77 / 2.01 / 1.50, outperforming the base architecture for each class of entities. However, this ablation showed significantly weaker performance than the base architecture in reconstructing LTO, highlighting the usefulness of various synthetic masks during training. While the ablation applied to the MAT-MAE 806 with a shallow autoencoder outperformed the base architecture in reconstructing defenses with LTO, it underperformed the base architecture on all other metrics.
[0169] right Figure 8 The system was further quantitatively studied. Figure 10 A visualization of the three possessions is provided in .
[0170] Figure 10Depicted is an example visualization 1000 of predicted trajectories for a basketball possession using a diffusion neural network (e.g., MAT-MAE 806) according to one or more embodiments. Visualization 1000 includes a first example 1002 for an STO scenario, a second example 1004 for an LTO scenario, and a third example 1006 for an additional LTO scenario. Each column shows the first eight seconds of each of the different possession segments. Ground truth is depicted, along with predictions from the graph interpolator and MAT-MAE 806 for the example scenario. Each example shows the accompanying L2 reconstruction loss for each method over time.
[0171] Qualitatively, based on Figure 10 In the experiments depicted in , MAT-MAE 806 demonstrates impressive performance in generating realistic trajectory boundaries, particularly in the presence of LTO. This is exemplified in the second example 1004 and the third example 1006, where the initial positions of the actors are generated for the first ≥ 7 seconds of the trajectory without appearance information. In these cases, the MAT-MAE 806 method is able to realistically interpolate the initial states of the actors by fusing long-term inter-actor dependencies in both the trajectory and semantic information streams. Since the selected baseline cannot natively interpolate the start / end of the trajectory, a rule-based approach is used to interpolate these parts of the trajectory. While this rule-based approach performs only slightly worse than MAT-MAE for STO (as depicted in the first example 1002), its performance degrades significantly for LTO (as depicted in the second example 1004 and the third example 1006).
[0172] Another notable difference between the performance of MAT-MAE 806 and the baselines is the MAT-MAE approach's ability to generate trajectories exhibiting semantic actions based on the event detection data stream. This is evident in the second example 1004 and the third example 1006, where in the ground truth, the ball clearly swaps between two offensive players, indicating a pass. This same general behavior is exhibited in MAT-MAE's output, reflecting its ability to generate multi-agent trajectories conditional on discrete event data. This capability is achieved by MAT-MAE's use of a Transformer-based approach, which is capable of performing long-term attention with partially occluded trajectory information and sparse semantic information. In contrast, this unique behavior is not produced in the output of the graph interpolator, where the ball instead smoothly transitions from its initial position to its first visible position. At a high level, each of the baselines provides a method for smoothly fusing forward and backward context. As a result, these approaches may be unable to generate actions where the high-level intent of an entity abruptly changes in the occluded portion of the trajectory, such as in a pass.
[0173] The second experiment on MAT-MAE 806's performance performed broadcast tracking interpolation. This second experiment demonstrates the practicality of this interpolation method in downstream applications. Since the system operates in the sports domain, the experiment uses two metrics meaningful for health and performance evaluation: total distance traveled and average speed. These values are also reported across both teams.
[0174] The experiment implemented a baseline that predicts the player's speed based on the average speed of the entity during the visible part of the game. This average is calculated separately for each type of entity (offense, defense, ball). Figure 11 Depicted are exemplary results 1100 of tracking one or more players using a diffusion neural network (eg, MAT-MAE 806 ) according to experiments described herein.
[0175] Examining the results from the experiments, at the team level, MAT-MAE 806 predicted the total distance and average speed of the University of Michigan team to be 39,089 feet and 5.90 feet per second, respectively, and the total distance and average speed of the Penn State team to be 39,628 feet and 5.98 feet per second, respectively. These metrics were significantly more accurate than the average speed baseline. Furthermore, at the individual player level, MAT-MAE 806 was able to more accurately predict the total distance for 16 of the 22 players when compared to the baseline. These strong results reflect potential downstream performance analysis tasks that would be infeasible without this interpolation method.
[0176] MAT-MAE 806 performs multi-agent trajectory interpolation. By leveraging sparse semantic information, MAT-MAE 806 is able to generate high-fidelity behaviors in the presence of various forms of LTO. MAT-MAE 806 demonstrates an impressive ability to reconstruct semantic events in trajectory space and predict the true initial state of the agent when the beginning of the trajectory is heavily occluded. Both quantitatively and qualitatively, MAT-MAE 806 outperforms a range of baseline interpolation methods. Using naturally occluded basketball broadcast tracking data from a single game, MAT-MAE 806 is able to significantly increase the proportion of fully visible frames from 59.48% to 86.23%. MAT-MAE 806 has demonstrated practicality in downstream domain-specific applications in sports.
[0177] Figure 12 An exemplary flowchart 1200 is depicted for a method for tracking one or more individuals during a sporting event according to one or more embodiments. The flowchart 1200 may be, for example, Figure 2 The system 200 may be executed by the system 200. Flowchart 1200 may depict a method for tracking one or more individuals during a sporting event.
[0178] At step 1202, sports broadcast footage of a sporting event may be received as input.
[0179] At step 1204, tagged event data for sports broadcast footage may be received as input. The tagged event data may include a sequential stream of one or more major events from the entire sports event, where the major event includes at least one of a pass, a shot, a steal, a foul, a turnover, a free throw, a goal, a score, or a substitution from the sports event. The event data may be represented as a two-dimensional space-time grid representing a stack of events for each player.
[0180] At step 1206, multi-object tracking may be performed on one or more subjects of the received sports broadcast footage to determine one or more vectors. The one or more vectors may include at least one of two-dimensional coordinates of the subject on the field of the sporting event, the subject's location, the subject's team, an indicator indicating that the subject is the ball, or player visibility information. At step 1208 , the marker event data and one or more vectors may be input into a diffusion model.
[0181] At step 1210, a diffusion model may be used to determine one or more trajectory sequences for one or more actors. The diffusion model may apply spatiotemporal axial attention to the received event data and one or more vectors, wherein self-attention is applied to the temporal axis and the spatial axis, respectively. The diffusion model may include: an event encoder; and a tracking decoder, wherein the event encoder encodes the labeled event data and the tracking decoder conditionally decodes the set of trajectories. The event encoder may embed the event data, embedding the event data further comprising: tokenizing the labeled event data using a linear projection; applying a sinusoidal position embedding to specify the temporal occurrence of the event data; processing the event data with a stacked encoder; and outputting the event embedding. The tracking decoder may use attention to embed and fuse the one or more vectors with the event embedding. The diffusion model may also include: a second tracking decoder; and a transposed temporal convolution, the temporal convolution being configured to extend the trajectory to its original temporal dimension.
[0182] Overview of Neural Network Training and Computation Systems
[0183] Figure 13 Depicted is a flow chart for training a machine learning model according to one aspect. Figure 13As shown in flowchart 1310 of , training data 1312 may include one or more of stage inputs 1314 and known results 1318 associated with the machine learning model to be trained. Stage inputs 1314 may come from any applicable source, including the components or collections shown in the figures provided herein. For machine learning models generated based on supervised or semi-supervised training, known results 1318 may be included. Supervised machine learning models may be trained without using known results 1318. Known results 1318 may include known or expected outputs for future inputs similar or of the same type as stage inputs 1314, for which there are no corresponding known outputs.
[0184] The training data 1312 and the training algorithm 1320 can be provided to a training component 1330, which can apply the training data 1312 to the training algorithm 1320 to generate a trained machine learning model 1350. According to one embodiment, a comparison result 1316 can be provided to the training component 1330, which compares the previous output of the corresponding machine learning model to apply the previous results to retrain the machine learning model. The training component 1330 can use the comparison result 1316 to update the corresponding machine learning model. The training algorithm 1320 can utilize machine learning networks and / or models, including but not limited to deep learning networks such as deep neural networks (DNNs), convolutional neural networks (CNNs), fully convolutional networks (FCNs), and recurrent neural networks (RCNs), probabilistic models (such as Bayesian networks and graphical models), and / or discriminative models (such as decision forests and maximum margin methods). The output of the flowchart 1310 can be a trained machine learning model 1350.
[0185] The machine learning models disclosed herein can be trained by adjusting one or more weights, layers and / or biases during a training phase. During the training phase, historical or simulated data can be provided as input to the model. The model can adjust one or more of its weights, layers and / or biases based on such historical or simulated information. The adjusted weights, layers and / or biases can be configured with a production version of the machine learning model (e.g., a trained model) based on the training. Once trained, the machine learning model can output the output of the machine learning model according to the subject matter disclosed herein. According to one embodiment, one or more machine learning models disclosed herein can be continuously updated based on the use or implementation of the output of the machine learning model.
[0186] It will be understood that the aspects of the present invention are merely exemplary, and that other aspects may include various combinations of features from other aspects, as well as additional or fewer features.
[0187] Generally, any process or operation discussed in the present invention that is understood to be computer-implementable, such as the process shown in the flowchart disclosed herein, can be performed by one or more processors of a computer system, such as any one of the systems or devices in the exemplary environment disclosed herein, as described above. The process or process steps performed by one or more processors can also be referred to as operations. One or more processors can be configured to perform these processes by accessing instructions (e.g., software or computer-readable code) that, when executed by one or more processors, cause one or more processors to perform the process. Instructions can be stored in a memory of a computer system. The processor can be a central processing unit (CPU), a graphics processing unit (GPU), or a processing unit of any suitable type.
[0188] A computer system, such as the system or device implementing the processes or operations in the examples above, may include one or more computing devices, such as one or more of the systems or devices disclosed herein. The one or more processors of the computer system may be included in a single computing device or distributed across multiple computing devices. The memory of the computer system may include a corresponding memory for each of the multiple computing devices.
[0189] Figure 14 14 is a simplified functional block diagram of a computer 1400 that can be configured as an apparatus for performing the methods disclosed herein according to exemplary aspects of the present invention. For example, according to exemplary aspects of the present invention, the computer 1400 can be configured as a system. In various aspects, any of the systems herein can be a computer 1400 that includes, for example, a data communication interface 1420 for packet data communication. The computer 1400 can also include a central processing unit ("CPU") 1402 in the form of one or more processors for executing program instructions. Although the computer 1400 can receive programming and data via network communications, the computer 1400 can include an internal communication bus 1408 and a storage unit 1406 (such as, ROM, HDD, SDD, etc.) that can store data on a computer readable medium 1422.
[0190] Although the instructions 1424 may be temporarily or permanently stored in other modules of the computer 1400 (e.g., the processor 1402 and / or the computer-readable medium 1422), the computer 1400 may also have stored therein instructions for performing the techniques presented herein, such as those related to Figure 13Memory 1404 (such as RAM) for instructions 1424 of the described methods. Computer 1400 may also include input and output ports 1412 and / or display 1410 to connect to input and output devices such as a keyboard, mouse, touch screen, monitor, display, etc. Various system functions can be implemented in a distributed manner on multiple similar platforms to distribute the processing load. Alternatively, the system can be implemented by appropriately programming a computer hardware platform.
[0191] The programmatic aspects of the technology can be considered a "product" or "article of manufacture," which typically takes the form of executable code and / or associated data carried on or embodied in a type of machine-readable medium. "Storage"-type media include computers, processors, etc., or their associated modules, such as any or all of various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. Sometimes, all or part of the software can be communicated via the Internet or various other telecommunications networks. For example, such communication can enable software to be loaded from one computer or processor to another, such as from a management server or host of a mobile communication network to a server's computer platform and / or from a server to a mobile device. Therefore, another type of medium that can carry software elements includes optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and fiber optic route networks, and through various air links. Physical elements that carry such waves, such as wired or wireless links, fiber optic links, etc., can also be considered as media that carry software. As used herein, unless restricted to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0192] Although the disclosed methods, apparatuses, and systems are described with exemplary reference to transmitting data, it should be understood that the disclosed aspects may be applicable to any environment, such as a desktop or laptop computer, a car entertainment system, a home entertainment system, etc. In addition, the disclosed aspects may be applicable to any type of Internet protocol.
[0193] It should be understood that in the above description of exemplary aspects of the invention, various features of the invention are sometimes grouped together in a single aspect, figure, or description thereof to simplify the invention and aid in understanding one or more of the various inventive aspects. However, this approach to the invention should not be interpreted as reflecting an intention that the claimed embodiments require more features than expressly recited in each claim. On the contrary, as reflected in the following claims, inventive aspects exist in features that are less than all of the features of a single aforementioned disclosed aspect. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim existing independently as a separate aspect of the invention.
[0194] Furthermore, although some aspects described herein include some features included in other aspects and not other features, combinations of features from different aspects should be within the scope of the present invention and form different aspects, as will be understood by those skilled in the art. For example, in the following claims, any of the claimed aspects may be used in any combination.
[0195] Thus, while certain aspects have been described, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the invention, and it is intended that all such changes and modifications fall within the scope of the invention. For example, functionality may be added or deleted from the block diagrams, and operations may be interchanged between functional blocks. Operations may be added or deleted from the methods described within the scope of the invention.
[0196] The subject matter disclosed above should be considered illustrative, not restrictive, and the appended claims are intended to cover all such modifications, enhancements, and other embodiments that fall within the true spirit and scope of the invention. Therefore, to the maximum extent permitted by law, the scope of the present invention should be determined by the broadest permissible interpretation of the following claims and their equivalents, and should not be restricted or limited by the foregoing detailed description. Although various embodiments of the present invention have been described, it will be apparent to those skilled in the art that more embodiments may be within the scope of the present invention. Therefore, the present invention is not limited except in accordance with the appended claims and their equivalents. Figure 1 112A-N: Behavior 106: Site 103: Camera 102: Tracking system 105: Network 108: Client device 109: Application 110: Display 104: Computing System 114: Client application server 116: Processor 118: Data storage device 126: Predictor 128A-N: Machine Learning Model 122: Predictive Analytics Engine Figure 2 202: Video input data 206: Broadcast tracking 208: Event data 210: Diffuser 212: Sports Tracking Figure 4 PLAYER OCCLUSION DURATIONS: Duration of player occlusion DURATION: duration NUMBER OF OCCULUSIONS: Number of occlusions Figure 5A 508: Event Encoder 514: Tracking Decoder 524: Denoised Trajectory 522: Temporal Deconvolution 520: Tracking Decoder 506: Temporal Convolution 518: Time Convolution 519: Time Convolution 504: Event Stream 510: Broadcast Tracking 516: Disturbance Tracking Figure 5B 534: Residual Connection and Layer Normalization 532: Feedforward 528: Self-Attention 530: Residual Connections and Layer Normalization 526: Temporal Attention Figure 5C 544: Residual Connections and Layer Normalization 542: Feedforward 540: Time x Attention 538: Spatial Attention 536: Temporal Attention Figure 6A RAW BROADCAST TRACKING: Raw Broadcast Tracking EVENT DATA: event data CONTROL: control PASS: pass GROUND-TRUTH: Ground truth SOCCER DIFFUSER(OURS): Soccer Diffuser (ours) Figure 6B RAW BROADCAST TRACKING: Raw Broadcast Tracking EVENT DATA: event data CONTROL: control PASS: pass GROUND-TRUTH: Ground truth TAKE-ON: vs. TOUCH: Touch SOCCER DIFFUSER(OURS): Soccer Diffuser (ours) Figure 6C RAW BROADCAST TRACKING: Raw Broadcast Tracking EVENT DATA: event data OUT: out of bounds PASS: pass GROUND-TRUTH: Ground truth TAKE-ON: vs. SOCCER DIFFUSER(OURS): Soccer Diffuser(OURS) Figure 7A 、 Figure 7B PLAYER: Player BALL: Ball Figure 7C OBSERVED: Observed OCCLUDED: OPFENSIVE: Offensive DEFENSIVE: defense BALL: ball CURRENT TIMESTEP: Current time step Figure 8 ENTITY: entity TIME: time MULTI-AGENT MASKED AUTOENCODER: Multi-Agent Masked Autoencoder ENCODER: Encoder DECODER: Decoder RELATIONAL AEGNT ENCODING: Relational Behavior Subject Encoding Figure 9 POSSESSION START: Possession starts POSSESSION END: Possession ends PLAYER: player BALL: ball RANDOM: random TIMESTEP: time step BLOCK: block START: start END: End Figure 10 EXAMPLE: Example GROUND-TRUTH: Ground truth GRAPH IMPUTER: OBSERVED: Observed OFFENSIVE: Offense DEFENSIVE: Defense BALL: Ball IMPUTED: Imputed LOSS: Loss TIME: Time SECONDS: seconds Figure 11 TOTAL DISTANCE: Total distance PLAYER: Player BASELINE: Baseline GROUND-TRUTH: Ground truth UNIVERSITY OF MICHIGAN: University of Michigan PENIN STATE: Pennsylvania State University TOTAL: Total Figure 12 1202: Receive geospatial data of sports events as input 1204: Receives tagged event data as input 1206: Perform multi-object tracking on one or more entities of received geospatial data to determine one or more vectors 1208: Inputting the labeled event data and one or more vectors into the diffusion model 1210: Using diffusion models to determine the trajectory of one or more actors Figure 13 1350: Trained Model 1312: Training Data 1314: Stage Input 1318: Known results 1316: Comparison results 1320: Training algorithm 1330: Training Components Figure 14 1402: Processor 1424: Instruction 1404: Memory 1406: driving unit 1422: computer readable medium 1410: display 1412: User input / output device 1420: Communication interface 105: Network
Claims
1. A computer-implemented method for tracking one or more individuals during a sporting event, the method comprising: receiving geospatial data of sports events as input; receiving as input tagged event data based on sports broadcast footage of the sporting event; performing multi-object tracking on one or more subjects of the received geospatial data to determine one or more vectors; inputting the marker event data and one or more vectors into a diffusion model; as well as One or more trajectory sequences of the one or more agents are determined using the diffusion model.
2. The method of claim 1 , wherein the tagged event data comprises a continuous stream of one or more major events throughout a sporting event, the major events comprising at least one of a pass, a shot, a steal, a foul, a turnover, a free throw, a goal, a score, or a substitution from the sporting event, and wherein the geospatial data comprises one or more of the sports broadcast footage, in-field footage, global positioning system (GPS) data, near field communication (NFC) data, or radio frequency identification (RFID) data.
3. The method of claim 1 , wherein the one or more vectors include at least one of the following: two-dimensional coordinates of an actor on a sports field, a position of an actor, a team of an actor, an indicator indicating that the actor is a ball, or player visibility information.
4. The method of claim 1 , wherein the event data is represented as a two-dimensional space-time grid, the grid representing a stack of events for each player.
5. The method according to claim 1, wherein the diffusion model applies spatiotemporal axial attention to the received event data and one or more vectors, wherein self-attention is applied to the time axis and the space axis respectively.
6. The method of claim 1 , wherein the diffusion model comprises: Event encoder; as well as Tracking decoder, The event encoder encodes the marker event data, and the track decoder conditionally decodes a track set.
7. The method of claim 6, wherein the event encoder embeds the event data, embedding the event data further comprising: tokenizing the labeled event data using a linear projection; applying a sinusoidal position embedding to specify a temporal occurrence of the event data; processing the event data with a stacked encoder; as well as Output event embedding.
8. The method according to claim 7, wherein: The tracking decoder uses attention to embed and fuse the one or more vectors with the event embedding.
9. The method according to claim 6, further comprising: a second tracking decoder; as well as Transposed temporal convolution, which is configured to extend the trajectory to its original time dimension.
10. A system for tracking one or more individuals during a sporting event, the system comprising: a non-transitory computer-readable medium configured to store processor-readable instructions; as well as a processor operatively connected to the memory and configured to execute the instructions to perform operations comprising: receiving geospatial data of sports events as input; receiving as input tagged event data based on sports broadcast footage of the sporting event; performing multi-object tracking on one or more subjects of the received geospatial data to determine one or more vectors; inputting the marker event data and one or more vectors into a diffusion model; and One or more trajectory sequences of the one or more agents are determined using the diffusion model.
11. The system of claim 10, wherein the tagged event data comprises a continuous stream of one or more major events throughout a sporting event, the major events comprising at least one of a pass, a shot, a steal, a foul, a turnover, a free throw, a goal, a score, or a substitution from the sporting event, and wherein the geospatial data comprises one or more of the sports broadcast footage, in-field footage, global positioning system (GPS) data, near field communication (NFC) data, or radio frequency identification (RFID) data.
12. The system of claim 10, wherein the one or more vectors include at least one of the following: two-dimensional coordinates of an actor on a sports field, a position of an actor, a team of an actor, an indicator indicating that the actor is a ball, or player visibility information.
13. The system of claim 10, wherein the event data is represented as a two-dimensional space-time grid, the grid representing a stack of events for each player.
14. The system of claim 10, wherein the diffusion model applies spatiotemporal axial attention to the received event data and one or more vectors, wherein self-attention is applied to the time axis and the space axis, respectively.
15. The system of claim 10, wherein the diffusion model comprises: Event encoder; as well as Tracking decoder, The event encoder encodes the marker event data, and the track decoder conditionally decodes a track set.
16. The system of claim 15, wherein the event encoder embeds the event data, embedding the event data further comprising: tokenizing the labeled event data using a linear projection; applying a sinusoidal position embedding to specify a temporal occurrence of the event data; processing the event data with a stacked encoder; as well as Output event embedding.
17. The system according to claim 16, wherein: The tracking decoder uses attention to embed and fuse the one or more vectors with the event embedding.
18. The system of claim 15, further comprising: a second tracking decoder; as well as Transposed temporal convolution, which is configured to extend the trajectory to its original time dimension.
19. A non-transitory computer-readable medium configured to store processor-readable instructions, wherein when executed by a processor, the instructions perform operations comprising: receiving geospatial data of sports events as input; receiving as input tagged event data based on sports broadcast footage of the sporting event; performing multi-object tracking on one or more subjects of the geospatial data to determine one or more vectors; inputting the marker event data and one or more vectors into a diffusion model; as well as One or more trajectory sequences of the one or more agents are determined using the diffusion model.
20. The non-transitory computer-readable medium of claim 19, wherein the tagged event data comprises a continuous stream of one or more major events throughout a sporting event, the major events comprising at least one of a pass, a shot, a steal, a foul, a turnover, a free throw, a goal, a score, or a substitution from the sporting event, and wherein the geospatial data comprises one or more of the sports broadcast footage, in-field footage, global positioning system (GPS) data, near field communication (NFC) data, or radio frequency identification (RFID) data.