Production platform for providing commentary based on event awareness in content media

A machine learning-based simulated commentator generates synchronized or asynchronous commentary for audio and visual content, enhancing viewer engagement and understanding by addressing the time constraints of human analysts.

WO2026039705A1PCT designated stage Publication Date: 2026-02-19TUVIX LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/042102
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-14
Filing Date
2025-08-14
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Real-life commentators and analysts have limited time and cannot engage with all available content, leading to a suboptimal viewer experience in understanding and enjoying audio and visual content.

Method used

Implementing a processor-implemented method using machine learning models to generate audio and/or visual commentary by a simulated commentator based on events in audio and/or visual content media, which can be synchronized or asynchronous with the content playback.

Benefits of technology

Enables continuous and engaging commentary that enhances viewer experience by providing real-time or synchronized analysis and suggestions, addressing the limitations of human commentators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025042102_19022026_PF_FP_ABST
    Figure US2025042102_19022026_PF_FP_ABST
Patent Text Reader

Abstract

A system includes at least one processor, and at least one memory having stored thereon instructions. The instructions, when executed by the at least one processor, cause the system at least to perform: receiving, via an application programming interface, information regarding events captured in an audio and / or visual (AV) content media; and generating, by at least one machine learning model, based on the information regarding the events, at least one of: audio or video, of a simulated commentator which presents commentary regarding the events captured in the AV content media.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. 3766-2 PCTPRODUCTION PLATFORM FOR PROVIDING COMMENTARY BASED ON EVENT AWARENESS IN CONTENT MEDIACROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This patent application claims priority to and / or receives benefit from U.S. Provisional Application No. 63 / 683,019, filed on August 14, 2024. The U.S. Provisional Application is hereby incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] This disclosure relates generally to producing and delivering audio and / or visual commentary, and more specifically, to producing and delivering audio and / or visual commentary by a simulated commentator based on awareness of events in audio and / or visual content media.BACKGROUND

[0003] Commentators and analysts play an important role in shaping an audience’s experience of viewed content, whether in sports, news coverage, esports, entertainment broadcasts, or cultural celebrations, among other content. Their responsibility is to provide narration, context, and analysis that help viewers follow the action and understand its significance. Skilled commentators and analysts blend factual reporting with personality, storytelling, and audience engagement to make the commentary both informative and entertaining. However, real-life commentators and analysts have limited time and cannot engage even a tiny fraction of all content that is available for viewing.SUMMARY

[0004] The present disclosure relates to producing and delivering audio and / or visual commentary by a simulated commentator based on awareness of events in a content media. A content media refers to and may include, for example, one or more streams of content and / or one or more stored clips / files of content. As examples, and without limitation, a content media may include one or more streams of live content (e.g., live sporting event, live gaming event, other live event, etc.), one or more streams of non-live content (e.g., recorded sporting event, recorded gaming event, other recorded event, etc.), and / or one or more files containing contentAttorney Docket No. 3766-2 PCT(e.g., one or more MPEG-4 files containing sporting event, gaming event, other event, etc.), and / or combinations of the foregoing.

[0005] In aspects of the present disclosure, a processor-implemented method includes: receiving, via an application programming interface, information regarding events captured in an audio and / or visual (AV) content media; and generating, by at least one machine learning model, based on the information regarding the events, at least one of: audio or video, of a simulated commentator which presents commentary regarding the events captured in the AV content media.

[0006] In aspects of the present disclosure, a system includes: at least one processor, and at least one memory having stored thereon instructions. The instructions, when executed by the at least one processor, cause the system at least to perform: receiving, via an application programming interface, information regarding events captured in an audio and / or visual (AV) content media; and generating, by at least one machine learning model, based on the information regarding the events, at least one of: audio or video, of a simulated commentator which presents commentary regarding the events captured in the AV content media.

[0007] In aspects of the present disclosure, a non-transitory processor-readable medium has stored thereon instructions which, when executed by at least one processor of a system, cause the system at least to perform: receiving, via an application programming interface, information regarding events captured in an audio and / or visual (AV) content media; and generating, by at least one machine learning model, based on the information regarding the events, at least one of: audio or video, of a simulated commentator which presents commentary regarding the events captured in the AV content media.

[0008] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the simulated commentator is a simulation of a real person, audio of the simulated commentator, if generated, simulates speech of the real person, and video of the simulated commentator, if generated, simulates look of the real person.

[0009] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the simulated commentator is a simulation of a fictional character, audio of the simulated commentator, if generated, simulates speech of the fictional character, and where video of the simulated commentator, if generated, simulates look of the fictional character.

[0010] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the AV content media is streamed or played, and the at least one machine learning model processes the AV content media in real time to generate atAttorney Docket No. 3766-2 PCT least one of audio or video of the simulated commentator while the AV content media is streamed or played. The system, processor-implemented method, or non-transitory processor- readable medium further include: playing the simulated commentator presenting the commentary regarding the events in real time as the AV content media is streamed or played.

[0011] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the AV content media captures one of: a live sporting event, or a live video gaming event.

[0012] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the at least one machine learning model processes the AV content media offline to generate at least one of audio or video of the simulated commentator presenting the commentary.

[0013] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the system, processor-implemented method, or non- transitory processor-readable medium further include: synchronizing the at least one of audio or video of the simulated commentator with the AV content media; and playing the at least one of audio or video of the simulated commentator presenting the commentary in time synchronization with the AV content media.

[0014] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the system, processor-implemented method, or non- transitory processor-readable medium further include: playing the at least one of audio or video of the simulated commentator presenting the commentary separately from the AV content media.

[0015] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the AV content media includes a live stream of a video game played by a user, and the commentary regarding the events captured in the AV content media include live gameplay suggestions for the user.

[0016] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the information regarding events captured in the AV content media includes a priority level for an event.

[0017] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the generating, by the at least one machine learning model, based on the information regarding the events, includes: based on the priority level for the event being relatively lower, generating at least one of: less commentary regarding theAttorney Docket No. 3766-2 PCT event, or no commentary regarding the event; and based on the priority level for the event being relatively higher, generating more commentary regarding the event.

[0018] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, priority level of a frequent event is lower than priority level for an infrequent event.

[0019] In various embodiments of the system, processor-implemented method, or non- transitory processor-readable medium, the information regarding the events includes event descriptions, where the event descriptions are generated by a machine learning model which processed the AV content media to identify the events.

[0020] Any of the embodiments and / or aspects disclosed above or below herein may be combined with one or more or embodiments and / or aspects disclosed herein. All combinations of embodiments and / or aspects are contemplated to be within the scope of the present disclosure.

[0021] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the techniques described in this disclosure will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. To facilitate this description, like reference numerals designate like structural elements. Embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings.

[0023] FIG. 1 illustrates an example of a graphical user interface of a produced episode of commentary for game play, according to some embodiments of the disclosure.

[0024] FIG. 2 illustrates an example of a system involving a production platform, one or more digital double models, a game application, an events application programming interface, and one or more players, according to some embodiments of the disclosure.

[0025] FIG. 3 illustrates an example of a game events manifest, according to some embodiments of the disclosure.

[0026] FIG. 4 illustrates exemplary implementations of the production platform, according to some embodiments of the disclosure.

[0027] FIG. 5 illustrates exemplary implementations of the digital double model, according to some embodiments of the disclosure.Attorney Docket No. 3766-2 PCT

[0028] FIG. 6 illustrates an example of generating different produced episodes for different end users, according to some embodiments of the disclosure.

[0029] FIG. 7 depicts a flow chart illustrating an example of a method for generating a produced episode, according to some embodiments of the disclosure.

[0030] FIG. 8 is a block diagram of an exemplary computing device, according to some embodiments of the disclosure.

[0031] FIG. 9 is a block diagram of an example of an operation for receiving information about events and generating a simulated commentator, according to some embodiments of the disclosure.DETAILED DESCRIPTION

[0032] Overview

[0033] The present disclosure relates to producing and delivering audio and / or visual commentary by a simulated commentator based on awareness of events in a content media. As mentioned above, a content media refers to and may include, for example, one or more streams of content and / or one or more stored clips / files of content. As examples, and without limitation, a content media may include one or more streams of live content (e.g., live sporting event, live gaming event, other live event, etc.), one or more streams of non-live content (e.g., recorded sporting event, recorded gaming event, other recorded event, etc.), and / or one or more stored clips / files of content (e.g., one or more MPEG-4 files containing sporting event, gaming event, other event, etc.), and / or combinations of the foregoing. As used herein the term “AV content media” refers to an audio and / or visual content media, which may include only audio content, only video content, or audio and video content.

[0034] Aspects of the present disclosure apply generative machine learning models / generative artificial intelligence (GenAI) to provide a simulated commentator which presents commentary (which may include analysis, recommendations, etc., among other commentary) regarding an AV content media. In various embodiments, the simulated commentator may be a simulation of a real-life person or may be a simulation of a fictional character (e.g., The Incredible Hulk, Captain America, etc.). In various embodiments, the simulated commentator may be a simulation of a famous persona or may be a simulation of a non-famous persona.

[0035] In various embodiments, the GenAI may provide a simulated commentator which presents commentary at the same time the AV content media is played by audiences, i.e., in real time. For example, as a live event (e.g., live sports event, live esports event, live news reel, etc.) is being live streamed, the GenAI may provide the simulated commentator which presentsAttorney Docket No. 3766-2 PCT commentary at the same time the live event is being live streamed. As another example, as a previous event is being streamed (e.g., previous sports event, previous esports event, previous news reel, etc.), the GenAI may provide a simulated commentator which presents commentary at the same time the previous event is being streamed. As another example, as an AV content media (e.g., containing any content, e.g., movie, TV episode, documentary, etc.) is being played, the GenAI may provide the simulated commentator which presents commentary at the same time the AV content media is being played.

[0036] In various embodiments, the GenAI may provide a simulated commentator which presents commentary that is coupled with and that is time synchronized with the content in the AV content media, such that the simulated commentator presents the commentary in time synchronization with the content in the AV content media as the AV content media is played or streamed by audiences. For example, any content (e.g., sports event, esports event, news reel, movie, TV episode, documentary, etc.) may be recorded and / or stored and may be streamed or played later. The GenAI may provide a simulated commentator which presents commentary that is time synchronized with the content in the content media. As the content media is being played or streamed (e.g., sports event, esports event, news reel, movie, TV episode, documentary, etc.), the time synchronized simulated commentator may present the commentary in time synchronization with the content that is being played or streamed.

[0037] In various embodiments, the GenAI may provide a simulated commentator that presents commentary that is not time synchronized with the content in the AV content media. For example, the simulated commentator may present the commentary separately from the playback or streaming of the AV content media, such that the commentary may be viewed or heard separately from the playback of the AV content media. Such and other embodiments are contemplated to be within the scope of the present disclosure.

[0038] Large language models (LLMs) can be used to create conversant agents having different persona characteristics. Other GenAI, such as Synthesia, among others, can create videos of digital personas. Digital doubles (e.g., digital replicas or digital simulations of real people or fictional characters) can be created for real persons or fictional characters, such as popular or famous personae. As used herein, the terms “digital double” and “simulated commentator” may be used interchangeably. Thus, any mention of a digital double in a description shall be treated as though the description also referred to simulated commentator, and any mention of a simulated commentator in a description shall be treated as though the description also referred to a digital double.Attorney Docket No. 3766-2 PCT

[0039] End users can experience an almost-real-life interaction with digital doubles. Digital doubles include avatars or digital recreations of the real person or fictional character and text-to-speech audio to provide a full audio visual experience for the end user. Avatars can be animated with realistic movement, emotion, and lip synching. Audio can be synthesized to mimic the real person’ s voice and tone or fictional character’ s voice and tone. Digital doubles offer an opportunity for real people or fictional character enterprises to expand online presence and audio visual content creation capabilities. Digital doubles enable famous personae, for example, to create multiple instances of realistic, interactive and personalized experiences with many end users at the same time and without having to be physically present. A produced experience with a digital double or simulated commentator for an end user (e.g., audio visual content having a fixed time duration) can be referred to herein as a produced episode.

[0040] While a digital double can converse with an end user in an asynchronous, isolated chat session naturally, it is not trivial to incorporate a digital double in other experiences, such as video games, interactive experiences, and community experiences. One technical task is to make the digital double aware of the context, such as the game, the experience, and the end user profile. Another technical task is to ensure the digital double has real-time awareness of the experience and can react in real-time. Yet another technical task is to ensure the digital double is prompted to respond to what is most important at a certain point in time. Yet another technical task is to ensure the digital double is prompted to respond in a manner that is aligned with the objectives of the experience. Yet another technical task is to implement intelligent dialogue state tracking to ensure the conversation is as natural and realistic as possible.

[0041] A digital double can be added to complement a game play streaming episode. In some cases, more than one digital double / simulated commentator can be added to complement a game play streaming episode. Users can tune in to the game play stream at a certain time, or users can request the game play stream on demand. The avatar of the digital double with lip synched audio can be displayed on the screen next to the game play. A chat window can be displayed on the screen next to the game play. The game play can be pre-recorded or live. The game play video, the avatar, and the chat window form a produced episode with the digital double. In some embodiments, a produced episode can be produced with the digital double / simulated commentator interacting with many end users via a group chat. In some embodiments, many produced episodes can be produced with the digital double with many end users, where each produced episode can be individualized to the end user. Many produced episodes individualized to the end user can be produced with a different digital double for the same game play.Attorney Docket No. 3766-2 PCT

[0042] A digital double can be added to individual end user game play to create a produced episode that has an enriched interactive user experience with the digital double. The enriched interactive user experience can include a graphical user interface. While an individual end user is playing a game, another application can be running to display an avatar with lip synched audio of the digital double and / or allow the audio of the digital double to be played. The application can display a chat window if desired. The avatar and optionally the chat window form a produced episode with the digital double.

[0043] Besides producing individualized experiences through produced episodes, the interaction with the digital double can be configured to achieve one or more objectives associated with a specific type of interactive experience. In one example, a digital double can offer suggestions, recommendations, and / or training to teach or coach the end user how to play the game, or how to complete a particular level of the game. For example, the digital double can make suggestions for actions to be done in the game. The digital double can compliment or acknowledge the end user for following the suggestions that the digital double has made.

[0044] U sing a digital double in a produced episode that results in a high quality and natural end user experience can have several technical challenges. To address some of these challenges, a production platform may implement an application programming interface (API) for game applications, event hosts, streaming services, and / or playback services, among others, to send event information about events in a content media to the production platform. A game application may provide game event information, which will be used herein as a primary example of the disclosed technology. However, other use cases, such as API for use by event hosts, streaming services, and / or playback services, among others, are contemplated to be within the scope of the present disclosure. Any description of a game application and game events herein shall be understood to be applicable to other use cases and events for the use cases.

[0045] The game events can enable game awareness. The production platform can run separately from the game application to avoid intrusive changes to the game application. The API is designed to allow game developers to easily call the function to send game events to the production platform without significant alterations to the game code. The API can offer abstraction, which means that the API can hide the internal complexities of the production platform and the game application, exposing only the bare minimum functions and data structures. The API allows developers of the production platform and the game application to use the module without understanding its inner workings of the production platform or the digital double model. The API offers a standardized way for the game application to send gameAttorney Docket No. 3766-2 PCT events easily to the production platform. The API separates the production platform from the game application, which means that the production platform and the game application can innovate and change independently without one developer needing to request changes from another developer.

[0046] The production platform can use the game events to enable game awareness. Game awareness can include a game state, which may include one or more of: a chronological history of events, player status, world state, game settings, progression, active quests, economic resource statuses, etc. In some cases, game events have associated priority levels (e.g., high, medium, or low) to allow for filtering and / or prioritization of game events. In some cases, game events have associated categorizations or classes to allow for filtering game events.

[0047] Besides game awareness, the production platform can be aware of the overall experience for the end user, referred to as experience awareness, which may include having awareness of the game (game awareness), awareness of the produced episode (episode awareness), and awareness of the end user (end user awareness). In some cases, the experience awareness may include awareness of a group of end users, such as a group of end users tuned into a streaming game play episode with a digital double. Episode awareness can include one or more of: dialogue states (e.g., conversation states with the end user), time of day, seasonality, relevant promotions, current or trending news events, etc. End user awareness can include one or more of: user profile information, user demographics, user payment history, whether user is logged in, usage patterns, historical user information extracted from the current produced episode or one or more past episodes, social network information, etc.

[0048] The production platform includes logic to maintain and leverage experience awareness to produce suitable digital double inputs to the digital double model. In some cases, experience awareness informs what digital double inputs to generate and send to the digital double model. Using a digital double model can be costly (in terms of system resources, cost per input token, cost per output token, and / or processing time) if every end user input or game event results in a digital double input. Also, producing digital double inputs based on what is salient or appropriate for a particular moment in time can result in digital double outputs that would offer a more natural and realistic interaction with the digital double. Digital double inputs may be derived from end user input (e.g., text messages input by an end user). Digital double inputs may be derived from one or more of: end user state, episode state, and game state.

[0049] The production platform interacts with one or more digital double models. A digital double model can include an LLM and / or other GenAI which generates a video and / or audioAttorney Docket No. 3766-2 PCT of a simulated commentator. In some cases, the LLM can be replaced with one or more suitable natural language processing models, machine learning models, and generative machine learning models. The digital double model can include interaction configurations, which can allow the digital double to respond in a manner that is consistent with one or more objectives and one or more corresponding responsive operations. The digital double model can include static configurations, which can allow the digital double to respond according to personae characteristics. The production platform can send a digital double input to the digital double model. Based on the interaction configurations, static configurations and the digital double input, a prompt generator of the digital double model can generate a prompt to the LLM. The LLM can produce text in response to the prompt.

[0050] In some embodiments, the digital double model may include an avatar generator. Based on the text and an output of the avatar generator, the digital double model can produce a digital double output having, e.g., text, audio, an emotion, and lip synching information. In some cases, the digital double model can produce the audio visual assets for the avatar, and the audio visual assets are included in the digital double output. The digital double output can be passed back to the production platform. The production platform can deliver the digital double output to an end user.

[0051] In one implementation, the production platform may generate digital double outputs to a configured digital double supported by the Inworld platform’ s character creation platform. A digital double model supported by the Inworld platform may be configured for a specific personae through character profiles that guide the digital double’s responses, including backstory, traits, knowledge base, and behavioral tendencies. The digital double model may support maintenance of conversation history and context. The digital double model may support multimodal interaction, such as voice synthesis, facial animation, and avatar generation. The digital double model may recognize and respond to user emotions or simulate appropriate emotional responses based on the character profile. The digital double model may support task-oriented dialogue to handle task-specific interactions.

[0052] In one implementation, the digital double model includes an Inworld character and an NVIDIA instance to generate an avatar (an animated character) based on digital double outputs. The NVIDIA instance may include one or more processes to implement: 3D modeling with detailed 3D meshes of faces and bodies, texture synthesis to generate realistic skin textures, hair, and clothing, facial animation to enable expressive facial movements for more lifelike avatars, style transfer to apply different artistic styles to avatars, and customization to adjust features like hair, skin tone, etc.Attorney Docket No. 3766-2 PCT

[0053] Exemplary graphical user interface of a produced episode

[0054] FIG. 1 illustrates an example of a graphical user interface of a produced episode, according to some embodiments of the disclosure. The illustrated graphical user interface includes an area displaying game play, a chat window, and an avatar that is overlaid on top of the game play in an upper left comer. The avatar can be synthesized or rendered based on the digital double output having one or more of: text, audio, an emotion, and lip synching information. Comments of the digital double can be shown in the chat window. An end user can input chat messages for the digital double and the chat messages can be displayed in the chat window. The end user can use the chat to ask questions about the digital double, the game, or about any topic. The digital double can comment on the game play, comment about a specific subject, and / or converse with an end user.

[0055] Exemplary system for generating audio visual content with game-aware digital doubles

[0056] FIG. 2 illustrates an example of a system involving production platform 202, one or more digital double models (e.g., digital double model 234), game application 250, events API 252, and one or more players (e.g., end user 274, ghost player 272, and other player 270), according to some embodiments of the disclosure. One or more digital double models can include digital double model 234. One or more users can include end user 274, ghost player 272, and other player 270.

[0057] One or more users (e.g., end user 274, ghost player 272, and other player 270) may play a game using game application 250. Game application 250 may provide an interactive virtual environment where one or more users can participate in gameplay sessions. In some cases, game application 250 may be part of a client-server model. In some cases, game application 250 may be a standalone application running on a computing device (e.g., such as computing device 800 of FIG. 8). Game application 250 can include an interactive software system that presents the one or more users with entertainment content through a graphical user interface. Game application 250 can include a game engine that manages a state of the game, tracks user inputs, and generates audiovisual outputs corresponding to gameplay events. Game application 250 can processes user commands received through various input mechanisms from the one or more users to modify game variables and advance game progression according to predefined rulesets and victory conditions. Visual elements of game application 250 can be rendered through a graphics pipeline that generates two-dimensional or three-dimensional representations of game objects and environments. The application maintains data structures for tracking game statistics, user preferences, and achievement metrics while providingAttorney Docket No. 3766-2 PCT feedback through visual, audio and / or haptic signals that correspond to events occurring in a game. During game play, game application 250 may utilize events API 252 to transmit one or more game events to production platform 202.

[0058] Game application 250 communicates game events to production platform 202 via events API 252. Game events communicated via events API 252 implements a core function to enable production platform 202 to be game-aware and to enable digital double model 234 to be game-aware as well.

[0059] A developer / producer of game application 250 and / or production platform 202 can identify a list of game events. A game event can describe an occurrence or happening. A game event can involve a change or an action taking place at a certain point in time.

[0060] A game event can have a corresponding priority or priority level. Priority level may be associated with different labels, and / or different numerical values. Priority levels can be defined as a fixed set of named constants.

[0061] A game event can have a corresponding category or classification.

[0062] A game event can have one or more parameters or variables associated with the game event. A game event can have an object to which the event applies, or to which the event affects. A game event can have a subject, which may be the entity that is performing the action or being described in the event.

[0063] An exemplary list of game events (listed as game event descriptions) and corresponding priority levels are depicted in FIG. 3.

[0064] Referring back to FIG. 2, once the list of game events is established, the game developer / producer can modify the code of game application 250 such that when those game events occur, the code can call an API function, e.g., a function with name “sendDDEventTrigger”, to transmit the game event to production platform 202. The API function, illustrated as events API 252 can include one or more input parameters, such as a string describing the game event, the subject, the object, the class / category, priority, timestamp, (unique) event / experience identifier, game application identifier, etc.

[0065] In some cases, game application 250 can provide one or more of a string describing the game event, the subject, the object, the class / category, and priority as input parameters of the API function to supply information about a game event. A software development kit can be used in conjunction with events API 252 to insert or add the timestamp and (unique) event / experience identifier and send the game event to production platform 202.

[0066] The software development kit can keep track of whether game application 250 was launched with a (unique) event / experience identifier. If game application 250 is launched withAttorney Docket No. 3766-2 PCT the unique event / experience identifier, the software development kit can package / format the information submitted by game application 250 via events API 252 into a game event and send the game event to production platform 202. If game application 250 is not launched with a unique event / experience identifier, the software development kit does not perform any operation (e.g., quick return) even if game application 250 submits information via events API 252.

[0067] Utilizing a string description of the game event in events API 252 to track events occurring in game application 250 can be beneficial because the string description can be directly passed on to digital double model 234 by production platform 202 as a digital double input. The string description can be used readily as part of the digital double input, since the string description describes the game event in natural language. Digital double model 234 can leverage the internal LLM in digital double model 234 to interpret the game event being described using natural language, and understand the contextual cues and nuances of the game event.

[0068] In some cases, production platform 202 may store a game events manifest that associates certain game events with one or more attributes or parameters. For example, production platform 202 can store a game events manifest that associates different game events to different priorities or associates different game events to different classes / categories. Including a game events manifest can mean that game application 250 can supply less information via events API 252 and reduce bandwidth for transmitting game events.

[0069] An exemplary API documentation for the function call “sendDDEventTrigger” for events API 252, based on a REST POST API may include the following:Attorney Docket No. 3766-2 PCT

[0070] Another exemplary API documentation for the function call “sendDDEventTrigger” for events API 252, based on a REST POST API may include the following:Attorney Docket No. 3766-2 PCTAttorney Docket No. 3766-2 PCT

[0071] Yet another exemplary APT documentation for the function call “sendDDEventTrigger” for events API 252, based on a REST POST API may include the following:Attorney Docket No. 3766-2 PCTAttorney Docket No. 3766-2 PCT

[0072] FIG. 4 illustrates exemplary implementations of production platform 202, according to some embodiments of the disclosure.

[0073] Production platform 202 may include admin tool 480. Admin tool 480 may implement a user interface to allow a platform administrator or authorized user to provide one or more inputs: experience start time, estimated duration, name of the experience, one or more identifiers of one or more digital doubles to which the experience corresponds. Admin tool 480 may create a unique experience identifier for the experience.

[0074] Production platform 202 may include experiences publisher and registration manager 404. Experiences publisher and registration manager 404 may implement a user interface to allow one or more end users to view a list of available experiences. Experiences publisher and registration manager 404 may implement a user interface to allow one or more end users to register for an upcoming experience.

[0075] Production platform 202 may include media server 418. At a scheduled experience time, media server 418 can fetch the appropriate video file for the experience if the game play is pre-recorded and begin streaming the file as if it was a live stream. Media server 418 can also stream live game play of game application 250 to one or more end users.

[0076] Production platform 202 may include media production 406. Media production 406 may include studio software to record audio visual content. Media production 406 may include game play recorder 408. Game play recorder 408 may record full screen game play of game application 250. The full screen game play of game application 250 may be generated using ghost player 272 or other player 270.

[0077] Game play recorder 408 may include sync tool 414. Sync tool 414 may receive an experience identifier. Sync tool 414 may receive port information for game play recorder 408 to be able to communicate with game play recorder 408. Sync tool 414 may implement a user interface to allow an authorized user, such as a user that created the experience having theAttorney Docket No. 3766-2 PCT experience identifier, to perform one or more operations, e.g., start recording, stop recording, and finalize recording. A start recording operation can start the game play recording (by sending a command to game play recorder 408 using the port information) and mark the time that the recording started (e.g., experience start time). A stop recording operation can stop the game play recording (by sending a command to game play recorder 408 using the port information) and mark the time that the recording ended or stopped (e.g., experience end time). A finalize recording operation can include downloading the game events associated with the experience identifier from game events 432. A finalize recording operation can include filtering out any events that occurred before the experience start time, and after the experience end time). A finalize recording operation can include adjusting all timestamps of the game events based on the experience start time. Video recordings can start from time 0. Adjusting a timestamp of a game event can include subtracting the experience start time from the timestamp. A finalize recording operation can produce a set of game events with timestamps that are relative to the game play video that was recorded by game play recorder 408. In some cases, a finalize recording operation can include updating the estimated duration of the experience to be the actual duration of the game play recording.

[0078] Production platform 202 may include episode recorder 410 to record a produced episode delivered by web application 412 to an end user. Recorded produced episodes can be stored in produced episodes recordings 488. The recorded game play video can be uploaded to game play recordings 428. In game play recordings 428, the recorded game play video may be tagged with the experience identifier. The time-adjusted game events created by sync tool 414 (also associated with the experience identifier) can be stored with the recorded game play video in game play recordings 428. In game play recordings 428, the time-adjusted game events may be tagged with the experience identifier.

[0079] Production platform 202 may include web application 412 to deliver the audio visual content of a produced episode to an end user. An end user may consume the audio visual content of the produced episode via a web browser (an exemplary user interface of web application 412 for an end user is depicted in FIG. 1).

[0080] If the produced episode involves live game play, web application 412 may connect to media server 418 and deliver an appropriate live stream to the web browser. If the produced episode involves pre-recorded game play, web application 412 may cause media server 418 to fetch and deliver the game play recording stored in game play recordings 428 to the web browser. Web application 412 may fetch time-adjusted game events associated with the experience identifier from game play recordings 428. Each time-adjusted game event mayAttorney Docket No. 3766-2 PCT include a call back. The call back can be called at the appropriate time relative to the game play video, so the game events are in sync with the game play video being watched. Call backs are effective for handling buffering and stalls. The call backs can synchronize game events to be sent to and processed by experience awareness 420.

[0081] Web application 412 may instantiate an instance of an avatar generator (e.g., avatar generator 542) for the digital double associated with the experience. Web application 412 can position a small chat window in the upper right corner on top of game play video, as an example. Web application 412 may include chat management 416, which can include a chat window for the end user, e.g., end user 274 to interact with the digital double. Chat management 416 can send one or more chat messages from an end user to experience awareness 420. Chat management 416 may receive one or more chat messages generated by digital double model 234 and display the received chat messages in the chat window.

[0082] In some cases, web application 412 may generate a produced episode to accompany individual game play. Web application 412 may cause experience awareness 420 to fetch live game events associated with the experience identifier so that game events can be processed by experience awareness 420 in real-time. Web application 412 may instantiate an instance of an avatar generator (e.g., avatar generator 542) for the digital double associated with the experience. Web application 412 can position a small chat window in the upper right corner on top of game play video, as an example. Web application 412 may include chat management 416, which can include a chat window for the end user, e.g., end user 274 to interact with the digital double. Chat management 416 can send one or more chat messages from an end user to experience awareness 420. Chat management 416 may receive one or more chat messages generated by digital double model 234 and display the received chat messages in the chat window.

[0083] Production platform 202 may include game controller 430. Game controller 430 may be used to trigger a game play recording for a particular experience. Game controller 430 may pass in the experience identifier to the game application 250 with command line parameters. A software development kit (SDK) of game application 250 can read the command line parameters and begin sending game events to production platform 202 using events API 252. The SDK can add the timestamp and the experience identifier (provided in the command line parameters) to the game events being sent from game application 250 to production platform 202. Game events received by production platform 202 may be stored in game events 432.Attorney Docket No. 3766-2 PCT

[0084] Production platform 202, such as episode recorder 410) may record a video, such as a video recording of game play, or a video recording of a produced episode, in one or more formats. Examples of video formats may include, MPEG-4 Part 14 (MP4), Audio Video Interleave (AVI), QuickTime File Format (MOV), Windows Media Video (WMV), Flash Video Format (FLV), Web Media Format (WebM), Matroska Multimedia Container (MKV), Moving Picture Experts Group (MPEG), MPEG Transport Stream (TS), 3rd Generation Partnership Project (3GP), etc.

[0085] Production platform 202, such as media server 418, may stream multimedia content, such as game play or a produced episode, to one or more end users using one or more streaming protocols. Examples of video streaming transport protocols include, Real-Time Messaging Protocol (RTMP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming (HLS), Dynamic Adaptive Streaming over HTTP (DASH), Real-Time Streaming Protocol (RTSP), Web Real-Time Communication (WebRTC), Secure Reliable Transport (SRT), User Datagram Protocol (UDP), Transmission Control Protocol (TCP), Microsoft Smooth Streaming (MSS), Adobe HTTP Dynamic Streaming (HDS),

[0086] In some cases, game application 250 may determine whether the game event is associated with an active experience identifier. If the game event is not associated with an active experience identifier, game application 250 does not make the API function call using events API 252. If the game event is associated with an active experience identifier, then game application 250 can make the API function call using events API 252 to send the game event to production platform 202. Game application 250 can insert the current timestamp and the experience identifier and stores the game event in game events 432 in production platform 202.

[0087] In some cases, production platform 202 may determine whether the received game event is associated with an active experience identifier. If the received game event is not associated with an active experience identifier, the received game event is discarded. If the game event is associated with an active experience identifier, then production platform 202 can insert the current timestamp and the experience identifier and stores the game event in game events 432 in production platform 202.

[0088] Production platform 202 includes experience awareness 420. Experience awareness 420 includes one or more of: end user awareness 422, episode awareness 424, and game awareness 426. Experience awareness 420 can track information such as end user input, end user information, historical information, digital double inputs, digital double outputs, and game events, to maintain and track contextual information about a particular produced episode. Experience awareness 420 can take information and produce one or more digital double inputs,Attorney Docket No. 3766-2 PCT which can be passed to digital double model 234 for processing. Experience awareness 420 can receive and process one or more digital double outputs from digital double model 234.

[0089] Experience awareness 420 can include game awareness 426 to track and maintain a game state, which may include one or more of: a chronological history of game events, player status, world state, game settings, progression, economic resource statuses, etc. Chronological history of game events may include a log of received game events associated with the experience identifier. Player status may include health / lives information, score information, inventory information, current location information, etc. World state can include position(s) of other players, state of interactive objects in the game, game clock, world environment conditions, etc. Game settings can include difficulty level, active mods or cheats, graphics settings, audio settings, etc. Progression can include completed quests or missions, active quests or missions, unlocked levels or areas, achieved milestones or checkpoints, saved slots, play time, last checkpoint or autosave, experience points, skills or abilities unlocked, character level, etc. Economic resource statuses can include in-game currency, resources collected, etc.

[0090] In some cases, game events have associated priority levels (e.g., high, medium, or low) to allow for filtering and / or prioritization of game events by game awareness 426. In some cases, game events have associated categorizations or classes to allow for filtering game events by game awareness 426.

[0091] Experience awareness 420 can include episode awareness 424 to track and maintain episode state. Episode state can include one or more of: dialogue states (e.g., conversation states with the end user), time of day, seasonality, relevant promotions, current or trending news events, etc.

[0092] Experience awareness 420 can include end user awareness 422 to track and maintain end user state. End user state can include one or more of: user profile information, user demographics, user payment history, whether user is logged in, usage patterns, historical user information extracted from the current produced episode or one or more past episodes, social network information, etc.

[0093] Experience awareness 420 can implement logic to produce one or more digital double inputs based on information maintained in one or more of: user awareness 422, episode awareness 424, and game awareness 426. Experience awareness 420 may extract salient information based on the information maintained in one or more of: user awareness 422, episode awareness 424, and game awareness 426. In one example, experience awareness 420 may extract salient information based on priority level or priority information associated withAttorney Docket No. 3766-2 PCT game events, which may be maintained or managed by game awareness 426. The digital double inputs (e.g., in the form of one or messages) can be sent to one or more digital double models.

[0094] Experience awareness 420 can receive one or more digital double outputs from the one or more digital double models, and use the one or more digital double outputs in maintaining awareness of the experience (e.g., as part of dialogue state in episode awareness 424).

[0095] FIG. 5 illustrates exemplary implementations of the digital double model, according to some embodiments of the disclosure.

[0096] Digital double model 234 can include interaction configurations 536. Interaction configurations 536 may include objectives 538 and responses 540 that correspond to each objective. Interaction configurations 536 enable prompts to be generated by prompt generator 548 to trigger LLM 546 to respond in a manner that is consistent with interaction configurations 536. Interaction configurations 536 offers a flexible way to configure digital double model 234 to interact in a manner that is tailored to the particular produced experience and / or game application. Interaction configurations 536 enable prompts to be generated by prompt generator 548 to trigger LLM 546 to respond in a manner that is consistent with the tailored experience and / or game application.

[0097] Digital double model 234 can include static configurations 544. Static configurations 544 enable prompts to be generated by prompt generator 548 to trigger LLM 546 to respond in a manner that is consistent with specific personae characteristics of the digital double (e.g., mannerisms, catch phrases, backstory, personality, tone of voice, demographic, interests, etc.).

[0098] Examples of objectives 538 and a corresponding response in responses 540 may include:Attorney Docket No. 3766-2 PCT

[0099] Referring briefly back to FIG. 4, experience awareness 420 of production platform 202 can generate a digital double input based on information such as end user state, episode state, and game state. The digital double input can be sent to the digital double model.

[0100] Referring to FIG. 5, based on interaction configurations 536 and / or static configurations 544, and the digital double input, prompt generator 548 of digital double model 234 can generate a prompt to LLM 546. LLM 546 can produce text or other output modalities in response to the prompt. Prompt generator 548 can craft prompts to elicit optimal responses from LLM 546. In some cases, prompt generator 548 may utilize a prompt template, which may include instructions, context, and examples to guide prompt generator 548 to generate outputs in a specific format and / or to include certain content. Prompt generator 548 can implement prompt chaining, where prompt generator 548 may connect several prompts in aAttorney Docket No. 3766-2 PCT sequential workflow based on interaction configurations 536 and / or static configurations 544 where an output of LLM 546 can serve as part of the next input prompt to LLM 546.

[0101] In some embodiments, prompt generator 548 can implement one or more techniques when generating input prompts to LLM 546 to enable LLM 546 to consider the context or history of the experience. For example, to maintain contextual state with LLM 546, prompt generator 548 can reference previous messages using phrases like "based on what we discussed earlier about X" or "continuing our analysis of Y.” Prompt generator 548 can include relevant context summaries at the beginning of a prompt like "Given that we’ve established [previous conclusions].” Prompt generator 548 can implement use role-based prompting such as "As we continue our conversation where you're acting as a coach that teaches the user how to play the game.” Prompt generator 548 can implement system prompts that define persistent behavioral characteristics throughout the conversation, based on interaction configurations 536 and / or static configurations 544.

[0102] In some embodiments, prompt generator 548 can chain multiple exchanges of prompts and outputs together by crafting prompts that build upon previous outputs generated by LLM 546, such as "Using the analysis you just provided about X, now evaluate how it applies to Y. "

[0103] In some embodiments, prompt generator 548 can maintain state by periodically summarizing key points in a prompt (e.g., based on interaction configurations 536 and / or static configurations 544) and asking LLM 546 to incorporate them into its subsequent responses.

[0104] In some embodiments, prompt generator 548 can use explicit memory tokens like "Remember that earlier we decided X" in a prompt to carry forward important context of the experience.

[0105] In some implementations, functionalities of prompt generator 548, interaction configurations 536, and / or static configurations 544 can be implemented by experience awareness 420 (e.g., using user awareness 422, episode awareness 424, game awareness 426) of FIG. 4.

[0106] Digital double model 234 may include avatar generator 542. In some cases, based on outputs produced by LLM 546, digital double model 234 can produce a digital double output having, e.g., text, audio, an emotion, and lip synching information. In some cases, avatar generator 542 can produce audio visual assets for the avatar, and the audio visual assets are included in the digital double output. The digital double output can be passed back to production platform 202.Attorney Docket No. 3766-2 PCT

[0107] In some cases, LLM 546 implements functionalities associated with avatar generator 542 and produces multimodal digital double outputs having, e.g., text, audio, an emotion, and lip synching information.

[0108] Exemplary use cases and applications of produced episodes with game-aware digital doubles

[0109] FIG. 6 illustrates generating different produced episodes for different end users, according to some embodiments of the disclosure. The data flow illustrates data used to produce a plurality of produced episodes involving pre-recorded game play, and live game play.

[0110] End user inputs 602 from different end users can be provided as input to production platform 202. Production platform 202 may maintain one or more of: end user state 604, episode state 606, and game state 608.

[0111] Production platform 202 may receive game events 612 from a game application. Production platform 202 may maintain game events 612 in game state 608.

[0112] In some embodiments, production platform 202 may record and / or store prerecorded game video 616 along with associated pre-recorded game events 614.

[0113] In some embodiments, production platform 202 may receive game events 612 in real-time and store them as live game events 618. Production platform 202 may receive and / or stream live game video 620 of game play.

[0114] Production platform 202 may include game events manifest 610 to obtain information associated with individual game events, such as priority level, classification, category, etc. In some embodiments, game events manifest 610 is optional, if the information associated with individual game events are included in received game events 612. The information associated with individual game events may be maintained in or used to update game state 608.

[0115] Production platform 202 may produce digital double inputs 630 based on one or more of the end user state 604, episode state 606, and game state 608.

[0116] Different instances of end user state 604 may be maintained for different end users. Digital double inputs 630 generated by production platform 202 may differ for the different end users to tailor the experience to the specific end user.

[0117] Different instances of episode state 606 may be maintained for different end users based on one or more of: end user inputs 602, digital double inputs 630, and digital double outputs 650. Digital double inputs 630 generated by production platform 202 may differ for the different produced episodes 690 to tailor each experience / episode differently.Attorney Docket No. 3766-2 PCT

[0118] Different instances of game state 608 may be maintained for different game plays (based on game events received for a given game play event / experience). Digital double inputs 630 generated by production platform 202 may differ for different game plays to tailor the experience to the actual happenings, occurrences, and state of the game.

[0119] Digital double inputs 630 may be sent to digital double model 234. Based on the digital double configurations 640 (e.g., static configurations and interaction configurations), digital double model 234 may produce digital double outputs 650. In some cases, the digital double outputs 650 may include avatar audio visual assets 660. Digital double outputs 650 may be sent to production platform 202.

[0120] Production platform 202 can produce individualized produced episodes 690 for different end users based on the digital double outputs 650.

[0121] In some embodiments, production platform 202 can produce a produced episode having more than one digital double that is aware of the experience. The digital doubles may have different interaction configurations and / or static configurations, and thus may respond differently to the same digital double inputs. Depending on the experience, the same or different digital double inputs may be sent to the different digital doubles.

[0122] In some embodiments, production platform 202 can produce a produced episode that is streamed to a plurality of end users who have tuned into the produced episode. The produced episode may include a chat with the plurality of end users. The end users may submit chat messages to the group chat. Chat messages can be directed to the digital double, or to one or more other end users. Experience awareness in production platform 202 may maintain group state 688, which may determine or derive one or more salient topics being discussed in the group chat between the digital double and the plurality of end users. Based on group state 688, a digital double input may be generated and sent to digital double model 234 as part of digital double inputs 630. Digital double model 234 may generate a digital double output as part of digital double outputs 650, which may be incorporated by production platform 202 into the produced episode that is being streamed to the plurality of end users. The (common) digital double output may be output to all end users that have tuned in.

[0123] In some scenarios, experience awareness may maintain a plurality of end user states associated with different end users (e.g., different instances of end user state 604). End user state 604 can be used to derive one or more individualized salient topics associated with a particular end user. Based on end user state 604, a digital double input may be generated and sent to digital double model 234 as part of digital double inputs 630. Digital double model 234 may generate a digital double output as part of digital double outputs 650, which may beAttorney Docket No. 3766-2 PCT incorporated by production platform 202 into the produced episode that is being streamed to particular end user. The individualized digital double output may be sent to the particular end user as a private message, or as a group chat message that tags the particular end user, via the chat for a mixed group and individualized experience.

[0124] FIG. 7 depicts a flow chart illustrating method 700 for generating a produced episode, according to some embodiments of the disclosure. Method 700 can be performed by computing device 800 of FIG. 8. Method 700 can be performed by one or more components illustrated in FIG. 2, e.g., production platform 202, and digital double model 234.

[0125] In 702, a game event is received via an application programming interface from a game application. In 704, the game event is maintained in a game state. In 706, a digital double input based on the game state. In 708, the digital double input is transmitted to a digital double model. In 710, a digital double output is received from the digital double model. In 712, a graphical user interface output is generated based on the digital double output. In 714, the graphical user interface output is transmitted to an end user device.

[0126] In some embodiments, method 700 further includes determining a priority associated with the game event and maintaining the priority in the game state.

[0127] In some embodiments, method 700 further includes receiving an end user input, and maintaining the end user input, and a timestamp corresponding to receiving the end user input in an episode state. Determining the digital double input can be further based on the episode state.

[0128] In some embodiments, method 700 further includes receiving an end user profile information and maintaining the end user profile information in an end user state. Determining the digital double input can be further based on the end user state.

[0129] In some embodiments, method 700 further includes storing a video recording of game play, a start time of the video recording, and an end time of the video recording, in a game play recordings data store. Method 700 further includes storing the game event and a timestamp corresponding to the game event in a game events data store. Method 700 can further include subtracting the timestamp by the start time of the video recording.

[0130] In some embodiments, method 700 further includes determining that the game event is associated with a produced experience and associating the game event and the timestamp corresponding to the game event to an identifier associated with the produced experience in the game events data store.

[0131] In some embodiments, the game event comprises an identifier associated with a produced experience. In some embodiments, the game event comprises a timestamp. In someAttorney Docket No. 3766-2 PCT embodiments, the game event comprises a string that describes the game event. In some embodiments, the game event comprises one or more parameters for the game event. In some embodiments, the game event comprises an enumerated value that indicates a priority associated with the game event.

[0132] Exemplary computing device

[0133] FIG. 8 is a block diagram of an exemplary computing device 800, according to some embodiments of the disclosure. One or more computing devices 800 may be used to implement the functionalities described with FIG. 8 and herein (e.g., operations described in FIGS. 1 -7 and / or 9). A number of components are illustrated in the FIGS, as included in the computing device 800, but any one or more of these components may be omitted or duplicated, as suitable for the application. In some embodiments, some or all of the components included in the computing device 800 may be attached to one or more motherboards. In some embodiments, some or all of these components are fabricated onto a single system on a chip (SoC) die. Additionally, in various embodiments, the computing device 800 may not include one or more of the components illustrated in FIG. 8, and the computing device 800 may include interface circuitry for coupling to the one or more components. For example, the computing device 800 may not include a display device 806, and may include display device interface circuitry (e.g., a connector and driver circuitry) to which a display device 806 may be coupled. In another set of examples, the computing device 800 may not include an audio input device 818 or an audio output device 808 and may include audio input or output device interface circuitry (e.g., connectors and supporting circuitry) to which an audio input device 818 or audio output device 808 may be coupled.

[0134] The computing device 800 may include a processing device 802 (e.g., one or more processing devices, one or more of the same type of processing device, one or more of different types of processing device). The processing device 802 may include electronic circuitry that processes electronic data from data storage elements (e.g., registers, memory, resistors, capacitors, quantum bit cells) to transform that electronic data into other electronic data that may be stored in registers and / or memory. Examples of processing device 802 may include a central processing unit (CPU), a graphical processing unit (GPU), a quantum processor, a machine learning processor, an artificial intelligence processor, a neural network processor, an artificial intelligence accelerator, an application specific integrated circuit (ASIC), an analog signal processor, an analog computer, a microprocessor, a digital signal processor, a field programmable gate array (FPGA), a tensor processing unit (TPU), a data processing unit (DPU), etc.Attorney Docket No. 3766-2 PCT

[0135] The computing device 800 may include a memory 804, which may itself include one or more memory devices such as volatile memory (e.g., DRAM), nonvolatile memory (e.g., read-only memory (ROM)), high bandwidth memory (HBM), flash memory, solid state memory, and / or a hard drive. Memory 804 includes one or more non-transitory computer- readable storage media. In some embodiments, memory 804 may include memory that shares a die with the processing device 802.

[0136] In some embodiments, memory 804 includes one or more non-transitory computer- readable media and / or processor-readable media storing instructions executable to perform operations described with FIGS. 1-6 and / or 9. Memory 804 may store instructions that encode one or more exemplary parts described with FIGS. 2 and 4-6. Memory 804 may store instructions that encode one or more exemplary operations described with method 700 of FIG. 7. The instructions stored in the one or more non-transitory computer-readable media may be executed by processing device 802.

[0137] In some embodiments, memory 804 may store data, e.g., data structures, binary data, bits, metadata, files, blobs, etc., as described with the FIGS, and herein.

[0138] In some embodiments, memory 804 may store one or more machine learning models (and or parts thereof) that are used in a digital double model. Memory 804 may store training data for training the one or more machine learning models. Memory 804 may store input data (e.g., input tokens), output data (e.g., output tokens), intermediate outputs, intermediate inputs of one or more machine learning models. Memory 804 may store instructions to perform one or more operations of the machine learning model. Memory 804 may store one or more parameters used by the machine learning model. Memory 804 may store information that encodes how processing units of the machine learning model are connected with each other.

[0139] In some embodiments, the computing device 800 may include a communication device 812 (e.g., one or more communication devices). For example, the communication device 812 may be configured for managing wired and / or wireless communications for the transfer of data to and from the computing device 800. The term "wireless" and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data through the use of modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. The communication device 812 may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEEAttorney Docket No. 3766-2 PCT802.10 family), IEEE 802.16 standards (e.g., IEEE 802.16-2005 Amendment), Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as "3GPP2"), etc.). IEEE 802.16 compatible Broadband Wireless Access (BWA) networks are generally referred to as WiMAX networks, an acronym that stands for worldwide interoperability for microwave access, which is a certification mark for products that pass conformity and interoperability tests for the IEEE 802. 16 standards. The communication device 812 may operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. The communication device 812 may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). The communication device 812 may operate in accordance with Codedivision Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. The communication device 812 may operate in accordance with other wireless protocols in other embodiments. The computing device 800 may include an antenna 822 to facilitate wireless communications and / or to receive other wireless communications (such as radio frequency transmissions). The computing device 800 may include receiver circuits and / or transmitter circuits. In some embodiments, the communication device 812 may manage wired communications, such as electrical, optical, or any other suitable communication protocols (e.g., the Ethernet). As noted above, the communication device 812 may include multiple communication chips. For instance, a first communication device 812 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second communication device 812 may be dedicated to longer-range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some embodiments, a first communication device 812 may be dedicated to wireless communications, and a second communication device 812 may be dedicated to wired communications.

[0140] The computing device 800 may include power source / power circuitry 814. The power source I power circuitry 814 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 800 to an energy source separate from the computing device 800 (e.g., DC power, AC power, etc.).Attorney Docket No. 3766-2 PCT

[0141] The computing device 800 may include a display device 806 (or corresponding interface circuitry, as discussed above). The display device 806 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display, for example.

[0142] The computing device 800 may include an audio output device 808 (or corresponding interface circuitry, as discussed above). The audio output device 808 may include any device that generates an audible indicator, such as speakers, headsets, or earbuds, for example.

[0143] The computing device 800 may include an audio input device 818 (or corresponding interface circuitry, as discussed above). The audio input device 818 may include any device that generates a signal representative of a sound, such as microphones, microphone arrays, or digital instruments (e.g., instruments having a musical instrument digital interface (MIDI) output).

[0144] The computing device 800 may include a GPS device 816 (or corresponding interface circuitry, as discussed above). The GPS device 816 may be in communication with a satellite-based system and may receive a location of the computing device 800, as known in the art.

[0145] The computing device 800 may include a sensor 830 (or one or more sensors). The computing device 800 may include corresponding interface circuitry, as discussed above). Sensor 830 may sense physical phenomenon and translate the physical phenomenon into electrical signals that can be processed by, e.g., processing device 802. Examples of sensor 830 may include: capacitive sensor, inductive sensor, resistive sensor, electromagnetic field sensor, light sensor, camera, imager, microphone, pressure sensor, temperature sensor, vibrational sensor, accelerometer, gyroscope, strain sensor, moisture sensor, humidity sensor, distance sensor, range sensor, time-of-flight sensor, pH sensor, particle sensor, air quality sensor, chemical sensor, gas sensor, biosensor, ultrasound sensor, a scanner, etc.

[0146] The computing device 800 may include another output device 810 (or corresponding interface circuitry, as discussed above). Examples of the other output device 810 may include an audio codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, haptic output device, gas output device, vibrational output device, lighting output device, home automation controller, or an additional storage device.Attorney Docket No. 3766-2 PCT

[0147] The computing device 800 may include another input device 820 (or corresponding interface circuitry, as discussed above). Examples of the other input device 820 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a bar code reader, a Quick Response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.

[0148] The computing device 800 may have any desired form factor, such as a handheld computer system, mobile computer system, a cell phone, a smart phone, a mobile internet device, a music player, a tablet computer, a laptop computer, a desktop computer system, a server or other networked computing component, a set-top box, an entertainment control unit, a vehicle control unit, a virtual reality system, an augmented reality system, a mixed reality system, a smart watch, smart glasses, and a wearable computer system. In some embodiments, the computing device 800 may be an electronic device or system that processes data and can output audio visual content.

[0149] Large language models

[0150] Various embodiments of the digital double models described herein involve one or more large language models. A large language model is a type of artificial intelligence system that uses deep learning techniques, specifically transformers and self-attention mechanisms, to process and generate human- like text based on patterns learned from vast amounts of training data. These models are trained on massive datasets, often having billions of words, allowing them to learn complex patterns in language usage, context, and meaning.

[0151] A large language model has a transformer-based architecture. The transformer is one of the building blocks of a large language model. The transformer is a type of neural network that uses self- attention mechanisms to capture long-range dependencies in sequential data, such as text. The transformer architecture includes an encoder and a decoder, both having multiple (multi-head) attention layers and feedforward neural network layers.

[0152] A large language model may include an embeddings layer, an encoder, a decoder, and output layer.

[0153] Embeddings layer converts the input text into numerical vector representations called embeddings. These embeddings represent the semantic and syntactic properties of words, allowing the large language model to understand the meaning and context of the input. Since the transformer architecture does not have an inherent notion of word order, positional encodings can be added to the input embeddings to provide the model with information about the position of each word in the sequence.Attorney Docket No. 3766-2 PCT

[0154] The encoder processes the input sequence and creates a context-aware representation. The encoder includes multiple attention layers and feedforward neural network layers. The encoder may include layer normalization and residual connections.

[0155] The decoder takes the encoded input representation from the encoder and generates the output sequence, token by token. The decoder can autoregressively generate output tokens one by one, attending to the encoded input and the previous output. The decoder includes multiple attention layers and feedforward neural network layers. The decoder may include layer normalization and residual connections.

[0156] The output layer takes the representations from the decoder and can output probability distributions over the vocabulary for the next token in the sequence. The output layer may include a softmax activation function to convert logits into probabilities.

[0157] The attention layers allow the model to weigh different parts of the input sequence when producing the output. The attention mechanism enables the model to focus on the most relevant parts of the input for a given task, such as generating a coherent and contextually appropriate response. Multi-head attention is a technique that allows the large language model to attend to different representations of the input simultaneously. Multi-head attention may include several attention heads, each of which learns to attend to different aspects of the input, improving the model's ability to capture complex relationships and patterns. Multi-head attention applies multiple attention operations in parallel. Each attention head can learn to focus on different aspects of the input, such as syntactic structure, semantic relationships, or long- range dependencies. The outputs of these heads are then concatenated and linearly transformed to produce the final attention output.

[0158] Feedforward neural network layers apply non-linear transformations to the output of the attention layers, allowing the model to learn more complex representations of the input data. A feedforward neural network layer may include two linear transformations with a nonlinear activation function in between. The feedforward neural network layer may allow the model to leam more complex, non-linear representations of the data, enhancing its ability to capture intricate patterns and relationships in language.

[0159] The input text, or a sequence of input tokens, received and processed by a large language model is referred to as a prompt. A prompt may include a sequence of words and characters. The words and characters may be converted by the large language model into a sequence of tokens. In some cases, a prompt can range from a single word to multiple paragraphs, depending on the task and model capabilities. The prompt is first tokenized into a sequence of tokens, which are then converted into input embeddings.Attorney Docket No. 3766-2 PCT

[0160] The description above discussed a gaming application as a primary example of the disclosed technology. As mentioned above, the disclosed technology is applicable to other use cases. Across use cases, the uses may apply a similar operational flow, which will now be described in connection with FIG. 9.

[0161] FIG. 9 is a block diagram of an example of an operational flow. The operational flow involves use of an application programming interface (API) 920, such as the API described above herein (e.g., FIG. 2, 252). Aspects of an API were described above, e.g., in connection with FIG. 2, among other figures. The API 920 is used to convey information about events in an audio and / or visual (AV) content media to a platform that generates a simulated commentator. Aspects of information about events were described above herein, e.g., in connection with FIG. 3, among other figures. The information about events may include, for example, event description 930 of an event, event time information 932 for an event, and commentary priority level 934 for an event.

[0162] In accordance with aspects of the present disclosure, the API 920 may be used via programmed use 910, via manual use 912, and / or via artificial intelligence (Al) use 914.

[0163] Programmed use 910 of the API refers to programming uses of the API into an application. Examples of programmed use 910 of the API include, without limitation, programming API functionalities into an application (e.g., gaming application, streaming application, playback application, etc.). Examples relating to programming API functionalities into a gaming application were described above herein. As an example of programming API functionalities into a streaming or playback application, a streaming or playback application may be programmed to, for example, utilize the API 920 when triggers are encountered in the content (e.g., metadata in content may act as triggers). In various embodiments, a streaming or playback application may be programmed to, for example, utilize the API 920 at regular intervals, among other possibilities. Other scenarios for programmed use 910 of the API are contemplated to be within the scope of the present disclosure.

[0164] Manual use 912 of the API refers to uses that are manually triggered by a person. As an example, personnel for a live stream may manually trigger use of the API 920 by, e.g., blogging in real-time about events in the live stream. Each blog entry may trigger use of the API 920 to convey information about an event in the live stream, e.g., as reflected in the blog entry, for example. Other scenarios for manual use 912 of the API are contemplated to be within the scope of the present disclosure.

[0165] Artificial intelligence (Al) use 914 of the API refers to Al models which can recognize events occurring in AV content and automatically convey information about theAttorney Docket No. 3766-2 PCT recognized events via the API. Examples of artificial intelligence (Al) use 914 of the API include, for example, Al models which process images, video, audio, and / or metadata of the content in a stream or a playback to recognize what is occurring in the content of the stream or playback and which generate descriptors or descriptions of the events. Examples of Al models that may recognize what is occurring in the content of a stream or playback include object detection machine learning models and scene analysis machine learning models, among other possibilities. Examples of such models include, without limitation, YOLO, EfficientDet, RetinaNet, Detection Transformer, and Azure Al Video Indexer, among other possibilities.

[0166] In accordance with aspects of the present disclosure, combinations of programmed use 910 of the API, manual use 912 of the API, and / or Al use 914 of the API may be implemented and are contemplated to be within the scope of the present disclosure.

[0167] As described above, the API 920 is used to convey information about events in AV content media, and the event information may include event description 930, event time information 932, and commentary priority level 934, among other event information. In various embodiments, event description 930 may be generated by programmed event descriptions, manually entered event descriptions, and / or ALgenerated event descriptions, among other possibilities. In various embodiments, event time information 932 may be generated by programming, by manual entry, and / or by Al-inference. In various embodiments, commentary priority level for an event may be a programmed priority level, a manually entered priority level, or an Al-inferred priority level. Such and other embodiments are contemplated to be within the scope of the present disclosure.

[0168] In accordance with aspects of the present disclosure, event information (e.g., 930- 934) are input to a large language model 940, which processes the event information to generate text commentary about the event. Such text commentary may include, for example, a discussion of the event, analysis of the event, and / or recommendations relating to the event. Events which have lower priority level may result in little discussion, if any. Events which have higher priority may result in greater discussion. In various embodiments, the large language model 940 may be trained to simulate the word choices, sentence structure, and language style, etc., of a persona, such as a real person or a fictional character.

[0169] In accordance with aspects of the present disclosure, the text commentary generated by the large language model 940 is input to a Gen Al model 950 that generates audio and / or video of a simulated commentator to present the commentary. An example of such a GenAI model is provided by Synthesia, but other machine learning models are contemplated to be within the scope of the present disclosure. The output of the GenAI model 950 may includeAttorney Docket No. 3766-2 PCT only audio (e.g., commentary for radio audience about a sport game), only video (e.g., video of a sign language commentator), or audio and video. In various embodiments, the GenAI model 950 may be trained to simulate the look and / or speech of a persona, such as a real person or a fictional character. Such and other embodiments are contemplated to be within the scope of the present disclosure.

[0170] Accordingly, the operation flow of FIG. 9 allows various uses of an API to convey information about events, which are then used to generate a simulated commentator which presents commentary about the events. The illustration in FIG. 9 and the description of FIG. 9 are merely examples, and variations are contemplated to be within the scope of the present disclosure.

[0171] The following are examples of use cases of aspects of the disclosed technology.

[0172] In aspects of the present disclosure, the AV content media may include a live stream of a video game gameplay. The information regarding events in the video game gameplay may be generated by programmed event descriptions in the gaming application. One or more machine learning models may process the information regarding events in the video game replay in real time to provide audio and / or video of a simulated commentator which provides commentary regarding the events in the video game gameplay. The audio and / or video of the simulated commentator may be played in real time during the live stream of the video game gameplay. In various embodiments, the commentary may include live gameplay recommendations for a user of the video game.

[0173] In aspects of the present disclosure, the AV content media may include a live stream of a sports game (e.g., baseball, football, etc.). The information regarding events in the sports game may be generated by manual event descriptions and / or by Al-generated event descriptions. One or more machine learning models may process the information regarding events in the sports game in real time to provide audio and / or video of a simulated commentator which provides commentary regarding the events in the sports game. The audio and / or video of the simulated commentator may be played in real time during the live stream of the sports event.

[0174] In aspects of the present disclosure, the AV content media may include a previously captured event, such as a previous video game gameplay or a previous sports game, among other possibilities. The information regarding events in the previously captured event may be generated by programmed event descriptions, manual event descriptions, and / or by AI- generated event descriptions. One or more machine learning models may process the information regarding events offline to provide audio and / or video of a simulated commentatorAttorney Docket No. 3766-2 PCT which provides commentary regarding the events in the previously captured event. The audio and / or video of the simulated commentator may be time synchronized with the content of the AV content media and may be played in time synchrony with the content of the AV content media.

[0175] In aspects of the present disclosure, the AV content media may include a previously captured event, such as a previous video game gameplay or a previous sports game, among other possibilities. The information regarding events in the previously captured event may be generated by programmed event descriptions, manual event descriptions, and / or by AI- generated event descriptions. One or more machine learning models may process the information regarding events offline to provide audio and / or video of a simulated commentator which provides commentary regarding the events in the previously captured event. The audio and / or video of the simulated commentator is not synchronized with the content of the AV content media and may be played separately from the content of the AV content media.

[0176] In aspects of the present disclosure, where the commentary of the simulated commentator is intended for a radio audience, only audio of the simulated commentator may be generated.

[0177] In aspects of the present disclosure, where the commentary of the simulated commentator is intended for a sign language audience, only video of sign language of the simulated commentator may be generated.

[0178] The uses cases above are merely examples, and other use cases are contemplated to be within the scope of the present disclosure.

[0179] The following describe Examples of various aspects of the present disclosure.

[0180] Example 1.1. A system, comprising: at least one processor; and at least one memory having stored thereon instructions which, when executed by the at least one processor, cause the system at least to perform: receiving, via an application programming interface, information regarding events captured in an audio and / or visual (AV) content media; and generating, by at least one machine learning model, based on the information regarding the events, at least one of: audio or video, of a simulated commentator which presents commentary regarding the events captured in the AV content media.

[0181] Example 1.2. The system of Example 1.1, wherein the simulated commentator is a simulation of a real person,Attorney Docket No. 3766-2 PCT wherein audio of the simulated commentator, if generated, simulates speech of the real person, and wherein video of the simulated commentator, if generated, simulates look of the real person.

[0182] Example 1.3. The system of Example 1.1, wherein the simulated commentator is a simulation of a fictional character, wherein audio of the simulated commentator, if generated, simulates speech of the fictional character, and wherein video of the simulated commentator, if generated, simulates look of the fictional character.

[0183] Example 1.4. The system of any one of Example 1.1- Example 1.3, wherein the AV content media is streamed or played, wherein the at least one machine learning model processes the AV content media in real time to generate at least one of audio or video of the simulated commentator while the AV content media is streamed or played, and wherein the instructions, when executed by the at least one processor, further cause the system at least to perform: playing the simulated commentator presenting the commentary regarding the events in real time as the AV content media is streamed or played.

[0184] Example 1.5. The system of Example 1.4, wherein the AV content media captures one of: a live sporting event, or a live video gaming event.

[0185] Example 1.6. The system of any one of Example 1.1- Example 1.3, wherein the at least one machine learning model processes the AV content media offline to generate at least one of audio or video of the simulated commentator presenting the commentary.

[0186] Example 1.7. The system of Example 1.6, wherein the instructions, when executed by the at least one processor, further cause the system at least to perform: synchronizing the at least one of audio or video of the simulated commentator with the AV content media; and playing the at least one of audio or video of the simulated commentator presenting the commentary in time synchronization with the AV content media.

[0187] Example 1.8. The system of Example 1.6, wherein the instructions, when executed by the at least one processor, further cause the system at least to perform: playing the at least one of audio or video of the simulated commentator presenting the commentary separately from the AV content media.

[0188] Example 1.9. The system of any one of Example 1.1- Example 1.5,Attorney Docket No. 3766-2 PCT wherein the AV content media comprises a live stream of a video game played by a user, and wherein the commentary regarding the events captured in the AV content media comprise live gameplay suggestions for the user.

[0189] Example 1.10. The system of any one of Example 1.1- Example 1.9, wherein the information regarding events captured in the AV content media comprises a priority level for an event.

[0190] Example 1.1 1. The system of Example 1.10, wherein the generating, by the at least one machine learning model, based on the information regarding the events, comprises: based on the priority level for the event being relatively lower, generating at least one of: less commentary regarding the event, or no commentary regarding the event; and based on the priority level for the event being relatively higher, generating more commentary regarding the event.

[0191] Example 1.12. The system of Example 1.11, wherein priority level of a frequent event is lower than priority level for an infrequent event.

[0192] Example 1.13. The system of any one of Example 1.1- Example 1.12, wherein the information regarding the events comprises event descriptions, wherein the event descriptions are generated by a machine learning model which processed the AV content media to identify the events.

[0193] Example 2.1. A non-transitory processor-readable medium having stored thereon instructions which, when executed by at least one processor of a system, cause the system at least to perform: receiving, via an application programming interface, information regarding events captured in an audio and / or visual (AV) content media; and generating, by at least one machine learning model, based on the information regarding the events, at least one of: audio or video, of a simulated commentator which presents commentary regarding the events captured in the AV content media.

[0194] Example 2.2. The non-transitory processor-readable medium of Example 2.1, wherein the simulated commentator is a simulation of a real person, wherein audio of the simulated commentator, if generated, simulates speech of the real person, and wherein video of the simulated commentator, if generated, simulates look of the real person.

[0195] Example 2.3. The non-transitory processor-readable medium of Example 2.1,Attorney Docket No. 3766-2 PCT wherein the simulated commentator is a simulation of a fictional character, wherein audio of the simulated commentator, if generated, simulates speech of the fictional character, and wherein video of the simulated commentator, if generated, simulates look of the fictional character.

[0196] Example 2.4. The non-transitory processor-readable medium of any one of Example 2.1- Example 2.3, wherein the AV content media is streamed or played, wherein the at least one machine learning model processes the AV content media in real time to generate at least one of audio or video of the simulated commentator while the AV content media is streamed or played, and wherein the instructions, when executed by the at least one processor, further cause the system at least to perform: playing the simulated commentator presenting the commentary regarding the events in real time as the AV content media is streamed or played.

[0197] Example 2.5. The non-transitory processor-readable medium of Example 2.4, wherein the AV content media captures one of: a live sporting event, or a live video gaming event.

[0198] Example 2.6. The non-transitory processor-readable medium of any one of Example 2.1- Example 2.3, wherein the at least one machine learning model processes the AV content media offline to generate at least one of audio or video of the simulated commentator presenting the commentary.

[0199] Example 2.7. The non-transitory processor-readable medium of Example 2.6, wherein the instructions, when executed by the at least one processor, further cause the system at least to perform: synchronizing the at least one of audio or video of the simulated commentator with the AV content media; and playing the at least one of audio or video of the simulated commentator presenting the commentary in time synchronization with the AV content media.

[0200] Example 2.8. The non-transitory processor-readable medium of Example 2.6, wherein the instructions, when executed by the at least one processor, further cause the system at least to perform: playing the at least one of audio or video of the simulated commentator presenting the commentary separately from the AV content media.Attorney Docket No. 3766-2 PCT

[0201] Example 2.9. The non-transitory processor-readable medium of any one of Example 2.1- Example 2.5, wherein the AV content media comprises a live stream of a video game played by a user, and wherein the commentary regarding the events captured in the AV content media comprise live gameplay suggestions for the user.

[0202] Example 2. 10. The non-transitory processor- readable medium of any one ofExample 2.1- Example 2.9, wherein the information regarding events captured in the AV content media comprises a priority level for an event.

[0203] Example 2.11. The non-transitory processor-readable medium of Example 2.10, wherein the generating, by the at least one machine learning model, based on the information regarding the events, comprises: based on the priority level for the event being relatively lower, generating at least one of: less commentary regarding the event, or no commentary regarding the event; and based on the priority level for the event being relatively higher, generating more commentary regarding the event.

[0204] Example 2. 12. The non-transitory processor-readable medium of Example 2.11 , wherein priority level of a frequent event is lower than priority level for an infrequent event.

[0205] Example 2.13. The non-transitory processor-readable medium of any one of Example 2.1- Example 2.12, wherein the information regarding the events comprises event descriptions, wherein the event descriptions are generated by a machine learning model which processed the AV content media to identify the events.

[0206] Example 3.1. A processor-implemented method comprising: receiving, via an application programming interface, information regarding events captured in an audio and / or visual (AV) content media; and generating, by at least one machine learning model, based on the information regarding the events, at least one of: audio or video, of a simulated commentator which presents commentary regarding the events captured in the AV content media.

[0207] Example 3.2. The processor-implemented method of Example 3.1, wherein the simulated commentator is a simulation of a real person, wherein audio of the simulated commentator, if generated, simulates speech of the real person, andAttorney Docket No. 3766-2 PCT wherein video of the simulated commentator, if generated, simulates look of the real person.

[0208] Example 3.3. The processor-implemented method of Example 3.1, wherein the simulated commentator is a simulation of a fictional character, wherein audio of the simulated commentator, if generated, simulates speech of the fictional character, and wherein video of the simulated commentator, if generated, simulates look of the fictional character.

[0209] Example 3.4. The processor-implemented method of any one of Example 3.1- Example 3.3, wherein the AV content media is streamed or played, wherein the at least one machine learning model processes the AV content media in real time to generate at least one of audio or video of the simulated commentator while the AV content media is streamed or played, and the processor-implemented method further comprising: playing the simulated commentator presenting the commentary regarding the events in real time as the AV content media is streamed or played.

[0210] Example 3.5. The processor-implemented method of Example 3.4, wherein the AV content media captures one of: a live sporting event, or a live video gaming event.

[0211] Example 3.6. The processor-implemented method of any one of Example 3.1- Example 3.3, wherein the at least one machine learning model processes the AV content media offline to generate at least one of audio or video of the simulated commentator presenting the commentary.

[0212] Example 3.7. The processor-implemented method of Example 3.6, further comprising: synchronizing the at least one of audio or video of the simulated commentator with the AV content media; and playing the at least one of audio or video of the simulated commentator presenting the commentary in time synchronization with the AV content media.

[0213] Example 3.8. The processor-implemented method of Example 3.6, further comprising: playing the at least one of audio or video of the simulated commentator presenting the commentary separately from the AV content media.Attorney Docket No. 3766-2 PCT

[0214] Example 3.9. The processor-implemented method of any one of Example 3.1- Example 3.5, wherein the AV content media comprises a live stream of a video game played by a user, and wherein the commentary regarding the events captured in the AV content media comprise live gameplay suggestions for the user.

[0215] Example 3.10. The processor-implemented method of any one of Example 3.1- Example 3.9, wherein the information regarding events captured in the AV content media comprises a priority level for an event.

[0216] Example 3.11. The processor-implemented method of Example 3.10, wherein the generating, by the at least one machine learning model, based on the information regarding the events, comprises: based on the priority level for the event being relatively lower, generating at least one of: less commentary regarding the event, or no commentary regarding the event; and based on the priority level for the event being relatively higher, generating more commentary regarding the event.

[0217] Example 3.12. The processor-implemented method of Example 3.11, wherein priority level of a frequent event is lower than priority level for an infrequent event.

[0218] Example 3.13. The processor-implemented method of any one of Example 3.1-Example 3.12, wherein the information regarding the events comprises event descriptions, wherein the event descriptions are generated by a machine learning model which processed the AV content media to identify the events.

[0219] Example 3.14. A system comprising: at least one processor; and at least one memory having stored thereon instructions which, when executed by the at least one processor, cause the system at least to perform a method as in any one of Example 3.1- Example 3.13.

[0220] Example 3.15. A non-transitory processor-readable medium having stored thereon instructions which, when executed by the at least one processor, cause the system at least to perform a method as in any one of Example 3.1- Example 3.13.

[0221] Example 4.1. A method, comprising: receiving a game event via an application programming interface from a game application; maintaining the game event in a game state;Attorney Docket No. 3766-2 PCT determining a digital double input based on the game state; transmitting the digital double input to a digital double model; receiving a digital double output from the digital double model; generating a graphical user interface output based on the digital double output; and transmitting the graphical user interface output to an end user device.

[0222] Example 4.2. The method of Example 4. 1 , further comprising: determining a priority associated with the game event; and maintaining the priority in the game state.

[0223] Example 4.3. The method of Example 4.1 or Example 4.2, further comprising: receiving an end user input; and maintaining the end user input, and a timestamp corresponding to receiving the end user input in an episode state; wherein determining the digital double input is further based on the episode state.

[0224] Example 4.4. The method of any one of Example 4.1 - Example 4.3, further comprising: receiving an end user profile information; and maintaining the end user profile information in an end user state; wherein determining the digital double input is further based on the end user state.

[0225] Example 4.5. The method of any one of Example 4.1- Example 4.4, further comprising: storing a video recording of game play, a start time of the video recording, and an end time of the video recording, in a game play recordings data store; and storing the game event and a timestamp corresponding to the game event in a game events data store.

[0226] Example 4.6. The method of Example 4.5, further comprising: subtracting the timestamp by the start time of the video recording.

[0227] Example 4.7. The method of Example 4.5 or Example 4.6, further comprising: determining that the game event is associated with a produced experience; and associating the game event and the timestamp corresponding to the game event to an identifier associated with the produced experience in the game events data store.Attorney Docket No. 3766-2 PCT

[0228] Example 4.8. The method of any one of Example 4.1- Example 4.7, wherein the game event comprises an identifier associated with a produced experience.

[0229] Example 4.9. The method of any one of Example 4.1- Example 4.8, wherein the game event comprises a timestamp.

[0230] Example 4.10. The method of any one of Example 4.1- Example 4.9, wherein the game event comprises a string that describes the game event.

[0231] Example 4. 11. The method of any one of Example 4.1 - Example 4.10, wherein the game event comprises a parameter for the game event.

[0232] Example 4.12. The method of any one of Example 4.1- Example 4.11, wherein the game event comprises an enumerated value that indicates a priority associated with the game event.

[0233] Example 4. 13. One or more non-transitory computer-readable media storing instructions which, when executed by one or more processors of a system, cause the system to perform a method as in any one of Example 4.1- Example 4.12.

[0234] Example 4.14. A system, comprising: one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform a method as in any one of Example 4. 1- Example 4.12.

[0235] Variations and other notes

[0236] Although the operations of the example methods shown in and described with reference to the FIGS, are illustrated as occurring once each and in a particular order, it will be recognized that the operations may be performed in any suitable order and repeated as desired. Additionally, one or more operations may be performed in parallel. Furthermore, the operations illustrated in the FIGS, may be combined or may include more or fewer details than described.

[0237] The description of illustrated implementations of the disclosure, including what is described in the Abstract, is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. While specific implementations of, and examples for, the disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure, as those skilled in the relevant art will recognize. These modifications may be made to the disclosure in light of the detailed description.

[0238] For purposes of explanation, specific numbers, materials and configurations are set forth in order to provide a thorough understanding of the illustrative implementations. However, it will be apparent to one skilled in the art that the present disclosure may be practicedAttorney Docket No. 3766-2 PCT without the specific details and / or that the present disclosure may be practiced with only some of the described aspects. In other instances, well known features are omitted or simplified in order not to obscure the illustrative implementations.

[0239] Further, references are made to the accompanying drawings that form a part hereof, and in which are shown, by way of illustration, embodiments that may be practiced. It is to be understood that other embodiments may be utilized, and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense.

[0240] Various operations may be described as multiple discrete actions or operations in turn, in a manner that is most helpful in understanding the disclosed subject matter. However, the order of description should not be construed as to imply that these operations are necessarily order dependent. In particular, these operations may not be performed in the order of presentation. Operations described may be performed in a different order from the described embodiment. Various additional operations may be performed or described operations may be omitted in additional embodiments.

[0241] For the purposes of the present disclosure, the phrase “A or B” or the phrase "A and / or B" means (A), (B), or (A and B). For the purposes of the present disclosure, the phrase “A, B, or C” or the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). The term "between," when used with reference to measurement ranges, is inclusive of the ends of the measurement ranges.

[0242] The description uses the phrases "in an embodiment" or "in embodiments," which may each refer to one or more of the same or different embodiments. The terms "comprising," "including," "having," and the like, as used with respect to embodiments of the present disclosure, are synonymous. The disclosure may use perspective-based descriptions such as "above," "below," "top," "bottom," and "side" to explain various features of the drawings, but these terms are simply for ease of discussion, and do not imply a desired or required orientation. The accompanying drawings are not necessarily drawn to scale. Unless otherwise specified, the use of the ordinal adjectives “first,” “second,” and “third,” etc., to describe a common object, merely indicates that different instances of like objects are being referred to and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking or in any other manner.

[0243] In the following detailed description, various aspects of the illustrative implementations will be described using terms commonly employed by those skilled in the art to convey the substance of their work to others skilled in the art.Attorney Docket No. 3766-2 PCT

[0244] The terms “substantially,” “close,” “approximately,” “near,” and “about,” generally refer to being within + / - 20% of a target value as described herein or as known in the art. Similarly, terms indicating orientation of various elements, e.g., “coplanar,” “perpendicular,” “orthogonal,” “parallel,” or any other angle between the elements, generally refer to being within + / - 5-20% of a target value as described herein or as known in the art.

[0245] In addition, the terms “comprise,” “comprising,” “include,” “including,” “have,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a method, process, or device, that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such method, process, or device. Also, the term “or” refers to an inclusive “or” and not to an exclusive “or.”

[0246] The systems, methods and devices of this disclosure each have several innovative aspects, no single one of which is solely responsible for all desirable attributes disclosed herein. Details of one or more implementations of the subject matter described in this specification are set forth in the description and the accompanying drawings.

Claims

Attorney Docket No. 3766-2 PCTWhat is Claimed Is:

1. A processor-implemented method comprising: receiving, via an application programming interface, information regarding events captured in an audio and / or visual (AV) content media; and generating, by at least one machine learning model, based on the information regarding the events, at least one of: audio or video, of a simulated commentator which presents commentary regarding the events captured in the AV content media.

2. The processor-implemented method of claim 1 , wherein the simulated commentator is a simulation of a real person, wherein audio of the simulated commentator, if generated, simulates speech of the real person, and wherein video of the simulated commentator, if generated, simulates look of the real person.

3. The processor-implemented method of claim 1, wherein the simulated commentator is a simulation of a fictional character, wherein audio of the simulated commentator, if generated, simulates speech of the fictional character, and wherein video of the simulated commentator, if generated, simulates look of the fictional character.

4. The processor-implemented method of any one of claims 1-3, wherein the AV content media is streamed or played, wherein the at least one machine learning model processes the AV content media in real time to generate at least one of audio or video of the simulated commentator while the AV content media is streamed or played, and the processor-implemented method further comprising: playing the simulated commentator presenting the commentary regarding the events in real time as the AV content media is streamed or played.

5. The processor- implemented method of claim 4, wherein the AV content media captures one of: a live sporting event, or a live video gaming event.Attorney Docket No. 3766-2 PCT6. The processor-implemented method of any one of claims 1-3, wherein the at least one machine learning model processes the AV content media offline to generate at least one of audio or video of the simulated commentator presenting the commentary.

7. The processor-implemented method of claim 6, further comprising: synchronizing the at least one of audio or video of the simulated commentator with the AV content media; and playing the at least one of audio or video of the simulated commentator presenting the commentary in time synchronization with the AV content media.

8. The processor-implemented method of claim 6, further comprising: playing the at least one of audio or video of the simulated commentator presenting the commentary separately from the AV content media.

9. The processor-implemented method of any one of claims 1-5, wherein the AV content media comprises a live stream of a video game played by a user, and wherein the commentary regarding the events captured in the AV content media comprise live gameplay suggestions for the user.

10. The processor-implemented method of any one of claims 1-9, wherein the information regarding events captured in the AV content media comprises a priority level for an event.

11. The processor-implemented method of claim 10, wherein the generating, by the at least one machine learning model, based on the information regarding the events, comprises: based on the priority level for the event being relatively lower, generating at least one of: less commentary regarding the event, or no commentary regarding the event; and based on the priority level for the event being relatively higher, generating more commentary regarding the event.

12. The processor-implemented method of claim 11 , wherein priority level of a frequent event is lower than priority level for an infrequent event.Attorney Docket No. 3766-2 PCT13. The processor-implemented method of any one of claims 1-12, wherein the information regarding the events comprises event descriptions, wherein the event descriptions are generated by a machine learning model which processed the AV content media to identify the events.

14. A system comprising: at least one processor; and at least one memory having stored thereon instructions which, when executed by the at least one processor, cause the system at least to perform a method as in any one of claims 1- 13.

15. A non-transitory processor-readable medium having stored thereon instructions which, when executed by the at least one processor, cause the system at least to perform a method as in any one of claims 1-13.

Citation Information

Patent Citations

  • Method for implementing event real-time commentation and medium

    CN108337573A

  • Information processing device, information processing method, and information processing system

    EP4325376A1

  • Driving virtual influencers based on predicted gaming activity and spectator characteristics

    US20210299575A1

  • Intelligent commentary generation and playing methods, apparatuses, and devices, and computer storage medium

    US20230362457A1