Artificial intelligence (AI) controlled camera perspective generator and AI broadcaster
By using artificial intelligence to generate scene renderings and commentary for video games, identifying areas of interest for viewers and providing personalized commentary, the problem of a lack of professional coverage of ordinary players' gaming sessions is solved, enhancing the viewer experience and the appeal of game streaming.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SONY INTERACTIVE ENTERTAINMENT LLC
- Filing Date
- 2020-09-09
- Publication Date
- 2026-06-02
AI Technical Summary
In existing video game streaming services, the lack of detailed coverage and engaging commentary from professional broadcasters during player sessions results in a poor viewing experience, fails to meet the needs of audiences from different languages and regions, and current technologies cannot effectively address this issue.
By using artificial intelligence to generate renderings and narrations of game scenes, it identifies areas of interest for viewers and provides personalized comments. It also uses AI models to analyze game status data and user data to generate the most interesting camera angles and narrations.
It enhances the viewer experience of video game streaming, provides personalized game scene rendering and commentary, meets the interests and needs of different viewers, and enhances the attractiveness and interactivity of live games.
Smart Images

Figure CN116474378B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application 202080081541.5 entitled “AI-controlled camera perspective generator and AI broadcaster”, filed on September 9, 2020. Technical Field
[0002] This disclosure relates to video games or game applications. Among other things, this disclosure describes one or more views of a scene for selecting and / or generating a game session of a video game, and methods and systems for broadcasting the generation of said one or more views to one or more viewers. Background Technology
[0003] Video game streaming is becoming increasingly popular. Viewers can access both high-viewership tournaments featuring professional gamers and gameplay videos posted by ordinary gamers. Viewers can access live streams of game sessions, as well as streams of previous game sessions.
[0004] The production of popular tournaments includes detailed coverage of corresponding game sessions provided by professional broadcasters. For example, tournament broadcasters are fed multiple broadcasts of a game session along with various facts associated with it. Broadcasts can be obtained from game views generated for one or more players within the session. Facts may be presented by the ongoing gameplay or by a human assistant constantly searching for interesting facts about the players and / or the game session. Multiple broadcasts of a game session can point to different players within the session. The broadcasters can choose when to show which broadcasts and provide commentary using facts related to the game session. Thus, viewers of these live tournaments experience professionally produced programming, including the game session itself and the broadcasters' engaging and optimal perspectives that make the viewing experience even more exciting.
[0005] Viewers can also enjoy streaming game sessions from ordinary gamers. Some of these streams don't include any commentary and may only consist of a view of the game session displayed on the player's screen. Other game session streams may include commentary, such as comments provided by the player while they are playing the game application. In a sense, the player provides live commentary during the game session. However, the popularity of streams featuring ordinary gamers' game sessions may be limited. For example, viewers of these streams may tire of watching game sessions without any commentary, or may avoid streams that don't offer commentary. Additionally, viewers of these streams may tire of live commentary provided by the player in the game session, as the player is only focused on his or her gameplay. Viewers may want to hear information about other players in the game session, or they may want to hear additional insights into the game session that the commenting player may not be aware of. For example, viewers may want background information on the game session, such as background information provided by professional broadcasters of live sports or gaming events. Or viewers may want to see scenes showing other players' gameplay along with interesting facts corresponding to their scenes. Essentially, it is logically and economically impractical for viewers to expect professional human broadcasters to stream ordinary gamers' game sessions.
[0006] It is against this backdrop that the proposed implementation scheme has emerged. Summary of the Invention
[0007] The embodiments disclosed herein relate to generating one or more renderings of scenes for a game application using artificial intelligence (AI), and further relate to generating narration and / or generating scene renderings for scenes of a game application using AI.
[0008] In one embodiment, a method for generating a view of a game is disclosed. The method includes receiving game state data and user data of the one or more players participating in a game session of a video game being played. The method includes identifying a region of interest (ROI) in the game session, the ROI having a scene of a virtual game world of the video game, the scene being viewable from one or more camera perspectives within the virtual game world. The method includes identifying a first camera perspective based on a first AI model of the ROI, the first AI model being trained to generate one or more corresponding camera perspectives of the scene corresponding to the ROI.
[0009] In another embodiment, a non-transitory computer-readable medium is disclosed that stores computer programs for generating views of a game. The computer-readable medium includes program instructions for receiving game state data and user data of the one or more players participating in a game session of a video game being played. The computer-readable medium includes program instructions for identifying a region of interest (ROI) in the game session, the ROI having a scene of a virtual game world of the video game, the scene being viewable from one or more camera perspectives within the virtual game world. The computer-readable medium includes program instructions for identifying a first camera perspective based on a first AI model trained to generate one or more corresponding camera perspectives of a scene corresponding to the ROI.
[0010] In yet another embodiment, a computer system is disclosed, comprising a processor and a memory, wherein the memory is coupled to the processor and stores instructions therein, which, when executed by the computer system, cause the computer system to perform a method for generating a view of a game. The method includes receiving game state data and user data of the one or more players participating in a game session of a video game being played. The method includes identifying a region of interest (ROI) in the game session, the ROI having a scene of a virtual game world of the video game, the scene being viewable from one or more camera perspectives within the virtual game world. The method includes identifying a first camera perspective of the ROI based on a first AI model trained to generate one or more corresponding camera perspectives of the scene corresponding to the ROI.
[0011] In another embodiment, a method for generating a broadcast is disclosed. The method includes receiving game state data and user data of one or more players participating in a game session of a video game being played. The method includes identifying a region of interest (ROI) in the game session, the ROI having a scene of a virtual game world of the video game, viewable from one or more camera perspectives within the virtual game world. The method includes generating statistics and facts for the game session based on the game state data and the user data using a first AI model trained to isolate game state data and user data that one or more viewers might be interested in. The method includes generating narration for the scene in the ROI using a second AI model configured to select statistics and facts from the statistics and facts generated using the first AI model, the selected statistics and facts having the highest potential viewer interest determined by the second AI model, and the second AI model configured to generate narration using the selected statistics and facts.
[0012] In another embodiment, a non-transitory computer-readable medium is disclosed storing computer programs for generating broadcasts. The computer-readable medium includes program instructions for receiving game state data and user data of the one or more players participating in a game session of a video game being played. The computer-readable medium includes program instructions for identifying a region of interest (ROI) in the game session, the ROI having a scene of a virtual game world of the video game, viewable from one or more camera perspectives within the virtual game world. The computer-readable medium includes program instructions for generating statistics and facts for the game session based on the game state data and the user data using a first AI model trained to isolate game state data and user data that might be of interest to one or more viewers. The computer-readable medium includes program instructions for generating narration for the scene in the ROI using a second AI model configured to select statistics and facts from those generated using the first AI model, the selected statistics and facts having the highest potential viewer interest determined by the second AI model, and the second AI model configured to generate narration using the selected statistics and facts.
[0013] In another embodiment, a computer system is disclosed, comprising a processor and a memory, wherein the memory is coupled to the processor and stores instructions therein, which, when executed by the computer system, cause the computer system to perform a method for generating a broadcast. The method includes receiving game state data and user data of the one or more players participating in a game session of a video game being played. The method includes identifying a region of interest (ROI) in the game session, the ROI having a scene of a virtual game world of the video game, the scene being viewable from one or more camera perspectives within the virtual game world. The method includes generating statistics and facts for the game session based on the game state data and the user data using a first AI model, the first AI model being trained to isolate game state data and user data that one or more viewers might be interested in. The method includes generating narration for the scene in the ROI using a second AI model, the second AI model being configured to select statistics and facts from the statistics and facts generated using the first AI model, the selected statistics and facts having the highest potential viewer interest determined by the second AI model, the second AI model being configured to generate narration using the selected statistics and facts.
[0014] Other aspects of this disclosure will become apparent from the following detailed description taken in conjunction with the accompanying drawings, which illustrate the principles of this disclosure by way of example. Attached Figure Description
[0015] This disclosure can be best understood by referring to the following description in conjunction with the accompanying drawings, in which:
[0016] Figure 1A A system according to one embodiment of the present disclosure is shown, the system being used for AI-controlled generation of camera views of scenes in a game session of a video game, and for narration of scenes and / or AI-controlled generation of camera views of scenes.
[0017] Figure 1B An exemplary neural network for constructing an artificial intelligence (AI) model according to one embodiment of the present disclosure is shown, which includes an AI camera view model and an AI narration model.
[0018] Figure 2A system is shown according to one embodiment of the present disclosure for providing game control to one or more users playing one or more game applications, said one or more game applications being executed locally on the corresponding user or on a backend cloud gaming server in one or more game sessions, wherein the game sessions are used to train camera view models and narration models via AI, the AI models being used to identify interesting camera views of scenes in the game sessions and to provide narration for those scenes in the game sessions that are most likely to interact with the audience.
[0019] Figure 3A A system configured to train an AI model according to an embodiment of the present disclosure is shown, the AI model being able to identify regions of interest (ROIs) in multiple game sessions of multiple game applications to identify interesting camera views of those identified ROIs and to provide narration for the scenes of those identified ROIs.
[0020] Figure 3B A system for training a viewer region of interest (ROI) model using AI is shown according to one embodiment of the present disclosure, the ROI model being used to identify viewer regions of interest that may be of interest to the viewer.
[0021] Figure 3C A system for training via an AI camera view model is shown according to one embodiment of the present disclosure. The AI camera view model can be configured to identify camera views of scenes that have been previously identified by a region of interest AI model as regions of interest of interest to the viewer, wherein the identified camera views are also identified as regions of interest to the viewer.
[0022] Figure 3D A system for training a broadcast / narration model using AI according to one embodiment of the present disclosure is shown, the broadcast / narration model being configured to generate narration for scenes previously identified by a region of interest AI model as regions of interest of interest to the audience, wherein the narration can be customized for camera views identified by an AI camera view model.
[0023] Figure 4 A user interface configured for selecting camera views and regions of interest (ROIs) for a game session of a game application, according to one embodiment of the present disclosure, is shown. The user interface includes one or more ROIs identified by an ROI AI model as being of interest to the viewer, and one or more camera views generated for those identified ROIs, the camera views being identified as being of interest to the viewer using an AI camera view model.
[0024] Figure 5This is a flowchart illustrating the steps in a method for identifying a viewer's potential region of interest using one or more AI models, according to one embodiment of the present disclosure.
[0025] Figure 6 This is a flowchart illustrating steps in a method for constructing and / or generating narration for a scene identified as a viewer's region of interest using an AI model, according to one embodiment of the present disclosure, wherein the narration may be customized for one or more camera views of the scene, which is also identified using an AI model.
[0026] Figure 7 The invention illustrates the generation of highlight segments of a game session of one or more players, including a game application, according to an embodiment of the present disclosure. The highlight segments are generated using an AI model that can identify regions of interest (ROIs) in multiple game sessions of multiple game applications to identify interesting camera angles in those ROIs and provide narration for the scenes in those ROIs.
[0027] Figure 8 Components of an exemplary apparatus that can be used to carry out various embodiments of the present disclosure are shown. Detailed Implementation
[0028] While the following detailed description contains numerous specific details for illustrative purposes, those skilled in the art will recognize that many variations and modifications of these details are within the scope of this disclosure. Therefore, in clarifying the various aspects of this disclosure described below, the generality of the appended claims is not diminished, and no limitation is imposed on these claims.
[0029] Generally, various embodiments of this disclosure describe systems and methods for generating one or more renderings of scenes for a game application using artificial intelligence (AI), and further relate to generating narration and / or generating scene renderings for game application scenes using AI. In some embodiments, one or more renderings of the game are generated by artificial intelligence (AI). The AI can control camera angles to capture the most interesting actions and can clip between angles to merge actions that do not fit into a single view and show actions from different perspectives. The AI can consider many aspects of the game and many aspects that players are interested in to determine the actions that are most interesting to the audience. Different streams can be generated for different audiences. The AI can include action replays in its streams so that viewers can see exciting actions again from different perspectives and see actions that did not occur in the live camera view of the game. The AI can select camera angles to include a complete view of suspicious actions, such as simultaneously showing the angles of both the sniper and the target the sniper is aiming at. The AI can add commentary to the game streams it generates. In some embodiments, the AI commentary has a compelling personality. The commentary can convey to the audience why the AI considers the view it has chosen important. For example, if the AI notices a new player is in a position that will eliminate an experienced player, the AI can point out this fact and entice viewers to see if the new player can actually use this situation to eliminate the more experienced player. In some implementations, the AI providing commentary can interpret the actions of esports games, making them understandable and exciting for visually impaired viewers. In some implementations, the AI providing commentary generates a stream of commentary aimed at viewers in multiple locations (i.e., taking into account local and / or regional customs), and may include commentary in different languages (including sign language). In some implementations, highlight reporting is available after the game ends (such as when generating highlight clips). Such reporting can be automatically generated based on AI selection. Such reporting can be based on what viewers choose to watch or rewatch while watching a live stream of gameplay. In yet another implementation, game reporting may include the ability to view a map of the game area. Such a map may include metadata such as the positions of all players still in the game and the locations where players are eliminated. One way to highlight player eliminations on the map view is to flash a bright red dot at the location where the elimination occurred.
[0030] Throughout this specification, the references to "video game" or "game application" are intended to refer to any type of interactive application that is initiated by executing input commands. For illustrative purposes only, interactive applications include applications for games, word processing, video processing, video game processing, etc. Furthermore, the terms video game and game application are used interchangeably.
[0031] Building upon the above general understanding of the various implementation schemes, exemplary details of the implementation schemes will now be described with reference to the accompanying drawings.
[0032] Figure 1A A system 10 according to one embodiment of the present disclosure is illustrated, the system being used for AI-controlled generation of camera views of scenes in a game session of a video game, and for narration of the scene and / or AI-controlled generation of camera views of the scene. System 10 can also be configured to train an AI model for generating camera views of a scene in a game session, and for generating narration for the scene and / or generating camera views for the scene.
[0033] like Figure 1A As shown, the video game (e.g., game logic 177) can be executed locally on user 5's client device 100, or the video game (e.g., game logic 277) can be executed at a backend game execution engine 211, which operates at a backend game server 205 of a cloud gaming network or game cloud system. Game execution engine 211 can operate within one of a plurality of game processors 201 of game server 205. Furthermore, the video game can be executed in single-player or multiplayer mode, wherein embodiments of the invention provide multiplayer enhancements (e.g., assistance, communication, game segment generation, etc.) for both operating modes. In either case, the cloud gaming network is configured to generate one or more renders of the video game scene via artificial intelligence (AI), and to generate narration and / or scene rendering of the video game scene via AI. More specifically, the cloud gaming network of system 10 can be configured to construct one or more AI models for generating one or more renders of a scene captured through an identified camera perspective and for generating narration of the scene and / or the identified camera perspective.
[0034] In some implementations, cloud gaming network 210 may include multiple virtual machines (VMs) running on a host hypervisor, wherein one or more VMs are configured to utilize the hardware resources available to the host hypervisor supporting single-player or multi-player video games to execute game processor module 201. In other implementations, cloud gaming network 210 is configured to support multiple local computing devices supporting multiple users, wherein each local computing device can execute an instance of a video game, such as in a single-player or multi-player video game. For example, in multi-player mode, while the video game is running locally, the cloud gaming network simultaneously receives information (e.g., game state data) from each local computing device and distributes that information accordingly among one or more of the local computing devices, enabling each user to interact with other users within the game environment of the multi-player video game (e.g., through a corresponding character in the video game). In this way, the cloud gaming network coordinates and combines the gameplay of each user within the multi-player game environment.
[0035] As shown in the figure, system 10 includes a game server 205 that executes a game processor 201, which provides access to multiple interactive video games. Game server 205 can be any type of server computing device available in the cloud and can be configured to run one or more virtual machines on one or more hosts, as previously described. For example, game server 205 can manage virtual machines that support game processor 201. Game server 205 is also configured to provide additional services and / or content to user 5.
[0036] Client device 100 is configured to request access to a video game via network 150 (such as the Internet) and to render an instance of a video game or game application executed by game server 205 and delivered to display device 12 associated with user 5. For example, user 5 may be interacting with an instance of a video game executed on game processor 201 via client device 100. As previously described, client device 100 may also include a game execution engine 111 configured for local execution of the video game. Client device 100 can receive input from various types of input devices, such as game controller 6, tablet computer 11, and keyboard, as well as gestures captured by a camera, mouse, touchpad, etc. Client device 100 can be any type of computing device that has at least a memory and processor module capable of connecting to game server 205 via network 150. Some examples of client device 100 include personal computers (PCs), game consoles, home theater systems, general-purpose computers, mobile computing devices, tablet computers, telephones, or any other type of computing device that can interact with game server 205 to execute an instance of a video game.
[0037] The game logic of video games is built upon a game engine or game name processing engine. The game engine includes core functionalities that the game logic can use to construct the game environment of the video game. For example, some functionalities of a game engine may include a physics engine for simulating physical forces and collisions on objects in the game environment, a rendering engine for 2D or 3D graphics, collision detection, sound, animation, artificial intelligence, networking, streaming, etc. In this way, the game logic does not have to be built from scratch using the core functionalities provided by the game engine. Whether on the client-side or at a cloud gaming server, the game logic, combined with the game engine, is executed by the CPU and GPU, which can be configured within Accelerated Processing Units (APUs). That is, the CPU and GPU, along with shared memory, can be configured as a rendering pipeline to generate video frames for game rendering, such that the rendering pipeline outputs the game-rendered images as video or image frames suitable for display, including corresponding color information for each pixel in the target and / or virtualized display.
[0038] Network 150 may include one or more generations of network topologies. As newer topologies come online, increased bandwidth is provided to handle higher resolution rendered images and / or media. For example, 5G networks (fifth generation) providing broadband access via digital cellular networks will replace 4G LTE mobile networks. 5G wireless devices connect to local cells via local antennas. These antennas are then connected to telephone networks and the internet via high-bandwidth connections (e.g., fiber optics). 5G networks are expected to operate at higher speeds and with better quality of service. Therefore, higher resolution video frames and / or media can be provided in connection with video games through these newer network topologies.
[0039] Client device 100 is configured to receive and display rendered images on display 12. For example, via a cloud-based service, the rendered image may be delivered to user 5 in association with an instance of a video game executed on game execution engine 211 of game server 205. In another example, via local game processing, the rendered image may be delivered by local game execution engine 111 (e.g., for playing or watching). In either case, client device 100 is configured to interact with execution engine 211 or 111 in association with user 5's game playing process, such as through input commands to drive the game playing process.
[0040] Furthermore, client device 100 is configured to interact with game server 205 to capture and store metadata of user 5's gameplay while playing video games, wherein each piece of metadata includes information related to the gameplay process (e.g., game state, etc.). More specifically, game processor 201 of game server 205 is configured to generate and / or receive metadata of user 5's gameplay while playing video games. For example, metadata may be generated by a local game execution engine 111 on client device 100, and the output may be sent to game processor 201 via network 150. Alternatively, metadata may be generated by game execution engine 211 within game processor 201, such as by an instance of a video game executed on engine 211. Additionally, other game processors of game server 205 associated with other virtual machines are configured to execute instances of video games associated with gameplay processes of other users and capture metadata during those gameplay processes.
[0041] As described below, according to embodiments of this disclosure, metadata can be used to train one or more AI models, which are configured for generating camera views of scenes in a game session of a video game and for AI control generation of scene narration and / or scene camera views.
[0042] More specifically, the metadata also includes game state data that defines the state of the game at that point. For example, game state data may include game characters, game objects, game object attributes, game properties, game object states, graphics overlays, etc. In this way, game state data allows for the generation of a game environment that exists at the corresponding point in the video game. Game state data may also include the state of each device used to render the gameplay process, such as the state of the CPU, GPU, memory, register values, program counter values, programmable DMA state, DMA buffer data, audio chip state, CD-ROM state, etc. Game state data can also identify which parts of the executable code need to be loaded to start executing the video game from that point. Not all game state data needs to be captured and stored; only the data sufficient to allow the executable code to start the game at the corresponding point needs to be captured and stored. Game state data may be stored in the game state database 145 of data storage area 140. In some embodiments, some game state data may be stored only in memory.
[0043] Metadata also includes user-saved data. Typically, user-saved data includes information that personalizes the video game for the corresponding user. This includes information associated with the user's character, allowing the video game to be rendered using a character that is likely unique to that user (e.g., shape, appearance, clothing, weapons, etc.). In this way, user-saved data enables the generation of a character for the corresponding user's gameplay, where the character has states corresponding to points in the video game associated with the metadata. For example, user-saved data may include the game difficulty, game level, character attributes, character position, remaining lives, total possible lives, armor, prizes, time counter values, and other asset information selected by user 5 while playing the game. For example, user-saved data may also include user profile data that identifies user 5. The user-saved data is stored in database 141 of data storage area 140.
[0044] Additionally, the metadata includes random seed data that can be generated using artificial intelligence. This random seed data may not be part of the original game code but can be added to the overlay to make the game environment appear more realistic and / or more appealing to the user. In other words, random seed data provides additional features to the game environment at corresponding points in the user's gameplay. For example, AI characters can be randomly generated and provided in the overlay. These AI characters are not associated with any user playing the game but are placed in the game environment to enhance the user experience. Game state data can include learning performed by the AI, such as allowing the AI to remember previous interactions with the player so that these interactions can influence future decisions made by the AI. For example, these AI characters might randomly walk on streets in a city scene. Furthermore, other objects can be generated and rendered in the overlay. For example, clouds in the background and birds flying through space can be generated and rendered in the overlay. Random seed data is stored in a random seed database 143 in data storage area 140.
[0045] In this way, system 10 can be used to train one or more AI models configured for generating camera views of scenes in a video game session and for AI-controlled generation of scene narration and / or scene camera views. Specifically, metadata collected during the execution of one or more video games during a game session or during replay of a game session and a recorded replay of a game session can be used to train one or more AI models. For example, a cloud gaming system including game server 205 can be configured to use the deep learning engine 190 of AI server 260 to train a statistical information and fact generation model, an AI camera view model 321, and an AI narration model 331. Moreover, the deep learning engine can be used to identify regions of interest (ROIs) that viewers may be interested in. The identified ROIs are used to generate camera views and narrations that viewers may want to watch using artificial intelligence.
[0046] Therefore, the cloud gaming system can be accessed later by viewers to generate a real-time rendered view of the scene from the current and / or previously played game session (e.g., from a camera perspective) and provide commentary for that scene, wherein the rendered view and commentary are generated using artificial intelligence. Specifically, the camera perspective engine 230 is configured to implement an AI camera perspective model 321 to generate a rendered view of the scene from the corresponding game session, wherein the rendered view is generated from one or more camera perspectives identified as potentially of great interest to the viewer. Moreover, the broadcaster / commentator engine 220 is configured to implement an AI commentary model 331 to generate comments and / or commentary on the scene and / or the camera perspective of the scene, wherein the comments are generated using templates as well as statistics and facts, which are generated using an AI statistics / facts model 341 based on metadata collected for the scene during the game session.
[0047] Figure 1B An exemplary neural network for constructing an artificial intelligence (AI) model according to one embodiment of this disclosure is shown, comprising an AI camera view model and an AI narration model. In this manner, given one or more game sessions, including single-player and multiplayer sessions, as input, the AI model is configured to identify regions of interest (ROIs) for scenes in those game sessions that the audience may be potentially interested in. The AI model is configured to generate views or camera perspectives of those scenes that the audience may be interested in, and / or generate separately provided narration or supplementary camera views or camera perspectives of these scenes. Thus, the AI model can be configured to generate rendered views (e.g., camera perspectives) of scenes in a video game game session, and generate narration and / or camera perspectives of the scenes, wherein the rendered views and / or narration can be streamed to the audience during or after the game session.
[0048] More specifically, according to one embodiment of this disclosure, a deep learning engine 190 is used to train and / or build an AI camera view model and an AI narration model. In one embodiment, the neural network 190 may be implemented within an AI server 260 at a backend server. Specifically, the deep learning engine 190 (e.g., using an AI modeler) is configured to learn which regions of interest (ROIs) of a scene in a game session are of most interest to the viewer, wherein, while training the corresponding AI model, the scene of the virtualized game environment within the game session can be viewed from one or more camera views. For example, the ROI might focus on the expert, or it could be defined as a David and Goliath situation where a beginner has the opportunity to deliver a fatal blow to the expert, or it could be defined as a situation where the expert aims at the victim with his or her scope. Moreover, the deep learning engine 190 may be configured to learn, while training the AI camera view model 321, which views of the scene corresponding to the identified ROIs are of most interest to the viewer. The views can be defined by one or more camera angles acquired within the game environment showing the scene. Furthermore, the deep learning engine 190 can be configured to learn which statistics and facts are of interest to the audience when constructing commentary for scenes corresponding to identified regions of interest, wherein the commentary can be provided separately or supplemented to the view of the scene corresponding to the identified regions of interest during streaming to the audience. In this way, one or more viewers can watch the most exciting scenes of a specific game session supported by commentary constructed to generate maximum interest for the viewers, wherein a view can be provided in real-time for a live game session, or a view can be generated for a replay of the corresponding game session, which can be streamed during or after the game session.
[0049] Specifically, a deep learning or machine learning engine 190 (e.g., collaborating with an AI modeler, not shown) is configured to analyze training data collected during one or more game sessions of one or more video games, during audience viewing of one or more game sessions, during replays of one or more game sessions, during live broadcasts with live commentary of one or more game sessions, etc. During the learning and / or modeling phases, the deep learning engine 190 utilizes artificial intelligence, including deep learning algorithms, reinforcement learning, or other AI-based algorithms, to build one or more trained AI models (e.g., generating significant interest, etc.) using the training data and success criteria. The trained AI models are configured to predict and / or identify regions of interest (ROIs) corresponding to scenes of significant interest to the audience during the game session, identify views and camera angles that can be generated for those scenes of significant interest to the audience, and identify or generate statistics and facts for constructing commentary corresponding to the identified ROIs and the identified camera angles. The deep learning engine 190 can be configured to continuously improve the trained AI models given any updated training data. The improvement is based on determining which training datasets can be used for training, and this determination is based on how these sets perform within the deep learning engine 190 according to the corresponding success criteria.
[0050] Additional training data can be collected and / or analyzed from human input. For example, experts in the field of broadcasting professional gaming events can provide highly valuable input data (i.e., data associated with high success criteria). These experts can help define which regions of interest (ROIs) should be tracked during a gaming session, which views and / or camera perspectives should be generated for those identified ROIs, and what statistics and facts should be identified and / or generated to construct the scene commentary corresponding to the identified ROIs and the identified camera perspectives.
[0051] The final AI model for the video game can be used by a camera view engine 230, a broadcaster / narrator engine 220, and / or a highlight engine 240 to generate the most exciting scenes of a particular game session given any set of input data (e.g., a game session), which is supported by narration built to generate maximum interest for the audience. This can provide a view in real time for a live game session or generate a view for a replay of the corresponding game session, which can be streamed during or after the game session.
[0052] Neural network 190 represents an example of an automated analysis tool used to analyze datasets to determine audience regions of interest, views and / or camera angles of the scene corresponding to the identified audience regions of interest, and explanations including statistical information and facts collected or generated during the corresponding game session that are of great interest to one or more audiences. Different types of neural networks 190 are possible. In one example, neural network 190 supports deep learning, which can be implemented by a deep learning engine 190. Thus, deep neural networks, convolutional deep neural networks, and / or recurrent neural networks trained in a supervised or unsupervised manner can be implemented. In another example, neural network 190 includes a deep learning network that supports reinforcement learning or reward-based learning (e.g., through the use of success criteria, success metrics, etc.). For example, neural network 190 is configured to support Markov decision processes (MDPs) using reinforcement learning algorithms.
[0053] Typically, a neural network (190) represents a network of interconnected nodes, such as an artificial neural network. Each node learns some information from data. Knowledge can be exchanged between nodes through interconnections. The input to a neural network (190) activates a set of nodes. In turn, this set of nodes activates other nodes, thereby propagating knowledge about the input. This activation process is repeated on other nodes until an output is provided.
[0054] As shown in the figure, neural network 190 includes a hierarchical structure of nodes. At the lowest level of the hierarchy, there is an input layer 191. Input layer 191 includes a set of input nodes. For example, each of these input nodes is mapped to an instance of playing a video game, wherein the instance includes one or more features defining the instance (e.g., controller input, game state, outcome data, etc.). The intermediate predictions of the model are determined by a classifier that creates labels (e.g., output, features, nodes, classification, etc.).
[0055] At the highest hierarchical level, there exists an output layer 193. Output layer 193 comprises a set of output nodes. For example, output nodes represent decisions (e.g., regions of interest, camera viewpoints, statistics and facts about a given set of input data, etc.) related to one or more components of the trained AI model 160. As previously described, output nodes can identify predicted or anticipated actions or learned actions for a given set of inputs, where the inputs can be one or more single-player or multi-player game sessions of one or more video games. These results can be compared with predetermined and real results or learned actions and outcomes (e.g., human inputs driving successful outcomes, etc.) obtained from current and / or previous game sessions and / or broadcasts of those game sessions used to collect training data, in order to improve and / or modify the parameters used by the deep learning engine 190 to iteratively determine the appropriate responses and / or actions predicted or anticipated for a given set of inputs. In other words, the nodes in the neural network 190 learn the parameters of a trained AI model (e.g., an AI camera view model 321, an AI interpretation model 331, a statistical information and fact model 341, etc.), which can be used to make such decisions when improving the parameters.
[0056] Specifically, hidden layer 192 exists between input layer 191 and output layer 193. Hidden layer 192 comprises "N" hidden layers, where "N" is an integer greater than or equal to one. Each hidden layer, in turn, includes a set of hidden nodes. Input nodes are interconnected with hidden nodes. Similarly, hidden nodes are interconnected with output nodes, such that input nodes are not directly interconnected with output nodes. If multiple hidden layers exist, input nodes are interconnected with the hidden nodes of the lowest hidden layer. These, in turn, are interconnected with the hidden nodes of the next highest hidden layer, and so on. The hidden nodes of the next highest hidden layer are interconnected with the output node. Interconnections link two nodes. Interconnections have learnable numerical weights, allowing neural network 190 to adapt to the input and learn.
[0057] Typically, hidden layer 192 allows knowledge about the input node to be shared among all tasks corresponding to the output node. To this end, in one implementation, a transformation f is applied to the input node via hidden layer 192. In one example, the transformation f is non-linear. Different non-linear transformations f can be used, for example, including the rectified function f(x) = max(0,x).
[0058] Neural Network 190 also uses a cost function c to find the optimal solution. The cost function measures the deviation between the prediction output by Neural Network 190 and the ground truth or target value y (e.g., the expected outcome) for a given input x, defined as f(x). The optimal solution represents the situation where no solution has a cost lower than the cost of the optimal solution. For such ground truth labels, an example of a cost function is the mean squared error between the prediction and the ground truth. During the learning process, Neural Network 190 can use backpropagation to learn model parameters (e.g., the weights of the interconnections between nodes in hidden layer 192) that minimize the cost function, employing different optimization methods. An example of such optimization methods is stochastic gradient descent.
[0059] In one example, the training dataset for neural network 190 can come from the same data domain. For instance, neural network 190 is trained to learn predicted or expected responses and / or actions to be performed for a given set of inputs or input data (e.g., a game session). In this illustration, the data domain includes game session data collected through multiple gameplay sessions by multiple users, audience views of the game session, human input provided by experts selecting regions of interest, camera perspectives, and statistics and facts used to define the baseline input data. In another example, the training dataset comes from a different data domain to include input data outside the baseline. Based on these predictions, neural network 190 can also define a trained AI model for determining those outcomes and / or actions to be performed given a set of inputs (e.g., a game session) (e.g., predicted regions of interest, camera perspectives, and statistics and facts used to construct the commentary). Therefore, the AI model is configured to generate the most exciting scenes of a particular game session given any set of input data (e.g., a game session), which is supported by commentary built to generate maximum interest for the audience, where a view can be provided in real time for a live game session, or a view can be generated for a replay of the corresponding game session, where the replay can be streamed during the game session or after the game session ends.
[0060] Figure 2A system 200, according to one embodiment of this disclosure, provides game control to one or more users playing one or more video games, said one or more video games being executed locally on the corresponding user or at a backend cloud gaming server within one or more game sessions of said one or more video games. The one or more game sessions are used to train an AI camera view model and an AI narration model using artificial intelligence. The AI model is used to identify the camera view of the game session and provide narration for scenes in those game sessions most likely to be viewed by a viewer. In other embodiments, the AI model is used to generate rendered views of scenes in single-player or multi-player game sessions using AI-recognized statistics and facts that are then streamed to one or more viewers, said game sessions accompanied by narration. In one embodiment, system 200 and... Figure 1A The system 10 works together to train and / or implement AI models configured to generate camera perspectives for game sessions and to provide narration for scenes in those game sessions most likely to be viewed by the audience. Referring now to the accompanying drawings, the same reference numerals denote the same or corresponding parts.
[0061] like Figure 2 As shown, multiple users 115 (e.g., user 5A, user 5B, ..., user 5N) play multiple video games in one or more game sessions. Each of the game applications can be executed locally on the corresponding user's client device 100 (e.g., a game console) or at the backend cloud gaming system 210. Additionally, each of the multiple users 115 accesses a display 12 or device 11, each configured to display rendered images of the game session used for training, or rendered images of interesting scenes from the current or previous game sessions constructed using an AI model.
[0062] Specifically, according to one embodiment of this disclosure, system 200 provides game control to multiple users 115 playing one or more locally executed video games. For example, user 5A can play a first video game on a corresponding client device 100, wherein an instance of the first video game is executed by corresponding game logic 177A (e.g., executable code) and game name execution engine 111A. For illustrative purposes, the game logic can be delivered to the corresponding client device 100 via portable media (e.g., flash drive, optical disc, etc.) or via a network (e.g., downloaded from a game provider via the Internet 150). Furthermore, user 115N plays an Nth video game on the corresponding client device 100, wherein an instance of the Nth video game is executed by corresponding game logic 177N and game name execution engine 111N.
[0063] Additionally, according to one embodiment of this disclosure, system 200 provides game control to multiple users 115 playing one or more video games executed on cloud gaming system 210. Cloud gaming system 210 includes game server 205, which provides access to multiple interactive video games or game applications. Game server 205 can be any type of server computing device available in the cloud and can be configured to execute one or more virtual machines on one or more hosts. For example, game server 205 can manage virtual machines that support game processors, which are instances of game applications instantiated by users. Thus, multiple game processors associated with multiple virtual machines on game server 205 are configured to execute multiple instances of game applications associated with the gameplay of multiple users 115. In this way, the backend server supports providing media streams (e.g., video, audio, etc.) of the gameplay of multiple video games to multiple corresponding users (e.g., players, viewers, etc.). In some embodiments, the cloud gaming network can be a game cloud system 210, which includes multiple virtual machines (VMs) running on a host hypervisor, wherein one or more VMs are configured to utilize hardware resources available to the host hypervisor to execute game processors. One or more users can access the game cloud system 210 via network 150 to remotely process video games using client device 100 or 100', wherein client device 100' can be configured as a thin client (e.g., including a game execution engine 211) to interface with a backend server providing computing functionality. For example, user 5B plays a second video game on the corresponding client device 100', wherein an instance of the second video game is executed by the corresponding game name execution engine 211 of the cloud gaming system 200. Game logic (e.g., executable code) that implements the second video game is executed in cooperation with the game name processing engine 211 to execute the second video game. User 5B has at least accessed device 11 configured to display rendered images.
[0064] Client devices 100 or 100' can receive input from various types of input devices (such as game controllers, tablets, keyboards), and gestures captured by cameras, mice, touchpads, etc. Client devices 100 or 100' can be any type of computing device with at least a memory and processor module capable of connecting to game server 205 via network 150. Furthermore, client devices 100 and 100' corresponding to the user are configured to generate rendered images executed by a game name execution engine 111, either locally or remotely, and to display these rendered images on a monitor. For example, client devices 100 or 100' are configured to interact with an instance of a corresponding video game executed locally or remotely to enable the corresponding user to play the game, such as through input commands used to drive the gameplay process.
[0065] In one implementation, the client device 100 operates in single-player mode for the corresponding user who is playing a game application.
[0066] In another implementation, multiple client devices 100 operate in multiplayer mode for their respective users playing a specific game application. In that case, multiplayer functionality can be provided via backend server support, such as through a multiplayer processing engine 119, via game server 205. Specifically, the multiplayer processing engine 119 is configured to control multiplayer game sessions for a specific game application. For example, the multiplayer processing engine 130 communicates with a multiplayer session controller 116, which is configured to establish and maintain a communication session with each user and / or player participating in the multiplayer game session. In this way, users in the session can communicate with each other under the control of the multiplayer session controller 116.
[0067] Furthermore, the multiplayer processing engine 119 communicates with the multiplayer logic 118 to enable user interaction within each user's corresponding game environment. Specifically, the state sharing module 117 is configured to manage the state of each user in a multiplayer game session. For example, state data may include game state data, which defines the state of the corresponding user's gameplay process at a specific point in the game application. For example, game state data may include game characters, game objects, game object attributes, game properties, game object states, graphics overlays, etc. In this way, game state data allows the generation of game environments that exist at corresponding points in the game application. Game state data may also include the state of each device used to render the gameplay process, such as the state of the CPU, GPU, memory, register values, program counter values, programmable DMA state, DMA buffer data, audio chip state, CD-ROM state, etc. Game state data may also identify which parts of the executable code need to be loaded to start executing the video game from that point. Game state data may be stored in... Figure 1A It is located in database 140 and can be accessed by state sharing module 117.
[0068] Furthermore, the state data may include user-saved data, which includes information that personalizes the video game for the corresponding player. This includes information associated with the character played by the user, allowing the video game to be rendered using a character that is likely unique to that user (e.g., location, shape, appearance, clothing, weapons, etc.). In this way, the user-saved data enables the generation of a character for the corresponding user's gameplay, wherein the character has a state corresponding to the point the corresponding user is currently experiencing in the game application. For example, the user-saved data may include the game difficulty selected by the corresponding user, game level, character attributes, character position, remaining lives, total possible lives, armor, prizes, time counter values, etc. The user-saved data may also include user profile data that identifies the corresponding user. The user-saved data may be stored in database 140.
[0069] In this way, the multiplayer processing engine 119, using state-shared data 117 and multiplayer logic 118, can overwrite / insert objects and roles into each of the game environments of users participating in a multiplayer game session. For example, the first user's role is overwritten / inserted into the second user's game environment. This allows users in a multiplayer game session to interact via each of their respective game environments (e.g., as displayed on the screen).
[0070] Additionally, the backend server of game server 205 supports the generation of AI-rendered views (e.g., camera perspectives) of scenes within a video game session, and the generation of scene commentary and / or scene camera perspectives, wherein the rendered views and / or commentary can be streamed to viewers during or after the game session. In this way, one or more viewers can watch the most exciting scenes of a specific game session, supported by commentary built to generate maximum viewer interest, where views can be provided in real-time for a live game session, or views can be generated for replays of the corresponding game session using an AI model implemented by AI server 260. Specifically, the AI model is used to identify regions of interest (ROIs) within one or more game sessions that viewers may be potentially interested in. For a specific identified ROI corresponding to a scene in a video game session, the camera perspective engine uses an AI camera perspective model to generate or identify one or more camera perspectives of the scene of interest to the viewer. Furthermore, for a specific identified ROI corresponding to a scene, the broadcaster / commentator engine 220 uses an AI commentary model to identify and / or generate statistics and facts based on game state data and other information about the game session corresponding to the scene. Comment templates are used to weave together generated statistics and facts to produce comments, which can be streamed individually or streamed together with one or more camera views of the scene to one or more viewers.
[0071] Furthermore, the highlight engine 240 of the cloud gaming system 210 is configured to use an AI model to generate AI-rendered views of the current or previous game sessions to generate one or more rendered views (e.g., camera views) corresponding to one or more scenes in one or more game sessions and to provide narration for these scenes. Highlight clips (including AI-rendered views and AI-generated narration) can be generated by the highlight engine 240 and streamed to one or more viewers.
[0072] Figure 3A A system 300A according to an embodiment of this disclosure is illustrated, the system being configured to train one or more AI models to identify regions of interest (ROIs) for viewers in multiple game sessions of multiple video games, identify interesting camera perspectives of those identified ROIs, and generate commentary and / or narration for scenes of those identified ROIs. System 300 is also configured to use the AI models to generate rendered views (e.g., camera perspectives) of scenes identified by the AI in game sessions of the video games, and to generate narration and / or narration of the scenes, wherein the rendered views and / or narration may be streamed to viewers during or after the corresponding game session. In some embodiments, one or more renders of the game are generated by artificial intelligence (AI), such as using an AI-trained camera perspective model 321. In one embodiment, the AI camera perspective model 321 can control camera angles to capture the most interesting actions and can clip between angles to merge actions that are not suitable for a single view and show actions from different perspectives. The AI camera perspective model 321 can select camera angles to include a complete view of suspicious actions, such as simultaneously showing the angles of both the sniper and the target the sniper is aiming at. In another implementation, the AI region of interest model 311 can consider many aspects of the game and the many aspects that players are interested in to determine the actions that the audience is most interested in. Different streams can be generated for different viewers. In yet another implementation, the AI commentary model 331 can add commentary to the game stream it generates. In yet another implementation, the highlight engine can include action replays in its stream, allowing viewers to see exciting actions again from different perspectives and to see actions that did not occur in the game's live camera view.
[0073] For training, system 300A receives metadata, player statistics, and other information as input from one or more game sessions of one or more video games, where the game sessions can be single-player or multi-player game sessions. The one or more game sessions can be generated from one or more players playing the video game (e.g., for supervised learning purposes), or from one or more AI players playing the video game (e.g., for unsupervised learning), or from a combination of one or more human players and / or one or more AI players. Additionally, human input can be provided as training data, such as expert choices and / or definitions of viewer regions of interest (ROIs), camera perspectives, and / or statistics and facts used for narration. For example, player 1's game space 301A is associated with a game session of the corresponding video game, where metadata 302A and player statistics 303A are provided as input to the AI viewer ROI trainer 310 to train the viewer ROI model 311 using artificial intelligence. Similar data is collected from player 2's game space 301B and up to player N's game spaces 301N and provided as input. Therefore, given a set of inputs (e.g., one or more game sessions), the viewer region of interest model 311 is trained to output one or more identified regions of interest 312 that are potentially of great interest to the viewer. Figure 3B The document provides a more detailed discussion of the training and implementation of the AI audience region of interest trainer 310 and region of interest model 311.
[0074] Each of the identified regions of interest (ROIs) is associated with a corresponding scene in a video game performed for a game session. Different camera perspectives can be generated for the game (e.g., for the player) or newly generated using artificial intelligence leveraging game state data. One advantage of generating a new camera angle rendering at a later point in time is that specific actions can be identified to be shown in playback, and then the camera angle that best shows that action can be calculated, and then a rendering of the action from that camera angle can be generated. This allows actions to be played back from a more favorable camera angle than any camera angle identified when the action occurred (e.g., initially generated by the game). The newly identified and generated camera angles can then be shown in the playback. This can be useful, for example, when events that produce actions of interest initially do not seem interesting, i.e., when these events occur before they are known to produce interesting actions (e.g., actions of interest), in being implemented to show these events (e.g., during playback). Typically, the AI camera perspective trainer 320 is configured to receive one or more camera perspectives of the identified ROIs output from the ROI model 311 as input. Based on success criteria, the AI camera view trainer is configured to train a camera view model 321 to output at least one identified view or camera view 322 for a corresponding scene of a selected and / or identified viewer region of interest, wherein one or more viewers are of great interest to the camera view. Figure 3C The document provides a more detailed discussion of the training and implementation of the AI camera view trainer 320 and the camera view model 321.
[0075] Furthermore, for each scene corresponding to the identified audience interest region, artificial intelligence can be used to construct commentary. Specifically, the AI broadcaster / commentator trainer 330 receives one or more statistical information and facts as input from the corresponding game session. Based on success criteria, the AI broadcaster / commentator trainer 330 is configured to train the commentary model 331 to identify statistical information and facts that are of great interest to the audience, and to weave those identified statistical information and facts into commentary and / or commentary, such as using commentary templates 332. Figure 3D The document provides a more detailed discussion of the training and implementation of the AI broadcaster / narrator trainer 330 and the narration model 331.
[0076] Figure 3B A system 300B according to one embodiment of the present disclosure is shown for training a viewer region of interest (ROI) model 311 using artificial intelligence, the ROI model being used to identify regions of interest that a viewer may be potentially interested in. System 300B may be implemented within a cloud gaming system 210, and more specifically within an AI server 260 of the cloud gaming system 210.
[0077] For training, system 300B receives metadata, player statistics, and other information as input from one or more game sessions of one or more video games, where the game sessions can be single-player or multi-player game sessions. The one or more game sessions can be generated from one or more players playing the video game (e.g., for supervised learning purposes), or from one or more AI players playing the video game (e.g., for unsupervised learning), or from a combination of one or more human players and / or one or more AI players. For example, player 1's game space 301A is associated with a corresponding video game game session, where metadata 302A and player statistics 303A are provided as input to the AI audience region of interest trainer 310 to train the audience interest model 311 using artificial intelligence. Similarly, player 2's game space 301B is associated with a corresponding video game game session, where metadata 302B and player statistics 303B are provided as input to the AI audience region of interest trainer 310 to train the audience interest model 311. Additional input data is provided. For example, player N's game space 301N is associated with the game session of the corresponding video game, where metadata 302N and player statistics 303N are provided as input to the AI audience region of interest trainer 310 to train the audience interest model 311. Additionally, human input can be provided as training data, such as expert choices and / or the definition of the audience region of interest, camera perspectives, and / or statistics and facts used for narration.
[0078] As previously mentioned, metadata is collected during the corresponding game session used for training and may include game state and other user information. For example, game state data defines the state of the game at a point in the game environment used to generate the video game and may include game characters, game objects, game object attributes, game properties, game object states, graphics overlays, etc. Game state may include the state of each device used to render the gameplay (e.g., CPU, GPU, memory, etc.). For example, game state may also include random seed data generated using artificial intelligence. Game state may also include user-saved data, such as data used to personalize the video game. For example, user-saved data can personalize a particular user's character, allowing the video game to be rendered with a character (e.g., shape, appearance, clothing, weapons, etc.) that is likely unique to that user. User-saved data may include user-defined game difficulty, game levels, character attributes, character positions, remaining lives, asset information, etc. User-saved data may include user profile data.
[0079] Furthermore, player statistics may include user (e.g., player 1) profile information, or other statistics typically applicable to player 1, such as player name, games played, game types played, skill level, etc. Player statistics may also include or be based on user profile data.
[0080] As previously described, the region of interest trainer 310 implements the deep learning engine 190 to perform analysis using artificial intelligence training data collected during one or more game sessions of one or more video games, during audience viewing of one or more game sessions, during replays of one or more game sessions, during live broadcasts with live commentary of one or more game sessions, human input, etc., to identify regions of interest that are of great interest to the audience. Specifically, the deep learning engine 190 utilizes artificial intelligence (including deep learning algorithms, reinforcement learning, or other AI-based algorithms) during the learning and / or modeling phases to construct a region of interest model 311 using training data and success criteria.
[0081] For example, a feedback loop can be provided to train the region of interest (ROI) model 311. Specifically, with each iteration, the feedback loop can help improve the success criteria used by the ROI trainer 310 when outputting one or more identified ROIs 312. As shown, the AI ROI trainer 310 can provide one or more potential ROIs as outputs for a given set of inputs through the trained ROI model 311. The inputs can be associated with one or more game sessions of one or more video games. For ease of illustration, the feedback process is described in terms of a set of inputs provided from a game session (such as a single-player game session or a multi-player game session). Additionally, the game session may be live or may have already occurred. In this way, historical game data corresponding to previous game sessions can be used to train the ROI model 311 to better predict how to focus on actions in the game session and select content of interest to the audience.
[0082] Specifically, the provided output 313 may include past, current, and future regions of interest (ROIs) of the game session provided as a set of inputs. The current ROI provided as output 313 can provide real-time tracking of interesting ROIs in an ongoing game session or in a game session that has ended but is being replayed using recorded game state data. The past ROI provided as output 313 can evaluate previous gameplay in the current game session or for a game session that has ended. The future ROI provided as output 313 can evaluate potentially interesting scenarios (live or ended scenarios) that may occur in the game session, allowing the ROI model 311 to be configured to focus on actions in a live or replayed game session using game state data.
[0083] Features of regions of interest (ROIs) can be defined (e.g., through game state data or statistics in a game session) to help identify content that one or more viewers may be interested in. For example, an ROI can be defined by an event, situation, or context (e.g., David and Goliath, loser, reaching a goal, completing a level, etc.). In another example, an ROI can be defined by a scene in a video game, such as a crucial scene in the game (e.g., reaching the apex of a quest that allows a character to evolve to a higher state). In yet another example, an ROI can be defined by a portion or segment of a video game, such as reaching a specific geographic area and / or region of the game world, or when reaching a boss level where the character is about to fight a boss. In yet another example, an ROI can be defined by a specific action of interest, such as when a character is about to climb a mountain in a video game to reach a goal. In yet another example, an ROI can be defined by a schadenfreude cluster, where one or more viewers might enjoy watching a player's destruction. Other definitions are also supported in embodiments of this disclosure.
[0084] In another example, the zone of interest (ROI) can be defined by the relationship between two players in a game session. In one scenario, two skilled players (e.g., professionals) might be geographically within each other in the game world of a video game, where they are likely to meet, such as in a head-to-head match. In another scenario, a lower-level player might set up a situation where they eliminate a professional player with a single, decisive blow (e.g., David vs. Goliath or a loser scenario), where the lower-level player only has one chance before the professional eliminates them. In yet another scenario, the lower-level player is within the scope of a professional (e.g., someone with a high sniping rate). In yet another example, the ROI can be defined by the player, where players popular with one or more viewers can be tracked within the ROI. In yet another example, a player might have specific characteristics that compel viewers to watch, such as having a high sniping rate that instructs the player to seek out sniping situations that viewers might turn to. The ROI can include temporal characteristics, where a player with a high sniping rate might have a target within his or her scope. The ROI can include geographical characteristics, where a player with a high sniping rate might be within a certain area of the game world where there might be a large number of targets. Other definitions are also supported in embodiments of this disclosure.
[0085] In another example, the region of interest (ROI) can be characterized by predicted player actions or outcomes. For instance, an ROI can be defined as a region or area in a video game world that has a high level-up rate or a high character mortality rate (e.g., performing difficult tasks or traversing challenging terrain). In other instances, an ROI can be defined as game conditions (task execution) with low or high success rates. Other definitions are also supported in embodiments of this disclosure.
[0086] In other examples, priority can be given to popular players, which defines zones of interest. Furthermore, priority can also be given to players more likely to be eliminated, such as those in zones with many other players nearby, rather than hidden players without other players nearby. Players to focus on can be partially or completely determined from the most popular broadcasts for viewers. Viewers can vote directly or indirectly by directly selecting which broadcasts to watch or which players to follow.
[0087] The regions of interest provided as output 311 during training can be further filtered and fed back to the AI audience region of interest trainer 310 to help define success criteria. For example, regions of interest can be selected by humans in the loop or by active and / or passive measurements from other audiences.
[0088] Specifically, viewers' viewing choices influence which regions of interest (ROIs) are selected using the AI model. For example, each of the ROIs provided as output 313 may correspond to one or more game session segments that viewers can choose to watch. This selection may help define success criteria for what is defined as a preferred ROI that viewers are interested in. In some cases, viewers may actively choose to watch a live game session from one or more live game session options, or actively choose to watch a replay of a game session that can be selected from one or more replay options. In some implementations, viewers may set criteria for which replays or replays will be automatically shown to them as options they can choose to watch. In some cases, viewers may select content that reinforces the definition of the AI ROI model 311 by choosing to watch replay streams controlled by humans and / or by AI-selected and generated replay streams controlled by artificial intelligence.
[0089] Examples of actions 360 taken by viewers that help define which areas of interest (IOIs) viewers like (such as defined viewer-selected IIOs 315 provided as positive feedback to the AI viewer IIO trainer 310) are provided below and are intended to be illustrative rather than exhaustive. For example, a viewer jumping into a count 360A exceeding a threshold indicates that the viewer likes that IIO. Moreover, especially when there is a lot of positive feedback, positive viewer feedback 360B provided as comments or textual feedback during viewing a portion of a game session that includes the corresponding viewer IIO can help define the corresponding viewer IIO. Furthermore, active or passive viewer actions 360C during viewing a portion of a game session that includes the corresponding viewer IIO can help define the corresponding viewer IIO. For example, passive actions may include biofeedback (e.g., sweating, smiling, etc.). In another example, active actions may include selecting icons that indicate a positive interest in watching the game session. Viewers can vote directly or indirectly by directly selecting which broadcast to watch or which player to follow. For example, a viewer replay count 360D indicating the popularity of a replay of a portion of a game session that includes the corresponding viewer IIO can help define the corresponding viewer IIO. The area of interest (OI) that a viewer chooses to follow can serve as an indicator of how interesting that OI is. For example, when a viewer switches to follow an OI that is currently focused on a particular player, it can serve as an indicator of the viewer's interest in that OI. In some cases, viewers can choose their own camera angle or viewing angle. The selected camera angle can serve as an indicator of where the action of interest occurs, which can be used to determine the OI or a specific camera angle. In another example, an indicator of a live game session includes the viewer streaming view count 360E, which represents the popularity of the replay of a portion of the game session corresponding to the viewer's OI, which can help define the corresponding OI that the viewer likes. Other definitions for providing viewer feedback are also supported in embodiments of this disclosure.
[0090] Viewers can also be given the ability to simply choose reactions such as applause, boos, or yawns. These can be aggregated and displayed to players in the league, and collected to determine the level of attention viewers pay to areas of interest or actions. In some cases, viewers can provide this feedback in association with overall actions, specific actions, non-player character activities, or specific player activities. This information can be used to determine actions players want to see, which can be used to select actions shown during patches or to design game content to increase the likelihood that users will enjoy what happens during a match. In some cases, viewers can provide feedback on how well they perceive the gameplay broadcast to be completed, which can be associated with specific action choices or replays to be shown, the chosen camera angle, information included in the commentary, the wording of the commentary, the rendered appearance of the commentator, or the voice characteristics of the rendered commentary. Viewer feedback can be used by AI to aggregate and generate broadcasts that are more engaging to a wider audience. AI can use viewer feedback from a specific viewer to specifically generate broadcasts that are more tailored to that viewer's enjoyment.
[0091] Other inputs related to the game session and / or the user can also be provided. For example, human input 314 can be provided that may inherently have great value or define high success criteria when used for training. For example, human experts can be used to help identify regions of interest (ROIs) from one or more game sessions watched by the expert and provide them as feedback to the AI ROI trainer 310. In one use case, a game session can be live-streamed while an expert is watching a live stream. In another use case, the expert may be watching a game session that can be recorded or replayed using the game state. In either case, the expert can actively select and / or define preferred ROIs that, according to the definition, have or translate into high success criteria when used for training. In some implementations, the AI generates multiple renders from camera angles in one or more ROIs, which are presented to the human expert as options that can be selected to be included in the broadcast. Renders from human-controlled camera angles can also be presented to the human expert, which can also be selected. In some cases, the human expert may provide feedback to the AI to influence the camera angles of the renders presented by the AI, such as zooming in or out, or panning to show different angles.
[0092] In one implementation, the weighting module 350 provides weights for the feedback, which are provided as human-in-loop selected region of interest 314 and audience-selected region of interest 315.
[0093] Once trained, the AI region of interest model 311 is able to identify one or more regions of interest for a given input. For example, when providing a game session (e.g., single-player or multiplayer), the AI region of interest model 311 provides one or more identified regions of interest 312 within that game session as output. Additional actions, such as those indicated by bubble A, can be taken. Figure 3C Further details are provided below.
[0094] Figure 3C A system 300C according to one embodiment of the present disclosure is illustrated for training via an AI camera view model, which can be configured to recognize camera views of scenes previously identified by a region of interest AI model as regions of interest of interest to the viewer, wherein the identified camera views are also identified as being of interest to the viewer. System 300C can be implemented within a camera view engine 230 of a cloud gaming system 210, and more specifically, within the camera view engine 230, utilizing an AI server 260 of the cloud gaming system 210.
[0095] In some implementations, one or more renders of the game are generated by artificial intelligence (AI), such as using an AI-trained camera perspective model 321. In one implementation, the AI camera perspective model 321 can control the camera angle to capture the most interesting action and can clip between angles to merge actions that don't fit into a single view and show the action from different perspectives. Different streams with different camera perspectives can be generated for different viewers. The AI camera perspective model 321 can select camera angles to include a complete view of suspicious action, such as simultaneously showing the angles of both the sniper and the target the sniper is aiming at.
[0096] For training, system 300C receives one or more identified regions of interest (ROIs) 312 that are of great interest to the audience as input. Specifically, AI camera view trainer 320 receives one or more identified ROIs 312 as input. Additionally, other input data may be provided, such as metadata 302, player statistics 303, and other information from one or more game sessions of one or more video games previously provided to AI ROI trainer 310, where the game sessions may be single-player or multi-player game sessions.
[0097] As previously described, the AI camera view trainer 320 implements a deep learning engine 190 to perform analysis using training data collected during one or more game sessions of one or more video games performed by human players and / or AI players, during viewers watching one or more game sessions, during replays of one or more game sessions, during live broadcasts with live commentary or commentary on game sessions, human input, etc., to identify camera views corresponding to rendered views of scenes with identified regions of interest. Specifically, the deep learning engine 190 utilizes artificial intelligence (including deep learning algorithms, reinforcement learning, or other AI-based algorithms) during the learning and / or modeling phases to construct a camera view model 321 using training data and success criteria.
[0098] For example, a feedback loop can be provided to train the camera view model 321. Specifically, with each iteration, the feedback loop can help improve the success criteria used by the camera view trainer 320 when outputting at least one camera view 322 of the corresponding scene with a selected region of interest.
[0099] As shown in the figure, the AI camera perspective trainer 320 can provide one or more potential camera perspectives as outputs 323 for a given set of inputs via a trained camera perspective model 321. The inputs can be associated with one or more game sessions of one or more video games performed by human players and / or AI players. For ease of illustration, the feedback process is described in terms of a set of inputs provided from a game session (such as a single-player game session or a multi-player game session). Additionally, the game session may be live or may have already occurred. In this way, historical game data corresponding to previous game sessions can be used to train the camera perspective model 321 to better predict how to focus on actions in the game session and select content of interest to the audience (the camera perspective for viewing).
[0100] As previously described, one or more identified regions of interest (ROIs) may be associated with a live, recorded, or replayable game session using game state data. The AI camera view trainer 320 is configured to generate one or more camera views for each of the identified ROIs. These camera views, provided as output 323, may be game-generated, such as by acquiring video generated for players in the game session. Some camera views provided as output 323 may be newly generated by the video game within the game session, where the new camera views may be AI-generated or suggested by a human-in-the-loop expert (e.g., a professional broadcaster). For example, a new camera view may include a close-up view of a target character seen through another player's scope.
[0101] During training, the camera perspectives provided as output 323 can be further filtered and fed back to the AI camera perspective trainer 320 to help define success criteria. For example, camera perspectives can be selected from output 323 by human-in-the-loop or by active and / or passive measurements by the audience.
[0102] As previously mentioned, viewer viewing options influence which camera angles can be selected through filtering. For example, each of the camera angles provided as output 323 may correspond to one or more game session segments that viewers can select to watch. This selection may help define success criteria that are defined as preferred camera angles of interest to viewers. For example, viewers may actively choose to watch a live game session or a recorded game session, or (e.g., using game status) a replay of a game session selected from one or more viewing options (of one or more game sessions), as previously described. Figure 4 A user interface 400 configured to select camera views and regions of interest (ROIs) for a game session of a game application, according to one embodiment of the present disclosure, is shown. The user interface 400 includes one or more view regions of interest (ROIs) 312 identified by an ROI AI model 311 as of viewer interest, and one or more camera views 323 generated for those identified ROIs 312, which are identified as of viewer interest using the AI camera view model 321. Thus, the viewer can select ROIs (312-1, 312-2, or 312-3) for the game session, where each ROI can focus on a different part of the corresponding game world (e.g., focusing on player 1 in one ROI, player 2 in another ROI, etc.). Once an ROI is selected, the viewer can select the corresponding camera view. For example, the viewer can select one of camera views 323-1A, 323-1B, or 323-1C for view region of interest 312-1. Similarly, the viewer can select one of camera views 323-2A, 323-2B, or 323-2C for view region of interest 312-2. Furthermore, viewers can select one of camera perspectives 323-3A, 323-3B, or 323-3C for their area of interest 312-3. In some implementations, viewers can control the camera perspective to create custom camera perspectives that are not presented to them as predefined options. In some cases, another viewer can also view a custom camera perspective controlled by the viewer.
[0103] Compared with previous information Figure 3BThe same action 360 described can be used to define a viewer-selected camera view 325 through filtering. For example, action 360 may include viewer jump-in count 360A, positive viewer feedback provided as comments or text feedback 360B, active or passive viewer action 360C, viewer replay count 360D, viewer streaming view count 360E, and other viewer feedback determined during viewing of the selected camera view provided as output 323 can help define the viewer's preferred corresponding camera view.
[0104] Other inputs related to the game session and / or the user can also be provided. For example, human input 324, which may inherently have great value or define high success criteria when used for training, can be provided. For example, a human expert can be used to help identify camera angles from one or more game sessions watched by the expert and provide them as feedback to the AI viewer camera angle trainer 320. That is, the camera position with the corresponding angle can be selected by the human expert. In one use case, a game session can be live-streamed while the expert is watching a live stream. In another use case, the expert may be watching a game session that can be recorded or replayed using the game state. In either case, the expert can actively select and / or define a favorite camera angle, which, when used for training, has or is transformed into a high success criterion according to the definition. For example, the previously described user interface 400 can be used by a human expert to select one or more camera angles for corresponding identified regions of interest. The user interface 400 includes one or more viewer regions of interest 312 identified by the region of interest AI model 311 as being of interest to the viewer, and one or more camera angles 323 generated for those identified viewer regions of interest 312, which are identified as being of interest to the viewer using the AI camera angle model 321.
[0105] In one implementation, the weighting module 350 provides weights for the feedback, which are provided as the human-in-the-loop selected camera viewpoint 324 and the viewer-selected camera viewpoint 325.
[0106] Once trained, the AI camera view 321 is able to identify one or more viewpoints for a given input. For example, when a game session (e.g., single-player or multiplayer) is provided, the AI camera view model 321 provides one or more identified camera views 322 as output for a corresponding scene within a selected region of interest within that game session. The region of interest within the game application may correspond to a scene generated by a video game within a game world, which can be viewed through one or more rendered views obtained from different camera views within that game world. For example, the AI camera view model 321 can generate multiple view renders, which may be a fixed camera position within the game world, and / or a camera position with corresponding perspectives focusing on individual players or teams, and / or a camera position with corresponding perspectives controlled by AI to focus on various aspects of the action. As an example, identified camera views 322 within the game world can be obtained from the perspective of an enemy in the game session. This would allow the viewer to observe the player from the perspective of an enemy (another player or generated game) fighting the observed player. Additional actions can be taken, as indicated by bubble B, and... Figure 3D Further details are provided below.
[0107] Figure 3D A system 300D according to one embodiment of this disclosure is shown for training an AI-based commentary model 331, which can be configured to generate commentary and / or narration for scenes previously identified by the region of interest AI model 311 as regions of interest of interest to the audience, wherein the commentary can be tailored to a camera perspective identified by the AI camera perspective model 321. System 300D can be implemented within the broadcaster / narrator engine 220 of a cloud gaming system 210, and more specifically within the broadcaster / narrator engine 220 utilizing the AI server 260 of the cloud gaming system 210.
[0108] In some implementations, commentary is generated for scenes in identified regions of interest (ROIs) previously identified by artificial intelligence as of interest to one or more viewers. The commentary is generated using statistics and facts based on metadata and player statistics collected and / or generated for the corresponding game session performed by human and / or AI players. That is, the AI commentary model 331 can add commentary to the scene corresponding to the selected ROI, where the scene's commentary and / or camera view can be streamed to one or more users. In this way, the most exciting scenes of a specific game session supported by commentary built to generate maximum audience interest can be streamed to one or more viewers, providing a view for a live game session or generating a view for a replay of the corresponding game session, which can be streamed during or after the game session.
[0109] For training, system 300D receives one or more identified camera views as input for a corresponding scene within a selected region of interest. Specifically, AI broadcaster / narrator trainer 330 receives one or more identified camera views 322 as input. For illustrative purposes, system 300D can be trained using a specific camera view 322-A selected by AI and associated with a scene within a region of interest identified or selected by AI.
[0110] Additionally, the AI broadcaster / narrator trainer 330 can receive one or more identified regions of interest 312 that are of great interest to the audience, either alone or in conjunction with one or more identified camera viewpoints 322. That is, commentary and / or narration can be generated based on the identified regions of interest, allowing commentary to be streamed as audio without any video, or as a supplement to video.
[0111] Additionally, other input data, such as game and player statistics and facts 342 generated by artificial intelligence, can be provided to the AI broadcaster / commentator trainer 330. In one implementation, the game and player statistics and facts 342 are not filtered, where the AI commentary model 331 is trained to filter the statistics and facts 342 to use artificial intelligence to identify one or more statistics and facts that are of great interest to the audience. For example, metadata 302, player statistics 303, and other information from one or more game sessions of one or more video games previously provided to the AI audience region of interest trainer 310 can also be provided to the AI statistics and facts trainer 340. In this way, the statistics and facts trainer 340 implements a deep learning engine 190 to analyze training data collected during one or more game sessions to use artificial intelligence to identify statistics and facts about those game sessions. Specifically, the deep learning engine 190 can utilize artificial intelligence (such as deep learning algorithms, reinforcement learning, etc.) to build a statistics / fact model 341 for identifying and / or generating statistics and facts for corresponding game sessions.
[0112] For example, a feedback loop can be provided to train the AI commentary model 331. Specifically, with each iteration, the feedback loop can help improve the success criteria used by the AI broadcaster / commentator trainer 330 when outputting one or more selected and / or identified statistics and facts 333. As shown, the AI broadcaster / commentator trainer 330 can provide one or more statistics and facts 333 as outputs for a given set of inputs (e.g., an identified scene and / or the camera view of that scene) through the trained statistics / facts model 331. The inputs can be associated with one or more game sessions. For ease of illustration, the feedback process is described in terms of a set of inputs provided from a game session (e.g., a live, recorded, or replayable game session using game states). In this way, historical game data corresponding to previous game sessions can be used to train the AI commentary model 331 to better predict how to focus on actions in the game session and select content (e.g., statistics and facts) of interest to the audience when constructing the corresponding commentary.
[0113] As shown in the figure, the AI broadcaster / commentator trainer is configured to select statistics and facts 333 from a set of statistics and facts 342 generated and / or identified using an AI-powered statistical information / fact model 341 that utilizes game sessions. Specifically, the AI broadcaster / commentator trainer 330 is configured to use AI to identify one or more statistics and facts that are of great interest to the audience. In this way, the AI broadcaster / commentator trainer 330 implements a deep learning engine 190 to analyze training data collected during one or more game sessions, during live or recorded or replayed game sessions watched by the audience, live commentary of the game sessions, and human input to identify statistics and facts that are of great interest to the audience. Specifically, the deep learning engine 190 can utilize AI (such as deep learning algorithms, reinforcement learning, etc.) to build the AI commentary model 331. In this way, given a set of inputs (e.g., selected audience interest regions with corresponding scenes and / or corresponding camera perspectives), the AI commentary model 331 is configured to output one or more game and player statistics and facts that are of great interest to the audience as output 331, which can then be used to train or build commentary 332.
[0114] The selected statistics and facts provided as output 333 can be further filtered and provided as feedback to help define success criteria. For example, statistics and facts can be selected from output 333 by human intervention or by active and / or passive audience measurement. That is, filtering can define expert-selected statistics and facts 333A and AI-selected statistics and facts 333B, which are reinforced by audience actions.
[0115] As previously mentioned, viewer viewing options influence which camera perspectives can be selected through filtering. For example, the statistics and facts provided as output 333 may correspond to one or more game session segments that viewers can choose to watch. This selection may help define success criteria for statistics and facts that are defined as of the viewer's interest. For example, viewers may actively choose to watch a live game session or a recorded game session, or (e.g., using game status) a replay of a game session selected from one or more viewing options (of one or more game sessions), as previously described. The selected view also corresponds to unique statistics and facts that can help define the statistics and facts 333B selected by the AI. Additionally, viewer actions (e.g., jump-in count 360A, positive viewer feedback 360B, active or passive viewer actions 360C, replay count 360D, streaming view count 360E, etc.) can be used to define the statistics and facts 333B selected by artificial intelligence. For example, statistics and facts associated with game session segments that can be viewed by one or more viewers through selection may be defined as of viewer's high interest because the viewed game session is also identified as of viewer's high interest.
[0116] Additional inputs related to the game session and / or the user can also be provided. For example, human inputs that may inherently have significant value or define high success criteria for use in training can be provided. For instance, human experts can be used to help identify statistics and facts 333A from one or more game sessions watched by the expert and provide them as feedback to the AI broadcaster / commentator trainer 330. That is, certain statistics and facts provided as output 333 can be selected by the human expert watching the corresponding game session, which can be a live stream, a recorded game session, or a game session using game state replay. In this way, the expert can actively select and / or define preferred statistics and facts that, according to definition, have or are transformed into high success criteria when used for training.
[0117] In one implementation, the weighting module 350 provides weights for the feedback, which are provided as human-in-the-loop selected statistics and facts 333A and AI-generated and audience-enhanced statistics and facts 333B.
[0118] In one implementation, filtered statistics and facts used to build the AI commentary model 331 can be sent to the AI game and player statistics and facts trainer 340. In this way, the trained statistics / facts model 341 can also be configured to identify one or more statistics and facts that are of great interest to the audience. In one implementation, instead of outputting all statistics and facts for a particular game session, the statistics / facts model 341 is configured to filter those statistics and facts to output only those that are of great interest to the audience.
[0119] In one implementation, the expert-selected statistical information and facts 33A are collected from the live commentary 334 provided by the expert. That is, the live commentary 334 includes expert-selected statistical information and facts 333A, which can be parsed in one implementation or provided through active selection by the expert in another. The live commentary can also be provided as feedback to the broadcaster / commentator trainer 330 to help identify the commentary type (e.g., for template construction) and one or more corresponding expert-selected facts that are of great interest to the audience, wherein the feedback can be used to construct templates for generating commentary.
[0120] In another implementation, the statistical information and facts 333B selected by the AI are used to construct the AI-generated commentary 335. For example, the commentary 335 can be constructed using commentary templates 338 woven from the statistical information and facts 333B selected by the AI. Audience actions 360 can be used to help select those commentary templates that are most interesting to the audience. Each of the templates can reflect a specific style preferred by one or more audiences, such as a template for a single audience member or a template with a style for a group of audience members.
[0121] In another implementation, the AI-generated narration 335 can be further filtered using culture and / or language, as well as other filters. For example, the AI-generated narration 335 can be formatted for a specific language using a language filter 337. In this way, a single scene and / or the camera perspective of a scene can be supported by one or more narrations 335 formatted in one or more languages. Furthermore, the narration 335 can be further filtered for cultural characteristics (e.g., sensitivities, preferences, likes, etc.) using a culture filter 336. Cultural characteristics can be defined for specific geographic regions (e.g., countries, states, borders, etc.). For example, different cultural characteristics may describe a scene differently, such as ignoring objects or related themes in one cultural characteristic, or positively addressing objects or related themes in another. Some cultural characteristics may prefer engaging commentary, while others may prefer more moderate commentary. In this way, the AI-generated narration and / or commentary 335 can be generated using appropriate templates 338 further customized for culture and language. In some implementations, the AI-generated narration and / or commentary 335 may be more personalized for a specific audience or group of viewers, where a cultural filter 336 is designed to apply filters specific to the audience or group of viewers. In this way, the AI narration model 331 is configured to generate one or more streams for viewers in multiple locations, including commentary in different languages that take into account cultural and gender considerations.
[0122] In one implementation, a weighting module 350 provides weights for the feedback, which are provided as human-in-the-loop live commentary 334 and AI-generated commentary 335. In some implementations, the broadcast includes commentary from one or more human commentators and one or more AI commentators. When determining what to say, the AI commentator may combine the content spoken by the human commentator. In some cases, an AI commentator in a sporting event broadcast may combine the content spoken by the human commentator with a separate broadcast of the same event when determining what to say. In some implementations, the AI commentator may combine the content spoken by one or more commentators and / or one or more viewers when determining what to say, wherein the combined statements are not included in the broadcast. These statements can be accessed through the broadcast system or through other channels such as social media. In some cases, the statements may be non-verbal, such as emojis, likes, shares, or favorites.
[0123] Once trained, the AI commentary model 331 is configured to construct commentary 332 for a given input using AI-selected statistics and facts. For example, when provided with a game session (e.g., single-player or multi-player) along with its corresponding data (e.g., metadata, player statistics, etc.), the AI commentary model 331 provides commentary 332 as output, which may be further filtered for language and / or cultural characteristics (applicable to one or more audiences). That is, the AI commentary model 331 may include or access a cultural filter 336, a language filter 337, and a commentary template 338.
[0124] In some implementations, the AI-generated commentary and / or narration 332 possesses a compelling personality. For example, the AI narration model can be configured to select statistics and facts that can convey to the audience why the view streamed to them is important. This can be achieved by selecting an appropriate commentary template 338. For instance, the AI-generated commentary and / or narration 332 can provide notification that a new or low-level player is in a position to eliminate an experienced player (a specific zone of interest). In this way, notification can be provided to the audience in an attempt to entice them to watch the scene to see if a new player can actually take advantage of the situation to eliminate a more experienced player.
[0125] In other implementations, AI voice-over can be provided as commentary and / or narration 332. In one implementation, the voice-over can explain the actions of esports games to make them understandable and exciting for visually impaired viewers. In another implementation, the voice-over is provided as detailed reporting and commentary on events that can stand alone, or as a view of the event.
[0126] Based on a detailed description of the various modules of the game server and client device that communicate over a network, according to one embodiment of this disclosure, the following now pertains to... Figure 5 Flowchart 500 describes a method for identifying camera views from which viewers may potentially be interested using one or more AI models. Specifically, flowchart 500 illustrates the operational processes and data flows involved at a backend AI server for generating one or more renders of a game using artificial intelligence, wherein the renders can be streamed to one or more viewers. For example, the method of flowchart 500 can be at least partially derived from… Figure 1A , Figure 2 , Figure 3A and Figure 3C The camera perspective engine 230 is executed at the cloud gaming server 210.
[0127] At 510, the method includes receiving game state data and user data of one or more players participating in a game session of a video game. The game session can be a single-player or multi-player session. Metadata (which partially includes game state and user data) is received in association with the game session at a cloud gaming server. User data of one or more players participating in the game session can be received. For example, game state data can define the state of gameplay at a corresponding point in the gameplay process (e.g., game state data includes game characters, game objects, object attributes, graphical overlays, character assets, character skill settings, character's quest achievement history in the game application, character's current geographical location in the game environment, character's current state of gameplay, etc.). Game state can allow the generation of the game environment (e.g., a virtual game world) existing at the corresponding point in the gameplay process. User-saved data can be used to personalize the game application for the corresponding user (e.g., player), said data can include information for personalizing characters during gameplay (e.g., shape, appearance, clothing, weapons, game difficulty, game levels, character attributes, etc.). As previously mentioned, other information may include random seed data that may be associated with the game state.
[0128] At 520, the method includes identifying a region of interest (ROI) for the viewer during a game session. The ROI is associated with a scene within the virtual game world of the video game. For example, a scene may include a virtual game world located at a specific location or in the general vicinity of a character, which may be viewed from one or more camera perspectives within the virtual game world. For example, camera perspectives may include views from a character's perspective, views from another player, top-down views of the scene, close-up views of objects or areas within the scene, scope views, etc.
[0129] In one implementation, artificial intelligence can be used to identify regions of interest (ROIs) for viewers. Specifically, more ROIs can be identified using an AI model trained to isolate one or more ROIs that may be viewed by one or more viewers, as previously described. For example, an ROI could indicate when a popular player enters a game session, or when two expert players are geographically within a certain distance from each other in the game world and are likely to meet. Each identified ROI is associated with a corresponding scene in the virtual game world within the game session. In one implementation, a first ROI may have the highest rating of interest among potential viewers and may be associated with a first scene (e.g., focusing on the most popular expert player). Within the scene, one or more camera perspectives can be generated, such as a view from one player, or a view from a second player, a top-down view, a close-up view, a scope view, etc.
[0130] In another implementation, viewers' regions of interest (ROIs) can be identified through viewer input. For example, viewers can actively select the ROIs they wish to view. For instance, a selected ROI could include a specific player's gameplay action within a game session. As another example, a selected ROI could include gameplay actions of one or more players interacting with a specific part of the virtual game world. Each selected and identified ROI is associated with a corresponding scene within the virtual game world of the game session. Within the scene, one or more camera perspectives can be generated, such as a view from one player, or a view from a second player, a top-down view, a close-up view, a scope view, etc.
[0131] At 530, the method includes identifying a first camera viewpoint corresponding to a viewer's region of interest. The first camera viewpoint is identified using an AI model (e.g., camera viewpoint model 321) trained to generate one or more camera views corresponding to the viewer's region of interest. The first camera viewpoint can be determined to be of potential greatest interest (e.g., highest rating) to one or more viewers. For example, the first camera viewpoint could show a view from a player acting as an expert, where the expert has a target in his or her scope. In this way, the first camera viewpoint can be identified and streamed over a network to one or more viewers, wherein the streamed content includes views (e.g., camera views) determined by artificial intelligence to be of great interest to one or more viewers.
[0132] In one implementation, the stream may include one or more camera views of a scene. For example, the scene may be centered on an expert player who has a target in his or her scope. The stream may show different camera views of the scene, such as one view from the expert player and another view from the target player (e.g., an unsuspecting target), and yet another view including a close-up view of the target obtained through the scope. For example, a second camera view may be identified and generated for the scene based on an AI model trained to generate one or more camera views corresponding to a viewer's region of interest. The second camera view may already be game-generated (e.g., for the current gameplay of a player in a game session), or it may be newly generated using the game state and not previously generated by the video game being executed (e.g., a scope camera view, or a top-down view showing two players). In this way, the streamed content may include different views of the same scene.
[0133] In another implementation, first and second camera perspectives are generated for different streams. For example, a first camera perspective can be generated for a first stream sent over a network to a first group of one or more viewers. In one use case, the first camera perspective may focus on a first professional player, where the first group of viewers may want to focus on the first professional player in the first stream. A second camera perspective may focus on a second professional player, where the second group of viewers may want to focus on the second professional player in the second stream. The two streams can be delivered over the network simultaneously to provide different views of the same scene to different groups of viewers.
[0134] In one implementation, a first camera viewpoint is identified for use as a highlight of a streaming game session, where the game session can be a live stream showing the highlight during the game session, or where the game session may have ended and is being watched via recording or replay using game state. The highlight may include interesting scenes (e.g., tournament scenes) identified from regions of interest in previous gameplay that occurred during the game session. The camera viewpoint may have been previously generated by the video game, such as for players in the game session, or may be newly generated via artificial intelligence from a viewpoint identified as of great interest to one or more viewers.
[0135] In one implementation, the narration of a scene is generated based on an AI model (e.g., AI broadcast / narration model 331), which is trained to construct narration for scenes corresponding to those regions of interest using statistical information and facts relevant to those regions of interest (e.g., focusing on the most popular professional gamers). The statistical information and facts can be selected by artificial intelligence and used to construct narration and / or commentary, as follows regarding... Figure 6 A more comprehensive description.
[0136] Based on a detailed description of the various modules of the game server and client device that communicate over a network, according to one embodiment of this disclosure, the following now pertains to... Figure 6 Flowchart 600 describes a method for constructing and / or generating narration for a scene identified as a region of interest (ROI) of interest to a viewer using one or more AI models, wherein the narration can be customized for one or more camera views of the scene identified by artificial intelligence. Specifically, flowchart 600 illustrates the operational processes and operational data flows involved at a backend AI server for generating one or more renderings of a game using artificial intelligence, wherein the renderings can be streamed to one or more viewers. For example, the method of flowchart 600 can be at least partially derived from… Figure 1A , Figure 2 , Figure 3A and Figure 3D The broadcaster / narrator engine 220 is executed at cloud gaming server 210.
[0137] At 610, the method includes receiving game state data and user data of one or more players participating in a game session of a video game being played by one or more players participating in a single-player or multi-player session. As previously described, metadata (which partially includes game state and user data) is received in association with the game session at a cloud gaming server. User data of one or more players participating in the game session may be received. For example, game state data may define the state of gameplay at a corresponding point in the gameplay process (e.g., game state data includes game characters, game objects, object attributes, graphical overlays, character assets, character skill settings, character's quest achievement history in the game application, character's current geographic location in the game environment, character's current state of gameplay, etc.). Game state may allow the generation of the game environment (e.g., a virtual game world) existing at the corresponding point in the gameplay process. User-saved data may be used to personalize the game application for the corresponding user (e.g., player), wherein the data may include information for personalizing the character during gameplay (e.g., shape, appearance, clothing, weapons, game difficulty, game levels, character attributes, music, background sounds, etc.). As previously described, other information may include random seed data that may be associated with the game state.
[0138] At 620, the method includes identifying a region of interest (ROI) for the viewer during a game session. The ROI is associated with a scene within the virtual game world of the video game. For example, a scene may include a virtual game world located at a specific location or in the general vicinity of a character, where the scene can be viewed from one or more camera perspectives within the virtual game world. For example, camera perspectives may include views from a character's perspective, views from another player, top-down views of the scene, close-up views of objects or areas within the scene, scope views, etc.
[0139] In one implementation, artificial intelligence can be used to identify regions of interest (ROIs). Specifically, ROIs can be identified using an AI model trained to isolate one or more ROIs that may be viewed by one or more viewers. For example, an ROI could indicate when a popular player enters a game session, or when two expert players are geographically within range of each other in the game world and are likely to meet, or identify areas in the virtual game world where a large number of actions are seen (e.g., player failure, character accidents, character deaths, etc.). Each identified ROI is associated with a corresponding scene in the virtual game world within the game session. In the scene, one or more camera perspectives can be generated, such as a view from one player, or a view from a second player, a top-down view, a close-up view, a scope view, etc. For example, a first ROI can be associated with a first scene in the virtual game world.
[0140] In another implementation, viewers' regions of interest (ROIs) can be identified through viewer input. For example, viewers can actively select the ROIs they wish to view. For instance, a selected ROI could include the gameplay actions of a specific player within a game session. As another example, a selected ROI could include the gameplay actions of one or more players interacting with a specific part of the virtual game world. Each selected and identified ROI is associated with a corresponding scene in the virtual game world within the game session. Within the scene, one or more camera perspectives can be generated, such as a view from one player, or a view from a second player, a top-down view, a close-up view, a scope view, etc. Additionally, viewers may wish to include accompanying narration (e.g., broadcasts) within the view of the selected ROI, where the narration can be generated using artificial intelligence, as described more fully below.
[0141] At 630, the method includes generating statistics and facts for a game session based on game state data and user data using an AI model (e.g., statistics / fact model 341 and / or AI interpretation model 331), said AI model being trained to isolate game state data and user data that one or more viewers might be interested in. The isolated game state data and user data can be transformed or used to generate statistics and facts through artificial intelligence. Specifically, the statistics and facts are related to a corresponding identified region of interest, the scene associated with that region of interest, and possibly to one or more identified camera views of that scene. Additionally, the AI model is configured to identify one or more facts and statistics that are of great interest to one or more viewers, as determined by artificial intelligence.
[0142] At 640, the method includes generating narration for a scene in the audience's region of interest using another AI model (e.g., AI narration model 331), said AI model being trained to select statistics and facts previously identified for the scene. In another implementation, the AI model trained to select statistics and facts does not perform filtering, because a previously identified AI model trained to isolate one or more game state data and user data that might be of interest to the audience can also be trained to generate one or more statistics and facts that might be of interest to the audience using artificial intelligence. In one implementation, the AI model trained to select statistics and facts further filters the statistics and facts used for narration. That is, the AI model is trained to select statistics and facts from those generated using statistics and facts generated by the AI model trained to isolate one or more game state data and user data that might be of interest to the audience, wherein the selected statistics and facts have the highest potential audience interest, as determined by the AI model trained to select statistics and facts.
[0143] Additionally, the AI model, trained to select statistical information and facts, is configured to generate narration and / or commentary using the selected statistical information and facts. In one implementation, narration is generated for the scene using an appropriate template. The template may be based on the video game genre of the video game or the region of interest (ROI) type of the first viewer's ROI (i.e., the ROI defines a specific scene or situation corresponding to a particular template). In some implementations, the template considers the target or requesting group of viewers watching the scene with the AI-generated narration. For example, the template may consider the commentary style preferred by the viewer group (e.g., excited or calm, etc.).
[0144] In other implementations, further filtering of the narration is performed. For example, the narration may be filtered taking into account the cultural customs and habits of the target audience group. For example, depending on the group and the geographic region associated with the group, there may be cultural preferences that determine how the narration is structured, as previously mentioned. For example, one group might not have questions about including narration on the topic of the scene, while another group might object to including that topic in the narration. In some implementations, the filtering may apply a communication format filter, where the communication format can be spoken, non-spoken (sign language), etc. For example, the filtering may apply a language filter, such that the narration is generated for a specific language preferred by the audience group. In some implementations, when a non-spoken filter is applied, the filtered customs and habits should reflect the correct expression taking into account gender, sign language variations, geographic location, etc.
[0145] In one implementation, taking into account one or more filters, the same scene can be streamed to different audience groups. For example, a first narration tailored to one or more first customs of a first group of viewers and / or a first geographic region can be generated, wherein the first narration is generated in a first language preferred by the first group. Additionally, a second narration tailored to the same scene tailored to one or more second customs of a second group of viewers and / or a second geographic region can be generated, wherein the second narration is generated in a second language preferred by the second group.
[0146] In one implementation, the narration can be customized for the corresponding camera view that is also being streamed to the audience group. In this way, the narration can closely follow the rendered view that is also being streamed. For example, the narration can move from a first-person perspective of the scene obtained from the first player's view, and then the narration can shift content-wise to focus on a second-person perspective of the scene being streamed to the audience group. Thus, there is a strong correlation between the streaming camera view of the scene and the AI-generated narration.
[0147] In one implementation, commentary is generated for highlights of a streaming game session, where the game session can be a live stream showing highlights during the game session, or where the game session may have ended and is being watched via recording or replay using game state. Highlights can include interesting scenes (e.g., tournament scenes) identified from regions of interest in previous gameplay that occurred during the game session. Camera perspectives may have been previously generated by the video game, such as for players in the game session, or can be newly generated using artificial intelligence from perspectives identified as highly interesting to one or more viewers. Furthermore, the commentary is constructed for streaming highlights and can be customized for the scene and the camera perspectives for streaming the scene, where the commentary includes statistics and facts identified as highly interesting to one or more viewers using artificial intelligence. In this way, even low-traffic game sessions can have highlights including AI commentary that includes interesting statistics and facts generated and identified using artificial intelligence.
[0148] Figure 7 The invention illustrates the generation of highlight segments of a game session of one or more players, including a game application, according to an embodiment of the present disclosure. The highlight segments are generated using an AI model that can identify regions of interest (ROIs) in multiple game sessions of multiple game applications to identify interesting camera angles in those ROIs and provide narration for the scenes in those ROIs. Figure 7 The process flow shown can be partially implemented by the highlight engine 240.
[0149] Specifically, highlight clips can be generated for live game sessions (e.g., providing highlights during a game session) or for game sessions that have ended, where highlight clips can be generated from game session recordings or from game sessions using game state replays.
[0150] Therefore, highlight reports can be obtained after the game session ends. As mentioned earlier, such reports can be automatically generated based on AI options. For example, highlight reports that include areas of interest, camera angles, and statistical information and facts can be based on what viewers choose to watch or rewatch while watching the game session. For example, highlight reports could include near misses in the game session that did not result in any player elimination but were still very exciting for viewers to watch.
[0151] Additionally, highlight reports can include replays of scenes not shown in the live stream, or replays of camera angles of corresponding scenes that may not have been generated during the game session. For example, a highlight report could show interesting player eliminations in a multiplayer game session that weren't shown in the live stream. Alternatively, a highlight report could show camera angles that might be very popular with viewers but weren't generated during the game session. For instance, in a player elimination shown in the live stream, a new overhead camera angle could be generated, showing both the shooter and the eliminated player in a single view.
[0152] In one implementation, highlight coverage of a game session can last longer than the match itself, where the game session is match-specific (e.g., head-on clashes, final survivors, etc.). In some implementations, highlight coverage of earlier eliminations in a match or game session can continue after the game session ends. For example, in situations where the state of the game session is not available in real time, coverage of the game session can be extended to show each elimination, potentially extending the live broadcast of the game session beyond the duration of the corresponding match.
[0153] like Figure 7 As shown, metadata 702 and player statistics for players 1-N in the game session are fed to an AI audience region of interest (ROI) trainer 310, which is now configured to apply an ROI model 311. For example, the trainer 310 can be implemented as an AI server 260 or within an AI server. The ROI model 311 is configured to generate and / or identify one or more ROIs that are of great interest to one or more viewers. For example, an ROI could focus on a highly popular pro player, making it possible for one or more viewers to want to stream a pro player's view during the game session or while watching highlights.
[0154] The identified regions of interest (ROIs) can be provided as input to an AI camera view trainer 320, which is now configured to apply a camera view model 321. For example, trainer 320 can be implemented as an AI server 260 or within an AI server. Camera view model 321 is configured to generate one or more camera views for the corresponding ROIs. These camera views may be generated by the video game during a game session, or they may be newly generated to capture interesting rendered views that are not previously shown but are of great interest to the audience (e.g., a top-down view showing the shooting player and the eliminated player).
[0155] The identified and / or generated camera perspectives are provided as input to the AI commentary trainer 330, which is now configured to apply the broadcast / commentary model 331. For example, the trainer 330 may be implemented as an AI server 260 or within an AI server. The broadcast / commentary model 331 is configured to generate commentary for scenes corresponding to and identified regions of interest. Specifically, the broadcast / commentary model 331 is configured to select relevant statistics and facts for scenes in a game session using artificial intelligence. In one implementation, the statistics and facts are selected from a set of statistics and facts generated based on metadata and player statistics of the game session. For example, metadata 702 and player statistics for players 1-N in the game session are fed to the AI game and player statistics and facts trainer 340, which is now configured to apply the statistics / facts model 341 to generate all statistics and facts for the game session. For example, the trainer 340 may be implemented as an AI server 260 or within an AI server.
[0156] In one implementation, the statistical information / fact model 341 and / or the narration model 331 are configured individually or in combination to identify and / or generate statistical information and facts for corresponding scenes in regions of interest that are of great interest to one or more viewers. In this way, the selected statistical information and facts can be used to generate narration 335 using artificial intelligence. For example, the statistical information and facts selected by the AI can be woven into an appropriate template to generate narration 335, wherein the template can be tailored to a specific scene being viewed, or it can be tailored to a target audience (e.g., a preference for excellent narration). Additionally, one or more filtering processes can be performed on the narration 335. For example, a cultural filter can be applied to the narration, where the cultural filter is suitable for the target audience. Additionally, a language filter can be applied to the narration 335, enabling the commentary and / or narration 335 to be provided in a preferred language.
[0157] As shown in the figure, the highlight clip generator can package a rendered view of a scene within a game session's identified region of interest (ROI) along with AI-generated commentary. The highlight clip can be streamed over network 150 to one or more viewers 720. In this way, one or more viewers 720 can watch the most exciting scenes of a specific game session, supported by commentary built to maximize viewer interest. This can provide a live view of the game session or generate a view for a replay of the corresponding game session, which can be streamed during or after the game session.
[0158] In one implementation, a second highlight fragment can be generated from another first highlight fragment of the game session. That is, the first highlight fragment is provided as input to the highlight engine 240, and the second highlight fragment is generated as output. In other words, one or more replay renders (e.g., highlight fragments) can be made available to other viewers. Furthermore, other subsequent highlight fragments can be generated from the second highlight fragment, enabling the generation of a highlight fragment chain.
[0159] In one implementation, a rendered view of a game session or a highlight segment of a game session can be provided in the context of a live event, such as an esports tournament, where a multiplayer game session is performed live in front of an audience. Specifically, multiple viewers can watch the live event in a public location (e.g., an esports arena), where one or more monitors may be set up for multiple viewers to watch. Additionally, in addition to the screens available to multiple viewers, viewers may also have one or more personal screens for them to watch. Each viewer can actively control which event views are displayed on their personal screen using embodiments of this disclosure. That is, the viewer uses the previously described AI model to control the camera positioning used for rendering the live gameplay. For example, the camera perspective can be fixed in space, focus on individual players or teams, or be configured to capture various interesting aspects of the action through artificial intelligence.
[0160] Figure 8 Components of an exemplary apparatus 800 that can be used to perform various aspects of embodiments of the present disclosure are shown. For example, Figure 8 An exemplary hardware system is illustrated that is suitable for rendering one or more scenes of a game application generated by artificial intelligence (AI) and further suitable for narrating and / or rendering scenes of a game application generated by AI, according to embodiments of this disclosure. The block diagram illustrates apparatus 800, which may be combined with or may be a personal computer, server computer, game console, mobile device, or other digital device, each suitably suited for practicing embodiments of this disclosure. Apparatus 800 includes a central processing unit (CPU) 802 for running software applications and optionally running an operating system. CPU 802 may consist of one or more homogeneous or heterogeneous processing cores.
[0161] According to various implementations, CPU 802 is one or more general-purpose microprocessors having one or more processing cores. Other implementations may use one or more CPUs with microprocessor architectures particularly suited to highly parallel and computationally intensive applications configured for graphics processing during game execution, such as media and interactive entertainment applications.
[0162] Memory 804 stores applications and data used by CPU 802 and GPU 816. Storage device 806 provides non-volatile storage for applications and data and other computer-readable media, and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROMs, DVD-ROMs, Blu-ray discs, HD-DVDs, or other optical storage devices, as well as signal transmission and storage media. User input device 808 conveys user input from one or more users to device 800, examples of which may include a keyboard, mouse, joystick, touchpad, touchscreen, still or video recorder / camera, and / or microphone. Network interface 814 allows device 800 to communicate with other computer systems via electronic communication networks and may include wired or wireless communication over local area networks and wide area networks such as the Internet. Audio processor 812 is adapted to generate analog or digital audio output from instructions and / or data provided by CPU 802, memory 804, and / or storage device 806. The components of device 800 (including CPU 802, graphics subsystem including GPU 816, memory 804, data storage device 806, user input device 808, network interface 810, and audio processor 812) are connected via one or more data buses 822.
[0163] According to one embodiment of this disclosure, camera view engine 230 may be configured within CPU 802 or as separate hardware from CPU 802, and is further configured to identify camera views of scenes previously identified by the region of interest (ROI) AI model as ROIs of interest to the audience, wherein the camera views are also identified by the AI model as ROIs of interest to the audience. According to one embodiment of this disclosure, broadcaster / narrator engine 220 may be configured within CPU 802 or as separate hardware from CPU 802, and is further configured to generate narration using an AI model for scenes previously identified by the ROI AI model as ROIs of interest to the audience, wherein the generated narration may be customized for the camera views identified by the AI camera view model. AI server 260 may be configured for training and / or implementing the AI camera view model and for training and / or implementing the AI broadcast / narration model. According to embodiments of this disclosure, the highlight engine 240 may be configured within the CPU 802 or as separate hardware from the CPU 802, and is further configured to generate highlight segments of game sessions for one or more players of a game application, wherein the highlight segments are generated using an AI model that can identify regions of interest (ROIs) in multiple game sessions of multiple game applications to identify interesting camera angles of those ROIs and provide narration for the scenes of those ROIs.
[0164] The graphics subsystem 814 is further connected to the data bus 822 and components of the device 800. The graphics subsystem 814 includes a graphics processing unit (GPU) 816 and a graphics memory 818. The graphics memory 818 includes display memory (e.g., a frame buffer) for storing pixel data for each pixel of the output image. The graphics memory 818 may be integrated into the same device as the GPU 816, connected to the GPU 816 as a separate device, and / or implemented within memory 804. Pixel data may be provided directly from the CPU 802 to the graphics memory 818. Alternatively, the CPU 802 provides the GPU 816 with data and / or instructions defining the desired output image, and the GPU 816 generates pixel data for one or more output images based on the data and / or instructions. The data and / or instructions defining the desired output image may be stored in memory 804 and / or graphics memory 818. In one implementation, the GPU 816 includes 3D rendering capabilities for generating pixel data for an output image based on instructions and data that define the geometry, lighting, shading, texturing, motion, and / or camera parameters of a scene. The GPU 816 may also include one or more programmable execution units capable of executing shader programs.
[0165] The graphics subsystem 814 periodically outputs pixel data of an image from the graphics memory 818 for display on the display device 810 or for projection by the projection system 840. The display device 810 can be any device capable of displaying visual information in response to signals from the device 800, including CRT, LCD, plasma, and OLED displays. The device 800 can provide, for example, analog or digital signals to the display device 810.
[0166] Other implementations for optimizing the graphics subsystem 814 may include multi-tenant GPU operations that share GPU instances among multiple applications, and distributed GPUs that support a single game. The graphics subsystem 814 may be configured as one or more processing devices.
[0167] For example, in one implementation, the graphics subsystem 814 can be configured to perform multi-tenant GPU functionality, whereby the graphics subsystem can implement graphics and / or rendering pipelines for multiple games. That is, the graphics subsystem 814 is shared among multiple running games.
[0168] In other implementations, the graphics subsystem 814 includes multiple GPU devices that are combined to perform graphics processing on a single application running on a corresponding CPU. For example, the multiple GPUs may perform alternating frame rendering in consecutive frame cycles, where GPU 1 renders the first frame, GPU 2 renders the second frame, and so on, until the last GPU is reached, then the initial GPU renders the next video frame (e.g., if there are only two GPUs, GPU 1 renders the third frame). That is, the GPUs rotate while rendering frames. Rendering operations may overlap, where GPU 2 may begin rendering the second frame before GPU 1 has finished rendering the first frame. In another implementation, different shader operations may be assigned to the multiple GPU devices in the rendering and / or graphics pipeline. The main GPU is performing main rendering and compositing. For example, in a group comprising three GPUs, the primary GPU 1 can perform primary rendering (e.g., first shader operations) and composite the output from secondary GPUs 2 and 3, where secondary GPU 2 can perform second shader operations (e.g., fluid effects, such as rivers), and secondary GPU 3 can perform third shader operations (e.g., particle smoke). The primary GPU 1 composites the results from each of GPUs 1, 2, and 3. In this way, different GPUs can be assigned to perform different shader operations (e.g., waving flags, wind, smoke generation, fire, etc.) to render video frames. In yet another implementation, each of the three GPUs can be assigned to different objects and / or portions of the scene corresponding to a video frame. In the above implementations and methods, these operations can be performed in the same frame period (simultaneously in parallel) or in different frame periods (sequentially in parallel).
[0169] While specific implementations have been provided to demonstrate one or more renderings of scenes for game applications generated by artificial intelligence (AI), and / or narration and / or rendering of scenes for game applications generated by AI, those skilled in the art who read this disclosure will implement additional implementations that fall within the spirit and scope of this disclosure.
[0170] It should be noted that access services delivered over vast geographical areas (such as providing access to games in current implementations) often utilize cloud computing. Cloud computing is a computing paradigm in which dynamically scalable and often virtualized resources are provided as a service over the internet. Users do not need to be experts in the technical infrastructure supporting their “cloud.” Cloud computing can be categorized into different services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Cloud computing services typically provide commonly used applications (such as video games) online and accessible from a web browser, while the software and data are stored on servers in the cloud. Based on how the internet is depicted in computer network diagrams, the term cloud is used as a metaphor for the internet and is an abstract concept of the complex infrastructure it conceals.
[0171] A game processing server (GPS) (or simply "game server") is used by game clients to play single-player and multiplayer video games. Most video games played on the internet operate via a connection to a game server. Typically, games use a dedicated server application that collects data from players and distributes it to other players. This is more efficient and effective than a peer-to-peer arrangement, but it requires a separate server to host the server application. In another implementation, GPS establishes communication between players and their corresponding gaming devices to exchange information without relying on a centralized GPS.
[0172] Dedicated GPS servers operate independently of the client. These servers typically run on dedicated hardware located within a data center, providing greater bandwidth and dedicated processing power. For most PC-based multiplayer games, dedicated servers are the preferred method for hosting game servers. Large-scale multiplayer online games run on dedicated servers, which are often hosted by the software company that owns the game's title, allowing them to control and update content.
[0173] Users access remote services using client devices, which include at least a CPU, display, and I / O. Client devices can be PCs, mobile phones, netbooks, PDAs, etc. In one implementation, the network on the game server identifies the type of device used by the client and adjusts the communication method accordingly. In other cases, the client device uses standard communication methods (e.g., HTML) to access applications on the game server via the Internet.
[0174] The embodiments of this disclosure can be practiced with various computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and so on. This disclosure can also be practiced in distributed computing environments, where tasks are performed by remote processing devices linked via wired or wireless networks.
[0175] It should be understood that a given video game can be developed for a specific platform and a specific associated controller device. However, when such a game becomes available through the game cloud system presented herein, the user may be using a different controller device to access the video game. For example, the game may have been developed for a game console and its associated controller, while the user may be using a keyboard and mouse to access a cloud-based version of the game from a personal computer. In this case, input parameter configuration can define a mapping from inputs generated by the user's available controller device (in this case, a keyboard and mouse) to inputs acceptable for executing the video game.
[0176] In another example, a user can access the cloud gaming system via a tablet computing device, a touchscreen smartphone, or other touchscreen-driven device. In this case, the client device and controller device are integrated into the same device, where input is provided through detected touchscreen inputs / gestures. For such a device, input parameter configuration can define specific touchscreen inputs corresponding to the game inputs of the video game. For example, during the operation of the video game, buttons, steering wheels, or other types of input elements may be displayed or covered to indicate the location on the touchscreen that the user can touch to generate game inputs. Gestures such as swipes in a specific direction or specific touch actions can also be detected as game inputs. In one implementation, the user may be provided with instructions on how to provide input via the touchscreen to play the game before starting the gameplay of the video game, in order to familiarize the user with operating controls on the touchscreen.
[0177] In some implementations, the client device acts as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection to transmit input from the controller device to the client device. The client device can then process this input and transmit the input data to the cloud gaming server via a network (e.g., via a local networked device such as a router). However, in other implementations, the controller itself can be a networked device with the ability to transmit input directly to the cloud gaming server via the network, without first transmitting such input through the client device. For example, the controller can connect to a local networked device (e.g., the router mentioned above) to send and receive data from the cloud gaming server. Therefore, although the client device may still need to receive the video output from the cloud-based video game and render it on a local display, input latency can be reduced by allowing the controller to send input directly to the cloud gaming server via the network, thus bypassing the client device.
[0178] In one implementation, the networked controller and client device can be configured to send certain types of input directly from the controller to the cloud gaming server, and other types of input via the client device. For example, input whose detection does not rely on any additional hardware or processing outside the controller itself can be sent directly from the controller to the cloud gaming server via the network, bypassing the client device. Such inputs may include button inputs, joystick inputs, embedded motion detection inputs (e.g., accelerometers, magnetometers, gyroscopes), etc. However, inputs utilizing additional hardware or requiring processing by the client device can be sent to the cloud gaming server by the client device. These may include video or audio captured from the game environment, which can be processed by the client device before being sent to the cloud gaming server. Additionally, input from the controller's motion detection hardware can be processed by the client device in conjunction with the captured video to detect the controller's position and movement, which the client device then transmits to the cloud gaming server. It should be understood that the controller device according to various embodiments can also receive data (e.g., feedback data) from the client device or directly from the cloud gaming server.
[0179] It should be understood that the various implementation schemes defined herein can be combined or assembled into specific implementations using the various features disclosed herein. Therefore, the examples provided are merely some possible examples and are not limited to the various implementations that become possible by combining various elements to define more implementations. In some examples, some implementations may include fewer elements without departing from the spirit of the disclosed or equivalent implementations.
[0180] The embodiments of this disclosure can be practiced with various computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and so on. The embodiments of this disclosure can also be practiced in distributed computing environments where tasks are performed via remote processing devices based on wired or wireless network links.
[0181] In light of the above embodiments, it should be understood that embodiments of this disclosure can employ various computer-implemented operations involving data stored in a computer system. These operations are those that require the physical manipulation of physical quantities. Any operation described herein that forms part of embodiments of this disclosure is a useful machine operation. Embodiments of the invention also relate to means or apparatus for performing these operations. The apparatus may be specifically constructed for the desired purpose, or the apparatus may be a general-purpose computer selectively activated or configured by a computer program stored in a computer. Specifically, various general-purpose machines may be used with computer programs written in accordance with the teachings herein, or it may be more convenient to construct more specialized apparatus to perform the desired operations.
[0182] This disclosure can also be embodied as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that can store data that can subsequently be read by a computer system. Examples of computer-readable media include hard disk drives, network attached storage devices (NAS), read-only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. The computer-readable medium may include computer-readable tangible media distributed across network-coupled computer systems, enabling the distributed storage and execution of computer-readable code.
[0183] Although the method operations are described in a specific order, it should be understood that other housekeeping operations may be performed between operations, or operations may be adjusted so that they occur at slightly different times, or they may be distributed in a system that allows processing operations to occur at various intervals associated with the processing, as long as the processing of the covering operations is performed in the desired manner.
[0184] Although the foregoing disclosure has been described in some detail for clarity of understanding, it will be apparent that changes and modifications may be practiced within the scope of the appended claims. Therefore, this embodiment is to be considered illustrative rather than restrictive, and embodiments of this disclosure are not limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.
Claims
1. A method comprising: Receive game status data and user data from multiple players playing video games, representing multiple gameplay sessions; Statistical information and facts are generated for the game session of the video game based on the game status data and the user data, wherein the multiple game playing processes include the game playing process of the game session, and the game session is a live broadcast of the video game; The camera perspective of the game session scene is generated using artificial intelligence, wherein the camera perspective is selected by the viewer; as well as Based on the game state data and the user data, the artificial intelligence is used to automatically generate machine-based narration for the camera view of the scene in the game session in real time.
2. The method of claim 1, wherein generating the machine-based interpretation comprises: The machine-based commentary is generated using an AI model configured to select a subset of statistics and facts of potential audience interest from statistics and facts generated for the game session.
3. The method according to claim 2, further comprising: Select a subset of statistics and facts that have the highest potential audience interest from the statistics and facts generated for the game session.
4. The method according to claim 1, further comprising: Identify one or more regions of interest (ROIs) that may be viewed by at least one viewer, each of which is associated with one or more scenes in the game session. One or more camera perspectives provide one or more views of each of the one or more scenes.
5. The method according to claim 4, further comprising: A device that streams the camera view of the scene and the machine-based narration to the audience via a network.
6. The method according to claim 4, further comprising: The audience interest areas are selected and identified by the audience.
7. The method according to claim 4, further comprising: Highlights of the game session for the video game are generated. The highlights include at least one camera view from multiple image frames of the scene and the machine-based narration.
8. The method of claim 1, wherein generating the machine-based interpretation comprises: The context of the scenario in the game session is determined based on the statistical information and facts generated for the game session, as well as the game state data and user data from the multiple game play processes; as well as Use the narration template to construct the scenario.
9. A non-transitory computer-readable medium for performing a method, the computer-readable medium comprising: Program instructions for receiving game status data and user data from multiple players playing video games; Program instructions for generating statistics and facts for the game session of the video game based on the game state data and the user data, wherein the plurality of game playing processes include the game playing process of the game session; Program instructions for using artificial intelligence to generate camera views of the scene in the game session, wherein the camera views are selected by the viewer; as well as The program instructions are used to automatically generate machine-based narration for the camera view of the game session scene in real time, based on the game state data and the user data, using the artificial intelligence.
10. The non-transitory computer-readable medium of claim 9, wherein the program instructions for generating the machine-based explanation comprise: Program instructions for generating the machine-based commentary using an AI model, the AI model being configured to select a subset of statistics and facts of potential audience interest from statistics and facts generated for the game session for the machine-based commentary.
11. The non-transitory computer-readable medium of claim 10, further comprising: Program instructions for selecting a subset of statistics and facts that have the highest potential audience interest from the statistics and facts generated for the game session.
12. The non-transitory computer-readable medium of claim 9, further comprising: Program instructions for identifying one or more regions of interest that may be viewed by at least one viewer, each of the one or more regions of interest being associated with one or more scenes of the game session, wherein one or more camera perspectives provide one or more views of each of the one or more scenes; as well as Program instructions for a device used to stream the camera view of a scene and the machine-based narration to a viewer via a network.
13. The non-transitory computer-readable medium of claim 12, further comprising: Program instructions for generating highlights of the game session of the video game. The highlights include at least one camera view from multiple image frames of the scene and the machine-based narration.
14. The non-transitory computer-readable medium of claim 9, wherein the program instructions for generating the machine-based explanation comprise: Program instructions for determining the context of a scenario in the game session based on statistical information and facts generated for the game session, as well as game state data and user data from the multiple game play processes; as well as Program instructions used to construct the scenario using the narration template.
15. A computer system comprising: processor; as well as A memory coupled to the processor and storing instructions therein, the instructions causing the computer system to perform a method when executed by the computer system, the method comprising: Receive game status data and user data from multiple players playing video games, representing multiple gameplay sessions; Statistical information and facts are generated for the game session of the video game based on the game status data and the user data, wherein the multiple game playing processes include the game playing process of the game session; The camera perspective of the game session scene is generated using artificial intelligence, wherein the camera perspective is selected by the viewer; and Based on the game state data and the user data, the artificial intelligence is used to automatically generate machine-based narration for the camera view of the scene in the game session in real time.
16. The computer system of claim 15, wherein generating the machine-based interpretation in the method comprises: The machine-based commentary is generated using an AI model configured to select a subset of statistics and facts of potential audience interest from statistics and facts generated for the game session.
17. The computer system according to claim 16, the method further comprising: Select a subset of statistics and facts that have the highest potential audience interest from the statistics and facts generated for the game session.
18. The computer system according to claim 15, the method further comprising: Identify one or more regions of interest that may be viewed by at least one viewer, each of the one or more regions of interest being associated with one or more scenes of the game session, wherein one or more camera perspectives provide one or more views of each of the one or more scenes; as well as A device that streams the camera view of the scene and the machine-based narration to the audience via a network.
19. The computer system according to claim 18, the method further comprising: Highlights of the game session for the video game are generated. The highlights include at least one camera view from multiple image frames of the scene and the machine-based narration.
20. The computer system of claim 15, wherein generating the machine-based interpretation in the method comprises: The context of the scenario in the game session is determined based on the statistical information and facts generated for the game session, as well as the game state data and user data from the multiple game play processes; as well as Use the narration template to construct the scenario.