Text extraction to individually encode text and images for streaming during low connectivity periods
By separating the text and scene rendering of image frames in video games and using artificial intelligence model processing, the text unreadable problem caused by excessive image compression during low connectivity periods is solved, high-resolution text transmission and efficient image compression are achieved, and the gaming experience is improved.
Patent Information
- Application Number
- CN202480010210.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-02
- Filing Date
- 2024-01-25
- Publication Date
- 2025-08-29
AI Technical Summary
During the low connectivity period of video games, image information is overcompressed, which makes it impossible for players to interpret text information effectively, affecting the game experience.
By separating the text and scene rendering in the image frame on the server side and processing and encoding separately, the text is streamed in high resolution while the image is streamed in lower resolution, using artificial intelligence models to identify rendering commands to achieve this separation.
Maintain high-resolution transmission of text during low network connectivity periods, ensuring that players can clearly read text information in the game, while allowing images to be transmitted at a higher compression rate, improving the overall gaming experience.
Smart Images

Figure CN120569245A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to games, and more specifically, to separating text and image information (e.g., background scenes) presented for rendering for separate processing by a graphics processor and encoder, such that during periods of low connectivity between a server and a client device (e.g., a game streaming service), text information can be delivered with little or no compression, while image information can undergo additional compression to accommodate lower bandwidth. Furthermore, an artificial intelligence model can be constructed to identify which commands in a command buffer are used to render text and which are used to render images. Furthermore, the separation of text information provides for more targeted manipulation of either text information or image information using artificial intelligence. Background Art
[0002] Video games and / or gaming applications and their related industries (e.g., the video game industry) are extremely popular and represent a large percentage of the entertainment market worldwide. Video games are played anywhere and anytime using various types of platforms (including game consoles, desktop or laptop computers, mobile phones, etc.).
[0003] Video games can be streamed from a backend server, such as in a cloud gaming configuration, where the video game is executed on the backend server and the video game gameplay is streamed to the client device over the network. During the video game gameplay, in addition to the background scene presented as a sequence of image frames, there can also be a user interface that provides useful information or game-generated text or other text inserted into and / or overlaying the background scene. In this way, players can receive information related to the video game gameplay presented in various formats besides audio. For example, hearing-impaired players rely on closed captions when playing video games. In other examples, many players actively turn off the audio and play video games with only closed captions for various reasons (such as less distraction, faster understanding of the game environment, etc.).
[0004] During periods of low connectivity between the backend server and the client device, increased compression is applied to the data being transmitted to the client device. This information includes image frames, the user interface, and / or text provided within the user interface. In other words, compression is applied to all data. On the one hand, the reduced resolution of image frames for a period of time due to increased compression generally does not degrade the player experience, as the storyline and action typically presented during gameplay remain unaffected. On the other hand, however, the reduced text resolution can impair the player experience, as the player may be unable to decipher the text. Consequently, the information provided via text is compressed so much during periods of low connectivity that it is rendered useless to the player.
[0005] It is against this backdrop that embodiments of the present disclosure emerge. Summary of the Invention
[0006] Embodiments of the present disclosure relate to extracting text from a rendering pipeline to separate the rendering of text and the rendering of a scene in an image frame, wherein the rendered scene is encoded, wherein the original text and the encoded scene are streamed separately to a client device so that the text is streamed at a higher resolution than if the text were encoded together with the scene, especially when experiencing low connectivity between the server and the client device. In addition, an artificial intelligence model can be built to identify which commands in the command buffer are used to render text and which commands are used to render images. The separation of text information uses artificial intelligence that is specific to only the text information or the image information to provide more isolated operations.
[0007] In one embodiment, a method is disclosed. The method includes executing a video game on a server to generate multiple draw calls for execution by one or more graphics processing units (GPUs) to render an image frame. The method includes identifying a first command set of one or more draw calls for rendering text in the image frame. The method includes identifying a second command set of one or more draw calls for rendering a scene in the image frame. The method includes executing the first command set to render the text. The method includes executing the second command set to render the scene independently of rendering the text. The method includes encoding the rendered scene to generate an encoded scene. The method includes separately streaming the text and the encoded scene to a client device.
[0008] In another embodiment, a non-transitory computer-readable medium storing a computer program for implementing a method is disclosed. The computer-readable medium includes program instructions for executing a video game on a server to generate a plurality of draw calls for execution by one or more graphics processing units (GPUs) to render an image frame. The computer-readable medium includes program instructions for identifying a first command set of one or more draw calls for rendering text in an image frame. The computer-readable medium also includes program instructions for identifying a second command set of one or more draw calls for rendering a scene in an image frame. The computer-readable medium includes program instructions for executing the first command set to render the text. The computer-readable medium also includes program instructions for executing the second command set to render the scene independently of rendering the text. The computer-readable medium also includes program instructions for generating an encoded scene by encoding the rendered scene. The computer-readable medium also includes program instructions for streaming the text and the encoded scene separately to a client device.
[0009] In yet another embodiment, a computer system is disclosed, comprising a processor and a memory coupled to the processor and storing instructions that, if executed by the computer system, cause the computer system to perform a method. The method includes executing a video game on a server to generate a plurality of draw calls for execution by one or more graphics processing units (GPUs) to render an image frame. The method includes a first command set of one or more draw calls for rendering text in the image frame. The method includes identifying a second command set of one or more draw calls for rendering a scene in the image frame. The method includes executing the first command set to render the text. The method includes executing the second command set to render the scene independently of rendering the text. The method includes generating an encoded scene by encoding the rendered scene. The method includes separately streaming the text and the encoded scene to a client device.
[0010] Other aspects of the disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, illustrating by way of example the principles of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The present disclosure may best be understood by reference to the following description taken in conjunction with the accompanying drawings, in which:
[0012] Figure 1 A system including a text parser configured to extract text from a rendering pipeline used to generate images and text that may be included in an overlay is shown according to one embodiment of the present disclosure.
[0013] Figure 2 is a flowchart illustrating a method for extracting text from a rendering pipeline configured to generate an image sequence of a scene so as to encode text and images separately during a period of low network connectivity according to one embodiment of the present disclosure.
[0014] Figure 3 is a diagram illustrating a rendering and encoding process of extracting text from a rendering pipeline to separate an image for text and a scene including an object according to one embodiment of the present disclosure.
[0015] Figure 4A is a general representation of an Image Generation AI (IGAI) processing sequence according to one embodiment.
[0016] Figure 4B Additional processing that may be performed on the input provided to the IGAI processing sequence depicted in FIG. 6A is shown in accordance with one embodiment of the present invention.
[0017] Figure 4CIt is shown how the output of the encoder used is then fed into the latent space processing in the IGAI processing sequence according to one embodiment.
[0018] Figure 5 Components of an example device that may be used to perform aspects of various embodiments of the present disclosure are shown. DETAILED DESCRIPTION
[0019] Although the following detailed description includes many specific details for the purpose of illustration, those skilled in the art will appreciate that many variations and modifications to the following details are within the scope of the present disclosure. Accordingly, the various aspects of the present disclosure described below are set forth without losing the generality of, and without imposing limitations on, the claims appended hereto.
[0020] Generally speaking, various embodiments of the present disclosure describe systems and methods for extracting text from a rendering pipeline configured to generate an image of a scene including one or more objects. The text may overlay one or more images and may further be included within a user interface overlaying the one or more images. Extracting text from images of a scene is useful during periods of low network connectivity. For example, text and images including one or more objects can be rendered separately, where the images may be encoded while the text remains unencoded (i.e., raw) in preparation for streaming to a client device. In this manner, and advantageously, while the images can be streamed using higher compression to accommodate lower network connectivity, the text can be streamed at a higher resolution, allowing the user to fully understand the text even though the images may be at a lower resolution during the same period of lower network connectivity. Furthermore, artificial intelligence (AI) techniques can be implemented to build an AI model that identifies CPU-generated drawing commands, thereby separating drawing commands for text and / or overlays from drawing commands for the corresponding images and / or objects in the images. As another advantage, the separation of text information provides for more isolated operations using additional AI processes specific to either text or image information.
[0021] Throughout this specification, references to "game," "video game," or "gaming application" are intended to refer to any type of interactive application that is initiated by executing input commands. For illustrative purposes only, interactive applications include applications for gaming, word processing, video processing, video game processing, and the like. Furthermore, the terms "virtual world," "virtual environment," or "metaverse" are intended to refer to any type of environment generated by a corresponding application or applications for interaction between multiple users in a multi-player session or multi-player gaming session. Furthermore, the terms introduced above are interchangeable.
[0022] With the above general understanding of the various embodiments in mind, example details of the embodiments will now be described with reference to the various figures.
[0023] Figure 1 1 is a diagram of a system 100 for extracting text from a rendering pipeline according to one embodiment of the present disclosure, wherein the text is rendered separately from the corresponding image and / or image objects so that the encoding process of the text and the image can be handled separately during periods of low network connectivity. In this way, when the information is streamed to the client device, the image can be compressed more while the text remains uncompressed so that the information conveyed in the text is not lost.
[0024] In one embodiment, system 100 is configured to provide gaming between one or more cloud gaming servers over a network. Cloud gaming involves executing a video game on a server to generate game-rendered video frames, which are then transmitted to a client for display. While embodiments of the present disclosure are described for extracting text when streaming cloud gaming content, it should be understood that extracting text and / or overlay information from image information (including image objects) can be performed when streaming any information including text, such as when streaming video content including text and / or an overlay with text information (e.g., a movie).
[0025] Figure 1 The implementation of multiple graphics processing units (GPUs) when processing an application is illustrated, where the multiple GPUs collaborate to process images or data. It should also be understood that, in various embodiments, multi-GPU execution can be performed using physical GPUs, virtual GPUs, or a combination of both. For example, a hypervisor can be used on host hardware (e.g., located in a data center) that utilizes one or more components of a hardware layer (such as multiple CPUs, memory modules, GPUs, network interfaces, communication components, etc.) to create a virtual machine (e.g., an instance). These physical resources can be arranged in racks, such as CPU racks, GPU racks, memory racks, etc., where the physical resources in the racks can be accessed using a top-of-rack switch that facilitates the fabric for assembling and accessing components for the instance (e.g., when building virtualized components for the instance). Typically, the hypervisor can present multiple guest operating systems configured with multiple instances of virtual resources. That is, each operating system can be configured with a corresponding set of virtualized resources supported by one or more hardware resources (e.g., located in a corresponding data center). For example, each operating system can be supported by a virtual CPU, multiple virtual GPUs, virtual memory, virtualized communication components, etc.
[0026] According to one embodiment of the present disclosure, system 100 provides games via cloud gaming network 190, where the games are executed remotely from the corresponding users playing the games on client devices 110 (e.g., thin clients). System 100 can provide game control in single-player or multi-player modes to one or more users playing one or more games through cloud gaming network 190 via network 150. In some embodiments, cloud gaming network 190 may include multiple virtual machines (VMs) running on a host's hypervisor, where one or more VMs are configured to utilize hardware resources available to the host's hypervisor to execute game processor modules. Network 150 may include one or more communication technologies. In some embodiments, network 150 may include fifth-generation (5G) network technology with advanced wireless communication systems.
[0027] As shown, cloud gaming network 190 includes a game server 160 that provides access to multiple video games. Game server 160 can be any type of server computing device available in the cloud and can be configured as one or more virtual machines executing on one or more host computers. For example, game server 160 can manage virtual machines that support game processors that instantiate game instances for users. Thus, the multiple game processors of game server 160, associated with multiple virtual machines, are configured to execute multiple instances of one or more games associated with multiple users' gameplay. In this way, the backend server supports streaming of gameplay media (e.g., video, audio, etc.) for multiple game applications to multiple corresponding users. That is, game server 160 is configured to stream data (e.g., rendered images and / or frames of the corresponding gameplay) back to the corresponding client devices 110 via network 150. In this way, computationally complex game applications can be executed on the backend server in response to controller inputs received and forwarded by client devices 110. Each server can render images and / or frames, and then encode (eg, compress) and stream the images and / or frames to a corresponding client device for display.
[0028] For example, multiple users can access the cloud gaming network 190 via the communication network 150 using corresponding client devices 110 configured to receive streaming media. In one embodiment, the client devices 110 may be configured as thin clients that interface with a backend server (e.g., the cloud gaming network 190) configured to provide computing functionality (e.g., including the game title processing engine 111). In another embodiment, the client devices 110 may be configured with a game title processing engine and game logic for at least some local processing of the video game, and may further be configured to receive streaming content generated by the video game executed at the backend server, or other content provided by the backend server. For local processing, the game title processing engine includes basic processor-based functionality for executing the video game and services associated with the video game. In this case, the game logic may be stored locally on the client device 110 and used to execute the video game.
[0029] Each of the client devices 110 may request access to a different game from the cloud gaming network. For example, the cloud gaming network 190 may execute one or more game logics built on the game name processing engine 111, such as using the CPU resources 163 and GPU resources 165 of the game server 160. For example, game logic 115a that cooperates with the game name processing engine 111 may be executed on the game server 160 for one client, game logic 115b that cooperates with the game name processing engine 111 may be executed on the game server 160 for a second client, and game logic 115n that cooperates with the game name processing engine 111 may be executed on the game server 160 for an Nth client.
[0030] Specifically, a client device 110 of a corresponding user (not shown) is configured to request access to a game via a communication network 150 (such as the Internet) and to render images generated by a video game executed by a game server 160 for display. Encoded images (i.e., one or more compressed images) generated by an encoder 167 are delivered to the client device 110 for display in association with the corresponding user. For example, a user can interact with an instance of a video game executed on a game processor of the game server 160 via the client device 110. More specifically, the instance of the video game is executed by a game title processing engine 111. Game logic (e.g., executable code) 115 implementing the video game is stored and accessible via a data storage area (not shown) and used to execute the video game. The game title processing engine 111 is capable of supporting multiple video games using multiple game logics (e.g., game applications), each of which can be selected by the user.
[0031] For example, client device 110 is configured to interact with a game title processing engine 111, which is associated with the corresponding user's gameplay, such as through input commands used to drive gameplay. Specifically, client device 110 can receive input from various types of input devices, such as a game controller, tablet, keyboard, camera-captured gestures, mouse, touchpad, and the like. Client device 110 can be any type of computing device having at least memory and a processor module that can connect to game server 160 via network 150. The back-end game title processing engine 111 is configured to generate rendered images, which are delivered via network 150 for display on a corresponding display associated with client device 110. For example, via a cloud-based service, the game rendered images may be delivered by an instance of the corresponding game (e.g., game logic) executing on the game execution engine 111 of game server 160. That is, client device 110 is configured to receive encoded images (e.g., encoded from game rendered images generated by executing a video game) and display the rendered images on display 11. In one embodiment, display 11 comprises an HMD (e.g., displaying VR content). In some implementations, the rendered image can be streamed wirelessly or wired to a smartphone or tablet directly from a cloud-based service or via a client device 110 (eg, Play Station® Remote Play).
[0032] In one embodiment, game server 160 and / or game title processing engine 111 include basic processor-based functionality for executing games and services associated with gaming applications. For example, game server 160 includes central processing unit (CPU) resources 163 and graphics processing unit (GPU) resources 165 configured to perform processor-based functions, including 2D or 3D rendering, physics simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadowing, culling, transformations, artificial intelligence, and the like. Furthermore, the CPU and GPU combination may implement services for gaming applications, including, in part, memory management, multi-threading management, quality of service (QoS), bandwidth testing, social networking, management of social friends, communication with friends on social networks, communication channels, texting, instant messaging, and chat support. In one embodiment, one or more applications share specific GPU resources. In one embodiment, multiple GPU devices may be combined to perform graphics processing for a single application executing on a corresponding CPU.
[0033] Furthermore, system 100 includes a text extractor 120 configured to extract text from the rendering pipeline in order to separate the rendering of text from the rendering of images (e.g., objects in a scene generated by executing a video game). A network monitor 169 is configured to perform quality of service (QoS) monitoring between the game server and the corresponding client device. This allows network monitor 169 to detect and / or predict when the network connection between the server and the client device is performing poorly, or falls below a threshold metric, and trigger a text extraction mode accordingly. Furthermore, network monitor 169 is configured to determine when the network connection has returned to normal or has risen above the threshold metric and terminate text extraction mode, allowing encoding of text and images and / or image objects to proceed normally (i.e., encoding both text and images together). Specifically, during periods of low network connectivity between the cloud gaming network 190 and the corresponding client device 110, images can be encoded with higher compression than during normal network connectivity, while text can remain unencoded or lightly compressed. In this way, while images are received at a low resolution at the client device, text is still received at a high resolution, resulting in no loss in the transmission of information conveyed by the text.
[0034] pass Figure 1 A detailed description of the system 100, Figure 2 Flowchart 200 discloses a method for extracting text from a rendering pipeline configured to generate an image sequence of a scene so as to encode text and images separately during a period of low network connectivity according to one embodiment of the present disclosure. The operations performed in the flowchart may be performed by one or more of the previously described entity components and Figure 1 The system 100 described in
[15] (including, in part, the text extractor 120) can be implemented. Text extraction can be performed during periods of low network connectivity between the cloud gaming server and the client device, allowing images to be encoded at higher compression (e.g., when it is beneficial to the smoothness of the video, such as when all image frames of a video sequence are presented), while the text remains unencoded or lightly encoded when streaming the data to the client device. In this way, while the images can be streamed at a low resolution, the text is streamed at a high resolution, allowing the information in the text to be fully conveyed.
[0035] At 210, the method includes executing a video game on a server. For example, a request may be received over a network to establish a game session (e.g., a single-player or multi-player session) for a user playing the video game. During the game session, an instance of the video game is instantiated at the cloud gaming server, as previously described. Specifically, in embodiments of the present disclosure, the server strives to generate game-rendered image frames within one or more consecutive frame periods, and more specifically within a frame period.
[0036] For example, during the execution of a video game, image frames are generated in response to control information (e.g., user input commands) or game logic that is not driven by control information. More specifically, the CPU generates multiple draw calls for execution by one or more GPUs to render image frames, such as images comprising a scene and / or objects of the scene. The draw calls are stored in one or more command buffers that are executed by one or more selected GPUs, each of which implements a graphics pipeline or rendering pipeline.
[0037] Additionally, the image frame may include rendered text, wherein the text conveys information or a message. The text may be formatted in any language, such as providing words using the English alphabet, or words derived from pictograms (e.g., Chinese characters, Japanese Kanji, Islamic calligraphy, etc.). The text may also be included within a user interface, wherein the text and / or an overlay including the text is then overlaid onto the image for display to the user at the client device.
[0038] One or more GPUs, as scheduled by the CPU or a scheduler, can be implemented to render images (e.g., scenes and / or scene objects) and text. Typically, a draw call can be directed to a scene or scene object within an image frame, or to text and / or a user interface providing text. One or more assets required for a draw call are also loaded into system memory for use by the one or more GPUs to render the image frame. Generally speaking, the CPU can write a draw call to one or more command buffers for execution by one or more selected GPUs (i.e., selected by the scheduler), where the command buffers are contained within system memory or may be included within the corresponding GPU. In other words, a draw call is input into the rendering pipeline. Each command buffer is configured to store a corresponding draw call command. Furthermore, corresponding assets are also stored and made available to the GPU executing the draw call in the command buffer. Thus, the CPU uses the GPU API to write the draw call commands to the corresponding command buffer(s), where one or more command buffers can be used to render the image frame. Subsequent video frames are rendered using similarly configured command buffers.
[0039] In one embodiment, network connectivity between a server and a client device receiving a plurality of image frames generated by executing a video game is monitored. For example, the network connectivity is monitored with respect to a plurality of quality of service (QoS) metrics. Furthermore, the method includes determining when network connectivity is poor or below a QoS threshold. Furthermore, the method may include predicting when network connectivity is likely to drop below a QoS threshold. In this manner, when network connectivity falls below the QoS threshold or is predicted to fall below the QoS threshold, a text extraction mode may be triggered, causing text and scene (e.g., scene objects) to proceed along separate processing paths for rendering and / or encoding. Specifically, text from the rendering pipeline may be extracted for improved remote playback of the text (i.e., at the client device) during periods of low network connectivity. The text is not encoded (or lightly encoded) and is streamed (i.e., transmitted) to the client device separately, providing improved rendering and display of the text at the client device at a higher resolution than would typically be achieved if the text were encoded alongside images (e.g., scene and / or scene objects) during periods of low connectivity.
[0040] Additionally, the method may include determining when the network connection returns to normal (e.g., above a QoS threshold), wherein the text extraction mode may be exited and the image (e.g., a scene and / or objects of the scene) and text and / or a user interface including the text rendered and encoded normally for streaming to the client device.
[0041] At 220, during a text extraction mode, the method includes identifying a first command set of one or more draw calls for rendering text in an image frame. The draw calls may be identified before the CPU writes the draw calls to one or more command buffers for scheduling and execution by one or more selected GPUs. In some implementations, the draw calls may be identified after the CPU writes the command buffers for execution, in which case a new command buffer is written to separate the rendering process for the text and the scene (including objects of the scene).
[0042] For example, the text may include information presented in any of a variety of formats (e.g., English characters, Japanese Kanji, Chinese characters, Islamic calligraphy, etc.). The text may be provided within a user interface. For example, the user interface may provide health information for a character included within a status interface, or the user interface may include communications provided in thought bubbles, or may include information provided in a chat forum.
[0043] Information conveyed in the text may include player status information, health information, game status information, communications from the game as game-generated text, and communications from other players, including chat forums. In some cases, the text is presented as closed captions to supplement the gameplay. For example, some players may be hearing-impaired or simply prefer to play without sound in order to better focus and more quickly realize what is happening in the video game.
[0044] At 230, during text extraction mode, the method includes identifying a second command set of one or more draw calls used to render a scene (i.e., an image) or an object of a scene in an image frame. This can be determined passively, as the remaining draw calls that are not identified as text or providing coverage of text will be identified as being used to render the scene and / or objects of the scene. In another embodiment, the identification of draw calls used to render the scene and / or objects of the scene is determined actively.
[0045] At 240, the method includes executing a first command set to render text. During text extraction rendering mode, separate rendering pipelines are used to render text and images of a scene including objects of the scene. Specifically, draw calls for text are identified and sent along one path for storage in one or more corresponding command buffers for execution by one or more GPUs. Draw calls for images of the scene or objects of the scene in an image frame are sent along another path for storage in other corresponding command buffers for execution by one or more GPUs. In this manner, draw calls are input into the rendering pipeline and executed by the corresponding GPU to render text and / or a user interface including text in the image frame. In other words, the first command set is used to render the user interface and text in combination. The rendered text and / or user interface including text can be placed on a corresponding display or frame buffer. In some embodiments, the same GPU can still be used to render text and / or images of a scene (and / or objects of the scene), as long as the output data for the text can be kept separate from the data for the images of the scene and / or objects of the scene.
[0046] At 250, the method includes executing a second set of commands to render an image of a scene or an object of the scene independently of rendering text. During the text extraction rendering mode, draw calls for scenes (e.g., images) and / or objects of the scene of an image frame are sent on a different path than draw calls for text for storage in one or more corresponding command buffers for execution by one or more GPUs. Thus, the draw calls are input into the rendering pipeline and executed by the corresponding GPU to render the scene and / or objects of the scene of the image frame. The rendered images of the scene and / or objects of the scene are placed on a corresponding display or frame buffer.
[0047] At 260, the method includes generating an encoded image by encoding the rendered image of the scene or an object of the scene, as accessed from the display buffer. Specifically, the rendered image of the scene and / or the object of the scene is sent to an encoder for compression. Because network connectivity is determined to be low, a higher compression may be applied to the data. That is, the rendered image of the scene and / or the object of the scene may be encoded at a low resolution, where the rendered image of the scene and / or the object of the scene would typically be encoded at a high resolution.
[0048] More specifically, because the text is extracted and rendered separately from the image of the scene and / or the object of the scene, the rendered text and / or the user interface including the text can remain in its original format and not be encoded by the encoder. In some embodiments, the rendered text and / or the user interface including the text can be lightly encoded. In this way, the client device receives the rendered text and / or the user interface including high-resolution text. If encoded normally, the rendered text and / or the user interface including the text will be encoded similarly to the image of the rendered scene and / or the object of the scene (i.e., to result in high compression at low resolution).
[0049] At 270, during a text extraction mode implemented during a period of low connectivity, the method includes separately streaming the text and an encoded image of the scene or objects of the scene to the client device. That is, the rendered text and / or user interface including the text is not encoded (or lightly encoded) and is separately streamed to the client device. In addition, the rendered scene and / or objects of the scene are separately streamed to the client device after being encoded (i.e., highly compressed to a low resolution).
[0050] Thus, the client device is configured to reconstruct an image frame using the rendered text and / or a user interface including the text and the encoded image of the rendered scene and / or an object of the scene. Specifically, the encoded information is decoded to obtain an image of the rendered scene and / or an object of the scene (at a low resolution), and the text (at a high resolution) is superimposed on the image of the rendered scene and / or an object of the scene for display at the client device.
[0051] In one embodiment, timing information (timestamps, image frame identifiers, etc.) is delivered with the streamed data. That is, the timing information is included in the encoded and streamed rendered image of the scene and / or objects of the scene, and the timing information is also included in the streamed text and / or user interface including text. In this way, the client device can reconstruct the corresponding image frame using the appropriately rendered image of the scene and / or objects of the scene, with the objects of the scene overlaid with the appropriate text and / or user interface including text.
[0052] Figure 3 3 is a data flow diagram 300 illustrating a process of extracting text from a rendering pipeline to separate the rendering and encoding of text and images for a scene including an object according to one embodiment of the present disclosure. The operations shown in the data flow diagram 300 may be performed in Figure 1 The cloud gaming network 190 (eg, the game server 160, the text parser 120, etc.) is executed to implement Figure 2 200 of a method for extracting text from a rendering pipeline to separate and encode the rendering of text from the rendering of images (e.g., objects of a scene generated by executing a video game) during periods of low connectivity between a server and a corresponding client device.
[0053] As shown, CPU resources 163 are configured to generate multiple draw calls for execution by one or more GPUs to render one or more image frames. For example, an image frame may include an image containing a scene and / or objects of the scene. An image frame may also include text and / or a user interface including text, as previously described. The draw calls are stored in one or more command buffers for execution by one or more selected GPUs, each of which implements a graphics pipeline or rendering pipeline. Specifically, CPU resources 163 execute a video game to generate one or more draw calls for an image frame. For example, multiple draw call commands 315 are output by CPU resources 163. A draw call includes, for example, instructions and / or commands for execution by one or more corresponding GPUs implementing a graphics pipeline for rendering. For a particular image frame, there may be multiple draw calls generated by CPU resources 163 and executed by GPU resources 165. More specifically, as video games are designed to render increasingly complex scenes, this requires pushing more and more draw calls for each image frame to generate the scene during gameplay.
[0054] In one embodiment, when draw calls are generated by CPU resource 163, text parser 120 is configured to extract draw call commands for text and / or a user interface including text from draw call commands for a scene and / or an image of an object of the scene, as previously described. Specifically, text parser 120 is configured to output text / overlay draw call commands 320 (i.e., for rendering text and / or a user interface including text) and object / scene draw call commands 325 (i.e., for rendering an image of a scene and / or an object of the scene). Furthermore, text parser 120 may include a frame timer 125 that provides timing information (e.g., timestamps, image frame identifiers, etc.) for use in reconstructing various components of an image frame.
[0055] In some embodiments, text parser 120 includes artificial intelligence (AI), including a deep / machine learning engine 130. This deep / machine learning engine 130 is configured to build, train, and implement an AI model 135 for identifying draw calls, or commands within draw calls, that involve text and / or user interfaces that include text. Specifically, AI model 135 is configured to identify features within draw calls (i.e., commands) and further identify which draw calls within draw calls include features commonly used when rendering text. More specifically, AI model 135 is configured to identify a first set of commands used to render text and / or user interfaces that include text. In one embodiment, AI learning model 135 is a machine learning model configured to apply machine learning to learn and identify which commands within draw calls are used to render text and / or user interfaces that include text. In another embodiment, AI learning model 135 is a deep learning model configured to apply deep learning to learn and identify which commands within draw calls are used to render text and / or user interfaces that include text. Machine learning is a subcategory of artificial intelligence, and deep learning is a subcategory of machine learning.
[0056] Purely for illustration, according to one embodiment of the present disclosure, the deep / machine learning engine 190 can be configured as a neural network for training and / or implementing the AI model 135. Generally speaking, a neural network represents a network of interconnected nodes that responds to input (e.g., extracted features) and generates output (e.g., learning and identifying which commands in a draw call are used to render text and / or a user interface that includes text). In one embodiment, the AI neural network includes a hierarchical structure of nodes. For example, there can be an input layer of nodes, an output layer of nodes, and an intermediate or hidden layer of nodes. The input nodes are interconnected to hidden nodes in the hidden layer, and the hidden nodes are interconnected to the output nodes. The interconnections between the nodes can have numerical weights that can be used to link multiple nodes together between inputs and outputs, such as when defining rules for the AI model 135.
[0057] For example, the AI model 135 may be configured to analyze draw calls output by the CPU resource 163, or to analyze draw calls stored in a command buffer assigned to a specific GPU, to generate a set of command buffers for rendering text and / or a user interface including text for an image frame, and to generate another set of command buffers for rendering an image of a scene or an object of the scene in the image frame. In other words, the draw call text and / or the user interface including text are rendered on a separate pipeline compared to the draw call scene and / or the image of an object on the scene.
[0058] For example, in one embodiment, text extraction is performed at the kernel level in the operating system of a computing resource (e.g., a game server). Draw calls to be stored in a command buffer are intercepted, and draw calls for text or user interface elements and / or user interfaces that include text and other user interface elements are extracted. In one embodiment, training the AI model 135 in part using known draw calls for text in various embodiments allows these draw calls for text and / or user interface elements to be identified. The accuracy of the AI model 135 during training can be checked by making multiple draw calls and extracting draw calls for text from the draw calls for the image of the scene in the image frame. After rendering, the rendered text can be overlaid on the rendered image of the scene and compared to the original rendering of the multiple draw calls without text extraction and stored in the command buffer that generated the original image frame.
[0059] In this way, text in video games and other user interface elements are rendered separately from the action presented in the continuous image of the scene. That is, the image of the scene or its objects is rendered separately from the text and / or the user interface (i.e., user interface elements) that includes text. Because separate game binaries are used to render the game scene's images and text on different rendering pipelines, during periods of low connectivity, not only can the text be streamed at high resolution (i.e., unencoded) while the corresponding game scene images are encoded at high compression, but the binary streaming of both text and game scene images also provides more appropriate AI analysis. That is, AI analysis can focus directly on the text or images of the game scene.
[0060] System memory 330 may include multiple command buffers. Each command buffer is configured for a corresponding draw call command. Specifically, CPU resource 163 and / or text parser 120 uses the GPU API to write commands for a corresponding draw call to a corresponding command buffer, where one or more command buffers may be used to render an image frame. Subsequent video frames are rendered using similarly configured command buffers. As shown, due to the extraction of draw calls for rendering text, text / overlay draw calls 320 are stored in text command buffer 335 for execution by the corresponding GPU 340. For example, for an image frame, a first set of draw call commands may be stored in a corresponding text command buffer for execution by the GPU to render text and / or a user interface including text.
[0061] Additionally, the object / scene draw calls 325 are stored in an object / scene command buffer 337 for rendering by the corresponding GPU 340. Thus, for an image frame, a second command set of draw calls may be stored in a corresponding image command buffer or object / scene command buffer for execution by the GPU to render an image of the scene and / or objects of the scene.
[0062] One or more GPUs 340, as scheduled by the CPU and / or a text parser or a GPU scheduler, can be implemented to render images (e.g., scenes and / or objects of scenes) and text and / or user interfaces including text (e.g., user interface elements) corresponding to image frames. Thus, different rendering pipelines are used to render text and / or user interfaces including text, and images of scenes and / or objects of scenes, respectively. Specifically, each GPU 340 is configured to implement a graphics pipeline. For example, a graphics pipeline can execute a shader program on the vertices of objects within a scene to generate texture values for pixels of a display, with operations performed in parallel across the GPUs 340 for efficiency. Typically, a graphics pipeline receives input geometry (e.g., vertices of objects within a game world). A vertex shader constructs the polygons or primitives that make up the objects within the scene. Vertex shaders or other shader programs can perform lighting, shading, shading, and other operations on the polygons. Vertex shaders or other shader programs can also implement depth or z-buffering to determine which objects are visible in the scene rendered from a corresponding viewpoint. Rasterization is performed to project objects in the three-dimensional world onto a two-dimensional plane defined by the viewpoint. Pixel-sized fragments are generated for the objects, where one or more fragments can contribute to the color of a corresponding pixel when the image is displayed. The fragments are combined and / or blended to determine the combined color of each pixel, which is stored in the frame buffer for display.
[0063] GPU 340 outputs the rendered data to a display and / or frame buffer 370. Specifically, because text rendering and scene image rendering are performed on different rendering pipelines, the rendered text and / or a user interface including text (e.g., a user interface element) is stored in display buffer 370A. The rendered scene and / or scene object images are also stored in display buffer 370B.
[0064] As shown, during periods of low connectivity, encoder 350 is implemented to encode images of rendered scenes and / or objects of the scenes using higher-than-normal compression. The highly compressed (at low resolution) images of the rendered scenes and / or objects of the scenes are delivered to streamer 360, which manages the streaming of data to a client device (not shown). Additionally, rendered text and / or user interfaces including text bypass encoder 350 and are delivered directly to streamer 360, allowing the rendered text and / or user interfaces including text to be streamed to the client device without any encoding. In one embodiment, the rendered text and / or user interfaces including text are lightly compressed, or the compression level can be selected for optimal streaming during periods of low connectivity.
[0065] In one embodiment, the generation of output images, graphics, and / or 3D representations by an image generation AI (IGAI) may include one or more artificial intelligence processing engines and / or models. Typically, the AI model is generated using training data from a dataset. The dataset selected for training can be custom-curated for a specific desired output, and in some cases, the training dataset can include a wide range of general-purpose data that can be consumed from multiple sources via the internet. For example, the IGAI should have access to a large amount of data, such as images, videos, and 3D data. The IGAI uses this general-purpose data to gain an understanding of the type of content expected from the input. For example, if the input requests the generation of a tiger in the Sahara Desert, the dataset should contain a variety of images of tigers and deserts to access and utilize during the processing of the output image. On the other hand, a curated dataset can be more specific to a particular type of content, such as video game-related technology, videos, and other asset-related content. Even more specifically, a curated dataset may include images related to a specific scene in a game or an action sequence involving game assets (e.g., unique avatar characters, etc.). As described above, the IGAI can be customized to allow the input of a unique descriptive language statement to style the requested output image or content. The descriptive language statement can be text or other sensory input, such as inertial sensor data, input speed, emphasis statement, and other data that can be formed into an input request. An image, video, or image set can also be provided to the IGAI to define the context of the input request. In one embodiment, the input can be text describing the desired output and one or more images to convey the desired context scenario being requested as the output.
[0066] In one embodiment, IGAI is provided to implement text-to-image generation. The image generation is configured to implement a latent diffusion process in a latent space to synthesize text into the image processing. In one embodiment, a conditioning process helps shape the output toward the desired usage output (e.g., using structured metadata). The structured metadata can include information obtained from user input to guide the machine learning model to gradually denoise using cross-attention until the processed denoising is decoded back into pixel space. During the decoding stage, upscaling is applied to achieve higher quality images, videos, or 3D assets. Therefore, IGAI is a customized tool that is designed to process specific types of inputs and present specific types of outputs. When IGAI is customized, the machine learning and deep learning algorithms are adjusted to achieve specific customized outputs, such as, for example, unique image assets to be used in gaming technology, specific game titles, and / or movies.
[0067] In another configuration, IGAI can be a third-party processor, for example, such as a processor provided by Stable Diffusion or others (such as OpenAI's GLIDE, DALL-E, MidJourney, or Imagen). In some configurations, IGAI can be used online via one or more application programming interface (API) calls. It should be understood that references to available IGAI are for informational purposes only. For additional information related to IGAI technology, reference can be made to the paper entitled "High-Resolution Image Synthesis with Latent Diffusion Models" (Robin Rombach et al., pages 1-45) published by Ludwig Maximilian University of Munich. This paper is incorporated herein by reference.
[0068] Figure 4AThe following is a general representation of the processing sequence of Image Generation AI (IGAI) 402 according to one embodiment. As shown, input 406 is configured to receive input in the form of data, such as a text description with semantic descriptions or keywords. The text description can be in the form of a sentence, for example, containing at least a noun and a verb. The text description can also be in the form of a fragment or a simple single word. The text can also be in the form of multiple sentences, describing a scene, some action, or some characteristics. In some configurations, the input text can be entered in a specific order to influence the focus on one word over others, or even to de-emphasize words, letters, or sentences. Furthermore, the text input can be in any form, including characters, emoticons, icons, and foreign language characters (e.g., Japanese, Chinese, Korean, etc.). In one embodiment, text description is implemented through contrastive learning. The basic idea is to embed both the image and text in a latent space, so that the text corresponding to the image is mapped to the same region in the latent space as the image. This abstracts the structure of, for example, "what it means is a dog" from the visual and textual representations. In one embodiment, the goal of contrastive representation learning is to learn an embedding space in which similar pairs of examples remain close together, while dissimilar pairs are far apart. Contrastive learning can be applied in both supervised and unsupervised settings. When using unsupervised data, contrastive learning is one of the most powerful methods in self-supervised learning.
[0069] In addition to text, input can also include other content, such as images or even images that themselves have descriptive content. Image analysis can be used to interpret images to identify objects, colors, intent, characteristics, shading, textures, three-dimensional representations, depth data, and combinations thereof. Broadly speaking, input 406 is configured to convey the intention of a user who wishes to utilize IGAI to generate some digital content. In the context of gaming technology, the target content to be generated can be game assets used in a specific game scene. In this scenario, the data set used to train IGAI and input 406 can be used to customize the way artificial intelligence (e.g., a deep neural network) processes data to guide and tune the desired output image, data, or three-dimensional digital asset.
[0070] Input 406 is then passed to IGAI, where an encoder 408 takes the input data and / or pixel-space data and converts it into latent space data. The concept of "latent space" is central to deep learning, as feature data is reduced to a simplified data representation for the purpose of finding and exploiting patterns. Therefore, latent space processing 410 is performed on the compressed data, significantly reducing processing overhead compared to pixel-space processing and learning algorithms, which are much more expensive and require significantly more processing power and time to analyze and produce the desired image. The latent space is simply a representation of the compressed data, where similar data points are spatially closer together. Within the latent space, the processing is configured to learn relationships between data points that the machine learning system has been able to derive from the information it is fed (e.g., the dataset used to train IGAI). In latent space processing 410, a diffusion model is used to calculate the diffusion process. The latent diffusion model relies on an autoencoder to learn a low-dimensional representation of the pixel space. The latent representation is passed through a diffusion process, with noise added at each step, such as multiple stages. The output is then fed into a denoising network based on a U-Net architecture with a crisscross attention layer. A conditioning process is also applied to guide the machine learning model to remove noise and obtain an image that represents content that is closely related to the content requested via the user input. The decoder 412 then transforms the resulting output from the latent space back into the pixel space. The output 414 can then be processed to increase the resolution. The output 414 is then delivered as a result, which can be an image, a graphic, 3D data, or data that can be rendered into a physical or digital form.
[0071] Figure 4BThe following illustrates additional processing that can be performed on input 406 in one embodiment. User interface tools 420 can be used to enable a user to provide input request 404. As described above, input request 404 can be an image, text, structured text, or general data. In one embodiment, before providing the input request to encoder 408, the input can be processed by a machine learning process that generates a machine learning model 432 and learns from a training dataset 434. As an example, the input data can be processed by a context analyzer 426 to understand the context of the request. For example, if the input is "space rockets for flying to Mars," the input can be analyzed 426 to determine that the context relates to outer space and planets. Context analysis can use machine learning model 432 and training dataset 434 to find relevant images for this context or identify specific libraries of art, images, or videos. If the input request also includes an image of a rocket, feature extractor 428 can be used to automatically identify characteristic features in the rocket image, such as fuel tanks, length, color, location, edges, text, flames, etc. Feature classifier 430 can also be used to classify features and improve machine learning model 432. In one embodiment, input data 407 can be generated to produce structured information that can be encoded into a latent space by encoder 408. Additionally, structured metadata 422 can be extracted from the input request. Structured metadata 422 can be, for example, descriptive text that instructs IGAI 402 to modify a feature or make a change to the input image, or to change color, texture, or a combination thereof. For example, input request 404 can include an image of a rocket, and the text can say, "Make the rocket wider," "Add more flames," or "Make it stronger," or some other qualifier of the user's intent (e.g., semantically provided and contextually analyzed). Structured metadata 422 can then be used in subsequent latent space processing to adjust the output toward the user's intent. In one embodiment, structured metadata can be in the form of a semantic graph, text, image, or data designed to represent the user's intent regarding what changes or modifications should be made to the input image or content.
[0072] Figure 4CThe figure shows how the output of encoder 408 is then fed into latent space processing 410, according to one embodiment. A diffusion process is performed by diffusion process stage 440, where the input is processed through multiple stages to add noise to one or more input images associated with the input text. This is an incremental process, where noise is added at each stage, for example, 10-50 or more stages. Next, denoising is performed by denoising stage 442. Similar to the denoising stage, a reverse process is performed, where noise is gradually removed at each stage, and at each stage, machine learning is used to predict what the output image or content should look like, based on the input request intent. In one embodiment, structured metadata 422 can be used by machine learning model 444 at each denoising stage to predict how the resulting denoised image should look and how it should be modified. During these predictions, machine learning model 444 uses training dataset 446 and structured metadata 422 to increasingly closely resemble the output requested in the input. In one embodiment, a U-Net architecture with cross-attention layers can be used during denoising to improve predictions. After the final denoising stage, the output is provided to decoder 412, which transforms the output into pixel space. In one embodiment, the output is also upscaled to increase resolution. In one embodiment, the output of the decoder can optionally be run through a context adjuster 436. A context adjuster is a process that can use machine learning to examine the resulting output to make adjustments to make the output more realistic or to remove unrealistic or unnatural output. For example, if the input requires "a boy pushing a lawn mower" and the output shows a boy with three legs, the context adjuster can make adjustments using an underdrawing process or overlay to correct or prevent inconsistent or undesirable output. However, as the machine learning model 444 becomes smarter over time with more training, the context adjuster 436 will be less needed before the output is rendered in the user interface tool 420.
[0073] Figure 5Components of an example device 500 that can be used to perform aspects of various embodiments of the present disclosure are shown. This block diagram illustrates a device 500 suitable for practicing embodiments of the present disclosure, which may be incorporated into or may be a personal computer, video game console, personal digital assistant, server, or other digital device. Device 500 includes a central processing unit (CPU) 502 for running software applications and, optionally, an operating system. CPU 502 may include one or more homogeneous or heterogeneous processing cores. For example, CPU 502 may be one or more general-purpose microprocessors with one or more processing cores. Further embodiments may be implemented using one or more CPUs with a microprocessor architecture that is particularly well-suited for highly parallel and computationally intensive applications, such as interpreting queries, identifying context-sensitive resources, and immediately implementing and rendering context-sensitive resources within a video game. Device 500 can be located near the player playing the game segment (e.g., a game console), remotely from the player (e.g., a backend server processor), or in a game cloud system using one of many virtualized servers for remotely streaming gameplay to clients or for implementing additional services, such as supervisor functionality.
[0074] Specifically, according to one embodiment of the present disclosure, the CPU 502 may be configured to implement a text parser 120 having functionality configured to extract text from a rendering pipeline to separate the rendering and encoding processes for text (and / or user interface elements including a user interface that include text) and images of a scene including objects. Because the rendering of images of a scene or objects of a scene and the rendering of text and / or a user interface including text are performed on different rendering pipelines, the text and user interface elements can be streamed at high resolution (i.e., without compression), while the corresponding images of the scene and / or objects of the scene are encoded at high compression, such as during periods of low connectivity. In this way, after reconstruction at the client device, while the rendered images of the scene can be delivered at low resolution after decoding (e.g., to facilitate smooth playback without skipped frames), the text and / or user interface including text remains at high resolution, allowing the user to fully receive and / or understand the information conveyed by the text.
[0075] Memory 504 stores applications and data for use by CPU 502. Storage 506 provides non-volatile storage and other computer-readable media for applications and data and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. User input device 508 transmits user input from one or more users to device 500. Examples of device 500 may include a keyboard, mouse, joystick, touchpad, touch screen, still or video recorder / camera, tracking device for gesture recognition, and / or microphone. Network interface 514 allows device 500 to communicate with other computer systems via an electronic communications network and may include wired or wireless communications over local and wide area networks (such as the Internet). Audio processor 512 is adapted to generate analog or digital audio output from instructions and / or data provided by CPU 502, memory 504, and / or storage 506. The components of device 500 , including CPU 502 , memory 504 , data storage 506 , user input device 508 , network interface 510 , and audio processor 512 , are connected via one or more data buses 522 .
[0076] Graphics subsystem 520 further interfaces with data bus 522 and components of device 500. Graphics subsystem 520 includes a graphics processing unit (GPU) 516 and graphics memory 518. Graphics memory 518 comprises display memory (e.g., a frame buffer) for storing pixel data for each pixel of an output image. Graphics memory 518 may be integrated into the same device as GPU 516, interfaced with GPU 516 as a separate device, and / or implemented within memory 504. Pixel data may be provided directly to graphics memory 518 from CPU 502. Alternatively, CPU 502 provides data and / or instructions defining desired output images to GPU 516, which then generates pixel data for one or more output images from the data and / or instructions. The data and / or instructions defining the desired output images may be stored in memory 504 and / or graphics memory 518. In an embodiment, GPU 516 includes 3D rendering capabilities for generating pixel data for an output image from instructions and data defining a scene's geometry, lighting, shading, textures, motion, and / or camera parameters. GPU 516 may also include one or more programmable execution units capable of executing shader programs. In one embodiment, GPU 516 may be implemented within an AI engine (e.g., machine learning engine 190) to provide additional processing capabilities, such as for AI, machine learning, or deep learning functions.
[0077] Graphics subsystem 520 periodically outputs pixel data for an image from graphics memory 518 for display on display device 510. Display device 510 may be any device capable of displaying visual information in response to a signal from device 500, including CRT, LCD, plasma, and OLED displays. For example, device 500 may provide an analog or digital signal to display device 510.
[0078] In other embodiments, the graphics subsystem 520 includes multiple GPU devices that are combined to perform graphics processing for a single application executing on a corresponding CPU. For example, multiple GPUs can perform alternate forms of frame rendering, where GPU 1 renders the first frame in a consecutive frame cycle, GPU 2 renders the second frame in a consecutive frame cycle, and so on, until the last GPU is reached, at which point the initial GPU renders the next video frame (e.g., if there are only two GPUs, GPU 1 renders the third frame). In other words, the GPUs rotate when rendering frames. Rendering operations can overlap, where GPU 2 can begin rendering the second frame before GPU 1 finishes rendering the first frame. In another embodiment, different shader operations can be assigned to multiple GPU devices in the rendering and / or graphics pipeline. The primary GPU is performing primary rendering and compositing. For example, in a group of three GPUs, master GPU 1 can perform primary rendering (e.g., a first shader operation) and composite the outputs from slave GPUs 2 and 3. Slave GPU 2 can perform a second shader operation (e.g., a fluid effect such as a river) and slave GPU 3 can perform a third shader operation (e.g., particle smoke). Master GPU 1 composites the results from each of GPUs 1, 2, and 3. In this manner, different GPUs can be assigned to perform different shader operations (e.g., flag waving, wind, smoke generation, fire, etc.) to render a video frame. In yet another embodiment, each of the three GPUs can be assigned to a different object and / or portion of a scene corresponding to a video frame. In the above embodiments and implementations, these operations can be performed in the same frame cycle (simultaneously and in parallel) or in different frame cycles (sequentially and in parallel).
[0079] Thus, in various embodiments, the present disclosure describes systems and methods configured for extracting text from a rendering pipeline to separate a rendering and encoding process of text (and / or user interface elements of a user interface (including text)) and images of a scene including objects, such that, for example, during periods of low connectivity, the text and user interface elements can be streamed at high resolution (i.e., without compression) while corresponding images of the scene and / or objects of the scene are encoded at high compression.
[0080] It should be noted that access services delivered over a wide geographic area, such as providing access to games in the current implementation, often utilize cloud computing. Cloud computing is a style of computing in which dynamically scalable, often virtualized, resources are provided as services over the internet. Users do not need to be experts in the technical infrastructure supporting their "cloud." Cloud computing can be categorized into different services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). Cloud computing services typically provide common applications, such as video games, online, accessed from a web browser, while the software and data are stored on servers in the cloud. The term cloud is used as a metaphor for the internet, based on how it is depicted in computer network diagrams and as an abstraction of the complex infrastructure it hides.
[0081] In some embodiments, a game server can be used to operate a duration information platform for video game players. Most video games played over the internet operate via a connection to a game server. Typically, games utilize dedicated server applications that collect data from players and distribute it to other players. In other embodiments, video games may be executed by a distributed game engine. In these embodiments, a distributed game engine may execute on multiple processing entities (PEs), with each PE performing a functional segment of the given game engine on which the video game is running. Each processing entity is viewed by the game engine as a simple compute node. A game engine typically performs a variety of operations to execute the video game application and additional services that enhance the user experience. For example, a game engine implements game logic, performs game calculations, physics, geometric transformations, rendering, lighting, shading, audio, and additional in-game or game-related services. Additional services may include, for example, messaging, social utilities, audio communication, game play replay functionality, help functions, and the like. While a game engine may sometimes execute on an operating system virtualized by a specific server's hypervisor, in other embodiments, the game engine itself is distributed across multiple processing entities, each of which may reside on a different server unit in a data center.
[0082] Depending on the embodiment, the corresponding processing entity used to perform operations can be a server unit, a virtual machine, or a container, depending on the needs of each game engine segment. For example, if a game engine segment is responsible for camera transformations, that particular game engine segment may be equipped with a virtual machine associated with a graphics processing unit (GPU), as it will be performing a large number of relatively simple mathematical operations (e.g., matrix transformations). Other game engine segments requiring fewer but more complex operations may be equipped with processing entities associated with one or more higher-powered central processing units (CPUs).
[0083] By distributing the game engine, it is provided with elastic computing properties that are not constrained by the capacity of physical server units. Instead, the game engine can be provisioned with more or fewer computing nodes as needed to meet the needs of the video game. From the perspective of the video game and the video game player, a game engine distributed across multiple computing nodes is indistinguishable from a non-distributed game engine executed on a single processing entity, as the game engine manager or supervisor distributes the workload and seamlessly integrates the results to provide the video game output component to the end user.
[0084] A user accesses a remote service using a client device, which includes at least a CPU, a display, and I / O. The client device can be a PC, mobile phone, netbook, PDA, etc. In one embodiment, a network executed on the game server identifies the type of device used by the client and adjusts the communication method used. In other cases, the client device uses standard communication methods (such as HTML) to access the application on the game server via the Internet. It should be understood that a given video game or game application can be developed for a specific platform and a specific associated controller device. However, when such a game is made available via a game cloud system as presented herein, the user may be accessing the video game using different controller devices. For example, a game may have been developed for a game console and its associated controller, while the user may be accessing the cloud-based game version from a personal computer using a keyboard and mouse. In this case, the input parameter configuration can define a mapping from inputs that can be generated by the user's available controller devices (in this case, a keyboard and mouse) to inputs acceptable to the execution of the video game.
[0085] In another example, a user may access the cloud gaming system via a tablet computing device, a touchscreen smartphone, or other touchscreen driven device. In this case, the client device and the controller device are integrated together in the same device, where input is provided by detected touchscreen input / gestures. For such a device, the input parameter configuration may define a specific touchscreen input corresponding to a game input for a video game. For example, a button, directional pad, or other type of input element may be displayed or overlaid during the running of a video game to indicate a location on the touchscreen that a user can touch to generate game input. Gestures such as swipes in a specific direction or specific touch motions may also be detected as game inputs. In one embodiment, a tutorial may be provided to the user indicating how to provide input for gameplay via the touchscreen, for example, before starting gameplay of a video game, so as to acclimate the user to the operation of the controls on the touchscreen.
[0086] In some embodiments, the client device serves as a connection point for the controller device. That is, the controller device communicates with the client device via a wireless or wired connection to transmit input from the controller device to the client device. The client device can, in turn, process these inputs and then transmit the input data to the cloud gaming server via a network (e.g., accessed via a local networked device such as a router). However, in other embodiments, the controller itself can be a networked device with the ability to transmit input directly to the cloud gaming server via the network, without first requiring such input to be transmitted through the client device. For example, the controller can be connected to a local networked device (such as the aforementioned router) to send and receive data to and from the cloud gaming server. Thus, while the client device may still need to receive video output from the cloud-based video game and render it on a local display, input latency can be reduced by allowing the controller to send input directly to the cloud gaming server over the network, bypassing the client device.
[0087] In one embodiment, a networked controller and client device can be configured to send certain types of input directly from the controller to the cloud gaming server, and other types of input via the client device. For example, inputs whose detection does not rely on any additional hardware or processing beyond the controller itself can be sent directly from the controller to the cloud gaming server via the network, bypassing the client device. Such inputs can include button inputs, joystick inputs, embedded motion detection inputs (e.g., accelerometers, magnetometers, gyroscopes), and so on. However, inputs that utilize additional hardware or require client device processing can be sent by the client device to the cloud gaming server. These can include video or audio captured from the gaming environment, which can be processed by the client device before being sent to the cloud gaming server. Additionally, input from the controller's motion detection hardware can be processed by the client device in conjunction with captured video to detect the controller's position and motion, which is then transmitted by the client device to the cloud gaming server. It should be understood that controller devices according to various embodiments can also receive data (e.g., feedback data) from the client device or directly from the cloud gaming server.
[0088] Client devices access the cloud gaming network through a communication network that implements one or more communication technologies. In some embodiments, the network may include fifth-generation (5G) network technology with advanced wireless communication systems. 5G stands for fifth-generation cellular network technology. A 5G network is a digital cellular network in which the service area covered by a provider is divided into small geographic areas called cells. Analog signals representing sound and image are digitized in the phone, converted by an analog-to-digital converter, and transmitted as a bit stream. All 5G wireless devices in a cell communicate via radio waves with a local antenna array and low-power automated transceivers (transmitters and receivers) in the cell on a frequency channel allocated by the transceiver from a pool of frequencies reused in other cells. The local antennas connect to the telephone network and the internet via high-bandwidth fiber optic or wireless backhaul connections. As in other cellular networks, mobile devices crossing from one cell to another automatically transition to the new cell. It should be understood that 5G networks are merely an example type of communication network, and embodiments of the present disclosure can utilize earlier generations of wireless or wired communications, as well as later generations of wired or wireless technologies emerging after 5G.
[0089] In one embodiment, various technical examples can be implemented using a virtual environment via a head-mounted display (HMD). The HMD may also be referred to as a virtual reality (VR) headset. As used herein, the term "virtual reality" (VR) generally refers to user interaction with a virtual space / environment. This involves viewing the virtual space through the HMD (or VR headset) in a manner that responds in real time to the HMD's movements (e.g., controlled by the user), providing the user with the sensation of being within the virtual space, or metaverse. For example, a user may see a three-dimensional (3D) view of the virtual space when facing a given direction. When the user turns to one side, thereby also rotating the HMD, the view of that side of the virtual space is rendered on the HMD. The HMD can be worn similarly to glasses, goggles, or a helmet and is configured to display video games or other metaverse content to the user. The HMD can provide a highly immersive experience by providing a display mechanism that is in close proximity to the user's eyes. Thus, the HMD can provide a display area for each eye that occupies a large portion or even the entire field of view of the user and can also provide viewing with three-dimensional depth and perspective.
[0090] In one embodiment, the HMD may include a gaze-tracking camera configured to capture images of the user's eyes as the user interacts with the VR scene. The gaze information captured by the gaze-tracking camera may include information related to the user's gaze direction and the specific virtual objects and content items in the VR scene that the user is focused on or interested in interacting with. Thus, based on the user's gaze direction, the system can detect specific virtual objects and content items that may be focused on the user, such as game characters, game objects, game items, etc., with which the user is interested in interacting and engaging.
[0091] In some embodiments, the HMD may include (multiple) outward-facing cameras configured to capture images of the user's real-world space, such as the user's body movements and any real-world objects that may be located in the real-world space. In some embodiments, the images captured by the outward-facing cameras can be analyzed to determine the position / orientation of real-world objects relative to the HMD. Using the known position / orientation of the HMD, real-world objects, and inertial sensor data from the user's gestures and movements, the user's gestures and movements can be continuously monitored and tracked during the user's interaction with the VR scene. For example, when interacting with a scene in a game, a user can perform various gestures, such as pointing and walking towards specific content items in the scene. In one embodiment, gestures can be tracked and processed by the system to generate predictions of interactions with specific content items in the game scene. In some embodiments, machine learning can be used to facilitate or assist in the predictions.
[0092] During use of the HMD, various single-handed and two-handed controllers can be used. In some embodiments, the controllers themselves can be tracked by tracking lights included in the controllers or by tracking shapes, sensors, and inertial data associated with the controllers. Using these various types of controllers, or even simply gestures made and captured by one or more cameras, one can interface with, control, manipulate, interact with, and participate in the virtual reality environment or metaverse presented on the HMD. In some cases, the HMD can be wirelessly connected to a cloud computing and gaming system via a network. In one embodiment, the cloud computing and gaming system maintains and executes the video game being played by the user. In some embodiments, the cloud computing and gaming system is configured to receive input from the HMD and interactive objects via the network. The cloud computing and gaming system is configured to process the input to affect the game state of the executing video game. Output from the executing video game, such as video data, audio data, and haptic feedback data, is transmitted to the HMD and interactive objects. In other embodiments, the HMD can communicate wirelessly with the cloud computing and gaming system via alternative mechanisms or channels, such as a cellular network.
[0093] In addition, although the embodiments of the present disclosure may be described with reference to a head-mounted display, it should be understood that in other embodiments, non-head-mounted displays may be substituted, including but not limited to portable device screens (e.g., tablets, smartphones, laptops, etc.) or any other type of display that can be configured to render video and / or provide a display of an interactive scene or virtual environment in accordance with the present embodiment. It should be understood that the various features disclosed herein can be used to combine or assemble the various embodiments defined herein into specific embodiments. Therefore, the examples provided are merely some possible examples and are not limited to the various embodiments that are possible by combining various elements to define more embodiments. In some examples, some embodiments may include fewer elements without departing from the spirit of the disclosed or equivalent embodiments.
[0094] Embodiments of the present disclosure may be practiced with various computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. Embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked over a wired or wireless network.
[0095] Although the method operations are described in a particular order, it should be understood that other housekeeping operations may be performed between the operations, or the operations may be adjusted so that they occur at slightly different times or may be distributed in a system that allows processing operations to occur at various intervals associated with the processing, so long as the processing of telemetry and game state data used to generate the modified game state is performed in the desired manner.
[0096] In view of the above embodiments, it should be understood that embodiments of the present disclosure may employ various computer-implemented operations involving data stored in a computer system. These operations are operations requiring physical manipulation of physical quantities. Any operation described herein that forms part of an embodiment of the present disclosure is a useful machine operation. Embodiments of the present disclosure also relate to devices or apparatus for performing these operations. The apparatus may be specially constructed for the desired purpose, or the apparatus may be a general-purpose computer selectively activated or configured by a computer program stored in the computer. In particular, various general-purpose machines may be used with computer programs written according to the teachings herein, or more specialized apparatuses may be more conveniently constructed to perform the desired operations.
[0097] One or more embodiments may also be implemented as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that can store data that can thereafter be read by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tape, and other optical and non-optical data storage devices. Computer-readable media may include computer-readable tangible media distributed across a network-coupled computer system, allowing the computer-readable code to be stored and executed in a distributed fashion.
[0098] In one embodiment, a video game is executed locally on a game console, a personal computer, or on a server. In some cases, the video game is executed by one or more servers in a data center. When executing a video game, some instances of the video game may be simulations of the video game. For example, the video game may be executed by an environment or server that generates a simulation of the video game. In some embodiments, the simulation is an instance of the video game. In other embodiments, the simulation may be generated by an emulator. In either case, if the video game is represented as a simulation, the simulation can be executed to render interactive content that can be interactively streamed, executed, and / or controlled by user input.
[0099] Although the foregoing embodiments have been described in some detail for purposes of clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims. The present embodiments are therefore to be considered as illustrative and not restrictive, and the embodiments are not to be limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.
Claims
1. A method comprising: executing a video game on a server to generate a plurality of draw calls for execution by one or more graphics processing units (GPUs) to render image frames; identifying a first command set of one or more draw calls for rendering text in the image frame; a second command set identifying one or more draw calls for rendering a scene in the image frame; executing the first set of commands to render the text; executing the second set of commands to render the scene independently of rendering the text; generating an encoded scene by encoding the rendered scene; as well as The text and the encoded scene are streamed separately to a client device.
2. The method according to claim 1, in, The client device reconstructs the image frame by decoding the encoded scene to generate the rendered scene and overlaying the text onto the rendered scene.
3. The method according to claim 2, further comprising: including timing information in the rendered scene; as well as The timing information is included with the text.
4. The method according to claim 1, further comprising: Identifying when a network connection between the server and the client device falls below a quality of service (QoS) threshold.
5. The method according to claim 1, further comprising: Stream the original said text.
6. The method according to claim 1, wherein The text is one of the following: Player status information; and Closed captioning; and Player communications; and Game-generated text; and Chat communication.
7. The method according to claim 1, further comprising: Rendering a user interface associated with the recognized text, wherein the first command set is used to render the user interface and the text in combination, Wherein, the user interface and the text are streamed in combination to the client device.
8. The method according to claim 7, wherein: The user interface includes one of the following: Character status screen; and Thought bubbles; and Chat forum.
9. The method according to claim 1, further comprising: An artificial intelligence (AI) model is applied to identify the first set of commands for rendering the text.
10. The method according to claim 1, further comprising: storing the first command set in a text command buffer for execution by a first GPU to render the text; as well as The second command set is stored in a graphics command buffer for use by a second GPU in rendering the scene.
11. A non-transitory computer-readable medium storing a computer program for executing a method, the computer-readable medium comprising: program instructions for executing a video game on a server to generate a plurality of draw calls for execution by one or more graphics processing units (GPUs) to render image frames; program instructions for identifying a first command set of one or more draw calls for rendering text in the image frame; program instructions for identifying a second command set of one or more draw calls for rendering a scene in the image frame; program instructions for executing the first command set to render the text; program instructions for executing the second set of commands to render the scene independently of rendering the text; program instructions for generating an encoded scene by encoding the rendered scene; as well as program instructions for separately streaming said text and said encoded scene to a client device, The client device reconstructs the image frame by decoding the encoded scene to generate the rendered scene and superimposing the text on the rendered scene.
12. The non-transitory computer-readable medium of claim 11 , further comprising: Program instructions for identifying when a network connection between the server and the client device falls below a quality of service (QoS) threshold.
13. The non-transitory computer-readable medium of claim 11, wherein: In the method, the text is one of the following: Player status information; and Closed captioning; and Player communications; and Game-generated text; and Chat communication.
14. The non-transitory computer-readable medium of claim 11 , further comprising: program instructions for rendering a user interface associated with the recognized text, wherein the first command set is used to render the user interface and the text in combination, Wherein, the user interface and the text are streamed in combination to the client device.
15. The non-transitory computer-readable medium of claim 11 , further comprising: program instructions for storing the first set of commands in a text command buffer for execution by a first GPU to render the text; as well as Program instructions are provided for storing the second set of commands in a graphics command buffer for use by a second GPU in rendering the scene.
16. A computer system comprising: processor; a memory coupled to the processor and having stored therein instructions that, if executed by the computer system, cause the computer system to perform a method comprising: executing a video game on a server to generate a plurality of draw calls for execution by one or more graphics processing units (GPUs) to render image frames; identifying a first command set of one or more draw calls for rendering text in the image frame; a second command set identifying one or more draw calls for rendering a scene in the image frame; executing the first set of commands to render the text; executing the second set of commands to render the scene independently of rendering the text; generating an encoded scene by encoding the rendered scene; and separately streaming the text and the encoded scene to a client device, The client device reconstructs the image frame by decoding the encoded scene to generate the rendered scene and superimposing the text on the rendered scene.
17. The computer system of claim 16, wherein the method further comprises: Identifying when a network connection between the server and the client device falls below a quality of service (QoS) threshold.
18. The computer system according to claim 16, wherein: In the method, the text is one of the following: Player status information; and Closed captioning; and Player communications; and Game-generated text; and Chat communication.
19. The computer system of claim 16, wherein the method further comprises: Rendering a user interface associated with the recognized text, wherein the first command set is used to render the user interface and the text in combination, Wherein, the user interface and the text are streamed in combination to the client device.
20. The computer system of claim 16, wherein the method further comprises: storing the first command set in a text command buffer for execution by a first GPU to render the text; as well as The second command set is stored in a graphics command buffer for use by a second GPU in rendering the scene.
Citation Information
Cited By
AI streaming output rich text real-time rendering method based on minimum rendering unit recognition
CN120763417A