High speed scan-out of server display buffers for cloud gaming applications
By performing high-speed or early scan output operations at the cloud gaming server, video frames are scanned to the encoder line by line and encoding begins at the flip time, solving the waiting time problem between the server and the client in cloud gaming and improving the smoothness of client display and user experience.
Patent Information
- Application Number
- CN202080081921.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-31
- Filing Date
- 2020-09-29
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2040-09-29
AI Technical Summary
In cloud gaming, the round-trip and one-way waiting times between the server and the client are relatively long, which affects the user experience.
Perform high-speed or early scan output operations at the cloud gaming server, scan video frames one scan line at a time to the encoder, and start the encoding process at the flip time of the video frame to reduce one-way waiting time.
By scanning and outputting early or at high speed, the one-way waiting time in cloud gaming applications is reduced, improving the smoothness of video display on the client and the user experience.
Smart Images

Figure CN114828972B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to streaming systems configured for streaming content across a network, and more particularly, to performing a high-speed scan-out operation at a cloud gaming server and / or performing an early scan-out operation at the server to reduce latency between the cloud gaming server and a client, where the smoothness of the client displaying video can be improved by transmitting ideal display times to the client. BACKGROUND
[0002] In recent years, online services have been driving the push for allowing online or cloud gaming in a streaming format between a cloud gaming server and a client connected through a network. The streaming format is increasingly popular due to the on-demand availability of game titles, the ability for players to network for multiplayer gaming, the sharing of assets between players, the sharing of instant experiences between players and / or spectators, allowing friends to watch friends play video games, letting friends join friends in games in progress, and so on. Unfortunately, this demand is also challenging the limits of network connectivity capabilities and the processing performed at the server and the client, the responsiveness of which should be sufficient to render high-quality images delivered to the client. For example, the results of all game activities performed on the server need to be compressed and transmitted back to the client with low millisecond latency for the best user experience. Round-trip latency can be defined as the total time between a user’s controller input and the display of a video frame at the client; it can include the processing and transmission of control information from the controller to the client, the processing and transmission of control information from the client to the server, the use of this input at the server to generate a video frame in response to the input, the processing of the video frame and passing the video frame to an encoding unit (e.g., scan-out), the encoding of the video frame, the transmission of the encoded video frame back to the client, the reception and decoding of the video frame, and any processing or staging of the video frame before its display. One-way latency can be defined as the portion of the round-trip latency consisting of the time from the start of passing the video frame to the encoding unit (e.g., scan-out) at the server to the start of displaying the video frame at the client. The portion of the round-trip latency and one-way latency is associated with the time it takes for data streams to be sent from the client to the server and from the server to the client through the communication network. The other portion is associated with the processing at the client and the server; improvements in these operations, such as advanced strategies related to frame decoding and display, can significantly reduce the round-trip latency and one-way latency between the server and the client and provide a higher quality experience for users of the cloud gaming service.
[0003] It is in this context that embodiments of the present disclosure arise. SUMMARY
[0004] Embodiments of the present disclosure relate to streaming systems configured for streaming content (e.g., games) across a network, and more particularly, to performing a high-speed scan-out operation or performing scan-out earlier, such as before the occurrence of the next system VSYNC signal or at the flip time of the corresponding video frame, for delivering a modified video frame to an encoder.
[0005] Embodiments of the present disclosure disclose a method for cloud gaming. The method includes generating a video frame when a video game is executed at a server. The method includes performing a scan-out process by scanning the video frame and one or more user interface features into one or more input frame buffers on a scan-line-by-scan-line basis and compositing and blending the video frame and the one or more user interface features into a modified video frame. The method includes, during the scan-out process, scanning the modified video frame to an encoder at the server on a scan-line-by-scan-line basis. The method includes, during the scan-out process, starting to scan the video frame and the one or more user interface features into the one or more input frame buffers at a corresponding flip time of the video frame.
[0006] In another embodiment, a non-transitory computer-readable medium storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating a video frame when a video game is executed at a server. The computer-readable medium includes program instructions for performing a scan-out process by scanning the video frame and one or more user interface features into one or more input frame buffers on a scan-line-by-scan-line basis and compositing and blending the video frame and the one or more user interface features into a modified video frame. The computer-readable medium includes program instructions for, during the scan-out process, scanning the modified video frame to an encoder at the server on a scan-line-by-scan-line basis. The computer-readable medium includes program instructions for, during the scan-out process, starting to scan the video frame and the one or more user interface features into the one or more input frame buffers at a corresponding flip time of the video frame.
[0007] In another embodiment, a computer system includes a processor and a memory coupled to the processor and having stored therein instructions that, if executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating a video frame when a video game is executed at a server. The method includes performing a scan-out process by scan line by scan line scanning the video frame and one or more user interface features to one or more input frame buffers and compositing and blending the video frame and the one or more user interface features into a modified video frame. The method includes, during the scan-out process, scan line by scan line scanning the modified video frame to an encoder at the server. The method includes, during the scan-out process, starting to scan the video frame and the one or more user interface features to the one or more input frame buffers at a corresponding flip time of the video frame.
[0008] In another embodiment, a method for cloud gaming is disclosed. The method includes generating a video frame when a video game is executed at a server. The method includes performing a scan-out process to deliver the video frame to an encoder configured to compress the video frame, wherein the scan-out process starts at a flip time of the video frame. The method includes transmitting the compressed video frame to a client. The method includes determining a target display time of the video frame at the client. The method includes scheduling a display time of the video frame at the client based on the target display time.
[0009] In another embodiment, a non-transitory computer-readable medium storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating a video frame when a video game is executed at a server. The computer-readable medium includes program instructions for performing a scan-out process to deliver the video frame to an encoder configured to compress the video frame, wherein the scan-out process starts at a flip time of the video frame. The computer-readable medium includes program instructions for transmitting the compressed video frame to a client. The computer-readable medium includes program instructions for determining a target display time of the video frame at the client. The computer-readable medium includes program instructions for scheduling a display time of the video frame at the client based on the target display time.
[0010] In another implementation, a computer system includes a processor and a memory coupled to the processor and having stored therein instructions that, if executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating a video frame when a video game is executed at a server. The method includes performing a scan-out process to deliver the video frame to an encoder configured to compress the video frame, where the scan-out process begins at a flip time of the video frame. The method includes transmitting the compressed video frame to a client. The method includes determining a target display time of the video frame at the client. The method includes scheduling a display time of the video frame at the client based on the target display time.
[0011] In another implementation, a method for cloud gaming is disclosed. The method includes generating a video frame when a video game is executed at a server. The method includes performing a scan-out process to deliver the video frame to an encoder configured to compress the video frame, where the scan-out process includes scanning the video frame and one or more user interface features into one or more input frame buffers line-by-line, and compositing and blending the video frame and the one or more user interface features into a modified video frame, where the scan-out process begins at a flip time of the video frame. The method includes transmitting the compressed modified video frame to a client. The method includes determining a target display time of the modified video frame at the client. The method includes scheduling a display time of the modified video frame at the client based on the target display time.
[0012] In another implementation, a non-transitory computer-readable medium storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating a video frame when a video game is executed at a server. The computer-readable medium includes program instructions for performing a scan-out process to deliver the video frame to an encoder configured to compress the video frame, where the scan-out process includes scanning the video frame and one or more user interface features into one or more input frame buffers line-by-line, and compositing and blending the video frame and the one or more user interface features into a modified video frame, where the scan-out process begins at a flip time of the video frame. The computer-readable medium includes program instructions for transmitting the compressed modified video frame to a client. The computer-readable medium includes program instructions for determining a target display time of the modified video frame at the client. The computer-readable medium includes program instructions for scheduling a display time of the modified video frame at the client based on the target display time.
[0013] In another embodiment, a computer system includes a processor and a memory coupled to the processor and having instructions stored therein that, if executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating a video frame when a video game is executed at a server. The method includes performing a scan-out process to deliver the video frame to an encoder configured to compress the video frame, wherein the scan-out process includes scanning the video frame and one or more user interface features into one or more input frame buffers line-by-line, and compositing and blending the video frame and the one or more user interface features into a modified video frame, wherein the scan-out process begins at a flip time of the video frame. The method includes transmitting the modified video frame that is compressed to a client. The method includes determining a target display time of the modified video frame at the client. The method includes scheduling a display time of the modified video frame at the client based on the target display time.
[0014] In another embodiment, a method for cloud gaming is disclosed. The method includes generating a video frame when a video game is executed at a server, wherein the video frame is stored in a frame buffer. The method includes determining a maximum pixel clock of a chipset including a scan-out block. The method includes determining a frame rate setting based on the maximum pixel clock and an image size of a target display of a client. The method includes determining a speed setting value of the chipset. The method includes scanning the video frame from the frame buffer into the scan-out block. The method includes scan-out the video frame from the scan-out block to the encoder at the speed setting value.
[0015] In another embodiment, a non-transitory computer-readable medium storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions to generate a video frame when a video game is executed at a server, wherein the video frame is stored in a frame buffer. The computer-readable medium includes program instructions to determine a maximum pixel clock of a chipset including a scan-out block. The computer-readable medium includes program instructions to determine a frame rate setting based on the maximum pixel clock and an image size of a target display of a client. The computer-readable medium includes program instructions to determine a speed setting value of the chipset. The computer-readable medium includes program instructions to scan the video frame from the frame buffer into the scan-out block. The computer-readable medium includes program instructions to scan-out the video frame from the scan-out block to the encoder at the speed setting value.
[0016] In another embodiment, a computer system includes a processor and a memory coupled to the processor and having instructions stored therein that, if executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating video frames when executing a video game at a server, wherein the video frames are stored in a frame buffer. The method includes determining a maximum pixel clock of a chip set including a scan-out block. The method includes determining a frame rate setting based on the maximum pixel clock and an image size of a target display of a client. The method includes determining a speed setting value of the chip set. The method includes scanning the video frames from the frame buffer into the scan-out block. The method includes scanning out the video frames from the scan-out block to the encoder at the speed setting value.
[0017] Other aspects of the disclosure will become apparent from consideration of the following detailed description taken in conjunction with the accompanying drawings, which are by way of illustration. BRIEF DESCRIPTION OF DRAWINGS
[0018] The disclosure can best be understood by referring to the following description in conjunction with the accompanying drawings, in which:
[0019] FIG. 1A is a graphical representation of a VSYNC signal at the beginning of a frame period according to one embodiment of the disclosure.
[0020] FIG. 1B is a frequency plot of a VSYNC signal according to one embodiment of the disclosure.
[0021] FIG. 2A is a graphical representation of a system for providing a game between one or more cloud gaming servers and one or more client devices in various configurations over a network, wherein VSYNC signals can be synchronized and offset to reduce one-way latency according to one embodiment of the disclosure.
[0022] FIG. 2B is a graphical representation of a system for providing a game between two or more peer devices, wherein VSYNC signals can be synchronized and offset to achieve optimal timing of receiving controllers and other information between the devices according to one embodiment of the disclosure.
[0023] FIG. 2C various network configurations that benefit from proper synchronization and offset of VSYNC signals between a source device and a target device are shown according to one embodiment of the disclosure.
[0024] FIG. 2DA multi-tenant configuration that benefits from correct synchronization and offset of VSYNC signals between a source device and a target device is shown between a cloud game server and multiple clients according to one embodiment of the disclosure.
[0025] FIG. 3 Variations in one-way latency between a cloud game server and a client due to clock drift when streaming video frames generated from a video game executing on a server are shown according to one embodiment of the disclosure.
[0026] FIG. 4 A network configuration including a cloud game server and a client when streaming video frames generated from a video game executing on a server, the VSYNC signals between the server and the client are synchronized and offset to allow operations at the server and the client to overlap, and reduce one-way latency between the server and the client are shown.
[0027] FIG. 5A-1 An accelerated processing unit (APU) configured for performing a high-speed scan-out operation when streaming content from a video game executing at a cloud game server across a network for delivery to an encoder, or alternatively a CPU and GPU connected through a bus (e.g., PCI Express) according to one embodiment of the disclosure is shown.
[0028] FIG. 5A-2 A chipset 540B configured for performing a high-speed scan-out operation when streaming content from a video game executing at a cloud game server across a network for delivery to an encoder, where user interface features are integrally formed into the video frames of the game rendering according to one embodiment of the disclosure is shown.
[0029] FIG. 5B-1 , FIG. 5B-2 and FIG. 5B-3 A scan-out operation performed when streaming content from a video game executing at a cloud game server across a network to a client to generate modified video frames for delivery to an encoder according to one embodiment of the disclosure is shown.
[0030] FIG. 5C-FIG. 5D An exemplary server configuration having one or more input frame buffers used when performing a high-speed scan-out operation when streaming content from a video game executing at a cloud game server across a network for delivery to an encoder according to embodiments of the disclosure is shown.
[0031] FIG. 6is a flowchart showing a method for cloud gaming according to one embodiment of the present disclosure, where an early scan-out process is performed to initiate the encoding process earlier to reduce one-way latency between the server and the client.
[0032] FIG. 7A is a process for generating and transmitting video frames at a cloud gaming server according to one embodiment of the present disclosure, where the process is optimized to perform high speed and / or early scan-out to the encoder to reduce one-way latency between the cloud gaming server and the client.
[0033] FIG. 7B is timing when performing a scan-out process at a cloud gaming server according to one embodiment of the present disclosure, where the scan-out is performed at high speed and / or early to enable video frames to be scanned to the encoder earlier to reduce one-way latency between the cloud gaming server and the client.
[0034] FIG. 7C is a time period to perform scan-out at high speed according to one embodiment of the present disclosure to enable video frames to be scanned to the encoder earlier to reduce one-way latency between the cloud gaming server and the client.
[0035] FIG. 8A is a flowchart showing a method for cloud gaming according to one embodiment of the present disclosure, where the video displayed by the client can be smoothed in the cloud gaming application, where high speed and / or early scan-out operations can be performed at the server to reduce one-way latency between the cloud gaming server and the client.
[0036] FIG. 8B is a timing diagram of server and client operations performed during execution of a video game at a server to generate video frames of a game render that are then sent to a client for display according to one embodiment of the present disclosure.
[0037] FIG. 9 is a component of an example device that can be used to perform aspects of various embodiments of the present disclosure. DETAILED DESCRIPTION
[0038] While the following detailed description contains many specifics, those skilled in the art will appreciate that many variations and alterations in the details are within the scope of the present disclosure. Thus, the following description is presented by way of example only and should not be taken as limiting the present description.
[0039] In general, various embodiments of the present disclosure describe methods and systems configured to reduce latency and / or latency instability between a source device and a target device when streaming media content (e.g., streaming audio and video from a video game). Latency instability can be introduced in one-way latency between a server and a client due to additional time required to generate complex frames (e.g., scene changes) at the server, increased time to encode / compress complex frames at the server, variable communication paths on the network, and increased time to decode complex frames at the client. Latency instability can also be introduced due to clock differences at the server and client, which can cause drift between the server VSYNC signal and the client VSYNC signal. In embodiments of the present disclosure, one-way latency between a server and a client can be reduced in a cloud gaming application by performing a high-speed scan-out of the cloud gaming display buffer. In another embodiment, one-way latency can be reduced by performing an early scan-out of the cloud gaming display buffer. In another embodiment, when addressing latency issues, the smoothness of the client displaying video in a cloud gaming application can be improved by transmitting an ideal display time to the client.
[0040] In particular, in some embodiments of the disclosure, one-way latency in a cloud gaming application can be reduced by starting the encoding process earlier. For example, in certain architectures for streaming media content (e.g., streaming audio and video from a video game) from a cloud gaming server to a client, the scan-out of the server display buffer includes performing additional operations on a video frame to generate one or more layers that will then be combined and scanned to a unit that performs video encoding. By performing scan-out at a high rate (120 Hz or even higher), it can be possible to start the encoding process earlier, and thus reduce one-way latency. Also, in some embodiments of the disclosure, one-way latency in a cloud gaming application can be reduced by performing an early scan-out process at the cloud gaming server. In particular, in certain architectures for streaming media content (e.g., streaming audio and video from a video game) from a cloud gaming server to a client, an application program (e.g., a video game) running on the server requests a “flip” of the server display buffer to occur at the completion of rendering a video frame. Rather than performing the scan-out operation at the subsequent occurrence of the server VSYNC signal, the scan-out operation begins at the flip time, where the scan-out of the server display buffer includes performing additional operations on a video frame to generate one or more layers that are then combined and scanned to a unit that performs video encoding. By scan-out at the flip time (rather than the next VSYNC), it can be possible to start the encoding process earlier, and thus reduce one-way latency. Because there is no display attached to the cloud gaming server in practice, display timing is not affected. In some embodiments of the disclosure, when the server scan-out of the display buffer is performed at the flip time (rather than the subsequent VSYNC), the ideal display timing at the client depends on the time at which scan-out occurred and the intent of the game with respect to that particular display buffer (e.g., whether it was for the next VSYNC, or whether the game ran late and it was actually for the previous VSYNC). The strategy differs depending on whether the game is fixed frame rate or variable frame rate, and whether the information will be implicit (inferred from the scan-out timing) or explicit (the game provides the ideal timing via GPU API, possibly VSYNC or fractional time).
[0041] With the above general understanding of the various embodiments, example details of the embodiments will now be described with reference to the various drawings.
[0042] Throughout this specification, reference to “a game” or “a video game” or “a gaming application” is intended to mean any type of interactive application that is directed by the execution of input commands. For illustrative purposes only, interactive applications include applications for gaming, word processing, video processing, video game processing, etc. Moreover, the above-introduced terms are interchangeable.
[0043] Cloud gaming includes executing a video game at a server to generate video frames of game renders, which are then transmitted to a client for display. Operation timings at both the server and the client can be associated with respective vertical sync (VSYNC) parameters. When the VSYNC signals are properly synchronized and / or offset between the server and / or the client, the operations performed at the server (e.g., generating and transmitting video frames within one or more frame periods) are synchronized with the operations performed at the client (e.g., displaying video frames on a display at a display frame or refresh rate that corresponds to the frame period). In particular, a server VSYNC signal generated at the server and a client VSYNC signal generated at the client can be used to synchronize the operations at the server and the client. That is, when the server and client VSYNC signals are synchronized and / or offset, the server generating and transmitting video frames are synchronized with how the client displays these video frames.
[0044] VSYNC signaling and vertical blanking intervals (VBIs) have been incorporated for generating video frames and displaying these video frames when streaming media content between a server and a client. For example, the server strives to generate a video frame of a game render within one or several frame periods defined by a corresponding server VSYNC signal (e.g., generating one video frame per frame period results in 60 Hz operation and generating one video frame per two frame periods results in 30 Hz operation if the frame period is 16.7 ms), and then encodes and transmits this video frame to the client. At the client, the received encoded video frames are decoded and displayed, where the client displays each video frame that is rendered for display starting with a corresponding client VSYNC.
[0045] To illustrate, FIG. 1A A VSYNC signal 111 is shown as to how it can indicate the start of a frame period, where various operations can be performed at the server and / or the client during the corresponding frame period. When streaming media content, the server can use a server VSYNC signal to generate and encode video frames, and the client can use a client VSYNC signal to display video frames. The VSYNC signal 111 is generated at a defined frequency that corresponds to the defined frame period 110, as shown in FIG. 1B Additionally, a VBI 105 defines a period of time between when the last raster line of a previous frame period is drawn on a display and when the first raster line (e.g., top) is drawn to the display. As shown, after the VBI 105, the video frame that is rendered for display is displayed via raster scan lines 106 (e.g., left to right, raster line by raster line).
[0046] Additionally, various embodiments of the present disclosure are disclosed for reducing one-way latency and / or latency instability between a source device and a target device, such as when streaming media content (e.g., video game content). For purposes of illustration only, various embodiments for reducing one-way latency and / or latency instability are described within a server and client network configuration. However, it should be understood that the various techniques disclosed for reducing one-way latency and / or latency instability can be implemented within other network configurations and / or on a peer-to-peer network, as shown in FIG. 2A to FIG. 2D For example, various embodiments disclosed for reducing one-way latency and / or latency instability can be implemented between one or more of a server and client device in various configurations (e.g., server and client, server and server, server and multiple clients, server and multiple servers, client and client, client and multiple clients, etc.).
[0047] FIG. 2A is a diagram of a system 200A for providing games in various configurations between one or more cloud gaming networks 290 and / or servers 260 and one or more client devices 210 over a network 250 in accordance with an embodiment of the present disclosure, where server and client VSYNC signals can be synchronized and offset, and / or where dynamic buffering is performed on the client, and / or where encoding and transmission operations on the server can overlap, and / or where receiving and decoding operations at the client can overlap, and / or where decoding and display operations on the client can overlap to reduce one-way latency between the servers 260 and the clients 210. In particular, system 200A provides games via cloud gaming networks 290 in which games are executed on client devices 210 (e.g., thin clients) that are remote from the corresponding users that are playing the games. System 200A can provide game control to one or more users that are playing one or more games in single player or multi-player mode over cloud gaming networks 290 via network 250. In some embodiments, cloud gaming networks 290 can include a plurality of virtual machines (VMs) running on a hypervisor of a host, where one or more virtual machines are configured to execute game processor modules with hardware resources available to the hypervisor of the host. Network 250 can include one or more communication technologies. In some embodiments, network 250 can include a 5thGeneration (5G) network technology with an advanced wireless communication system.
[0048] In some embodiments, wireless technology can be used to facilitate communication. Such technology can include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. A 5G network is a digital cellular network in which a provider's service area is divided into small geographic areas called cells. Analog signals representing sounds and images are digitized, converted by an analog-to-digital converter, and transmitted as bit streams in the phone. All 5G wireless devices in a cell communicate with a local array of antennas and low-power automatic transceivers (transmitters and receivers) in the cell over radio waves on a frequency channel assigned to the device from a pool of frequencies that is reused in other cells. The local antennas are connected with the phone network and the Internet through high-bandwidth fiber-optic or wireless backhaul. As in other cell networks, a mobile device crossing from one cell to another is automatically transferred to the new cell. It should be understood that 5G networks are just one example type of communication network, and embodiments of the present disclosure can utilize wireless or wired communication of a previous generation, as well as a next generation of wired or wireless technology after 5G.
[0049] As shown, the cloud gaming network 290 includes game servers 260 that provide access to a plurality of video games. The game servers 260 can be any type of server computing device available in the cloud and can be configured as one or more virtual machines executing on one or more hosts. For example, the game servers 260 can manage virtual machines that support game processors that instantiate game instances for users. As such, a plurality of game processors of the game servers 260 associated with a plurality of virtual machines are configured to execute a plurality of instances of one or more games associated with game play of a plurality of users. In this manner, the backend servers support streaming of media (e.g., video, audio, etc.) of game play of a plurality of game applications to a plurality of corresponding users. That is, the game servers 260 are configured to stream data (e.g., rendered images and / or frames of corresponding game play) back to corresponding client devices 210 over the network 250. In this manner, computationally complex game applications can be executed at the backend servers in response to controller inputs received and forwarded by the client devices 210. Each server is capable of rendering images and / or frames, which are then encoded (e.g., compressed) and streamed to corresponding client devices for display
[0050] For example, multiple users can access a cloud gaming network 290 via a communication network 250 using corresponding client devices 210 configured to receive streaming media. In one implementation, a client device 210 can be configured as a thin client that provides an interface with a back-end server (e.g., a game server 260 of a cloud gaming network 290) that is configured to provide computing functionality (e.g., including a game title processing engine 211). In another implementation, a client device 210 can be configured with a game title processing engine and game logic for at least some local processing of a video game, and can also be utilized to receive streaming content generated by a video game executing at a back-end server, or other content provided in support by the back-end server. For local processing, a game title processing engine includes basic processor-based functionality for executing a video game and services associated with the video game. Game logic is stored on a local client device 210 and is used to execute a video game.
[0051] In particular, a client device 210 of a corresponding user (not shown) is configured to request access to a game over a communication network 250, such as the Internet, and to render display images generated by a video game executing at a game server 260, with encoded images being delivered to the client device 210 for display in association with the corresponding user. For example, a user can interact with an instance of a video game executing on a game processor of a game server 260 through a client device 210. More particularly, the instance of the video game is executed by a game title processing engine 211. Corresponding game logic (e.g., executable code) 215 implementing the video game is stored and accessible through a data store (not shown), and is used to execute the video game. The game title processing engine 211 is capable of supporting multiple video games using multiple game logic, each of which can be selected by a user.
[0052] For example, the client device 210 is configured to interact with a game title processing engine 211 associated with game play of a corresponding user, such as through input commands for driving game play. In particular, the client device 210 can receive input from various types of input devices, such as game controllers, tablets, keyboards, gestures captured by a video camera, mice, touchpads, etc. The client device 210 can be any type of computing device having at least a memory and a processor module, capable of connecting to a game server 260 through a network 250. The backend game title processing engine 211 is configured for generating rendered images, which are delivered through the network 250 for display on a corresponding display associated with the client device 210. For example, through a cloud-based service, the game rendered images can be delivered by an instance of the corresponding game executing on a game execution engine 211 of the game server 260. That is, the client device 210 is configured for receiving encoded images (e.g., encoded by game rendered images generated by executing a video game) and for displaying as rendered images on a display 11. In one embodiment, the display 11 includes an HMD (e.g., displaying VR content). In some embodiments, the rendered images can be streamed wirelessly or wired directly from the cloud-based service or via the client device 210 (e.g., Remote Play) to a smartphone or tablet. Remote Play) to a smartphone or tablet.
[0053] In one embodiment, the game server 260 and / or game title processing engine 211 includes processor-based functions for executing games and services associated with the game application. For example, the processor-based functions include 2D or 3D rendering, physics, physics simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadowing, culling, transformation, artificial intelligence, etc. In addition, the services for the game application include memory management, multi-threading management, quality of service (QoS), bandwidth testing, social networking, social friend management, social network communication with friends, communication channels, text messaging, instant messaging, chat support, etc.
[0054] In one embodiment, the cloud game network 290 is a distributed game server system and / or architecture. In particular, a distributed game engine that executes game logic is configured as a corresponding instance of a corresponding game. Generally, the distributed game engine takes each of the functions of a game engine and distributes those functions for execution by multiple processing entities. The various functions can also be distributed across one or more processing entities. The processing entities can be configured in different configurations, including physical hardware, and / or configured as virtual components or virtual machines, and / or configured as virtual containers, where a container is different from a virtual machine in that it virtualizes an instance of a game application running on a virtualized operating system. The processing entities can utilize and / or rely on servers on one or more servers (compute nodes) of the cloud game network 290, where the servers can be located on one or more racks. Coordination, assignment, and management of these functions to the various processing entities is performed by a distribution synchronization layer. In this way, execution of these functions is controlled by the distribution synchronization layer to support generation of media (e.g., video frames, audio, etc.) for the game application in response to player controller inputs. The distribution synchronization layer is able to efficiently execute (e.g., through load balancing) these functions across the distributed processing entities such that critical game engine components / functions are distributed and reassembled for more efficient processing.
[0055] The game title processing engine 211 includes a central processing unit (CPU) and a group of graphics processing units (GPUs) that can be configured to perform multi-tenant GPU functionality. In another embodiment, multiple GPU devices are combined to perform graphics processing for a single application executing on a corresponding CPU.
[0056] FIG. 2B is a diagram for providing a game between two or more peer devices according to one embodiment of the disclosure, where VSYNC signals can be synchronized and offset to achieve optimal timing of receiving controller and other information between devices. For example, a head-to-head game can be performed using two or more peer devices connected through a network 250 or directly through peer-to-peer communication (e.g., Bluetooth, local area network, etc.).
[0057] As shown, a game is executed locally on each of the corresponding users' client devices 210 (e.g., game consoles) that are playing a video game, where the client devices 210 communicate through a peer-to-peer network. For example, an instance of the video game is executed by the game title processing engine 211 of the corresponding client device 210. Game logic 215 (e.g., executable code) that implements the video game is stored on the corresponding client device 210 and used to execute the game. For purposes of illustration, the game logic 215 can be delivered to the corresponding client device 210 through a portable medium (e.g., optical medium) or through a network (e.g., downloaded from a game provider over the Internet).
[0058] In one embodiment, the game title processing engine 211 of a corresponding client device 210 includes processor-based functionality for executing a game and services associated with the game application. For example, the processor-based functionality includes 2D or 3D rendering, physics, physics simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadowing, culling, transformation, artificial intelligence, etc. In addition, the services for the game application include memory management, multi-threading management, quality of service (QoS), bandwidth testing, social networking, social friend management, social network communication with friends, communication channels, text messaging, instant messaging, chat support, etc.
[0059] The client device 210 can receive input from various types of input devices, such as game controllers, tablets, keyboards, gestures captured by a video camera, mice, touchpads, etc. The client device 210 can be any type of computing device having at least a memory and a processor module and configured for generating rendered images executed by the game title processing engine 211 and for displaying the rendered images on a display (e.g., display 11 or a display 11 including a head-mounted display - HMD, etc.). For example, the rendered images can be associated with an instance of a game executing locally on the client device 210 for a corresponding user’s game play, such as through input commands for driving the game play. Some examples of the client device 210 include a personal computer (PC), a game console, a home theater device, a general purpose computer, a mobile computing device, a tablet, a phone, or any other type of computing device that can execute an instance of a game.
[0060] FIG. 2C Various network configurations that benefit from correct synchronization and offset of VSYNC signals between a source device and a target device are shown in accordance with embodiments of the present disclosure, including those configurations shown in FIG. 2A to FIG. 2B Particularly, the various network configurations benefit from correct alignment of the frequency of server and client VSYNC signals, as well as timing offset of the server and client VSYNC signals, to reduce one-way latency and / or latency variability between the server and the client. For example, one network device configuration includes a cloud game server (e.g., source) to client (target) configuration. In one embodiment, the client can include a WebRTC client configured for providing audio and video communication within a web browser. Another network configuration includes a client (e.g., source) to server (target) configuration. Yet another network configuration includes a server (e.g., source) to server (e.g., target) configuration. Another network device configuration includes a client (e.g., source) to client (target) configuration, where the clients can each be a game console to provide, for example, head-to-head gaming.
[0061] In particular, alignment of the VSYNC signals can include synchronizing the frequency of the server VSYNC signal and the client VSYNC signal, and can also include adjusting the timing offset between the client VSYNC signal and the server VSYNC signal for removing drift, and / or maintaining a desired relationship between the server VSYNC signal and the client VSYNC signal to reduce one-way latency and / or latency variability. In one embodiment, to achieve proper alignment, the server VSYNC signal can be tuned to achieve proper alignment between the server 260 and the pair of clients 210. In another embodiment, the client VSYNC signal can be tuned to achieve proper alignment between the server 260 and the pair of clients 210. Once the client VSYNC signal and the server VSYNC signal are aligned, the server VSYNC signal and the client VSYNC signal occur at substantially the same frequency and are offset from each other by a timing offset that can be adjusted from time to time. In another embodiment, alignment of the VSYNC signals can include synchronizing the frequency of the VSYNC of two clients, and can also include adjusting the timing offset between their VSYNC signals to remove drift and / or achieve optimal reception timing of the controller and other information; either VSYNC signal can be tuned to achieve such alignment. In another embodiment, alignment can include synchronizing the frequency of the VSYNC of multiple servers, and can also include synchronizing the frequency of the server VSYNC signal and the client VSYNC signal, and adjusting the timing offset between the client VSYNC signal and the server VSYNC signal, e.g., for a head-to-head cloud gaming. In server-to-client and client-to-client configurations, alignment can include synchronization of the frequency between the server VSYNC signal and the client VSYNC signal, and providing a proper timing offset between the server VSYNC signal and the client VSYNC signal. In server-to-server configurations, alignment can include synchronization of the frequency between the server VSYNC signal and the client VSYNC signal without setting a timing offset.
[0062] FIG. 2D A multi-tenant configuration between a cloud gaming server 260 and one or more clients 210 that benefits from proper synchronization and offset of VSYNC signals between a source device and a target device is shown in accordance with one embodiment of the disclosure. In server-to-client configurations, alignment can include synchronization of the frequency between the server VSYNC signal and the client VSYNC signal, and providing a proper timing offset between the server VSYNC signal and the client VSYNC signal. In multi-tenant configurations, in one embodiment, the client VSYNC signal is tuned at each client 210 to achieve proper alignment between the server 260 and the pair of clients 210.
[0063] For example, in one embodiment, a graphics subsystem can be configured to perform multi-tenancy GPU functionality, where one graphics subsystem can implement graphics and / or rendering pipelines for multiple games. That is, the graphics subsystem is shared among multiple games that are being executed. In particular, in one embodiment, a game title processing engine can include a CPU and GPU bank that can be configured to perform multi-tenancy GPU functionality, where one CPU and GPU bank can implement graphics and / or rendering pipelines for multiple games. That is, the CPU and GPU bank is shared among multiple games that are being executed. The CPU and GPU bank can be configured as one or more processing devices. In another embodiment, multiple GPU devices are combined to perform graphics processing for a single application executing on a corresponding CPU.
[0064] FIG. 3A general procedure is shown in which a video game is executed at a server to generate video frames of game renderings and send these video frames to a client for display. Conventionally, several operations at the game server 260 and client 210 are performed within a frame period as defined by the respective VSYNC signals. For example, the server 260 strives to generate a video frame of game rendering at 301 in one or more frame periods as defined by the corresponding server VSYNC signal 311. The video frame is generated by the game in response to control information (e.g., input commands of a user) delivered from the input device at operation 350 or game logic that is not driven by control information. There can be transmission jitter 351 when sending control information to the server 260, where the jitter 351 measures variations in network latency from the client to the server (e.g., when sending input commands). As shown, the thick arrow shows the current delay when sending control information to the server 260, but due to jitter, there can be a range of arrival times (e.g., bounded by the dashed arrows) for the control information at the server 260. At flip time 309, the GPU hits a flip command, which indicates that the corresponding video frame has been completely generated and put into a frame buffer at the server 260. Thereafter, the server 260 performs scan-out / scan-in (operation 302), in which for this video frame, the scan-out can be aligned with the VSYNC signal 311 (VBI is omitted for clarity) within a subsequent frame period defined by the server VSYNC signal 311. Subsequently, the video frame is encoded (operation 303) (e.g., encoding starts after the occurrence of the VSYNC signal 311, and the end of encoding can not be aligned with the VSYNC signal) and transmitted (operation 304, where the transmission can not be aligned with the VSYNC signal 311) to the client 210. At the client 210, the encoded video frame is received (operation 305, where the reception can not be aligned with the client VSYNC signal 312), decoded (operation 306, where the decoding can not be aligned with the client VSYNC signal 312), buffered, and displayed (operation 307, where the start of the display can be aligned with the client VSYNC signal 312). In particular, the client 210 displays each video frame that is rendered for display, starting with the corresponding occurrence of the client VSYNC signal 312.
[0065] One-way latency 315 can be defined as the latency from the start of the delivery of a video frame to the encoding unit at the server (e.g., scan-out 302) to the start of the display 307 of the video frame at the client. That is, one-way latency is the time from server scan-out to client display, accounting for client buffering. Individual frames have a latency from scan-out 302 to decode 306 completion that can vary from frame to frame due to server operations such as encoding 303 and transmission 304, network transmission between server 260 and client 210 with jitter 352, and high variability in client reception 305. As shown, the straight, thick arrow shows the current latency when the corresponding video frame is sent to the client 210, but due to jitter 352, there can be a range of arrival times (e.g., bounded by the dashed arrows) for the video frame at the client 210. Because one-way latency must be relatively stable (e.g., remain fairly consistent) to achieve a good gaming experience, traditionally the result of buffering 320 is that the display of individual frames with low latency (e.g., from scan-out 302 to decode 306 completion) is delayed by several frame periods. That is, if there is network instability or unpredictable encoding / decoding times, additional buffering is needed to keep the one-way latency consistent.
[0066] According to one embodiment of the present disclosure, when streaming video frames generated from a video game executing on a server, the one-way latency between the cloud gaming server and the client can vary due to clock drift. That is, a difference in frequency of the server VSYNC signal 311 and the client VSYNC signal 312 can cause the client VSYNC signal to drift relative to the frames arriving from the server 260. The drift can be due to a very slight difference in the crystal oscillators used in each of the respective clocks at the server and client. Further, embodiments of the present disclosure reduce one-way latency by performing one or more of synchronization and offset of the VSYNC signals between the server and the client for alignment; providing dynamic buffering on the client; overlapping encoding and transmission of video frames at the server; overlapping reception and decoding of video frames at the client; and overlapping decoding and display of video frames at the client.
[0067] FIG. 4 A data flow through a network configuration including a highly optimized cloud gaming server 260 and a highly optimized client 210 when streaming video frames generated from a video game executing on a server is shown, according to embodiments of the present disclosure, in which overlapping server operations and client operations reduce one-way latency, and synchronization and offset of the VSYNC signals between the server and the client reduce one-way latency and reduce variability of the one-way latency between the server and the client. In particular,FIG. 4 The desired alignment between the server VSYNC signal and the client VSYNC signal is shown. In one embodiment, tuning of the server VSYNC signal 311 is performed to obtain the correct alignment between the server VSYNC signal and the client VSYNC signal, such as in a server and client network configuration. In another embodiment, tuning of the client VSYNC signal 312 is performed to obtain the correct alignment between the server VSYNC signal and the client VSYNC signal, such as in a multi-tenant server to multi-client network configuration. For purposes of illustration, FIG. 4 Tuning of the server VSYNC signal 311 to synchronize the frequency of the server VSYNC signal and the client VSYNC signal, and / or adjust the timing offset between the corresponding client VSYNC signal and the server VSYNC signal, is described in the Background of the Invention, but it is understood that the client VSYNC signal 312 can also be used for tuning. In the context of this patent, "synchronize" is understood to mean tuning the signals so that their frequencies match, but the phases can be different; "offset" is understood to mean a time delay between the signals, such as the time between when one signal reaches its maximum value and when the other signal reaches its maximum value.
[0068] As shown, in embodiments of the disclosure, FIG. 4 An improved process is shown for executing a video game at a server to generate rendered video frames and sending those video frames to a client for display. The process is shown with respect to generating and displaying a single video frame at the server and the client. In particular, the server generates a game rendered video frame at 401. For example, the server 260 includes a CPU configured to execute the game (e.g., game title processing engine 211). The CPU generates one or more draw calls for the video frame, where the draw calls include commands placed into a command buffer for a corresponding GPU of the server 260 to execute in a graphics pipeline. The graphics pipeline can include one or more shader programs that operate on vertices of objects within a scene to generate texture values as rendered for the video frame for display, where the operations are performed in parallel by the GPU for efficiency. At flip time 409, the GPU hits a flip command in the command buffer that indicates that the corresponding video frame has been completely generated and / or rendered and is placed into a frame buffer at the server 260.
[0069] At 402, the server performs a scan-out of a video frame of a game rendering to an encoder. In particular, the scan-out is performed line-by-line or in groups of consecutive lines, where a line refers to a single horizontal line, e.g., from a screen edge to a screen edge of a display. These lines or groups of consecutive lines are sometimes referred to as slices, and are referred to as screen slices in this specification. In particular, the scan-out 402 can include several processes that modify the frame of the game rendering, including covering it with another frame buffer, or shrinking it so as to surround it with information from another frame buffer. During the scan-out 402, the modified video frame is then scanned into the encoder for compression. In one embodiment, the scan-out 402 is performed at the occurrence 311a of the VSYNC signal 311. In other embodiments, the scan-out 402 can be performed before the occurrence of the VSYNC signal 311, such as at the flip time 409.
[0070] At 403, the video frame of the game rendering, which can have undergone modification, is encoded at the encoder on a per-encoder slice basis to generate one or more encoded slices, where an encoded slice is independent of a scan line or screen slice. As such, the encoder generates one or more encoded (e.g., compressed) slices. In one embodiment, the encoding process begins before the scan-out process 402 for the corresponding video frame has completed. Further, the beginning and / or end of the encoding 403 can or can not be aligned with the server VSYNC signal 311. The boundaries of the encoded slices are not limited to a single scan line, and can include a single scan line or multiple scan lines. Additionally, the end of an encoded slice and / or the beginning of the next encoder slice can not necessarily occur at an edge of the display screen (e.g., can occur somewhere in the middle of the screen or in the middle of a scan line), such that the encoded slice does not necessarily traverse the edge-to-edge of the display screen. As shown, one or more encoded slices can be compressed and / or encoded, including the compressed, hash-tagged "encoded slice A."
[0071] At 404, the encoded video frame is transmitted from the server to the client, where the transmission can occur on a per-encoded slice basis, where each encoded slice is an encoder slice that has been compressed. In one embodiment, the transmission process 404 begins before the encoding process 403 for the corresponding video frame has completed. Further, the beginning and / or end of the transmission 404 can or can not be aligned with the server VSYNC signal 311. As shown, the compressed, encoded slice A is transmitted to the client independently of other compressed, encoder slices of the rendered video frame. The encoder slices can be transmitted one at a time, or in parallel.
[0072] At 405, the client again receives compressed video frames on a per- encoded slice basis. Again, the start and / or end of the reception 405 can or can not be aligned with the client VSYNC signal 312. As shown, the client receives a compressed encoded slice A. There can be transmission jitter 452 between the server 260 and the client 210, where the jitter 452 measures the variation in network latency from the server 260 to the client 210. Lower jitter values represent a more stable connection. As shown, the straight, thick arrow shows the current latency when the corresponding video frame is sent to the client 210, but due to jitter, there can be a range of arrival times (e.g., bounded by the dashed arrows) at the client 210 regarding the video frame. The variation in latency can also be due to one or more operations at the server, such as the encoding 403 and the transmission 404, as well as network issues that introduce latency when transmitting the video frame to the client 210.
[0073] At 406, the client again decodes the compressed video frames on a per- encoded slice basis, resulting in a decoded slice A (shown without hash marks) that is now ready for display. In one embodiment, the decoding process 406 begins before the reception process 405 of the corresponding video frame is completely finished. Again, the start and / or end of the decoding 406 can or can not be aligned with the client VSYNC signal 312. At 407, the client displays the decoded rendered video frame on a display at the client. That is, for example, the decoded video frame is placed in a display buffer that is streamed out to a display device on a per-scan line basis. In one embodiment, the display process 407 (i.e., streaming out to the display device) begins after the decoding process 406 of the corresponding video frame has completely finished, i.e., the decoded video frame is completely resident in the display buffer. In another embodiment, the display process 407 begins before the decoding process 406 of the corresponding video frame has completely finished. That is, the streaming out to the display device begins from an address in the display buffer at which time only a portion of the decoded frame buffer is resident in the display buffer. The display buffer is then timely updated or populated with the remaining portion of the corresponding video frame for display, such that the updating of the display buffer is performed before these portions are streamed out to the display. Again, the start and / or end of the display 407 is aligned with the client VSYNC signal 312.
[0074] In one embodiment, the one-way latency 416 between the server 260 and the client 210 can be defined as the time elapsed between the start of the scan-out 402 and the start of the display 407. Embodiments of the present disclosure are capable of aligning the VSYNC signals between the server and the client (e.g., synchronizing the frequencies and adjusting the offset) to reduce the one-way latency between the server and the client and to reduce the variability of the one-way latency between the server and the client. For example, embodiments of the present disclosure are capable of calculating the optimal adjustment to the offset 430 between the server VSYNC signal 311 and the client VSYNC signal 312 such that the decoded rendered video frame can be available in time for the display process 407 even in the case of near worst-case time required for server processing (such as encoding 403 and transmission 404), near worst-case network latency between the server 260 and the client 210, and near worst-case client processing (such as reception 405 and decoding 406). That is, it is not necessary to determine the absolute offset between the server VSYNC and the client VSYNC; it is sufficient to adjust the offset such that the decoded rendered video frame can be available in time for the display process.
[0075] In particular, the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 can be aligned by synchronization. Synchronization is achieved by tuning the server VSYNC signal 311 or the client VSYNC signal 312. For purposes of illustration, tuning is described with respect to the server VSYNC signal 311, but it is understood that tuning can instead be performed on the client VSYNC signal 312. For example, as shown in FIG. 4B, the server VSYNC signal 311 is tuned to be faster than the client VSYNC signal 312. As a result, the server frame period 410 (e.g., the time between two occurrences 311c and 31 Id of the server VSYNC signal 311) is shorter than the client frame period 415 (e.g., the time between two occurrences 312a and 312b of the client VSYNC signal 312). FIG. 4 As shown in FIG. 4B, the server frame period 410 (e.g., the time between two occurrences 311c and 31 Id of the server VSYNC signal 311) is substantially equal to the client frame period 415 (e.g., the time between two occurrences 312a and 312b of the client VSYNC signal 312), which indicates that the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 are also substantially equal.
[0076] To maintain frequency synchronization of the server VSYNC signal and the client VSYNC signal, the timing of the server VSYNC signal 311 can be manipulated. For example, the vertical blanking interval (VBI) in the server VSYNC signal 311 can be increased or decreased over a period of time, such as to account for drift between the server VSYNC signal 311 and the client VSYNC signal 312. Manipulation of the vertical blanking (VBLANK) line in the VBI provides for adjusting the number of scan lines used for VBLANK for one or more frame periods of the server VSYNC signal 311. Decreasing the number of scan lines of VBLANK decreases the corresponding frame period (e.g., time interval) between occurrences of the server VSYNC signal 311. Conversely, increasing the number of scan lines of VBLANK increases the corresponding frame period (e.g., time interval) between occurrences of the VSYNC signal 311. In this manner, the frequency of the server VSYNC signal 311 is adjusted to align the frequency between the client VSYNC signal 312 and the server VSYNC signal 311 to be substantially the same frequency. Also, the offset between the server VSYNC signal and the client VSYNC signal can be adjusted by increasing or decreasing the VBI for a short period of time, and then returning the VBI to its original value. In one embodiment, the server VBI is adjusted. In another embodiment, the client VBI is adjusted. In another embodiment, instead of there being two devices (a server and a client), there are multiple connected devices, each of which can have a corresponding VBI that is adjusted. In one embodiment, each of the multiple connected devices can be independent peer devices (e.g., there is no server device). In another embodiment, the multiple devices can include one or more server devices and / or one or more client devices, arranged in one or more server / client architectures, multi-tenant server / client architectures, or some combination thereof.
[0077] Alternatively, in one embodiment, the server’s pixel clock (e.g., located in the server’s southbridge, or in the case of a discrete GPU, it would use its own hardware to generate the pixel clock itself) can be manipulated to perform a coarse and / or fine tuning of the frequency of the server’s VSYNC signal 311 over a period of time to bring the frequency synchronization back into alignment between the server’s VSYNC signal 311 and the client’s VSYNC signal 312. Specifically, the server’s pixel clock in the southbridge can be overclocked or underclocked to adjust the overall frequency of the server’s VSYNC signal 311. In this way, the frequency of the server’s VSYNC signal 311 is adjusted to bring the frequency alignment between the client’s VSYNC signal 312 and the server’s VSYNC signal 311 to substantially the same frequency. The offset between the server’s VSYNC and the client’s VSYNC can be adjusted by increasing or decreasing the client server pixel clock for a short period of time and then returning the pixel clock to its original value. In one embodiment, the server’s pixel clock is adjusted. In another embodiment, the client’s pixel clock is adjusted. In another embodiment, instead of there being two devices (a server and a client), there are multiple connected devices, each of which can have a corresponding pixel clock that is adjusted. In one embodiment, each of the multiple connected devices can be independent peer devices (e.g., there is no server device). In another embodiment, the multiple connected devices can include one or more server devices and one or more client devices arranged in one or more server / client architectures, multi-tenant server / client architectures, or some combination thereof.
[0078] FIG. 5A-1 A chip set 540 configured for performing a high speed scan-out operation for delivery to an encoder when streaming content from a video game executing at a cloud gaming server is shown according to one embodiment of the present disclosure. In addition, the chip set 540 can be configured to perform the scan-out operation earlier, such as before the occurrence of the next system VSYNC signal or at the flip time of the corresponding video frame. In particular, FIG. 5A-1 How the speed of the scan-out block 550 is determined for the target display of the client is shown in one embodiment.
[0079] Chipset 540 is configured to operate at a maximum pixel clock 515. The pixel clock defines the rate at which the chipset can process pixels, such as by scan out block 550. The rate of the pixel clock is typically expressed in a value of megahertz, representing the number of pixels that can be processed. In particular, pixel clock calculator 510 is configured to determine the maximum pixel clock 515 based on chip compute settings 501 and / or self-diagnostic test 505. For example, chipset 540 can be designed with a particular maximum pixel clock, which is included in chip compute settings 501. However, once built, chipset 540 can be able to operate at a higher pixel clock, or can not actually operate at the designed pixel clock determined from chip compute settings 501. As such, test 505 can be performed to determine a self-diagnostic pixel clock 505. Pixel clock calculator 510 can be configured to define the maximum pixel clock 515 of chipset 540 based on the higher of the designed pixel clock determined from either chip compute settings 501 or self-diagnostic pixel clock 505. For illustrative purposes, an exemplary maximum pixel clock can be 300 megapixels per second (Mpps).
[0080] Scan out block 550 operates at a speed corresponding to the target display of client 210. In particular, frame rate calculator 520 determines frame rate settings 525 based on various inputs, including maximum pixel clock 515 of chipset 540 and requested image size 521. Information in requested image size 521 can be taken from values 522, including regular display values (e.g., 480p, 720p, 1080p, 4K, 8K, etc.), as well as other defined values. For the same maximum pixel clock, there can be different frame rate settings depending on the target display of the client, where the frame rate setting is determined by dividing the maximum pixel clock 515 by the number of pixels of the target display. For example, at a maximum pixel clock of 300 megapixels per second, the frame rate setting for a 480p display (e.g., approximately 300k pixels such as used in mobile phones) is approximately 1000 Hz. Also, at the same maximum pixel clock of 300 megapixels per second, the frame rate setting for a 1080p display (approximately 2 megapixels) is approximately 150 Hz. And for illustration, at a maximum pixel clock of 300 megapixels per second, the frame rate setting for a 4k display is approximately 38 Hz.
[0081] The frame rate setting 525 is input to a scan-out setting converter 530, which is configured to determine a speed setting value 535 formatted for a chipset 540. For example, the chipset 540 can operate at a bit rate. In one embodiment, the speed setting value 535 can be the frame rate setting 525 (e.g., frames per second). In some embodiments, the speed setting value 535 can be determined as a multiple of a base frame rate. For example, the speed setting value 535 can be set to a multiple of 30 frames per second (e.g., 30 Hz), such as 30 Hz, 60 Hz, 90 Hz, 120 Hz, 150 Hz, etc. The speed setting value 535 is input to a cache 545 of the chipset 540 for access by a corresponding scan-out block 550 in the chipset to determine its operating speed for the target display of the client 210.
[0082] The chipset 540 includes a game title processing engine 211 configured to execute video game logic 215 of a video game to generate video frames of game renderings to stream back to the client 210. As shown, the game title processing engine 211 includes a CPU 501 and a GPU 502 (e.g., configured to implement a graphics pipeline). In one embodiment, the CPU 501 and GPU 502 are configured as an accelerated processing unit (APU) configured to integrate a CPU and GPU onto the same chip or die using the same bus for faster communication and processing. In another embodiment, the CPU 501 and GPU 502 can be connected through a bus, such as PCI-Express, Gen-Z, etc. A plurality of video frames of game renderings are generated for the video game and placed into a buffer 555 (e.g., a display buffer or frame buffer) that includes one or more game buffers, such as game buffer 0 and game buffer 1. Game buffer 0 and game buffer 1 are driven by a flip control signal to determine which game buffer will store which video frame output from the game title processing engine 211. The game title processing engine is operating at a particular speed defined by the video game. For example, the game title processing engine 211 can output video frames at 30 Hz or 60 Hz, etc.
[0083] Additional information can optionally be generated for inclusion in the video frames of the game rendering. In particular, the feature generation block 560 includes one or more feature generation units, where each unit is configured to generate a feature. Each feature generation unit includes a feature processing engine and a buffer. For example, the feature generation unit 560-A includes a feature processing engine 503. In one implementation, the feature processing engine 503 executes (e.g., on other threads) on the CPU 501 and GPU 502 of the game title processing engine 211. The feature processing engine 503 can be configured to generate a plurality of user interface (UX) features, such as user interfaces, messaging, etc. In one implementation, the UX features can be presented as an overlay. The plurality of UX features generated for the video game are placed into a buffer (e.g., a display buffer or frame buffer) that includes one or more UX buffers, such as UX buffer 0 and UX buffer 1. The UX buffer 0 and UX buffer 1 are driven by a corresponding flip control signal to determine which UX buffer will store which feature output from the feature processing engine 503. Also, the feature processing engine 503 is operating at a particular speed that can be defined by the video game. For example, the feature processing engine 503 can output video frames at 30 Hz or 60 Hz, etc. The feature processing engine 503 can also operate at a speed that is independent of the speed at which the game title processing engine 211 can output video frames (i.e., at a rate different than 30 Hz or 60 Hz, etc.).
[0084] The video frames of the game rendering scanned from the buffer 555 and the optional features scanned from the buffer of the feature generation unit (e.g., unit 560-A) are scanned to the scan out block 550 at a rate X. In one implementation, the rate X for scanning the game buffer 555 holding the video frames of the game rendering and / or the UX buffers holding the features can not correspond to the speed setting value 535, such that the information is scanned out of the buffers as fast as possible. In another implementation, the rate X does correspond to the speed setting value 535.
[0085] As previously described, the scan-out block 550 operates at a speed corresponding to the target display of the client 210. In the case where there can be multiple clients with multiple target displays (e.g., mobile phone, television display, computer display, etc.), there can be multiple scan-out blocks, each supporting a corresponding display, and each operating at a different speed setting value. For example, scan-out block A (550-A) receives the game rendered video frames from the buffer 555 and the feature overlay from the feature generation block 560. The scan-out block A (550-A) operates through the corresponding speed setting value (such as a corresponding frame rate setting) in the cache-A (545-A). As such, for the target display, the scan-out block 550 outputs the modified video frames to the encoder 570 at a rate defined by the speed setting value (e.g., 120Hz). That is, the rate at which the modified video frames are output to the encoder 570 is higher than the rate at which the video frames are generated and / or encoded, where the rate is based on the maximum pixel clock of the chipset 540 including the scan-out block 550 and the image size of the target display.
[0086] In one implementation, the encoder 570 can be part of the chipset 540. In other implementations, the encoder 570 is separate from the chipset 540. The encoder 570 is partially configured to compress the modified video frames for streaming to the client 210. For example, the modified video frames are encoded on the encoder on a slice-by-slice basis to generate one or more encoded slices for a corresponding modified video frame. The one or more encoded slices of the corresponding modified video frame including the additional feature overlay are then streamed over the network to the target display of the client 210. The encoder outputs the one or more encoded slices at a rate that is independent of the speed setting value and can be bound to the synchronized and offset server and client VSYNC signals, as previously described. For example, the one or more encoded slices can be output at 60Hz.
[0087] FIG. 5A-2 A chipset 540B configured for performing a high-speed scan-out operation for delivery to an encoder when streaming content from a video game executing at a cloud gaming server across a network is shown, according to one embodiment of the disclosure, where optional user interface features can be integrally formed into the game rendered video frames. Additionally, the chipset 540 can be configured to perform the scan-out operation earlier, such as before the occurrence of the next system VSYNC signal or at the flip time of the corresponding video frame. FIG. 5A-2 Some of the components shown in FIG. 5A-1 are similar to those of FIG. 5A-2 , where similar features have similar functionality. In the corresponding chipset FIG. 5A-1the difference between FIGS. 5A2 and FIG. 5A-1 FIG. 5A-2 The configuration of the chip set 540B differs from that of FIG. 5A2 in that there is no separate feature generation block. As such, one or more optional UX features can be generated by the CPU 501 and / or GPU 502 and integrally formed into the game-rendered video frames that are put into the buffer 555, as previously described. That is, the features need not be provided as overlays, as they are integrally formed into the rendered video frames. The game-rendered video frames can optionally be scanned from the buffer 555 to the scan-out block 550, which includes one or more scan-out blocks 550-B for one or more target displays of the client. As previously described, the corresponding scan-out block 550-B operates at the speed of the target display. As such, for the target display, the corresponding scan-out block 550-B outputs the video frames to the encoder 570 at the rate defined by the speed setting value. In some embodiments, because the features are integrally formed into the game-rendered video frames, whereby only the buffer 555 is needed, the rendered video frames can be scanned directly to the encoder and bypass the scan-out block 550. In this case, for example, the additional operations performed during scan-out can be performed by the CPU 501 and / or GPU 502.
[0088] FIG. 5B-1 FIG. 5A2 illustrates scan-out operations performed on game-rendered video frames that optionally include one or more additional features (e.g., layers) for delivery to an encoder when content from a video game executed at a cloud gaming server is streamed across a network to a client, according to one embodiment of the present disclosure. For example, FIG. 5B-1 FIG. 5A2 illustrates scan-out operations performed on game-rendered video frames that optionally include one or more additional features (e.g., layers) for delivery to an encoder when content from a video game executed at a cloud gaming server is streamed across a network to a client, according to one embodiment of the present disclosure. For example, FIG. 5A-1 FIG. 5A2 illustrates scan-out operations performed on game-rendered video frames that optionally include one or more additional features (e.g., layers) for delivery to an encoder when content from a video game executed at a cloud gaming server is streamed across a network to a client, according to one embodiment of the present disclosure. For example,
[0089] In particular, the scan-out block A (550-A) receives game-rendered video frames from the buffer 555 and receives a feature overlay from the feature generation block 560, which is provided to the input buffer 580. As previously described, the scan-out block A (550-A) operates through the corresponding speed setting value in the cache-A (545-A), such as the corresponding frame rate setting for the target display of the client 210. For example, a plurality of game-rendered video frames are output from the game buffer 0 and the game buffer 1 to the input frame buffer 580-A of the scan-out block A (550-A) under control of the flip control signal.
[0090] Additionally, scan-out block A (550-A) can optionally receive one or more UX features (e.g., as an overlay). For example, a plurality of UX features are output from buffer 560-A under control of a corresponding flip control signal, which includes UX buffer 0 and UX buffer 1. The plurality of UX features are scanned to input frame buffer 580-B of scan-out block-A (550-A). Other feature overlays can be provided, where exemplary UX features can include user interfaces, system user interfaces, text messages, messaging, menus, communications, additional game viewpoints, e-sports information, etc. For example, an additional plurality of UX features can be output from buffers 560A-560N under control of a corresponding flip control signal, each of which includes UX buffer 0 and UX buffer 1. To illustrate, a plurality of UX features are output from buffer 560-N to input frame buffer 580-N.
[0091] The information in input frame buffers 580 is output to combiner 585, which is configured to compose the information. For example, for each corresponding video frame generated by the video game, combiner 585 combines the game-rendered video frame from input frame buffer 580-A with each of the optional UX features provided in input frame buffers 580-B through 580-N.
[0092] The game-rendered video frame combined with one or more optional UX features is then provided to block 590, where additional operations can be performed to generate a modified video frame suitable for display. In block 590, the additional operations performed during the scan-out process can include one or more operations such as decompression of DCC-compressed surfaces, resolution scaling to a target display, color space conversion, de-gamma, HDR expansion, gamut remapping, LUT shaping, tone mapping, blend gamma, blend, etc.
[0093] In other implementations, the additional operations outlined in block 590 are performed at each of input frame buffers 580 to generate a corresponding layer of modified video frames. For example, the input frame buffers can be used to store and / or generate game-rendered video frames of a video game, as well as one or more optional UX features (e.g., as an overlay) such as user interfaces (UIs), system UIs, text, messaging, etc. The additional operations can include decompression of DCC-compressed surfaces, resolution scaling, color space conversion, de-gamma, HDR expansion, gamut remapping, LUT shaping, tone mapping, blend gamma, etc. After these operations are performed, one or more layers of input frame buffers 580 are composed and blended, optionally put into a display buffer, and then scanned to an encoder (e.g., scanned from the display buffer).
[0094] As such, for the target display, the scan-out block 550-A outputs the plurality of modified video frames to the encoder 570 at a rate defined by the speed setting value (e.g., 120 Hz). That is, the rate at which the modified video frames are output to the encoder 570 is higher than the rate at which the video frames are generated and / or encoded, where the rate is based on the maximum pixel clock of the chipset 540 including the scan-out block 550 and the image size of the target display. As previously mentioned, the encoder 570 compresses each of the modified video frames. For example, the corresponding modified video frames can be compressed into one or more encoded slices (compressed encoder slices), which can also be packetized for network streaming. The modified video frames that have been compressed and / or packetized into encoded slices are then stored into a buffer 580 (e.g., a first-in-first-out or FIFO buffer). The streamer 575 is configured to transmit the encoded slices to the client 210 over the network 250. As previously mentioned, the streamer device can be configured to operate at the application layer of the Transmission Control Protocol / Internet Protocol (TCP / IP) computer network model. In embodiments, assuming an IP-based network (e.g., home / internet), TCP / IP or UDP can be used. For example, a cloud gaming service can use UDP. TCP / IP guarantees all data arrives; however, the “arrival guarantee” comes at the cost of retransmission, which introduces additional latency. On the other hand, UDP-based protocols provide the best latency performance but at the cost of packet loss, which incurs data loss.
[0095] FIG. 5B-2 The scan-out operation is shown performed on a video frame of a game render that can optionally include one or more additional features (e.g., layers) for delivery to an encoder when content from a video game executed at a cloud gaming server is streamed across a network to a client, in accordance with one embodiment of the present disclosure. For example, FIG. 5B-2 The operation of the scan-out block 550-A2 is shown. The scan-out block A2 (550-A2) receives the video frame of the game render scan line by scan line. FIG. 5B-2 The configuration of the scan-out block 550-A2 of FIG. 5B-1 The scan-out block 550-A2 of FIG. 5B-2 The scan-out block 550-A2 of FIG. 5B-1 The scan-out block 550-A2 of
[0096] As shown in the figure, information in each of the input frame buffers 580 is delivered to the corresponding block 590, where additional operations are performed. Specifically, the additional operations outlined in block 590 are performed for each of the input frame buffers 580 to generate the corresponding layer. These additional operations may include decompressing DCC-compressed surfaces, resolution scaling, color space conversion, gamma removal, HDR extension, gamut remapping, LUT shaping, tone mapping, gamma blending, etc. After these operations are performed, one or more modified layers are delivered individually to encoder 570. The encoder delivers each layer individually to a client, where the client can composite and blend the layers to generate modified video frames for display.
[0097] FIG. 5B-3 The scan output operation according to one embodiment of the present disclosure is illustrated, wherein the scan output operation is performed on video frames rendered by the game to be delivered to an encoder when content from a video game executed at a cloud gaming server is streamed across a network to a client. For example, FIG. 5B-3 It shows FIG. 5A-2 The operation of scan output block 550B, where scan output block 550-B does not have combiner functionality. Some components of scan output block 550-B are related to... FIG. 5B-1 The scan output block 550-A is similar, with similar features having similar functionality. FIG. 5B-3 The scan output block 550-B is different. FIG. 5B-1The scan output block 550-A has only a single input frame buffer because it lacks a combiner (e.g., for performing compositing and blending) and because there is no separate feature generation. Specifically, scan output block B (550-B) receives game-rendered video frames scan-line by scan from buffer 555. Optionally, user interface features may be integrally formed into the game-rendered video frames generated by the CPU and / or GPU. For example, multiple game-rendered video frames are output from game buffer 0 and game buffer 1 to the input frame buffer 580 of scan output block B (550-B) under the control of a flip control signal. The game-rendered video frames are then provided to block 590, where additional operations (e.g., decompressing DCC-compressed surfaces, scaling the resolution to the target display, color space conversion, etc.) may be performed to generate modified video frames suitable for display, as previously described. These additional operations may not necessarily require compositing and / or blending, as optional UX features are already integrally formed into the game-rendered video frames. In some implementations, the additional operations outlined in block 590 may be performed at input frame buffer 580. Accordingly, for the target display, scan output block 550-B outputs multiple modified video frames (e.g., at a rate defined by corresponding speed setting values) to encoder 570. As previously described, encoder 570 compresses each of the modified video frames, such as compressing them into one or more encoded slices (compressed encoder slices), which may also be packaged for network streaming. The modified video frames that have been compressed and / or packaged into encoded slices are then stored in buffer 580. As previously described, streamer 575 is configured to transmit the encoded slices to client 210 over the network.
[0098] FIG. 5C-FIG. 5D An exemplary server configuration according to an embodiment of this disclosure is shown, including a scan output block having one or more input frame buffers, which are used when performing a high-speed scan output operation to deliver to an encoder during the streaming of content from a video game executed at a cloud gaming server across a network. Specifically, FIG. 5C-FIG. 5D It shows the use of FIG. 5A-1 The scan output block 550-A and / or FIG. 5A-2 An exemplary configuration of the scan output block 550B includes one or more input frame buffers for generating composite video frames to be displayed on a high-definition display or a virtual reality (VR) display (e.g., a head-mounted display). In one implementation, the input frame buffers may be implemented in hardware.
[0099] FIG. 5CA scan output block 550-A' is shown, comprising four input frame buffers that can be used to generate composite video frames for high-definition displays. By way of example only, three input frame buffers (e.g., FB0, FB1, and FB2) are dedicated to video games and can be used to store and / or generate corresponding layers including at least one of video frames, UI, esports UI, and text layers. The input frame buffers for video games can generate game-rendered video frames from one or more viewpoints in the game environment. Another input frame buffer, FB3, is dedicated to the system and can be used to generate system overlays (e.g., UI), such as friend notifications.
[0100] FIG. 5D A scan output block 550-A” is shown, comprising four input frame buffers that can be used to generate composite video frames for a VR display. By way of example only, two input frame buffers (e.g., FB0 and FB1) are dedicated to video games and can be used to store and / or generate corresponding layers including at least one of video frames taken from different viewpoints of the game environment, a UI, an esports UI, and a text layer. Two other input frame buffers (FB2 and FB3) are dedicated to the system and can be used to generate system overlay layers (e.g., UIs), such as those including friend notifications or esports UIs.
[0101] In embodiments of this disclosure, high-speed and / or early scan output / scan input can be performed at the server without regard to display requirements and / or parameters, since no physical display is connected to the server. Specifically, the server can perform scan output / scan input against a target virtual display, which can be defined by the user to operate at a selected frequency (e.g., 93Hz, 120Hz).
[0102] Through the FIG. 2A-FIG. 2D Detailed description of various client devices 210 and / or cloud gaming networks 290 (e.g., in game server 260), FIG. 6 Flowchart 600 illustrates a method for cloud gaming according to one embodiment of the present disclosure, wherein high-speed and / or early scan output operations can be performed to reduce one-way waiting time between the cloud gaming server and the client.
[0103] At 610, the method includes generating a video frame when executing the video game at the server. For example, the server can execute the video game in a streaming mode such that the CPU of the server executes the video game in response, in part, to input commands from a user or game logic that is not driven by control information from the user in order to generate a video frame of the game rendering using a graphics pipeline available for streaming. In particular, the CPU and GPU graphics pipelines that cooperatively execute the video game are configured to generate a plurality of video frames. In cloud gaming, the video frames generated by the game are typically rendered for display on a virtual display. The server can perform additional operations on the video frames generated by the game in a scan-out process. For example, one or more overlay layers can be added to the corresponding video frames generated by the game, such as during the scan-out process.
[0104] At 620, the method includes performing the scan-out process by scanning the plurality of screen slices of the video frame to one or more input frame buffers line-by-line to perform one or more operations that modify the plurality of screen slices. As previously described, UX features (e.g., overlay layers) can be scanned to the one or more input frame buffers. As such, the one or more input frame buffers can be used to store and / or generate video frames of the game rendering of the video game, as well as one or more optional UX features (e.g., as overlay layers), such as user interfaces (UIs), system UIs, text, messaging, etc. The scan-out process generates modified video frames that are composited and blended to include the one or more optional UX features, such as those implemented by overlay layers. In one implementation, the UX features (e.g., as overlay layers) are composited first and then additional operations are performed, as previously described. For example, the additional operations can include decompression of DCC compressed surfaces, resolution scaling, color space conversion, de-gamma, HDR expansion, color gamut remapping, LUT shaping, tone mapping, blend gamma, etc. In another implementation, the additional operations are performed on each of the UX features before compositing and blending, as previously described.
[0105] At 630, after the modified video frame is generated, the plurality of screen slices of the modified video frame are scanned to an encoder line-by-line in the scan-out process. As such, the modified video frames generated by the game (e.g., modified with optional layers of UX features) are scanned into the encoder for compression in preparation for streaming the modified video frames to a client, such as when streaming content from a video game executed at a cloud gaming server to a client across a network.
[0106] In particular, at 640, the method includes early starting the scan-out process. In one embodiment, multiple screen slices of a game-generated video frame are scanned to one or more input frame buffers at a corresponding flip time of the video frame. That is, instead of waiting for the next occurrence of the server VSYNC signal to start the scan-out process, the modified video frame is scanned to the corresponding input frame buffer earlier (i.e., before the next server VSYNC signal). The flip time can be included in a command in the command buffer that, when the GPU is executing in the graphics pipeline, instructs the GPU that it has completed executing the multiple commands in the command buffer and that the game-rendered video frame is fully loaded to the server’s display buffer. This game-rendered video frame is then scanned to the corresponding input frame buffer during the scan-out process. Additionally, one or more optional UX features (e.g., overlays) are also scanned to the one or more input frame buffers at corresponding flip times generated for the UX features.
[0107] In another embodiment, in accordance with one embodiment of the disclosure, the scan-out process is performed at a high speed when streaming content from a video game executing at a cloud gaming server across a network. For example, the scan-out process operates at a speed / rate corresponding to a target display of the client and is based on a maximum pixel clock of the server and a requested image size of the target display, as previously described. For example, the scan-out process includes receiving a game-rendered video frame and a feature overlay, then compositing, where additional operations can be performed on the composited video frame, such as scaling, color scaling, blending, etc. As previously described, the scan-out process outputs the modified video frame at a scan-out rate based on a speed setting value (e.g., 120 Hz), where the speed setting value is based on the maximum pixel clock of the server and the requested image size of the target display. In one implementation, the speed setting value is a frame rate. As such, the scan-out rate at which the modified video frame is output to the encoder can be higher than the rate at which the video frame is generated and / or encoded.
[0108] Each modified video frame can be partitioned into one or more encoder slices, which are then compressed into one or more encoded slices. In particular, the encoder receives the modified video frame and encodes the modified video frame on a slice-by-slice basis to generate one or more encoded slices. As previously mentioned, the boundaries of the encoded slices are not limited to a single scan line and can include a single scan line or multiple scan lines. Additionally, the end of an encoded slice and / or the beginning of the next encoded slice can not necessarily occur at the edge of the display screen (e.g., can occur somewhere in the middle of the screen or in the middle of a scan line). In one embodiment, because the server VSYNC signal and the client VSYNC signal are synchronized and offset, the operations at the encoder can overlap. In particular, the encoder is configured to generate a first encoded slice of the modified video frame, where the modified video frame can include multiple encoded slices. The encoder can be configured to begin compressing the first encoded slice before the modified video frame is completely received. That is, the first encoded slice can be encoded (e.g., compressed) before multiple screen slices of the modified video frame are completely received, where the screen slices are delivered on a scan line-by-scan line basis. In some embodiments, depending on the number of processors or hardware, multiple slices can be encoded simultaneously (e.g., in parallel). For example, some game consoles can generate four encoded slices in parallel. More particularly, due to hardware pipelining, a hardware encoder can be configured to compress multiple encoder slices in parallel (e.g., to generate one or more encoded slices).
[0109] FIG. 7AA process for generating and transmitting modified video frames at a cloud gaming server is shown according to one embodiment of the disclosure, where the process is optimized to perform a high-speed and / or early scan-out to the encoder to reduce one-way latency between the cloud gaming server and the client. The process is shown with respect to the generation and transmission of a single modified video frame at the server with additional UX features (e.g., overlay) modification. The operations at the server include generating a game-rendered video frame 490 at operation 401. The scan-out process 402 includes delivering the game-rendered video frame 490 to one or more input frame buffers of a scan-out block to generate a composited overlay. That is, the game-rendered video frame 490 is composited with optional UX features (e.g., overlay). Additional operations (e.g., blending, resolution scaling, color space conversion, etc.) are performed on the composited video frame to generate a modified video frame (e.g., game-rendered video frame modified with additional UX feature overlay). During the scan-out process, the modified video frame is scanned to the encoder. At operation 403, the modified video frame is encoded (e.g., compression performed) into an encoded video frame on the encoder on a slice-by-slice basis. At operation 404, the compressed encoded video frame is transmitted from the server to the client.
[0110] As previously described, the scan-out process 402 is shown as being performed early, before the server VSYNC signal 311 occurs. Typically, scan-out begins at the next occurrence of the server VSYNC signal. In one embodiment, early scan-out is performed at the flip time 701, where the flip time occurs when the GPU has completed generating the rendered frame 490, as previously described.
[0111] By performing the early scan-out process, one-way latency between the server and the client can be reduced, as the remaining server operations (e.g., encoding, transmission, etc.) can also begin early and / or overlap. In particular, by performing early scan-out, additional time 725 is gained, where the additional time is defined between the flip time 701 and the next occurrence of the server VSYNC signal. This additional time 725 can offset any adverse latency changes experienced during other operations, such as encoding 403 or transmission 404. For example, if the encoding process 403 takes longer than a frame period, when this encoding process 403 is started early (e.g., not synchronized to begin at the VSYNC signal), the additional time gained can be sufficient for the video frame to be encoded before the next server VSYNC signal. Similarly, the additional time gained by performing the early scan-out operation can be used to reduce latency changes when delivering the video frame to the client (e.g., increased delivery time over the network).
[0112] FIG. 7BTiming is shown when a scan-out process is performed at a cloud game server, in accordance with one embodiment of the present disclosure, in which scan-out is performed early and / or at high speed, such that a video frame can be scanned to an encoder early at the end of the scan-out process, reducing one-way latency between the cloud game server and the client. Typically, an application running on the server (e.g., a video game) requests a “flip” of the display buffer to occur when rendering is complete. The flip occurs during a flip time 701 during a frame period 410, where a flip command is executed by a graphics processing unit (GPU). The flip command is one of a plurality of commands put into a command buffer by a central processing unit (CPU) when executing the application, where the commands in the command buffer are used by the GPU to render a corresponding video frame. As such, the flip indicates that the GPU has completed executing the commands in the command buffer to generate a rendered video frame, and the rendered video frame has been fully loaded to the display buffer of the server. There is a latency period 725 after which the scan-out process 402a is performed on a subsequent occurrence of the server VSYNC signal 311f. That is, in a typical process, the scan-out 402a is performed after the latency period 725, where the modified video frame in the display buffer (e.g., a game-rendered video frame synthesized and blended with optional UX feature overlays) is scanned to an encoder to perform video encoding. That is, the scan-out process typically occurs at the next VSYNC signal and after the latency period, even though the display buffer is full at an earlier time.
[0113] Embodiments of the present disclosure provide for early scan-out 402b of a display buffer to an encoder, such as in a cloud gaming application. As shown in FIG. 7B The scan-out process 402b is triggered early at the flip time 701, rather than at the next occurrence of the server VSYNC signal 311f, as shown in the middle. This allows the encoder to start encoding early while the operations overlap, rather than waiting for the next server VSYNC signal to perform scan-out to deliver to the encoder for encoding / compression. The display time is not affected, as there is no display attached to the server in practice. As previously mentioned, early encoding reduces one-way latency between the server and the client, as there is less chance of missing one or more VSYNCs intended for delivery to the client and / or for display at the client, for a complex video frame.
[0114] FIG. 7CA time period in which scan-out is performed at a high speed is shown in accordance with one embodiment of the present disclosure, such that a video frame can be scanned to an encoder earlier, reducing one-way latency between a cloud gaming server and a client. In particular, when streaming content from a video game executing at a cloud gaming server across a network, a scan-out process can be performed at a high speed, where the scan-out process operates at a speed / rate corresponding to a target display of a client, and based on a server's maximum pixel clock and a requested image size of the target display, as previously described. As such, the scan-out rate at which modified video frames are output to an encoder can be higher than the rate at which video frames are generated and / or encoded. That is, the scan-out rate can not correspond to the rate at which video frames are generated by a video game. For example, the scan-out rate (e.g., frame rate setting) is higher than the frequency of a server VSYNC signal used to generate video frames when a video game is executed at a server.
[0115] In another embodiment, the scan-out speed can not correspond to a refresh rate of a display device of a client (e.g., 60 Hz, etc.). That is, the display rate of a display device at a client and the scan-out speed can not be the same rate. For example, the display rate of a display device at a client can be 60 Hz, or a variable refresh rate, etc., where the scan-out rate is a different rate (e.g., 120 Hz, etc.).
[0116] Generally, a scan-out process of a video frame is performed over an entire frame period (e.g., 16.6 ms at 60 Hz). For example, one representative frame period 410 is shown between two server VSYNC signals 311c and 311d. In embodiments of the present disclosure, rather than performing a scan-out process of a rendered video frame over an entire frame period, the scan-out is performed at a higher rate. By performing a scan-out process (e.g., including scanning to an encoder) at a rate higher than the rate at which frames are processed (e.g., 60 Hz), it can be possible to start an encoding process earlier, such as when waiting for a scan-out process 402 to end before starting encoding 403, or when overlapping scan-out 402 and encoding 403. For example, the scan-out process 402 can be performed in a period 730 (e.g., approximately 8 ms) that is less than the full frame period 410 (e.g., 16.6 ms at 60 Hz).
[0117] In some cases, encoding can begin earlier, such as before the next server VSYNC signal. Specifically, the encoder can begin processing as soon as the minimum amount of data from the corresponding modified video frame (e.g., a game-rendered video frame modified with one or more optional UX features as an overlay) is delivered to the encoder (e.g., 16 or 64 scan lines), and then process any additional data as soon as it arrives. This reduces one-way latency because the chance of missing one or more VSYNCs is smaller when processing complex video frames intended for delivery to the client and / or for display at the client. One-way latency can be due to network jitter and / or increased processing time at the server. For example, a modified video frame with a large amount of data (e.g., scene changes) may take more than one frame cycle to encode. The faster the scan output process, the more time is left for encoding, and the more likely a modified video frame with a large amount of data is to complete the encoding process before the server VSYNC signal intended for delivery to the client.
[0118] In another implementation, the encoding process can be further optimized to ensure the minimum amount of time used for encoding by limiting the encoding resolution to the resolution required by the client display, so that no time is wasted encoding video frames at a higher resolution than the client display can handle or requests at a given moment.
[0119] Through the FIG. 2A-FIG. 2D Detailed description of various client devices 210 and / or cloud gaming networks 290 (e.g., in game server 260), FIG. 8A Flowchart 800A illustrates a method for cloud gaming according to one embodiment of the present disclosure, wherein video displayed on a client can be smoothed in a cloud gaming application, and high-speed and / or early scan output operations can be performed at the server to reduce one-way latency between the cloud gaming server and the client.
[0120] At 810, the method includes generating video frames when a video game is executed at a server. For example, a cloud gaming server may execute a video game in streaming mode, such that the CPU executes the video game in response to input commands from the user, thereby generating video frames for game rendering using a graphics pipeline.
[0121] The server can perform additional operations on the game-generated video frames during the scanout process. For example, one or more overlay layers can be added to the corresponding game-generated video frames, such as during the scanout process. In particular, at 820, the method includes performing a scanout process to generate and deliver a modified video frame to an encoder configured to compress the video frame. The scanout process includes scanning the video frame and one or more user interface features into one or more input frame buffers line-by-line, and compositing and blending the video frame and the one or more user interface (UX) features (e.g., as an overlay layer including user interface (UI), system UI, text, messaging, etc.) into a modified video frame, where the scanout process begins at a flip time of the video frame. As such, the scanout process generates a modified video frame that is composited and blended to include one or more optional UX features, such as those implemented by overlay layers.
[0122] At 830, the method includes transmitting the compressed modified video frame to the client. In particular, each modified video frame can be split into one or more encoder slices, which are then compressed by the encoder into one or more encoded slices. That is, the encoder receives the modified video frame and encodes the modified video frame on a slice-by-slice basis on the encoder to generate one or more encoded slices, which are then packaged and delivered to the client over a network.
[0123] At 840, the method includes determining a target display time of the modified video frame at the client. In particular, when the scanout of the server display buffer occurs at a flip time, but not the next occurrence of the server VSYNC signal, ideal display timing on the client side can be performed based on the time at which the scanout occurred at the server and the game's intent regarding a particular display buffer (e.g., target display buffer VSYNC). The game intent determines whether the frame is intended for the next client VSYNC or actually for the previous VSYNC of the client, as the game ran late in processing the frame.
[0124] At 850, the method includes scheduling a display time of the modified video frame at the client based on the target display time. The client-side strategy for choosing when to display the frame can depend on whether the game is designed for a fixed frame rate or a variable frame rate, and whether the VSYNC timing information is implicit or explicit, as will be further described below with respect to FIG. 8B
[0125] FIG. 8B A timing diagram showing server and client operations in accordance with one embodiment of the disclosure are shown, the operations being performed during the execution of a video game at the server 260 to generate rendered video frames, which are then sent to the client 210 for display. Because the client knows the various timing parameters associated with each of the rendered video frames generated at the server, which can be used to indicate and / or determine the ideal display time, the client can decide when to display these video frames based on one or more policies. In particular, the indication of the ideal display time for a corresponding rendered video frame generated at the server indicates when the game application executing on the server intends to display the rendered video frame with reference to the target occurrence of the server VSYNC signal. This target server VSYNC signal can be converted to a target client VSYNC signal, especially when the server VSYNC signal and the client VSYNC signal are synchronized (e.g., frequency and timing) and aligned using an appropriate offset.
[0126] FIG. 8B The desired synchronization and alignment between the server VSYNC signal and the client VSYNC signal is shown in the middle. In particular, the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 are synchronized so that they have the same frequency and corresponding frame periods. For example, the frame period 410 of the server VSYNC signal 311 is substantially equal to the frame period 415 of the client VSYNC signal 312. In addition, the server VSYNC signal and the client VSYNC signal can be aligned with an offset 430. The timing offset can be determined so that a predetermined number (e.g., 99.99%) of the received video frames arrive at the client for display at the next timely occurrence of the client VSYNC signal. More particularly, the offset is set so that the video frames received within the predetermined number and having the highest variability in terms of the one-way latency between the server and the client arrive just before the next timely occurrence of the client VSYNC signal for display purposes. The correct synchronization and alignment allow the use of the ideal display time for the video frames generated at the server, which can be converted between the server and the client.
[0127] In one embodiment, the timing parameters include an ideal display time, at which the corresponding video frame is intended to be displayed. The ideal display time can refer to a target occurrence of a server VSYNC signal. That is, the ideal display time is explicitly provided in the timing parameters. In one embodiment, the timing parameters can be delivered from the server to the client via some mechanism within one of the data packets used to deliver the encoded video frames. For example, the timing parameters can be added to the data packet header, or the timing parameters can be part of the encoded frame data of the data packet. In another embodiment, the timing parameters can be delivered from the server to the client using the GPU API used to send the data control packets. The GPU API can be configured to send the data control packets from the server to the client over the same data channel used to transmit the compressed rendered video frames. The data control packets are formatted such that the client understands what type of information is included and understands the correct reference to the corresponding rendered video frame. In one implementation, the communication protocol for the GPU API, the format for the data control packets, etc. can be defined in the corresponding software development kit (SDK) for the video game, the signaling information provides the client with notification of the data control packets (e.g., provided in a header, provided in a packet with a flag, etc.). In one implementation, the data control packets bypass the encoding process because their size is minimal.
[0128] In another embodiment, the timing parameters include a flip time and a simulation time delivered from the server to the client, as previously described. The client can use the flip time and the simulation time to determine the ideal display time. That is, the ideal display time is implicitly provided in the timing parameters. The timing parameters can include other information that can be used to infer the ideal display time. In particular, the flip time indicates when the flip of the display buffer occurs, thereby indicating that the corresponding rendered video frame is ready for transmission and / or display. In one embodiment, the scan-out / scan-in process also occurs early in the flip time. The simulation time refers to the time it takes to render a video frame through the CPU and GPU pipelines. The determination of the ideal display time for the corresponding video frame depends on whether the game is executed at a fixed frame rate or a variable frame rate.
[0129] For fixed frame rate games, the client can implicitly determine the target VSYNC timing information from the scan output / scan input timing (e.g., flip timestamps) and the corresponding analog time. For example, the server records the scan output / scan input times for a corresponding video frame and sends them to the client. The client can infer the target occurrence time of the server's VSYNC signal from the scan output / scan input timing and the corresponding analog time, and convert that target occurrence time to the target occurrence time of the client's VSYNC signal. When the game provides ideal display timing (e.g., via a GPU API), the client can explicitly determine the target VSYNC timing information, which may be an integer VSYNC timing or a fractional VSYNC timing. Fractional VSYNC timing can be implemented when the frame processing time exceeds the frame period, where the ideal display timing can be specified by analog time or based on analog time.
[0130] For games with variable frame rates, the client can implicitly determine the ideal target VSYNC timing information from the scan output / scan input timing and analog time of the corresponding video frame. For example, the server records the scan output time and analog time of the corresponding frame and sends them to the client. The client can infer the target appearance time of the server's VSYNC signal for displaying the corresponding video frame from the scan output / scan input timing and analog time, where the target VSYNC signal can be converted into a corresponding target appearance time for the client's VSYNC signal. Alternatively, when the game provides ideal timing via the GPU API, the client can explicitly determine the target VSYNC timing information. In this case, fractional VSYNC timing can be specified by the game, such as providing analog time or display time.
[0131] like FIG. 8B As shown, the server VSYNC signal 311 and the client VSYNC signal 312 occur at a timing of 60 Hz. The server VSYNC signal 311 is synchronized (e.g., at substantially equal frequencies) and aligned (e.g., with an offset) with the client VSYNC signal 312. For example, the occurrence of the server VSYNC signal may be aligned with the occurrence of the client VSYNC signal. Specifically, the occurrence of the server VSYNC signal 311a corresponds to the occurrence of the client VSYNC signal 312a, the server VSYNC signal 311c corresponds to the client VSYNC signal 312c, the server VSYNC signal 311d corresponds to the client VSYNC signal 312d, the server VSYNC signal 311e corresponds to the client VSYNC signal 312e, and so on.
[0132] For purposes of illustration, the server 260 is executing a video game that runs at 30 Hz, such that rendered video frames are generated at 30 Hz (e.g., corresponding to 30 frame periods per second) during a frame period (33.33 milliseconds). As such, the video game can render at most 30 frames per second. Also shown is ideal display timing for the corresponding video frames. The ideal display timing can reflect the intent of the game to display the video frames. As previously noted, the ideal display timing can be determined from the flip times of each frame, also shown. The client can use the ideal display time to determine when to display the video frames according to the employed policy, as described below. For example, video frame A is rendered and ready for display at flip time 0.6 (e.g., 0.6 / 60 at 60 Hz). Moreover, the ideal display timing for video frame A is intended to be displayed at the occurrence of server VSYNC signal 311a, which translates to being intended to be displayed at the client at the occurrence of client VSYNC signal 312a. Similarly, video frame B is rendered and ready for display at flip time 2.1 (e.g., 2.1 / 60 at 60 Hz). The ideal display timing for video frame B is intended to be displayed at the occurrence of server VSYNC signal 311c, which translates to being intended to be displayed at the client at the occurrence of client VSYNC signal 312c. Moreover, video frame C is rendered and ready for display at flip time 4.1 (e.g., 4.1 / 60 at 60 Hz). The ideal display timing for video frame C is intended to be displayed at the occurrence of server VSYNC signal 311e, which translates to being intended to be displayed at the client at the occurrence of client VSYNC signal 312e. Moreover, video frame D is rendered and ready for display at flip time 7.3 (e.g., 7.3 / 60 at 60 Hz). The ideal display timing for video frame D is intended to be displayed at the occurrence of server VSYNC signal 311g, which translates to being intended to be displayed at the client at the occurrence of client VSYNC signal 312g.
[0133] FIG. 8B One problem illustrated in the middle is that video frame D took longer to generate than expected, such that the flip time for video frame D occurs at 7.3, which is after the target occurrence of server VSYNC signal 311g. That is, the server 260 should have completed rendering of video frame D before the occurrence of server VSYNC signal 311g. However, because the ideal display time for video frame D is known or determinable, the client can still display video frame D at the occurrence of client VSYNC signal 312g that aligns with the ideal display time (e.g., server VSYNC signal 311g), even though the server missed its time to generate the video frame.
[0134] FIG. 8BAnother issue illustrated in FIG. 3 is that while video frame B and video frame C are generated at server 260 with proper timing (e.g., intended to be displayed at different server VSYNC signals), due to additional latency experienced during transmission, video frame B and video frame C are received at the client in the same frame period, such that both appear to be intended to be displayed at the client at the same client VSYNC signal 312d. For example, transmission delays cause video frame B and video frame C to arrive in the same frame period. However, with proper buffering and knowledge of the intended display timing for both video frame B and video frame C, the client can determine how and when to display these video frames depending on which strategy is implemented, including following game intent, preferring latency, preferring smoothness, or adjusting client-side VBI settings for a variable refresh rate display.
[0135] For example, one strategy is to follow the game intent determined during execution on the server. Intent can be inferred from the timing of the flip time of the corresponding video frame, such that video frames A, B, and C are intended for display at the next server VSYNC signal. Intent can be explicitly communicated by the video game such that video frame D is intended for display at the previous server VSYNC signal 311e, even though it finishes rendering after that VSYNC signal. In addition, ambiguity of video frame B and video frame C arriving similarly at the client (e.g., arriving in the same frame period) will be resolved by following the game’s intent. As such, with proper buffering, the client can display the video frames in the following order at 60Hz (16.66ms of display per frame): A— A— A— B— C— C— D— D, etc.
[0136] A second strategy is to prefer latency over frame display smoothness, such that the goal is to reduce latency as much as possible and use the least amount of buffering. That is, the video frames are displayed to quickly resolve latency by displaying the latest received video frame at the next client VSYNC signal. As such, ambiguity of video frame B and video frame C arriving similarly at the client (e.g., arriving in the same frame period) will be resolved by discarding video frame B and displaying only video frame C at the next client VSYNC signal. This sacrifices frame smoothness during display, as video frame B is skipped in the sequence of displayed video frames, which can draw the viewer’s attention. As such, with proper buffering, the client can display the video frames in the following order at 60Hz (16.66ms of display per frame): A— A— A— C— C— C— D— D, etc.
[0137] The third strategy favors frame display smoothness over latency. In this case, the additional latency is not a factor and can be addressed with appropriate buffering. That is, the video frames are displayed in a manner that gives the viewer the best viewing experience. The client uses the time between the target VSYNCs as a guide, e.g., the time between B target 312c and C target 312e is two VSYNCs, so no matter what the times are for B and C to arrive at the client, B should be displayed for two frames; the time between C target 312e and D target 312g is two VSYNCs, so no matter what the times are for C and D to arrive at the client, C should be displayed for two frames, and so on. As such, with appropriate buffering, the client can display the video frames at 60 Hz (each frame displayed for 16.66 ms) in the following order: A— A— A— B— B— C— C— D— D, etc.
[0138] The fourth strategy provides for adjusting the client-side VBI timing for displays that support variable refresh rates. That is, a variable refresh rate display allows for increasing or decreasing the VBI interval when displaying a video frame to achieve an instantaneous frame rate for displaying a video frame rendered at the client for display. For example, instead of displaying a video frame rendered for display at the client at every client VSYNC signal, which can require displaying the video frame twice while waiting for the delayed video frame, the refresh rate of the display can be dynamically adjusted for each video frame rendered for display. As such, when a video frame is received, decoded, and rendered at the client for display, the video frame can be displayed to adjust for the variability of the latency. FIG. 9 In the example shown in FIG. 6B, while video frame B and video frame C are generated at the server 260 with appropriate timing (e.g., intended to be displayed at different server VSYNC signals), video frame B and video frame C are received at the client in the same frame period due to the additional latency experienced during transmission. In this case, video frame B can be displayed for a shorter period of time than intended (e.g., less than a frame period) to enable video frame C to be rendered at the determined client and target client VSYNC signal. By way of example, video frame C can have a target occurrence of a server VSYNC signal that is then converted to a target client VSYNC signal, especially when the server VSYNC signal and the client VSYNC signal are synchronized (e.g., frequency and timing) and aligned using an appropriate offset.
[0139] FIG. 9 Components of an example device 900 that can be used to perform aspects of the various embodiments of the present disclosure are shown. For example, An exemplary hardware system suitable for streaming media content and / or receiving streamed media content in accordance with embodiments of the present disclosure is shown, including performing a high-speed scan-out operation or performing scan-out earlier, such as before the occurrence of the next system VSYNC signal or at the flip time of the corresponding video frame, when streaming content from a video game executing at a cloud gaming server, to deliver a modified video frame to an encoder. The block diagram shows a device 900, which can incorporate or be a personal computer, a server computer, a game console, a mobile device, or other digital device, each of which is suitable for practicing embodiments of the present invention. The device 900 includes a central processing unit (CPU) 902 for running software applications and optionally an operating system. The CPU 902 can include one or more homogeneous or heterogeneous processing cores.
[0140] According to various embodiments, the CPU 902 is one or more general-purpose microprocessors having one or more processing cores. Additional embodiments can be implemented using one or more CPUs having microprocessor architectures specifically adapted for highly parallel and computationally intensive applications, such as media and interactive entertainment applications, configured for graphics processing during execution of a game.
[0141] The memory 904 stores applications and data for use by the CPU 902 and the GPU 916. The storage 906 provides non-volatile storage for applications and data and other computer readable media, and can include fixed or removable magnetic disks, flash memory, and CD-ROM, DVD-ROM, Blu-ray discs, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. The user input devices 908 communicate information from one or more users to the device 900, examples of which can include a keyboard, a mouse, a joystick, a touchpad, a touchscreen, a still or video camera, and / or a microphone. The network interface 909 allows the device 900 to communicate with other computer systems via an electronic communications network, and can include wired or wireless communication over local- and wide-area networks such as the Internet. The audio processor 912 is adapted to generate analog or digital audio output from instructions and / or data provided from the CPU 902, the memory 904, and / or the storage 906. The components of the device 900 are connected via one or more data buses 922, which include the CPU 902, the graphics subsystem 914 including the GPU 916 and the GPU cache 918, the memory 904, the data storage 906, the user input devices 908, the network interface 909, and the audio processor 912.
[0142] Graphics subsystem 914 also interfaces with data bus 922 and the components of device 900. Graphics subsystem 914 includes a graphics processing unit (GPU) 916 and graphics memory 918. Graphics memory 918 includes a display memory (e.g., a frame buffer) for storing pixel data for each of the pixels that make up an output image. Graphics memory 918 can be integrated in the same device as GPU 916, connected as a separate device with GPU 916, and / or implemented within memory 904. Pixel data can be provided to graphics memory 918 directly from CPU 902. Alternatively, CPU 902 provides data and / or instructions to GPU 916 for processing, and GPU 916 generates pixel data for one or more output images in accordance with the data and / or instructions. The data and / or instructions can be stored in memory 904 and / or graphics memory 918. In embodiments, GPU 916 includes 3D rendering capability for generating pixel data for output images from instructions and data that define the geometry, lighting, shading, texturing, motion, and / or camera parameters for a scene. GPU 916 can also include one or more programmable execution units capable of executing shader programs.
[0143] Graphics subsystem 914 periodically outputs pixel data for an image from graphics memory 918 to display device 910 or to be projected by a projection system (not shown). Display device 910 can be any device capable of displaying visual information in response to a signal from device 900, including CRT, LCD, plasma, and OLED displays. Device 900 can provide the signal to display device 910 in the form of analog or digital signals.
[0144] Other embodiments for optimizing graphics subsystem 914 can include multi-tenancy GPU operations in which GPU instances are shared among multiple applications, and distributed GPUs that support a single game. Graphics subsystem 914 can be configured as one or more processing devices.
[0145] For example, in one embodiment, graphics subsystem 914 can be configured to perform multi-tenancy GPU functionality in which one graphics subsystem can implement graphics and / or rendering pipelines for multiple games. That is, graphics subsystem 914 is shared among multiple games that are being executed.
[0146] In other embodiments, graphics subsystem 914 includes multiple GPU devices that are combined to perform graphics processing for a single application executing on a corresponding CPU. For example, multiple GPUs can perform alternate forms of frame rendering, where in a sequential frame period, GPU 1 renders a first frame, and GPU 2 renders a second frame, and so on until the last GPU is reached, whereupon the initial GPU renders the next video frame (e.g., GPU 1 renders a third frame if there are only two GPUs). That is, the GPUs rotate when rendering frames. The rendering operations can overlap, where GPU 2 can begin rendering a second frame before GPU 1 completes rendering a first frame. In another implementation, different shader operations can be assigned to multiple GPU devices in the rendering and / or graphics pipeline. A master GPU is performing master rendering and compositing. For example, in a group including three GPUs, master GPU 1 can perform master rendering (e.g., a first shader operation) and compositing output from slave GPU 2 and slave GPU 3, where slave GPU 2 can perform a second shader (e.g., fluid effects, such as rivers) operation, and slave GPU 3 can perform a third shader (e.g., particle smoke) operation, where master GPU 1 composites the results from each of GPU 1, GPU 2, and GPU 3. In this way, different GPUs can be assigned to perform different shader operations (e.g., waving flags, wind, smoke generation, fire, etc.) to render a video frame. In another embodiment, each of the three GPUs can be assigned to different objects and / or portions of a scene corresponding to a video frame. In the above embodiments and implementations, these operations can be performed in the same frame period (simultaneous parallelism) or in different frame periods (sequential parallelism).
[0147] Accordingly, the present disclosure describes methods and systems configured for streaming media content and / or receiving streamed media content, including performing a high-speed scan-out operation or performing scan-out earlier (such as before the occurrence of the next system VSYNC signal or at the flip time of the corresponding video frame) to deliver a modified video frame to an encoder when streaming content from a video game executing at a cloud gaming server.
[0148] It should be appreciated that various embodiments defined herein can be combined or assembled into specific implementations using the various features disclosed herein. Accordingly, provided examples are just some of the possible implementations and are not limited to the various implementations possible by combining various elements. In some examples, some implementations can include fewer elements than disclosed, without departing from the spirit of the disclosed or equivalent implementations.
[0149] Embodiments of the disclosure can be practiced with various computer system configurations including hand-held devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers and the like. Embodiments of the disclosure can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a wire- or wireless-based network.
[0150] With the above embodiments in mind, it should be understood that the embodiments of the disclosure can employ various computer-implemented operations that manipulate data stored in computer systems. This manipulation is often referred to in terms, such as "processing," "computing," "calculating," "determining," or "displaying" to name a few. Any of the operations described herein that form part of the embodiments of the disclosure are useful machine operations. The embodiments of the disclosure also relate to a device or an apparatus for performing these operations. The apparatus can be specially constructed for the required purposes, or it can be a general purpose computer selectively activated or configured by a computer program stored in the computer. In particular, various general purpose machines can be used with computer programs written in accordance with the teachings herein, or it can be more convenient to construct a more specialized apparatus to perform the required operations.
[0151] The present disclosure can also be embodied as computer readable code on a computer readable medium. Any of the data storage devices described herein qualify as computer readable media. The computer readable medium is any data storage device that can store data which can thereafter be read by a computer system. Examples of computer readable media include hard drives, network attached storage devices (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. The computer readable medium can include computer readable tangible medium distributed over a network-coupled computer system so that the computer readable code is stored and executed in a distributed fashion.
[0152] Although the method operations were described in a specific order, it should be understood that other housekeeping operations can be performed in between operations, or operations can be adjusted so that they occur at slightly different times, or can be distributed in different order, or distributed across different systems, without departing from the spirit of the process. Likewise, the various tasks can be distributed amongst numerous computers and / or computer systems.
[0153] While the foregoing disclosure has been described in some detail for purposes of clarity and the specific embodiments described are shown by way of example, it is not intended that the application's embodiments should in any way be construed as limited to the examples disclosed. Therefore, the scope of the embodiments of this application should be determined by the appended claims and equivalents thereof.
Claims
1. A cloud gaming method comprising; generating video frames when a video game is executed at a server, wherein the video frames are stored in a frame buffer; determining a maximum pixel clock of a chipset comprising the frame buffer; determining a frame rate setting based on the maximum pixel clock and an image size of a target display of a client; determining a speed setting value of the chipset based on the frame rate setting; scanning out the video frames to an encoder at the speed setting value.
2. The method of claim 1, further comprising: scanning the video frames from the frame buffer into a scan-out block, wherein the chipset comprises the scan-out block; and scanning out the video frames from the scan-out block at the speed setting value.
3. The method of claim 2, wherein a rate at which the video frames are scanned out from the scan-out block to the encoder is higher than a rate at which the video frames are generated.
4. The method of claim 2, wherein the maximum pixel clock is based on a chip compute setting of the chipset, or a self-diagnostic test, wherein the speed setting value is a bit rate or a frame rate setting.
5. The method of claim 2, further comprising: generating a feature at a feature processing engine and storing the feature in a user interface buffer, wherein the feature is configured as an overlay; scanning the feature from the user interface buffer into the scan-out block; modifying the video frames using the feature to generate modified video frames; and scanning out the modified video frames from the scan-out block to the encoder at the speed setting value.
6. The method of claim 5, wherein the modifying the video frames comprises: combining the feature and the video frames at the scan-out block at the speed setting value; performing one or more additional operations on the combined feature and video frames to generate the modified video frames.
7. The method of claim 6, wherein the one or more additional operations comprise: decompressing a DCC compressed surface; resolution scaling; performing a color space conversion; performing de-gamma; performing HDR expansion; performing gamut remapping; performing LUT reshaping; tone mapping; blending gamma.
8. The method of claim 5, further comprising: encoding the modified video at the encoder.
9. A non-transitory computer-readable medium storing a computer program for cloud gaming, the computer-readable medium comprising: program instructions for generating video frames when a video game is executed at a server, wherein the video frames are stored in a frame buffer; program instructions for determining a maximum pixel clock of a chipset comprising the frame buffer; program instructions for determining a frame rate setting based on the maximum pixel clock and an image size of a target display of a client; program instructions for determining a speed setting value of the chipset based on the frame rate setting; program instructions for scanning out the video frames to an encoder at the speed setting value.
10. The non-transitory computer-readable medium of claim 9, further comprising: program instructions for scanning the video frames from the frame buffer into a scan-out block, wherein the chipset comprises the scan-out block; and program instructions for scan-out of the video frames from the scan-out block at the speed setting value.
11. The non-transitory computer-readable medium of claim 10, wherein in the computer program for cloud gaming, a rate at which the video frames are scan-out from the scan-out block to the encoder is higher than a rate at which the video frames are generated.
12. The non-transitory computer-readable medium of claim 10, wherein in the computer program for cloud gaming, the maximum pixel clock is based on a chip calculation setting of the chipset, or a self-diagnostic test, wherein in the computer program for cloud gaming, the speed setting value is a bit rate or a frame rate setting.
13. A computer system comprising: a processor; and a memory coupled to the processor and having instructions stored therein that, if executed by the computer system, cause the computer system to perform a method for cloud gaming, the method comprising: generating video frames when executing a video game at a server, wherein the video frames are stored in a frame buffer; determining a maximum pixel clock of a chipset comprising the frame buffer; determining a frame rate setting based on the maximum pixel clock and an image size of a target display of a client; determining a speed setting value of the chipset based on the frame rate setting; scan-out of the video frames to an encoder at the speed setting value.
14. The computer system of claim 13, the method further comprising: scanning the video frames from the frame buffer into a scan-out block, wherein the chipset comprises the scan-out block; and scan-out of the video frames from the scan-out block at the speed setting value.
15. The computer system of claim 14, wherein in the method, a rate at which the video frames are scan-out from the scan-out block to the encoder is higher than a rate at which the video frames are generated.
16. The computer system of claim 14, wherein in the method, the maximum pixel clock is based on a chip calculation setting of the chipset, or a self-diagnostic test, wherein in the method, the speed setting value is a bit rate or a frame rate setting.
17. The computer system of claim 14, the method further comprising: generating a feature at a feature processing engine and storing the feature in a user interface buffer, wherein the feature is configured as an overlay; scanning the feature from the user interface buffer into the scan-out block; modifying the video frames using the feature to generate modified video frames; and scan-out of the modified video frames from the scan-out block to the encoder at the speed setting value.
18. The computer system of claim 17, wherein in the method, the modifying the video frames comprises: combining the feature and the video frames at the scan-out block at the speed setting value; performing one or more additional operations on the combined feature and the video frame to generate the modified video frame.
19. The computer system of claim 18, wherein in the method the one or more additional operations include: decompressing a DCC compressed surface; resolution scaling; performing a color space conversion; performing de-gamma; performing HDR expansion; performing gamut remapping; performing LUT shaping; tone mapping; blending gamma.
20. The computer system of claim 17, the method further comprising: encoding the modified video at the encoder.
Citation Information
Patent Citations
Video frame rate compensation through adjustment of vertical blanking
CN104917990A