Cloud game method
By performing high-speed or early scan output operations at the cloud gaming server, video frames are scanned to the encoder line by line, solving the problem of excessively long waiting time between the server and the client in cloud gaming, and improving user experience and video display smoothness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-29
- Publication Date
- 2026-04-03
AI Technical Summary
In cloud gaming, the round-trip and one-way waiting times between the server and the client are relatively long, which affects the user experience.
Perform high-speed or early scan output operations at the cloud gaming server, scan video frames one scan line at a time to the encoder, and start the scan output process at the flip time of the video frame to reduce one-way waiting time.
By reducing one-way waiting time, the user experience quality of cloud gaming is improved, and the smoothness of video display on the client is enhanced.
Smart Images

Figure CN121775437A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with national application number 202080081921.9, international application date of September 29, 2020, entry into the Chinese national phase date of May 25, 2022, and invention title "High-speed scan output of server display buffer for cloud gaming applications". Technical Field
[0002] This disclosure relates to a streaming system configured for streaming content across a network, and more specifically, to performing high-speed scan output operations at a cloud gaming server and / or performing early scan output operations at the server to reduce latency between the cloud gaming server and the client, wherein smoothness of video display on the client can be improved by transmitting an ideal display time to the client. Background Technology
[0003] In recent years, online services have been continuously advancing, enabling online or cloud gaming in streaming formats between cloud gaming servers and clients connected via a network. Streaming formats are gaining popularity due to the on-demand availability of game titles, the ability for players to network for multiplayer games, asset sharing between players, instant experience sharing between players and / or spectators, allowing friends to watch friends play video games, and letting friends join games they are playing, among other things. Unfortunately, this demand also challenges the limitations of network connectivity and the processing capabilities performed at both the server and client ends, where the processing response should be sufficient to render high-quality images delivered to the client. For example, the results of all game activities performed on the server need to be compressed and transmitted back to the client with low millisecond latency for optimal user experience. Round-trip latency can be defined as the total time between the user's controller input and the display of a video frame at the client; it may include the processing and transmission of control information from the controller to the client, the processing and transmission of control information from the client to the server, the generation of a video frame at the server using the input in response to the input, processing the video frame and passing it to an encoding unit (e.g., a scan output), encoding the video frame, transmitting the encoded video frame back to the client, receiving and decoding the video frame, and any processing or grading of the video frame before its display. One-way latency can be defined as a portion of the round-trip latency, consisting of the time from the start of passing the video frame to the encoding unit (e.g., a scan output) at the server to the start of displaying the video frame at the client. Both round-trip latency and a portion of one-way latency are associated with the time it takes for data to flow through the communication network from the client to the server and from the server to the client. Another portion is associated with processing at both the client and server; improvements to these operations (such as advanced strategies related to frame decoding and display) can significantly reduce round-trip latency and one-way latency between the server and client, providing a higher quality experience for users of cloud gaming services.
[0004] It is against this backdrop that the proposed implementation plan was developed. Summary of the Invention
[0005] Embodiments of this disclosure relate to streaming systems configured for streaming content (e.g., games) across a network, and more specifically, to performing high-speed scan output operations or performing scan output earlier (such as before the next system VSYNC signal occurs or at the flip time of the corresponding video frame) for delivering modified video frames to an encoder.
[0006] This disclosure discloses a method for cloud gaming. The method includes generating video frames when a video game is executed at a server. The method includes performing a scan output process by scanning the video frame and one or more user interface features scan line-by-scan into one or more input frame buffers, and by synthesizing and blending the video frame and one or more user interface features into a modified video frame. The method includes scanning the modified video frame scan line-by-scan into an encoder at the server during the scan output process. The method also includes initiating scanning the video frame and one or more user interface features into one or more input frame buffers at corresponding flip times during the scan output process.
[0007] In another embodiment, a non-transitory computer-readable medium storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating video frames when a video game is executed at a server. The computer-readable medium includes program instructions for performing a scan-output process by scanning the video frames and one or more user interface features scanned scan-line by scan into one or more input frame buffers, and by combining and mixing the video frames and one or more user interface features into a modified video frame. The computer-readable medium includes program instructions for scanning the modified video frames scanned scan-line by scan into an encoder at the server during the scan-output process. The computer-readable medium includes program instructions for initiating scanning of the video frames and one or more user interface features into one or more input frame buffers at corresponding flip times of the video frames during the scan-output process.
[0008] In another embodiment, the computer system includes a processor and a memory coupled to the processor and storing instructions therein that, if executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating video frames when a video game is executed at a server. The method includes performing a scan-output process by scanning the video frame and one or more user interface features scan line-by-scan into one or more input frame buffers and combining and blending the video frame and one or more user interface features into a modified video frame. The method includes scanning the modified video frame scan line-by-scan into an encoder at the server during the scan-output process. The method includes initiating scanning the video frame and one or more user interface features into one or more input frame buffers at corresponding flip times of the video frame during the scan-output process.
[0009] In another embodiment, a method for cloud gaming is disclosed. The method includes generating video frames when a video game is executed at a server. The method includes performing a scan-out process to deliver the video frames to an encoder configured to compress the video frames, wherein the scan-out process begins at the flip time of the video frames. The method includes transmitting the compressed video frames to a client. The method includes determining a target display time for the video frames at the client. The method includes scheduling the display time of the video frames at the client based on the target display time.
[0010] In another embodiment, a non-transitory computer-readable medium storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating video frames when a video game is executed at a server. The computer-readable medium includes program instructions for performing a scan-output process to deliver video frames to an encoder configured to compress video frames, wherein the scan-output process begins at the flip time of the video frames. The computer-readable medium includes program instructions for transmitting the compressed video frames to a client. The computer-readable medium includes program instructions for determining a target display time for the video frames at the client. The computer-readable medium includes program instructions for scheduling the display time of the video frames at the client based on the target display time.
[0011] In another embodiment, the computer system includes a processor and a memory coupled to the processor and storing instructions that, if executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating video frames when a video game is executed at a server. The method includes performing a scan output process to deliver the video frames to an encoder configured to compress the video frames, wherein the scan output process begins at the flip time of the video frames. The method includes transmitting the compressed video frames to a client. The method includes determining a target display time for the video frames at the client. The method includes scheduling the display time of the video frames at the client based on the target display time.
[0012] In another embodiment, a method for cloud gaming is disclosed. The method includes generating video frames when a video game is executed at a server. The method includes performing a scan-output process to deliver the video frames to an encoder configured to compress the video frames, wherein the scan-output process includes: scanning the video frames and one or more user interface features scan line-by-scan into one or more input frame buffers, and combining and blending the video frames and one or more user interface features into a modified video frame, wherein the scan-output process begins at the flip time of the video frame. The method includes transmitting the compressed modified video frame to a client. The method includes determining a target display time for the modified video frame at the client. The method includes scheduling the display time of the modified video frame at the client based on the target display time.
[0013] In another embodiment, a non-transitory computer-readable medium storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating video frames when a video game is executed at a server. The computer-readable medium includes program instructions for performing a scan-output process to deliver video frames to an encoder configured to compress video frames, wherein the scan-output process includes: scanning the video frames and one or more user interface features scan line-by-scan into one or more input frame buffers, and combining and mixing the video frames and one or more user interface features into a modified video frame, wherein the scan-output process begins at the flip time of the video frame. The computer-readable medium includes program instructions for transmitting the compressed modified video frame to a client. The computer-readable medium includes program instructions for determining a target display time for the modified video frame at the client. The computer-readable medium includes program instructions for scheduling the display time of the modified video frame at the client based on the target display time.
[0014] In another embodiment, the computer system includes a processor and a memory coupled to the processor and storing instructions therein that, if executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating video frames when a video game is executed at a server. The method includes performing a scan-out process to deliver the video frames to an encoder configured to compress the video frames, wherein the scan-out process includes: scanning the video frames and one or more user interface features scanned scan-line by scan into one or more input frame buffers, and combining and mixing the video frames and one or more user interface features into a modified video frame, wherein the scan-out process begins at the flip time of the video frame. The method includes transmitting the compressed modified video frame to a client. The method includes determining a target display time for the modified video frame at the client. The method includes scheduling the display time of the modified video frame at the client based on the target display time.
[0015] In another embodiment, a method for cloud gaming is disclosed. The method includes generating video frames when a video game is executed at a server, wherein the video frames are stored in a frame buffer. The method includes determining a maximum pixel clock of a chipset including a scan output block. The method includes determining a frame rate setting based on the maximum pixel clock and an image size of a target display of a client. The method includes determining a speed setting value for the chipset. The method includes scanning the video frames from the frame buffer into the scan output block. The method includes scanning the video frames from the scan output block to the encoder at the speed setting value.
[0016] In another embodiment, a non-transitory computer-readable medium storing computer programs for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating video frames when a video game is executed at a server, wherein the video frames are stored in a frame buffer. The computer-readable medium includes program instructions for determining a maximum pixel clock of a chipset including a scan output block. The computer-readable medium includes program instructions for determining a frame rate setting based on the maximum pixel clock and the image size of a target display on the client. The computer-readable medium includes program instructions for determining a speed setting value for the chipset. The computer-readable medium includes program instructions for scanning video frames from the frame buffer into the scan output block. The computer-readable medium includes program instructions for scanning video frames from the scan output block to an encoder at the speed setting value.
[0017] In another embodiment, the computer system includes a processor and a memory coupled to the processor and storing instructions therein that, if executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating video frames when a video game is executed at a server, wherein the video frames are stored in a frame buffer. The method includes determining a maximum pixel clock of a chipset including a scan output block. The method includes determining a frame rate setting based on the maximum pixel clock and an image size of a target display on a client. The method includes determining a speed setting value for the chipset. The method includes scanning the video frames from the frame buffer into the scan output block. The method includes scanning the video frames from the scan output block to the encoder at the speed setting value.
[0018] In another embodiment, a method for cloud gaming is disclosed. The method includes generating a video frame within a frame period when a video game is executed at a server; generating one or more user interface features for the video frame; performing a scan output process by scanning the video frame and the one or more user interface features scan line-by-scan into one or more input frame buffers, and synthesizing and blending the video frame and the one or more user interface features into a modified video frame; during the scan output process, scanning the modified video frame scan line-by-scan into an encoder at the server; and during the scan output process, at a flip time of the video frame occurring within the video frame, starting to scan the video frame and begin scanning the one or more user interface features into the one or more input frame buffers.
[0019] In another embodiment, a method for cloud gaming is disclosed. The method includes generating video frames while a video game is executed at a server; performing a scan-output process to deliver the video frames to an encoder configured to compress the video frames, wherein the scan-output process begins at a flip time of the video frames; transmitting the compressed video frames to a client; determining a target display time for the video frames at the client; and scheduling the display time of the video frames at the client based on the target display time.
[0020] In another embodiment, a method for cloud gaming is disclosed. The method includes receiving video frames from a frame buffer; determining a maximum pixel clock of a chipset including the frame buffer, wherein the maximum pixel clock is a static value; determining a frame rate setting based on the maximum pixel clock and a target display of a client device; and scanning the video frames to an encoder at the frame rate setting before the next occurrence of a VSYNC signal used for timing the generation of multiple video frames, wherein the frame rate setting is determined independently of the rate at which the multiple video frames are generated.
[0021] In another embodiment, a method for cloud gaming is disclosed. The method includes generating a video frame within a frame period when a video game is executed at a server, wherein generating the video frame includes a flip time occurring before the next occurrence of the server's vertical synchronization (VSYNC) signal; executing a command indicating that the video frame has been fully rendered; and performing a scan output process on the video frame to deliver the video frame to an encoder configured to compress the video frame, wherein performing the scan output process on the video frame begins at the flip time.
[0022] In another embodiment, a method for cloud gaming is disclosed. The method includes aligning a server VSYNC signal of a server and a client VSYNC signal of a client; receiving an encoded video frame from the server at the client; receiving timing information for the video frame from the server at the client; determining a target display time for the video frame based on the timing information, wherein the target display time is a target occurrence of the server VSYNC signal; converting the target display time into a target occurrence of the client VSYNC signal; and scheduling the display time of the video frame based on the target occurrence of the client VSYNC signal.
[0023] In another embodiment, a method for cloud gaming is disclosed. The method includes, at a client, receiving an encoded video frame from a server; decoding the encoded video frame; receiving timing information for a target occurrence of a server VSYNC signal, wherein the video frame is targeted for display when the target occurrence of the server VSYNC signal occurs; converting the target occurrence of the server VSYNC signal into a target occurrence of a client VSYNC signal; and displaying the decoded video frame when the target occurrence of the client VSYNC signal occurs. Other aspects of this disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate the principles of this disclosure by way of example. Attached Figure Description
[0024] This disclosure is best understood by referring to the following description taken in conjunction with the accompanying drawings, in which:
[0025] Figure 1A This is a diagram of the VSYNC signal at the beginning of a frame period according to one embodiment of the present disclosure.
[0026] Figure 1B This is a frequency diagram of the VSYNC signal according to one embodiment of the present disclosure.
[0027] Figure 2A This is an illustration of a system according to one embodiment of the present disclosure for providing games over a network between one or more cloud gaming servers and one or more client devices in various configurations, wherein the VSYNC signal can be synchronized and offset to reduce one-way latency.
[0028] Figure 2B This is an illustration of an embodiment of the present disclosure for providing a game between two or more peer devices, wherein the VSYNC signal can be synchronized and offset to achieve optimal timing for receiving controller and other information between the devices.
[0029] Figure 2CVarious network configurations that benefit from proper synchronization and offset of the VSYNC signal between the source and target devices are shown according to one embodiment of this disclosure.
[0030] Figure 2D A multi-tenant configuration is shown for a cloud gaming server and multiple clients according to one embodiment of the present disclosure, which benefits from proper synchronization and offset of the VSYNC signal between the source device and the target device.
[0031] Figure 3 The illustration shows the variation in one-way latency between the cloud gaming server and the client due to clock drift when streaming video frames generated from a video game executed on a server, according to one embodiment of the present disclosure.
[0032] Figure 4 The diagram illustrates the network configuration for cloud gaming servers and clients when streaming video frames generated from a video game running on a server. The VSYNC signal between the server and client is synchronized and offset to allow overlapping operations at the server and client, and to reduce one-way latency between the server and client.
[0033] Figure 5A-1 An accelerated processing unit (APU) according to one embodiment of the present disclosure is shown, the APU being configured to perform high-speed scan output operations to deliver to an encoder, or alternatively, a CPU and GPU connected via a bus (e.g., PCI Express), when content from a video game executed at a cloud gaming server is streamed across a network.
[0034] Figure 5A-2 A chipset 540B according to one embodiment of the present disclosure is shown, which is configured to perform high-speed scan output operations to deliver to an encoder when content from a video game executed at a cloud gaming server is streamed across a network, wherein user interface features are integrally formed into the video frames rendered by the game.
[0035] Figure 5B-1 , Figure 5B-2 and Figure 5B-3 The scan output operation according to one embodiment of the present disclosure is shown, which is performed when content from a video game executed at a cloud gaming server is streamed across a network to a client to generate modified video frames for delivery to an encoder.
[0036] Figures 5C-5DAn exemplary server configuration with one or more input frame buffers according to an embodiment of the present disclosure is shown, which is used when performing a high-speed scan output operation to deliver to an encoder during the streaming of content from a video game executed at a cloud gaming server across a network.
[0037] Figure 6 This is a flowchart illustrating a method for cloud gaming according to one embodiment of the present disclosure, wherein an early scan output process is performed to initiate the encoding process earlier, thereby reducing the one-way waiting time between the server and the client.
[0038] Figure 7A A process for generating and transmitting video frames at a cloud gaming server according to one embodiment of the present disclosure is shown, wherein the process is optimized to perform high-speed and / or early scan output to the encoder to reduce one-way latency between the cloud gaming server and the client.
[0039] Figure 7B The timing of the scan output process performed at a cloud gaming server according to one embodiment of the present disclosure is shown, wherein the scan output is performed at high speed and / or early so that video frames can be scanned to the encoder earlier, thereby reducing the one-way waiting time between the cloud gaming server and the client.
[0040] Figure 7C The illustration shows a time period for high-speed scanning output according to one embodiment of the present disclosure, so that video frames can be scanned to the encoder earlier, thereby reducing the one-way waiting time between the cloud gaming server and the client.
[0041] Figure 8A The flowchart illustrates a method for cloud gaming according to one embodiment of the present disclosure, wherein video displayed on a client can be smoothed in a cloud gaming application, and high-speed and / or early scan output operations can be performed at the server to reduce one-way latency between the cloud gaming server and the client.
[0042] Figure 8B A timing diagram of server and client operations according to one embodiment of the present disclosure is shown, wherein the operations are performed at the server while a video game is being executed to generate game-rendered video frames, which are then sent to the client for display.
[0043] Figure 9 Components of an example apparatus that can be used to carry out various embodiments of this disclosure are shown. Detailed Implementation
[0044] While the following detailed description contains numerous specific details for illustrative purposes, those skilled in the art will recognize that many variations and modifications of these details are within the scope of this disclosure. Therefore, aspects of this disclosure described below are set forth without diminishing the generality of the claims or imposing any limitations on them.
[0045] Generally, various embodiments of this disclosure describe methods and systems configured to reduce latency and / or latency instability between a source device and a target device when streaming media content (e.g., streaming audio and video from a video game). Latency instability can be introduced into the one-way latency between the server and client due to additional time required to generate complex frames (e.g., scene changes) at the server, increased time for encoding / compressing complex frames at the server, variable communication paths on the network, and increased time for decoding complex frames at the client. Latency instability can also be introduced due to clock differences at the server and client, causing drift between the server's VSYNC signal and the client's VSYNC signal. In embodiments of this disclosure, one-way latency between the server and client can be reduced in cloud gaming applications by performing a high-speed scan output of the cloud gaming display buffer. In another embodiment, one-way latency can be reduced by performing an early scan output of the cloud gaming display buffer. In yet another embodiment, when addressing latency issues, the smoothness of video displayed on the client in a cloud gaming application can be improved by transmitting an ideal display time to the client.
[0046] In particular, in some embodiments of this disclosure, one-way latency in cloud gaming applications can be reduced by starting the encoding process earlier. For example, in some architectures for streaming media content (e.g., streaming audio and video from a video game) from a cloud gaming server to a client, the scan output of the server display buffer includes performing additional operations on video frames to generate one or more layers, which are then combined and scanned to the unit performing video encoding. By performing the scan output at a high speed (120Hz or even higher), it is possible to start the encoding process earlier, and thus reduce one-way latency. Moreover, in some embodiments of this disclosure, one-way latency in cloud gaming applications can be reduced by performing an early scan output process at the cloud gaming server. In particular, in some architectures for streaming media content (e.g., streaming audio and video from a video game) from a cloud gaming server to a client, an application running on the server (e.g., a video game) requests a "flip" of the server display buffer that occurs when rendering a video frame is complete. Instead of performing the scan output operation upon the subsequent occurrence of the server VSYNC signal, the scan output operation begins at the flip time, where the scan output of the server display buffer includes performing additional operations on video frames to generate one or more layers, which are then combined and scanned to the unit performing video encoding. By performing the scan output at the flip time (rather than the next VSYNC), it is possible to start the encoding process earlier and thus reduce one-way latency. Since no display is actually attached to the cloud gaming server, display timing is unaffected. In some embodiments of this disclosure, when the server scan output of the display buffer is performed at the flip time (rather than the subsequent VSYNC), the ideal display timing at the client depends on the timing of the scan output and the game's intent regarding that specific display buffer (e.g., whether it is for the next VSYNC, or whether the game is running late and it is actually for the previous VSYNC). The strategy varies depending on whether the game has a fixed or variable frame rate, and whether the information will be implicit (inferred from the scan output timing) or explicit (the game provides the ideal timing via the GPU API, which may be VSYNC or fractional time).
[0047] Building upon the above general understanding of the various implementation schemes, exemplary details of the implementation schemes will now be described with reference to the various accompanying drawings.
[0048] Throughout this specification, references to "game," "video game," or "game application" are intended to refer to any type of interactive application that is initiated by executing input commands. For illustrative purposes only, interactive applications include applications for games, word processing, video processing, video game processing, etc. Furthermore, the terms used above are interchangeable.
[0049] Cloud gaming involves executing a video game on a server to generate game-rendered video frames, which are then sent to a client for display. The timing of operations at both the server and client can be associated with corresponding vertical synchronization (VSYNC) parameters. When the VSYNC signals are properly synchronized and / or offset between the server and / or client, operations performed on the server (e.g., generating and transmitting video frames within one or more frame periods) are synchronized with operations performed on the client (e.g., displaying video frames on a display at a display frame or refresh rate corresponding to the frame period). Specifically, the server-generated VSYNC signal and the client-generated VSYNC signal can be used to synchronize operations at the server and client. That is, when the server and client VSYNC signals are synchronized and / or offset, the server's generation and transmission of video frames are synchronized with how the client displays those video frames.
[0050] VSYNC signaling and Vertical Blanking Interval (VBI) have been incorporated for generating and displaying video frames when streaming media content between a server and a client. For example, the server attempts to generate video frames for game rendering within one or more frame cycles defined by the corresponding server VSYNC signal (e.g., generating one video frame per frame cycle results in 60Hz operation if the frame cycle is 16.7ms, and generating one video frame every two frame cycles results in 30Hz operation), and then encodes and transmits the video frame to the client. At the client, the received encoded video frames are decoded and displayed, with the client displaying each video frame rendered for display, starting with the corresponding client VSYNC.
[0051] For the purpose of explanation, Figure 1A This illustrates how the VSYNC signal 111 can indicate the start of a frame period, during which various operations can be performed at the server and / or client. When streaming media content, the server can use the server VSYNC signal to generate and encode video frames, and the client can use the client VSYNC signal to display video frames. The VSYNC signal 111 is generated at a defined frequency corresponding to the defined frame period 110, as follows: Figure 1B As shown in the figure. Additionally, VBI 105 defines the time period between when the last raster line of the previous frame period is drawn on the display and when the first raster line (e.g., the top) is drawn on the display. As shown, after VBI 105, the rendered video frames for display are displayed via raster scan lines 106 (e.g., raster lines from left to right).
[0052] Furthermore, various embodiments of this disclosure are disclosed for reducing one-way latency and / or latency instability between a source device and a target device, such as in streaming media content (e.g., video game content). For illustrative purposes only, various embodiments for reducing one-way latency and / or latency instability are described within a server and client network configuration. However, it should be understood that the various techniques disclosed for reducing one-way latency and / or latency instability can be implemented in other network configurations and / or on peer-to-peer networks, such as... Figures 2A to 2D As shown in the illustration. For example, various disclosed implementations for reducing one-way wait times and / or wait time instability can be implemented between one or more server and client devices in various configurations (e.g., server and client, server and server, server and multiple clients, server and multiple servers, client and client, client and multiple clients, etc.).
[0053] Figure 2A This is an illustration of a system 200A, according to one embodiment of the present disclosure, for providing games in various configurations between one or more cloud gaming networks 290 and / or servers 260 and one or more client devices 210 via network 250, wherein server and client VSYNC signals can be synchronized and offset, and / or wherein dynamic buffering is performed on the client, and / or wherein encoding and transmission operations on the server can overlap, and / or wherein receiving and decoding operations at the client can overlap, and / or wherein decoding and display operations on the client can overlap, to reduce one-way latency between server 260 and client 210. Specifically, according to one embodiment of the present disclosure, system 200A provides games via cloud gaming network 290, wherein the game is executed remotely by client device 210 (e.g., a thin client) of the corresponding user playing the game. System 200A can provide game control via network 250 to one or more users playing one or more games via cloud gaming network 290 in single-player or multi-player mode. In some implementations, cloud gaming network 290 may include multiple virtual machines (VMs) running on a host hypervisor, wherein one or more VMs are configured to utilize hardware resources available to the host hypervisor to execute a game processor module. Network 250 may include one or more communication technologies. In some implementations, network 250 may include 5G network technology with advanced wireless communication systems.
[0054] In some implementations, wireless technologies can be used to facilitate communication. Such technologies may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. A 5G network is a digital cellular network in which the service area covered by the provider is divided into small geographical areas called cells. Analog signals representing voice and images are digitized in a telephone call, converted by an analog-to-digital converter, and transmitted as a bit stream. All 5G wireless devices in a cell communicate with the local antenna array and low-power automatic transceivers (transmitters and receivers) in the cell via radio waves on frequency channels assigned by the transceiver from a frequency pool that is reused in other cells. The local antennas are connected to the telephone network and the Internet via high-bandwidth fiber optic or wireless backhaul connections. As in other cellular networks, mobile devices moving from one cell to another are automatically transferred to the new cell. It should be understood that 5G networks are merely example types of communication networks, and embodiments of this disclosure may utilize previous generations of wireless or wired communication, as well as subsequent generations of wired or wireless technologies after 5G.
[0055] As shown in the figure, the cloud gaming network 290 includes a game server 260 that provides access to multiple video games. The game server 260 may be any type of server computing device available in the cloud and may be configured to execute one or more virtual machines on one or more hosts. For example, the game server 260 may manage virtual machines that support game processors that instantiate game instances for users. Accordingly, multiple game processors associated with multiple virtual machines on the game server 260 are configured to execute multiple instances of one or more games associated with multiple users. In this way, the backend server supports streaming media (e.g., video, audio, etc.) of multiple game applications to multiple corresponding users. That is, the game server 260 is configured to stream data (e.g., rendered images and / or frames of the corresponding game) back to the corresponding client device 210 via network 250. In this way, computationally complex game applications can be executed at the backend server in response to controller input received and forwarded by the client device 210. Each server is capable of rendering images and / or frames, then encoding (e.g., compressing) them and streaming them to the corresponding client device for display.
[0056] For example, multiple users can access the cloud gaming network 290 via a communication network 250 using a corresponding client device 210 configured to receive streaming media. In one embodiment, the client device 210 may be configured as a thin client, providing an interface to a backend server (e.g., game server 260 of the cloud gaming network 290) configured to provide computational functionality (e.g., including a game title processing engine 211). In another embodiment, the client device 210 may be configured with a game title processing engine and game logic for at least some local processing of the video game, and may also be used to receive streaming content generated by the video game executed at the backend server, or other content supported by the backend server. For local processing, the game title processing engine includes basic processor-based functionality for executing the video game and services associated with the video game. The game logic is stored on the local client device 210 and used to execute the video game.
[0057] Specifically, the client device 210 corresponding to a user (not shown) is configured to request access to the game via a communication network 250 such as the Internet, and to render display images generated by the video game executed by the game server 260, wherein encoded images are delivered to the client device 210 for display in association with the corresponding user. For example, the user can interact with an instance of the video game executed on the game processor of the game server 260 via the client device 210. More specifically, the instance of the video game is executed by the game title processing engine 211. The corresponding game logic (e.g., executable code) 215 implementing the video game is stored and accessible through a data storage area (not shown), and is used to execute the video game. The game title processing engine 211 is capable of supporting multiple video games using multiple game logics, each of which can be selected by the user.
[0058] For example, client device 210 is configured to interact with a game title processing engine 211 associated with the game of the corresponding user, such as through input commands used to drive the game. Specifically, client device 210 can receive input from various types of input devices (e.g., game controllers, tablets, keyboards, gestures captured by cameras, mice, touchpads, etc.). Client device 210 may be any type of computing device, having at least a memory and processor module, and is capable of connecting to game server 260 via network 250. The backend game title processing engine 211 is configured to generate rendered images, which are then delivered via network 250 for display on a corresponding display associated with client device 210. For example, through a cloud-based service, the game-rendered image may be delivered by an instance of the corresponding game running on the game execution engine 211 of game server 260. That is, client device 210 is configured to receive encoded images (e.g., encoded by game-rendered images generated by executing a video game) and to display them as images rendered for display 11. In one embodiment, display 11 includes an HMD (e.g., displaying VR content). In some implementations, the rendered image can be streamed wirelessly or wired directly from a cloud-based service or via a client device 210 (e.g., PlayStation® Remote Play) to a smartphone or tablet.
[0059] In one implementation, the game server 260 and / or the game title processing engine 211 include basic processor-based functions for executing the game and services associated with the game application. For example, processor-based functions include 2D or 3D rendering, physics, physics simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadows, culling, transformation, artificial intelligence, etc. Additionally, services for the game application include memory management, multithreading management, Quality of Service (QoS), bandwidth testing, social networks, social friend management, social network communication with friends, communication channels, SMS, instant messaging, chat support, etc.
[0060] In one implementation, the cloud gaming network 290 is a distributed game server system and / or architecture. Specifically, a distributed game engine executing game logic is configured as a corresponding instance of the corresponding game. Generally, the distributed game engine takes each of the functions of the game engine and distributes these functions for execution by multiple processing entities. Individual functions may also be distributed across one or more processing entities. Processing entities may be configured in different configurations, including physical hardware, and / or configured as virtual parts or virtual machines, and / or configured as virtual containers, where a container differs from a virtual machine because it virtualizes an instance of a game application running on a virtualized operating system. Processing entities may utilize and / or rely on servers on one or more servers (compute nodes) of the cloud gaming network 290 and their underlying hardware, where servers may reside on one or more racks. The coordination, assignment, and management of the execution of these functions by the various processing entities are performed by a distributed synchronization layer. In this way, the execution of these functions is controlled by the distributed synchronization layer to support the generation of media (e.g., video frames, audio, etc.) for the game application in response to player controller input. The distributed synchronization layer enables these functions to be performed efficiently across distributed processing entities (e.g., through load balancing), allowing critical game engine components / functions to be distributed and reassembled for more efficient processing.
[0061] The game title processing engine 211 includes a central processing unit (CPU) and a graphics processing unit (GPU) group that can be configured to perform multi-tenant GPU functionality. In another embodiment, multiple GPU devices are combined to perform graphics processing for a single application running on a corresponding CPU.
[0062] Figure 2B This is an illustration of an embodiment of the present disclosure for providing a game between two or more peer devices, wherein the VSYNC signal can be synchronized and offset to achieve optimal timing for receiving controller and other information between the devices. For example, a head-mounted game can be performed using two or more peer devices connected via a network 250 or directly via peer-to-peer communication (e.g., Bluetooth, LAN, etc.).
[0063] As shown in the figure, the game is executed locally on each of the client devices 210 (e.g., game consoles) of the corresponding user playing the video game, where the client devices 210 communicate via a peer-to-peer network. For example, the instance of the video game is executed by the game title processing engine 211 of the corresponding client device 210. The game logic 215 (e.g., executable code) that implements the video game is stored on the corresponding client device 210 and used to execute the game. For illustrative purposes, the game logic 215 may be delivered to the corresponding client device 210 via portable media (e.g., optical media) or via a network (e.g., downloaded from a game provider via the Internet).
[0064] In one implementation, the game title processing engine 211 corresponding to client device 210 includes basic processor-based functions for executing the game and services associated with the game application. For example, processor-based functions include 2D or 3D rendering, physics, physics simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadows, culling, transformation, artificial intelligence, etc. Additionally, services for the game application include memory management, multithreading management, Quality of Service (QoS), bandwidth testing, social networks, social friend management, social network communication with friends, communication channels, SMS, instant messaging, chat support, etc.
[0065] Client device 210 can receive input from various types of input devices, such as game controllers, tablets, keyboards, gestures captured by cameras, mice, touchpads, etc. Client device 210 may be any type of computing device, having at least a memory and processor module, and is configured to generate rendered images executed by game title processing engine 211, and to display the rendered images on a display (e.g., display 11 or display 11 including a head-mounted display – HMD, etc.). For example, the rendered images may be associated with an instance of a game running locally on client device 210 to enable gameplay for the corresponding user, such as through input commands used to drive the game. Some examples of client device 210 include personal computers (PCs), game consoles, home theater systems, general-purpose computers, mobile computing devices, tablets, telephones, or any other type of computing device capable of executing game instances.
[0066] Figure 2C Various network configurations that benefit from proper synchronization and offset of the VSYNC signal between the source and target devices according to embodiments of this disclosure are shown, including Figures 2A to 2B The configurations shown are examples of those described above. In particular, various network configurations benefit from proper alignment of the frequencies of the server and client VSYNC signals, and timing offsets of the server and client VSYNC signals, to reduce one-way latency and / or latency variability between the server and client. For example, one network device configuration includes a cloud gaming server (e.g., source) to client (target) configuration. In one embodiment, the client may include a WebRTC client configured to provide audio and video communication within a web browser. Another network configuration includes a client (e.g., source) to server (target) configuration. Yet another network configuration includes a server (e.g., source) to server (e.g., target) configuration. Another network device configuration includes a client (e.g., source) to client (target) configuration, wherein the clients may each be game consoles to provide, for example, head-mounted games.
[0067] Specifically, VSYNC signal alignment may include synchronizing the frequencies of the server VSYNC signal and the client VSYNC signal, and may also include adjusting the timing offset between the client VSYNC signal and the server VSYNC signal to remove drift and / or maintain an ideal relationship between the server VSYNC signal and the client VSYNC signal to reduce one-way latency and / or latency variability. In one embodiment, to achieve proper alignment, the server VSYNC signal may be tuned to achieve proper alignment between the server 260 and client 210 pair. In another embodiment, the client VSYNC signal may be tuned to achieve proper alignment between the server 260 and client 210 pair. Once the client VSYNC signal and the server VSYNC signal are aligned, the server VSYNC signal and the client VSYNC signal occur at substantially the same frequency and are offset from each other by a timing offset that can be adjusted from time to time. In another embodiment, VSYNC signal alignment may include synchronizing the frequencies of the VSYNC signals of two clients, and may also include adjusting the timing offset between their VSYNC signals to remove drift and / or achieve optimal reception timing for controller and other information; either VSYNC signal may be tuned to achieve this alignment. In another embodiment, alignment may include synchronizing the frequencies of the VSYNC signals of multiple servers, and may also include synchronizing the frequencies of the server VSYNC signals and client VSYNC signals, and adjusting the timing offset between the client VSYNC signals and server VSYNC signals, for example, for cloud gaming with a header. In server-to-client and client-to-client configurations, alignment may include synchronizing the frequencies of the server VSYNC signals and client VSYNC signals, and providing the correct timing offset between the server VSYNC signals and client VSYNC signals. In a server-to-server configuration, alignment may include synchronizing the frequencies of the server VSYNC signals and client VSYNC signals without setting a timing offset.
[0068] Figure 2D A multi-tenant configuration is illustrated between a cloud gaming server 260 and one or more clients 210 according to one embodiment of the present disclosure, benefiting from proper synchronization and offset of VSYNC signals between the source and target devices. In a server-to-client configuration, alignment may include frequency synchronization between the server's VSYNC signal and the client's VSYNC signal, as well as providing a correct timing offset between the server's VSYNC signal and the client's VSYNC signal. In a multi-tenant configuration, in one embodiment, the client's VSYNC signal is tuned at each client 210 to achieve proper alignment between the server 260 and client 210 pairs.
[0069] For example, in one embodiment, a graphics subsystem may be configured to perform multi-tenant GPU functionality, whereby the graphics subsystem may implement graphics and / or rendering pipelines for multiple games. That is, the graphics subsystem is shared among multiple games being executed. Specifically, in one embodiment, a game title processing engine may include a group of CPUs and GPUs that may be configured to perform multi-tenant GPU functionality, whereby the CPU and GPU group may implement graphics and / or rendering pipelines for multiple games. That is, the CPU and GPU group is shared among multiple games being executed. The CPU and GPU group may be configured as one or more processing devices. In another embodiment, multiple GPU devices are combined to perform graphics processing for a single application running on a corresponding CPU.
[0070] Figure 3This illustrates the general process of executing a video game at a server to generate game-rendered video frames and sending these frames to a client for display. Traditionally, several operations at game server 260 and client 210 are performed within frame periods, as defined by corresponding VSYNC signals. For example, server 260 attempts to generate game-rendered video frames at 301 within one or more frame periods, as defined by corresponding server VSYNC signal 311. Video frames are generated by the game in response to control information delivered from an input device at operation 350 (e.g., user input commands) or by game logic not driven by control information. Transmission jitter 351 may exist when sending control information to server 260, where jitter 351 measures the variation in network latency from client to server (e.g., when sending input commands). As shown, the thick arrows indicate the current latency when sending control information to server 260, but due to jitter, there may be a range of arrival times for control information at server 260 (e.g., the range defined by the dashed arrows). At flip time 309, the GPU triggers a flip command indicating that the corresponding video frame has been fully generated and placed in the frame buffer at server 260. Thereafter, server 260 performs a scan output / scan input (operation 302), where, for this video frame, the scan output is aligned with VSYNC signal 311 for subsequent frame periods defined by server VSYNC signal 311 (VBI omitted for clarity). Subsequently, the video frame is encoded (operation 303) (e.g., encoding begins after VSYNC signal 311 appears, and the end of encoding may not be aligned with VSYNC signal 311), and transmitted (operation 304, where transmission may not be aligned with VSYNC signal 311) to client 210. At client 210, the encoded video frames are received (operation 305, where reception may not be aligned with client VSYNC signal 312), decoded (operation 306, where decoding may not be aligned with client VSYNC signal 312), buffered, and displayed (operation 307, where the start of display may be aligned with client VSYNC signal 312). Specifically, client 210 displays each video frame rendered for display, starting with the corresponding occurrence of client VSYNC signal 312.
[0071] One-way latency 315 can be defined as the time from the start of transmitting a video frame to the encoding unit at the server (e.g., scan output 302) to the start of displaying the video frame 307 at the client. That is, one-way latency is the time from the server's scan output to the client's display, taking into account client buffering. Each frame has a latency from the start of scan output 302 to the completion of decoding 306. This latency may vary from frame to frame due to server operations such as encoding 303 and transmission 304, network transmission between server 260 and client 210 accompanied by jitter 352, and altitude variations at client reception 305. As shown in the figure, the straight, thick arrows indicate the current latency when the corresponding video frame is sent to client 210, but due to jitter 352, there may be a range of arrival times for video frames at client 210 (e.g., a range defined by the dashed arrows). Because one-way latency must be relatively stable (e.g., fairly consistent) to achieve a good gaming experience, the traditional result of implementing buffer 320 is that the display of individual frames with low latency (e.g., from the start of scan output 302 to the completion of decoding 306) is delayed by several frame cycles. That is, if there is network instability or unpredictable encoding / decoding time, additional buffering is needed to keep the one-way latency consistent.
[0072] According to one embodiment of this disclosure, when streaming video frames generated from a video game executed on a server, the one-way latency between the cloud gaming server and the client may vary due to clock drift. Specifically, the frequency difference between the server's VSYNC signal 311 and the client's VSYNC signal 312 may cause the client's VSYNC signal to drift relative to frames arriving from the server 260. This drift may be due to minute differences in the crystal oscillators used in each of the corresponding clocks at the server and client. Furthermore, embodiments of this disclosure reduce the one-way latency by: performing one or more of synchronization and offset of the VSYNC signals for alignment between the server and client; providing dynamic buffering on the client; overlapping the encoding and transmission of video frames at the server; overlapping the reception and decoding of video frames at the client; and overlapping the decoding and display of video frames at the client.
[0073] Figure 4 The illustration shows a data stream via a network configuration including a highly optimized cloud gaming server 260 and a highly optimized client 210 when streaming video frames generated from a video game executed on a server, according to an embodiment of the present disclosure. Overlapping server and client operations reduces one-way latency, and synchronizing and offsetting VSYNC signals between the server and client further reduces one-way latency and the variability of one-way latency between the server and client. Specifically, Figure 4 The desired alignment between the server VSYNC signal and the client VSYNC signal is illustrated. In one embodiment, tuning of the server VSYNC signal 311 is performed to achieve proper alignment between the server and client VSYNC signals, such as in a server-client network configuration. In another embodiment, tuning of the client VSYNC signal 312 is performed to achieve proper alignment between the server and client VSYNC signals, such as in a multi-tenant server-to-multi-client network configuration. For illustrative purposes, Figure 4 The document describes the tuning of the server VSYNC signal 311 to synchronize the frequencies of the server VSYNC signal and the client VSYNC signal, and / or adjust the timing offset between the corresponding client VSYNC signal and the server VSYNC signal. However, it is understood that the client VSYNC signal 312 can also be used for tuning. In the context of this patent, "synchronization" should be understood as tuning the signals so that their frequencies match, but their phases may differ; "offset" should be understood as representing the time delay between signals, such as the time between when one signal reaches its maximum value and when another signal reaches its maximum value.
[0074] As shown in the figure, in the implementation scheme of this disclosure, Figure 4 An improved process is illustrated for executing a video game at a server to generate rendered video frames and sending these video frames to a client for display. This process is shown in contrast to generating and displaying a single video frame at both the server and client. Specifically, the server generates the game-rendered video frames at 401. For example, server 260 includes a CPU (e.g., game title processing engine 211) configured to execute the game. The CPU generates one or more draw calls for the video frames, wherein the draw calls include commands placed in a command buffer for execution in the graphics pipeline by the corresponding GPU of server 260. The graphics pipeline may include one or more shader programs that manipulate the vertices of objects within the scene to generate texture values as rendered for the video frames used for display, wherein the operations are performed in parallel by the GPU for efficiency. At flip time 409, the GPU touches a flip command in the command buffer indicating that the corresponding video frame has been fully generated and / or rendered and placed in the frame buffer at server 260.
[0075] At 402, the server executes scan output of the game-rendered video frames to the encoder. Specifically, the scan output is executed scan-line by scan or in groups of consecutive scan lines, where a scan line refers to a single horizontal line, such as from one edge of the display screen to the other. These scan lines or groups of consecutive scan lines are sometimes referred to as slices, and are referred to as screen slices in this specification. Specifically, scan output 402 may include several processes that modify the game-rendered frames, including overlaying them with another frame buffer or shrinking them to surround them with information from another frame buffer. During scan output 402, the modified video frames are then scanned into the encoder for compression. In one embodiment, scan output 402 is executed at the occurrence 311a of the VSYNC signal 311. In other embodiments, scan output 402 may be executed before the occurrence of the VSYNC signal 311 (e.g., at toggle time 409).
[0076] At 403, the game-rendered video frame (which may have been modified) is encoded slice by slice at the encoder to generate one or more encoded slices, wherein the encoded slices are independent of scan lines or screen slices. Thus, the encoder generates one or more encoded (e.g., compressed) slices. In one embodiment, the encoding process begins before the scan output process 402 of the corresponding video frame has been fully completed. Furthermore, the start and / or end of encoding 403 may or may not be aligned with the server VSYNC signal 311. The boundaries of the encoded slices are not limited to a single scan line and may include a single scan line or multiple scan lines. Additionally, the end of the encoded slice and / or the start of the next encoder slice may not necessarily occur at the edge of the display (e.g., it may occur somewhere in the middle of the screen or in the middle of a scan line) so that the encoded slices do not have to completely traverse the edge-to-edge of the display. As shown, one or more encoded slices may be compressed and / or encoded, including a compressed “encoded slice A” with a hash tag.
[0077] At 404, the encoded video frame is transmitted from the server to the client, whereby the transmission may occur on a per-encoded slice basis, where each encoded slice is a compressed encoder slice. In one embodiment, transmission 404 begins before the encoding process 403 of the corresponding video frame has been fully completed. Furthermore, the start and / or end of transmission 404 may or may not be aligned with the server's VSYNC signal 311. As shown, compressed encoded slice A is transmitted to the client independently of other compressed encoder slices of the rendered video frame. Encoder slices may be transmitted one at a time or in parallel.
[0078] At 405, the client again receives compressed video frames on a per-wave encoded slice basis. Furthermore, the start and / or end of reception 405 may or may not be aligned with the client's VSYNC signal 312. As shown, the client receives compressed encoded slice A. Transmission jitter 452 may exist between server 260 and client 210, where jitter 452 measures the variation in network latency from server 260 to client 210. Lower jitter values indicate a more stable connection. As shown, the straight, thick arrows indicate the current latency when the corresponding video frame is sent to client 210, but due to jitter, there may be a range of arrival times for video frames at client 210 (e.g., the range defined by the dashed arrows). Variations in latency may also be due to one or more operations at the server, such as encoding 403 and transmission 404, and network problems that introduce latency when transmitting video frames to client 210.
[0079] At 406, the client again decodes the compressed video frame on a per-scan encoded slice basis, resulting in a decoded slice A (shown without hash markers) now ready for display. In one embodiment, the decoding process 406 begins before the receiving process 405 of the corresponding video frame is fully completed. Furthermore, the start and / or end of decoding 406 may or may not be aligned with the client's VSYNC signal 312. At 407, the client displays the decoded rendered video frame on its display. That is, for example, the decoded video frame is placed in a display buffer that is streamed to the display device on a per-scan-line basis. In one embodiment, the display process 407 (i.e., the streaming output to the display device) begins after the decoding process 406 of the corresponding video frame has been fully completed, i.e., the decoded video frame is fully residing in the display buffer. In another embodiment, the display process 407 begins before the decoding process 406 of the corresponding video frame has been fully completed. That is, the streaming output to the display device begins at the address of the display buffer, at which point only a portion of the decoded frame buffer resides in the display buffer. The display buffer is then updated or populated with the remainder of the corresponding video frame in a timely manner for display, ensuring that the update of the display buffer is performed before these portions are streamed to the display. Furthermore, the start and / or end of display 407 are aligned with the client's VSYNC signal 312.
[0080] In one embodiment, the one-way latency 416 between server 260 and client 210 can be defined as the elapsed time between the start of scan output 402 and the start of display 407. Embodiments of this disclosure can align the VSYNC signals between the server and client (e.g., synchronize frequencies and adjust offsets) to reduce the one-way latency between the server and client and reduce the variability of the one-way latency between them. For example, embodiments of this disclosure can calculate the optimal adjustment of the offset 430 between the server VSYNC signal 311 and the client VSYNC signal 312, such that even under near-worst-case times for server processing (e.g., encoding 403 and transmission 404), near-worst-case network latency between server 260 and client 210, and near-worst-case client processing (e.g., receiving 405 and decoding 406), the decoded rendered video frames are available in time for display process 407. That is, it is not necessary to determine the absolute offset between the server VSYNC and the client VSYNC; adjusting the offset to ensure that the decoded rendered video frames are available in time for display is sufficient.
[0081] Specifically, the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 can be aligned through synchronization. Synchronization is achieved by tuning either the server VSYNC signal 311 or the client VSYNC signal 312. For illustrative purposes, tuning is described with respect to the server VSYNC signal 311, but it is understood that tuning can alternatively be performed on the client VSYNC signal 312. For example, as... Figure 4 As shown, the server frame period 410 (e.g., the time between two occurrences 311c and 311d of the server VSYNC signal 311) is substantially equal to the client frame period 415 (e.g., the time between two occurrences 312a and 312b of the client VSYNC signal 312), which indicates that the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 are also substantially equal.
[0082] To maintain frequency synchronization between the server and client VSYNC signals, the timing of the server VSYNC signal 311 can be manipulated. For example, the vertical blanking interval (VBI) in the server VSYNC signal 311 can be increased or decreased over a period of time, such as to account for drift between the server VSYNC signal 311 and the client VSYNC signal 312. Manipulating the vertical blanking (VBLANK) lines in the VBI allows adjustment of the number of scan lines used for VBLANK for one or more frame periods of the server VSYNC signal 311. Decreasing the number of VBLANK scan lines reduces the corresponding frame period (e.g., time interval) between two occurrences of the server VSYNC signal 311. Conversely, increasing the number of VBLANK scan lines increases the corresponding frame period (e.g., time interval) between two occurrences of the VSYNC signal 311. In this way, the frequency of the server VSYNC signal 311 is adjusted so that the client VSYNC signal 312 is frequency-aligned with the server VSYNC signal 311 to substantially the same frequency. Furthermore, the offset between the server's VSYNC signal and the client's VSYNC signal can be adjusted by increasing or decreasing the VBI for a short period of time, and then the VBI can be returned to its original value. In one embodiment, the server's VBI is adjusted. In another embodiment, the client's VBI is adjusted. In yet another embodiment, instead of two devices (server and client), there are multiple connected devices, each of which may have a corresponding VBI that has been adjusted. In one embodiment, each of the multiple connected devices may be an independent peer device (e.g., no server device). In another embodiment, the multiple devices may include one or more server devices and / or one or more client devices, arranged in one or more server / client architectures, multi-tenant server / client architectures, or some combination thereof.
[0083] Alternatively, in one implementation, the server's pixel clock (e.g., the southbridge located in the server's northbridge / southbridge core logic chipset, or, in the case of a discrete GPU, generating its own pixel clock using its own hardware) can be manipulated to perform coarse and / or fine adjustments to the frequency of the server VSYNC signal 311 over a period of time, so that the frequency synchronization between the server VSYNC signal 311 and the client VSYNC signal 312 returns to alignment. Specifically, the pixel clock in the server's southbridge can be overclocked or underclocked to adjust the overall frequency of the server's VSYNC signal 311. In this way, the frequency of the server VSYNC signal 311 is adjusted so that the frequency alignment between the client VSYNC signal 312 and the server VSYNC signal 311 is substantially the same. The offset between the server VSYNC and the client VSYNC can be adjusted by increasing or decreasing the client-server pixel clock for a short period of time and then returning the pixel clock to its original value. In one implementation, the server pixel clock is adjusted. In another implementation, the client pixel clock is adjusted. In another embodiment, instead of two devices (server and client), there are multiple connected devices, each of which may have an adjusted corresponding pixel clock. In one embodiment, each of the multiple connected devices may be an independent peer device (e.g., no server device). In another embodiment, the multiple connected devices may include one or more server devices and one or more client devices, arranged in one or more server / client architectures, multi-tenant server / client architectures, or some combination thereof.
[0084] Figure 5A-1 A chipset 540 according to one embodiment of the present disclosure is shown, which is configured to perform a high-speed scan output operation to deliver to an encoder when content from a video game executed at a cloud gaming server is streamed across a network. Additionally, the chipset 540 can be configured to perform the scan output operation earlier, such as before the next system VSYNC signal or at the flip time of the corresponding video frame. Specifically, Figure 5A-1 This illustrates how, in one implementation, the speed of the scan output block 550 is determined for the target display of the client.
[0085] Chipset 540 is configured to operate at a maximum pixel clock 515. The pixel clock defines the rate at which the chipset can process pixels, such as by scanning output block 550. The pixel clock rate is typically expressed in megahertz values, representing the number of pixels that can be processed. Specifically, pixel clock calculator 510 is configured to determine the maximum pixel clock 515 based on chip computing settings 501 and / or self-diagnostic tests 505. For example, chipset 540 may be designed to have a specific maximum pixel clock included in chip computing settings 501. However, once built, chipset 540 may be able to operate at a higher pixel clock, or may not actually operate at the designed pixel clock determined according to chip computing settings 501. Accordingly, test 505 may be performed to determine the self-diagnostic pixel clock 505. Pixel clock calculator 510 may be configured to define the maximum pixel clock 515 of chipset 540 based on the higher of the designed pixel clock determined from chip computing settings 501 or self-diagnostic pixel clock 505. For illustrative purposes, an exemplary maximum pixel clock might be 300 megapixels per second (Mpps).
[0086] The scan output block 550 operates at a speed corresponding to the target display of the client 210. Specifically, the frame rate calculator 520 determines a frame rate setting 525 based on various inputs, including the maximum pixel clock 515 of the chipset 540 and the requested image size 521. Information in the requested image size 521 can be taken from values 522, including common display values (e.g., 480p, 720p, 1080p, 4K, 8K, etc.) as well as other defined values. For the same maximum pixel clock, different frame rate settings may exist depending on the client's target display, where the frame rate setting is determined by dividing the maximum pixel clock 515 by the number of pixels on the target display. For example, at a maximum pixel clock of 300 megapixels per second, the frame rate setting for a 480p display (e.g., approximately 300k pixels, such as those used in mobile phones) is approximately 1000Hz. And, at the same maximum pixel clock of 300 megapixels per second, the frame rate setting for a 1080p display (approximately 2 megapixels) is approximately 150Hz. Furthermore, to illustrate this point, at a maximum pixel clock speed of 300 megapixels per second, the frame rate setting for a 4K display is approximately 38Hz.
[0087] A frame rate setting 525 is input to a scan output setting converter 530, which is configured to determine a speed setting value 535 formatted for chipset 540. For example, chipset 540 may operate at bitrate. In one embodiment, the speed setting value 535 may be a frame rate setting 525 (e.g., frames per second). In some embodiments, the speed setting value 535 may be determined as a multiple of the base frame rate. For example, the speed setting value 535 may be set to a multiple of 30 frames per second (e.g., 30Hz), such as 30Hz, 60Hz, 90Hz, 120Hz, 150Hz, etc. The speed setting value 535 is input to a cache 545 of chipset 540 for access by the corresponding scan output block 550 in the chipset, thereby determining its operating speed for the target display of client 210.
[0088] Chipset 540 includes a game title processing engine 211 configured to execute video game logic 215 of the video game to generate game-rendered video frames to be streamed back to client 210. As shown, the game title processing engine 211 includes a CPU 501 and a GPU 502 (e.g., configured to implement a graphics pipeline). In one embodiment, the CPU 501 and GPU 502 are configured as accelerated processing units (APUs) that are integrated onto the same chip or die using the same bus for faster communication and processing. In another embodiment, the CPU 501 and GPU 502 can be connected via a bus (e.g., PCI-Express, Gen-Z, etc.). Multiple game-rendered video frames are generated for the video game and placed in buffers 555 (e.g., display buffers or frame buffers), which include one or more game buffers, such as game buffer 0 and game buffer 1. Game buffer 0 and game buffer 1 are driven by a toggle control signal to determine which game buffer will store which video frame output from the game title processing engine 211. The game title processing engine operates at a specific speed defined by the video game. For example, the game title processing engine 211 may output video frames at 30Hz or 60Hz, etc.
[0089] Optionally, additional information can be generated to be included in the video frames rendered by the game. Specifically, feature generation block 560 includes one or more feature generation units, each configured to generate features. Each feature generation unit includes a feature processing engine and a buffer. For example, feature generation unit 560-A includes feature processing engine 503. In one implementation, feature processing engine 503 executes on the CPU 501 and GPU 502 of the game title processing engine 211 (e.g., on other threads). Feature processing engine 503 can be configured to generate multiple user interface (UX) features, such as user interface, messaging, etc. In one implementation, UX features can be rendered as an overlay. Multiple UX features generated for the video game are placed in buffers (e.g., display buffers or frame buffers), including one or more UX buffers, such as UX buffer 0 and UX buffer 1. UX buffer 0 and UX buffer 1 are driven by corresponding toggle control signals to determine which UX buffer will store which feature output from feature processing engine 503. Furthermore, feature processing engine 503 operates at a specific speed that may be defined by the video game. For example, the feature processing engine 503 can output video frames at 30Hz or 60Hz, etc. The feature processing engine 503 can also operate independently of the speed at which the game title processing engine 211 can output video frames (i.e., at a rate different from 30Hz or 60Hz, etc.).
[0090] The game-rendered video frames scanned from buffer 555 and optional features scanned from the buffer of the feature generation unit (e.g., unit 560-A) are scanned to the scan output block 550 at a rate X. In one implementation, the rate X used to scan the game buffer 555 that holds the game-rendered video frames and / or the UX buffer that holds the features may not correspond to a speed setting value 535, in order to scan the output information from the buffer as quickly as possible. In another implementation, the rate X does correspond to the speed setting value 535.
[0091] As previously described, scan output block 550 operates at a speed corresponding to the target display of client 210. In cases where multiple clients may have multiple target displays (e.g., mobile phones, television displays, computer monitors, etc.), multiple scan output blocks may exist, each supporting a corresponding display, and each operating at a different speed setting value. For example, scan output block A (550-A) receives video frames rendered by the game from buffer 555 and a feature overlay from feature generation block 560. Scan output block A (550-A) operates via a corresponding speed setting value (such as a corresponding frame rate setting) in cache-A (545-A). Accordingly, for the target display, scan output block 550 outputs modified video frames to encoder 570 at a rate defined by the speed setting value (e.g., 120 Hz). That is, the rate at which modified video frames are output to encoder 570 is higher than the rate at which video frames are generated and / or encoded, where this rate is based on the maximum pixel clock of chipset 540, including scan output block 550, and the image size of the target display.
[0092] In one implementation, encoder 570 may be part of chipset 540. In other implementations, encoder 570 is separate from chipset 540. Encoder 570 is partially configured to compress modified video frames for streaming to client 210. For example, modified video frames are encoded on the encoder on a slice-by-slice basis to generate one or more encoded slices for the corresponding modified video frames. The one or more encoded slices of the corresponding modified video frames, including additional feature overlays, are then streamed over a network to a target display of client 210. The encoder outputs one or more encoded slices at a rate independent of the speed setting value and may be bound to synchronized and offset server and client VSYNC signals, as previously described. For example, one or more encoded slices may be output at 60Hz.
[0093] Figure 5A-2 A chipset 540B according to one embodiment of the present disclosure is shown. This chipset is configured to perform a high-speed scan output operation to deliver to an encoder when content from a video game executed at a cloud gaming server is streamed across a network, wherein optional user interface features may be integrally formed into the video frames rendered by the game. Additionally, the chipset 540 can be configured to perform the scan output operation earlier, such as before the next system VSYNC signal occurs or at the flip time of the corresponding video frame. Figure 5A-2 Some of the components shown in the image are related to Figure 5A-1 The components are similar, with similar features having similar functionality. This is shown in the corresponding chipsets. Figure 5A-2 and Figure 5A-1The differences between them. Specifically, in Figure 5A-2 and Figure 5A-1 between, Figure 5A-2 The difference in the configuration of chipset 540B is the absence of a separate feature generation block. Instead, one or more optional UX features can be generated by CPU 501 and / or GPU 502 and integrally formed into the game-rendered video frame placed in buffer 555, as previously described. That is, the features do not need to be provided as an overlay, as they are integrally formed into the rendered video frame. The game-rendered video frame can optionally be scanned from buffer 555 to scan output block 550, which includes one or more scan output blocks 550-B for one or more target displays for the client. As previously described, the corresponding scan output block 550-B operates at the speed of the target display. Accordingly, for the target display, the corresponding scan output block 550-B outputs the video frame to encoder 570 at a rate defined by a speed setting value. In some embodiments, because the features are integrally formed into the game-rendered video frame, thus requiring only buffer 555, the rendered video frame can be scanned directly into the encoder and bypass scan output block 550. In this case, for example, additional operations performed during the scan output can be executed by CPU 501 and / or GPU 502.
[0094] Figure 5B-1 The illustration shows a scan output operation according to one embodiment of the present disclosure, performed on video frames rendered by the game, optionally including one or more additional features (e.g., layers), to be delivered to an encoder when content from a video game executed at a cloud gaming server is streamed across a network to a client. For example, Figure 5B-1 It shows Figure 5A-1 The operation of the scan output block 550-A. The scan output block A (550-A) receives the video frames rendered by the game line by line.
[0095] Specifically, scan output block A (550-A) receives game-rendered video frames from buffer 555 and a feature overlay from feature generation block 560, which is provided to input buffer 580. As previously described, scan output block A (550-A) operates via corresponding speed settings (such as corresponding frame rate settings for the target display of client 210) in cache-A (545-A). For example, multiple game-rendered video frames are output from game buffer 0 and game buffer 1 to input frame buffer 580-A of scan output block A (550-A) under the control of a flip control signal.
[0096] Additionally, scan output block A (550-A) may optionally receive one or more UX features (e.g., as an overlay). For example, multiple UX features are output from buffer 560-A, which includes UX buffer 0 and UX buffer 1, under the control of corresponding flip control signals. Multiple UX features are scanned to the input frame buffer 580-B of scan output block A (550-A). Other feature overlays may be provided, wherein exemplary UX features may include user interface, system user interface, SMS, messaging, menus, communications, additional game viewpoints, esports information, etc. For example, additional multiple UX features may be output from buffers 560A-560N, each including UX buffer 0 and UX buffer 1, under the control of corresponding flip control signals. For illustration, multiple UX features are output from buffer 560-N to input frame buffer 580-N.
[0097] Information in input frame buffer 580 is output to combiner 585, which is configured to synthesize the information. For example, for each corresponding video frame generated by a video game, combiner 585 combines the game-rendered video frame from input frame buffer 580-A with each of the optional UX features provided in input frame buffers 580-B to 580-N.
[0098] The video frames rendered by the game, combined with one or more optional UX features, are then provided to block 590, where additional operations can be performed to generate modified video frames suitable for display. In block 590, the additional operations performed during the scan output process may include one or more operations such as decompressing DCC-compressed surfaces, scaling the resolution to the target display, color space conversion, degamma, HDR expansion, gamut remapping, LUT shaping, tone mapping, gamma blending, and blending.
[0099] In other implementations, the additional operations outlined in block 590 are performed at each location in the input frame buffer 580 to generate a corresponding layer of the modified video frame. For example, the input frame buffer may be used to store and / or generate video frames rendered for a video game, as well as one or more optional UX features (e.g., as overlays), such as a user interface (UI), system UI, text, messaging, etc. Additional operations may include decompressing DCC-compressed surfaces, resolution scaling, color space conversion, degamma, HDR expansion, gamut remapping, LUT shaping, tone mapping, gamma blending, etc. After these operations are performed, one or more layers of the input frame buffer 580 are composited and blended, optionally placed into a display buffer, and then scanned to the encoder (e.g., scanned from the display buffer).
[0100] Accordingly, for the target display, scan output block 550-A outputs multiple modified video frames to encoder 570 at a rate defined by a speed setting value (e.g., 120Hz). That is, the rate at which the modified video frames are output to encoder 570 is higher than the rate at which the video frames are generated and / or encoded, where this rate is based on the maximum pixel clock of chipset 540, including scan output block 550, and the image size of the target display. As previously described, encoder 570 compresses each of the modified video frames. For example, a corresponding modified video frame may be compressed into one or more encoded slices (compressed encoder slices), which may also be packaged for network streaming. The modified video frames, which have been compressed and / or packaged into encoded slices, are then stored in buffer 580 (e.g., a first-in-first-out or FIFO buffer). Streamer 575 is configured to transmit the encoded slices to client 210 via network 250. As previously mentioned, streaming devices can be configured to operate at the application layer of a Transmission Control Protocol / Internet Protocol (TCP / IP) computer network model. In implementations, assuming an IP-based network (e.g., home / Internet), either TCP / IP or UDP can be used. For example, cloud gaming services can use UDP. TCP / IP guarantees all data arrival; however, this "arrival guarantee" comes at the cost of retransmissions, introducing additional latency. On the other hand, UDP-based protocols offer optimal latency performance, but at the cost of packet loss, resulting in data loss.
[0101] Figure 5B-2 The illustration shows a scan output operation according to one embodiment of the present disclosure, performed on video frames rendered by the game, optionally including one or more additional features (e.g., layers), to be delivered to an encoder when content from a video game executed at a cloud gaming server is streamed across a network to a client. For example, Figure 5B-2 The operation of scan output block 550-A2 is shown. Scan output block A2 (550-A2) receives video frames rendered by the game line by line. Figure 5B-2 The configuration of the scan output block 550-A2 and Figure 5B-1 The scan output block 550-A is similar, with similar features having similar functionality. Figure 5B-2 The scan output block 550-A2 is different Figure 5B-1 The scan output block 550-A is used because there is no combiner 585. Therefore, the video frames rendered by the game and the UX feature overlay can be composited and blended on the client side.
[0102] As shown in the figure, information in each of the input frame buffers 580 is delivered to the corresponding block 590, where additional operations are performed. Specifically, the additional operations outlined in block 590 are performed for each of the input frame buffers 580 to generate the corresponding layer. These additional operations may include decompressing DCC-compressed surfaces, resolution scaling, color space conversion, gamma removal, HDR extension, gamut remapping, LUT shaping, tone mapping, gamma blending, etc. After these operations are performed, one or more modified layers are delivered individually to encoder 570. The encoder delivers each layer individually to a client, where the client can composite and blend the layers to generate modified video frames for display.
[0103] Figure 5B-3 The scan output operation according to one embodiment of the present disclosure is illustrated, wherein the scan output operation is performed on video frames rendered by the game to be delivered to an encoder when content from a video game executed at a cloud gaming server is streamed across a network to a client. For example, Figure 5B-3 It shows Figure 5A-2 The operation of scan output block 550B, where scan output block 550-B does not have combiner functionality. Some components of scan output block 550-B are related to... Figure 5B-1 The scan output block 550-A is similar, with similar features having similar functionality. Figure 5B-3 The scan output block 550-B is different. Figure 5B-1The scan output block 550-A has only a single input frame buffer because it lacks a combiner (e.g., for performing compositing and blending) and because there is no separate feature generation. Specifically, scan output block B (550-B) receives game-rendered video frames scan-line by scan from buffer 555. Optionally, user interface features may be integrally formed into the game-rendered video frames generated by the CPU and / or GPU. For example, multiple game-rendered video frames are output from game buffer 0 and game buffer 1 to the input frame buffer 580 of scan output block B (550-B) under the control of a flip control signal. The game-rendered video frames are then provided to block 590, where additional operations (e.g., decompressing DCC-compressed surfaces, scaling the resolution to the target display, color space conversion, etc.) may be performed to generate modified video frames suitable for display, as previously described. These additional operations may not necessarily require compositing and / or blending, as optional UX features are already integrally formed into the game-rendered video frames. In some implementations, the additional operations outlined in block 590 may be performed at input frame buffer 580. Accordingly, for the target display, scan output block 550-B outputs multiple modified video frames (e.g., at a rate defined by corresponding speed setting values) to encoder 570. As previously described, encoder 570 compresses each of the modified video frames, such as compressing them into one or more encoded slices (compressed encoder slices), which may also be packaged for network streaming. The modified video frames that have been compressed and / or packaged into encoded slices are then stored in buffer 580. As previously described, streamer 575 is configured to transmit the encoded slices to client 210 over the network.
[0104] Figures 5C-5D An exemplary server configuration according to an embodiment of this disclosure is shown, including a scan output block having one or more input frame buffers, which are used when performing a high-speed scan output operation to deliver to the encoder during the streaming of content from a video game executed at a cloud gaming server across a network. Specifically, Figures 5C-5D It shows the use of Figure 5A-1 The scan output block 550-A and / or Figure 5A-2 An exemplary configuration of the scan output block 550B includes one or more input frame buffers for generating composite video frames to be displayed on a high-definition display or a virtual reality (VR) display (e.g., a head-mounted display). In one implementation, the input frame buffers may be implemented in hardware.
[0105] Figure 5CA scan output block 550-A' is shown, comprising four input frame buffers that can be used to generate composite video frames for high-definition displays. By way of example only, three input frame buffers (e.g., FB0, FB1, and FB2) are dedicated to video games and can be used to store and / or generate corresponding layers including at least one of video frames, UI, esports UI, and text layers. The input frame buffers for video games can generate game-rendered video frames from one or more viewpoints in the game environment. Another input frame buffer, FB3, is dedicated to the system and can be used to generate system overlays (e.g., UI), such as friend notifications.
[0106] Figure 5D A scan output block 550-A'' is shown, comprising four input frame buffers that can be used to generate composite video frames for a VR display. By way of example only, two input frame buffers (e.g., FB0 and FB1) are dedicated to video games and can be used to store and / or generate corresponding layers including at least one of video frames taken from different viewpoints of the game environment, a UI, an esports UI, and a text layer. Two other input frame buffers (FB2 and FB3) are dedicated to the system and can be used to generate system overlay layers (e.g., UIs), such as those including friend notifications or esports UIs.
[0107] In embodiments of this disclosure, high-speed and / or early scan output / scan input can be performed at the server without regard to display requirements and / or parameters, since no physical display is connected to the server. Specifically, the server can perform scan output / scan input against a target virtual display, which can be defined by the user to operate at a selected frequency (e.g., 93Hz, 120Hz).
[0108] Through the Figures 2A-2D Detailed description of various client devices 210 and / or cloud gaming networks 290 (e.g., in game server 260), Figure 6 Flowchart 600 illustrates a method for cloud gaming according to one embodiment of the present disclosure, wherein high-speed and / or early scan output operations can be performed to reduce one-way waiting time between the cloud gaming server and the client.
[0109] At 610, the method includes generating video frames while the video game is being executed at the server. For example, the server may execute the video game in a streaming mode, such that the server's CPU executes the video game in partial response to input commands from the user or game logic driven by control information from the user, in order to generate game-rendered video frames using a graphics pipeline available for streaming. Specifically, the CPU and GPU graphics pipelines co-executing the video game are configured to generate multiple video frames. In cloud gaming, game-generated video frames are typically rendered and displayed on a virtual display. The server may perform additional operations on the game-generated video frames during the scan output process. For example, one or more overlays may be added to the corresponding game-generated video frames, such as during the scan output process.
[0110] At 620, the method includes performing a scan output process by scanning multiple screen slices of a video frame scan line-by-scan onto one or more input frame buffers to perform one or more operations modifying the multiple screen slices. As previously described, UX features (e.g., overlays) may be scanned onto one or more input frame buffers. Accordingly, the one or more input frame buffers may be used to store and / or generate video frames rendered for a video game, as well as one or more optional UX features (e.g., as overlays), such as user interface (UI), system UI, text, messaging, etc. The scan output process generates modified video frames, which are synthesized and blended to include one or more optional UX features, such as those implemented through overlays. In one implementation, UX features (e.g., as overlays) are first synthesized, and then additional operations are performed, as previously described. For example, additional operations may include decompressing DCC-compressed surfaces, resolution scaling, color space conversion, gamma demapping, HDR expansion, gamut remapping, LUT shaping, tone mapping, gamma blending, etc. In another implementation, additional operations are performed on each of the UX features prior to composition and blending, as described above.
[0111] At 630, after the modified video frame is generated, multiple screen slices of the modified video frame are scanned to the encoder line-by-line during the output scan. Thus, the modified game-generated video frame (e.g., modified with an optional UX feature layer) is scanned into the encoder for compression in preparation for streaming the modified video frame to a client, such as when streaming content from a video game running on a cloud gaming server to a client across a network.
[0112] Specifically, at 640, the method includes initiating the output scanning process early. In one implementation, multiple screen slices of the game-generated video frame are scanned into one or more input frame buffers at the corresponding flip time of the video frame. That is, instead of waiting for the next server VSYNC signal to occur before starting the output scanning process, the modified video frame is scanned into the corresponding input frame buffer earlier (i.e., before the next server VSYNC signal). The flip time can be included in a command in the command buffer that, when the GPU executes in the graphics pipeline, indicates that the GPU has completed executing multiple commands in the command buffer and that the game-rendered video frame has been fully loaded into the server's display buffer. The game-rendered video frame is then scanned into the corresponding input frame buffer during the output scanning process. Additionally, one or more optional UX features (e.g., overlays) are also scanned into one or more input frame buffers at the corresponding flip time generated for the UX feature.
[0113] In another embodiment, according to one embodiment of this disclosure, a scan output process is performed at high speed when streaming content from a video game executed at a cloud gaming server across a network. For example, the scan output process operates at a speed / rate corresponding to the target display of the client and is based on the server's maximum pixel clock and the requested image size of the target display, as previously described. For example, the scan output process includes receiving video frames rendered by the game and feature overlays, and then compositing them, wherein additional operations such as scaling, color scaling, blending, etc., can be performed on the composited video frames. As previously described, the scan output process outputs modified video frames at a scan output rate based on a speed setting value (e.g., 120Hz), where the speed setting value is based on the server's maximum pixel clock and the requested image size of the target display. In one implementation, the speed setting value is the frame rate. Accordingly, the scan output rate at which the modified video frames are output to the encoder may be higher than the rate at which the video frames are generated and / or encoded.
[0114] Each modified video frame can be segmented into one or more encoder slices, and then said one or more encoder slices are compressed into one or more encoded slices. Specifically, the encoder receives the modified video frame and encodes it on the encoder on a slice-by-slice basis to generate one or more encoded slices. As previously mentioned, the boundaries of the encoded slices are not limited to a single scan line and can include a single scan line or multiple scan lines. Additionally, the end of an encoded slice and / or the beginning of the next encoded slice may not necessarily occur at the edge of the display (e.g., it may occur somewhere in the middle of the screen or in the middle of a scan line). In one embodiment, because the server VSYNC signal and the client VSYNC signal are synchronized and offset, the operation at the encoder may overlap. Specifically, the encoder is configured to generate a first encoded slice of the modified video frame, wherein the modified video frame may include multiple encoded slices. The encoder can be configured to begin compressing the first encoded slice before the modified video frame is fully received. That is, the first encoded slice may be encoded (e.g., compressed) before multiple screen slices of the modified video frame are fully received, wherein said screen slices are delivered scan-line by scan. In some implementations, depending on the number of processors or hardware, multiple slices can be encoded simultaneously (e.g., in parallel). For example, some game consoles can generate four encoded slices in parallel. More specifically, due to hardware pipelines, hardware encoders can be configured to compress multiple encoder slices in parallel (e.g., to generate one or more encoded slices).
[0115] Figure 7AA process for generating and transmitting modified video frames at a cloud gaming server, according to one embodiment of this disclosure, is illustrated, wherein the process is optimized to perform high-speed and / or early scan output to the encoder to reduce one-way latency between the cloud gaming server and the client. This process is illustrated relative to the generation and transmission of a single modified video frame modified at the server with additional UX features (e.g., overlays). The operation at the server includes generating a game-rendered video frame 490 in operation 401. The scan output process 402 includes delivering the game-rendered video frame 490 to one or more input frame buffers of a scan output block to generate a composited overlay. That is, the game-rendered video frame 490 is composited with optional UX features (e.g., overlays). Additional operations (e.g., blending, resolution scaling, color space conversion, etc.) are performed on the composited video frame to generate a modified video frame (e.g., a game-rendered video frame modified with an additional UX feature overlay). During the scan output process, the modified video frame is scanned to the encoder. At operation 403, the modified video frame is encoded (e.g., compressed) into an encoded video frame on an encoder on a slice-by-slice basis. At operation 404, the compressed encoded video frame is transmitted from the server to the client.
[0116] As previously described, the scan output process 402 is performed early, prior to the occurrence of the server VSYNC signal 311. Typically, scan output begins the next time the server VSYNC signal occurs. In one embodiment, early scan output is performed at a flip time 701, where the flip time occurs after the GPU has finished generating the rendered frame 490, as previously described.
[0117] By performing an early scan output procedure, the one-way latency between the server and client can be reduced because remaining server operations (e.g., encoding, transmission, etc.) can also begin earlier and / or overlap. Specifically, an additional time 725 is gained by performing early scan output, where the additional time is defined between the flip time 701 and the next occurrence of the server VSYNC signal. This additional time 725 can offset any adverse latency variations experienced during other operations (such as encoding 403 or transmission 404). For example, if encoding 403 takes longer than the frame period, the additional time gained when encoding 403 begins early (e.g., not synchronized to begin at the VSYNC signal) may be sufficient to allow the video frame to be encoded before the next server VSYNC signal. Similarly, the additional time gained by performing early scan output can be used to reduce any variations in latency when delivering video frames to the client (e.g., increased delivery time over the network).
[0118] Figure 7BThe timing of a scan output process performed at a cloud gaming server according to one embodiment of the present disclosure is illustrated, wherein the scan output is performed at high speed and / or early so that video frames can be scanned to the encoder earlier at the end of the scan output process, thereby reducing the one-way wait time between the cloud gaming server and the client. Typically, an application running on the server (e.g., a video game) requests a "flip" of the display buffer when rendering is complete. The flip occurs during the execution of a flip command at flip time 701 during frame period 410, wherein the flip command is executed by the graphics processing unit (GPU). The flip command is one of a plurality of commands placed in a command buffer by the central processing unit (CPU) while executing the application, wherein the commands in the command buffer are used by the GPU to render the corresponding video frame. Accordingly, the flip indicates that the GPU has completed the execution of the commands in the command buffer to generate the rendered video frame, and that the rendered video frame has been fully loaded into the server's display buffer. There is a wait period 725, after which the scan output process 402a is executed upon the subsequent occurrence of the server's VSYNC signal 311f. In other words, in a typical process, scan output 402a is performed after a waiting period of 725, during which modified video frames (e.g., game-rendered video frames composited and blended with optional UX feature overlays) in the display buffer are scanned to the encoder for video encoding. That is, the scan output process typically occurs at the next VSYNC signal and after the waiting period, even if the display buffer was full earlier.
[0119] This disclosure provides an embodiment of early scan output 402b from the display buffer to the encoder, such as in cloud gaming applications. Figure 7B As shown, the scan output process 402b is triggered earlier at flip time 701, rather than at the next occurrence of the server VSYNC signal 311f. This allows the encoder to begin encoding earlier during operational overlap, instead of waiting for the next server VSYNC signal to perform the scan output for delivery to the encoder for encoding / compression. Display time is unaffected because no display is actually attached to the server. As previously mentioned, early encoding reduces the one-way latency between the server and the client because the chance of missing one or more VSYNCs when processing complex video frames is smaller; one or more VSYNCs are intended for delivery to the client and / or for display at the client.
[0120] Figure 7CThe illustration shows a time period for high-speed scan output according to one embodiment of the present disclosure, allowing video frames to be scanned to the encoder earlier, thereby reducing one-way latency between the cloud gaming server and the client. Specifically, the scan output process can be performed at high speed when streaming content from a video game executed at the cloud gaming server across the network, wherein the scan output process operates at a speed / rate corresponding to the target display of the client and is based on the server's maximum pixel clock and the requested image size of the target display, as previously described. Accordingly, the scan output rate for outputting modified video frames to the encoder may be higher than the rate at which video frames are generated and / or encoded. That is, the scan output rate may not correspond to the rate at which the video game generates video frames. For example, the scan output rate (e.g., frame rate setting) may be higher than the frequency of the server VSYNC signal used to generate video frames when the video game is executed at the server.
[0121] In another implementation, the scan output speed may not correspond to the refresh rate of the client's display device (e.g., 60Hz, etc.). That is, the display rate of the display device at the client and the scan output speed may not be the same. For example, the display rate of the display device at the client may be 60Hz, or a variable refresh rate, etc., where the scan output speed is a different rate (e.g., 120Hz, etc.).
[0122] Typically, the scan output process of a video frame is performed over the entire frame period (e.g., 16.6 ms at 60 Hz). For example, a representative frame period 410 is shown between two server VSYNC signals 311c and 311d. In embodiments of this disclosure, instead of performing the scan output process over the entire frame period, the scan output is performed at a higher rate. By performing the scan output process (e.g., including scan-to-encoder) at a rate higher than the frame processing rate (e.g., 60 Hz), such as waiting for the scan output process 402 to finish before starting encoding 403, or when overlapping scan output 402 and encoding 403. For example, the scan output process 402 may be performed over a period 730 (e.g., approximately 8 ms) that is less than the full frame period 410 (e.g., 16.6 ms at 60 Hz).
[0123] In some cases, encoding can begin earlier, such as before the next server VSYNC signal. Specifically, the encoder can begin processing as soon as the minimum amount of data from the corresponding modified video frame (e.g., a game-rendered video frame modified with one or more optional UX features as an overlay) is delivered to the encoder (e.g., 16 or 64 scan lines), and then process any additional data as soon as it arrives. This reduces one-way latency because the chance of missing one or more VSYNCs is smaller when processing complex video frames intended for delivery to the client and / or for display at the client. One-way latency can be due to network jitter and / or increased processing time at the server. For example, a modified video frame with a large amount of data (e.g., scene changes) may take more than one frame cycle to encode. The faster the scan output process, the more time is left for encoding, and the more likely a modified video frame with a large amount of data is to complete the encoding process before the server VSYNC signal intended for delivery to the client.
[0124] In another implementation, the encoding process can be further optimized to ensure the minimum amount of time used for encoding by limiting the encoding resolution to the resolution required by the client display, so that no time is wasted encoding video frames at a higher resolution than the client display can handle or requests at a given moment.
[0125] Through the Figures 2A-2D Detailed description of various client devices 210 and / or cloud gaming networks 290 (e.g., in game server 260), Figure 8A Flowchart 800A illustrates a method for cloud gaming according to one embodiment of the present disclosure, wherein video displayed on a client can be smoothed in a cloud gaming application, and high-speed and / or early scan output operations can be performed at the server to reduce one-way latency between the cloud gaming server and the client.
[0126] At 810, the method includes generating video frames when a video game is executed at a server. For example, a cloud gaming server may execute a video game in streaming mode, such that the CPU executes the video game in response to input commands from the user, thereby generating video frames for game rendering using a graphics pipeline.
[0127] The server can perform additional operations on the game-generated video frames during the scan output process. For example, one or more overlays can be added to the corresponding game-generated video frames, such as during the scan output process. Specifically, at 820, the method includes performing a scan output process to generate a modified video frame and delivering it to an encoder configured to compress the video frame. The scan output process includes: scanning the video frame and one or more user interface features scan line-by-scan into one or more input frame buffers, and synthesizing and blending the video frame and one or more user interface (UX) features (e.g., as overlays including user interface (UI), system UI, text, messaging, etc.) into a modified video frame, wherein the scan output process begins at the flip time of the video frame. Accordingly, the scan output process generates a modified video frame, which is synthesized and blended to include one or more optional UX features, such as those implemented through overlays.
[0128] At 830, the method includes transmitting the modified video frame to be compressed to the client. Specifically, each modified video frame may be segmented into one or more encoder slices, and then the encoder compresses the one or more encoder slices into one or more encoded slices. That is, the encoder receives the modified video frame and encodes the modified video frame on a slice-by-slice basis to generate one or more encoded slices, which are then packaged and delivered to the client over the network.
[0129] At 840, the method includes determining the target display time of the modified video frame at the client. Specifically, when the scan output of the server display buffer occurs at the flip time rather than the next occurrence of the server VSYNC signal, the ideal display timing on the client side can be performed based on the timing of the scan output at the server and the game's intent regarding a specific display buffer (e.g., the target display buffer VSYNC). The game intent determines whether the frame is for the next client VSYNC or actually for the previous client VSYNC, since the game runs later when processing that frame.
[0130] At 850, the method includes scheduling the display time of modified video frames at the client based on a target display time. The client-side strategy for selecting when to display frames may depend on whether the game is designed for a fixed or variable frame rate, and whether the VSYNC timing information is implicit or explicit, as will be discussed below. Figure 8B Further description.
[0131] Figure 8BA timing diagram of server and client operations according to one embodiment of this disclosure is shown, wherein the operation is performed while a video game is executed at server 260 to generate rendered video frames, which are then sent to client 210 for display. Because the client knows various timing parameters associated with each of the rendered video frames generated at the server, these timing parameters can be used to indicate and / or determine the ideal display time, allowing the client to decide when to display these video frames based on one or more strategies. Specifically, the ideal display time for a corresponding rendered video frame generated at the server indicates when the game application executing on the server intends to display the rendered video frame with reference to a target occurrence of the server VSYNC signal. This target server VSYNC signal can be converted into a target client VSYNC signal, particularly when the server VSYNC signal and the client VSYNC signal are synchronized (e.g., frequency and timing) and aligned using an appropriate offset.
[0132] Figure 8B The diagram illustrates the desired synchronization and alignment between the server VSYNC signal and the client VSYNC signal. Specifically, the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 are synchronized so that they have the same frequency and corresponding frame period. For example, the frame period 410 of the server VSYNC signal 311 is substantially equal to the frame period 415 of the client VSYNC signal 312. Additionally, the server and client VSYNC signals may be aligned with an offset 430. A timing offset can be determined such that a predetermined number (e.g., 99.99%) of the received video frames arrive at the client for display when the next timely client VSYNC signal appears. More specifically, an offset is set such that the video frames received within the predetermined number, and which have the highest variability in the one-way latency between the server and client, arrive precisely before the next timely client VSYNC signal appears for display purposes. Proper synchronization and alignment allow for the use of ideal display times for video frames generated at the server, which can be switched between the server and client.
[0133] In one implementation, the timing parameters include an ideal display time, at which the corresponding video frame is intended to be displayed. The ideal display time can be referenced to the target occurrence of the server's VSYNC signal. That is, the ideal display time is explicitly provided in the timing parameters. In one implementation, the timing parameters can be delivered from the server to the client via some mechanism within a data packet used to deliver the encoded video frame. For example, the timing parameters can be added to the data packet header, or the timing parameters may be part of the encoded frame data of the data packet. In another implementation, the timing parameters can be delivered from the server to the client using a GPU API used to send data control packets. The GPU API can be configured to send data control packets from the server to the client via the same data channel used to transmit compressed rendered video frames. The data control packets are formatted so that the client understands what type of information is included and understands the correct reference to the corresponding rendered video frame. In one implementation, the communication protocol for the GPU API, the format for the data control packets, etc., can be defined in the corresponding software development kit (SDK) of the video game, and signaling information provides notification of the data control packets to the client (e.g., provided in the header, provided in a tagged data packet, etc.). In one implementation, data control packets bypass the encoding process because they are minimal in size.
[0134] In another implementation, the timing parameters include the flip time and simulation time from server to client, as previously described. The client can use the flip time and simulation time to determine the ideal display time. That is, the ideal display time is implicitly provided in the timing parameters. The timing parameters may include other information that can be used to infer the ideal display time. Specifically, the flip time indicates when the flip of the display buffer occurs, thus indicating that the corresponding rendered video frame is ready for transmission and / or display. In one implementation, the scan output / scan input process also occurs early in the flip time. Simulation time refers to the time spent rendering the video frame through the CPU and GPU pipeline. The determination of the ideal display time for the corresponding video frame depends on whether the game is executed at a fixed frame rate or a variable frame rate.
[0135] For fixed frame rate games, the client can implicitly determine the target VSYNC timing information from the scan output / scan input timing (e.g., flip timestamps) and the corresponding analog time. For example, the server records the scan output / scan input times for a corresponding video frame and sends them to the client. The client can infer the target occurrence time of the server's VSYNC signal from the scan output / scan input timing and the corresponding analog time, and convert that target occurrence time to the target occurrence time of the client's VSYNC signal. When the game provides ideal display timing (e.g., via a GPU API), the client can explicitly determine the target VSYNC timing information, which may be an integer VSYNC timing or a fractional VSYNC timing. Fractional VSYNC timing can be implemented when the frame processing time exceeds the frame period, where the ideal display timing can be specified by analog time or based on analog time.
[0136] For games with variable frame rates, the client can implicitly determine the ideal target VSYNC timing information from the scan output / scan input timing and analog time of the corresponding video frame. For example, the server records the scan output time and analog time of the corresponding frame and sends them to the client. The client can infer the target appearance time of the server's VSYNC signal for displaying the corresponding video frame from the scan output / scan input timing and analog time, where the target VSYNC signal can be converted into a corresponding target appearance time for the client's VSYNC signal. Alternatively, when the game provides ideal timing via the GPU API, the client can explicitly determine the target VSYNC timing information. In this case, fractional VSYNC timing can be specified by the game, such as providing analog time or display time.
[0137] like Figure 8B As shown, the server VSYNC signal 311 and the client VSYNC signal 312 occur at a timing of 60 Hz. The server VSYNC signal 311 is synchronized (e.g., at substantially equal frequencies) and aligned (e.g., with an offset) with the client VSYNC signal 312. For example, the occurrence of the server VSYNC signal may be aligned with the occurrence of the client VSYNC signal. Specifically, the occurrence of the server VSYNC signal 311a corresponds to the occurrence of the client VSYNC signal 312a, the server VSYNC signal 311c corresponds to the client VSYNC signal 312c, the server VSYNC signal 311d corresponds to the client VSYNC signal 312d, the server VSYNC signal 311e corresponds to the client VSYNC signal 312e, and so on.
[0138] For illustrative purposes, server 260 is executing a video game running at 30Hz, so that rendered video frames are generated at 30Hz (e.g., corresponding to 30 frame cycles per second) during a frame period (33.33 milliseconds). Thus, the video game can render up to 30 frames per second. The ideal display timing for the corresponding video frames is also shown. The ideal display timing may reflect the game's intention to display the video frames. As previously mentioned, the ideal display timing can be determined based on the flip time of each frame, which is also shown. The client can use this ideal display timing to determine when to display the video frames according to the adopted strategy, as described below. For example, video frame A is rendered and ready for display at a flip time of 0.6 (e.g., 0.6 / 60 at 60Hz). Moreover, the ideal display timing of video frame A is intended to be displayed when the server's VSYNC signal 311a occurs, which translates to being intended to be displayed at the client's location when the client's VSYNC signal 312a occurs. Similarly, video frame B is rendered and ready for display at a flip time of 2.1 (e.g., 2.1 / 60 at 60Hz). The ideal display timing for video frame B is for display when the server's VSYNC signal 311c occurs, which translates to display at the client when the client's VSYNC signal 312c occurs. Furthermore, video frame C is rendered and ready for display at flip time 4.1 (e.g., 4.1 / 60 at 60Hz). The ideal display timing for video frame C is intended to be displayed when the server's VSYNC signal 311e occurs, which translates to display at the client when the client's VSYNC signal 312e occurs. Furthermore, video frame D is rendered and ready for display at flip time 7.3 (e.g., 7.3 / 60 at 60Hz). The ideal display timing for video frame D is intended to be displayed when the server's VSYNC signal 311g occurs, which translates to display at the client when the client's VSYNC signal 312g occurs.
[0139] Figure 8B One issue illustrated is that video frame D took longer to generate than expected, resulting in its flip time occurring at 7.3, after the target occurrence of server VSYNC signal 311g. That is, server 260 should have completed rendering video frame D before server VSYNC signal 311g occurred. However, because the ideal display time of video frame D is known or determinable, the client can still display video frame D when client VSYNC signal 312g, which aligns with the ideal display time (e.g., server VSYNC signal 311g), occurs, even if the server misses its time to generate the video frame.
[0140] Figure 8BAnother issue illustrated is that although video frames B and C are generated at server 260 with appropriate timing (e.g., intended to be displayed at different server VSYNC signals), they are received at the client within the same frame cycle due to additional latency during transmission, making them appear as if both were intended to be displayed at the client when the same client VSYNC signal 312d is present. For example, transmission delays cause video frames B and C to arrive within the same frame cycle. However, with proper buffering and knowledge of the ideal display timing for both video frames B and C, the client can determine how and when to display these video frames based on which strategy is implemented, including following game intent, preferred latency, preferred smoothness, or adjusting client-side VBI settings for variable refresh rate displays.
[0141] For example, one strategy is to follow the game intent determined during execution on the server. The intent can be inferred from the timing of the flip times of corresponding video frames, such that video frames A, B, and C are intended for display on the next server VSYNC signal. The intent can be explicitly assumed to be conveyed by the video game, such that video frame D is intended for display on the previous server VSYNC signal 311e, even if it completes rendering after that VSYNC signal. Furthermore, the ambiguity of video frames B and C arriving at the client in similar order (e.g., within the same frame period) will be resolved by following the game's intent. Accordingly, with appropriate buffering, the client can display video frames in the following order at 60Hz (16.66 ms per frame): A -- A -- A -- B -- C -- C -- D -- D, etc.
[0142] The second strategy prioritizes latency over frame display smoothness, aiming to minimize latency as much as possible using minimal buffering. That is, displaying video frames is intended to quickly resolve latency by showing the most recently received video frame in the next client VSYNC signal. Accordingly, the blurring of video frames B and C, which arrive at the client similarly (e.g., within the same frame period), is resolved by discarding video frame B and displaying only video frame C in the next client VSYNC signal. This sacrifices frame smoothness during display because video frame B is skipped in the displayed video frame sequence, potentially attracting the viewer's attention. With proper buffering, the client can then display video frames in the following order at 60Hz (each frame displayed for 16.66 ms): A -- A -- A -- C -- C -- C -- D -- D, etc.
[0143] The third strategy prioritizes frame display smoothness over latency. In this case, additional latency is not a factor and can be addressed with proper buffering. That is, video frames are displayed in a way that provides the viewer with the best viewing experience. The client uses the time between target VSYNCs as a guide; for example, the time between target B 312c and target C 312e is two VSYNCs, so regardless of the arrival times of B and C at the client, B should be displayed for two frames; the time between target C 312e and target D 312g is also two VSYNCs, so regardless of the arrival times of C and D at the client, C should be displayed for two frames, and so on. Accordingly, with proper buffering, the client can display video frames in the following order at 60Hz (each frame is displayed for 16.66 ms): A -- A -- A -- B -- B -- C -- C -- D -- D, etc.
[0144] The fourth strategy provides adjustment of the client-side VBI timing for displays that support variable refresh rates. In other words, variable refresh rate displays allow the VBI interval to be increased or decreased while displaying video frames to achieve an instantaneous frame rate for the video frames rendered for display at the client. For example, instead of displaying the video frame rendered for display at the client with every client VSYNC signal—which might require displaying the video frame twice while waiting for a delayed frame—the display's refresh rate can be dynamically adjusted for each video frame rendered for display. Accordingly, video frames can be displayed to adjust for the variability of the waiting time while they are received, decoded, and rendered at the client for display. Figure 8B In the example shown, although video frames B and C are generated at server 260 with appropriate timing (e.g., intended to be displayed at different server VSYNC signals), they are received at the client within the same frame period due to additional latency during transmission. In this case, video frame B may be displayed for a shorter time period than expected (e.g., less than the frame period) so that video frame C can be rendered at the determined client and at the target client VSYNC signal. For example, video frame C may have the target appearance of the server VSYNC signal and then be converted to the target client VSYNC signal, especially when the server VSYNC signal and the client VSYNC signal are synchronized (e.g., frequency and timing) and aligned with an appropriate offset.
[0145] Figure 9 Components of an example apparatus 900, which can be used to carry out various embodiments of the present disclosure, are shown. For example, Figure 9An exemplary hardware system suitable for streaming media content and / or receiving streamed media content, according to embodiments of the present disclosure, is illustrated, including performing high-speed scan output operations or performing scan output earlier (such as before the next system VSYNC signal or at the flip time of the corresponding video frame) when streaming content from a video game executed at a cloud gaming server across a network to deliver modified video frames to an encoder. The block diagram illustrates device 900, which may be incorporated into or may be a personal computer, server computer, game console, mobile device, or other digital device, each suitable for practicing embodiments of the present invention. Device 900 includes a central processing unit (CPU) 902 for running software applications and optionally an operating system. CPU 902 may include one or more homogeneous or heterogeneous processing cores.
[0146] According to various implementations, CPU 902 is one or more general-purpose microprocessors having one or more processing cores. Other implementations may use one or more CPUs having a microprocessor architecture particularly suited for highly parallel and computationally intensive applications such as media and interactive entertainment applications configured for graphics processing during game execution.
[0147] Memory 904 stores applications and data used by CPU 902 and GPU 916. Storage device 906 provides non-volatile storage and other computer-readable media for applications and data, and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROMs, DVD-ROMs, Blu-ray discs, HD-DVDs, or other optical storage devices, as well as signal transmission and storage media. User input device 908 conveys user input from one or more users to device 900, examples of which may include a keyboard, mouse, joystick, touchpad, touchscreen, still or video recorder / camera, and / or microphone. Network interface 909 allows device 900 to communicate with other computer systems via electronic communication networks and may include wired or wireless communication over local area networks and wide area networks such as the Internet. Audio processor 912 is adapted to generate analog or digital audio output from instructions and / or data provided by CPU 902, memory 904, and / or storage device 906. The components of the device 900 are connected via one or more data buses 922, including a CPU 902, a graphics subsystem 914 including a GPU 916 and a GPU cache 918, a memory 904, a data storage device 906, a user input device 908, a network interface 909, and an audio processor 912.
[0148] The graphics subsystem 914 is also connected to the data bus 922 and components of the device 900. The graphics subsystem 914 includes a graphics processing unit (GPU) 916 and a graphics memory 918. The graphics memory 918 includes display memory (e.g., a frame buffer) for storing pixel data for each pixel of the output image. The graphics memory 918 may be integrated into the same device as the GPU 916, connected to the GPU 916 as a separate device, and / or implemented within memory 904. Pixel data may be provided directly from the CPU 902 to the graphics memory 918. Alternatively, the CPU 902 provides the GPU 916 with data and / or instructions defining the desired output image, and the GPU 916 generates pixel data for one or more output images based on the data and / or instructions. The data and / or instructions defining the desired output image may be stored in memory 904 and / or graphics memory 918. In the implementation, the GPU 916 includes 3D rendering capabilities for generating pixel data for an output image based on instructions and data defining the geometry, lighting, shading, texturing, motion, and / or camera parameters of a scene. The GPU 916 may also include one or more programmable execution units capable of executing shader programs.
[0149] The graphics subsystem 914 periodically outputs pixel data of an image from the graphics memory 918 for display on the display device 910 or projection by a projection system (not shown). The display device 910 may be any device capable of displaying visual information in response to signals from the device 900, including CRT, LCD, plasma, and OLED displays. The device 900 may provide, for example, analog or digital signals to the display device 910.
[0150] Other implementations for optimizing the graphics subsystem 914 may include multi-tenant GPU operations shared among multiple applications by GPU instances, and distributed GPUs supporting a single game. The graphics subsystem 914 may be configured as one or more processing devices.
[0151] For example, in one implementation, the graphics subsystem 914 may be configured to perform multi-tenant GPU functionality, whereby one graphics subsystem can implement graphics and / or rendering pipelines for multiple games. That is, the graphics subsystem 914 is shared among multiple games being executed.
[0152] In other implementations, the graphics subsystem 914 includes multiple GPU devices that are combined to perform graphics processing for a single application running on a corresponding CPU. For example, the multiple GPUs may perform alternating frame rendering, where GPU 1 renders the first frame in a sequential frame cycle, and GPU 2 renders the second frame, and so on, until the last GPU is reached, at which point the initial GPU renders the next video frame (e.g., if there are only two GPUs, GPU 1 renders the third frame). That is, the GPUs take turns rendering frames. Rendering operations may overlap, where GPU 2 may begin rendering the second frame before GPU 1 has finished rendering the first frame. In another implementation, different shader operations may be assigned to the multiple GPU devices in the rendering and / or graphics pipeline. The main GPU is performing main rendering and compositing. For example, in a group comprising three GPUs, the main GPU 1 performs primary rendering (e.g., first shader operations) and composites outputs from the slave GPUs 2 and 3, where the slave GPU 2 performs second shader operations (e.g., fluid effects, such as rivers), and the slave GPU 3 performs third shader operations (e.g., particle smoke), with the main GPU 1 compositing the results from each of GPUs 1, 2, and 3. In this manner, different GPUs can be assigned to perform different shader operations (e.g., waving flags, wind, smoke generation, fire, etc.) to render video frames. In another embodiment, each of the three GPUs can be assigned to different objects and / or portions of a scene corresponding to a video frame. In the above embodiments and implementations, these operations can be performed in the same frame period (simultaneous parallelism) or in different frame periods (sequential parallelism).
[0153] Therefore, this disclosure describes methods and systems configured for streaming media content and / or receiving streaming media content, including performing high-speed scan output operations or performing scan output earlier (such as before the next system VSYNC signal or at the flip time of the corresponding video frame) when streaming content from a video game executed at a cloud gaming server across a network to deliver modified video frames to an encoder.
[0154] It should be understood that the various implementations defined herein can be combined or assembled into specific implementations using the various features disclosed herein. Therefore, the examples provided are merely some possible examples and are not limited to the various implementations that could be defined by combining various elements. In some examples, some implementations may include fewer elements without departing from the spirit of the disclosed or equivalent implementations.
[0155] The embodiments of this disclosure can be practiced with various computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, and the like. The embodiments of this disclosure can also be practiced in distributed computing environments where tasks are performed via remote processing devices based on wired or wireless network links.
[0156] In light of the above embodiments, it should be understood that embodiments of this disclosure can employ various computer-implemented operations involving data stored in a computer system. These operations are those that require the physical manipulation of physical quantities. Any of the operations described herein that form part of embodiments of this disclosure are useful machine operations. Embodiments of this disclosure also relate to means or apparatus for performing these operations. The apparatus may be specifically constructed for the desired purpose, or the apparatus may be a general-purpose computer selectively activated or configured by a computer program stored in a computer. Specifically, various general-purpose machines may be used with computer programs written in accordance with the teachings herein, or it may be more convenient to construct more specialized apparatus to perform the desired operations.
[0157] This disclosure can also be embodied as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device capable of storing data that can subsequently be read by a computer system. Examples of computer-readable media include hard disk drives, network attached storage devices (NAS), read-only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. The computer-readable medium may include computer-readable tangible media distributed across network-connected computer systems to enable distributed storage and execution of computer-readable code.
[0158] Although the method operations are described in a specific order, it should be understood that other housekeeping operations may be performed between operations, or operations may be adjusted so that they occur at slightly different times, or they may be distributed in a system that allows processing operations to occur at various intervals associated with the processing, provided that the processing of the overriding operations is performed in the desired manner.
[0159] While the foregoing disclosure has been described in considerable detail for the purposes of clarity, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims. Therefore, embodiments of the invention are to be considered illustrative rather than restrictive, and embodiments of this disclosure are not limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.
Claims
1. A cloud gaming method, comprising: When a video game is played on a server, video frames are generated within a frame period. Generate one or more user interface features for the video frame; The scanning output process is performed by scanning the video frame and the one or more user interface features into one or more input frame buffers one scan line at a time, and then synthesizing and mixing the video frame and the one or more user interface features into a modified video frame. During the scanning output process, the modified video frames are scanned line by line to the encoder at the server. as well as During the scanning output process, at the time of the video frame flip within the video frame, the scanning of the video frame begins and the scanning of the one or more user interface features into the one or more input frame buffers also begins.
2. The cloud gaming method as described in claim 1, further comprising: The first encoded slice of the video frame is generated at the encoder; as well as Before the modified video frames are fully received, compression of the first encoded slice begins at the encoder.
3. The cloud gaming method as described in claim 1, further comprising: The frame rate setting is determined based on the image size requested for the client's display and the maximum pixel clock. Determine the speed setting value; The scan output process is performed at the speed setting value; as well as The modified multiple screen slices are scanned to the encoder at the frame rate setting.
4. A cloud gaming method, comprising: Video frames are generated when a video game is played on the server. A scan output process is performed to deliver the video frame to an encoder configured to compress the video frame, wherein the scan output process begins at the flip time of the video frame; The compressed video frames are transmitted to the client; The target display time of the video frame is determined at the client. as well as The display time of the video frame is arranged at the client based on the target display time.
5. The cloud gaming method of claim 4, wherein determining the target display time at the client includes: The target display time is inferred based on statistical analysis of multiple arrival times of multiple video frames at the client.
6. A cloud gaming method, comprising: Receive video frames from the frame buffer; Determine the maximum pixel clock of the chipset including the frame buffer, wherein the maximum pixel clock is a static value; The frame rate setting is determined based on the maximum pixel clock and the target display of the client device; as well as Before the next occurrence of the VSYNC signal used to time the generation of multiple video frames, the video frames are scanned and output to the encoder at the frame rate set. The frame rate setting is determined independently of the rate at which multiple video frames are generated.
7. A cloud gaming method, comprising: When a video game is executed on the server, video frames are generated within a frame period, wherein generating the video frames includes a flip time that occurs before the next occurrence of the server's vertical synchronization (VSYNC) signal, and executing a command indicating that the video frame has been fully rendered; and A scan output process is performed on the video frame to deliver the video frame to an encoder configured to compress the video frame, wherein the scan output process is performed on the video frame at the flip time.
8. A cloud gaming method, comprising: Align the server's VSYNC signal with the client's VSYNC signal; At the client end, encoded video frames are received from the server; At the client end, timing information for the video frame is received from the server; based on the timing information, the target display time of the video frame is determined. The target display time is the target occurrence time of the server's VSYNC signal; Convert the target display time into the target occurrence of the client's VSYNC signal; and The display time of the video frame is scheduled based on the appearance of the target in the client's VSYNC signal.
9. A cloud gaming method, comprising: At the client end, encoded video frames are received from the server; Decode the encoded video frames; Receive timing information for the occurrence of a target in the server VSYNC signal, wherein the video frame is intended to be displayed when the target in the server VSYNC signal occurs; The target occurrence of the server's VSYNC signal is converted into the target occurrence of the client's VSYNC signal; as well as When the target of the client's VSYNC signal appears, the decoded video frame is displayed.
10. The cloud gaming method as described in claim 9, further comprising: The received video frames arrive at the client in a timely manner based on a predetermined number of frames so that they are displayed when the corresponding VSYNC signal on the client appears again, offset from the VSYNC signal on the server.