Encoder tuning to improve tradeoff between latency and video quality in cloud gaming applications

By tuning high-performance encoders and decoders and dynamically adjusting encoding parameters, the problem of excessive latency in cloud gaming has been solved, improving video quality and user experience.

CN114746157BActive Publication Date: 2026-04-28SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SONY INTERACTIVE ENTERTAINMENT LLC
Filing Date
2020-09-30
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Cloud gaming suffers from insufficient network connectivity and response speed, resulting in excessively long latency for high-quality image rendering and impacting user experience.

Method used

By tuning high-performance encoders and decoders, and monitoring client bandwidth, skipped frames, number of I-frames, and video frame size, encoder parameters are dynamically adjusted to reduce the total one-way latency between the cloud gaming server and the client.

Benefits of technology

This approach achieves a better balance between one-way latency and video quality in cloud gaming systems, reducing the total one-way latency between cloud gaming servers and clients and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114746157B_ABST
    Figure CN114746157B_ABST
Patent Text Reader

Abstract

A method for cloud gaming. The method includes generating a plurality of video frames when executing a video game at a cloud gaming server. The method includes encoding the plurality of video frames at a certain encoder bitrate, where the plurality of video frames to be compressed are transmitted from a streamer of the cloud gaming server to a client. The method includes measuring a maximum receiving bandwidth of the client. The method includes monitoring the encoding of the plurality of video frames at the streamer. The method includes dynamically tuning parameters of the encoder based on the monitoring of the encoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to streaming systems configured to stream content across networks, and more specifically, to high-performance encoders and decoders configured for cloud gaming systems, and to streaming systems that tune the encoder while knowing network speed and reliability and total latency targets. Background of the Invention

[0003] In recent years, there has been a continuous push for online services that allow online or cloud gaming to be streamed between cloud gaming servers and clients connected via a network. Streaming formats are gaining popularity because they enable on-demand use of game titles, the ability for players to network for multiplayer games, the sharing of assets between players, the sharing of instant experiences between players and / or spectators, allowing friends to watch friends play video games, and letting friends join games that friends are playing, among other things. Unfortunately, these demands are also exacerbating limitations such as network connectivity and the processing speed at both the server and client ends, which is sufficient to render the high-quality images delivered to the client. For example, for the best user experience, the results of all game activities performed on the server need to be compressed and transmitted back to the client with low millisecond latency. Round-trip latency can be defined as the total time between a user's controller input and the display of a video frame at the client; it can include the processing and transmission of control information from the controller to the client, the processing and transmission of control information from the client to the server, the generation of a video frame at the server in response to that input, processing the video frame and passing it to an encoding unit (e.g., a scan output), encoding the video frame, transmitting the encoded video frame back to the client, receiving and decoding the video frame, and any processing or grading of the video frame before its display. One-way latency can be defined as a portion of the round-trip latency, consisting of the time from the start of passing the video frame to the encoding unit (e.g., a scan output) at the server to the start of displaying the video frame at the client. A portion of the round-trip latency and one-way latency are associated with the time taken to send a data stream from the client to the server and from the server to the client via the communication network. Another portion is associated with the processing at the client and server; improvements to these operations, such as advanced strategies related to frame decoding and display, can result in substantially reduced round-trip and one-way latency between the server and client, and provide a higher quality experience for users of cloud gaming services.

[0004] It is against this backdrop that the proposed implementation scheme has emerged. Summary of the Invention

[0005] Implementations of this disclosure relate to streaming systems configured for streaming content (e.g., games) across a network, and more specifically, to streaming systems configured to provide encoder tuning to improve the trade-off between one-way latency and video quality in cloud gaming systems, wherein encoder tuning may be based on monitoring client bandwidth, skipped frames, the number of encoded I-frames, the number of scene changes, and / or the number of video frames exceeding a target frame size, wherein tuned parameters may include encoder bit rate, target frame size, maximum frame size, and quantization parameter (QP) values, wherein high-performance encoders and decoders contribute to reducing the total one-way latency between the cloud gaming server and the client.

[0006] This disclosure provides a method for cloud gaming. The method includes generating multiple video frames while a video game is executed at a cloud gaming server. The method includes encoding the multiple video frames at an encoder bit rate, wherein the compressed multiple video frames are transmitted from a streaming device of the cloud gaming server to a client. The method includes measuring the maximum receiving bandwidth of the client. The method includes monitoring the encoding of the multiple video frames at the streaming device. The method includes dynamically tuning encoder parameters based on the monitoring of the encoding.

[0007] In another embodiment, a non-transitory computer-readable medium storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating multiple video frames when a video game is executed at a cloud gaming server. The computer-readable medium includes program instructions for encoding the multiple video frames at a certain encoder bit rate, wherein the compressed multiple video frames are transmitted from a streaming device of the cloud gaming server to a client. The computer-readable medium includes program instructions for measuring the maximum receiving bandwidth of the client. The computer-readable medium includes program instructions for monitoring the encoding of the multiple video frames at the streaming device. The computer-readable medium includes program instructions for dynamically tuning encoder parameters based on the monitoring of the encoding.

[0008] In another embodiment, a computer system includes a processor and a memory coupled to the processor and storing instructions therein that, when executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating a plurality of video frames when a video game is executed at a cloud gaming server. The method includes encoding the plurality of video frames at a certain encoder bit rate, wherein the compressed plurality of video frames are transmitted from a streaming device of the cloud gaming server to a client. The method includes measuring the maximum receiving bandwidth of the client. The method includes monitoring the encoding of the plurality of video frames at the streaming device. The method includes dynamically tuning encoder parameters based on the monitoring of the encoding.

[0009] In another embodiment, a method for cloud gaming is disclosed. The method includes generating multiple video frames when a video game is executed at a cloud gaming server. The method includes predicting a scene change for a first video frame of the video game, wherein the scene change is predicted before generating the first video frame. The method includes generating a scene change cue that the first video frame is a scene change. The method includes sending the scene change cue to an encoder. The method includes delivering the first video frame to the encoder, wherein the first video frame is encoded as an I-frame based on the scene change cue. The method includes measuring the maximum receiving bandwidth of the client. The method includes determining whether to encode a second video frame received at the encoder based on the client's maximum receiving bandwidth and the target resolution of the client's display.

[0010] In another embodiment, a non-transitory computer-readable medium storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating a plurality of video frames when a video game is executed at a cloud gaming server. The computer-readable medium includes program instructions for predicting scene changes in a first video frame of the video game, wherein the scene changes are predicted before the first video frame is generated. The computer-readable medium includes program instructions for generating a scene change cue that the first video frame is a scene change. The computer-readable medium includes program instructions for sending the scene change cue to an encoder. The computer-readable medium includes program instructions for delivering the first video frame to the encoder, wherein the first video frame is encoded as an I-frame based on the scene change cue. The computer-readable medium includes program instructions for measuring the maximum receiving bandwidth of a client. The computer-readable medium includes program instructions for determining whether to encode a second video frame received at the encoder based on the client's maximum receiving bandwidth and the target resolution of the client's display.

[0011] In another embodiment, a computer system includes a processor and a memory coupled to the processor and storing instructions therein that, when executed by the computer system, cause the computer system to perform a method for cloud gaming. The method includes generating a plurality of video frames when a video game is executed at a cloud gaming server. The method includes predicting a scene change for a first video frame of the video game, wherein the scene change is predicted prior to generating the first video frame. The method includes generating a scene change cue indicating that the first video frame is a scene change. The method includes sending the scene change cue to an encoder. The method includes delivering the first video frame to the encoder, wherein the first video frame is encoded as an I-frame based on the scene change cue. The method includes measuring the maximum receiving bandwidth of a client. The method includes determining whether to encode a second video frame received at the encoder based on the client's maximum receiving bandwidth and the target resolution of the client's display.

[0012] Other aspects of this disclosure will become apparent from the following detailed description taken in conjunction with the accompanying drawings, which illustrate the principles of this disclosure by way of example. Attached Figure Description

[0013] This disclosure is best understood by referring to the following description taken in conjunction with the accompanying drawings, in which:

[0014] Figure 1A This is a diagram of the VSYNC signal at the beginning of a frame period according to one embodiment of the present disclosure.

[0015] Figure 1B This is a diagram showing the frequency of the VSYNC signal according to one embodiment of this disclosure.

[0016] Figure 2A This is a diagram of a system according to one embodiment of the present disclosure for providing games via a network between one or more cloud gaming servers and one or more client devices in various configurations, wherein the VSYNC signal can be synchronized and offset to reduce one-way latency.

[0017] Figure 2B This is a diagram of an embodiment of the present disclosure for providing a game between two or more peer devices, wherein the VSYNC signal can be synchronized and offset to achieve optimal timing for receiving controller and other information between the devices.

[0018] Figure 2C Various network configurations that benefit from appropriate synchronization and offset of the VSYNC signal between the source and target devices according to one embodiment of this disclosure are illustrated.

[0019] Figure 2DThe illustration depicts a multi-tenant configuration between a cloud gaming server and multiple clients that benefits from proper synchronization and offset of the VSYNC signal between the source and target devices, according to one embodiment of the present disclosure.

[0020] Figure 3 The illustration depicts the variation in one-way latency between the cloud gaming server and the client due to clock drift when streaming video frames generated from a video game executed on a server, according to one embodiment of the present disclosure.

[0021] Figure 4 The diagram illustrates the network configuration, including the cloud gaming server and client, when streaming video frames generated from a video game executed on the server. This configuration synchronizes and offsets the VSYNC signals between the server and client to allow overlap of operations at the server and client, and reduces one-way latency between the server and client.

[0022] Figure 5 This is a flowchart illustrating a method for cloud gaming according to one embodiment of the present disclosure, wherein encoding video frames includes tuning encoder parameters with knowledge of network transmission speed and reliability as well as a total latency target.

[0023] Figure 6 This is a diagram illustrating a method for measuring the bandwidth of a client by means of a streaming component operating at the application layer, according to one embodiment of the present disclosure, wherein the streaming component is configured to monitor and tune the encoder so that compressed video frames can be transmitted at a rate within the measured bandwidth of the client.

[0024] Figure 7A This is a graph illustrating the setting of encoder quantization parameters (QP) according to one embodiment of the present disclosure to optimize quality and buffer utilization at the client end.

[0025] Figure 7B This is a diagram illustrating the tuning target frame size, maximum frame size, and / or QP (e.g., minQP and / or maxQP) encoder settings according to one embodiment of the present disclosure to reduce the occurrence of I-frames exceeding the true target frame size supported by the client.

[0026] Figure 8 This is a flowchart illustrating a method for cloud gaming according to one embodiment of the present disclosure, wherein encoding video frames includes determining when to skip video frames or delay the encoding and transmission of video frames when the encoding run time is long or when the generated video frames are large (such as when encoding I-frames).

[0027] Figure 9AA sequence of video frames compressed by an encoder according to one embodiment of the present disclosure is illustrated, wherein when the client bandwidth is low for the target resolution of the client's display, the encoder abandons encoding the video frames after encoding the I-frames.

[0028] Figures 9B to 9C The illustration depicts a sequence of video frames compressed by an encoder according to one embodiment of the present disclosure, wherein each of the video frames in the sequence is encoded as an I-frame when the client bandwidth is moderate or high for the target resolution of the client's display, and subsequent video frames are encoded after a delay in encoding the I-frames.

[0029] Figure 10 Components of exemplary apparatus that can be used to carry out various embodiments of this disclosure are illustrated. Detailed Implementation

[0030] While the following detailed description contains many specific details for illustrative purposes, those skilled in the art will understand that many variations and modifications to these details are within the scope of this disclosure. Therefore, aspects of the disclosure described below are set forth without diminishing the generality of the claims following this description and without imposing limitations on the claims.

[0031] Generally, various embodiments of this disclosure describe methods and systems configured to reduce latency and / or latency instability between a source device and a target device when streaming media content (e.g., streaming audio and video from a video game). Latency instability can be introduced into the one-way latency between the server and the client due to factors such as the additional time required to generate complex frames (e.g., scene changes) at the server, the increased time required to encode / compress complex frames at the server, variable communication paths on the network, and the increased time required to decode complex frames at the client. Latency instability can also be introduced due to clock differences at the server and client, which can cause drift between the server's VSYNC signal and the client's VSYNC signal. In embodiments of this disclosure, one-way latency between the server and the client in cloud gaming applications can be reduced by providing high-performance encoding and decoding. When decompressing streaming media (e.g., streaming video, movies, clips, content), it is possible to buffer a large amount of decompressed video, thus potentially relying on average decoding capabilities and metrics (e.g., relying on an average amount of decoding resources to support 4K media at 60Hz) when displaying streaming content. However, for cloud gaming, increasing the time spent performing encoding and / or decoding operations, even for a single frame, results in correspondingly higher one-way latency. Therefore, it is beneficial for cloud gaming to provide more robust decoding and encoding resources, which appear unnecessary compared to the needs of streaming video applications and should be optimized for handling frames requiring longer or longest processing times. In other embodiments of this disclosure, encoder tuning can be performed to improve the trade-off between latency and video quality in cloud gaming applications. Encoder tuning is performed with knowledge of network transmission speed and reliability, as well as the overall latency target. In embodiments, when encoding runs are long or the generated data is large (e.g., for compressed I-frames, both are possible), a method is performed to determine whether to delay the encoding and transmission of subsequent frames or skip them. In embodiments, tuning of the quantization parameter (QP) value, target frame size, and maximum frame size is performed based on the network speed available to the client. For example, if the network speed is high, the QP can be reduced. In other embodiments, monitoring of the I-frame occurrence rate is performed and used to set the QP. For example, if I-frames are not frequent, the QP can be reduced (e.g., to give higher encoding precision or higher encoding quality), allowing the encoding of video frames to be skipped to keep one-way latency low, while sacrificing video playback quality. Therefore, high-performance encoding and decoding, along with encoder tuning, performed to improve the trade-off between latency and video quality in cloud gaming applications, result in reduced one-way latency, smoother frame rates, and more reliable and / or consistent one-way latency between cloud gaming servers and clients.

[0032] Based on the above general understanding of the various implementation schemes, exemplary details of the implementation schemes will now be described with reference to the various diagrams.

[0033] Throughout this specification, references to "game," "video game," or "game application" are intended to refer to any type of interactive application that is initiated by executing input commands. For illustrative purposes only, interactive applications include applications for games, word processing, video processing, video game processing, etc. Furthermore, the terms used above are interchangeable.

[0034] Cloud gaming involves executing a video game at a server to generate game-rendered video frames, and then sending those game-rendered video frames to a client for display. The timing of operations at both the server and client can be associated with corresponding vertical synchronization (VSYNC) parameters. When the VSYNC signals are properly synchronized and / or offset between the server and / or client, operations performed at the server (e.g., generating and transmitting video frames within one or more frame periods) are synchronized with operations performed at the client (e.g., displaying video frames on a display at a display frame rate or refresh rate corresponding to the frame period). Specifically, the server-side VSYNC signal generated at the server and the client-side VSYNC signal generated at the client can be used to synchronize operations at the server and client. That is, how the server and client display video frames is synchronized when the server-side and client-side VSYNC signals are synchronized and / or offset.

[0035] VSYNC signaling and Vertical Blanking Interval (VBI) have been incorporated to generate and display video frames when streaming media content between a server and a client. For example, the server attempts to generate game-rendered video frames within one or more frame cycles defined by the corresponding server VSYNC signal (e.g., generating one video frame per frame cycle results in 60Hz operation if the frame cycle is 16.7ms, and generating one video frame every two frame cycles results in 30Hz operation), and then encodes and transmits that video frame to the client. At the client, the received encoded video frame is decoded and displayed, with the client display rendering each video frame for display at the start of the corresponding client VSYNC signal.

[0036] For the sake of explanation, Figure 1AThis illustrates how the VSYNC signal 111 can indicate the start of a frame period, during which various operations can be performed at the server and / or client during the corresponding frame period. When streaming media content, the server can use the server VSYNC signal to generate and encode video frames, and the client can use the client VSYNC signal to display video frames. The VSYNC signal 111 is generated at a defined frequency corresponding to the defined frame period 110, as follows: Figure 1B As shown in the figure. Furthermore, VBI 105 defines the period between when the last raster line is drawn on the display during the previous frame cycle and when the first raster line (e.g., the top) is drawn onto the display. As shown, after VBI 105, the video frames rendered for display are displayed via raster scan lines 106 (e.g., raster line by raster line from left to right).

[0037] Furthermore, various embodiments of this disclosure are disclosed for reducing one-way latency and / or latency instability between a source device and a target device, such as when streaming media content (e.g., video game content). For illustrative purposes only, various embodiments for reducing one-way latency and / or latency instability are described in a server and client network configuration. However, it should be understood that the various techniques disclosed for reducing one-way latency and / or latency instability can be implemented in other network configurations and / or on peer-to-peer networks, such as in... Figures 2A to 2D As shown in the figure. For example, the various implementations disclosed for reducing one-way latency and / or latency instability can be implemented in various configurations between one or more of the server and client devices (e.g., server and client, server and server, server and multiple clients, server and multiple servers, client and client, client and multiple clients, etc.).

[0038] Figure 2AThis diagram illustrates a system 200A, according to one embodiment of the present disclosure, for providing games in various configurations between one or more cloud gaming networks 290 and / or servers 260 and one or more client devices 210 via network 250. This includes the possibility of synchronizing and offsetting server and client VSYNC signals, and / or implementing dynamic buffering on the client, and / or overlapping encoding and transmission operations on the server, and / or overlapping reception and decoding operations on the client, and / or overlapping decoding and display operations on the client, to reduce one-way latency between server 260 and client 210. Specifically, according to one embodiment of the present disclosure, system 200A provides games via cloud gaming network 290, wherein the game is executed remotely relative to a client device 210 (e.g., a thin client) of a corresponding user playing the game. System 200A can provide game control via network 250 to one or more users playing one or more games via cloud gaming network 290 in single-player or multi-player mode. In some implementations, cloud gaming network 290 may include multiple virtual machines (VMs) running on a host hypervisor, wherein one or more VMs are configured to execute a game processor module that utilizes hardware resources available to the host hypervisor. Network 250 may include one or more communication technologies. In some implementations, network 250 may include fifth-generation (5G) network technology with advanced wireless communication systems.

[0039] In some implementations, wireless technologies can be used to facilitate communication. Such technologies may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. A 5G network is a digital cellular network in which the service area covered by the provider is divided into small geographic areas, called cells. Analog signals representing voice and images are digitized in a telephone call, converted by an analog-to-digital converter, and transmitted as a bit stream. All 5G wireless devices in a cell communicate via radio waves through frequency channels assigned from a frequency pool by a transceiver, with the local antenna array and low-power automatic transceivers (transmitters and receivers) in that cell, the frequencies being reused in other cells. The local antennas are connected to the telephone network and the Internet via high-bandwidth fiber optic or wireless backhaul connections. As in other cellular networks, mobile devices moving from one cell to another are automatically transferred to the new cell. It should be understood that 5G networks are merely exemplary types of communication networks, and embodiments of this disclosure may utilize previous generations of wireless or wired communication, as well as subsequent generations of wired or wireless technologies after 5G.

[0040] As shown in the figure, the cloud gaming network 290 includes a game server 260 that provides access to multiple video games. The game server 260 can be any type of server computing device available in the cloud and can be configured to execute one or more virtual machines on one or more hosts. For example, the game server 260 can manage virtual machines that support game processors that instantiate instances of games for users. Thus, multiple game processors of the game server 260 associated with multiple virtual machines are configured to execute multiple instances of one or more games associated with multiple users playing games. In this way, the backend server supports streaming of media (e.g., video, audio, etc.) for multiple game applications to multiple corresponding users. That is, the game server 260 is configured to stream data (e.g., corresponding rendered images and / or frames of the game) back to the corresponding client device 210 via network 250. In this way, computationally complex game applications can be executed at the backend server in response to controller input received and forwarded by the client device 210. Each server is capable of rendering images and / or frames, then encoding (e.g., compressing) the images and / or frames and streaming them to the corresponding client device for display.

[0041] For example, multiple users can access the cloud gaming network 290 via a communication network 250 using a corresponding client device 210 configured to receive streaming media. In one embodiment, the client device 210 may be configured as a thin client, providing an interface to a backend server (e.g., game server 260 of the cloud gaming network 290) configured to provide computing functionality (e.g., including a game title processing engine 211). In another embodiment, the client device 210 may be configured with a game title processing engine and game logic for at least some local processing of the video game, and may be further configured to receive streaming content generated by the video game executed on the backend server, or for other content supported by the backend server. For local processing, the game title processing engine includes basic processor-based functionality for executing the video game and services associated with the video game. The game logic is stored on the local client device 210 and used to execute the video game.

[0042] Specifically, a client device 210 corresponding to a user (not shown) is configured to request access to a game via a communication network 250 such as the Internet, and to render display images generated by a video game executed by a game server 260, wherein encoded images are delivered to the client device 210 associated with the corresponding user for display. For example, a user can interact with an instance of a video game executed on a game processor of the game server 260 via the client device 210. More specifically, the instance of the video game is executed by a game title processing engine 211. Corresponding game logic (e.g., executable code) 215 implementing the video game is stored and accessible via a data storage device (not shown) and used to execute the video game. The game title processing engine 211 can use multiple game logics to support multiple video games, each of which can be selected by the user.

[0043] For example, client device 210 is configured to interact with game title processing engine 211 associated with the corresponding user's game, such as by input commands used to drive the game. Specifically, client device 210 can receive input from various types of input devices, such as game controllers, tablets, keyboards, gestures captured by cameras, mice, touchpads, etc. Client device 210 can be any type of computing device having at least memory and a processor module capable of connecting to game server 260 via network 250. Backend game title processing engine 211 is configured to generate rendered images, which are delivered via network 250 for display at a corresponding display associated with client device 210. For example, through a cloud-based service, the game-rendered images can be delivered by an instance of the corresponding game executed on game execution engine 211 of game server 260. That is, client device 210 is configured to receive encoded images (e.g., encoded from game-rendered images generated by executing a video game) and display them as images rendered by display 11. In one embodiment, display 11 includes HMD (e.g., displaying VR content). In some implementations, the rendered image can be wirelessly or wired directly from a cloud-based service or via client device 210 (e.g., Remote Play allows you to stream to your smartphone or tablet.

[0044] In one implementation, the game server 260 and / or the game title processing engine 211 include basic processor-based functions for executing the game and services associated with the game application. For example, processor-based functions include 2D or 3D rendering, physics, physics simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadows, culling, transformations, artificial intelligence, etc. Furthermore, the game application services include memory management, multithreading management, Quality of Service (QoS), bandwidth testing, social networks, social friend management, communication with friends' social networks, communication channels, text messaging, instant messaging, chat support, etc.

[0045] In one implementation, the cloud gaming network 290 is a distributed game server system and / or architecture. Specifically, a distributed game engine executing game logic is configured as a corresponding instance of a game. Generally, the distributed game engine employs each of the functions of a game engine and distributes those functions to multiple processing entities for execution. Individual functions may be further distributed across one or more processing entities. The processing entities may be configured in different configurations, including physical hardware, and / or as virtual components or virtual machines, and / or as virtual containers, where containers differ from virtual machines because containers virtualize instances of game applications running on a virtualized operating system. The processing entities may utilize and / or rely on their underlying hardware on one or more servers (compute nodes) of the cloud gaming network 290, wherein the servers may reside on one or more racks. Coordination, assignment, and management of the execution of these functions by the various processing entities are performed by a distributed synchronization layer. In that manner, the execution of those functions is controlled by the distributed synchronization layer, enabling the generation of media (e.g., video frames, audio, etc.) of the game application in response to player controller input. The distributed synchronization layer enables those functions to be executed (e.g., through load balancing) efficiently across distributed processing entities, allowing critical game engine components / functions to be distributed and reorganized for more efficient processing.

[0046] The game title processing engine 211 includes a central processing unit (CPU) and a graphics processing unit (GPU) group that can be configured to perform multi-tenant GPU functionality. In another embodiment, multiple GPU devices are combined to perform graphics processing for a single application running on a corresponding CPU.

[0047] Figure 2B This diagram illustrates an embodiment of the present disclosure for providing a game between two or more peer devices, wherein VSYNC signals can be synchronized and offset to achieve optimal timing for receiving controller and other information between the devices. For example, two or more peer devices connected via a network 250 or directly connected via peer-to-peer communication (e.g., Bluetooth, LAN, etc.) can be used to perform head-to-head games.

[0048] As shown, the game is executed locally on each of the client devices 210 (e.g., game consoles) of the corresponding user playing the video game, where the client devices 210 communicate via a peer-to-peer network. For example, an instance of the video game is executed by the game title processing engine 211 of the corresponding client device 210. The game logic 215 (e.g., executable code) that implements the video game is stored on the corresponding client device 210 and is used to execute the game. For illustrative purposes, the game logic 215 can be delivered to the corresponding client device 210 via a portable medium (e.g., optical medium) or via a network (e.g., downloaded from a game provider via the Internet).

[0049] In one implementation, the game title processing engine 211 of the corresponding client device 210 includes basic processor-based functions for executing the game and services associated with the game application. For example, processor-based functions include 2D or 3D rendering, physics, physics simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadows, culling, transformation, artificial intelligence, etc. Furthermore, the game application services include memory management, multithreading management, Quality of Service (QoS), bandwidth testing, social networks, social friend management, communication with friends' social networks, communication channels, text messaging, instant messaging, chat support, etc.

[0050] Client device 210 can receive input from various types of input devices, such as game controllers, tablet computers, keyboards, gestures captured by a camera, mice, touchpads, etc. Client device 210 can be any type of computing device having at least a memory and processor module, and is configured to generate rendered images executed by game title processing engine 211, and display the rendered images on a display (e.g., display 11, or display 11 including a head-mounted display (HMD), etc.). For example, the rendered images can be associated with an instance of a game running locally on client device 210, to implement gameplay for the corresponding user, such as through input commands used to drive gameplay. Some examples of client device 210 include personal computers (PCs), game consoles, home theater devices, general-purpose computers, mobile computing devices, tablet computers, telephones, or any other type of computing device capable of executing game instances.

[0051] Figure 2C Various network configurations that benefit from appropriate synchronization and offset of the VSYNC signal between the source and target devices according to embodiments of this disclosure are illustrated, including... Figures 2A to 2BThe configurations shown are examples of those described above. Specifically, various network configurations benefit from proper alignment of the frequencies of the server and client VSYNC signals to reduce one-way latency and / or latency variability between the server and client, as well as timing offsets between the server and client VSYNC signals. For example, one network device configuration includes a cloud gaming server (e.g., source) to client (target) configuration. In one embodiment, the client may include a WebRTC client configured to provide audio and video communication within a web browser. Another network configuration includes a client (e.g., source) to server (target) configuration. Yet another network configuration includes a server (e.g., source) to server (e.g., target) configuration. Yet another network device configuration includes a client (e.g., source) to client (target) configuration, wherein, for example, each client may be a game console used to provide head-to-head gameplay.

[0052] Specifically, VSYNC signal alignment may include synchronizing the frequencies of the server VSYNC signal and the client VSYNC signal, and may also include tuning the timing offset between the client VSYNC signal and the server VSYNC signal to eliminate drift, and / or maintain an ideal relationship between the server VSYNC signal and the client VSYNC signal to reduce unidirectional delay and / or delay variability. In one embodiment, to achieve proper alignment, the server VSYNC signal may be tuned to implement proper alignment between the server 260 and client 210 pair. In another embodiment, the client VSYNC signal may be tuned to implement proper alignment between the server 260 and client 210 pair. Once the client VSYNC signal and the server VSYNC signal are aligned, the server VSYNC signal and the client VSYNC signal occur at substantially the same frequency and are offset from each other by a timing offset that can be adjusted from time to time. In another embodiment, alignment of the VSYNC signal may include synchronizing the frequencies of the VSYNC signals of two clients, and may also include adjusting the timing offset between their VSYNC signals to eliminate drift and / or achieve optimal timing for receiving controller and other information; either VSYNC signal may be tuned to achieve this alignment. In yet another embodiment, alignment may include synchronizing the frequencies of the VSYNC signals of multiple servers, and may also include synchronizing the frequencies of the server VSYNC signals and client VSYNC signals, and adjusting the timing offset between the client VSYNC signals and server VSYNC signals, for example, for use in head-to-head cloud gaming. In server-to-client and client-to-client configurations, alignment may include synchronizing the frequencies of the server VSYNC signals and client VSYNC signals, and providing an appropriate timing offset between the server VSYNC signals and client VSYNC signals. In a server-to-server configuration, alignment may include synchronizing the frequencies of the server VSYNC signals and client VSYNC signals without setting a timing offset.

[0053] Figure 2D This illustration depicts a multi-tenant configuration between a cloud gaming server 260 and one or more clients 210, benefiting from proper synchronization and offset of VSYNC signals between a source device and a target device, according to one embodiment of this disclosure. In the server-to-client configuration, alignment may include frequency synchronization between the server's VSYNC signal and the client's VSYNC signal, as well as providing appropriate timing offset between the server's VSYNC signal and the client's VSYNC signal. In one embodiment, in the multi-tenant configuration, the client's VSYNC signal is tuned at each client 210 to implement proper alignment between the server 260 and client 210 pairs.

[0054] For example, in one embodiment, a graphics subsystem may be configured to perform multi-tenant GPU functionality, wherein the graphics subsystem may implement graphics and / or rendering pipelines for multiple games. That is, the graphics subsystem is shared among multiple games being executed. Specifically, in one embodiment, a game title processing engine may include a group of CPUs and GPUs configured to perform multi-tenant GPU functionality, wherein the CPU and GPU group may implement graphics and / or rendering pipelines for multiple games. That is, the CPU and GPU group is shared among multiple games being executed. The CPU and GPU group may be configured as one or more processing devices. In another embodiment, multiple GPU devices are combined to perform graphics processing for a single application executing on a corresponding CPU.

[0055] Figure 3The diagram illustrates the general process of executing a video game at a server to generate game-rendered video frames and sending those frames to a client for display. Traditionally, many operations at the game server 260 and client 210 are performed within frame periods defined by corresponding VSYNC signals. For example, server 260 attempts to generate game-rendered video frames at 301 within one or more frame periods defined by corresponding server VSYNC signals 311. The game generates video frames in response to control information delivered from an input device (e.g., user input commands) or game logic not driven by control information at operation 350. Transmission jitter 351 may exist when control information is sent to server 260, where jitter 351 measures the variation in network latency from the client to the server (e.g., when an input command is sent). As shown, the thick arrows indicate the current latency when control information is sent to server 260, but due to jitter, there may be a range of arrival times for control information at server 260 (e.g., the range defined by the dotted arrows). At flip time 309, the GPU receives a flip command indicating that the corresponding video frame has been fully generated and placed in the frame buffer of server 260. Thereafter, server 260 performs a scan output / scan input (operation 302, where the scan output may be aligned with VSYNC signal 311) for that video frame within the subsequent frame period defined by server VSYNC signal 311 (VBI omitted for clarity). The video frame is then encoded (operation 303) (e.g., encoding begins after VSYNC signal 311, and the end of encoding may not be aligned with VSYNC signal 311) and transmitted (operation 304, where transmission may not be aligned with VSYNC signal 311) to client 210. At client 210, encoded video frames are received (operation 305, where reception may not be aligned with client VSYNC signal 312), decoded (operation 306, where decoding may not be aligned with client VSYNC signal 312), buffered, and displayed (operation 307, where the start of display may be aligned with client VSYNC signal 312). Specifically, client 210 displays each video frame rendered to begin displaying upon the corresponding occurrence of client VSYNC signal 312.

[0056] One-way delay 315 can be defined as the delay from the start of transmitting a video frame to the encoding unit at the server (e.g., scan output 302) to the start of displaying the video frame at the client 307. That is, one-way delay is the time from the server scan output to the client display, taking into account client buffering. Individual frames have a delay from the start of scan output 302 to the completion of decoding 306, which can vary between frames due to server operations such as encoding 303 and transmission 304, network transmission between server 260 and client 210, and accompanying jitter 352, as well as height variations at client reception 305. As shown, the straight, thick arrows indicate the current delay when the corresponding video frame is sent to client 210, but due to jitter 352, the video frame at client 210 may have an arrival time range (e.g., a range defined by the dotted arrows). Because one-way latency must be relatively stable (e.g., remain fairly consistent) for a good playback experience, buffering 320 is traditionally performed. As a result, the display of individual frames with low latency (e.g., from the start of scan output 302 to the completion of decoding 306) is delayed by several frame cycles. That is, if there is network instability or unpredictable encoding / decoding time, additional buffering is needed to keep the one-way latency consistent.

[0057] According to one embodiment of this disclosure, when streaming video frames generated from a video game executed on a server, the one-way latency between the cloud gaming server and the client can vary due to clock drift. Specifically, the frequency difference between the server's VSYNC signal 311 and the client's VSYNC signal 312 may cause the client's VSYNC signal to drift relative to frames arriving from the server 260. This drift may be caused by a very small difference in the crystal oscillators used in each of the corresponding clocks at the server and client. Furthermore, embodiments of this disclosure reduce one-way latency by: performing synchronization and offset of one or more of the VSYNC signals for alignment between the server and client; providing dynamic buffering on the client; overlapping the encoding and transmission of video frames at the server; overlapping the reception and decoding of video frames at the client; and overlapping the decoding and display of video frames at the client.

[0058] Furthermore, during the encoding of video frames (operation 303), in the prior art, the encoder determines how much change exists between the current video frame being encoded and one or more previously encoded frames to determine whether a scene change exists (e.g., a complex image of the corresponding generated video frame). That is, scene change cues can be inferred from the differences between the current frame to be encoded and previously encoded frames. When content is streamed from a server to a client via a network, the encoder at the server can decide to encode video frames detected as scene changes with greater complexity. Otherwise, the encoder will encode video frames not detected as scene changes with less complexity. However, detecting scene changes at the encoder can take up to a frame cycle (e.g., adding jitter) because the video frame is initially encoded with less complexity (in the first frame cycle), but then re-encoded with greater complexity once a scene change is determined (in the second frame cycle). Moreover, scene change detection may be unnecessarily triggered (such as by slight explosions in the image) because even without a scene change, the difference between the currently encoded video frame and previously encoded video frames may exceed a threshold difference value. Therefore, when a scene change is detected at the encoder, an additional delay due to jitter is introduced at the encoder to accommodate the scene change detection and the more complex re-encoding of video frames.

[0059] Figure 4 The illustration depicts a data stream, according to an embodiment of this disclosure, during streaming of video frames generated from a video game executed on a server, via a network configuration including a highly optimized cloud gaming server 260 and a highly optimized client 210. Overlapping server and client operations reduce one-way latency, and VSYNC signal synchronization and offset between the server and client further reduce one-way latency and decrease the variability of one-way latency between the server and client. Specifically, Figure 4 The desired alignment between the server VSYNC signal and the client VSYNC signal is illustrated. In one embodiment, such as in a server and client network configuration, tuning of the server VSYNC signal 311 is performed to obtain proper alignment between the server VSYNC signal and the client VSYNC signal. In another embodiment, such as in a multi-tenant server to multiple client network configuration, tuning of the client VSYNC signal 312 is performed to obtain proper alignment between the server VSYNC signal and the client VSYNC signal. For illustrative purposes, Figure 4The document describes tuning the server VSYNC signal 311 to synchronize the frequencies of the server and client VSYNC signals, and / or adjust the timing offset between the corresponding client VSYNC signal and the server VSYNC signal. However, it should be understood that the client VSYNC signal 312 can also be used for tuning. In the context of this patent, "synchronization" should be understood as tuning the signals so that their frequencies match, but their phases may differ; "offset" should be understood as the time delay between signals, for example, the time between when one signal reaches its maximum value and when another signal reaches its maximum value.

[0060] As shown in the figure Figure 4 An improved process is illustrated in an embodiment of this disclosure for executing a video game at a server to generate rendered video frames and sending those video frames to a client for display. The process is illustrated with respect to generating and displaying individual video frames at both the server and client. Specifically, the server generates the game-rendered video frames at 401. For example, server 260 includes a CPU (e.g., game title processing engine 211) configured to execute the game. The CPU generates one or more draw calls for the video frames, wherein the draw calls include commands placed in a command buffer for execution in the graphics pipeline by the corresponding GPU of server 260. The graphics pipeline may include one or more shader programs on the vertices of objects within the scene for generating texture values ​​for rendering the video frames, wherein the operations are performed in parallel by the GPU for efficiency. At flip time 409, the GPU arrives at a flip command in the command buffer, the flip command indicating that the corresponding video frame has been fully generated and / or rendered and placed in the frame buffer at server 260.

[0061] At 402, the server performs a scan output of the game-rendered video frames to the encoder. Specifically, the scan output is performed scan-line by scan or in groups of consecutive scan lines, where a scan line refers to a single horizontal line from one edge of the screen to the other, for example, on a display. These scan lines or groups of consecutive scan lines are sometimes referred to as slices and are referred to as screen slices in this specification. Specifically, scan output 402 may include several processes that modify the game-rendered frames, including overlaying the game-rendered frames with another frame buffer or shrinking the game-rendered frames so that they are surrounded by information from another frame buffer. During scan output 402, the modified video frames are then scanned into the encoder for compression. In one embodiment, scan output 402 is performed at the occurrence 311a of the VSYNC signal 311. In other embodiments, scan output 402 may be performed before the occurrence of the VSYNC signal 311, for example, at flip time 409.

[0062] At 403, the video frames rendered by the game (which may have been modified) are encoded slice by slice at the encoder to generate one or more encoded slices, wherein the encoded slices are independent of scan lines or screen slices. Therefore, the encoder generates one or more encoded (e.g., compressed) slices. In one embodiment, the encoding process begins before the scan output 402 process for the corresponding video frame has been fully completed. Furthermore, the start and / or end of encoding 403 may or may not be aligned with the server VSYNC signal 311. The boundary of the encoded slice is not limited to a single scan line and may include a single scan line or multiple scan lines. Furthermore, the end of the encoded slice and / or the start of the next encoder slice may not necessarily occur at the edge of the display (e.g., it may occur somewhere in the middle of the screen or in the middle of a scan line), such that the encoded slice does not need to traverse completely from the edge of the display screen to the edge. As shown, one or more encoded slices may be compressed and / or encoded, including a compressed “encoded slice A” with a hash marker.

[0063] At 404, the encoded video frame is transmitted from the server to the client, wherein the transmission can be performed slice by slice, each slice being a compressed encoder slice. In one embodiment, the transmission process 404 begins before the encoding process 403 for the corresponding video frame has been fully completed. Furthermore, the start and / or end of transmission 404 may or may not be aligned with the server's VSYNC signal 311. As shown, the compressed encoded slice A is transmitted to the client independently of other compressed encoder slices of the rendered video frame. Encoder slices can be transmitted one at a time or in parallel.

[0064] At 405, the client again receives the compressed video frame slice by slice. Furthermore, the start and / or end of reception 405 may or may not be aligned with the client's VSYNC signal 312. As shown, the client receives the compressed, encoded slice A. Transmission jitter 452 may exist between server 260 and client 210, where jitter 452 measures the variation in network latency from server 260 to client 210. Lower jitter values ​​indicate a more stable connection. As shown, the thick, straight arrows indicate the current latency when the corresponding video frame is sent to client 210, but due to jitter, the video frame at client 210 may have an arrival time range (e.g., the range defined by the dotted arrows). The latency variation may also be due to one or more operations at the server, such as encoding 403 and transmission 404, and network problems that introduce latency when transmitting the video frame to client 210.

[0065] At 406, the client again decodes the compressed video frame slice by slice, producing a decoded slice A (shown without hash markers) ready for display. In one embodiment, the decoding process 406 begins before the reception process 405 for the corresponding video frame is fully completed. Furthermore, the start and / or end of decoding 406 may or may not be aligned with the client's VSYNC signal 312. At 407, the client displays the decoded rendered video frame on its display. That is, for example, the decoded video frame is placed in a display buffer that streams to the display device scan-line by scan. In one embodiment, the display process 407 (i.e., the stream output to the display device) begins after the decoding process 406 for the corresponding video frame is fully completed, i.e., the decoded video frame is fully residing in the display buffer. In another embodiment, the display process 407 begins before the decoding process 406 for the corresponding video frame is fully completed. That is, the stream output to the display device begins at the address of the display buffer while only a portion of the decoded frame buffer resides in the display buffer. Then, the display buffer is updated or filled with the remaining portion of the corresponding video frame in a timely manner for display, so that the update of the display buffer is performed before those portions are output to the display stream. Furthermore, the start and / or end of the display 407 are aligned with the client's VSYNC signal 312.

[0066] In one embodiment, the one-way delay 416 between server 260 and client 210 can be defined as the elapsed time between the start of scan output 402 and the start of display 407. Embodiments of this disclosure can align the VSYNC signals between the server and client (e.g., synchronize frequencies and adjust offsets) to reduce the one-way delay between the server and client, and to reduce the variability of the one-way delay between the server and client. For example, embodiments of this disclosure can calculate the optimal adjustment of the offset 430 between the server VSYNC signal 311 and the client VSYNC signal 312, such that even under near-worst-case conditions such as the time required for server processing (encoding 403 and transmission 404), near-worst-case network latency between server 260 and client 210, and near-worst-case conditions such as near-worst-case client processing (reception 405 and decoding 406), the decoded rendered video frames can be used in time for display process 407. That is, it is not necessary to determine the absolute offset between the server VSYNC and the client VSYNC; adjusting the offset so that the decoded rendered video frames can be used in time for display is sufficient.

[0067] Specifically, the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 can be aligned through synchronization. Synchronization is achieved by tuning either the server VSYNC signal 311 or the client VSYNC signal 312. For illustrative purposes, tuning is described in relation to the server VSYNC signal 311, but it should be understood that tuning the client VSYNC signal 312 can be performed alternatively. For example, as... Figure 4 As shown, the server frame period 410 (e.g., the time between two occurrences 311c and 311d of the server VSYNC signal 311) is substantially equal to the client frame period 415 (e.g., the time between two occurrences 312a and 312b of the client VSYNC signal 312), which indicates that the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 are also substantially equal.

[0068] To maintain frequency synchronization between the server and client VSYNC signals, the timing of the server VSYNC signal 311 can be manipulated. For example, the vertical blanking interval (VBI) in the server VSYNC signal 311 can be increased or decreased over a period of time to account for drift between the server VSYNC signal 311 and the client VSYNC signal 312. Manipulating the vertical blanking (VBLANK) lines in the VBI provides adjustment of the number of scan lines used for VBLANK during one or more frame periods of the server VSYNC signal 311. Discarding the number of scan lines for VBLANK reduces the corresponding frame period (e.g., time interval) between two occurrences of the server VSYNC signal 311. Conversely, increasing the number of scan lines for VBLANK increases the corresponding frame period (e.g., time interval) between two occurrences of the VSYNC signal 311. In that way, the frequency of the server VSYNC signal 311 is adjusted so that the frequencies of the client VSYNC signal 311 and the server VSYNC signal 312 are aligned to substantially the same frequency. Furthermore, the offset between the server's VSYNC signal and the client's VSYNC signal can be adjusted by briefly increasing or decreasing the VBI, and then the VBI can be returned to its original value. In one embodiment, the server's VBI is adjusted. In another embodiment, the client's VBI is adjusted. In yet another embodiment, instead of two devices (server and client), there are multiple connected devices, each of which may have a corresponding VBI that has been adjusted. In one embodiment, each of the multiple connected devices may be an independent peer device (e.g., without a server device). In another embodiment, the multiple devices may include one or more server devices and / or one or more client devices arranged in one or more server / client architectures, multi-tenant server / client architectures, or a combination thereof.

[0069] Alternatively, in one implementation, the server's pixel clock (e.g., located at the southbridge of the server's northbridge / southbridge core logic chipset, or, in the case of a discrete GPU, generating its own pixel clock using its own hardware) can be manipulated to perform coarse and / or fine adjustments to the frequency of the server VSYNC signal 311 over a period of time, so that the frequency synchronization between the server VSYNC signal 311 and the client VSYNC signal 312 is brought back into alignment. Specifically, the pixel clock in the server's southbridge can be overclocked or underclocked to adjust the overall frequency of the server's VSYNC signal 311. In that way, the frequency of the server VSYNC signal 311 is adjusted so that the frequencies of the client VSYNC signal 311 and the server VSYNC signal 312 are aligned to substantially the same frequency. The offset between the server VSYNC and the client VSYNC can be adjusted by increasing or decreasing the client-server pixel clock over a short period of time, and then the pixel clock is brought back to its original value. In one implementation, the server pixel clock is adjusted. In another implementation, the client pixel clock is adjusted. In another embodiment, instead of two devices (server and client), there are multiple connected devices, each of which may have an adjusted corresponding pixel clock. In one embodiment, each of the multiple connected devices may be an independent peer device (e.g., without a server device). In another embodiment, the multiple connected devices may include one or more server devices and one or more client devices arranged in one or more server / client architectures, multi-tenant server / client architectures, or a combination thereof.

[0070] In one implementation, high-performance codecs (e.g., encoders and / or decoders) can be used to further reduce one-way latency between the cloud gaming server and the client. In conventional streaming systems involving compressed media (e.g., streaming movies, TV shows, videos, etc.), a large amount of decompressed video may be buffered at the client to accommodate variations in encoding operations (e.g., longer encoding times), variations in transmission quality intrusion jitter, and variations in decoding operations (e.g., longer decoding times) when the streaming media is decompressed at the end target (e.g., the client). Therefore, in conventional streaming systems, it is possible to rely on average decoding capabilities and metrics (e.g., average decoding resources) because the buffering of the decoded content adapts to latency variability, allowing video frames to be displayed at the desired rate (e.g., supporting 4K media at 60Hz, or displaying video frames each time a client VSYNC signal is received).

[0071] However, buffering is extremely limited in cloud gaming environments (e.g., tending towards zero buffering), making real-time gaming difficult. Therefore, any variability introduced into the one-way latency between the cloud gaming server and client can adversely affect downstream operations. For example, spending longer times encoding and / or decoding complex frames, even for a single frame, results in higher one-way latency, ultimately increasing response time for the user and negatively impacting the user's real-time experience.

[0072] In one implementation, for cloud gaming, it is beneficial to provide more robust decoding and encoding resources that appear unnecessary compared to the needs of streaming video applications. Furthermore, encoder resources should be optimized for the time required to process the longest or most frequently processed frames, as will be described more fully below. That is, in one implementation, the encoder can be tuned to improve the trade-off between one-way latency and video quality in the cloud gaming system, wherein encoder tuning can be based on monitoring client bandwidth, skipped frames, the number of encoded I-frames, the number of scene changes, and / or the number of video frames exceeding the target frame size. The tuned parameters may include encoder bit rate, target frame size, maximum frame size, and quantization parameter (QP) values, where high-performance encoders and decoders help reduce the total one-way latency between the cloud gaming server and the client.

[0073] Through the Figures 2A to 2D Detailed description of various client devices 210 and / or cloud gaming networks 290 (e.g., in game server 260), Figure 5 Flowchart 500 illustrates a method for cloud gaming according to one embodiment of the present disclosure, wherein encoding video frames includes tuning encoder parameters with knowledge of network transmission speed and reliability, as well as a total latency target. A cloud gaming server is configured to stream content over a network to one or more client devices. This process provides smoother frame rates and more reliable latency, reducing and making more consistent one-way latency between the cloud gaming server and the client, thereby improving the smoothness of video display on the client.

[0074] At point 510, multiple video frames are generated when a video game is executed on a cloud gaming server. Generally, a cloud gaming server generates multiple video frames for game rendering. For example, the game logic of a video game is built on top of a game engine or game title processing engine. The game engine includes core functionalities that the game logic can use to construct the game environment of the video game. For example, some functionalities of the game engine may include a physics engine for simulating physical forces and collisions on objects in the game environment, a rendering engine for 2D or 3D graphics, collision detection, sound, animation, artificial intelligence, connectivity, streaming, etc. In that way, the game logic does not have to build the core functionalities provided by the game engine from scratch.

[0075] The game logic and game engine are executed by a combination of CPU and GPU, where the CPU and GPU can be configured within an Accelerated Processing Unit (APU). That is, the CPU and GPU, along with shared memory, can be configured as a rendering pipeline for generating video frames for game rendering, such that the rendering pipeline outputs game-rendered images as video or image frames suitable for display, including corresponding color information for each of the pixels in the target and / or virtualized display. Specifically, the CPU can be configured to generate one or more draw calls for video frames, each draw call including commands stored in a corresponding command buffer and executed by the GPU in the GPU pipeline. Generally, the graphics pipeline performs shader operations on the vertices of objects within the scene to generate texture values ​​for the pixels of the display. Specifically, the graphics pipeline receives input geometry (e.g., vertices of objects in the game environment), and vertex shaders construct primitives or polygons that constitute the objects. Vertex shader programs can perform lighting, shading, shading, and other operations on the primitives. Depth or z-buffering is performed to determine which objects are visible when rendered from the corresponding viewpoint. Rasterization is performed to project objects in the 3D game environment onto a 2D plane defined by the viewpoint. Pixel-sized fragments of the object are generated, where one or more fragments can contribute to the color of pixels in the image. Fragments may be merged and / or blended to determine the combined color of each of the pixels in the corresponding video, and may be stored in a frame buffer. Subsequent video frames are generated and / or rendered for display using a command buffer with a similar configuration, where multiple video frames are output from the GPU pipeline.

[0076] At 520, the method includes encoding multiple video frames at a certain encoder bit rate. Specifically, the multiple video frames are scanned into an encoder for compression and then streamed to a client using a streaming transmitter operating at the application layer. In one embodiment, each of the game-rendered video frames can be synthesized and blended with additional user interface features to form a corresponding modified video frame, which is then scanned into an encoder, where the encoder compresses the modified video frame for streaming to the client. For simplicity and clarity, Figure 5The method for tuning encoder parameters disclosed herein is described with reference to encoding multiple video frames, but should be understood as supporting the encoding of modified video frames. The encoder is configured to compress multiple video frames based on the described format. For example, when streaming media content from a cloud gaming server to a client, the Moving Picture Experts Group (MPEG) or H.264 standards may be implemented. Specifically, the encoder can perform compression by video frames or by encoder slices of video frames, where each video frame can be compressed into one or more coded slices, as previously described. Generally, when streaming media, video frames can be compressed into I-frames (intra-frames) or P-frames (predictive frames), each of which can be segmented into coded slices.

[0077] At 530, the maximum received bandwidth of the client is measured. In one implementation, the maximum bandwidth experienced by the client is determined through a feedback mechanism from the client. Figure 6 The illustration depicts a streaming transmitter 620 of a cloud gaming server measuring the bandwidth of client 210 according to one embodiment of the present disclosure. The streaming transmitter 620 is configured to monitor and tune encoder 610 such that compressed video frames can be transmitted at a rate within the client's measured bandwidth. As shown, compressed video frames, encoded slices, and / or packets are delivered from encoder 610 to buffer 630 (e.g., FIFO). The encoder delivers the compressed video frames at encoder fill rate 615. For example, the buffer can be filled as quickly as the encoder can generate compressed video frames, encoded slices 650, and / or encoded slice packets 655. Furthermore, compressed video frames are ejected from the buffer at buffer eject rate 635 for delivery to client 210 via network 250. In one embodiment, the buffer eject rate 635 is dynamically tuned based on the client's measured maximum receive bandwidth. For example, the buffer eject rate 635 can be adjusted to be approximately equal to the client's measured maximum receive bandwidth. In one implementation, packet encoding is performed at the same rate as the transmitted packets, such that both operations are dynamically tuned based on the maximum available bandwidth available to the client.

[0078] Specifically, the streaming transmitter 620, operating at the application layer, uses, for example, a bandwidth tester 625 to measure the maximum bandwidth of the client 210. This application layer is used within the User Datagram Protocol / Internet Protocol (UDP / IP) suite of protocols used to interconnect network devices via the Internet. For example, the application layer defines communication protocols and interface methods for communication between devices via IP communication networks. During testing, the streaming transmitter 620 provides additional buffered packets 640 (e.g., forward error correction-FEC packets) so that the buffer 630 can stream packets at a predefined bit rate (such as the maximum bandwidth being tested). In one embodiment, the client returns feedback 690 to the streaming transmitter 620 the number of packets it has received within an incremental sequence identifier (ID) range, such as a range of video frames. For example, the client might report something such as receiving 145 video frames out of 150 video frames for sequence IDs 100 to 250 (e.g., 150 video frames). Therefore, the streaming device 620 at server 260 can calculate the packet loss rate, and since the streaming device 620 knows the amount of bandwidth sent (e.g., as a test) during that packet sequence, it can dynamically determine the client's maximum bandwidth at a given moment. The client's measured maximum bandwidth can be delivered from the streaming device 620 to the buffer 630 as control information 627, allowing the buffer 630 to dynamically transmit packets at a rate approximately equal to the client's maximum bandwidth. Thus, the transmission rate of compressed video frames, encoded slices, and / or packets can be dynamically adjusted based on the client's currently measured maximum bandwidth.

[0079] At 540, the encoding process is monitored by a streaming transceiver. That is, the encoding of multiple video frames is monitored. In one embodiment, monitoring is performed at client 210, where feedback and / or tuning control signals are provided back to the encoder. In another embodiment, the monitoring is performed at cloud gaming server 260, such as by streaming transceiver 620. For example, monitoring of the encoding of video frames may be performed by monitoring and tuning unit 629 of streaming transceiver 620. Various encoding characteristics and / or operations can be tracked and / or monitored. For example, in one embodiment, the occurrence rate of I-frames within multiple video frames can be tracked and / or monitored. Furthermore, in one embodiment, the occurrence rate of scene changes within multiple video frames can be tracked and / or monitored. Moreover, in one embodiment, the number of video frames exceeding the target frame size can be tracked and / or monitored. Moreover, in one embodiment, the encoder bit rate used to encode one or more video frames can be tracked and / or monitored.

[0080] At 550, encoder parameters are dynamically tuned based on monitoring of video frame encoding. That is, monitoring of video frame encoding affects how the encoder operates when compressing current and future video frames received at the encoder. Specifically, the monitoring and tuning unit 629 is configured to determine which encoder parameters should be tuned in response to monitoring of video frame encoding and analysis of the monitored information. A control signal 621 is delivered from the monitoring and tuning unit 629 back to the encoder 610, used to configure the encoder. Encoder parameters used for tuning include quantization parameters (QP) (e.g., minQP, maxQP) or quality parameters, target frame size, maximum frame size, etc.

[0081] Tuning is performed with an understanding of network transmission speed and reliability, as well as overall latency targets. In one implementation, smoothness of video playback is prioritized over low latency or image quality. For example, skipping encoding of one or more video frames is disabled. Specifically, various encoder parameters are used to tune the balance between image resolution or quality (e.g., at 60Hz) and latency. Specifically, because the VSYNC signal at the cloud gaming server and client can be synchronized and offset, one-way latency between the cloud gaming server and client can be reduced, thereby reducing the need for skipped video frames to facilitate low latency. Synchronization and offset of the VSYNC signal also provide overlapping operations (scan output, encoding, and transmission) at the cloud gaming server; overlapping operations (reception, decoding, rendering, and display) at the client; and / or overlapping operations at the cloud gaming server and client, all of which contribute to reduced one-way latency, reduced variability in one-way latency, real-time generation and display of video content, and consistent video playback at the client.

[0082] In one implementation, the encoder bit rate is monitored to anticipate client bandwidth demands, taking into account upcoming frames and their complexity (e.g., predicted scene changes), and the encoder bit rate can be adjusted according to the anticipated demand. For example, when prioritizing smoothness in video playback, the encoder monitoring and tuning unit 629 can be configured to determine that the encoder bit rate used exceeds the measured maximum receive bandwidth. In response, the encoder bit rate can be reduced, which also reduces the frame size. When prioritizing smoothness, it is desirable to use an encoder bit rate lower than the maximum receive bandwidth (e.g., for a maximum receive bandwidth of 15 megabits per second, the encoder bit rate is 10 megabits per second). In that way, if the encoded frame rate spikes sharply above the maximum frame size, the encoded frames can still be transmitted within 60 Hz (Hertz). Specifically, the encoder bit rate can be converted to frame size. A given bit rate and target speed for a video game (e.g., 60 frames per second) will be converted to the average size of the encoded video frames. For example, at an encoder bit rate of 15 megabits per second and a given target speed of 60 frames per second, where 60 encoded frames share 15 megabits, each encoded frame has approximately 250k encoded bits. Therefore, controlling the encoder bit rate will also control the frame size of the encoded video frames, so that increasing the encoder bit rate provides more bits for encoding (higher precision), while decreasing the encoder bit rate provides fewer bits for encoding (lower precision). Similarly, when the encoder bit rate used to encode a set of video frames is within the measured maximum receiving bandwidth, increasing the encoder bit rate can also increase the frame size.

[0083] In one implementation, when prioritizing smoothness in video playback, the encoder monitoring and tuning unit 629 can be configured to determine that the encoder bit rate used to encode a set of video frames from multiple video frames exceeds a measured maximum receive bandwidth. For example, the encoder bit rate may be detected as 15 megabits per second (Mbps), while the maximum receive bandwidth may currently be 10 Mbps. In this way, the encoder pushes out more bits than can be transmitted to the client without increasing one-way latency. As previously described, when prioritizing smoothness, it may be necessary to use an encoder bit rate lower than the maximum receive bandwidth. In the example above, setting the encoder bit rate to equal to or lower than 10 megabits per second is acceptable for the maximum receive bandwidth of 10 megabits per second described above. In this way, if the number of encoded frames increases dramatically to above the maximum frame size, the encoded frames can still be transmitted within 60 Hz. In response, the QP value, where QP controls the precision used when compressing video frames, can be tuned with or without reducing the encoder bit rate. In other words, QP controls how much quantization is performed (e.g., compressing a variable range of values ​​in a video frame into a single quantum value). In H.264, the range of QP is "0" to "51". For example, a QP value of "0" means less quantization, less compression, higher precision, and higher quality. A QP value of "51" means more quantization, more compression, lower precision, and lower quality. Specifically, increasing the QP value can result in encoding video frames with lower precision.

[0084] In one implementation, when smoothness of video playback is prioritized, encoder monitoring performed by the monitoring and tuning unit 629 can be configured to determine that the encoder bit rate used to encode a group of video frames from multiple video frames is within the maximum receive bandwidth. As previously described, when smoothness is prioritized, an encoder bit rate lower than the maximum receive bandwidth may be required. Therefore, there is excess available bandwidth when transmitting the group of video frames. This excess bandwidth can be determined. In response, the QP value can be tuned, where QP controls the precision used when compressing video frames. Specifically, the QP value can be reduced based on the excess bandwidth, allowing encoding to be performed with higher precision.

[0085] In another implementation, the characteristics of the individual video game are considered when determining I-frame handling and QP settings, particularly when biased towards smooth video playback. For example, if the video game has infrequent "scene changes" (e.g., only camera cuts), it may be desirable to allow I-frames to be larger (low QP or higher encoder bit rate). That is, within a set of video frames from multiple compressed video frames, the number of video frames identified as having scene changes is determined to be below a threshold number of scene changes. In other words, the streaming system can handle said number of scene changes under current conditions (e.g., measured client bandwidth, required latency, etc.). In response, the QP value can be tuned, where QP controls the precision used when compressing video frames. Specifically, the QP value can be decreased so that encoding is performed with higher precision.

[0086] On the other hand, if a video game has frequent "scene changes" during gameplay, it is desirable to keep the I-frame size small (e.g., a high QP or low encoder bit rate). That is, within a set of video frames from multiple compressed video frames, the number of video frames identified as having scene changes meets or exceeds a threshold number for scene changes. In other words, the video game generates too many scene changes under current conditions (e.g., measured client bandwidth, required latency, etc.). In response, the QP value can be tuned, where QP controls the precision used when compressing video frames. Specifically, the QP value can be increased to perform encoding with lower precision.

[0087] In another implementation, the encoding mode can be considered when determining I-frame handling and QP settings, particularly when biased towards smooth video playback. For example, if the encoder does not frequently generate I-frames, it may be desirable to allow I-frames to be larger (low QP or higher encoder bit rate). That is, within a set of video frames from multiple compressed video frames, the number of video frames compressed into I-frames falls within or below a threshold number of I-frames. In other words, the streaming system can handle this number of I-frames under current conditions (e.g., measured client bandwidth, required latency, etc.). In response, the QP value can be tuned, where QP controls the precision used when compressing video frames. Specifically, the QP value can be decreased to perform encoding with higher precision.

[0088] If the encoder generates I-frames frequently, it is desirable to keep the I-frame size small (e.g., high QP, or a low encoder bit rate). That is, within a set of video frames from multiple compressed video frames, the number of video frames compressed into I-frames meets or exceeds a threshold number of I-frames. In other words, the video game generates too many I-frames under current conditions (e.g., measured client bandwidth, required latency, etc.). In response, the QP value can be tuned, where QP controls the precision used when compressing video frames. Specifically, the QP value can be increased to perform encoding with lower precision.

[0089] In another implementation, the encoding mode can be considered when rotating the encoder, especially when biased towards smooth video playback. For example, if the encoder is frequently below the target frame size, it may be desirable to allow the target frame size to become larger. That is, within a set of video frames from multiple compressed video frames transmitted at a certain transmission rate, the number of video frames is determined to be below a threshold. Each of these video frames is within the target frame size (i.e., equal to or less than the target frame size). In response, at least one of the target frame size and the maximum frame size is increased.

[0090] On the other hand, if the encoder size is frequently larger than the target frame size, it may be desirable to allow the target frame size to become smaller. That is, within a set of video frames from multiple compressed video frames transmitted at a certain transmission rate, the number of video frames is determined to meet or exceed a threshold. Each of these video frames exceeds the target frame size. In response, at least one of the target frame size and the maximum frame size is reduced.

[0091] Figure 7A This is a graph illustrating the setting of the encoder's quantization parameters (QP) according to one embodiment of this disclosure to optimize quality and buffer utilization at the client end. Graph 700A shows the frame size (in bytes) of each generated frame in the vertical direction, as shown in the horizontal direction. The target frame size and the maximum frame size are static. Specifically, line 711 shows the maximum frame size, and line 712 shows the target frame size, where the maximum frame size is greater than the target frame size. As shown in graph 700A, there are multiple peaks including compressed video frames exceeding the target frame size of line 712. Video frames exceeding the target frame size risk introducing playback jitter (e.g., increased one-way latency) because they may require more than one frame cycle for encoding and / or transmission from the cloud gaming server.

[0092] Graph 700B illustrates the encoder response after QP has been set based on the target frame size, maximum frame size, and QP range (e.g., minQP and maxQP) to optimize encoding quality and buffer utilization at the client. For example, QP can be adjusted and / or tuned based on encoder monitoring of encoder bit rate, scene change frequency, and I-frame generation frequency, as previously described. Graph 700B shows the frame size (in bytes) of each generated frame in the vertical direction, as shown in the horizontal direction. The target frame size at line 712 and the maximum frame size at line 711 remain at the same positions as in graph 700A. After QP tuning and / or adjustment, the number of peaks including compressed video frames exceeding the target frame size at line 712 is reduced compared to graph 700A. That is, QP has been tuned to optimize the encoding of video frames (i.e., falling within the target frame size) for current conditions (e.g., measured client bandwidth, required latency, etc.).

[0093] Figure 7B This is a diagram illustrating the tuning of the target frame size, maximum frame size, and / or QP (e.g., minQP and / or maxQP) encoder settings according to one embodiment of this disclosure to reduce the occurrence of I-frames exceeding the true target frame size supported by the client. For example, the QP can be adjusted and / or tuned based on encoder monitoring of encoder bit rate, scene change frequency, and I-frame generation frequency, as previously described.

[0094] Graph 720A shows the frame size (in bytes) of each generated frame in the vertical direction, as shown in the horizontal direction. For illustrative purposes, Figure 7B The curves of 720A and Figure 7A The curve 700A reflects similar encoder conditions and is used for encoder tuning. In curve 720A, the target frame size and maximum frame size are static. Specifically, line 711 shows the maximum frame size, and line 712 shows the target frame size, where the maximum frame size is higher than the target frame size. As shown in curve 720A, there are multiple peaks including compressed video frames exceeding the target frame size at line 712. Video frames exceeding the target frame size risk introducing playback jitter (e.g., increased one-way latency) because they may require more than one frame cycle for encoding and / or transmission from the cloud gaming server. For example, the peak reaching the maximum frame size at line 711 could be an I-frame that takes more than 16 ms to send to the client, causing playback jitter due to increased one-way latency between the cloud gaming server and the client.

[0095] Figure 720B illustrates the encoder response after at least one of the target frame size and / or maximum frame size has been tuned to reduce the occurrence of I-frames exceeding the true target frame size supported by the client. The true target frame size may be adjusted based on measured client bandwidth and / or encoder monitoring, including monitoring of encoder bit rate, frequency of scene changes, and frequency of I-frame generation, as previously described.

[0096] Graph 720B shows the frame size (in bytes) of each generated frame in the vertical direction, as shown in the horizontal direction. Compared to graph 720A, the values ​​of the target frame size at line 712' and the maximum frame size at line 711' have been reduced. For example, the target frame size at line 712' has been reduced from line 712, and the maximum frame size at line 711' has been reduced from line 711. After tuning the target frame size and / or the maximum frame size, the maximum size of the peak of compressed video frames exceeding the target frame size 712' is reduced for better transmission. Furthermore, compared to graph 700A, the number of peaks including compressed video frames exceeding the target frame size 712' has also been reduced. For example, only one peak is shown in graph 720B. That is, the target frame size and / or the maximum frame size have been tuned to optimize the encoding of video frames (i.e., falling within the target frame size) for current conditions (e.g., measured client bandwidth, required latency, etc.).

[0097] Through the Figures 2A to 2D Detailed description of various client devices 210 and / or cloud gaming networks 290 (e.g., in game server 260), Figure 8 Flowchart 800 illustrates a method for cloud gaming according to one embodiment of the present disclosure, wherein encoding video frames includes determining when to skip video frames or delay the encoding and transmission of video frames when the encoding run time is long or when the generated video frames are large (such as when encoding I-frames). Specifically, the decision to skip video frames is made based on an understanding of network transmission speed and reliability, as well as a total latency target. This process provides smoother frame rates and more reliable latency, thereby reducing and making more consistent one-way latency between the cloud gaming server and the client, and thus improving the smoothness of video display on the client.

[0098] At point 810, multiple video frames are generated when a video game is executed on a cloud gaming server operating in streaming mode. Generally, cloud gaming servers generate multiple video frames for game rendering. For example, already... Figure 5 At point 510, the generation of video frames for game rendering is described, and the generation is applicable to... Figure 8The generation of video frames in a video game. For example, the game logic of a video game is built on top of a game engine or game title processing engine. The combination of game logic and game engine is executed by the CPU and GPU, where the CPU and GPU, along with shared memory, can be configured as a rendering pipeline for generating video frames for game rendering, such that the rendering pipeline outputs game-rendered images as video or image frames suitable for display, including the corresponding color information of each of the pixels in the target and / or virtualized display.

[0099] At 820, a scene change in the first video frame of the video game is predicted. This scene change is predicted before the first video frame is generated. In one implementation, the game logic may be made aware of the scene change while the CPU is executing the video game. For example, the game logic or additional logic may include code that predicts the scene change when a video frame is generated (e.g., scene change logic), such as predicting that a series of video frames includes at least one scene change, or predicting that a particular video frame is a scene change. Specifically, the game logic or additional logic configured to perform scene change prediction analyzes game state data collected during the execution of the video game to determine and / or anticipate and / or predict when a scene change will occur, such as in the next X number of frames (e.g., a range) or for the identified video frames. For example, a scene change may be predicted when a character moves from one scene to another in a virtual game environment, or when a character has finished a level and transitions to another level in the video game, or when transitioning between scenes between two video frames (e.g., a scene change in a set of shots in a movie, or starting an interactive game after a series of menus), etc. Scene changes can be represented by video frames comprising large and complex scenes of a virtual game world or environment.

[0100] Game state data defines the state of the game at a given time and can include game characters, game objects, game object attributes, game attributes, game object states, graphical overlays, the character's position within the game world, the game scene or environment, the level of the game application, the character's assets (e.g., weapons, tools, bombs, etc.), equipment, the character's skill set, game levels, character attributes, character position, remaining lives, total possible available lives, armor, trophies, time counter values, and other asset information. In this way, game state data allows for the generation of the game environment existing at a corresponding point in the video game.

[0101] At 830, a scene change cue is generated and sent to the encoder, wherein the cue indicates that the first video frame is a scene change. This provides the encoder with notification of an upcoming scene change, allowing the encoder to adjust its encoding operations when compressing the identified video frames. The notification provided as a scene change cue can be delivered via an API for communication between components or between applications running on components of the cloud gaming server 260. In one embodiment, the API may be a GPU API. For example, the API may run on or be invoked by game logic and / or additional logic configured to detect scene changes to communicate with the encoder. In one embodiment, the scene change cue may be provided as a data control packet formatted such that all components receiving the data control packet can understand what type of information is included in the data control packet and understand the appropriate reference to the corresponding rendered video frame. In one embodiment, the communication protocol for the API and the formatting for the data control packet may be defined in the corresponding software development kit (SDK) for the video game.

[0102] At 840, the first video frame is delivered to the encoder. As previously described, game-generated video frames can be synthesized and blended with additional user interface features to form a modified video frame, which is then scanned to the encoder. The encoder is configured to compress the first video frame based on a desired format, such as MPEG or H.264 standards used for streaming media content from a cloud gaming server to a client. During streaming, video frames are encoded as P-frames until a scene change occurs or a keyframe is no longer referenced in the currently encoded frame (e.g., the previous I-frame), at which point the next video frame is then encoded as another I-frame. In this case, the first video frame is encoded as an I-frame based on a scene change cue, where the I-frame can be encoded without referencing any other video frames (e.g., independently as a keyframe).

[0103] At 850, the maximum received bandwidth of the client is measured. As previously described, the maximum bandwidth experienced by the client can be determined through a feedback mechanism from the client, such as... Figure 5 and Figure 6 As illustrated in operation 530. Specifically, the streaming server's streamer can be configured to measure the client's bandwidth.

[0104] At 860, the encoder receives the second video frame. That is, the second video frame is received after a scene change, and compressed after the first video frame has been compressed. Furthermore, the encoder decides whether to either not encode the second video frame (or subsequent video frames) or delay encoding the second video frame (or subsequent video frames). This decision is based on the client's maximum receiving bandwidth and the target resolution of the client's display. In other words, the decision to skip or delay encoding takes into account the client's available bandwidth. Generally, if the current bandwidth experienced by the client is sufficient to allow video frames generated and encoded for the client's target display to quickly return to low one-way latency after a latency hit (e.g., generating a large I-frame for a scene change), the second video frame (and / or subsequent video frames) may still be delayed in encoding. On the other hand, if the current bandwidth experienced by the client is insufficient, the second video frame (and / or subsequent video frames) may be skipped during the encoding process and not delivered to the client. Therefore, if the client's bandwidth exceeds the bandwidth required to support the target resolution of the display at the client, it is possible to have fewer skipped frames (and lower latency).

[0105] In one implementation, compressed video frames are transmitted from a server to a client via the network at a rate based on the maximum available bit rate or bandwidth at specific times. Therefore, the transmission rate of encoded slices and / or packets of encoded video frames is dynamically adjusted based on the currently measured maximum bandwidth. The video frame can be transmitted while it is being encoded, allowing transmission to occur immediately upon completion of encoding, without waiting for the next occurrence of the server's VSYNC signal and without waiting for the entire video frame to be encoded.

[0106] Additionally, in one implementation, packet encoding is performed at the same rate as the transmitted packets, such that both operations are dynamically tuned based on the maximum available bandwidth for the client. Furthermore, the encoder bit rate is monitored to anticipate client bandwidth demands, taking into account upcoming frames and their complexities (e.g., predicted scene changes), and the encoder bit rate can be adjusted according to these anticipated demands. Moreover, the encoder bit rate can be transmitted to the client, allowing the client to adjust its decoding speed accordingly to match the encoder bit rate.

[0107] In one implementation, when the transmission rate to the client is low for the target resolution of the client's display, the encoder skips the second video frame. That is, the second video frame is not encoded. Specifically, a set of compressed video frames is being transmitted to the client at a rate exceeding the maximum receiving bandwidth. For example, the transmission rate to the client might be 15 megabits per second (Mbps), but the client's measured receiving bandwidth might currently be between 5 Mbps and 10 Mbps. Thus, if all video frames are continuously pushed to the client, the one-way latency between the cloud gaming server and the client increases. To promote low latency, the encoder may skip the second and / or subsequent video frames.

[0108] Figure 9A A sequence of video frames 900A compressed by an encoder according to one embodiment of this disclosure is illustrated, wherein when the client bandwidth is low for the target resolution of the client's display, the encoder abandons encoding a second video frame 920 after encoding a first I-frame 905. Encoded blocks and transport blocks of the video frames are shown relative to the VSYNC signal 950. Specifically, if no spare bandwidth is available, a longer-encoded I-frame will result in one or more skipped frames in an attempt to maintain low one-way latency priority, where one-way latency may include the time it takes to display the video frame at the client. As shown, skipping one or more video frames after an I-frame allows an immediate return to low one-way latency (e.g., within one or two frame cycles). Otherwise, due to not skipping the encoding of the video frame, it would take several frame cycles to return to low one-way latency.

[0109] For example, video frame sequence 900A includes one encoded I-frame 905, while the remaining frames are encoded as P-frames. For illustration, coded blocks 901 and 902, which are P-frames, are encoded before coded block 905 is encoded as an I-frame. The encoder then compresses the video frames into P-frames until the next scene change, or until the video frame can no longer reference a previous keyframe (e.g., an I-frame). Typically, the encoding time for an I-frame block can be longer than that for a P-frame block. For example, the encoding time for I-frame block 905 can exceed one frame period. In some cases, the encoding time between P-frames and I-frames can often be roughly the same, especially when using a powerful encoder.

[0110] However, the transmission time between I-frames and P-frames differs significantly. As shown in the figure, various transmission times are illustrated relative to the corresponding encoded video frames. For example, transmission block 911 of encoded P-frame block 901 is shown to have low latency, allowing both encoding block 901 and transmission block 911 to be executed within one frame period. Furthermore, transmission block 912 of encoded P-frame block 902 is shown to have low one-way latency, enabling both encoding block 902 and transmission block 912 to be executed within one frame period.

[0111] On the other hand, the transport block 915A of the encoded I-frame block 905 exhibits a high one-way latency, causing the encoding block 905 and transport block 915A to occur within several frame periods, thus introducing jitter into the one-way latency between the cloud gaming server and the client. To provide a real-time experience for the user and achieve a low one-way latency, a buffer at the client end can be omitted to correct for jitter. In that case, the encoder can decide to skip the encoding of one or more video frames after encoding the I-frame. For example, the encoder discards video frame 920. In that case, such as after five subsequent video frames have been encoded into P-frames and transmitted to the client, the transmission of encoded video frames returns to one of the low one-way latency options around the salience region 910. That is, the fourth or fifth P-frame encoded after encoding the I-frame block 905 is also transmitted to the client within the same frame period, thus returning to the low one-way latency between the cloud gaming server and the client.

[0112] In one implementation, when the transmission rate to the client is high for the target resolution of the client's display, the encoder still compresses the second video frame after a certain delay (i.e., until the I-frame has been encoded). Specifically, the transmission rate of a set of compressed video frames to the client is within the maximum receiving bandwidth. For example, the transmission rate to the client may be 13 megabits per second (Mbps), but the measured receiving bandwidth at the client may currently be 15 Mbps. Therefore, there is no increase in one-way latency between the cloud gaming server and the client because there is no delay in receiving the encoded video frames at the client.

[0113] Furthermore, because the VSYNC signals at the cloud gaming server and client can be synchronized and offset, one-way latency between the cloud gaming server and client can be reduced, thereby compensating for any variability in latency introduced by jitter at the server, during transmission over the network, or at the client. Moreover, the synchronization and offset of the VSYNC signals also provides overlapping operations at the cloud gaming server (scanning output, encoding, and transmission); overlapping operations at the client (receiving, decoding, rendering, and display); and / or overlapping operations at both the cloud gaming server and client, all of which contribute to compensating for latency variability introduced by server, network, or client jitter, reducing one-way latency, reducing variability in one-way latency, real-time generation and display of video content, and consistent video playback at the client.

[0114] Figure 9BA sequence of video frames 900B compressed by an encoder according to one embodiment of this disclosure is illustrated, wherein the encoder takes into account the available bandwidth of the client, such that if the bandwidth exceeds the bandwidth required to support the target resolution of the client's display, it is possible to have no skipped frames or fewer skipped frames while still having low latency. Specifically, in sequence 900B, video frames are encoded as I-frames, and when the client bandwidth is moderate relative to the target resolution of the client's display, subsequent video frames are encoded normally after the delayed encoding of the I-frames. Because of the moderate bandwidth availability, a suitable amount of excess bandwidth can be used to compensate for latency variability (e.g., jitter) between the cloud gaming server and the client, making frame skipping avoidable and enabling a relatively fast (e.g., within two to four frame periods) return to low one-way latency. The encoded block and transport block of the video frame are shown relative to the VSYNC signal 950.

[0115] Video frame sequence 900B includes one encoded I-frame 905, while the remaining frames are encoded as P-frames. For illustrative purposes, encoding blocks 901 and 902, which are P-frames, are encoded before encoding block 905 is encoded as an I-frame. The encoder then compresses the video frames into P-frames until the next scene change, or until the video frame can no longer reference a previous keyframe (e.g., an I-frame). Generally, the encoding time for an I-frame block may be longer than that for a P-frame block, and the transmission of an I-frame may take more than one frame cycle. For example, the encoding and transmission time of I-frame block 905 exceeds one frame cycle. Furthermore, various transmission times are shown relative to the corresponding encoded video frames. For example, the encoding and transmission of video frames preceding I-frame block 905 are shown to have low one-way latency, allowing the corresponding encoding and transmission blocks to be executed within one frame cycle. However, the transport block 915B of the encoded I-frame block 905 exhibits a high one-way delay, causing the encoded block 905 and transport block 915B to occur within two or more frame periods, thereby introducing jitter in the one-way delay between the cloud gaming server and the client. As previously discussed, encoding time can be further reduced by tuning one or more encoder parameters (e.g., QP, target frame size, maximum frame size, encoder bit rate, etc.). That is, the second or subsequent video frames after the I-frame are encoded with lower precision when the transmission rate to the client is moderate for the target resolution of the client display, and with even lower precision than when the transmission rate is high for the target resolution.

[0116] Following I-frame block 905, the encoder continues to compress video frames, although they may be briefly delayed due to the encoding of the I-frames. Similarly, the synchronization and offset of the VSYNC signal provide overlapping operations at the cloud gaming server (scanning output, encoding, and transmission); overlapping operations at the client (receiving, decoding, rendering, and display); and / or overlapping operations at both the cloud gaming server and client. All of these facilitate compensation for variability in one-way latency introduced by server, network, or client jitter, reducing one-way latency, minimizing variability in one-way latency, real-time generation and display of video content, and consistent video playback at the client.

[0117] Because the client bandwidth is adequate relative to the target resolution of the client's display, the transmission of encoded video frames, such as after two or three subsequent video frames have been encoded into P-frames and transmitted to the client, returns to one of the low one-way latency options around the salience region 940. Within region 940, P-frames encoded after I-frame block 905 are also transmitted to the client within the same frame period, thus returning to low one-way latency between the cloud gaming server and the client.

[0118] Figure 9C A sequence of video frames 900C compressed by an encoder according to one embodiment of this disclosure is illustrated, wherein the encoder takes into account the available bandwidth of the client, such that if the bandwidth exceeds the bandwidth required to support the target resolution of the client's display, it is possible to have no skipped frames or fewer skipped frames while still having low one-way latency. Specifically, in sequence 900C, video frames are encoded as I-frames, and when the client bandwidth is high relative to the target resolution of the client's display, subsequent video frames are encoded normally after the delayed encoding of the I-frames. Because of the high bandwidth availability, a large amount of excess bandwidth can be used to compensate for the variability (e.g., jitter) in one-way latency between the cloud gaming server and the client, making frame skipping avoidable and enabling an immediate (e.g., within one to two frame periods) return to low one-way latency. The encoded block and transport block of the video frame are shown relative to the VSYNC signal 950.

[0119] Similar to Figure 9B , Figure 9CThe video frame sequence 900C in the diagram includes an encoded I-frame 905, while the remaining frames are encoded as P-frames. For illustrative purposes, encoding blocks 901 and 902, which are P-frames, are encoded before encoding block 905 is encoded as an I-frame. The encoder then compresses the video frames into P-frames until the next scene change, or until the video frame can no longer reference a previous keyframe (e.g., an I-frame). Typically, the encoding time for an I-frame block can be longer than that for a P-frame block. For example, the encoding time for I-frame block 905 can exceed one frame period. Furthermore, various transmission times are shown relative to the corresponding encoded video frames. For example, the encoding and transmission of video frames preceding I-frame block 905 are shown as having low latency, allowing the corresponding encoding and transmission blocks to be executed within one frame period. However, the transmission block 915C of encoded I-frame block 905 is shown as having higher latency, causing encoding block 905 and transmission block 915C to occur over two or more frame periods, thus introducing jitter in the one-way latency between the cloud gaming server and the client. As previously discussed, encoding time can be further reduced by tuning one or more encoder parameters (e.g., QP, target frame size, maximum frame size, encoder bit rate, etc.).

[0120] Following I-frame block 905, the encoder continues to compress video frames, although they may be briefly delayed due to the encoding of I-frames. Similarly, the synchronization and offset of the VSYNC signal provide overlapping operations at the cloud gaming server (scanning output, encoding, and transmission); overlapping operations at the client (receiving, decoding, rendering, and display); and / or overlapping operations at both the cloud gaming server and client. All of these contribute to compensating for latency variability introduced by server, network, or client jitter, reducing one-way latency, reducing variability in one-way latency, real-time generation and display of video content, and consistent video playback at the client. Because client bandwidth is high relative to the target resolution of the client display, transmission of encoded video frames, such as after one or two subsequent video frames have been encoded as P-frames and transmitted to the client, returns to one of the low one-way latency options around salience region 970. Within region 970, P-frames encoded after I-frame block 905 are also transmitted to the client within one frame period (but they may span between the two sides where the VSYNC signal appears), thus returning to low one-way latency between the cloud gaming server and client.

[0121] Figure 10 Components of an exemplary apparatus 1000, which can be used to perform various aspects of embodiments of this disclosure, are illustrated. For example, Figure 10An exemplary hardware system suitable for streaming media content and / or receiving streaming media content according to embodiments of this disclosure is illustrated, including providing encoder tuning to improve the trade-off between one-way latency and video quality in a cloud gaming system, so as to reduce latency and provide more consistent latency between the cloud gaming server and the client, and improve the smoothness of video display on the client. The encoder tuning may be based on monitoring client bandwidth, skipped frames, the number of encoded I-frames, the number of scene changes, and / or the number of video frames exceeding the target frame size. The tuned parameters may include encoder bit rate, target frame size, maximum frame size, and quantization parameter (QP) values. High-performance encoders and decoders contribute to reducing the total one-way latency between the cloud gaming server and the client. This block diagram illustrates apparatus 1000, which may be incorporated into or may be a personal computer, server computer, game console, mobile device, or other digital device, each of which is suitable for practicing embodiments of the invention. Apparatus 1000 includes a central processing unit (CPU) 1002 for running software applications and optionally an operating system. CPU 1002 may consist of one or more homogeneous or heterogeneous processing cores.

[0122] According to various implementations, CPU 1002 is one or more general-purpose microprocessors having one or more processing cores. Other implementations may be implemented using one or more CPUs with microprocessor architectures particularly suited to highly parallel and computationally intensive applications such as media and interactive entertainment applications, or applications configured for graphics processing during game execution.

[0123] Memory 1004 stores applications and data for use by CPU 1002 and GPU 1016. Storage device 1006 provides non-volatile storage and other computer-readable media for applications and data, and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROMs, DVD-ROMs, Blu-ray discs, HD-DVDs, or other optical storage devices, as well as signal transmission and storage media. User input device 1008 conveys user input from one or more users to device 1000, examples of which may include a keyboard, mouse, joystick, touchpad, touchscreen, still or video recorder / camera, and / or microphone. Network interface 1009 allows device 1000 to communicate with other computer systems via electronic communication networks and may include wired or wireless communication over local area networks and wide area networks such as the Internet. Audio processor 1012 is adapted to generate analog or digital audio output from instructions and / or data provided by CPU 1002, memory 1004, and / or storage device 1006. The components of the device 1000, including a CPU 1002, a graphics subsystem including a GPU 1016, a memory 1004, a data storage device 1006, a user input device 1008, a network interface 1009, and an audio processor 1012, are connected via one or more data buses 1022.

[0124] The graphics subsystem 1014 is further connected to the data bus 1022 and components of the device 1000. The graphics subsystem 1014 includes a graphics processing unit (GPU) 1016 and a graphics memory 1018. The graphics memory 1018 includes display memory (e.g., a frame buffer) for storing pixel data for each pixel of the output image. The graphics memory 1018 may be integrated into the same device as the GPU 1016, connected to the GPU 1016 as a separate device, and / or implemented within memory 1004. Pixel data may be provided directly from the CPU 1002 to the graphics memory 1018. Alternatively, the CPU 1002 provides the GPU 1016 with data and / or instructions defining the desired output image, and the GPU 1016 generates pixel data for one or more output images based on the data and / or instructions. The data and / or instructions defining the desired output image may be stored in memory 1004 and / or graphics memory 1018. In one implementation, GPU 1016 includes 3D rendering capabilities for generating pixel data of an output image based on instructions and data, the pixel data defining the geometry, lighting, shading, texturing, motion, and / or camera parameters of a scene. GPU 1016 may also include one or more programmable execution units capable of executing shader programs.

[0125] The graphics subsystem 1014 periodically outputs pixel data of an image from the graphics memory 1018 for display on the display device 1010 or projected by a projection system (not shown). The display device 1010 can be any device capable of displaying visual information in response to signals from the device 1000, including CRT, LCD, plasma, and OLED displays. For example, the device 1000 can provide analog or digital signals to the display device 1010.

[0126] Other implementations for optimizing the graphics subsystem 1014 may include multi-tenant GPU operation, where GPU instances are shared across multiple applications, and distributed GPUs support a single game. The graphics subsystem 1014 may be configured as one or more processing devices.

[0127] For example, in one implementation, graphics subsystem 1014 may be configured to perform multi-tenant GPU functionality, whereby a graphics subsystem may implement graphics and / or rendering pipelines for multiple games. That is, graphics subsystem 1014 is shared among multiple games being executed.

[0128] In other implementations, the graphics subsystem 1014 includes multiple GPU devices that are combined to perform graphics processing for a single application running on a corresponding CPU. For example, the multiple GPUs may perform alternating frame rendering, where in consecutive frame cycles, GPU 1 renders the first frame, GPU 2 renders the second frame, and so on, until the last GPU is reached, after which the initial GPU renders the next video frame (e.g., if there are only two GPUs, GPU 1 renders the third frame). That is, the GPUs rotate while rendering frames. Rendering operations may overlap, where GPU 2 may begin rendering the second frame before GPU 1 has finished rendering the first frame. In another implementation, different shader operations may be assigned to the multiple GPU devices in the rendering and / or graphics pipeline. The main GPU is performing main rendering and compositing. For example, in a group comprising three GPUs, the primary GPU 1 performs primary rendering (e.g., first shader operations) and synthesizes output from secondary GPUs 2 and 3, where secondary GPU 2 performs second shader operations (e.g., fluid effects, such as rivers), and secondary GPU 3 performs third shader operations (e.g., particulate smoke), with the primary GPU 1 synthesizing the results from each of GPUs 1, 2, and 3. In this manner, different GPUs can be assigned to perform different shader operations (e.g., waving flags, wind, smoke generation, fire, etc.) to render video frames. In yet another embodiment, each of the three GPUs can be assigned to different objects and / or portions of the scene corresponding to a video frame. In the above embodiments and implementations, these operations can be performed in the same frame cycle (simultaneously in parallel) or in different frame cycles (sequentially in parallel).

[0129] Therefore, this disclosure describes methods and systems configured to stream media content and / or receive streamed media content, including providing encoder tuning to improve the trade-off between one-way latency and video quality in a cloud gaming system, wherein encoder tuning may be based on monitoring client bandwidth, skipped frames, the number of encoded I-frames, the number of scene changes, and / or the number of video frames exceeding a target frame size, wherein tuned parameters may include encoder bit rate, target frame size, maximum frame size, and quantization parameter (QP) value, wherein high-performance encoders and decoders help reduce the total one-way latency between the cloud gaming server and the client.

[0130] It should be understood that the various embodiments defined herein can be combined or aggregated into specific implementations using the various features disclosed herein. Therefore, the examples provided are merely some possible examples and are not limited to the various implementations possible by combining various elements to define more implementations. In some examples, some implementations may include fewer elements without departing from the spirit of the disclosed or equivalent implementations.

[0131] The embodiments of this disclosure can be practiced using a variety of computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. Embodiments of this disclosure can also be practiced in distributed computing environments, where tasks are performed via remote processing devices connected by wired or wireless network links.

[0132] Having understood the above embodiments, it should be understood that embodiments of this disclosure can employ various computer-implemented operations involving data stored in a computer system. These operations are operations that require the physical manipulation of physical quantities. Any of the operations described herein that form part of embodiments of this disclosure is a useful machine operation. Embodiments of this disclosure also relate to means or apparatus for performing these operations. The apparatus may be specifically configured for the desired purpose, or the apparatus may be a general-purpose computer selectively activated or configured by a computer program stored in a computer. Specifically, various general-purpose machines may be used with computer programs written in accordance with the teachings herein, or it may be more convenient to construct more specialized apparatus to perform the desired operations.

[0133] This disclosure can also be embodied in computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device capable of storing data that can subsequently be read by a computer system. Examples of computer-readable media include hard disk drives, network attached storage devices (NAS), read-only memory, random access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. Computer-readable media may include computer-readable tangible media distributed across a network-coupled computer system, such that computer-readable code is stored and executed in a distributed manner.

[0134] Although the method operations are described in a specific order, it should be understood that other housekeeping operations may be performed between operations, or operations may be adjusted so that they occur at slightly different times, or they may be distributed throughout the system. This allows processing operations to occur at various intervals associated with processing, as long as the processing of the overriding operation is performed in the desired manner.

[0135] While the foregoing disclosure has been described in some detail for clarity, it will be understood that certain changes and modifications may be practiced within the scope of the appended claims. Therefore, this embodiment is to be regarded as illustrative rather than restrictive, and embodiments of this disclosure are not limited to the details given herein, but may be modified within the scope and equivalents of the appended claims.

Claims

1. A method for cloud gaming, the method comprising: When a video game is played on a cloud gaming server, multiple video frames are generated. The plurality of video frames are encoded using an encoder at a certain encoder bit rate, wherein the plurality of video frames to be compressed are transmitted from the streaming transmitter of the cloud gaming server to the client. Measure the client's maximum receive bandwidth; The encoding of the plurality of video frames, the number of encoded I-frames, the number of video frames with scene changes, and the number of video frames exceeding the target frame size are monitored at the streaming transmitter. as well as The encoder parameters are dynamically tuned based on the monitoring of the encoding, the number of encoded I-frames, the number of video frames with scene changes, and the number of video frames exceeding the target frame size, as well as the maximum receiving bandwidth of the client. The dynamic tuning of the parameters includes: Determine if the number of video frames compressed into I-frames in a set of compressed video frames from the plurality of compressed video frames satisfies or exceeds a threshold number of I-frames; and Increasing the value of the QP parameter results in encoding being performed with lower precision, where the parameter is the QP parameter.

2. The method of claim 1, wherein dynamically tuning the parameters comprises: or, Determine that the encoder bit rate used by a group of video frames encoded from the plurality of video frames exceeds the maximum receive bandwidth; as well as Increasing the value of the QP parameter results in encoding being performed with lower precision, where the parameter is the QP parameter.

3. The method of claim 1, wherein dynamically tuning the parameters comprises: or, The encoder bit rate used to encode a set of video frames from the plurality of video frames is determined to be within the maximum receive bandwidth; It was determined that there was excess bandwidth when sending the set of video frames; as well as The value of the QP parameter is reduced based on the excess bandwidth, so that encoding is performed with higher precision, wherein the parameter is the QP parameter.

4. The method of claim 1, wherein dynamically tuning the parameters comprises: or, It is determined that the number of video frames compressed into I-frames in a set of video frames from the plurality of compressed video frames is less than the threshold number of I-frames. as well as Decreasing the value of the QP parameter allows encoding to be performed with higher precision, where the parameter is the QP parameter.

5. The method of claim 1, wherein dynamically tuning the parameters comprises: or, Determining a set of video frames from the plurality of encoded video frames transmitted at a transmission rate comprises a certain number of video frames, the number of video frames satisfying or exceeding a threshold, wherein each of the video frames in that number exceeds a target frame size; and Reduce at least one of the target frame size and the maximum frame size for the parameters.

6. The method of claim 5, wherein the target frame size and the maximum frame size are equal.

7. The method of claim 1, wherein dynamically tuning the parameters comprises: or, Determining a set of video frames from the plurality of encoded video frames transmitted at a transmission rate comprises a certain number of video frames, the number of which is less than a threshold, wherein each of the video frames in that number is within a target frame size; and The parameter increases by at least one of the target frame size and the maximum frame size.

8. The method of claim 1, wherein dynamically tuning the parameters comprises: or, The number of video frames identified as having scene changes in a set of video frames from the plurality of compressed video frames meets or exceeds a threshold number of scene changes. as well as Increasing the value of the QP parameter results in encoding being performed with lower precision, where the parameter is the QP parameter.

9. The method of claim 1, wherein dynamically tuning the parameters comprises: or, The number of video frames identified as having scene changes in a set of video frames from the plurality of compressed video frames is determined to be below a threshold number of scene changes. as well as Decreasing the value of the QP parameter allows encoding to be performed with higher precision, where the parameter is the QP parameter.

10. The method of claim 1, further comprising: By disabling skipped encoding of video frames, playback smoothness is biased towards the client.

11. The method of claim 1, further comprising: The encoder bit rate at the encoder is dynamically adjusted based on the client's maximum receive bandwidth.

12. A non-transitory computer-readable medium storing a computer program for cloud gaming, the computer-readable medium comprising: Program instructions used to generate multiple video frames when a video game is played on a cloud gaming server; Program instructions used to measure the maximum receive bandwidth of a client; Program instructions for encoding the plurality of video frames at a certain encoder bit rate using an encoder, wherein the plurality of video frames to be compressed are transmitted from the streaming transmitter of the cloud gaming server to the client. Program instructions for monitoring the encoding of the plurality of video frames at the streaming transmitter, the number of encoded I-frames, the number of video frames with scene changes, and the number of video frames exceeding the target frame size. as well as Program instructions for dynamically tuning the encoder parameters based on the monitoring of the encoding, the number of encoded I-frames, the number of video frames with scene changes, and the number of video frames exceeding the target frame size, as well as the maximum receiving bandwidth of the client. The dynamic tuning of the parameters includes: Determine if the number of video frames compressed into I-frames in a set of compressed video frames from the plurality of compressed video frames satisfies or exceeds a threshold number of I-frames; and Increasing the value of the QP parameter results in encoding being performed with lower precision, where the parameter is the QP parameter.

13. The non-transitory computer-readable medium of claim 12, wherein the program instructions for dynamically tuning the parameters comprise: or, Program instructions for determining that the encoder bit rate used to encode a set of video frames from the plurality of video frames exceeds the maximum receive bandwidth; as well as The value of the QP parameter is used to increase the value of the encoded program instructions to execute them with lower precision, wherein the parameter is the QP parameter.

14. The non-transitory computer-readable medium of claim 12, wherein the program instructions for dynamically tuning the parameters comprise: or, Program instructions for determining the encoder bit rate used to encode a set of video frames from the plurality of video frames within the maximum receive bandwidth; Program instructions used to determine if there is excess bandwidth when sending the set of video frames; as well as The program instructions are used to reduce the value of the QP parameter based on the excess bandwidth so as to execute the encoded instructions with higher precision, wherein the parameter is the QP parameter.

15. The non-transitory computer-readable medium of claim 12, wherein the program instructions for dynamically tuning the parameters comprise: or, Program instructions for determining that the number of video frames compressed into I-frames in a set of compressed video frames is less than a threshold number of I-frames. as well as The value of the QP parameter is used to reduce the value of the encoded program instructions so that they are executed with higher precision, wherein the parameter is the QP parameter.

16. The non-transitory computer-readable medium of claim 12, wherein the program instructions for dynamically tuning the parameters comprise: or, Program instructions for determining a set of video frames from the plurality of encoded video frames transmitted at a transmission rate comprising a certain number of video frames, the number of video frames satisfying or exceeding a threshold, wherein each of the number of video frames exceeds a target frame size; and Program instructions for reducing at least one of the target frame size and the maximum frame size for the parameters.

17. The non-transitory computer-readable medium of claim 16, wherein in the computer program for cloud gaming, the target frame size and the maximum frame size are equal.

18. The non-transitory computer-readable medium of claim 12, wherein the program instructions for dynamically tuning the parameters comprise: or, Program instructions for determining a set of video frames from the plurality of encoded video frames transmitted at a transmission rate comprising a certain number of video frames, the number of video frames being less than a threshold, wherein each of the video frames in that number is within a target frame size; and Program instructions for increasing at least one of the target frame size and the maximum frame size as the parameters are stated.

19. The non-transitory computer-readable medium of claim 12, wherein the program instructions for dynamically tuning the parameters comprise: or, Program instructions for determining whether the number of video frames identified as having scene changes in a set of the plurality of compressed video frames meets or exceeds a threshold number of scene changes. as well as The value of the QP parameter is used to increase the value of the encoded program instructions to execute them with lower precision, wherein the parameter is the QP parameter.

20. The non-transitory computer-readable medium of claim 12, wherein the program instructions for dynamically tuning the parameters comprise: or, Program instructions for determining that the number of video frames identified as having scene changes in a set of compressed video frames is less than a threshold number of scene changes; and The value of the QP parameter is used to reduce the value of the encoded program instructions so that they are executed with higher precision, wherein the parameter is the QP parameter.

21. The non-transitory computer-readable medium of claim 12, wherein the non-transitory computer-readable medium further comprises: Program instructions used to bias playback smoothness at the client by disabling the skipping of video frame encoding.

22. The non-transitory computer-readable medium of claim 12, wherein the non-transitory computer-readable medium further comprises: Program instructions for dynamically adjusting the encoder bit rate speed at the encoder based on the maximum receive bandwidth of the client.

23. A computer system, the computer system comprising: processor; as well as A memory coupled to the processor and having instructions stored therein, which, when executed by the computer system, cause the computer system to perform a method for cloud gaming, the method comprising: When a video game is played on a cloud gaming server, multiple video frames are generated. The plurality of video frames are encoded using an encoder at a certain encoder bit rate, wherein the plurality of video frames to be compressed are transmitted from the streaming transmitter of the cloud gaming server to the client. Measure the client's maximum receive bandwidth; The encoding of the plurality of video frames, the number of encoded I-frames, the number of video frames with scene changes, and the number of video frames exceeding the target frame size are monitored at the streaming transmitter; and The encoder parameters are dynamically tuned based on the monitoring of the encoding, the number of encoded I-frames, the number of video frames with scene changes, and the number of video frames exceeding the target frame size, as well as the maximum receiving bandwidth of the client. The dynamic tuning of the parameters includes: Determine if the number of video frames compressed into I-frames in a set of compressed video frames from the plurality of compressed video frames satisfies or exceeds a threshold number of I-frames; and Increasing the value of the QP parameter results in encoding being performed with lower precision, where the parameter is the QP parameter.

24. The computer system of claim 23, wherein, in the method, dynamically tuning the parameters comprises: or, Determine that the encoder bit rate used by a group of video frames encoded from the plurality of video frames exceeds the maximum receive bandwidth; as well as Increasing the value of the QP parameter results in encoding being performed with lower precision, where the parameter is the QP parameter.

25. The computer system of claim 23, wherein, in the method, dynamically tuning the parameters comprises: or, The encoder bit rate used to encode a set of video frames from the plurality of video frames is determined to be within the maximum receive bandwidth; It was determined that there was excess bandwidth when sending the set of video frames; as well as The value of the QP parameter is reduced based on the excess bandwidth, so that encoding is performed with higher precision, wherein the parameter is the QP parameter.

26. The computer system of claim 23, wherein in the method, dynamically tuning the parameters comprises: or, It is determined that the number of video frames compressed into I-frames in a set of video frames from the plurality of compressed video frames is less than the threshold number of I-frames. as well as Decreasing the value of the QP parameter allows encoding to be performed with higher precision, where the parameter is the QP parameter.

27. The computer system of claim 23, wherein in the method, dynamically tuning the parameters comprises: or, Determining a set of video frames from the plurality of encoded video frames transmitted at a transmission rate comprises a certain number of video frames, the number of video frames satisfying or exceeding a threshold, wherein each of the video frames in that number exceeds a target frame size; and Reduce at least one of the target frame size and the maximum frame size for the parameters.

28. The computer system of claim 27, wherein in the method, the target frame size and the maximum frame size are equal.

29. The computer system of claim 23, wherein in the method, dynamically tuning the parameters comprises: or, Determining a set of video frames from the plurality of encoded video frames transmitted at a transmission rate comprises a certain number of video frames, the number of which is less than a threshold, wherein each of the video frames in that number is within a target frame size; and The parameter increases by at least one of the target frame size and the maximum frame size.

30. The computer system of claim 23, wherein in the method, dynamically tuning the parameters comprises: or, The number of video frames identified as having scene changes in a set of video frames from the plurality of compressed video frames meets or exceeds a threshold number of scene changes. as well as Increasing the value of the QP parameter results in encoding being performed with lower precision, where the parameter is the QP parameter.

31. The computer system of claim 23, wherein in the method, dynamically tuning the parameters comprises: or, The number of video frames identified as having scene changes in a set of video frames from the plurality of compressed video frames is determined to be below a threshold number of scene changes. as well as Decreasing the value of the QP parameter allows encoding to be performed with higher precision, where the parameter is the QP parameter.

32. The computer system of claim 23, wherein the method further comprises: By disabling skipped encoding of video frames, playback smoothness is biased towards the client.

33. The computer system of claim 23, wherein the method further comprises: The encoder bit rate at the encoder is dynamically adjusted based on the client's maximum receive bandwidth.

34. A method for cloud gaming, the method comprising: When a video game is played on a cloud gaming server, multiple video frames are generated. Predict scene changes in the first video frame of the video game, wherein the scene changes are predicted before the first video frame is generated; The first video frame is generated as a scene change prompt; Send the scene change notification to the encoder; The first video frame is delivered to the encoder, wherein the first video frame is encoded as an I-frame based on the scene change cue; Measure the client's maximum receive bandwidth; as well as Whether to encode the second video frame received at the encoder is determined based on the client's maximum receiving bandwidth and the target resolution of the client's display.

35. The method of claim 34, further comprising: The encoder bit rate at the encoder is dynamically adjusted based on the maximum receiving bandwidth of the client. as well as The video frame is transmitted to the client while it is being encoded.

36. The method of claim 34, wherein determining whether to encode the second video frame comprises: When the transmission rate to the client is low for the target resolution of the client's display, the encoding of the second video frame is skipped, such that the transmission rate of a set of video frames from the compressed plurality of video frames to the client exceeds the maximum receiving bandwidth.

37. The method of claim 34, wherein determining whether to encode the second video frame comprises: If the transmission rate to the client is high for the target resolution of the client's display, the second video frame is encoded normally such that the transmission rate of a set of video frames from the compressed plurality of video frames to the client is within the maximum receiving bandwidth.

38. The method of claim 37, wherein determining whether to encode the second video frame comprises: If the transmission rate to the client is appropriate for the target resolution of the client's display, the second video frame is encoded with lower precision.

39. The method of claim 34, wherein predicting scene changes in the first video frame comprises: The game logic built on the game engine of the video game is executed at the cloud gaming server to generate the multiple video frames; Execute scene change logic to predict the scene change of the first video frame, wherein the prediction is based on the game state collected during the execution of the game logic; The scene change prompt is generated using the scene change logic; as well as The scene change notification is sent before the encoder receives the first video frame.

40. The method of claim 39, wherein the scene change cue is delivered from the scene change logic to the encoder via an API.

41. The method of claim 34, wherein the second video frame is compressed after the encoder has compressed the first video frame.

42. A non-transitory computer-readable medium storing a computer program for cloud gaming, the computer-readable medium comprising: Program instructions used to generate multiple video frames when a video game is played on a cloud gaming server; Program instructions for predicting scene changes in a first video frame of the video game, wherein the scene changes are predicted before the first video frame is generated; Program instructions for generating a scene change prompt that the first video frame is a scene change; Program instructions used to send the scene change notification to the encoder; Program instructions for delivering the first video frame to the encoder, wherein the first video frame is encoded as an I-frame based on the scene change cue; Program instructions used to measure the maximum receive bandwidth of a client; as well as Program instructions for determining whether to encode a second video frame received at the encoder based on the client's maximum receiving bandwidth and the target resolution of the client's display.

43. The non-transitory computer-readable medium of claim 42, further comprising: Program instructions for dynamically adjusting the encoder bit rate speed at the encoder based on the maximum receive bandwidth of the client; as well as Program instructions for transmitting the video frame to the client while the video frame is being encoded.

44. The non-transitory computer-readable medium of claim 42, wherein the program instructions for determining whether to encode the second video frame include: Program instructions for the following operation: when the transmission rate to the client is low for the target resolution of the client's display, skip the encoding of the second video frame, such that the transmission rate of a set of video frames from the compressed plurality of video frames to the client exceeds the maximum receiving bandwidth.

45. The non-transitory computer-readable medium of claim 42, wherein the program instructions for determining whether to encode the second video frame include: Program instructions for the following operation: if the transmission rate to the client is high for the target resolution of the client's display, then the second video frame is encoded normally such that the transmission rate of a set of video frames from the compressed plurality of video frames to the client is within the maximum receiving bandwidth.

46. ​​The non-transitory computer-readable medium of claim 45, wherein the program instructions for determining whether to encode the second video frame include: The program instructions are used to encode the second video frame with lower precision if the transmission rate to the client is appropriate for the target resolution of the client's display.

47. The non-transitory computer-readable medium of claim 42, wherein the program instructions for predicting scene changes in the first video frame comprise: Program instructions for executing game logic built on the game engine of the video game at the cloud gaming server to generate the multiple video frames; Program instructions for executing scene change logic to predict the scene change of the first video frame, wherein the prediction is based on the game state collected during the execution of the game logic; Program instructions for generating the scene change prompt using the scene change logic; as well as Program instructions for sending the scene change notification before the encoder receives the first video frame.

48. The non-transitory computer-readable medium of claim 47, wherein in the computer program for cloud gaming, the scene change cue is delivered from the scene change logic to the encoder via an API.

49. The non-transitory computer-readable medium of claim 42, wherein in the computer program for cloud gaming, the second video frame is compressed after the encoder has compressed the first video frame.

50. A computer system, the computer system comprising: processor; as well as A memory coupled to the processor and having instructions stored therein, which, when executed by the computer system, cause the computer system to perform a method for cloud gaming, the method comprising: When a video game is played on a cloud gaming server, multiple video frames are generated. Predict scene changes in the first video frame of the video game, wherein the scene changes are predicted before the first video frame is generated; The first video frame is generated as a scene change prompt; Send the scene change notification to the encoder; The first video frame is delivered to the encoder, wherein the first video frame is encoded as an I-frame based on the scene change cue; Measure the client's maximum receive bandwidth; and Whether to encode the second video frame received at the encoder is determined based on the client's maximum receiving bandwidth and the target resolution of the client's display.

51. The computer system of claim 50, further comprising: The encoder bit rate at the encoder is dynamically adjusted based on the maximum receiving bandwidth of the client. as well as The video frame is transmitted to the client while it is being encoded.

52. The computer system of claim 50, wherein determining whether to encode the second video frame comprises: When the transmission rate to the client is low for the target resolution of the client's display, the encoding of the second video frame is skipped, such that the transmission rate of a set of video frames from the compressed plurality of video frames to the client exceeds the maximum receiving bandwidth.

53. The computer system of claim 50, wherein determining whether to encode the second video frame comprises: If the transmission rate to the client is high for the target resolution of the client's display, the second video frame is encoded normally such that the transmission rate of a set of video frames from the compressed plurality of video frames to the client is within the maximum receiving bandwidth.

54. The computer system of claim 53, wherein determining whether to encode the second video frame comprises: If the transmission rate to the client is appropriate for the target resolution of the client's display, the second video frame is encoded with lower precision.

55. The computer system of claim 50, wherein predicting scene changes in the first video frame comprises: The game logic built on the game engine of the video game is executed at the cloud gaming server to generate the multiple video frames; Execute scene change logic to predict the scene change of the first video frame, wherein the prediction is based on the game state collected during the execution of the game logic; The scene change prompt is generated using the scene change logic; as well as The scene change notification is sent before the encoder receives the first video frame.

56. The computer system of claim 55, wherein the scene change cue is delivered from the scene change logic to the encoder via an API.

57. The computer system of claim 50, wherein the second video frame is compressed after the encoder has compressed the first video frame.

Citation Information

Patent Citations

  • Bit rate control video compression method and device on basis of scene switching

    CN102630013A

  • System and method for encoding video content using virtual intra-frames

    CN104641638A

  • Video code rate processing method and device

    CN107872669A

  • Managing video adaptation algorithms

    US20100316066A1

  • Client side processing of player movement in a remote gaming environment

    US20140267429A1