Encoder tuning to improve the trade-off between latency and video quality in cloud gaming applications.
High-performance encoders and decoders in cloud gaming systems adjust parameters based on network speed and frame characteristics to balance latency and quality, reducing one-way latency and stabilizing frame rates for improved video streaming.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SONY INTERACTIVE ENTERTAINMENT LLC
- Filing Date
- 2026-01-07
- Publication Date
- 2026-04-10
AI Technical Summary
Cloud gaming systems face challenges in achieving a balance between one-way latency and video quality due to network connectivity limitations and processing constraints, leading to inconsistent and high latency in streaming high-quality video frames.
Implementing high-performance encoders and decoders that adjust encoder parameters based on network speed, client bandwidth, and video frame characteristics, such as scene changes and frame size, to optimize latency and quality.
Reduces one-way latency and stabilizes frame rate by dynamically adjusting encoder settings, ensuring smoother and more reliable video playback in cloud gaming applications.
Smart Images

Figure 2026062986000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a streaming system configured to stream content over a network, and more particularly, to a high-performance encoder and decoder for a cloud gaming system, and a streaming system configured for encoder adjustment recognizing network transmission speed and reliability, and overall latency targets.
Background Art
[0002] In recent years, online services enabling online or cloud gaming in a streaming format between a cloud gaming server and a client connected via a network have been continuously promoted. The streaming format is becoming increasingly popular due to, among other things, the availability of on-demand game titles, the ability to network between players for multiplayer games, the sharing of assets between players, the sharing of instant experiences between players and / or spectators, the ability for friends to watch friends play video games, and the ability for friends to have friends participate in ongoing gameplay. Unfortunately, this demand is also driven by the capabilities of network connectivity and the limitations of processing performed on both the server and the client, which must be responsive enough to render high-quality images once delivered to the client. For example, the results of all game activity performed on the server must be transmitted back to the client compressed with low millisecond latency for the best user experience. Round-trip latency can be defined as the total time between a user's controller input and the display of a video frame on the client, which may include processing and transmitting control information from the controller to the client, processing and transmitting control information from the client to the server, using that input on the server to generate a video frame in response to the input, processing and transferring the video frame to an encoding unit (e.g., scanout), encoding the video frame, transmitting the encoded video frame back to the client, receiving and decoding the video frame, and either processing or staging the video frame before display. One-way latency can be defined as part of round-trip latency, which consists of the time from the start of the transfer of a video frame to the encoding unit on the server (e.g., scanout) to the start of the display of the video frame on the client. Part of round-trip latency and one-way latency is related to the time it takes for a data stream to be sent from the client to the server and from the server to the client over the communication network. Another part is related to processing on the client and server, and improvements in these operations, such as advanced policies regarding frame decoding and display, can result in a substantial reduction in round-trip latency and one-way latency between the server and client, providing users of cloud gaming services with a high-quality experience.
[0003] The embodiments of this disclosure were made against this background. [Overview of the Initiative]
[0004] Embodiments of the present disclosure relate to a streaming system configured to stream content (e.g., games) over a network, and more specifically, to a streaming system configured to provide encoder tuning to improve the trade-off between one-way latency and video quality in a cloud gaming system, wherein the encoder tuning is obtained based on monitoring client bandwidth, skipped frames, the number of encoded I-frames, the number of scene changes, and / or the number of video frames exceeding the target frame size, and the tuned parameters may include encoder bitrate, target frame size, maximum frame size, and quantization parameter (QP) values, and high-performance encoders and decoders work to reduce the overall one-way latency between the cloud gaming server and the client.
[0005] Embodiments of this disclosure disclose a method for cloud gaming. The method includes generating a plurality of video frames when running a video game on a cloud gaming server. The method includes encoding the plurality of video frames at an encoder bitrate, and the compressed plurality of video frames are transmitted from a streamer on the cloud gaming server to a client. The method includes measuring the maximum receive bandwidth of the client. The method includes monitoring the encoding of the plurality of video frames in the streamer. The method includes dynamically adjusting the encoder parameters based on the encoding monitoring.
[0006] In another embodiment, a non-temporary computer-readable medium for storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating multiple video frames when a video game is run on a cloud gaming server. The computer-readable medium includes program instructions for encoding the multiple video frames at an encoder bitrate, and the compressed multiple video frames are transmitted from the streamer on the cloud gaming server to the client. The computer-readable medium includes program instructions for measuring the maximum receive bandwidth of the client. The computer-readable medium includes program instructions for monitoring the encoding of the multiple video frames in the streamer. The computer-readable medium includes program instructions for dynamically adjusting the encoder parameters based on the encoding monitoring.
[0007] In yet another embodiment, the computer system includes a processor and memory coupled to the processor, which stores instructions therein, causing the computer system to perform a cloud gaming method when executed by the computer system. The method includes generating a plurality of video frames when running a video game on a cloud gaming server. The method includes encoding the plurality of video frames at an encoder bitrate, and the compressed plurality of video frames are transmitted from the streamer on the cloud gaming server to the client. The method includes measuring the maximum receiving bandwidth of the client. The method includes monitoring the encoding of the plurality of video frames in the streamer. The method includes dynamically adjusting the encoder parameters based on the encoding monitoring.
[0008] In yet another embodiment, a method for cloud gaming is disclosed. This method includes generating a plurality of video frames when running a video game on a cloud gaming server. This method includes predicting a scene change in a first video frame of the video game, the scene change being predicted before the first video frame is generated. This method includes generating a scene change hint, where the first video frame is a scene change. This method includes sending the scene change hint to an encoder. This method includes delivering the first video frame to the encoder, where the first video frame is encoded as an I-frame based on the scene change hint. This method includes measuring the client's maximum receive bandwidth. This method includes determining whether or not to encode a second video frame received by the encoder, based on the client's maximum receive bandwidth and the target resolution of the client display.
[0009] In another embodiment, a non-temporary computer-readable medium for storing a computer program for cloud gaming is disclosed. The computer-readable medium includes program instructions for generating multiple video frames when a video game is run on a cloud gaming server. The computer-readable medium includes program instructions for predicting a scene change in a first video frame of the video game, where the scene change is predicted before the first video frame is generated. The computer-readable medium includes program instructions for generating a scene change hint, where the first video frame is a scene change. The computer-readable medium includes program instructions for sending the scene change hint to an encoder. The computer-readable medium includes program instructions for delivering the first video frame to the encoder, where the first video frame is encoded as an I-frame based on the scene change hint. The computer-readable medium includes program instructions for measuring the client's maximum receive bandwidth. The computer-readable medium includes program instructions for determining whether or not to encode a second video frame received by the encoder, based on the client's maximum receive bandwidth and the target resolution of the client display.
[0010] In yet another embodiment, the computer system includes a processor and memory coupled to the processor, which, when executed by the computer system, stores instructions that cause the computer system to perform a cloud gaming method. The method includes generating a number of video frames when running a video game on a cloud gaming server. The method includes predicting a scene change in a first video frame of the video game, where the scene change is predicted before the first video frame is generated. The method includes generating a scene change hint, where the first video frame is a scene change. This method includes sending a scene change hint to the encoder. This method includes delivering a first video frame to the encoder, which is encoded as an I-frame based on the scene change hint. This method includes measuring the client's maximum receiving bandwidth. This method includes determining whether or not to encode a second video frame received by the encoder, based on the client's maximum receiving bandwidth and the target resolution of the client display.
[0011] Other aspects of this disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, which illustrate the principles of this disclosure.
[0012] This disclosure can be best understood by referring to the following description in conjunction with the attached drawings. [Brief explanation of the drawing]
[0013] [Figure 1A] This is a diagram of the VSYNC signal at the start of a frame period according to one embodiment of the present disclosure. [Figure 1B] This is a diagram showing the frequency of the VSYNC signal according to one embodiment of the present disclosure. [Figure 2A] This diagram shows a system for providing games over a network between one or more cloud gaming servers and one or more client devices in various configurations according to one embodiment of the present disclosure, wherein VSYNC signals can be synchronized and offset to reduce one-way latency. [Figure 2B] This diagram shows an embodiment of the present disclosure for providing a game between two or more peer devices, where the VSYNC signal is synchronized and offset to achieve optimal timing for receiving other information between the controller and the devices. [Figure 2C] One embodiment of this disclosure illustrates various network configurations that benefit from proper synchronization and offset of VSYNC signals between a source device and a target device. [Figure 2D] One embodiment of this disclosure illustrates a multi-tenancy configuration between a cloud gaming server and multiple clients that benefits from proper synchronization and offset of VSYNC signals between a source device and a target device. [Figure 3] This embodiment of the present disclosure illustrates the fluctuation of one-way latency between a cloud gaming server and a client due to clock drift when streaming video frames generated from a video game running on a server. [Figure 4] This diagram illustrates a network configuration including a cloud gaming server and client when streaming video frames generated from a video game running on a server. The VSYNC signals between the server and client are synchronized and offset, allowing for overlapping operations on the server and client and reducing one-way latency between the server and client. [Figure 5] This is a flowchart illustrating a cloud gaming method according to one embodiment of the present disclosure, in which video frame encoding includes tuning encoder parameters with respect to network transmission speed and reliability, as well as an overall latency target. [Figure 6] This figure illustrates the measurement of client bandwidth by a streamer component operating at the application layer according to one embodiment of the present disclosure, the streamer being configured to monitor and adjust an encoder so that compressed video frames can be transmitted at a rate within the client's measured bandwidth. [Figure 7] Figure A shows the setting of encoder quantization parameters (QP) to optimize quality and buffer utilization in a client according to one embodiment of the present disclosure, and Figure B shows the adjustment of target frame size, maximum frame size, and / or QP (e.g., minimum QP and / or maximum QP) encoder settings to reduce the occurrence of I-frames that exceed the true target frame size supported by the client, according to one embodiment of the present disclosure. [Figure 8]This is a flowchart illustrating a cloud gaming method according to one embodiment of the present disclosure, wherein the encoding of video frames includes determining whether to skip video frames or delay the encoding and transmission of video frames when the encoding is long, such as when encoding I frames, or when the number of video frames being generated is large. [Figure 9] A shows a sequence of video frames being compressed by an encoder according to one embodiment of the present disclosure, in which the encoder drops encoding of video frames after encoding an I-frame when the client bandwidth is low relative to the target resolution of the client's display; B shows a sequence of video frames being compressed by an encoder according to one embodiment of the present disclosure, in which each sequence is encoded as an I-frame, and subsequent video frames are also encoded after a delay in encoding the I-frame when the client bandwidth is moderate or high relative to the target resolution of the client's display; and C shows a sequence of video frames being compressed by an encoder according to one embodiment of the present disclosure, in which each sequence is encoded as an I-frame, and subsequent video frames are also encoded after a delay in encoding the I-frame when the client bandwidth is moderate or high relative to the target resolution of the client's display. [Figure 10] This document illustrates exemplary device components that can be used to carry out various embodiments of the present disclosure. [Modes for carrying out the invention]
[0014] The following detailed description includes many specific details for illustrative purposes, but those skilled in the art will understand that many variations and modifications of the following details are within the scope of this disclosure. Accordingly, the aspects of this disclosure described below are described without loss of generality to and without limitation to the claims that follow this description.
[0015] Generally speaking, various embodiments of this disclosure describe methods and systems configured to reduce latency and / or latency instability between a source device and a target device when streaming media content (e.g., when streaming audio and video from a video game). Latency instability can occur in one-way latency between a server and a client due to the additional time required to generate complex frames (e.g., scene changes) on the server, the increased time required to encode / compress complex frames on the server, the variable communication path over the network, and the increased time required to decode complex frames on the client. Latency instability can also be introduced by differences in clocks between the server and the client, which cause drift between the server and client VSYNC signals. In embodiments of this disclosure, one-way latency between the server and client can be reduced in cloud gaming applications by providing high-performance encoding and decoding. When decompressing streaming media (e.g., streaming video, movies, clips, content), it is possible to buffer a significant portion of the decompressed video, and therefore, when displaying the streamed content, it is possible to rely on the average decoding capacity and metrics (e.g., to support 4K media at 60Hz, it is possible to rely on the average amount of decoding resources). However, in cloud gaming, the longer the time it takes to perform encoding and / or decoding operations (even for a single frame), the higher the one-way latency becomes. Therefore, in the case of cloud gaming, it is beneficial to supply powerful decoding and encoding resources that may seem unnecessary compared to the needs of the streaming video application, and the resources need to be optimized for the time it takes to process longer or the frames that require the longest processing. In other embodiments of this disclosure, encoder tuning may be performed to improve the trade-off between latency and video quality in cloud gaming applications. Encoder adjustment is performed in recognition of the network transmission speed and reliability, and in the overall latency target. In an embodiment, when encoding is executed for a long time or the generated data is large (e.g., both conditions can occur in a compressed I-frame), a method is executed to determine whether to delay or skip the encoding and transmission of subsequent frames. In an embodiment, the adjustment of quantization parameter (QP) values, target frame size, and maximum frame size is performed based on the network speed available to the client. For example, when the network speed is faster, the QP can be reduced. In other embodiments, monitoring of the I-frame generation rate is performed and used for QP setting. For example, when the frequency of I-frames is low, the QP can be reduced (e.g., resulting in higher encoding accuracy or higher quality of encoding), while sacrificing video playback quality, so that the encoding of video frames can be skipped to keep the one-way latency low. Thus, the encoder adjustment performed to improve the trade-off between high-performance encoding and decoding and the latency and video quality of cloud gaming applications leads to a reduction in one-way latency between the cloud gaming server and the client, a smoother frame rate, and a more reliable and / or consistent one-way latency.
[0016] With a general understanding of the various embodiments described above, exemplary details of the embodiments are described with reference to the various drawings.
[0017] Throughout the specification, references to "game" or "video game" or "gaming application" are meant to represent any type of interactive application that is directed through the execution of input commands. For purposes of illustration only, interactive applications include applications for games, word processing, video processing, video game processing, and the like. Further, the terms introduced above are compatible with each other.
[0018] Cloud gaming involves executing a video game on a server to generate video frames rendered in the game, which are then sent to a client for display. The timing of operations at both the server and the client can be associated with their respective vertical synchronization (VSYNC) parameters. When the VSYNC signals are properly synchronized and / or offset between the server and / or the client, operations executed at the server (e.g., generation and transmission of video frames over one or more frame periods) are synchronized with operations executed at the client (e.g., displaying video frames on a display at a display frame or refresh rate corresponding to the frame period). In particular, the server VSYNC signal generated at the server and the client VSYNC signal generated at the client can be used to synchronize operations at the server and the client. That is, when the VSYNC signals of the server and the client are synchronized and / or offset, the server generates and transmits video frames in synchronization with how the client displays those video frames.
[0019] VSYNC signals and vertical batten intervals (VBIs) are incorporated to generate video frames and display them when streaming media content between a server and a client. For example, the server attempts to generate video frames rendered for the game in one or more frame periods defined by the corresponding server VSYNC signal (for example, if the frame period is 16.7ms, generating a video frame every frame period results in a 60Hz operation, and generating one video frame every two frame periods results in a 30Hz operation), then encodes those video frames and transmits them to the client. The client decodes and displays the received encoded video frames, and the client displays each video frame rendered for display, starting with the corresponding client VSYNC.
[0020] For illustrative purposes, Figure 1A shows how the VSYNC signal 111 may indicate the start of a frame period, during which various operations may be performed on the server and / or client during the corresponding frame period. When streaming media content, the server can use the server VSYNC signal to generate and encode video frames, and the client can use the client VSYNC signal to display video frames. The VSYNC signal 111 is generated at a specified frequency corresponding to a specified frame period 110, as shown in Figure 1B. Furthermore, the VBI 105 defines the period between when the last raster line was drawn on the display for the previous frame period and when the first raster line (e.g., top) was drawn on the display. As illustrated, after the VBI 105, the video frame rendered for display is shown via the raster scanlines 106 (e.g., raster line by raster line from left to right).
[0021] Furthermore, various embodiments of this disclosure are disclosed for reducing unidirectional latency and / or latency instability between a source device and a target device, such as when streaming media content (e.g., video game content). For illustrative purposes only, various embodiments for reducing unidirectional latency and / or latency instability are described within a server-client network configuration. However, as shown in Figures 2A to 2D, it will be understood that various techniques disclosed for reducing unidirectional latency and / or latency instability may be implemented within other network configurations and / or on peer-to-peer networks. For example, various embodiments disclosed for reducing unidirectional latency and / or latency instability may be implemented between one or more server-client devices in various configurations (e.g., server-client, server-server, server-multiple clients, server-multiple servers, client-client, client-multiple clients, etc.).
[0022] Figure 2A is a diagram of a system 200A for providing games between one or more cloud gaming networks 290 and / or servers 260 and one or more client devices 210 via a network 250 in various configurations, wherein, according to one embodiment of the present disclosure, the VSYNC signals of the server and client can be synchronized and offset, and / or dynamic buffering is performed on the client, and / or encoding and transmission operations on the server can be duplicated, and / or receiving and decoding operations on the client can be duplicated, and / or decoding and display operations on the client can be duplicated, thereby reducing one-way latency between the server 260 and the client 210. In particular, according to one embodiment of the present disclosure, System 200A provides games via a cloud gaming network 290, and the games are run remotely from the client devices 210 (e.g., thin clients) of the corresponding users playing the games. System 200A can provide game control to one or more users playing one or more games via the cloud gaming network 290 through Network 250, in either single-player mode or multiplayer mode. In some embodiments, the cloud gaming network 290 may include a plurality of virtual machines (VMs) running on the hypervisor of a host machine, one or more of which are configured to run game processor modules utilizing the hardware resources available to the hypervisor of the host machine. Network 250 may include one or more communication technologies. In some embodiments, Network 250 may include fifth-generation (5G) network technology having an advanced wireless communication system.
[0023] In some embodiments, communication may be facilitated using wireless technology. Such technology may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. A 5G network is a digital cellular network, where the service area covered by a provider is divided into small geographical areas called cells. Analog signals representing sound and images are digitized by the telephone, converted by an analog-to-digital converter, and transmitted as a stream of bits. All 5G wireless devices within a cell communicate radio waves using local antenna arrays and low-power automatic transceivers (transmitters and receivers) within the cell, via frequency channels allocated by transceivers from a pool of frequencies reused by other cells. The local antennas are connected to the telephone network and the internet by high-bandwidth optical fiber or wireless backhaul connections. As with other cell networks, mobile devices moving from one cell to another are automatically transferred to the new cell. It should be understood that a 5G network is merely an exemplary type of communication network, and embodiments of this disclosure may utilize previous generations of wireless or wired communications, as well as subsequent generations of wired or wireless technologies following 5G.
[0024] As illustrated, the cloud gaming network 290 includes a game server 260 that provides access to multiple video games. The game server 260 can be any type of server computing device available in the cloud and can be configured as one or more virtual machines running on one or more hosts. For example, the game server 260 can manage virtual machines that support game processors that instantiate instances of games for users. Thus, multiple game processors of the game server 260 associated with multiple virtual machines are configured to run multiple instances of one or more games associated with the gameplay of multiple users. In this way, backend server support provides streaming of gameplay media (e.g., video, audio, etc.) for multiple game applications to multiple corresponding users. Specifically, the game server 260 is configured to stream data (e.g., rendered images and / or frames of the corresponding gameplay) back to the corresponding client devices 210 via the network 250. In this way, computationally complex game applications can be run on the backend server in response to controller inputs received and transmitted by the client devices 210. Each server can render images and frames, which can then be encoded (e.g., compressed) and streamed to the corresponding client devices for display.
[0025] For example, multiple users can access the cloud gaming network 290 via a communication network 250 using corresponding client devices 210 configured to receive streaming media. In one embodiment, the client device 210 may be configured as a thin client providing an interface with a backend server (e.g., a game server 260 of the cloud gaming network 290) configured to provide computing functions (e.g., including a game title processing engine 211). In another embodiment, the client device 210 may consist of a game title processing engine and game logic for at least some local processing of the video game, and may be further utilized to receive streaming content generated by the video game running on the backend, or for other content provided by backend server support. For local processing, the game title processing engine includes basic processor-based functions for running the video game and services related to the video game. The game logic is stored on the local client device 210 and used to run the video game.
[0026] In particular, a client device 210 of a corresponding user (not shown) is configured to request access to the game via a communication network 250 such as the Internet, and to render display images generated by the video game run by the game server 260. The encoded images are then delivered to the client device 210 for display in association with the corresponding user. For example, a user may interact with an instance of a video game running on the game processor of the game server 260 via a client device 210. More specifically, the instance of the video game is executed by a game title processing engine 211. The corresponding game logic (e.g., executable code) 215 that implements the video game is stored in a data store (not shown), accessible via the data store, and used to run the video game. The game title processing engine 211 can support multiple video games using multiple game logics, each of which is selectable by the user.
[0027] For example, the client device 210 is configured to interact with the corresponding user's gameplay in relation to the gameplay of the game title processing engine 211, such as through input commands used to drive gameplay. In particular, the client device 210 can receive input from various types of input devices such as game controllers, tablet computers, keyboards, gestures captured by video cameras, mice, and touchpads. The client device 210 may be any type of computing device having at least a memory and processor module that can connect to the game server 260 via the network 250. The backend game title processing engine 211 is configured to generate rendered images, which are delivered via the network 250 for display on a corresponding display associated with the client device 210. For example, through a cloud-based service, images rendered for a game may be delivered by a corresponding instance of the game running on the game execution engine 211 of the game server 260. That is, the client device 210 is configured to receive encoded images (e.g., encoded from game-rendered images generated through the execution of a video game) and display the rendered images for display 11. In one embodiment, the display 11 includes an HMD (e.g., for displaying VR content). In some embodiments, the rendered images may be streamed wirelessly or via a wired connection to a smartphone or tablet, either directly from the cloud-based service or via the client device 210 (e.g., PlayStation® Remote Play).
[0028] In one embodiment, the game server 260 and / or game title processing engine 211 include basic processor-based functions for performing services related to the game and game application. For example, processor-based functions include 2D or 3D rendering, physics, physics simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadowing, culling, transformation, artificial intelligence, etc. Furthermore, game application services include memory management, multithreading management, quality of service (QoS), bandwidth testing, social networking, social friend management, communication with friends' social networks, communication channels, text messaging, instant messaging, chat support, etc.
[0029] In one embodiment, the cloud gaming network 290 is a distributed game server system and / or architecture. In particular, a distributed game engine that executes game logic is configured as a corresponding instance of the corresponding game. Generally, the distributed game engine acquires each function of the game engine and distributes those functions for execution by a number of processing entities. Individual functions may be further distributed across one or more processing entities. Processing entities may be configured in various configurations, including physical hardware, and / or as virtual components or virtual machines, and / or as virtual containers, where a container differs from a virtual machine in that it virtualizes an instance of a game application running on a virtualized operating system. Processing entities can utilize and / or depend on servers and their underlying hardware on one or more servers (computing nodes) of the Cloud Gaming Network 290, and servers can be located in one or more racks. The coordination, allocation, and management of the execution of these functions to various processing entities are performed by a distributed synchronization layer. In this way, the execution of these functions is controlled by the distributed synchronization layer, making it possible to generate media for game applications (e.g., video frames, audio, etc.) in response to controller input from players. The distributed synchronization layer enables these functions to be executed efficiently across distributed processing entities (e.g., through load balancing), so that critical game engine components / functions are distributed and reconfigured for more efficient processing.
[0030] The game title processing engine 211 includes a central processing unit (CPU) and a group of graphics processing units (GPUs) which may be configured to perform multi-tenancy GPU functions. In another embodiment, multiple GPU devices are combined to perform graphics processing for a single application running on the corresponding CPU.
[0031] Figure 2B is a diagram illustrating an embodiment of the present disclosure for providing a game between two or more peer devices, where the VSYNC signal is synchronized and offset to achieve optimal timing for receiving other information between the controller and the devices. For example, a head-to-head game may be run using two or more peer devices directly connected via network 250 or via peer-to-peer communication (e.g., Bluetooth®, local area networking, etc.).
[0032] As illustrated, the game runs locally on each of the corresponding user's client devices 210 (e.g., game consoles) playing the video game, and the client devices 210 communicate via peer-to-peer networking. For example, an instance of the video game is executed by the game title processing engine 211 on the corresponding client device 210. The game logic 215 (e.g., executable code) that implements the video game is stored on the corresponding client device 210 and used to run the game. For illustrative purposes, the game logic 215 may be delivered to the corresponding client device 210 via a portable medium (e.g., optical medium) or via a network (e.g., downloaded from a game provider via the Internet).
[0033] In one embodiment, the game title processing engine 211 of the corresponding client device 210 includes basic processor-based functions for performing services related to the game and game application. For example, processor-based functions include 2D or 3D rendering, physics, physics simulation, scripting, audio, animation, graphics processing, lighting, shading, rasterization, ray tracing, shadowing, culling, transformation, artificial intelligence, etc. Furthermore, game application services include memory management, multithreading management, quality of service (QoS), bandwidth testing, social networking, social friend management, communication with friends' social networks, communication channels, text messaging, instant messaging, chat support, etc.
[0034] The client device 210 can receive input from various types of input devices, such as game controllers, tablet computers, keyboards, gestures captured by video cameras, mice, and touchpads. The client device 210 may be any type of computing device having at least memory and a processor module, and is configured to generate rendering images executed by the game title processing engine 211 and to display the rendering images on a display (e.g., display 11, or display 11 including a head-mounted display (HMD)). For example, a rendered image may be associated with an instance of a game running locally on a client device 210, implementing the corresponding user's gameplay through input commands used to drive the gameplay. Some examples of client devices 210 include personal computers (PCs), game consoles, home theater devices, general-purpose computers, mobile computing devices, tablets, phones, or any other type of computing device capable of running an instance of a game.
[0035] Figure 2C illustrates various network configurations, including those shown in Figures 2A and 2B, which benefit from proper synchronization and offset of VSYNC signals between source and target devices according to embodiments of the present disclosure. In particular, various network configurations benefit from proper frequency alignment of server and client VSYNC signals, and timing offset of server and client VSYNC signals for the purpose of reducing unidirectional latency and / or latency variability between the server and client. For example, one network device configuration includes a configuration from a cloud gaming server (e.g., source) to a client (target). In one embodiment, the client may include a web RTC client configured to provide audio and video communication within a web browser. Another network configuration includes a configuration from a client (e.g., source) to a server (target). Yet another network configuration includes a configuration from a server (e.g., source) to a server (e.g., target). Another network device configuration includes a configuration from a client (e.g., source) to a client (target), where each client could be, for example, a game console for providing a head-to-head game.
[0036] In particular, VSYNC signal alignment may include frequency synchronization of the server VSYNC signal and the client VSYNC signal, and may also include adjusting the timing offset between the client VSYNC signal and the server VSYNC signal to maintain an ideal relationship between the server and client VSYNC signals for the purpose of eliminating drift and / or reducing unidirectional latency and / or latency variability. In order to achieve proper alignment, in one embodiment, the server VSYNC signal may be adjusted to perform proper alignment between the server 260 and client 210 pair. In another embodiment, the client VSYNC signal may be adjusted to perform proper alignment between the server 260 and client 210 pair. Once the client and server VSYNC signals are aligned, the server VSYNC signal and the client VSYNC signal occur at substantially the same frequency and are offset from each other by a timing offset that can be adjusted as needed. In another embodiment, VSYNC signal alignment may include synchronizing the frequencies of the VSYNC signals of two clients, adjusting the timing offset between those VSYNC signals for the purpose of eliminating drift, and / or achieving proper timing for receiving controller and other information, and either VSYNC signal may be adjusted to achieve this alignment. In yet another embodiment, alignment may include synchronizing the VSYNC frequencies of multiple servers, and may also include synchronizing the frequencies of the server VSYNC signal and the client VSYNC signal, and adjusting the timing offset between the client VSYNC signal and the server VSYNC signal, for example, for head-to-head cloud gaming. In server-to-client and client-to-client configurations, alignment may include both synchronizing the frequencies between the server VSYNC signal and the client VSYNC signal, and providing an appropriate timing offset between the server VSYNC signal and the client VSYNC signal. In a server-to-server configuration, alignment may include synchronizing the frequencies between the server VSYNC signal and the client VSYNC signal without setting a timing offset.
[0037] Figure 2D shows a multi-tenancy configuration between a cloud gaming server 260 and one or more clients 210, according to one embodiment of the present disclosure, which benefits from proper synchronization and offset of VSYNC signals between a source device and a target device. In a server-to-client configuration, alignment may include both frequency synchronization between the server VSYNC signal and the client VSYNC signal, and providing a proper timing offset between the server VSYNC signal and the client VSYNC signal. In a multi-tenancy configuration, in one embodiment, the client VSYNC signal is tuned at each client 210 to implement proper alignment between the server 260 and the client 210 pair.
[0038] For example, a graphics subsystem may be configured to perform multi-tenancy GPU functions, and in one embodiment, one graphics subsystem may be the graphics implementation and / or the rendering of a pipeline for multiple games. That is, the graphics subsystem is shared among multiple games being run. In particular, a game title processing engine may include a group of CPUs and GPUs configured to perform multi-tenancy GPU functions, and in one embodiment, one group of CPUs and GPUs may be the graphics implementation and / or the rendering of a pipeline for multiple games. That is, the group of CPUs and GPUs is shared among multiple games being run. The group of CPUs and GPUs can be configured as one or more processing devices. In another embodiment, multiple GPU devices are combined to perform graphics processing for a single application running on the corresponding CPUs.
[0039] Figure 3 illustrates a typical process in which a server runs a video game, generates video frames rendered for the game, and sends those video frames to a client for display. Conventionally, several operations in the game server 260 and client 210 are performed within frame durations defined by their respective VSYNC signals. For example, server 260 attempts to generate video frames rendered for the game in 301 within one or more frame durations, as defined by the corresponding server VSYNC signal 311. Video frames are generated by the game in response to either control information delivered from the input device in operation 350 (e.g., user input commands) or game logic not driven by control information. Transmission jitter 351 may be present when sending control information to the server 260, and jitter 351 measures fluctuations in network latency from the client to the server (e.g., when sending input commands). As illustrated, the thick arrows indicate the current delay when sending control information to server 260, but due to jitter, there may be a range of arrival times for the control information at server 260 (e.g., the range enclosed by the dotted arrows). At flip time 309, the GPU receives a flip command indicating that the corresponding video frame has been fully generated and placed in the frame buffer of server 260. Server 260 then performs scan-out / scan-in (operation 302, where the scan-out may be aligned with VSYNC signal 311) on that video frame over a subsequent frame duration defined by the server VSYNC signal 311 (VBI is omitted for clarity). Subsequently, the video frame is encoded (operation 303) (e.g., encoding starts after the occurrence of VSYNC signal 311, and the end of encoding does not have to be aligned with VSYNC signal) and transmitted to client 210 (operation 304, where the transmission does not have to be aligned with VSYNC signal 311). In client 210, the encoded video frame is received (operation 305, where reception does not necessarily have to be aligned with the client VSYNC signal 312), decoded (operation 306, where decoding does not necessarily have to be aligned with the client VSYNC signal 312), buffered, and displayed (operation 307, where the start of display does not necessarily have to be aligned with the client VSYNC signal 312). In particular, client 210 displays each rendered video frame for display, which begins with the corresponding occurrence of the client VSYNC signal 312.
[0040] One-way latency 315 can be defined as the latency from the start of the transfer of the video frame to the encoding unit (e.g., scanout 302) at the server to the start of the display of the video frame at the client 307. In other words, one-way latency is the time from the server's scanout to the client's display, taking client buffering into account. Each individual frame has latency from the start of scanout 302 to the completion of decoding 306, and this latency may vary from frame to frame due to significant variations in server operations such as encoding 303 and transmission 304, network transmission with jitter 352 between the server 260 and the client 210, and client reception 305. As illustrated, the thick straight arrows indicate the current latency when sending the corresponding video frame to client 210, but due to jitter 352, there can be a range in the arrival time of the video frame at client 210 (e.g., the range demarcated by the dotted arrows). For a good playback experience, the one-way latency needs to be relatively stable (e.g., fairly consistent), so conventionally, buffering 320 is performed with the result that the display of individual frames with low latency (e.g., from the start of scanout 302 to the completion of decoding 306) is delayed over a period of several frames. In other words, if there is network instability or unpredictable encoding / decoding time, additional buffering is required to keep the one-way latency constant.
[0041] According to one embodiment of the present disclosure, one-way latency between a cloud gaming server and a client can vary due to clock drift when streaming video frames generated from a video game running on the server. Specifically, a difference in frequency between the server VSYNC signal 311 and the client VSYNC signal 312 can result in a client VSYNC signal drifting relative to frames arriving from the server 260. This drift may be due to very small differences in the crystal oscillators used for the respective clocks in the server and the client. Furthermore, embodiments of the present disclosure reduce one-way latency by performing one or more synchronization and offset of the VSYNC signals for alignment between the server and the client, by providing dynamic buffering at the client, by duplicating the encoding and transmission of video frames at the server, by duplicating the reception and decoding of video frames at the client, and by duplicating the decoding and display of video frames at the client.
[0042] Furthermore, during video frame encoding (operation 303), in previous techniques, the encoder would determine how much change there is between the current video frame being encoded and one or more previously encoded frames to determine if there is a scene change (e.g., a complex image in the corresponding generated video frame). In other words, a scene change hint can be inferred from the difference between the current frame being encoded and the previously encoded frames. When streaming content from a server to a client over a network, the encoder on the server may decide to encode video frames that are detected as scene changes with complexity. Otherwise, the encoder encodes video frames that are not detected as scene changes with lower complexity. However, scene change detection in an encoder can take up to one frame duration (e.g., added jitter) for the video frame to be initially encoded with lower complexity (in the first frame duration), and then, if a scene change is detected, re-encoded with higher complexity (in the second frame duration). Similarly, scene change detection can be unnecessarily triggered even if there is no scene change, simply because the difference between the currently encoded video frame and a previously encoded video frame exceeds a threshold difference value (e.g., via small bursts in the image). Thus, when a scene change is detected in the encoder, additional latency due to jitter is introduced to accommodate the execution of scene change detection and the re-encoding of the video frame with higher complexity.
[0043] Figure 4 shows the data flow through a network configuration including a highly optimized cloud gaming server 260 and a highly optimized client 210 when streaming video frames generated from a video game running on a server. According to embodiments of this disclosure, the overlap of server and client operations reduces unidirectional latency, the synchronization and offset of VSYNC signals between the server and client reduces unidirectional latency, and simultaneously reduces the variability of unidirectional latency between the server and client. In particular, Figure 4 shows the desired alignment between the server VSYNC signal and the client VSYNC signal. In one embodiment, the adjustment of the server VSYNC signal 311 is performed to obtain appropriate alignment between the server VSYNC signal and the client VSYNC signal in the server and client network configuration, etc. In another embodiment, the adjustment of the client VSYNC signal 312 is performed to obtain proper alignment between the server VSYNC signal and the client VSYNC signal, such as in a network configuration from a multi-tenant server to multiple clients. For illustrative purposes, the adjustment of the server VSYNC signal 311 is illustrated in Figure 4 for the purpose of synchronizing the frequencies of the server and client VSYNC signals and / or adjusting the timing offset between the corresponding client and server VSYNC signals, but it is understood that the client VSYNC signal 312 may also be used for adjustment. In the context of this patent application, “synchronization” should be interpreted as meaning adjusting the signals so that their frequencies match but their phases may differ, and “offset” should be interpreted as meaning a time delay between signals, such as the time between when one signal reaches its maximum value and when the other signal reaches its maximum value.
[0044] As illustrated, Figure 4 shows an improved process in an embodiment of the present disclosure for running a video game on a server to generate rendered video frames and sending those video frames to a client for display. The process is shown with respect to the generation and display of a single video frame on the server and the client. In particular, the server generates video frames rendered for the game in 401. For example, server 260 includes a CPU (e.g., a game title processing engine 211) configured to run the game. The CPU generates one or more draw calls for the video frame, which include commands placed in a command buffer for execution by the corresponding GPU of server 260 in the graphics pipeline. The graphics pipeline includes one or more shader programs at the vertices of objects in the scene, which can generate rendered texture values for video frames for display. This operation is performed in parallel via the GPU for increased efficiency. At flip time 409, the GPU reaches the flip command in the command buffer, indicating that the corresponding video frame has been fully generated and / or rendered and placed in the frame buffer on server 260.
[0045] In 402, the server performs a scanout of the video frames rendered for the game to the encoder. In particular, the scanout is performed scanline by scanline or in groups of consecutive scanlines, where a scanline refers to, for example, a single horizontal line from one edge of the screen to the other. These scanlines or groups of consecutive scanlines may also be referred to as slices, and in this specification are referred to as screen slices. In particular, the scanout 402 may include several processes, which include modifying the frames rendered for the game and overlaying them with another frame buffer, or shrinking them to surround them with information from another frame buffer. During the scanout 402, the modified video frames are then scanned to the encoder for compression. In one embodiment, the scanout 402 is performed at the occurrence 311a of the VSYNC signal 311. In other embodiments, the scanout 402 may be performed before the occurrence of the VSYNC signal 311, such as at the flip time 409.
[0046] In 403, the video frames rendered for the game (which may have been modified) are encoded by the encoder into one encoder slice per encoder slice, generating one or more encoded slices, where the encoded slices are independent of scanlines or screen slices. Thus, the encoder generates one or more encoded (e.g., compressed) slices. In one embodiment, the encoding process begins before the scanout process 402 is fully completed for the corresponding video frames. Furthermore, the start and / or end of encoding 403 may or may not be aligned with the server VSYNC signal 311. The boundaries of an encoded slice are not limited to a single scanline, but may consist of a single scanline or multiple scanlines. Furthermore, the end of an encoded slice and / or the start of the next encoder slice do not necessarily occur at the edge of the display screen (for example, they may occur in the center of the screen or in the center of a scanline), and an encoded slice does not need to completely traverse the display screen from edge to edge. As illustrated, one or more encoded slices may be compressed and / or encoded, including a compressed "encoded slice A" having a hash mark.
[0047] In 404, the encoded video frame is transmitted from the server to the client, and this transmission can be performed slice by slice, each encoded slice being a compressed encoder slice. In one embodiment, the transmission process 404 begins before the encoding process 403 is fully completed for the corresponding video frame. Furthermore, the start and / or end of transmission 404 may or may not be aligned with the server VSYNC signal 311. As illustrated, the compressed encoded slice A is transmitted to the client independently of other compressed encoder slices for the rendered video frame. Encoder slices may be transmitted one at a time or in parallel.
[0048] At 405, the client receives the compressed video frame, which is also done for each encoded slice. Furthermore, the start and / or end of reception 405 may or may not be aligned with the client VSYNC signal 312. As illustrated, the compressed encoded slice A is received by the client. Transmission jitter 452 may exist between the server 260 and the client 210, and the jitter 452 measures the variation in network latency from the server 260 to the client 210. Lower jitter values indicate a more stable connection. As illustrated, the thick straight arrows indicate the current latency when sending the corresponding video frame to client 210, but jitter can result in a range of arrival times for the video frame at client 210 (e.g., the range enclosed by the dotted arrows). Latency fluctuations can also be due to one or more operations on the server, such as encoding 403 or transmission 404, and similarly, network issues that result in latency when transmitting the video frame to client 210.
[0049] In 406, the client decodes the compressed video frame again for each encoded slice to produce a decoded slice A (shown without a hash mark) that is now ready for display. In one embodiment, the decoding process 406 is started before the receiving process 405 is fully completed for the corresponding video frame. Furthermore, the start and / or end of decoding 406 may or may not be aligned with the client VSYNC signal 312. In 407, the client displays the decoded rendered video frame on a display in the client. That is, the decoded video frame is placed in a display buffer that is streamed out to the display device, for example, scanline by scanline. In one embodiment, the display process 407 (i.e., streaming out to a display device) begins after the decoding process 406 has been fully completed for the corresponding video frame, i.e., after the decoded video frame has fully resided in the display buffer. In another embodiment, the display process 407 begins before the decoding process 406 is fully completed for the corresponding video frame. That is, the stream out to the display device begins at the address of the display buffer when only a portion of the decoded frame buffer resides in the display buffer. The display buffer is then updated or filled with the remainder of the corresponding video frame in time for display, and the update of the display buffer is performed before those portions are streamed out to the display. Furthermore, the start and / or end of the display 407 are aligned with the client VSYNC signal 312.
[0050] In one embodiment, the unidirectional latency 416 between the server 260 and the client 210 may be defined as the elapsed time between when the scanout 402 is initiated and when the display 407 is initiated. Embodiments of this disclosure can align the VSYNC signals between the server and the client (e.g., synchronize the frequencies and adjust the offset) to reduce the unidirectional latency between the server and the client and reduce the variability of the unidirectional latency between the server and the client. For example, embodiments of the present disclosure can calculate an optimal adjustment for the offset 430 between the server VSYNC signal 311 and the client VSYNC signal 312, so that even if the near worst-case time required for server processing such as encoding 403 and transmission 404 occurs, even if the near worst-case network latency between the server 260 and the client 210 occurs, and even if the near worst-case client processing such as receiving 405 and decoding 406 occurs, the decoded and rendered video frame is available in time for the display process 407. In other words, it is not necessary to determine an absolute offset between the server VSYNC and the client VSYNC; it is sufficient to adjust the offset so that the decoded and rendered video frame is available in time for the display process.
[0051] In particular, the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 can be synchronized. Synchronization is achieved by adjusting either the server VSYNC signal 311 or the client VSYNC signal 312. For illustrative purposes, the adjustment is described in relation to the server VSYNC signal 311, but it should be understood that the adjustment may instead be performed on the client VSYNC signal 312. For example, as shown in Figure 4, the server frame period 410 (e.g., the time between 311c and 311d, which are two occurrences of the server VSYNC signal 311) is substantially equal to the client frame period 415 (e.g., the time between 312a and 312b, which are two occurrences of the client VSYNC signal 312), which indicates that the frequencies of the server VSYNC signal 311 and the client VSYNC signal 312 are also substantially equal.
[0052] To maintain frequency synchronization between the server and client VSYNC signals, the timing of the server VSYNC signal 311 can be manipulated. For example, the vertical blanking interval (VBI) of the server VSYNC signal 311 can be increased or decreased over a period of time to check for drift between the server VSYNC signal 311 and the client VSYNC signal 312. Manipulating the vertical blanking (VBLANK) in the VBI provides a way to adjust the number of scan lines used for VBLANK for one or more frame periods of the server VSYNC signal 311. Decreasing the number of scan lines in VBLANK shortens the corresponding frame period (e.g., time interval) between two occurrences of the server VSYNC signal 311. Conversely, increasing the number of scan lines in VBLANK lengthens the corresponding frame period (e.g., time interval) between two occurrences of the VSYNC signal 311. In this way, the frequency of the server VSYNC signal 311 is adjusted so that the frequencies between the client and server VSYNC signals 311 and 312 are substantially the same. Additionally, the offset between the server and client VSYNC signals can be adjusted by briefly increasing or decreasing the VBI before returning the VBI to its original value. In one embodiment, the server VBI is configured. In another embodiment, the client VBI is configured. In yet another embodiment, instead of two devices (server and client), there may be multiple connected devices, each of which may have a configured corresponding VBI. In one embodiment, each of the multiple connected devices may be an independent peer device (e.g., without a server device). In another embodiment, the multiple devices may include one or more server devices and / or one or more client devices arranged in one or more server / client architectures, multitenant server / client(s) architectures, or some combination thereof.
[0053] In another form, the server's pixel clock (for example, located in the southbridge of the server's northbridge / southbridge core logic chipset, or, in the case of a discrete GPU, generating its own pixel clock using its own hardware) can, in one embodiment, be operated to perform coarse and / or fine adjustments of the frequency of the server VSYNC signal 311 over a period of time, restoring the frequency synchronization between the server VSYNC signal 311 and the client VSYNC signal 312 to an aligned state. Specifically, the pixel clock of the server's southbridge can be overclocked or underclocked to adjust the overall frequency of the server's VSYNC signal 311. In this way, the frequency of the server VSYNC signal 311 is adjusted so that the frequencies between the client and server VSYNC signals 311 and 312 are substantially the same. The offset between the server and client VSYNC can be adjusted by briefly increasing or decreasing the client-server pixel clocks before returning the pixel clocks to their original values. In one embodiment, the server pixel clock is adjusted. In another embodiment, the client pixel clock is regulated. In yet another embodiment, instead of two devices (server and client), there may be multiple connected devices, each of which may have a corresponding regulated pixel clock. In one embodiment, each of the multiple connected devices may be an independent peer device (e.g., without a server device). In another embodiment, the multiple connected devices may include one or more server devices and one or more client devices arranged in one or more server / client architectures, multitenant server / client(s) architectures, or some combination thereof.
[0054] In one embodiment, high-performance codecs (e.g., encoders and / or decoders) may be used to further reduce one-way latency between the cloud gaming server and the client. In conventional streaming systems involving the streaming of compressed media (e.g., streaming movies, TV shows, videos), when the streaming media is decompressed at the end target (e.g., the client), a significant portion of the decompressed video can be buffered at the client to accommodate jitter intruding into transmission quality during encoding operations (e.g., longer encoding times) and variations in decoding operations (e.g., longer decoding times). Thus, conventional streaming systems can rely on average decoding capabilities and metrics (e.g., average decoding resources) because the decoded content handles latency variability, allowing video frames to be displayed at the desired speed (e.g., supporting 4K media at 60Hz or displaying video frames every time a client VSYNC signal is generated).
[0055] However, buffering is severely limited in cloud gaming environments (e.g., transitioning to zero buffering), enabling real-time gaming. As a result, the variability introduced into the one-way latency between the cloud gaming server and client can negatively impact downstream operations. For example, if encoding or decoding complex frames takes a long time (even for a single frame), the one-way latency increases accordingly, ultimately leading to longer response times for the user and negatively impacting the user's real-time experience.
[0056] In one embodiment, for cloud gaming, it is beneficial to provide more powerful decoding and encoding resources that may seem unnecessary compared to the requirements of a streaming video application. Furthermore, as will be explained in more detail below, encoder resources need to be optimized for the time required to process long or the longest processing frames. In other words, in an embodiment where the encoder can be tuned to improve the trade-off between one-way latency and video quality in a cloud gaming system, the encoder tuning is obtained based on monitoring client bandwidth, skipped frames, the number of encoded I-frames, the number of scene changes, and / or the number of video frames exceeding the target frame size, and the tuned parameters may include encoder bitrate, target frame size, maximum frame size, and quantization parameter (QP) values, and high-performance encoders and decoders help reduce overall one-way latency between the cloud gaming server and the client.
[0057] Along with a detailed description of various client devices 210 and / or cloud gaming network 290 (e.g., within a game server 260) in Figures 2A to 2D, the flowchart 500 in Figure 5 illustrates a cloud gaming method according to one embodiment of the present disclosure, in which video frame encoding includes adjustment of encoder parameters that are aware of network transmission speed and reliability, as well as overall latency targets. The cloud gaming server is configured to stream content to one or more client devices over the network. This process results in smoother frame rates and more reliable latency, reducing and making the one-way latency between the cloud gaming server and the client more consistent, thereby improving the smoothness of the client display of video.
[0058] In 510, when a video game is run on a cloud gaming server, multiple video frames are generated. Generally, a cloud gaming server generates video frames rendered for multiple games. For example, the game logic of a video game is built upon a game engine or game title processing engine. The game engine contains core functions that can be used by the game logic to build the game environment of the video game. For example, some functions of a game engine may include a physics engine to simulate physical forces and collisions on objects in the game environment, a rendering engine for 2D or 3D graphics, collision detection, sound, animation, artificial intelligence, networking, streaming, etc. In this way, the game logic does not need to build the core functions provided by the game engine from scratch.
[0059] The game logic, combined with the game engine, is executed by the CPU and GPU, which can be configured within an Accelerator Processing Unit (APU). In other words, the CPU and GPU, along with shared memory, can be configured as a rendering pipeline to generate video frames rendered for the game, and the rendering pipeline outputs the game-rendered image as a video or image frame suitable for display, containing color information corresponding to each pixel of the targeted and / or virtualized display. In particular, the CPU may be configured to generate one or more draw calls for a video frame, each draw call containing a command stored in a corresponding command buffer that is executed by the GPU in the GPU pipeline. Generally, the graphics pipeline can perform shader operations on the vertices of objects in the scene to generate texture values for the pixels of the display. Specifically, the graphics pipeline receives input geometry (e.g., the vertices of objects in a gaming environment), and the vertex shader constructs the primitives or polygons that make up the object. The vertex shader program can perform lighting, shading, shadowing, and other operations on the primitives. Depth buffering, or Z-buffering, is performed to determine which objects are visible when rendered from a corresponding viewpoint. Rasterization is performed to project objects in a 3D game environment onto a 2D plane defined by the viewpoint. Pixel-sized fragments are generated for objects, and one or more fragments can contribute to the color of pixels in an image. Fragments can be merged and / or blended to determine the combined color of each pixel in the corresponding video and can be stored in a frame buffer. Subsequent video frames are generated and / or rendered for display using a similarly configured command buffer, and multiple video frames are output from the GPU pipeline.
[0060] In 520, the method includes encoding multiple video frames at an encoder bitrate. Specifically, the multiple video frames are scanned into an encoder to be compressed before being streamed to the client using a streamer operating in the application layer. In one embodiment, each video frame rendered for a game may be combined and blended with a corresponding modified video frame with additional user interface functionality, and then scanned into an encoder, which compresses the modified video frame and streams it to the client. For brevity and clarity, the method for adjusting the encoder parameters disclosed in Figure 5 is described in relation to encoding multiple video frames, but is understood to support encoding modified video frames. The encoder is configured to compress multiple video frames based on the described format. For example, when streaming media content from a cloud gaming server to a client, the Motion Picture Expert Group (MPEG) or H.264 standard may be implemented. In particular, the encoder can perform compression by video frame or by encoder slices of video frames, and as mentioned above, each video frame can be compressed as one or more encoded slices. Generally, when streaming media, video frames are compressed as I-frames (intra-frames) or P-frames (predictive frames), and each of them can be divided into encoded slices.
[0061] At 530, the client's maximum receiving bandwidth is measured. In one embodiment, the maximum bandwidth experienced by the client is determined by a feedback mechanism from the client. Figure 6 shows a measurement of the bandwidth of client 210 by a streamer of a cloud gaming server according to one embodiment of the present disclosure, where streamer 620 is configured to monitor and adjust encoder 610 so that compressed video frames can be transmitted at a speed within the range of the client's measured bandwidth. As illustrated, compressed video frames, encoded slices, and / or packets are delivered from encoder 610 to buffer 630 (e.g., first-in, first-out - FIFO). The encoder delivers compressed video frames at encoder filling rate 615. For example, the buffer may be filled at the same rate that the encoder can generate compressed video frames, encoded slices 650, and / or packets 655 of the encoded slices. Furthermore, the compressed video frames are discharged from the buffer at buffer discharge rate 635 for delivery to client 210 over network 250. In one embodiment, the buffer discharge rate 635 is dynamically adjusted to the client's measured maximum receive bandwidth. For example, the buffer discharge rate 635 may be adjusted to be approximately equal to the client's measured maximum receive bandwidth. In one embodiment, packet encoding is performed at the same rate at which they are transmitted, and both operations are dynamically adjusted to the maximum available bandwidth available to the client.
[0062] In particular, the streamer 620, operating at the application layer, measures the maximum bandwidth of the client 210, for example, by using the bandwidth tester 625. The application layer is used in the User Datagram Protocol / Internet Protocol (UDP / IP) suite of protocols used to interconnect network devices over the internet. For example, the application layer defines the communication protocols and interface methods used for communication between devices over an IP communication network. During testing, the streamer 620 provides additional buffered packets 640 (e.g., forward error correction (FEC) packets), allowing the buffer 630 to stream packets from a predefined bitrate, such as the maximum bandwidth being tested. In one embodiment, the client returns to the streamer 620 as feedback 690 the number of packets received over a range of incremental sequence identifiers (IDs), such as a range of video frames. For example, the client might report something like 145 of the 150 video frames received in sequence ID 100-250 (e.g., 150 video frames). In this way, the streamer 620 on the server 260 can calculate packet loss, and since the streamer 620 knows the amount of bandwidth transmitted (e.g., tested) during the sequence of packets, it can dynamically determine what the client's maximum bandwidth is at a given point in time. The client's measured maximum bandwidth may be delivered from the streamer 620 to the buffer 630 as control information 627, allowing the buffer 630 to dynamically transmit packets at a rate approximately equal to the client's maximum bandwidth. In this way, the transmission rates of compressed video frames, encoded slices, and / or packets can be dynamically adjusted according to the currently measured client maximum bandwidth.
[0063] In 540, the encoding process is monitored by the streamer; that is, the encoding of multiple video frames is monitored. In one embodiment, monitoring is performed on the client 210, and feedback and / or adjustment control signals are provided to be returned to the encoder. In another embodiment, monitoring is performed in the cloud gaming service 260, such as by the streamer 620. For example, monitoring of video frame encoding may be performed by the monitoring and adjustment unit 629 of the streamer 620. Various encoding characteristics and / or operations can be tracked and / or monitored. For example, in one embodiment, the occurrence rate of I-frames within multiple video frames can be tracked and / or monitored. Furthermore, in one embodiment, the occurrence rate of scene changes within multiple video frames can be tracked and / or monitored. Also, in one embodiment, the number of video frames exceeding the target frame size can be tracked and / or monitored. Also, in one embodiment, the encoder bitrate used to encode one or more video frames can be tracked and / or monitored.
[0064] In the 550, encoder parameters are dynamically adjusted based on monitoring of the video frame encoding. In other words, monitoring the video frame encoding affects how the encoder behaves when compressing current and future video frames received by the encoder. Specifically, the monitoring and adjustment unit 629 is configured to monitor the encoding of video frames and, in response to analysis performed on the monitored information, determine which encoder parameters should be adjusted. A control signal 621 is sent back from the monitoring and adjustment unit 629 to the encoder 610, which is used to configure the encoder. Encoder parameters for adjustment include quantization parameters (QP) (e.g., minimum QP, maximum QP) or quality parameters, target frame size, maximum frame size, etc.
[0065] Adjustments are made with consideration for network transmission speed and reliability, as well as overall latency targets. In one embodiment, smoothness of video playback takes precedence over low latency or image quality. For example, skipping the encoding of one or more video frames is disabled. Specifically, the balance between image resolution or image quality (e.g., 60Hz) and latency is adjusted using various encoder parameters. In particular, since VSYNC signals in the cloud gaming server and client can be synchronized and offset, unidirectional latency between the cloud gaming server and client can be reduced, thereby reducing the need to skip video frames to facilitate low latency. Synchronization and offsetting of VSYNC signals also provide redundant operations in the cloud gaming server (scan out, encode, and transmit), redundant operations in the client (receive, decode, render, display), and / or redundant operations between the cloud gaming server and client, all of which facilitate reduced unidirectional latency, reduced unidirectional latency variability, real-time generation and display of video content, and consistent video playback in the client.
[0066] In one embodiment, the encoder bitrate is monitored to predict the demand on client bandwidth, taking into account the subsequent frames and their complexity (e.g., anticipated scene changes), and the encoder bitrate can be adjusted according to the expected demand. For example, when prioritizing smooth video playback, the encoder monitoring and adjustment unit 629 may be configured to determine that the encoder bitrate being used exceeds the maximum received bandwidth being measured. Accordingly, the encoder bitrate can be reduced, and frame sizing can also be reduced. When smoothness is a priority, it is desirable to use an encoder bitrate lower than the maximum receiving bandwidth (for example, an encoder bitrate of 10 megabits / second for a maximum receiving bandwidth of 15 megabits / second). In this way, even if the encoded frames spike beyond the maximum frame size, the encoded frames can still be transmitted within 60 Hz. In particular, the encoder bitrate can be converted to frame size. A given bitrate and target speed of a video game (e.g., 60 frames per second) is converted to the average size of the encoded video frames. For example, with an encoder bitrate of 15 megabits / second and a given target speed of 60 frames / second, 60 encoded frames will share 15 megabits, and each encoded frame will have approximately 250,000 encoded bits. Thus, controlling the encoder bitrate also controls the frame size of the encoded video frames. As a result, increasing the encoder bitrate allows for more bits to be encoded (higher precision), while decreasing the encoder bitrate allows for fewer bits to be encoded (lower precision). Similarly, when the encoder bitrate used to encode a group of video frames is within the measured maximum receiving bandwidth, the encoder bitrate can be increased, and the frame size can also be increased.
[0067] In one embodiment, when smooth video playback is a priority, the encoder monitoring and adjustment unit 629 may be configured to determine whether the encoder bitrate used to encode a group of video frames from multiple video frames exceeds the measured maximum receive bandwidth. For example, the encoder bitrate may be detected as 15 megabits per second (Mbps), while the maximum receive bandwidth is currently 10 Mbps. In this way, the encoder pushes out more bits than can be sent to the client without increasing one-way latency. As mentioned earlier, when prioritizing smoothness, it may be desirable to use an encoder bitrate lower than the maximum receive bandwidth. In the example above, it may be acceptable to have an encoder bitrate set to 10 megabits / second or less, relative to the maximum receive bandwidth of 10 megabits / second mentioned above. In this way, even if the encoded frames spike above the maximum frame size, the encoded frames can still be transmitted within 60Hz. Accordingly, the QP value can be adjusted with or without reducing the encoder bitrate, and QP controls the precision used when compressing video frames. In short, QP controls the amount of quantization performed (for example, compressing a variable range of values within a video frame into a single quantum value). In H.264, the range of QP is from "0" to "51". For example, a QP value of "0" means less quantization, less compression, higher precision, and higher quality. For example, a QP value of "51" means less quantization, less compression, higher precision, and higher quality. Specifically, the QP value can be increased so that the encoding of the video frame is performed with lower precision.
[0068] In one embodiment, when prioritizing smooth video playback, encoder monitoring by the monitoring and adjustment unit 629 may be configured to determine that the encoder bitrate used to encode a group of video frames from multiple video frames is within the maximum receive bandwidth. As previously mentioned, when prioritizing smoothness, it may be desirable to use an encoder bitrate lower than the maximum receive bandwidth. Thus, there is excess bandwidth available when transmitting a group of video frames. The excess bandwidth can be determined. Accordingly, the QP value can be adjusted, where QP controls the precision used when compressing the video frames. In particular, the QP value can be reduced based on the excess bandwidth so that encoding is performed more accurately.
[0069] In another embodiment, the characteristics of individual video games are taken into consideration when determining I-frame processing and QP settings, particularly when prioritizing smooth video playback. For example, if video game "scene changes" are infrequent (e.g., only camera cuts), it may be desirable to have larger I-frames (lower QP or higher encoder bitrate). That is, within a group of video frames from multiple video frames being compressed, the number of video frames identified as having a scene change is determined to be less than the threshold number of scene changes. In other words, the streaming system can handle the number of scene changes in the current state (e.g., measured client bandwidth, required latency, etc.). Accordingly, the QP value can be adjusted, where QP controls the precision used when compressing video frames. In particular, the QP value can be reduced so that encoding is performed with higher precision.
[0070] On the other hand, if "scene changes" occur frequently during gameplay in a video game, it may be desirable to keep the I-frame size small (e.g., by increasing the QP or lowering the encoder bitrate). In other words, the number of video frames identified as having a scene change within a group of video frames being compressed is determined to meet or exceed the scene change threshold. That is, the video game is generating too many scene changes for the current state (e.g., measured client bandwidth, required latency, etc.). Accordingly, the QP value can be adjusted, where QP controls the precision used when compressing video frames. In particular, the QP value can be increased so that encoding is performed with lower precision.
[0071] In another embodiment, encoding patterns may be considered when determining I-frame processing and QP settings, particularly when prioritizing smooth video playback. For example, if the encoder generates I-frames infrequently, it may be desirable to increase the number of I-frames (lower QP or higher encoder bitrate). That is, within a group of video frames from multiple video frames being compressed, the number of video frames compressed as I-frames is within or below a threshold number of I-frames. In other words, the streaming system can handle the number of I-frames given the current state (e.g., measured client bandwidth, required latency, etc.). Accordingly, the QP value can be adjusted, where QP controls the precision used when compressing video frames. In particular, the QP value can be reduced so that encoding is performed with higher precision.
[0072] If the encoder frequently generates I-frames, it may be desirable to keep the I-frame size small (e.g., by increasing the QP or lowering the encoder bitrate). That is, within a group of video frames from multiple video frames being compressed, the number of video frames compressed as I-frames is either within or exceeding the I-frame threshold. In other words, the video game is generating too many I-frames for the current state (e.g., measured client bandwidth, required latency, etc.). Accordingly, the QP value can be adjusted, where QP controls the precision used when compressing video frames. In particular, the QP value can be increased so that encoding is performed with lower precision.
[0073] In another embodiment, when running the encoder, particularly when prioritizing smooth video playback, encoding patterns may be considered. For example, if the encoder frequently falls below the target frame size, it may be desirable to increase the target frame size. That is, within a group of video frames from multiple video frames compressed and transmitted at the transmission rate, it is determined that the number of video frames falls below a threshold. Each of the video frames falls within the target frame size (i.e., equal to or less than the target frame size). Accordingly, at least one of the target frame size and the maximum frame size is increased.
[0074] On the other hand, if the encoder frequently exceeds the target frame size, it may be desirable to reduce the target frame size. In other words, it is determined whether the number of video frames in a group of video frames from multiple video frames that are compressed and transmitted at the transmission rate meets or exceeds a threshold. Each of the number of video frames exceeds the target frame size. Accordingly, at least one of the target frame size and the maximum frame size is reduced.
[0075] Figure 7A shows the settings of encoder quantization parameters (QP) to optimize quality and buffer utilization at the client, according to one embodiment of the present disclosure. Graph 720A shows the vertical frame size (in bytes) for each generated frame, as shown horizontally. The target frame size and maximum frame size are static. In particular, line 711 shows the maximum frame size and line 712 shows the target frame size, with the maximum frame size being larger than the target frame size. As shown in Graph 720A, there are several peaks containing compressed video frames that exceed the target frame size at line 712. Video frames exceeding the target frame size may require multiple frame durations for encoding and / or transmission from the cloud gaming server, which carries the risk of introducing playback jitter (e.g., increased one-way latency).
[0076] Graph 700B shows the encoder response after QP has been set based on the target frame size, maximum frame size, and QP range (such as minimum and maximum QP) to optimize encoding quality and buffer utilization at the client. For example, QP can be adjusted and / or tuned based on encoder monitoring of the encoder bitrate, the frequency of scene changes, and the frequency of I-frame generation, as described above. Graph 700B shows the vertical frame size (in bytes) for each generated frame, as shown horizontally. The target frame size for line 712 and the maximum frame size for line 711 remain at the same positions as in graph 700A. After QP adjustment and / or adjustment, compared to graph 700A, the number of peaks containing compressed video frames exceeding the target frame size for line 712 decreases. In other words, QP is adjusted to optimize the encoding of video frames for the current conditions (e.g., measured client bandwidth, required latency, etc.) (i.e., to fit within the target frame size).
[0077] Figure 7B shows the adjustment of target frame size, maximum frame size, and / or QP (e.g., minimum QP and / or maximum QP) encoder settings to reduce the occurrence of I-frames exceeding the true target frame size supported by the client, according to one embodiment of the present disclosure. For example, QP may be adjusted and / or adjusted based on encoder monitoring of the encoder bitrate, the frequency of scene changes, and the frequency of I-frame generation, as described above.
[0078] Graph 720A shows the vertical frame size (in bytes) for each generated frame, as shown horizontally. For illustrative purposes, graphs 720A in Figure 7B and 700A in Figure 7A can reflect similar encoder states and are used for encoder tuning. In graph 720A, the target frame size and maximum frame size are static. Specifically, line 711 shows the maximum frame size and line 712 shows the target frame size, where the maximum frame size is larger than the target frame size. As shown in Graph 720A, there are multiple peaks, including compressed video frames that exceed the target frame size at line 712. Video frames exceeding the target frame size may require multiple frame durations for encoding and / or transmission from the cloud gaming server, which carries the risk of introducing playback jitter (e.g., increased one-way latency). For example, the peak reaching the maximum frame size at line 711 could be an I-frame that takes more than 16 milliseconds to be sent to the client, which causes playback jitter by increasing the one-way latency between the cloud gaming server and the client.
[0079] Graph 720B shows the encoder response after at least one of the target frame size and / or maximum frame size has been adjusted to reduce the occurrence of I-frames exceeding the true target frame size supported by the client. The true target frame size can be adjusted based on the measured client bandwidth and / or encoder monitoring, including monitoring of the encoder bitrate, the frequency of scene changes, and the frequency of I-frame generation, as described above.
[0080] Graph 720B shows the vertical frame size (in bytes) for each generated frame, as shown horizontally. Compared to Graph 720A, the target frame size values for line 712' and the maximum frame size values for line 711' are lower. For example, the target frame size for line 712' is smaller than that for line 712, and the maximum frame size for line 711' is smaller than that for line 711. After adjusting the target frame size and / or maximum frame size, the maximum peak size of video frames compressed beyond the target frame size of 712' has been reduced for better transmission. Furthermore, compared to graph 700A, the number of peaks containing compressed video frames exceeding the target frame size of 712' has also decreased. For example, there is only one peak shown in graph 720B. In other words, the target frame size and / or maximum frame size have been adjusted to optimize the encoding of video frames for the current conditions (e.g., measured client bandwidth, required latency, etc.) (i.e., to fit within the target frame size).
[0081] Along with a detailed description of various client devices 210 and / or cloud gaming networks 290 (e.g., within a game server 260) in Figures 2A to 2D, the flowchart 800 in Figure 8 illustrates a cloud gaming method according to one embodiment of the present disclosure, in which video frame encoding includes determining when to skip video frames or when to delay encoding and transmitting video frames if the encoding is long or the resulting video frames are large (e.g., when encoding I-frames). In particular, the decision to skip video frames is made considering network transmission speed and reliability, as well as the overall latency target. This process results in a smoother frame rate and more reliable latency, reducing and making the one-way latency between the cloud gaming server and client more consistent, thereby improving the smoothness of the video display to the client.
[0082] In 810, when a video game is run on a cloud gaming server operating in streaming mode, multiple video frames are generated. Generally, a cloud gaming server generates video frames rendered for multiple games. For example, the generation of video frames rendered for a game is illustrated in 510 of Figure 5 and is applicable to the generation of video frames in Figure 8. For example, the game logic of a video game is built upon a game engine or game title processing engine. The game logic, combined with the game engine, is executed by the CPU and GPU, and together with shared memory, the CPU and GPU may be configured as a rendering pipeline to generate video frames rendered for the game, and the rendering pipeline outputs the rendered images for the game as display-ready video or image frames containing color information corresponding to each pixel of the target and / or virtualized display.
[0083] In 820, scene changes are predicted for a first video frame of a video game, and scene changes are predicted before the first video frame is generated. In one embodiment, game logic can cause the CPU to recognize scene changes while it is running the video game. For example, game logic or add-on logic may include code (e.g., scene change logic) that predicts scene changes when generating video frames, such as predicting that a range of video frames will contain at least one scene change, or predicting that a particular video frame is a scene change. In particular, game logic or add-on logic configured for scene change prediction analyzes game state data collected during video game execution to determine and / or anticipate and / or predict when a scene change will occur, such as within the next X number of frames (e.g., a range) or for identified video frames. For example, a scene change can be predicted as when a character moves from one scene to another in a virtualized game environment, or in a video game, when a character finishes a level and moves to another, or when a transition occurs between two video frames (e.g., a scene cut in a cinematic sequence, or the start of interactive gameplay after a series of menus). Scene changes can be represented by video frames, including large and complex scenes within the virtualized game world or environment.
[0084] Game state data can define the state of the game at a given time and may include game characters, game objects, game object attributes, game attributes, game object states, graphic overlays, the location of the character in the game world during the player's gameplay, the gameplay scene or game environment, the level of the game application, character assets (e.g., weapons, tools, bombs, etc.), loadouts, character skill sets, game level, character attributes, character location, remaining lives, total available lives, armor, trophies, time counter values, and other asset information, etc.
[0085] In 830, a scene change hint is generated and sent to the encoder, indicating that the first video frame is a scene change. In this way, notification of the next scene change can be provided to the encoder, which can adjust its encoding operation when compressing the identified video frame. Notifications provided as scene change hints can be delivered via APIs used for communication between components or between applications running on components of the cloud gaming server 260. In one embodiment, the API may be a GPU API. For example, the API may run on or be called by game logic and / or add-on logic configured to detect scene changes in order to communicate with an encoder. In one embodiment, scene change hints may be provided as data control packets formatted so that all components receiving the data control packets can understand what type of information is contained in the data control packets and understand the appropriate criteria for the corresponding rendered video frames. In one embodiment, the communication protocol used for the API and the format for data control packets may be defined in the corresponding software development kit (SDK) for the video game.
[0086] In step 840, the first video frame is delivered to the encoder. As previously mentioned, the video frame generated by the game may be combined and blended with additional user interface features to become a modified video frame that is scanned by the encoder. The encoder is configured to compress the first video frame based on the desired format, such as the MPEG or H.264 standard used for streaming media content from the cloud gaming server to the client. When streaming, video frames are encoded as P-frames until there is a scene change or until the currently encoded frame can no longer reference a keyframe (e.g., a previous I-frame), and the next video frame is then encoded as another I-frame. In this case, the first video frame is encoded as an I-frame based on a scene change hint, and the I-frame can be encoded without referencing any other video frame (e.g., a standalone key image).
[0087] In step 850, the client's maximum received bandwidth is measured. As previously mentioned, the maximum bandwidth experienced by the client can be determined by means of a feedback mechanism from the client, as shown in operation 530 in Figures 5 and 6. In particular, the streamer of the cloud gaming server can be configured to measure the client's bandwidth.
[0088] In the 860, the encoder receives the second video frame. That is, the second video frame is received after a scene change and compressed after the first video frame has been compressed. The encoder also decides whether to either not encode the second video frame (or subsequent video frame) or to delay the encoding of the second video frame (or subsequent video frame). This decision is based on the client's maximum receiving bandwidth and the target resolution of the client display. In other words, the decision to skip or delay encoding takes into account the bandwidth available to the client. Generally, if the current bandwidth experienced by the client is sufficient, video frames generated and encoded for the client's target display can quickly return to low unidirectional latency after experiencing a latency hit (e.g., generating large I-frames for scene changes), and second video frames (and / or subsequent video frames) can still be encoded with delay. On the other hand, if the current bandwidth experienced by the client is insufficient, the second video frame (and / or subsequent video frames) may be skipped during the encoding process and not delivered to the client. Thus, if the bandwidth to the client exceeds the bandwidth required to support the target resolution of the display on the client, it is possible to have fewer skipped frames (and lower latency).
[0089] In one embodiment, compressed video frames are transmitted from the server to the client at a rate based on the maximum bitrate or bandwidth available over the network at a given point in time. Thus, the transmission rate of encoded slices and / or packets of encoded slices of the compressed video frame is dynamically adjusted according to the currently measured maximum bandwidth. Video frames may be transmitted as they are being encoded, without waiting for the next occurrence of the server VSYNC signal, and without waiting for the entire video frame to be encoded, with transmission occurring as soon as encoding is complete.
[0090] Furthermore, in one embodiment, packet encoding is performed at the same speed as they are transmitted, and both operations are dynamically adjusted to the maximum available bandwidth available to the client. The encoder bitrate can also be monitored, taking into account the next frame and its complexity (e.g., anticipated scene changes) to predict client bandwidth demand, and the encoder bitrate can be adjusted according to the anticipated demand. Additionally, the encoder bitrate can be communicated to the client so that the client can adjust the decoding speed accordingly to match the encoder bitrate.
[0091] In one embodiment, when the transmission rate to the client is low relative to the target resolution of the client display, the second video frame is skipped by the encoder. That is, the second video frame is not encoded. In particular, the transmission rate to the client for a group of compressed video frames exceeds the maximum receiving bandwidth. For example, the transmission rate to the client may be 15 megabytes / second (Mbps), but the client's measured receiving bandwidth may currently be 5-10 Mbps. In this way, if all video frames are continuously pushed to the client, the one-way latency between the cloud gaming server and the client increases. To promote low latency, the second and subsequent video frames can be skipped by the encoder.
[0092] Figure 9A shows a sequence of video frames 900A being compressed by an encoder according to one embodiment of the present disclosure, in which the encoder encodes a first I-frame 905 and then drops encoding a second video frame 920 when the client bandwidth is low for the target resolution of the client's display. The encoded and transmitted blocks of the video frame are shown in relation to the VSYNC signal 950. In particular, when extra bandwidth is unavailable, the time-consuming I-frame encoding is reduced by one or more skipped frames, which can help keep the unidirectional latency low, and this unidirectional latency may include the time it takes for the client to display the video frame. As illustrated, skipping one or more video frames after an I-frame allows for a quick return to low unidirectional latency (e.g., within a one or two-frame duration). Otherwise, by not skipping the encoding of the video frames, it would take several frames to return to low unidirectional latency.
[0093] For example, video frame sequence 900A contains one encoded I-frame 905, with the remaining frames encoded as P-frames. For illustrative purposes, encoded blocks 901 and 902 as P-frames are encoded before encoded block 905, which is encoded as an I-frame. The encoder then compresses the video frames as P-frames until the next scene change or until the video frame can no longer reference previous keyframes (such as I-frames). Generally, encoding time for I-frame blocks can be longer than for P-frame blocks. For example, encoding time for I-frame block 905 can exceed the duration of one frame. In some cases, encoding times between P-frames and I-frames can generally be almost the same, especially when using a high-power encoder.
[0094] However, the transmission times between I-frames and P-frames differ significantly. As illustrated, various transmission times are shown in relation to the corresponding encoded video frames. For example, transmission block 911 of encoded P-frame block 901 is shown with low latency, so that encoding block 901 and transmission block 911 can be executed within one frame duration. Similarly, transmission block 912 of encoded P-frame block 902 is shown with low unidirectional latency, so that encoding block 902 and transmission block 912 can also be executed within one frame duration.
[0095] On the other hand, the transmission block 915A of the encoded I-frame block 905 is shown with higher unidirectional latency, as the encoded block 905 and transmission block 915A occur over several frame durations, thereby introducing jitter into the unidirectional latency between the cloud gaming server and the client. To provide the user with a real-time experience with less unidirectional latency, buffers on the client cannot be used to correct the jitter. In that case, the encoder may decide to skip encoding one or more video frames after the I-frame has been encoded. For example, video frame 920 is dropped by the encoder. In that case, the transmission of the encoded video frame returns to one of the low one-way latencies around the highlighted region 910, as if five subsequent video frames had been encoded as P-frames and sent to the client. That is, the fourth or fifth P-frame encoded after the I-frame block 905 was encoded is also sent to the client within the same frame duration, thereby returning to a low one-way latency between the cloud gaming server and the client.
[0096] In one embodiment, if the transmission rate to the client is high relative to the target resolution of the client display, the second video frame is still compressed by the encoder after the delay (i.e., waits until the I-frame is encoded). In particular, the transmission rate to the client of the compressed group of video frames is within the range of the maximum receiving bandwidth. For example, the transmission rate to the client may be 13 megabytes / second (Mbps), while the client's measured receiving bandwidth may now be 15 Mbps. Thus, there is no delay in the reception of the encoded video frame at the client, and therefore there is no increase in one-way latency between the cloud gaming server and the client.
[0097] Furthermore, since VSYNC signals can be synchronized and offset between the cloud gaming server and the client, one-way latency between the cloud gaming server and the client can be reduced, thereby compensating for latency variability caused by jitter in either the server or the client during transmission over the network. Furthermore, VSYNC signal synchronization and offsetting provide redundant operations on the cloud gaming server (scan out, encode, and transmit), redundant operations on the client (receive, decode, render, and display), and / or redundant operations between the cloud gaming server and client. All of this compensates for latency variability caused by server, network, or client jitter, reduces one-way latency, reduces one-way latency variability, enables real-time generation and display of video content, and facilitates consistent video playback on the client.
[0098] Figure 9B shows a sequence of 900B of video frames being compressed by the encoder, which takes into account the bandwidth available to the client, and as a result, according to one embodiment of the present disclosure, if the bandwidth exceeds the bandwidth required to support the target resolution of the client display, it is possible to have lower latency and have no or fewer skipped frames. In particular, in sequence 900B, video frames are encoded as I-frames, and subsequent video frames are also encoded successfully, after a delay in I-frame encoding, when client bandwidth is moderate relative to the target resolution of the client display. Due to the availability of moderate bandwidth, a moderate amount of excess bandwidth is available to compensate for latency variability (e.g., jitter) between the cloud gaming server and the client, thereby preventing frame skipping and allowing a return to low unidirectional latency to be achieved relatively quickly (e.g., within 2-4 frame periods). The encoding and transmission blocks of the video frames are shown in relation to the VSYNC signal 950.
[0099] The video frame sequence 900B includes one encoded I-frame 905, with the remaining frames encoded as P-frames. For illustrative purposes, encoding blocks 901 and 902 as P-frames are encoded before encoding block 905 is encoded as an I-frame. The encoder then compresses the video frames as P-frames until the next scene change or until the video frame can no longer reference previous keyframes (such as I-frames). In general, the encoding time for an I-frame block may be longer than that for a P-frame block, and the transmission of an I-frame may take longer than one frame duration. For example, the encoding and transmission time for I-frame block 905 exceeds one frame duration. Furthermore, various transmission times are displayed in relation to the corresponding encoded video frame. For example, the encoding and transmission of the video frame preceding I-frame block 905 is shown with low unidirectional latency, resulting in the corresponding encoding and transmission block being able to be completed within one frame duration. However, transmission block 915B of the encoded I-frame block 905 is shown with higher unidirectional latency, resulting in encoding block 905 and transmission block 915B occurring over two or more frame durations, thereby introducing jitter into the unidirectional latency between the cloud gaming server and the client. As mentioned above, encoding time can be further reduced by adjusting one or more encoder parameters (e.g., QP, target frame size, maximum frame size, encoder bitrate, etc.). In other words, the second or subsequent video frames after the I-frame are encoded with lower precision when the transmission rate to the client is moderate relative to the target resolution of the client display, and with lower precision when the transmission rate is higher relative to the target resolution.
[0100] After I-frame block 905, the encoder continues to compress the video frames, but due to I-frame encoding, they may be temporarily delayed. In this case as well, the synchronization and offset of the VSYNC signal provide redundant operations on the cloud gaming server (scan out, encode, and transmit), redundant operations on the client (receive, decode, render, and display), and / or redundant operations between the cloud gaming server and client, all of which compensate for unidirectional latency variability caused by server or network or client jitter, reduce unidirectional latency, reduce unidirectional latency variability, enable real-time generation and display of video content, and facilitate consistent video playback on the client.
[0101] Because the client bandwidth is moderate with respect to the target resolution of the client display, the transmission of the encoded video frames returns to one of the low unidirectional latencies around the highlighted region 940, such as after two or three subsequent video frames have been encoded as P-frames and transmitted to the client. Within region 940, the P-frames encoded after the I-frame block 905 have been encoded are also transmitted to the client within the same frame duration, thereby returning to a low unidirectional latency between the cloud gaming server and the client.
[0102] Figure 9C shows a sequence 900C of video frames being compressed by an encoder, which takes into account the bandwidth available to the client. According to one embodiment of the present disclosure, if the bandwidth exceeds the bandwidth required to support the target resolution of the client display, it is possible to have no or fewer skipped frames while still having lower unidirectional latency. In particular, in sequence 900C, video frames are encoded as I-frames, and subsequent video frames are also successfully encoded, after a delay in encoding the I-frames if the client bandwidth is high relative to the target resolution of the client display. Due to the high bandwidth availability, a large amount of excess bandwidth is available to compensate for unidirectional latency variability (e.g., jitter) between the cloud gaming server and client, thereby preventing frame skipping and allowing for a rapid return to low unidirectional latency (e.g., within one to two frame durations). The encoding and transmission blocks of the video frame are shown in relation to the VSYNC signal 950.
[0103] Similar to Figure 9B, the video frame sequence 900C in Figure 9C contains one encoded I-frame 905, with the remaining frames encoded as P-frames. For illustrative purposes, encoding blocks 901 and 902 as P-frames are encoded before encoding block 905 is encoded as an I-frame. The encoder then compresses the video frames as P-frames until the next scene change or until the video frame can no longer reference previous keyframes (such as I-frames). Generally, encoding time for an I-frame block can be longer than for a P-frame block. For example, the encoding time for I-frame block 905 may exceed the duration of one frame. Furthermore, various transmission times are displayed in relation to the corresponding encoded video frame. For example, the encoding and transmission of the video frame preceding I-frame block 905 is shown with low latency, resulting in the corresponding encoding and transmission block being able to be completed within one frame duration. However, transmission block 915C of the encoded I-frame block 905 is shown with higher latency, resulting in encoding block 905 and transmission block 915C occurring over two or more frame durations, thereby introducing jitter into the one-way latency between the cloud gaming server and the client. As mentioned above, encoding times can be further reduced by adjusting one or more encoder parameters (e.g., QP, target frame size, maximum frame size, encoder bitrate, etc.).
[0104] After I-frame block 905, the encoder continues to compress video frames, but due to I-frame encoding, they may be temporarily delayed. In this case as well, the synchronization and offset of the VSYNC signal provide redundant operations on the cloud gaming server (scan out, encode, and transmit), redundant operations on the client (receive, decode, render, display), and / or redundant operations between the cloud gaming server and client, all of which compensate for latency variability caused by server or network or client jitter, reduce one-way latency, reduce one-way latency variability, enable real-time generation and display of video content, and facilitate consistent video playback on the client. Because the client bandwidth is high relative to the target resolution of the client display, the transmission of the encoded video frames returns to one of the low unidirectional latencies around the highlighted region 970, such as after one or two subsequent video frames are encoded as P-frames and transmitted to the client. Within region 970, the P-frames encoded after the I-frame block 905 is encoded are also transmitted to the client within a single frame duration (although these span both sides of the VSYNC signal generation), thereby returning to a low unidirectional latency between the cloud gaming server and the client.
[0105] Figure 10 shows components of an exemplary device 1000 that can be used to perform various embodiments of the present disclosure. For example, Figure 10 shows an exemplary hardware system suitable for streaming media content and / or receiving streamed media content, which includes providing encoder tuning to improve the trade-off between one-way latency and video quality in a cloud gaming system for the purpose of reducing latency and providing more consistent latency between clouds and to improve the smoothness of the video client display, wherein the encoder tuning is obtained based on monitoring client bandwidth, skipped frames, the number of encoded I-frames, the number of scene changes, and / or the number of video frames exceeding the target frame size, and the tuned parameters may include encoder bitrate, target frame size, maximum frame size, and quantization parameter (QP) values, and according to embodiments of the present disclosure, high-performance encoders and decoders work to reduce overall one-way latency between the cloud gaming server and the client. This block diagram shows device 1000, which may or may not be a personal computer, server computer, game console, mobile device, or other digital device, each of which is suitable for carrying out embodiments of the present invention. Device 1000 includes a central processing unit (CPU) 1002 for running software applications and, optionally, an operating system. The CPU 1002 may consist of one or more homogeneous or heterogeneous processing cores.
[0106] According to various embodiments, the CPU 1002 is one or more general-purpose microprocessors having one or more processing cores. Further embodiments can be implemented using one or more CPUs with a microprocessor architecture, which is specifically adapted for highly parallel and computationally intensive applications such as media and interactive entertainment applications, applications configured for graphics processing during game execution.
[0107] Memory 1004 stores applications and data for use by the CPU 1002 and GPU 1016. Storage 1006 provides non-volatile storage and other computer-readable media for applications and data, and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROMs, DVD-ROMs, Blu-ray®, HD-DVDs, or other optical storage devices, as well as signal transmission and storage media. User input device 1008 communicates user input from one or more users to device 1000, and examples of device 1000 may include a keyboard, mouse, joystick, touchpad, touchscreen, still image or video recorder / camera, and / or microphone. The network interface 1009 enables device 1000 to communicate with other computer systems via an electronic communication network, which may include wired or wireless communication via a local area network and a wide area network such as the Internet. The audio processor 1012 is adapted to generate analog or digital audio output from instructions and / or data provided by the CPU 1002, memory 1004, and / or storage 1006. The components of device 1000, including the CPU 1002, the graphics subsystem including the GPU 1016, memory 1004, data storage 1006, user input device 1008, network interface 1009, and audio processor 1012, are connected via one or more data buses 1022.
[0108] The graphics subsystem 1014 is further connected to the data bus 1022 and components of device 1000. The graphics subsystem 1014 includes a graphics processing unit (GPU) 1016 and graphics memory 1018. The graphics memory 1018 includes display memory (e.g., a frame buffer) used to store pixel data for each pixel of the output image. The graphics memory 1018 can be integrated into the same device as the GPU 1016, connected to the GPU 1016 as a separate device, and / or implemented within memory 1004. Pixel data can be provided directly from the CPU 1002 to the graphics memory 1018. Alternatively, the CPU 1002 provides the GPU 1016 with data and / or instructions defining a desired output image, from which the GPU 1016 generates pixel data for one or more output images. The data and / or instructions defining a desired output image can be stored in memory 1004 and / or graphics memory 1018. In one embodiment, the GPU 1016 includes a 3D rendering function for generating pixel data for an output image from instructions and data defining geometry, lighting, shading, texturing, motion, and / or camera parameters for a scene. The GPU 1016 may further include one or more programmable execution units capable of executing shader programs.
[0109] The graphics subsystem 1014 periodically outputs pixel data from the graphics memory 1018 of an image that is displayed on the display device 1010 or projected by a projection system (not shown). The display device 1010 may be any device capable of displaying visual information in response to signals from device 1000, including CRTs, LCDs, plasma displays, and OLED displays. Device 1000 may, for example, provide analog or digital signals to the display device 1010.
[0110] Other embodiments for optimizing the graphics subsystem 1014 may include multi-tenancy GPU operations in which GPU instances are shared among multiple applications and GPUs supporting a single game are distributed. The graphics subsystem 1014 can be configured as one or more processing devices.
[0111] For example, the graphics subsystem 1014 may be configured to perform multi-tenancy GPU functions, and in one embodiment, one graphics subsystem may implement graphics and / or rendering pipelines for multiple games. In other words, the graphics subsystem 1014 is shared among multiple games that are running.
[0112] In other embodiments, the graphics subsystem 1014 includes multiple GPU devices, which are combined to perform graphics processing for a single application running on the corresponding CPU. For example, multiple GPUs can perform an alternative form of frame rendering, where GPU1 renders the first frame, GPU2 renders the second frame in a consecutive frame period, and so on until the last GPU is reached, resulting in the first GPU rendering the next video frame (for example, if there are only two GPUs, GPU1 renders the third frame). In other words, the GPUs cycle when rendering frames. Rendering operations can overlap, and GPU2 may start rendering the second frame before GPU1 finishes rendering the first frame. In another embodiment, different shader operations can be assigned to multiple GPU devices in the rendering pipeline and / or graphics pipeline. The master GPU performs the main rendering and compositing. For example, in a group of three GPUs, master GPU1 can perform the main rendering (e.g., the first shader operation) and compositing the outputs from slave GPU2 and slave GPU3, slave GPU2 can perform the second shader operation (e.g., fluid effects such as rivers), slave GPU3 can perform the third shader operation (e.g., particle smoke), and master GPU1 composites the results from each of GPU1, GPU2, and GPU3. In this way, various GPUs can be assigned to perform various shader operations (e.g., flag waving, wind, smoke generation, fire) to render video frames. In yet another embodiment, each of the three GPUs can be assigned to a different object and / or portion of the scene corresponding to a video frame. In the embodiments and implementations described above, these operations can be performed in the same frame duration (concurrently and simultaneously) or in different frame durations (concurrently and sequentially).
[0113] Accordingly, this disclosure describes a method and system configured for streaming and / or receiving media content, including providing encoder tuning to improve the trade-off between one-way latency and video quality in a cloud gaming system, wherein the encoder tuning is obtained based on monitoring client bandwidth, skipped frames, number of encoded I-frames, number of scene changes, and / or number of video frames exceeding the target frame size, and the tuned parameters may include encoder bitrate, target frame size, maximum frame size, and quantization parameter (QP) values, and high-performance encoders and decoders work to reduce overall one-way latency between the cloud gaming server and the client.
[0114] It should be understood that the various embodiments defined herein may be combined or assembled into a particular embodiment using the various features disclosed herein. Therefore, the provided embodiments are only a few possible embodiments and are not limited to the various embodiments that can be defined by combining different elements. In some examples, a particular embodiment may include fewer elements without departing from the spirit of the disclosed or equivalent embodiments.
[0115] Embodiments of the Disclosure can be implemented in a variety of computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, and mainframe computers. Embodiments of the Disclosure can also be implemented in a distributed computing environment in which tasks are performed by remote processing devices linked via a wired or wireless network.
[0116] With the embodiments described above in mind, it should be understood that embodiments of the present disclosure can employ a variety of computer implementation operations, including data stored in a computer system. These operations require the physical manipulation of physical quantities. Any of the operations described herein that form part of embodiments of the present disclosure are useful mechanical operations. Embodiments of the disclosure also relate to devices or apparatus for performing these operations. The apparatus may be built specifically for a required purpose, or the apparatus may be a general-purpose computer selectively started or configured by a computer program stored in the computer. Specifically, various general-purpose machines can be used with computer programs written in accordance with the teachings of this specification, or it may be more convenient to build an apparatus more specialized to perform the required operations.
[0117] This disclosure can also be embodied as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that can store data which can then be read by a computer system. Examples of computer-readable mediums include hard drives, network-attached storage (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tapes, and other optical and non-optical data storage devices. A computer-readable medium may include computer-readable tangible media distributed on a network-connected computer system, in which computer-readable code is stored and executed in a distributed manner.
[0118] Although the operations of this method are described in a specific order, other housekeeping operations may be performed between operations, or the operations may be coordinated to occur at slightly different times, or they may be distributed throughout the system. It should be understood that this allows for the occurrence of processing operations at various intervals related to the processing, as long as the processing of overlapping operations is performed in the desired manner.
[0119] While the foregoing disclosure has been described in some detail for the purpose of clarifying understanding, it will be apparent that certain changes and modifications can be implemented within the scope of the appended claims. Accordingly, these embodiments should be considered illustrative rather than restrictive, and the embodiments of this disclosure are not limited to the details provided herein and may be modified within the scope of the appended claims and equivalents.
Claims
1. Cloud gaming method, When running a video game on a cloud gaming server, multiple video frames are generated. The plurality of video frames are encoded at the encoder bitrate, and the compressed plurality of video frames are transmitted from the streamer of the cloud gaming server to the client. Measure the client's maximum receive bandwidth, In the streamer, the encoding of the plurality of video frames is monitored, A method for dynamically adjusting the parameters of an encoder based on the monitoring of the encoding.
2. In the dynamic adjustment of the aforementioned parameters, It is determined that the encoder bitrate used to encode the group of video frames from the plurality of video frames exceeds the maximum receiving bandwidth. The method according to claim 1, wherein the value of the QP parameter is increased so that the encoding is performed with lower precision, the parameter being the QP parameter.
3. In the dynamic adjustment of the aforementioned parameters, Determine that the encoder bitrate used to encode the group of video frames from the plurality of video frames is within the maximum receiving bandwidth. When transmitting the group of video frames, it is determined that there is excess bandwidth, Reducing the value of the QP parameter based on the excess bandwidth so that encoding can be performed with higher accuracy, wherein the parameter is the QP parameter, and the reduction The method according to claim 1, including the method described in claim 1.
4. In the dynamic adjustment of the aforementioned parameters, Determine whether the number of video frames compressed as I-frames from the group of video frames from the compressed plurality of video frames satisfies or exceeds the threshold number of I-frames. Increasing the value of the QP parameter so that the encoding is performed with lower precision, wherein the parameter is the QP parameter, and the increase The method according to claim 1, including the method described in claim 1.
5. In the dynamic adjustment of the aforementioned parameters, It is determined that the number of video frames compressed as I-frames from the group of video frames from the compressed plurality of video frames is lower than the threshold number of I-frames. The method according to claim 1, wherein the value of the QP parameter is reduced so that encoding is performed with higher accuracy, and the parameter is the QP parameter.
6. In the dynamic adjustment of the aforementioned parameters, A group of video frames from the plurality of video frames encoded and transmitted at the transmission rate includes the number of video frames, and it is determined whether the number of video frames satisfies or exceeds a threshold, and each of the number of video frames exceeds the target frame size. The method according to claim 1, wherein at least one of the target frame size and the maximum frame size is reduced for the parameter.
7. The method according to claim 6, wherein the target frame size and the maximum frame size are equal.
8. In the dynamic adjustment of the aforementioned parameters, A group of video frames from the plurality of video frames encoded and transmitted at the transmission rate includes the number of video frames, and it is determined that the number of video frames is below a threshold, and each of the number of video frames is within the target frame size. The method according to claim 1, wherein at least one of the target frame size and the maximum frame size is increased as the parameter.
9. In the dynamic adjustment of the aforementioned parameters, Determine whether the number of video frames identified as having a scene change from the group of video frames from the compressed plurality of video frames satisfies or exceeds the scene change threshold number. The method according to claim 1, wherein the value of the QP parameter is increased so that the encoding is performed with lower precision, the parameter being the QP parameter.
10. In the dynamic adjustment of the aforementioned parameters, It is determined that the number of video frames identified as having a scene change from the group of video frames from the compressed plurality of video frames is lower than the threshold number for scene changes. The method according to claim 1, wherein the value of the QP parameter is reduced so that encoding is performed with higher accuracy, and the parameter is the QP parameter.
11. The method according to claim 1, further comprising prioritizing smooth playback on the client by disabling the skipping of video frame encoding.
12. The method according to claim 1, further comprising dynamically adjusting the encoder bitrate speed based on the client's maximum receiving bandwidth in the encoder.
13. A non-temporary computer-readable medium for storing cloud gaming computer programs, When running a video game on a cloud gaming server, it has program instructions for generating multiple video frames. It has a program instruction for measuring the maximum receive bandwidth of a client, The program has instructions for encoding the plurality of video frames at the encoder bitrate, and the plurality of video frames to be compressed are transmitted from the streamer of the cloud gaming server to the client. The streamer has program instructions for monitoring the encoding of the plurality of video frames, A non-temporary computer-readable medium having program instructions for dynamically adjusting the parameters of the encoder based on the monitoring of the encoding.
14. The program instructions for dynamically adjusting the aforementioned parameters are: The program has instructions for determining whether the encoder bitrate used to encode a group of video frames from the plurality of video frames exceeds the maximum receive bandwidth. A non-temporary computer-readable medium according to claim 13, having a program instruction for increasing the value of a QP parameter so that encoding is performed with lower precision, wherein the parameter is the QP parameter.
15. The program instructions for dynamically adjusting the aforementioned parameters are: The program has instructions for determining that the encoder bitrate used to encode a group of video frames from the plurality of video frames is within the maximum receive bandwidth. The program has instructions for determining whether there is excess bandwidth when transmitting the group of video frames, A non-temporary computer-readable medium according to claim 13, having a program instruction for reducing the value of a QP parameter based on the excess bandwidth so that encoding is performed with higher accuracy, wherein the parameter is the QP parameter.
16. The program instructions for dynamically adjusting the aforementioned parameters are: The program has instructions for determining whether the number of video frames identified as having a scene change from a group of video frames from the compressed plurality of video frames satisfies or exceeds a scene change threshold number. A non-temporary computer-readable medium according to claim 13, having a program instruction for increasing the value of a QP parameter so that encoding is performed with lower precision, wherein the parameter is the QP parameter.
17. The program instructions for dynamically adjusting the aforementioned parameters are: The program has instructions for determining whether the number of video frames compressed as I-frames from a group of video frames from the compressed plurality of video frames is lower than a threshold number of I-frames. The program has a program instruction for reducing the value of a QP parameter so that the encoding of the programming instruction is executed with higher precision, wherein the parameter is the QP parameter. , the non-temporary computer-readable medium according to claim 13.
18. The program instructions for dynamically adjusting the aforementioned parameters are: A group of video frames from the plurality of video frames encoded and transmitted at the transmission rate includes a number of video frames, and has a program instruction for determining whether the number of video frames satisfies or exceeds a threshold, wherein each of the number of video frames exceeds a target frame size. The non-temporary computer-readable medium according to claim 13, having a program instruction for reducing at least one of the parameters of the target frame size and the maximum frame size.
19. The non-temporary computer-readable medium according to claim 18, wherein the target frame size and the maximum frame size are equal within the computer program for cloud gaming.
20. The program instructions for dynamically adjusting the aforementioned parameters are: A program instruction for determining whether a group of video frames from a plurality of video frames encoded and transmitted at the transmission rate includes the number of video frames, and the number of video frames is below a threshold, wherein each of the number of video frames is within a target frame size. The non-temporary computer-readable medium according to claim 13, having a program instruction for increasing at least one of the target frame size and the maximum frame size as the parameter.
21. The program instructions for dynamically adjusting the aforementioned parameters are: The program has instructions for determining whether the number of video frames identified as having a scene change from a group of video frames from the compressed plurality of video frames satisfies or exceeds a scene change threshold number. A non-temporary computer-readable medium according to claim 13, having a program instruction for increasing the value of a QP parameter so that encoding is performed with lower precision, wherein the parameter is the QP parameter.
22. The program instructions for dynamically adjusting the aforementioned parameters are: The program has instructions for determining whether the number of video frames identified as having a scene change from a group of video frames from the compressed plurality of video frames is lower than a threshold number for scene changes. A non-temporary computer-readable medium according to claim 13, having a program instruction for reducing the value of a QP parameter so that encoding is performed with higher precision, wherein the parameter is the QP parameter.
23. The non-temporary computer-readable medium according to claim 13, further comprising program instructions for prioritizing smooth playback on the client by disabling the skipping of video frame encoding.
24. The non-temporary computer-readable medium according to claim 13, further comprising program instructions for dynamically adjusting the encoder bitrate speed in the encoder based on the client's maximum receiving bandwidth.
25. A computer system, Processor and The system includes a memory coupled to the processor and having instructions stored therein, wherein, when executed by the computer system, the instructions cause the computer system to execute a cloud gaming method, and the cloud gaming method is When running a video game on a cloud gaming server, multiple video frames are generated. The plurality of video frames are encoded at the encoder bitrate, and the compressed plurality of video frames are transmitted from the streamer of the cloud gaming server to the client. Measure the client's maximum receive bandwidth, In the streamer, the encoding of the plurality of video frames is monitored, A computer system that dynamically adjusts the parameters of the encoder based on the monitoring of the encoding process.
26. In the method for dynamically adjusting the aforementioned parameters, It is determined that the encoder bitrate used to encode a group of video frames from the plurality of video frames exceeds the maximum receiving bandwidth. The computer system according to claim 25, wherein the value of the QP parameter is increased so that encoding is performed with lower precision, the parameter being the QP parameter.
27. In the method for dynamically adjusting the aforementioned parameters, Determine that the encoder bitrate used to encode the group of video frames from the plurality of video frames is within the maximum receiving bandwidth. When transmitting the group of video frames, it is determined that there is excess bandwidth, The computer system according to claim 25, wherein the value of the QP parameter is reduced based on the excess bandwidth so that encoding is performed with higher accuracy, the parameter being the QP parameter.
28. In the method for dynamically adjusting the aforementioned parameters, Determine whether the number of video frames compressed as I-frames from the group of video frames from the compressed plurality of video frames satisfies or exceeds the threshold number of I-frames. The computer system according to claim 25, wherein the value of the QP parameter is increased so that encoding is performed with lower precision, the parameter being the QP parameter.
29. In the method for dynamically adjusting the aforementioned parameters, It is determined that the number of video frames compressed as I-frames from the group of video frames from the compressed plurality of video frames is lower than the threshold number of I-frames. The computer system according to claim 25, wherein the value of a QP parameter is reduced so that encoding is performed with higher accuracy, the parameter being the QP parameter.
30. In the method for dynamically adjusting the aforementioned parameters, A group of video frames from the plurality of video frames that have been encoded and transmitted at the transmission rate includes the number of video frames, and it is determined whether the number of video frames satisfies or exceeds a threshold, and each of the number of video frames exceeds the target frame size. The computer system according to claim 25, wherein at least one of the target frame size and the maximum frame size is reduced for the parameter.
31. The computer system according to claim 30, wherein the target frame size and the maximum frame size are equal in the method described above.
32. In the method for dynamically adjusting the aforementioned parameters, A group of video frames from the plurality of video frames that have been encoded and transmitted at the transmission rate includes the number of video frames, and it is determined that the number of video frames is below a threshold, and each of the number of video frames is within the target frame size. The computer system according to claim 25, wherein at least one of the target frame size and the maximum frame size is increased as the parameter.
33. In the method for dynamically adjusting the aforementioned parameters, Determine whether the number of video frames identified as having a scene change from the group of video frames from the compressed plurality of video frames satisfies or exceeds the scene change threshold number. The computer system according to claim 25, wherein the value of the QP parameter is increased so that encoding is performed with lower precision, the parameter being the QP parameter.
34. In the method for dynamically adjusting the aforementioned parameters, It is determined that the number of video frames identified as having a scene change from the group of video frames from the compressed plurality of video frames is lower than the threshold number for scene changes. The computer system according to claim 25, wherein the value of the QP parameter is reduced so that encoding is performed with higher accuracy, and the parameter is the QP parameter.
35. The above method further, The computer system according to claim 25, which prioritizes smooth playback on the client by disabling the skipping of video frame encoding.
36. The above method further, The computer system according to claim 25, wherein the encoder dynamically adjusts the encoder bitrate based on the client's maximum receiving bandwidth.
37. Cloud gaming method, When running a video game on a cloud gaming server, multiple video frames are generated. A scene change is predicted for a first video frame for the video game, wherein the scene change is predicted before the first video frame is generated. A scene change hint is generated indicating that the first video frame represents a scene change. The scene change hint is sent to the encoder. The first video frame is distributed to an encoder, and the first video frame is encoded as an I-frame based on the scene change hint. Measure the client's maximum receive bandwidth, A method for determining whether or not to encode a second video frame received by the encoder, based on the client's maximum receiving bandwidth and the client display's target resolution.
38. Furthermore, the encoder dynamically adjusts the encoder bitrate based on the client's maximum receiving bandwidth. The method according to claim 37, wherein, once the video frame has been encoded, the video frame is transmitted to the client.
39. In determining whether to encode or not encode the second video frame, The method according to claim 37, wherein when the transmission rate to the client is low relative to the target resolution of the client display, encoding of the second video frame is skipped so that the transmission rate to the client of the group of video frames from the compressed plurality of video frames exceeds the maximum receiving bandwidth.
40. In determining whether to encode or not encode the second video frame, The method according to claim 37, wherein, when the transmission speed to the client is high relative to the target resolution of the client display, the second video frame is successfully encoded such that the transmission speed to the client is within the maximum receiving bandwidth for a group of video frames from the compressed plurality of video frames.
41. In determining whether to encode or not encode the second video frame, The method according to claim 40, wherein the transmission speed to the client is moderate with respect to the target resolution of the client display, the second video frame is encoded with lower precision.
42. In predicting scene changes for the first video frame, In order to generate the aforementioned multiple video frames, the cloud gaming server executes the game logic built on the video game's game engine. Scene change logic is executed to predict the scene change for the first video frame, and the prediction is based on the game state collected during the execution of the game logic. Using the aforementioned scene change logic, generate the scene change hint. The method according to claim 37, wherein the encoder transmits the scene change hint before receiving the first video frame.
43. The method according to claim 42, wherein the scene change hint is delivered from the scene change logic to the encoder via an API.
44. The method according to claim 37, wherein the second video frame is compressed after the first video frame has been compressed by the encoder.
45. A non-temporary computer-readable medium for storing cloud gaming computer programs, When running a video game on a cloud gaming server, it has program instructions for generating multiple video frames. The program has a program instruction for predicting a scene change for a first video frame for the video game, wherein the scene change is predicted before the first video frame is generated. The program has a program instruction for generating a scene change hint, where the first video frame is a scene change. The program has instructions for sending the scene change hint to the encoder, The program has a program instruction for delivering the first video frame to an encoder, wherein the first video frame is encoded as an I-frame based on the scene change hint. It has a program instruction for measuring the maximum receive bandwidth of a client, A non-temporary computer-readable medium having program instructions for determining whether or not to encode a second video frame received by the encoder, based on the client's maximum receiving bandwidth and the target resolution of the client display.
46. Furthermore, the system has program instructions for dynamically adjusting the encoder bitrate speed in the encoder based on the client's maximum receiving bandwidth. The non-temporary computer-readable medium according to claim 45, having program instructions for transmitting the video frame to the client once the video frame has been encoded.
47. The program instruction for determining whether or not to encode the second video frame is: The non-temporary computer-readable medium according to claim 45, having a program instruction for skipping encoding of the second video frame such that, when the transmission rate to the client is low relative to the target resolution of the client display, the transmission rate to the client of a group of video frames from the compressed plurality of video frames exceeds the maximum receiving bandwidth.
48. The program instruction for determining whether or not to encode the second video frame is: The non-temporary computer-readable medium according to claim 45, comprising program instructions for successfully encoding the second video frame such that the transmission speed to the client is within the maximum receiving bandwidth for a group of video frames from the compressed plurality of video frames, when the transmission speed to the client is high relative to the target resolution of the client display.
49. The program instruction for determining whether or not to encode the second video frame is: The non-temporary computer-readable medium according to claim 48, having program instructions for encoding the second video frame with lower precision when the transmission speed to the client is moderate with respect to the target resolution of the client display.
50. The program instruction for predicting scene changes for the first video frame is: To generate the aforementioned plurality of video frames, the cloud gaming server has program instructions for executing game logic built on the game engine of the video game, The program has a program instruction that executes scene change logic to predict the scene change for the first video frame, wherein the prediction is based on the game state collected during the execution of the game logic. The program has instructions for generating the scene change hint using the scene change logic, The non-temporary computer-readable medium according to claim 45, having a program instruction for transmitting the scene change hint before the encoder receives the first video frame.
51. The non-temporary computer-readable medium according to claim 50, wherein the scene change hint is delivered from the scene change logic to the encoder via an API in the computer program for cloud gaming.
52. The non-temporary computer-readable medium according to claim 45, wherein in the computer program for cloud gaming, the second video frame is compressed after the first video frame has been compressed by the encoder.
53. A computer system, Processor and The computer system has a memory coupled to the processor and having instructions stored therein, and when the instructions are executed by the computer system, the computer system causes the computer system to execute a cloud gaming method, and the cloud gaming method is When running a video game on a cloud gaming server, multiple video frames are generated. A scene change is predicted for a first video frame for the video game, wherein the scene change is predicted before the first video frame is generated. The first video frame is a scene change, generate a scene change hint, The scene change hint is sent to the encoder. The first video frame is distributed to an encoder, and the first video frame is encoded as an I-frame based on the scene change hint. Measure the client's maximum receive bandwidth, A computer system that determines whether or not to encode a second video frame received by the encoder, based on the client's maximum receiving bandwidth and the target resolution of the client display.
54. Furthermore, in the encoder, the encoder bitrate speed is dynamically adjusted based on the client's maximum receiving bandwidth. The computer system according to claim 53, wherein when the video frame is encoded, the video frame is transmitted to the client.
55. In determining whether to encode or not encode the second video frame, The computer system according to claim 53, wherein when the transmission rate to the client is low relative to the target resolution of the client display, the encoding of the second video frame is skipped so that the transmission rate to the client of the group of video frames from the compressed plurality of video frames exceeds the maximum receiving bandwidth.
56. In determining whether to encode or not encode the second video frame, The computer system according to claim 53, which, when the transmission rate to the client is high relative to the target resolution of the client display, successfully encodes the second video frame such that the transmission rate to the client is within the maximum receiving bandwidth for a group of video frames from the compressed plurality of video frames.
57. In determining whether to encode or not encode the second video frame, The computer system according to claim 56, wherein the transmission speed to the client is moderate with respect to the target resolution of the client display, the second video frame is encoded with lower precision.
58. In predicting scene changes for the first video frame, In order to generate the aforementioned multiple video frames, the cloud gaming server executes the game logic built on the video game's game engine. Executing scene change logic to predict the scene change for the first video frame, wherein the prediction is based on the game state collected during the execution of the game logic, Using the aforementioned scene change logic, generate the scene change hint. The computer system according to claim 53, wherein the encoder transmits the scene change hint before receiving the first video frame.
59. The computer system according to claim 58, wherein the scene change hint is delivered from the scene change logic to the encoder via an API.
60. The computer system according to claim 53, wherein the second video frame is compressed after the first video frame has been compressed by the encoder.