Cloud Gaming GPU with Integrated NIC and Shared Frame Buffer Access for Low Latency
The integration of a GPU with an integrated NIC and frame buffers in cloud gaming systems addresses latency issues by enabling direct encoding and streaming of video content, improving the efficiency and reducing latency in cloud gaming systems.
Patent Information
- Application Number
- JP2022019573
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-15
- Filing Date
- 2022-02-10
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2042-02-10
AI Technical Summary
Existing cloud gaming systems experience latency issues due to the need to copy and transfer data rendered and encoded by the GPU to the main PC or device memory and then to the network card, creating a bottleneck.
Integration of a graphics processing unit (GPU) with an integrated network interface controller (NIC) and frame buffers, allowing direct encoding and streaming of video content without CPU involvement, using a video codec and packetizing the encoded content for transmission.
Reduces latency by eliminating the need for CPU-mediated data transfer, enhancing the efficiency of cloud gaming by directly streaming encoded video content over a network.
Smart Images

Figure 0007800991000001 
Figure 0007800991000002 
Figure 0007800991000003
Abstract
Description
[Technical Field]
[0001] Cloud gaming is a type of online gaming in which video games are run on remote servers in a data center (also known as the "cloud") and streamed as video content to player devices via local client software. The local client software is used to render the video content and provide player input to the remote server. This contrasts with existing gaming solutions in which games are run locally on a user's video game console, personal computer, or mobile device. [Background technology]
[0002] Latency is one of the most important criteria for the success of cloud gaming and interactive in-home streaming (e.g., using a PC (personal computer) for rendering but playing on a tablet in another room). One approach used today is to render video content and prepare encoded image data to be streamed using a separate graphics card with a graphics processing unit (GPU), and then stream the image data over a network to the player's device using the platform's central processing unit (CPU) and network card. However, data rendered and / or encoded by the GPU must first be copied to the main PC or device memory and then transferred to the network card for the image data to be sent out, resulting in a bottleneck. Summary of the Invention
[0003] The present invention provides an apparatus comprising: The graphics processing unit (GPU) is Multiple frame buffers; an integrated encoder / decoder coupled to the at least one frame buffer and including embedded logic for encoding and decoding at least one of image data and video content; and an integrated network interface controller (NIC) coupled to the integrated encoder / decoder.
[0004] The present invention also provides a method implemented in a graphics card including a graphics processing unit (GPU) with an integrated network interface controller (NIC), comprising: generating video frame content and buffering the video frame content in one or more frame buffers on the GPU; encoding the video frame content using a video codec integrated in the GPU to generate encoded video content; packetizing the encoded video content using the integrated NIC to generate a stream of packets and transmitting the stream of packets outbound to a network operatively coupled to an output port on the integrated NIC; Or, encoding tiles of video frame content using an image tile encoder integrated with the GPU to generate encoded video tiles; packetizing the encoded video tiles using the integrated NIC to generate a stream of packets and transmitting the stream of packets outbound to a network operatively coupled to an output port on the integrated NIC; The present invention provides a method comprising:
[0005] The present invention further provides a cloud game server, a first board having a plurality of expansion slots or connectors; a central processing unit (CPU) mounted on a first board or a second board, the second board being installed in an expansion slot or coupled to a mating connector on the first board; a main memory comprising one or more memory devices communicatively coupled to the CPU; a) one or more network adapter cards installed in each expansion slot, or b) one or more network interface controller (NIC) chips mounted on the first board or the second board; a plurality of graphics cards installed in respective expansion slots or coupled to respective mating connectors on the first board; Each graphics card is a graphics processing unit (GPU), one or more frame buffers; an integrated encoder / decoder coupled to the at least one frame buffer and including embedded logic for encoding and decoding at least one of image data and video content; an integrated network interface controller (NIC) coupled to the integrated encoder / decoder; a graphics memory coupled to or integrated with the GPU; At least one network port coupled to an integrated NIC; Input / output (I / O) interfaces connected to the GPU It provides a cloud gaming server equipped with the above. [Brief explanation of the drawings]
[0006] The foregoing aspects and many of the attendant advantages of this invention will become more readily understood as the same becomes better understood by reference to the following detailed description, taken in conjunction with the accompanying drawings, in which, unless otherwise specified, like reference numerals refer to like parts throughout the various views.
[0007] [Figure 1] 1 is a schematic diagram of a graphics card including a GPU with an integrated video codec and an integrated NIC according to one embodiment. [Figure 1a] 1 is a schematic diagram of a graphics card including a GPU with an integrated video codec directly coupled to a NIC according to one embodiment. [Figure 1b] 1 is a schematic diagram of a graphics card including a GPU with an integrated video codec and a NIC combined on a multi-chip module according to one embodiment. [Figure 1c] 1 is a schematic diagram of a graphics card including a GPU with an integrated tile encoder / decoder and an integrated NIC according to one embodiment. [Figure 2] 2 is a schematic diagram illustrating the use of the graphics card of FIG. 1 in a game server and a game client device, according to one embodiment. [Figure 2a] FIG. 1B is a schematic diagram illustrating the use of the graphics card of FIG. 1A in a game server and a game client device, according to one embodiment. [Figure 3] 2 is a schematic diagram illustrating the use of the graphics card of FIG. 1 in a gaming server according to one embodiment, illustrating a gaming laptop client including a GPU with an integrated network interface for communicating directly with a WiFi chip. [Figure 4] FIG. 1 is a schematic diagram illustrating an exemplary frame encoding and display scheme consisting of I-frames, P-frames, and B-frames. [Figure 5] 2 is a schematic diagram illustrating an example of end-to-end image data flow between a game server 200 and a desktop game client 202 according to one embodiment. [Figure 6] 6 is a flowchart illustrating operations performed to facilitate the end-to-end image data flow scheme of FIG. 5 according to one embodiment. [Figure 7a] FIG. 1 illustrates the generation, encoding, and streaming of game tiles using a GPU with an integrated tile encoder, according to one embodiment. [Figure 7b] FIG. 1 illustrates handling of a stream of game tiles received at a game client, including tile decoding and playback using a GPU with an integrated tile decoder, according to one embodiment. [Figure 8] 1 is a schematic diagram of a game server including multiple graphics cards and one or more network cards installed in expansion slots on a mainboard, according to one embodiment. [Figure 8a] FIG. 1 is a schematic diagram of a game server including multiple graphics cards installed in expansion slots on a mainboard with a NIC chip, according to one embodiment. [Figure 8b] 1 is a schematic diagram of a gaming server including multiple graphics cards and blade servers installed in slots or mating connectors in a backplane, midplane, or baseplane, according to one embodiment. [Figure 9] FIG. 1 is a schematic diagram illustrating an integrated NIC according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0008] Embodiments of a method and apparatus for a cloud gaming GPU with an integrated network interface controller (NIC) and shared frame buffer access for lower latency are described herein. In the following description, numerous specific and specific details are set forth to provide a thorough understanding of embodiments in accordance with the present invention. However, those skilled in the art will recognize that the present invention may be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the present invention.
[0009] References throughout this specification to "one embodiment" or "an embodiment" mean that at least one embodiment of the invention includes the particular feature, structure, or characteristic described in connection with that embodiment. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0010] For clarity, individual components in the figures herein may be referred to by their labeling in the figures rather than by a specific reference number. Furthermore, reference numbers referring to a particular type of component (as opposed to a particular component) may be followed by "(typ)," meaning "typical." It is understood that these component configurations are typical among similar components that may be present but are not shown in the figures for brevity and clarity, or among other similar components that are not labeled with a separate reference number. Conversely, "(typ)" should not be interpreted to mean that the component, element, etc. is typically employed for its disclosed function, apparatus, purpose, etc.
[0011] According to aspects of the embodiments disclosed herein, a GPU is provided that has an integrated encoder and an integrated NIC. The GPU includes an integrated encoder / decoder and one or more frame buffers that provide shared access to other GPU components. The GPU is configured to process outbound and inbound game image content that is encoded and decoded using a video codec or a game tile encoder and decoder. For example, when implemented in a cloud gaming server for a local video game host, video game frames generated by the GPU and buffered in the frame buffer are encoded by the integrated encoder and transferred directly to the NIC to be packetized and streamed using a media streaming protocol. The inbound streamed media content is depacketized by the NIC and decoded by the integrated decoder, which writes the decoded content to the frame buffer and plays the video game frames on the gaming client device. The video game frames are then displayed on the client device display.
[0012] Generally, a GPU may be implemented on a graphics card or on the main board of a gaming client device, such as a laptop or notebook computer or mobile device. Graphics cards reduce latency when generating outbound game content and when processing inbound game content because the processing path does not involve transferring encoded data to or from a CPU.
[0013] 1 shows a graphics card 100 including a GPU 102 having a frame buffer 104. The frame buffer 104 is accessed by an H.264 / H.265 video codec 106 via an interface 108. The H.264 / H.265 video codec 106 includes an I / O interface 110 coupled to a NIC 112, which in this embodiment is onboard the GPU 102. The GPU 102 is coupled to graphics memory 114, such as GDDR5 memory, in the illustrated embodiment. Other types of graphics memory may be used as well. Furthermore, all or a portion of the graphics memory may reside on the GPU.
[0014] GPU 102 and graphics card 100 have additional interfaces, including a Peripheral Component Interconnect Express (PCIe) interface 116 coupled to GPU 102, a graphics output 118 on GPU 102 coupled to one or more graphics ports 120, such as a DisplayPort or HDMI port, and an Ethernet port 122 coupled to NIC 112. NIC 112 may also communicate with a host CPU (not shown) via PCIe interface 116, as indicated by data path 123.
[0015] In addition to having a GPU with an on-chip NIC, a graphics card may include a GPU coupled to an off-chip NIC, as shown in graphics card 100a of FIG. 1a. Under this configuration, GPU 102a includes an I / O interface 124 coupled to NIC 112a. I / O interface 124 is coupled to I / O interface 110 on H.264 / H.265 video codec 106. NIC 112a is also coupled to PCIe interface 116, as shown by link 125.
[0016] In addition to integrating the NIC onto the GPU (e.g., as embedded circuitry and / or as a circuit die on the same substrate as the GPU circuitry), multi-chip modules or packages containing GPU and NIC chips may also be used. An example of this is shown in Figure 1b, where graphics card 100b includes GPU 102b and NIC 112b, which are part of multi-chip module 126.
[0017] In another embodiment (not shown), the CPU and GPU 100 or 100a may be integrated into a system on a chip (SoC), or the CPU, GPU 100b, and NIC 112b may be implemented in a multi-chip module or package, or a CPU+GPU SoC, with the NIC chip implemented in the multi-chip module or package.
[0018] Typically, graphics cards 100, 100a, and 100b are installed in PCIe slots in servers and the like, implemented as mezzanine cards in servers, daughterboards in blade servers or server modules, etc. As described and exemplified below, similar components may be implemented as graphics chipsets, etc. for devices having other form factors, such as laptops, notebooks, tablets, and mobile phones.
[0019] Frame information may be obtained from the frame buffer 104, such as the pixel resolution of the frame, the frame buffer format (e.g., RGBA 8-bit or RGBA 32-bit, etc.), and access to the frame buffer pointer, which may change over time if double or triple buffering is used for rendering. Furthermore, for some implementations, depth data from the GPU buffer may be obtained in addition to color data information. For example, in scenarios such as stereoscopic gaming, it may be advantageous to stream depth data to the client along with color data.
[0020] 2 illustrates one embodiment of a cloud gaming implementation, including a game server 200 coupled to a desktop game client 202 via a network 204. Each of the game server 200 and the desktop game client 202 includes a respective instance of the graphics card 100 of FIG. 1, as represented by graphics cards 100-1 and 100-2. The game server 200 includes a CPU 206, which comprises a multi-core processor coupled to a main memory 208, in which game server software 210 is loaded to execute on one or more cores of the CPU 206. The CPU 206 is coupled to the graphics card 100-1 via a PCIe interface 116, and an Ethernet port 122-1 is coupled to a network 204, which represents multiple interconnected networks, such as the Internet.
[0021] In practice, the cloud gaming server may distribute content via a distribution network (CDN) 228. As shown in Figure 2, the CDN 228 is located between the game server 200 and the network 204.
[0022] Desktop game client 202 generally represents various types of game clients that may be implemented using a desktop computer or the like. In the illustrated embodiment, graphics card 100-2 is a PCIe graphics card installed in a PCIe slot of the desktop computer. In some cases, the PCIe graphics card may connect via a single PCIe slot that occupies multiple expansion slots for the desktop computer. Desktop game client 202 includes CPU 212, which is a multi-core processor coupled to main memory 214 into which client-side game software 216 is loaded and executed by one or more cores on CPU 212. Ethernet port 122-2 of graphics card 100-2 is coupled to network 204. In the case of a typical game player, desktop game client 202 is coupled to a local area network (LAN), which includes a switch coupled to a cable modem or similar wide area network (WAN) access device. This is coupled to Internet Service Provider (ISP) network 218, which is then coupled to network 204.
[0023] FIG. 2a illustrates one embodiment of a cloud gaming implementation that includes a game server 200a coupled to a desktop game client 202a via a network 204 and an ISP network 218. Each of the game server 200a and the desktop game client 202a has a respective instance of the graphics card 100a of FIG. 1a, shown as graphics cards 100a-1 and 100a-2. In general, like-numbered components and blocks in FIGS. 2 and 2a are similar and perform similar operations. As explained in more detail below, the difference between the cloud gaming implementations of FIGS. 2 and 2a is the outgoing video control inputs and non-image data processing. At the same time, the processing and transfer of image data in the embodiments of FIGS. 2 and 2a is similar.
[0024] 3 shows a cloud gaming implementation including a game server 200 coupled to a laptop game client 301 via a CDN 228, a network 303, and an ISP network 318. The game server 200 has the same configuration as described above and shown in FIG. 2. Optionally, the game server 200 may be replaced by the game server 200a of FIG. 2.
[0025] Laptop gaming client 301 includes main board 300, which includes GPU 302 coupled to graphics memory 314, CPU 326 coupled to main memory 328 and GPU 302, Wi-Fi chip 315, DisplayPort and / or HDMI port 320, and USB-C interface 332. As before, client-side game software is loaded into main memory 328 and runs on one or more cores on CPU 326. GPU 302 includes frame buffer 304, which is accessed by H.264 / H.265 video codec 306 via interface 308. H.264 / H.265 video codec 306 includes I / O interface 310, which is coupled to network interface 313, which is coupled to hardware-based network stack 317. Generally, hardware-based network stack 317 may be integrated on Wi-Fi chip 315 or may include separate components. Generally, the laptop gaming client includes a mobile chipset (not shown) coupled to the CPU 326 that supports various communication ports and I / O interconnects, such as USB-C, USB 3.0, USB 2.0, and PCIe interconnects.
[0026] Under the illustrated configuration for laptop gaming client 301, wireless communication is facilitated by wireless access point 324 and antenna 319. As previously mentioned, the wireless access point is connected to a cable modem or similar ISP access means that provides connection to ISP network 318. Optionally, an Ethernet adapter may be connected to USB-C interface 332, allowing laptop gaming client 301 to use an Ethernet link to ISP network 318 (via an Ethernet switch and cable modem).
[0027] The main board 300 is housed within a laptop housing to which is coupled a display 334. Generally, the display is driven by applicable circuitry either integrated into the GPU 302 or implemented on a separate component coupled to the GPU 302, as illustrated by LCD driver 336.
[0028] In embodiments herein, the NIC may be configured via software running on the CPU (such as an operating system and / or a NIC driver), platform / server firmware, and / or a GPU that receives configuration information from software running on the CPU or platform / server firmware. For example, in some embodiments, the NIC is implemented as a PCIe endpoint and is part of the PCIe hierarchy of a PCIe interconnect structure managed by the CPU. In other embodiments, software on the GPU provides instructions to the GPU on how to configure the NIC.
[0029] Video Encoding and Streaming Guide
[0030] Under aspects of the embodiments disclosed herein, techniques are provided for streaming video game image data to end user devices operated by players (also known as player devices) in a manner that reduces latency. Aspects of streaming video game image data, when using frame encoding and decoding, may use the same codecs (coder / decoders) used for video streaming. Therefore, to better understand how embodiments may be implemented, a description of basic aspects of video compression and decompression technology is first provided. In addition to the details herein, further details regarding how video compression and decompression are implemented are available from many online sources, including the source of much of the description that follows. http: / / www.eetimes.com / document.asp?doc_id=1275437 Includes an article from EETimes.com titled "How Video Compression Works," available at
[0031] At a basic level, streaming video content is played back on a display as a sequence of "frames" or "pictures." Each frame, when rendered, consists of an array of pixels whose dimensions correspond to the playback resolution. For example, Full HD (High Definition) video has a resolution of 1920 horizontal pixels by 1080 vertical pixels, commonly known as 1080p (progressive) or 1080i (interlaced). The frames are then displayed at the frame rate, with the frame data being refreshed (and re-rendered, if applicable) at the frame rate. For many years, standard-definition (SD) televisions used a refresh rate of 30i (30 frames per second (fps) interlaced), which corresponds to updating two fields of interlaced video content alternately every 1 / 30th of a second. This created the illusion of a frame rate of 60 frames per second. It should also be noted that SD content has historically been analog video, which uses raster scanning rather than pixel scanning for display. SD video on digital displays has a resolution of 480 lines, while the analog signal used for decades actually had about 525 scan lines. As a result, DVD content has historically been encoded at 480i or 480p in NTSC (National Television System Committee) markets such as the United States.
[0032] Cable and satellite television providers stream video content over fiber optic or wired cable, or over the air (long-distance wireless). Terrestrial broadcasts are similarly transmitted over the air and have traditionally been broadcast as analog signals, but since about 2010 all high-power television stations have been required to broadcast exclusively using digital signals. Digital television broadcast signals in the United States commonly include 480i, 480p, 720p (1280 x 720 pixel resolution), and 1080i.
[0033] Blu-ray Disc (BD) video content was introduced in Japan in 2003 and officially released in 2006. Blu-ray Disc supports video playback at up to 1080p, which corresponds to 1920 x 1080 at 60 (59.94) fps. While BD supports up to 60 fps, much BD content (especially recent BD content) is actually encoded at 24 fps progressive (also known as 1080 / 24p), the frame rate historically used for film. The conversion from 24 fps to 60 fps is typically done using a 3:2 "pulldown" technique, under which frame content is repeated in a 3:2 pattern, which can result in various types of video artifacts, especially when playing content with a lot of movement. Furthermore, newer "smart" TVs have refresh rates of 120 Hz or 240 Hz, each an even multiple of 24. As a result, these TVs support a 24 fps "movie" or "cinema" mode using an HDMI (High Definition Multimedia Interface) digital video signal, displaying the 24 fps content at 120 fps or 240 fps by repeating extracted frame content with 5:5 or 10:10 pulldown to match the TV's refresh rate. More recently, smart TVs from manufacturers such as Sony and Samsung support a playback mode that creates multiple interpolated frames between the actual 24 fps frames to create a smoothing effect.
[0034] A compatible Blu-ray Disc playback device is required that supports three video encoding standards: H.262 / MPEG-2 Part 2, H.264 / MPEG-4 AVC, and VC-1. Each of these video encoding standards operates in a similar manner, with the following caveats: there are some differences between these standards.
[0035] In addition to video content encoded on DVDs and Blu-ray® Discs, a significant amount of video content is distributed using video streaming technologies. The encoding technologies used for streaming media such as movies and television programs may generally be the same as or similar to those used for Blu-ray content. For example, Netflix® and Amazon® Instant Video each use VC-1 (in addition to other streaming formats depending on the capabilities of the playback device). VC-1 is a proprietary video format originally developed by Microsoft and released as a SMPTE (Society of Motion Picture and Television Engineers) video codec standard in 2006. Meanwhile, YouTube® uses the same mixture of video encoding standards commonly used to record uploaded video content, most of which is recorded using consumer-level video recording equipment (e.g., camcorders, mobile phones, and digital cameras) rather than the professional-level equipment used to record original television content and some recent (re-release) movies.
[0036] As an example of how much video content is being streamed, recent measurements show that during peak consumption periods, Netflix streaming used more than one-third of the bandwidth of Comcast's cable internet service. In addition to Full HD (1080p) streaming since 2011, Netflix, Amazon, Hulu®, and others are streaming an ever-increasing amount of video content in 4K video (3840x2160), also known as Ultra-High Definition or UHD.
[0037] More advanced smart TVs globally support playback of streaming media delivered over IEEE 802.11-based wireless networks (commonly referred to as Wi-Fi). Furthermore, most new Blu-ray players support Wi-Fi streaming of video content, as do all smartphones. Additionally, many modern smartphones and tablets support wireless video streaming methods using Wi-Fi Direct, wireless Mobile High-Direction Link (MHL), and similar standards, allowing videos to be played back by the smartphone or tablet and viewed on the smart TV. Furthermore, the data service bandwidth available on LTE (Long-Term Evolution) and fifth-generation (5G) mobile networks has made services such as Internet Protocol Television (IPTV) a viable means of watching television and other video content over mobile networks.
[0038] At 1080p resolution, each frame contains approximately 2.1 million pixels. Using only 8-bit pixel encoding would require a data streaming rate of nearly 17 million bits per second (mbps) to support a frame rate of just 1 frame per second if the video content were delivered as raw pixel data. Because this would be impractical, video content is encoded in highly compressed formats.
[0039] Still images, such as those displayed using an internet browser, are typically encoded using JPEG (Joint Photographic Experts Group) or PNG (Portable Network Graphics) encoding. The original JPEG standard defines a "lossy" compression scheme, in which pixels in the decoded image may differ from the original image. In contrast, PNG uses a "lossless" compression scheme. Because lossless video is impractical on many levels, various video compression standards organizations, such as the Motion Photographic Expert Group (MPEG), which defined the original MPEG-1 compression standard (1993), use lossy compression techniques that involve still-image encoding of intraframes ("I-frames") (also known as "key" frames) in combination with motion estimation techniques used to generate other types of frames, such as predicted frames ("P-frames") and bidirectional frames ("B-frames").
[0040] Because digitized video content consists of a sequence of frames, video compression algorithms use concepts and techniques used in still image compression. Still image compression uses a combination of block encoding and advanced mathematics to substantially reduce the number of bits used to encode an image. For example, JPEG divides an image into 8x8 pixel blocks and uses a discrete cosine transform (DCT) to convert each block into a frequency domain representation. In general, other block sizes other than 8x8 and algorithms other than the DCT may be used for the block transform operation in other standard-based, proprietary, and other compression schemes.
[0041] The DCT transform is used to facilitate frequency-based compression techniques. Human vision is more sensitive to information contained in low frequencies (corresponding to larger image features) than to information contained in high frequencies (corresponding to smaller features). The DCT helps distinguish perceptually more important information from perceptually less important information.
[0042] After the block transform, the transform coefficients of each block are compressed using quantization and coding. Quantization reduces the precision of the transform coefficients in a biased manner, i.e., more bits are used for low-frequency coefficients and fewer bits are used for high-frequency coefficients. This takes advantage of the fact that, as mentioned above, human vision is more sensitive to low-frequency information, so high-frequency information is more approximate.
[0043] The number of bits used to represent the quantized DCT coefficients is then reduced by "coding" that exploits some of the coefficients' statistical properties. After quantization, many of the DCT coefficients (often the majority of high-frequency coefficients) are zero. A technique called "run-length coding" (RLC) exploits this fact by grouping consecutive zero-valued coefficients ("runs") and encoding the number of coefficients ("length"), instead of encoding each individual zero-valued coefficient.
[0044] Run-length coding is typically followed by variable-length coding (VLC), in which commonly occurring symbols (representing quantized DCT coefficients or runs of zero-valued quantized coefficients) are represented using codewords containing only a few bits, while less common symbols are represented with longer codewords. By using fewer bits for the most common symbols, VLC reduces the average number of bits required to encode a symbol, thereby reducing the number of bits required to encode the entire image.
[0045] At this stage, all of the aforementioned techniques operate on each 8x8 block independently of other blocks. Because images typically contain features much larger than an 8x8 block, more efficient compression can be achieved by considering similarities between neighboring blocks within the image. To exploit such similarities between blocks, a prediction step is often added before quantization of the transform coefficients. In this step, the codec attempts to predict the image information within a block using information from surrounding blocks. Some codecs (e.g., MPEG-4) perform this step in the frequency domain by predicting DCT coefficients. Other codecs (e.g., H.264 / AVC) perform this step in the spatial domain, predicting pixels directly. This latter approach is called "intra prediction."
[0046] In this operation, the encoder attempts to predict the values of some of the DCT coefficients (if done in the frequency domain) or pixel values (if done in the spatial domain) of each block based on the coefficients or pixels of surrounding blocks. The encoder then calculates the difference between the actual value and the predicted value and encodes the difference instead of the actual value. At the decoder, the coefficients are reconstructed by performing the same prediction and then adding the difference transmitted by the encoder. Because this difference tends to be small compared to the actual coefficient value, this technique reduces the number of bits required to represent the DCT coefficients.
[0047] When predicting the DCT coefficients or pixel values of a particular block, the decoder has access only to the values of surrounding blocks that have already been decoded. Therefore, the encoder must predict the DCT coefficients or pixel values of each block based only on values from previously encoded surrounding blocks. JPEG uses a very basic DCT coefficient prediction scheme in which only the lowest frequency coefficient (the "DC coefficient") is predicted using simple differential coding. MPEG-4 video uses a more refined scheme to attempt to predict the first DCT coefficient of each row and each column of an 8x8 block.
[0048] In contrast to MPEG-4, in H.264 / AVC, prediction is performed directly on pixels, and an integer transform such as the DCT always processes the residual, either from motion estimation or intra-prediction. In H.264 / AVC, pixel values are never directly transformed, as in JPEG or MPEG-4 I-frames. As a result, the decoder must decode the transform coefficients and perform an inverse transform to obtain the residual, which is added to the predicted pixels.
[0049] Another widely used video codec is High Efficiency Video Coding (HEVC), also known as H.265 (as used herein) and MPEG-H Part 2. Compared to H.264 / AVC, HEVC offers 25% to 50% better data compression at the same level of video quality, or substantially improved video quality at the same bitrate. It supports resolutions up to 8192x4320, including 8K UHD, and unlike primarily 8-bit AVC, HEVC's high-fidelity Main10 profile is built into nearly all supporting hardware. HEVC uses integer DCT and DST transforms with block sizes of 4x4 to 32x32.
[0050] Color images are typically represented using several "color planes." For example, an RGB color image includes a red plane, a green plane, and a blue plane. When overlaid and blended, the three planes make up the full-color image. To compress a color image, the still image compression techniques described above may be applied to each color plane in turn.
[0051] Imaging and video applications often use color schemes in which color planes do not correspond to specific colors. Instead, one color plane contains luminance information (the overall brightness of each pixel in a color image), and two or more color planes contain color (chrominance) information that, when combined with the luminance, may be used to derive specific levels of the red, green, and blue components of each image pixel. Such color schemes are useful because the human eye is more sensitive to luminance than to color, and therefore the chrominance planes can be stored and / or encoded at a lower image resolution than the luminance information. In many video compression algorithms, the chrominance planes are encoded at half the horizontal resolution and half the vertical resolution of the luminance plane. Thus, for every 16-pixel-by-16-pixel region in the luminance plane, each chrominance plane contains one 8-pixel-by-8-pixel block. In typical video compression algorithms, a "macroblock" is a 16-by-16 region in a video frame containing four 8-by-8 luminance blocks and two corresponding 8-by-8 chrominance blocks.
[0052] Video compression algorithms and still-image compression algorithms share many compression techniques, but a key difference is how they handle motion. One extreme approach is to encode each frame using JPEG or a similar still-image compression algorithm and then have the player decode the resulting JPEG frames. JPEG and similar still-image compression algorithms can produce good-quality images at compression ratios of approximately 10:1, while more advanced compression algorithms can produce similar quality at compression ratios as high as 30:1. While 10:1 and 30:1 are fairly large compression ratios, video compression algorithms can provide good-quality video at compression ratios of up to approximately 200:1. This is achieved using video-specific compression techniques, such as motion estimation and motion compensation, in combination with still-image compression techniques.
[0053] For each macroblock in the current frame, motion estimation attempts to find a close matching region in an already encoded frame (called a "reference frame"). The spatial offset between the current block and a block selected from the reference frame is called a "motion vector." The encoder calculates the pixel-by-pixel difference between the selected block from the reference frame and the current block and sends this "prediction error" along with the motion vector. In most video compression standards, motion-based prediction may be bypassed if the encoder does not find a good match for a macroblock. In this case, the macroblock itself is encoded instead of the prediction error.
[0054] It should be noted that a reference frame is not necessarily the immediately preceding frame in a sequence of displayed video frames. Rather, video compression algorithms typically encode frames in an order different from the order in which they are displayed. An encoder may skip forward several frames, encode a future video frame, and then skip backward and encode the next frame in the display sequence. This is done so that motion estimation can be performed backward in time using the encoded future frame as a reference frame. Also, video compression algorithms typically allow the use of two reference frames: a previously displayed frame and a previously encoded future frame.
[0055] Video compression algorithms periodically encode intra-frames using only still image coding techniques, without relying on previously encoded frames. If a frame in the compressed bitstream is corrupted by an error (e.g., a dropped packet or other transmission error), the video decoder may "restart" at the next I-frame, which does not require a reference frame for reconstruction.
[0056] 4 shows an exemplary frame encoding and display scheme consisting of an I-frame 400, a P-frame 402, and a B-frame 404. As mentioned above, I-frames are encoded cyclically in a manner similar to still images and do not depend on other frames. P-frames (predictive frames) are encoded using only previously displayed reference frames, as indicated by previous frame 406. B-frames (bidirectional frames), on the other hand, are encoded using both future and previously displayed reference frames, as indicated by previous frame 408 and future frame 410.
[0057] The bottom of Figure 4 shows an exemplary frame encoding sequence (progressing downward) and the corresponding display playback order (progressing rightward). In this example, each P frame is followed by three B frames in encoding order. However, in display order, each P frame is displayed after three B frames, indicating that the encoding order and display order are not the same. It should be further noted that the occurrence of P and B frames generally varies depending on the amount of motion present in the captured video. The use of one P frame followed by three B frames here is for simplicity and easy understanding of how I, P, and B frames are implemented.
[0058] One factor that complicates motion estimation is that the displacement of an object from a reference frame to a current frame may be a non-integer number of pixels. To handle such situations, modern video compression standards allow motion vectors to have non-integer values, resulting in motion vector resolutions of, for example, one-half or one-quarter of a pixel. To support block-matching searches in fractional pixel displacements, encoders employ interpolation to estimate pixel values in the reference frame at non-integer positions.
[0059] Due in part to processor limitations, motion estimation algorithms use various methods to select a limited number of promising candidate motion vectors (approximately 10-100 vectors in most cases) and evaluate only the 16x16 region (or up to a 32x32 region in the case of H.265) corresponding to these candidate vectors. One approach is to select candidate motion vectors in several stages and then select the best motion vector as a result. Another approach analyzes previously selected motion vectors for surrounding macroblocks in the current and previous frames to attempt to predict motion within the current macroblock. Based on this analysis, a small number of candidate motion vectors are selected and only these vectors are evaluated.
[0060] By selecting a small number of candidate vectors instead of exhaustively scanning the search area, the computational demands of motion estimation can be significantly reduced, sometimes by more than two orders of magnitude. However, there is a trade-off between processing load and image quality or compression efficiency. Generally, by searching more candidate motion vectors, the encoder can find blocks in the reference frame that better match each block in the current frame, thereby reducing the prediction error. The smaller the prediction error, the fewer bits required to encode the image. Therefore, increasing the number of candidate vectors allows for a reduction in the compression bit rate, at the expense of performing more computation. Alternatively, by increasing the number of candidate vectors while keeping the compressed bit rate constant, the prediction error can be encoded with greater precision, improving image quality.
[0061] Some codecs (including H.264 and H.265) allow for the subdivision of 16x16 macroblocks into smaller blocks (e.g., various combinations of 8x8, 4x8, 8x4, and 4x4 blocks) to lower prediction errors. Each of these smaller blocks may have its own motion vector. For such schemes, the motion estimation search begins by finding a good position for the entire 16x16 block (or 32x32 block). If the match is close enough, no further subdivision is necessary. However, if the match is insufficient, the algorithm starts from the best position found so far and further subdivides the original block into 8x8 blocks. For each 8x8 block, the algorithm searches for the best position near the position selected by the 16x16 search. Depending on how quickly a good match is found, the algorithm may continue the process using smaller blocks such as 8x4, 4x8, etc.
[0062] During playback, the video decoder performs motion compensation using motion vectors encoded in the compressed bitstream to predict the pixels of each macroblock. If the horizontal and vertical components of the motion vector are both integer-valued, the predicted macroblock is simply a copy of a 16-pixel by 16-pixel region of the reference frame. If any component of the motion vector is not integer-valued, interpolation is used to estimate the image at non-integer pixel locations. The prediction error is then decoded and added to the predicted macroblock to reconstruct the actual macroblock pixels. As mentioned above, for codecs such as H.264 and H.265, a 16x16 (or up to 32x32) macroblock may be divided into smaller sections with independent motion vectors.
[0063] Ideally, lossy image and video compression algorithms discard only perceptually insignificant information, so that the reconstructed image or video sequence appears identical to the original uncompressed image or video to the human eye. However, in practice, some artifacts may be visible, especially in scenes with greater motion, such as when a scene is panned. This may occur due to poor encoder implementation, video content that is particularly difficult to encode, or a selected bitrate that is too low for the video sequence, resolution, and frame rate. The latter case is particularly common, as many applications trade off video quality for reduced storage and / or bandwidth requirements.
[0064] Two types of artifacts commonly occur in video compression applications: "blocking" and "ringing." Blocking artifacts result from the fact that compression algorithms divide each frame into 8x8 blocks. Each block is reconstructed with some small error, and the error at the edge of a block often contrasts with the error at the edge of an adjacent block, making the block boundary visible. In contrast, ringing artifacts appear as distortion around the edges of image features. Ringing artifacts result from the encoder discarding too much information when quantizing high-frequency DCT coefficients.
[0065] To reduce blocking and ringing artifacts, video compression applications often use filters after decompression. These filtering steps are known as "deblocking" and "deringing," respectively. Alternatively, deblocking and / or deringing may be integrated into the video decompression algorithm. This approach, sometimes called "loop filtering," uses filtered reconstructed frames as reference frames for decoding future video frames. For example, H.264 includes an "in-the-loop" deblocking filter, sometimes referred to as a "loop filter."
[0066] End-to-end image data flow example
[0067] FIG. 5 illustrates an example of end-to-end image data flow between a game server 200 and a desktop game client 202, according to one embodiment. Further, related operations are illustrated in a flowchart 600 shown in FIG. 6. The example of FIG. 5 illustrates communication between a server graphics card 100-1 in the game server 200 and a client graphics card 100-2 in the desktop game client 202. In general, the communication shown in FIG. 5 may be communication between any type of device that generates game image content and any type of device that has a client that receives and processes the game image content, such as a cloud game server and a game device operated by a player. In this example, audio content is depicted as being transferred between the server graphics card 100-1 and the client graphics card 100-2. In some implementations, audio content is transferred through separate network interfaces (e.g., separate NICs or network cards) on the server and / or client (not shown). In some implementations using separate network interfaces, streaming session communications and control communications are transmitted over separate communication paths not shown in FIG. 5.
[0068] As shown in block 602 of flowchart 600, the process begins by establishing a streaming session between a server and a client. In general, any type of existing or future streaming session may be used, and the teachings and principles disclosed herein are generally agnostic to the particular type of streaming session. The types of streaming protocols that can be used include, but are not limited to, existing streaming protocols such as RTMP (Real-Time Messaging Protocol), RTSP (Real-Time Streaming Protocol) / RTP (Real-Time Transport Protocol), and HTTP-based adaptive protocols such as Apple HLS (HTTP Live Streaming), Low-Latency HLS, MPEG-DASH (Moving Picture Expert Group Dynamic Adaptive Streaming over HTTP), Low-Latency CMAF for DASH (Common Media Application Format for DASH), Microsoft Smooth Streaming, and Adobe HDS (HTTP Dynamic Streaming). Additionally, newer technologies such as SRT (Secure Reliable Transport) and webRTC (Web Real-Time Communications) may also be used. In one embodiment, an HTTP or HTTPS streaming session is established to support one of the HTTP-based adaptation protocols.
[0069] 5 illustrates two network communications between NIC 112-1 and NIC 112-2: a Transmission Control Protocol over Internet Protocol (TCP / IP) connection 500 and a Universal Datagram Protocol over IP (UDP / IP) stream 502. For simplicity, these are commonly referred to as TCP and UDP. TCP is a reliable connection protocol, and TCP packets 504 are transmitted from a transmitter to a receiver. The receiver acknowledges receipt of packets by sending ACKnowledgments (ACKs) 506 indicating successfully received frame sequences. Sometimes, TCP packets / frames are dropped or otherwise received with errors, as indicated by TCP packet 508. In response to detecting a missing or erroneous packet, the receiver sends a negative ACK (NACK) 510 containing information identifying the missing or erroneous packet. The missing or erroneous packet is then retransmitted, as indicated by retransmitted packet 508R. As further shown, TCP / IP connection 500 may be used to receive game control input 512 from desktop game client 202.
[0070] HTTP streaming sessions are set up using TCP. However, depending on the streaming protocol, video and / or audio content may use UDP. UDP, widely used for live streaming, is a connectionless, unreliable protocol. UDP uses a "best effort" transport, which means that packets may be dropped or erroneous packets may be received. In either case, the missing or erroneous packets are ignored by the receiver. The stream of UDP packets 514 shown in Figure 5 is used to illustrate the packets of video and (optionally) audio content. Some hybrid media transport schemes employ a combination of TCP and UDP transport.
[0071] Returning to flowchart 600, at block 604, a sequence of raw video frames is generated by GPU 102-1 in server graphics card 100-1 via execution of game software on game server 200, as shown by frame 605. As the sequence of raw video frames is generated, the contents of the individual frames are copied to frame buffer 104, with multiple individual frames stored in frame buffer 104 at any given time. At block 606, the frames are encoded using an applicable video codec, for example, an H.264 or H.265 codec, to generate a video stream. This is performed by H.264 / H.265 codec 106-1, which reads the raw video frame contents from frame buffer 104 and generates an encoded video stream, as shown by video stream 516. As will be appreciated by those skilled in the art, and as described in the above-mentioned guides, the generated video stream includes compressed and encoded content corresponding to a sequence of I, P, and B frames ordered to enable decoding and playback of the raw video frame content at the desktop game client 202.
[0072] As shown by block 607 and audio stream generation block 518 in Figure 5, game audio content is encoded in a streaming format in parallel with the generation and encoding of game image frames. Typically, audio content is generated by game software running on the game server CPU, and encoding of the audio content may be performed using either software or hardware. In this example, encoding is performed external to server graphics card 100-1. In some embodiments, either GPU 102-1 or other circuitry on server graphics card 100-1 (not shown) may be used to encode the audio content.
[0073] In block 608, the video stream content is packetized by the NIC 112-1. Optionally, the audio stream may also be packetized by the NIC 112-1. In one approach, the video stream and the audio stream are sent as separate streams (in parallel). There is information in one of the streams that is used to synchronize the audio and video content via playback functionality in the game client. In another approach, the video and audio content are combined and sent as a single packetized stream. In general, any existing or future video and audio streaming packetization scheme may be used.
[0074] As shown in block 610, AV (audio and video) content is streamed over a network from server graphics card 100-1 to client graphics card 100-2. As shown in Figure 5, the corresponding content is streamed via UDP packets 514, which represent one or more UDP streams used to transmit the AV content.
[0075] The receive-side operations are performed by client graphics card 100-2. When one or more UDP streams are received, the audio and video content is buffered in one or more UDP buffers 520 in NIC 112-2 and then depacketized, as indicated by block 612 in flowchart 600. As indicated by block 618 in flowchart 600 and audio decode and sync block 522 in FIG. 5, in embodiments where audio processing is not handled by the GPU or client graphics card, the depacketized audio content is separated and transferred to the host CPU for processing and output of the audio content. Optionally, audio processing may be performed by applicable functionality (not shown) on client graphics card 100-2.
[0076] In block 614, the video (game image) frames are decoded using the applicable video codec. In the example of FIG. 5, this is performed by the H.264 / H.265 codec 106-2. Various mechanisms may be used to transfer the depacketized encoded video content from the NIC 112-2 to the I / O interface 110 (111) on the H.264 / H.265 codec 106-2. For example, a work descriptor scheme may be used, in which the NIC writes a work descriptor to a memory location on the GPU 102-2, and then writes the corresponding "work" (e.g., encoded video data segments) to a location on the GPU 102-2 or to graphics memory (not shown) on the client graphics card 102-2. In another embodiment, a "doorbell" scheme may be used in which NIC 112-2 depacketizes available encoded video segments and the H.264 / H.265 codec 106-2 posts a doorbell when it reads the encoded video segment from NIC 112-2. Additionally, other types of queuing mechanisms may also be used. In some embodiments, a circular first-in-first-out (FIFO) buffer or queue, such as a circular FIFO, is used.
[0077] As shown in Figure 5, H.264 / H.265 codec 106-2 performs video stream decode processing 524 and frame generation (playback) 526 within decode and reassembly block 527. As indicated by block 616 and video frame 605 in Figure 6, playback frames may be written to GPU frame buffer 104-2 and then output to a display for desktop game client 202. For example, GPU 102-2 may generate game image frames and output a corresponding video signal via an applicable video interface (e.g., HDMI, DisplayPort, USB-C) for viewing on a monitor or other type of display. As indicated by audio / video output block 528, audio and video content are output to speakers and a display, respectively, for desktop game client 202.
[0078] Typically, when a NIC integrated on a GPU or coupled to a GPU on a client graphics card is used to process TCP traffic, received TCP packets are buffered in one or more TCP buffers 530. As described and illustrated below, each of NIC 112-1 and NIC 112-2 has the functionality to implement a complete network stack in hardware. Typically, received TCP packets are packetized and forwarded to the host CPU for further processing. The forwarding may be by current means, such as Direct Memory Access (DMA), using PCIe write transactions.
[0079] Tile-based games
[0080] Many popular games employ tiles and associated tilemaps, which can offer performance gains compared to using video encoding / decoding techniques.
[0081] Diagrams 700a and 700b in Figures 7a and 7b, respectively, illustrate operations performed on a game server and a game client for a tile-based game, according to one embodiment. As shown in Figure 7a, a tile 702, a full-frame image, is composed of multiple tiles 702 arranged in an XY grid. During game play, game software running on the game server generates tiles, as indicated by tile generation block 704. The tiles are written to one or more tile buffers 705. A tile encoder 706 encodes the tiles using an image compression algorithm to generate encoded tiles 708, after which the image data in the encoded tiles is packetized by packetization logic 712 on a NIC 710. The NIC 710 then transmits a stream of encoded tiles 714 over a network where they are delivered to the game clients.
[0082] Referring now to diagram 700b in FIG. 7b, a game client receives a stream of encoded tiles 714 at a NIC 716. The NIC 716 performs depacketization 718 to output encoded tiles 708. The encoded tiles are then written to a tile buffer 720 (or some other memory space located on or accessible to the GPU). A decode and regenerate tile block 722 is used to read the encoded tile content from the tile buffer 720 and decode the encoded tile content to regenerate the original tiles, shown as regenerated tiles 702R.
[0083] FIG. 1c illustrates one embodiment of a graphics card 100c including a GPU 102c configured to support the server-side and client-side operations shown in diagrams 700a and 700b. As indicated by the like-numbered components and blocks in FIGS. 1 and 1c, the configurations of graphics cards 100 and 100c are similar. The difference is that GPU 102c includes a tile encoder and decoder 706 having an I / O interface 111. The tile encoder and decoder is configured to perform the encoding operations of tile encoder 706 of FIG. 7a and at least the decoding operations for decoding and reproducing tile blocks 722 of FIG. 7b. In one embodiment, the complete logic for decoding and reproducing tile blocks 722 is implemented within tile encoder and decoder 706. Optionally, portions of the tile reproduction logic and other logic related to reassembling game tiles may be implemented in separate blocks (not shown).
[0084] Cloud-based game server with multiple graphics cards
[0085] In one approach, a cloud gaming server includes multiple graphics cards, as shown in cloud gaming server 800 of FIG. 8. Cloud gaming server 800 includes m graphics cards 100 (denoted as graphics cards 100-1, 100-2, 100-3, 100-4, 100-m), each occupying a respective PCIe slot (also known as an expansion slot) on the server's main board. The server's main board further includes one or more CPUs 806 coupled to main memory 808 into which game software 810 is loaded. Cloud gaming server 800 also includes one or more network adapter cards 812 installed in respective PCIe slots, each including a NIC chip 814, a PCIe interface 816, and one or more Ethernet ports, as shown by Ethernet ports 818 and 820.
[0086] In the embodiment of cloud gaming server 800a of Figure 8a, a NIC chip 815 including a PCIe interface 817 is mounted on the server's main board and coupled to CPU 806 via an applicable interconnect structure. For example, CPU 806 may include a PCIe root port (RP) 821 to which PCIe interface 817 is coupled via PCIe link 823.
[0087] 8b, the CPU 806, main memory 808, and NIC chip 815 are mounted on a mainboard within a blade server 824, which includes a PCIe interface 826. The cloud gaming server 800b includes a backplane, midplane, or baseplane 828 having a number of expansion slots or connectors, as indicated by slots / connectors 830 and 832. Each of the blade server 824 and m graphics cards 100-1, 100-2, 100-3, 100-4, 100-m is installed in a corresponding expansion slot or includes a connector that mates with a mating connector on the backplane, midplane, or baseplane 828.
[0088] The cloud gaming server is configured to scale game hosting capacity by using graphics card 100 to generate and stream game image data, while using one or more network adapter cards 812 or NICs 815 to process game control inputs and set up and manage streaming connections. Thus, the integrated NIC on graphics card 100 is not burdened with processing I / O traffic related to real-time control inputs, streaming configuration, and management traffic; rather, the integrated NIC only needs to process outbound image data traffic. Furthermore, latency is reduced because the data path flows directly from the image data encoder (which in this example is an H.264 / H.265 codec, but which may be a tile encoder / decoder in other embodiments). Furthermore, game audio content may be streamed using NIC 815 or network adapter card 812. In other embodiments, audio content is streamed using graphics card 100, as described above.
[0089] 9 illustrates block-level components implemented within an integrated NIC 900, according to one embodiment. The NIC 900 includes a NIC processor 902 coupled to memory 904, one or more network ports 906 (e.g., Ethernet ports) including a receive (RX) port 908 and a transmit (TX) port 910, a host I / O interface 912, a codec I / O interface 914, and embedded logic for implementing a network stack 916. The network port 906 includes circuitry and logic for implementing the physical layer (PHY Layer 1) and the media access channel (MAC) (Layer 2) of the Open Systems Interconnection (OSI) model.
[0090] The RX port 908 and the TX port 910 include respective RX and TX buffers in which received packets (e.g., packets A, B, C, and D) and packets to be transmitted (e.g., packets Q, R, S, and T) are buffered. Received packets are processed by an inbound packet processing block 918 and buffered in an upstream packet queue 920. Outbound packets are queued in a downstream packet queue 922 and processed using an outbound packet processing block 924.
[0091] Flow rules 926 are stored in memory 904 and are used to determine where received packets should be forwarded. For example, inbound video packets may be forwarded to a video codec or tile decoder, while game control and session management packets may be forwarded to the host CPU. NIC 900 may include optional DMA logic 928 that allows the NIC to write packet data to main memory (via host I / O interface 912) and / or directly to graphics memory.
[0092] The host I / O interface includes an input FIFO queue 930 and an output FIFO queue 932. Similarly, the codec I / O interface 914 includes an input FIFO queue 934 and an output FIFO queue 936. The counterpart host I / O on a GPU or graphics card and the counterpart codec I / O interface in a video codec include similar input and output FIFO queues (not shown).
[0093] In one embodiment, NIC 900 includes embedded logic for implementing Network Layer 3 and Transport Layer 4 of the OSI model. For example, Network Layer 3 is commonly used for the Internet Protocol (IP), and Transport Layer 4 is used for both the TCP and UDP protocols. In one embodiment, NIC 900 further includes embedded logic for implementing Session Layer 5, Presentation Layer 6, and Application Layer 7. This enables the NIC to facilitate functions associated with these layers, such as establishing HTTP and HTTPS streaming sessions and / or implementing the various media streaming protocols mentioned above. In implementations in which these operations are handled by the host CPU, it is not necessary to include Session Layer 5, Presentation Layer 6, and Application Layer 7.
[0094] NIC processor 902 executes firmware instructions 938 to perform the functions indicated by the various blocks in Figure 9. The firmware instructions may be stored in an optional firmware storage unit 940 on NIC 900, or may be stored somewhere external to the NIC. For example, if NIC 900 is an integrated NIC on a GPU used in a graphics card, the graphics card may include a storage unit or device on which the firmware is installed. In other configurations, such as when installed on a game server, all or part of the firmware instructions may be loaded from the host during a boot operation.
[0095] In general, the functionality of the blocks illustrated for NIC 900 may be implemented using some form of embedded logic. Embedded logic generally includes logic implemented in circuitry, such as using a Field Programmable Gate Array (FPGA), or using pre-programmed or fixed hardware logic (or a combination of pre-programmed / hard-coded / programmable logic), and firmware executing on one or more embedded processors, processor elements, engines, microcontrollers, etc. For purposes of illustration, but not limitation, an example of firmware execution on NIC processor 902 is shown in FIG. 9. NIC processor 902 is a form of embedded processor that may include multiple processing elements, such as cores or microengines.
[0096] NIC 900 may also include embedded "accelerator" hardware, etc., used to perform packet processing operations such as flow control, encryption, decryption, etc. For example, NIC 900 may include one or more cipher blocks configured to perform encryption and decryption in connection with HTTPS traffic. NIC 900 may also include a hash unit to accelerate hash key matching in connection with packet flow lookups.
[0097] In the embodiments shown herein, the H.264 / H.265 codec is shown for illustrative purposes and is non-limiting. In general, existing and future video codecs may be integrated onto the GPU and used in a similar manner as shown. In addition to H.264 and H.265, such video codecs include, but are not limited to, Versatile Video Coding (VVC) / H.266, AOMedia Video (AV1), VP8, and VP9.
[0098] In addition to use in cloud gaming environments, the GPUs and graphics cards described and illustrated herein may be used in other use cases, including, but not limited to: -Live game streaming with Twitch®. Game images can be streamed directly from the graphics card, resulting in even lower latency game streams. -Can accelerate other video streaming tasks, such as various YouTube live streams. Today, video streams are encoded either on the CPU or on the GPU, depending on which makes more sense for a particular use case. When encoding is done on the GPU, the encoded frames can be sent directly through the NIC, bypassing the readback step and PC memory. -Video telephony applications such as Skype® and Zoom®. -GPUs are often used as general-purpose accelerators because they are typically much faster and have better scalability for floating-point operations compared to CPUs. In the scientific world and in stock market analysis, calculations are often performed on GPUs. For example, data can be sent out to the network faster using an integrated NIC when needed elsewhere, for example for stock market software to make buying and selling decisions. - Crypto coin mining platform using GPU as accelerator.
[0099] In addition to using the PCI interfaces, interconnects, and protocols described and illustrated herein, other interconnect structures and protocols may be used, including, but not limited to, Compute Express Link (CXL), Cache Coherent Interconnect for Accelerators (CCIX), Open Coherent Accelerator Processor Interface (OpenCAPI), and Gen-Z interconnects.
[0100] Although some embodiments have been described with reference to particular implementations, other implementations are possible according to some embodiments. Moreover, the arrangement and / or order of elements or other features illustrated in the drawings and / or described herein need not be arranged in the particular manner illustrated and described. Many other configurations are possible according to some embodiments.
[0101] In each system shown in the figures, elements may have the same or different reference numbers in some cases to indicate that the depicted elements are different and / or similar. However, elements may be flexible enough to have different implementations and operate with some or all of the systems shown or described herein. The various elements shown in the figures may be the same or different. It is arbitrary to refer to any one as a first element or any one as a second element.
[0102] In the specification and claims, the terms “coupled” and “connected,” along with their derivatives, may be used. It is understood that these terms are not intended as synonyms for each other. Rather, in particular embodiments, “connected” may be used to indicate that two or more elements are in direct physical or electrical contact with each other. “Coupled” may mean that two or more elements are in direct physical or electrical contact. However, “coupled” may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other. Furthermore, “communicatively connected” means that two or more elements, which may or may not be in direct contact with each other, are enabled to communicate with each other. For example, if component A is connected to component B, which is connected to component C, component A may be communicatively connected to component C using component B as an intermediate component.
[0103] An embodiment is an implementation or example of the invention. References in the specification to "an embodiment," "one embodiment," "an embodiment," or "another embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least some, but not necessarily all, embodiments of the invention. Various appearances of "one embodiment," "one embodiment," or "an embodiment" do not necessarily all refer to the same embodiment.
[0104] Not all components, features, structures, characteristics, etc. described and illustrated herein need be present in a particular embodiment or embodiments. For example, if the specification describes a component, feature, structure, or characteristic as "may include," "might include," "can include," or "could include," it is not necessary to include that particular component, feature, structure, or characteristic. When the specification or claims refer to "a" or "an" element, this does not mean that only one of the element is present. When the specification or claims refer to "additional" elements, it does not exclude the presence of two or more of the additional elements.
[0105] An algorithm is here, and generally, conceived to be a self-consistent sequence of acts or operations leading to a desired result. This involves physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals that may be stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like. It should be understood, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities.
[0106] In the foregoing detailed description, letters (in italics), such as "m" and "n," are used to represent integers, and the use of a particular letter is not limited to a particular embodiment. Furthermore, different claims may use the same letter or different letters to represent different integers. Furthermore, the use of a particular letter in the detailed description may or may not match the letter used in claims relating to the same subject matter in the detailed description.
[0107] As noted above, various aspects of the embodiments herein may be facilitated by corresponding software and / or firmware components and applications, such as software and / or firmware executed by an embedded processor or the like. Accordingly, embodiments of the present invention may be used as or to support software programs, software modules, firmware, and / or distributed software executing on some form of processor, processor core, or embedded logic, within a virtual machine running on the processor or core, or within a virtual machine embodied or realized on or in a non-transitory computer-readable or machine-readable storage medium. A non-transitory computer-readable or machine-readable storage medium includes any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer). For example, a non-transitory computer-readable or machine-readable storage medium includes any mechanism that provides (i.e., stores and / or transmits) information in a form accessible by a computer or computing device (e.g., computing device, electronic system, etc.), such as recordable / non-recordable media (e.g., read-only memory (ROM), random-access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.). The content may be directly executable ("object" or "executable" format), source code, or difference code ("delta" or "patch" code). A non-transitory computer-readable or machine-readable storage medium may also include a store or database from which content can be downloaded. A non-transitory computer-readable or machine-readable storage medium may also include a device or product having content stored thereon at the time of sale or distribution. Accordingly, distributing a device having stored content or offering content for download via a communications medium may be understood as providing an article of manufacture that includes a non-transitory computer-readable or machine-readable storage medium having such content as described herein.
[0108] The operations and functions performed by the various components described herein may be performed by software executing on a processing element, by embedded hardware, etc., or by any combination of hardware and software. Such components may be implemented as software modules, hardware modules, special purpose hardware (e.g., application specific hardware, ASICs, DSPs, etc.), embedded controllers, hardware-implemented circuits, hardware logic, etc. Software content (e.g., data, instructions, configuration information, etc.) may be provided via an article of manufacture including a non-transitory computer-readable or machine-readable storage medium, which provides content representing executable instructions. The content may result in a computer performing the various functions / operations described herein.
[0109] As used herein, a list of items joined by the term "at least one of" may mean any combination of the listed terms. For example, the phrase "at least one of A, B, or C" may mean A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.
[0110] The above description of exemplary embodiments of the invention, including what is described in the Abstract, is not intended to be exhaustive or to limit the invention to the precise form disclosed. For purposes of illustration, specific embodiments of, and examples for, the invention are described herein, but those skilled in the art will recognize that various equivalent modifications are possible within the scope of the invention.
[0111] These modifications to the invention may be made in light of the above detailed description. The terms used in the following claims should not be construed to limit the invention to the specific embodiments disclosed in the specification and drawings. Rather, the scope of the invention is defined entirely by the following claims, which are to be construed in accordance with established doctrines of claim interpretation.
Claims
1. 1. An apparatus comprising a graphics processing unit (GPU), the graphics processing unit (GPU) comprising: Multiple frame buffers; an integrated encoder / decoder coupled to the at least one frame buffer and including embedded logic for encoding and decoding at least one of image data and video content; an integrated network interface controller (NIC) coupled to the integrated encoder / decoder; The GPU is configured to encode video game frame content using the integrated encoder / decoder to generate encoded video game content; configured to packetize the encoded video content using the integrated NIC to generate a stream of packets; The GPU further comprises: configured to encode tiles of video game frame content using the integrated encoder / decoder to generate encoded video game tiles; configured to packetize the encoded video game tiles using the integrated NIC to generate a stream of packets; The integrated encoder / decoder comprises: compression means configured to perform compression on said image data or said video content; filtering means configured to reduce artifacts caused by compression by said compression means.
2. the device includes a graphics card having an input / output (I / O) interface, the integrated NIC including an interface coupled to the I / O interface on the graphics card; The apparatus of claim 1 , wherein the integrated NIC is configured to receive data transmitted from the I / O interface over a network and to forward data received from the network to the I / O interface.
3. 3. The apparatus of claim 1, wherein the integrated NIC includes embedded logic that implements at least layers 3 and 4 of the Open Systems Interconnection (OSI) model.
4. 4. The apparatus of claim 1, wherein the integrated NIC includes embedded logic that implements layers 3 through 7 of the Open Systems Interconnection (OSI) model.
5. 5. The apparatus of claim 1, wherein the integrated encoder / decoder comprises a video codec.
6. The GPU is generating video game frame content and buffering the video game frame content in one or more of the frame buffers; encoding the video game frame content using the video codec to generate encoded video game content; The apparatus of claim 5 , configured to transmit the stream of packets outbound to a network using the integrated NIC.
7. The GPU includes a video output, and the device further includes a graphics card, the graphics card including: a graphics memory coupled to or integrated with the GPU; a network port coupled to the integrated NIC; an input / output (I / O) interface coupled to the GPU; a video port on the GPU coupled to the video output; The graphics card is receiving a stream of packets comprising streaming media content from a network coupled to said network port; configured to use the integrated NIC to depacketize the stream of packets and extract encoded video content; a) writing the depacketized encoded video content to a buffer accessible to said video codec; or b) reading the depacketized encoded video content buffered on the integrated NIC via the video codec; and configured to perform one of the following: The graphics card is Decoding the encoded video content using the video codec to play video game frame content; buffering the reproduced video game frame content in at least one frame buffer; 7. An apparatus according to claim 5 or 6, configured to output display content comprising video game frames via the video port.
8. The apparatus of claim 1 , wherein the integrated encoder / decoder is an image tile encoder and decoder.
9. The GPU is generating video game frame content and buffering the video game frame content in one or more of the frame buffers; encoding tiles of video game frame content using the image tile encoder to generate encoded video game tiles; 10. The apparatus of claim 8, configured to use the integrated NIC to send the stream of packets outbound to a network coupled to an Ethernet port on the graphics card.
10. The GPU includes a video output, and the device further includes a graphics card, the graphics card including: a graphics memory coupled to or integrated with the GPU; a network port coupled to the integrated NIC; an input / output (I / O) interface coupled to the GPU; a video port on the GPU coupled to the video output; The graphics card is receiving a stream of packets comprising streamed video game tiles from a network coupled to an Ethernet port on the graphics card; configured to depacketize the stream of packets using the integrated NIC to extract encoded video game tiles; a) writing the depacketized encoded video game tiles to a buffer accessible to an image tile decoder; or b) reading the depacketized encoded image tiles buffered on the integrated NIC via the image tile decoder; and configured to perform one of the following: The graphics card is using the image tile decoder to decode the encoded image tiles to reproduce game image tiles; Arraying the rendered game image tiles in a frame buffer to generate a video game frame; 10. An apparatus according to claim 8 or 9, configured to output display content comprising video game frames via the video port.
11. 1. A method implemented in a graphics card including a graphics processing unit (GPU) having an integrated network interface controller (NIC), comprising: generating video game frame content and buffering the video game frame content in one or more frame buffers on the GPU; encoding the video game frame content using a video codec integrated into the GPU to generate encoded video content; using the integrated NIC to packetize the encoded video content to generate a stream of packets and transmit the stream of packets outbound to a network operatively coupled to an output port on the integrated NIC; encoding tiles of video game frame content using an image tile encoder integrated into the GPU in addition to the video codec to generate encoded video tiles; using the integrated NIC to packetize the encoded video tiles to generate a stream of packets, and transmitting the stream of packets outbound to a network operatively coupled to an output port on the integrated NIC; encoding the video game frame content to generate encoded video content, compressing the video game frame content with the video codec; and reducing artifacts caused by said compression.
12. moreover, receiving encoded audio content at an input / output (I / O) interface of the graphics card; transferring the encoded audio content via an interconnect on the GPU coupling the I / O interface and the integrated NIC; 12. The method of claim 11, further comprising: using the integrated NIC to packetize the encoded audio content to generate a stream of audio packets; and transmitting the stream of audio packets outbound to the network via an output port of the integrated NIC.
13. moreover, establishing a streaming media session between the graphics card and a client device; and transferring the stream of packets to the client device using the streaming media session using at least one protocol associated with the streaming media session.
14. the graphics card is installed in a host and communicates with the host via an input / output (I / O) interface on the graphics card; and receiving video control input data from the network; detecting that the video control input data is transferred to the host via the integrated NIC; and transferring the video control input data to the host via an interconnect coupled between the integrated NIC and the input / output (I / O) interface of the graphics card.
15. Furthermore, receiving game control data at an input / output (I / O) interface of the graphics card; 15. The method of claim 11, further comprising: transferring the game control data over an interconnect on the GPU that couples the I / O interface and the integrated NIC; and transferring the game control data to a game client device using a reliable transport protocol.
16. a first board having a plurality of expansion slots or connectors; a central processing unit (CPU) mounted on the first board or the second board, the second board being installed in an expansion slot or coupled to a mating connector on the first board; a main memory comprising one or more memory devices communicatively connected to the CPU; a) one or more network adapter cards installed in each expansion slot, or b) one or more network interface controller (NIC) chips mounted on the first board or the second board; a plurality of graphics cards installed in respective expansion slots or coupled to respective mating connectors on the first board; Each graphics card is a graphics processing unit (GPU), the graphics processing unit (GPU) comprising: one or more frame buffers; an integrated encoder / decoder coupled to the at least one frame buffer and including embedded logic for encoding and decoding at least one of image data and video content; an integrated network interface controller (NIC) coupled to the integrated encoder / decoder; a graphics memory coupled to or integrated with the GPU; at least one network port coupled to the integrated NIC; an input / output (I / O) interface coupled to the GPU; Equipped with The GPU in at least one graphics card comprises: configured to encode video game frame content using the integrated encoder / decoder to generate encoded video game content; configured to packetize the encoded video content using the integrated NIC in the GPU to generate a stream of packets; The GPU of at least one graphics card further comprises: configured to encode tiles of video game frame content using the integrated encoder / decoder to generate encoded video game tiles; configured to packetize the encoded video game tiles using the integrated NIC to generate a stream of packets; The integrated encoder / decoder comprises: compression means configured to perform compression on said image data or said video content; and filtering means configured to reduce artifacts caused by compression by said compression means.
17. the integrated encoder / decoder for at least one graphics card includes a video codec; The GPU in at least one of the graphics cards generating video game frame content and buffering the video game frame content in one or more frame buffers on the GPU; encoding the video game frame content using the video codec to generate encoded video game content; 17. The cloud gaming server of claim 16, configured to use the integrated NIC in the GPU to send the stream of packets outbound to a network coupled to a network port on the graphics card.
18. the integrated encoder / decoder is an image tile encoder and decoder; The GPU of at least one graphics card comprises: generating video game frame content and buffering the video game frame content in one or more frame buffers on the GPU; encoding tiles of video game frame content using the image tile encoder to generate encoded video game tiles; 18. The cloud gaming server of claim 16 or 17, configured to use the integrated NIC to send the stream of packets outbound to a network coupled to a network port on the graphics card.
19. The cloud game server may further include a game software residing in at least one of the main memory and the storage device in the cloud game server, and executing the game software may include the cloud game server performing the following steps: establishing a streaming media session with a plurality of gaming client devices coupled to a network using network communications using at least one network adapter card or NIC; 19. The cloud gaming server of claim 16, wherein the cloud gaming server is capable of using multiple streaming media sessions and transferring streams of packets to multiple game client devices using multiple graphics cards.
20. 20. The cloud gaming server of claim 19, wherein the cloud gaming server is further configured to generate audio content for instances of games hosted on the cloud gaming server and stream the audio content to a plurality of the game client devices.
Citation Information
Patent Citations
Game providing server
JP2015195977A
Method, apparatus, and system for controlling power consumption of unused hardware of a link interface
JP2017511529A
Integrated GPU, NIC and Compression Hardware for Hosted Graphics
US20100013839A1
System and Method for Selecting a Video Encoding Format Based on Feedback Data
US20100166062A1
Network-enabled graphics processing module
US20180063555A1