Rack Assembly and Network Architecture
The PCIe fabric with array-level switches and Ethernet connectivity addresses bandwidth limitations in network architectures, enabling high-speed storage access for cloud gaming environments.
Patent Information
- Application Number
- JP2023035545
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-07-31
- Filing Date
- 2023-03-08
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2041-02-12
AI Technical Summary
Existing network architectures struggle to provide sufficient bandwidth for high-speed network storage access in cloud gaming environments, particularly with the increasing demand for gigabit connections.
A network architecture utilizing a PCI Express (PCIe) fabric with array-level PCIe switches and an Ethernet fabric to facilitate direct access to network storage from computational nodes, enabling bandwidth of over 4 gigabytes per second per compute node.
This solution provides high-speed network storage access, supporting computationally intensive applications like cloud gaming with reduced latency and improved performance.
Smart Images

Figure 0007724046000001 
Figure 0007724046000002 
Figure 0007724046000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to network storage, and more particularly to high-speed network storage access to compute nodes located on compute sleds of a streaming array of rack assemblies using PCI-Express. [Background technology]
[0002] In recent years, there has been a continuous push for online services that enable streaming online or cloud gaming between cloud gaming servers and clients connected via a network. The streaming format has become increasingly popular due to the availability of on-demand game titles, the ability to run more complex games, the ability to network among players for multiplayer games, the ability to share assets or properties among players, the ability to share instant experiences among players and / or spectators, the ability for friends to watch video games being played by friends, and the ability for friends to join other friends in gameplay while a friend is playing.
[0003] Unfortunately, demand is pushing the limits of network connection capabilities. For example, previous generation streaming network architectures provided network storage using Gigabit Ethernet communications connections (e.g., 40 Gigabit per second Ethernet connections). However, new generation streaming network architectures require better (faster) bandwidth performance (e.g., gigabit connections).
[0004] It is against this background that the embodiments of the present disclosure have been made. Summary of the Invention
[0005] Embodiments of the present disclosure relate to providing high-speed access to network storage, such as in a rack assembly, capable of providing network storage bandwidth (eg, access) of over 4 gigabytes per second (GB / s) per compute node.
[0006] An embodiment of the present disclosure discloses a network architecture. The network architecture includes network storage. The network architecture includes a plurality of streaming arrays, each including a plurality of computational threads, each including one or more computational nodes. The network architecture includes a PCI Express (PCIe) fabric configured to provide direct access to the network storage from each computational node of the plurality of streaming arrays. The PCIe fabric includes a plurality of array-level PCIe switches, each array-level PCIe switch communicatively coupled to the computational nodes of the computational threads of a corresponding streaming array and to a storage server. The network storage is shared by the plurality of streaming arrays.
[0007] An embodiment of the present disclosure discloses a network architecture. The network architecture includes network storage. The network architecture includes multiple streaming arrays, each including multiple computational threads, each including one or more computational nodes. The network architecture includes a PCI Express (PCIe) fabric configured to provide direct access to the network storage from each computational node of the multiple streaming arrays. The PCIe fabric includes multiple array-level PCIe switches, each array-level PCIe switch communicatively coupled to the computational nodes of the computational threads of a corresponding streaming array and to a storage server. The network architecture includes an Ethernet fabric configured to communicatively couple the computational nodes of the computational threads of the multiple streaming arrays to the network storage for streaming computational thread and computational node management information. The network storage is shared by the multiple streaming arrays.
[0008] Other aspects of the present disclosure will become apparent from the following detailed description, taken in conjunction with the accompanying drawings, illustrated by way of example of the principles of the disclosure.
[0009] The present disclosure is best understood by reference to the following detailed description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram of a game cloud system for providing games over a network among one or more computing nodes located in one or more data centers, according to one embodiment of the present disclosure. [Figure 2] FIG. 1 is a diagram of multiple rack assemblies including multiple computing nodes in a representative data center of a gaming cloud system, according to one embodiment of the present disclosure. [Figure 3]1 is a diagram of a rack assembly configured to provide compute nodes with high-speed access to network storage using PCIe communication, according to one embodiment of the present disclosure. [Figure 4] FIG. 1 is a diagram of a streaming array including multiple compute nodes arranged in a rack assembly configured to provide the compute nodes with high-speed access to network storage using PCIe communication, according to one embodiment of the present disclosure. [Figure 5] FIG. 1 is a diagram of a compute sled including multiple compute nodes arranged in a rack assembly configured to provide the compute nodes with high-speed access to network storage using PCIe communication, according to one embodiment of the present disclosure. [Figure 6] FIG. 1 is a diagram of a sled-level PCIe switch disposed within a rack assembly configured to provide compute nodes with high-speed access to network storage using PCIe communication, according to one embodiment of the present disclosure. [Figure 7] 1 illustrates components of an exemplary device that can be used to implement aspects of various embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0011] Although the following detailed description includes many specific details for purposes of illustration, those skilled in the art will appreciate that many variations and modifications to the following details are within the scope of the present disclosure. Accordingly, the aspects of the disclosure described below are set forth without loss of generality to, and without imposing limitations on, the claims that follow this description.
[0012] Generally speaking, embodiments of the present disclosure provide high-speed access to network storage, such as within a rack assembly, capable of providing network storage bandwidth (e.g., access) of greater than 4 gigabytes per second (GB / s) per compute node (e.g., of a rack assembly) at Non-Volatile Memory express (NVMe) latency.
[0013] With the above general understanding of the various embodiments, details of example embodiments will now be described with reference to the various drawing figures.
[0014] Throughout this specification, references to an "application" or a "game" or a "video game" or a "game application" are meant to refer to any type of interactive application that is directed through the execution of input commands. For purposes of explanation only, interactive applications include applications for games, word processing, video processing, video game processing, etc. Furthermore, these terms are interchangeable.
[0015] 1 is a diagram of a system 100 for providing games over a network 150 between one or more computing nodes located in one or more data centers, according to one embodiment of the present disclosure. According to one embodiment of the present disclosure, the system is configured to provide games over a network between one or more cloud gaming servers, and more particularly, to provide high-speed access from the computing nodes to network storage, such as in a rack assembly. Cloud gaming involves running a video game on a server to generate game-rendered video frames, which are then transmitted to clients for display.
[0016] It is also understood that cloud gaming can be executed in various embodiments (e.g., within a cloud gaming environment or a standalone system) using physical machines (e.g., central processing units—CPUs—and graphics processing units—GPUs), or virtual machines, or a combination of both. For example, virtual machines (e.g., instances) can be created using a hypervisor on host hardware (e.g., located in a data center) that utilizes one or more components of a hardware layer, such as multiple CPUs, memory modules, GPUs, network interfaces, communication components, etc. These physical resources can be arranged in racks, such as a rack of CPUs, a rack of GPUs, a rack of memory, etc., and the physical resources in the racks can be accessed using top of rack switches that facilitate the assembly and access fabric of components used for the instances (e.g., when building the virtualized components of the instances). Typically, a hypervisor can present multiple guest operating systems in multiple instances configured with virtual resources. That is, each of the operating systems can be configured with a corresponding set of virtualized resources supported by one or more hardware resources (e.g., located in a corresponding data center). For example, each operating system can be supported by a virtual CPU, multiple virtual GPUs, virtual memory, virtualized communication components, etc. Furthermore, the configuration of an instance can be transferred from one data center to another to reduce latency. Instant usage defined for a user or game can be used when saving a user's game session. The instantaneous usage may include any number of configurations described herein to optimize fast rendering of video frames for a game session. In one embodiment, the instantaneous usage defined for a game or user may be transferred between data centers as configurable settings. The ability to transfer instantaneous usage allows for efficient migration of game play from data center to data center when users connect to play games from different geographic locations.
[0017] System 100 includes a gaming cloud system 190 implemented across one or more data centers (e.g., data centers 1 through N). As shown, an instance of gaming cloud system 190 may be located in data center N providing management functionality, and the management functionality of gaming cloud system 190 may be distributed across multiple instances of gaming cloud system 190 at each data center. In some implementations, gaming cloud system management functionality may be located outside of any of the data centers.
[0018] The gaming cloud system 190 includes an assigner 191 configured to assign each of the client devices (e.g., 1-N) to corresponding resources in a corresponding data center. In particular, when the client device 110 logs into the gaming cloud system 190, the client device 110 may connect to an instance of the gaming cloud system 109 at data center N, which may be geographically closest to the client device 110. The assigner 191 may perform diagnostic tests to determine available transmit and receive bandwidth to the client device 110. Based on the tests, the assigner 191 may assign resources to the client device 110 in a very specific or undifferentiated manner. For example, the assigner 191 may assign a particular data center to the client device 110. Furthermore, the assigner 191 may assign a particular compute thread, a particular streaming array, or a particular compute node in a particular rack assembly to the client device 110. The allocation is performed based on knowledge of assets (e.g., games) available on the compute nodes. Previously, client devices were typically assigned to data centers and not further assigned to rack assemblies. In this manner, the assigner 191 can assign client devices requesting execution of a particular compute-intensive game application to compute nodes that may not be running the compute-intensive application. Additionally, the assigner 191 can perform load management of the allocation of compute-intensive game applications in response to requests from clients. For example, the same compute-intensive game application requested for a short period of time may be distributed across different compute nodes in different compute threads within a rack assembly or different rack assemblies to reduce the load on a particular compute node, compute thread, and / or rack assembly.
[0019] In some embodiments, the allocation may be performed based on machine learning. In particular, resource demand may be predicted for a particular data center and its corresponding resources. For example, if a data center can predict that it will soon handle many clients running computationally intensive gaming applications, assigner 191 can use that knowledge to assign client device 110 resources that may not currently be utilizing all of its resource capacity. In another case, assigner 191 may switch client device 110 from gaming cloud system 190 in data center N to resources available in data center 3 in anticipation of increased load at data center N. Further, prospective clients may be allocated resources in a distributed manner such that resource load and demand is distributed across the gaming cloud system, across multiple data centers, across multiple rack assemblies, across multiple computational threads, and / or across multiple computational nodes. For example, client device 110 may be allocated resources from both data center N (e.g., via path 1) and data center 3 (e.g., via path 2).
[0020] Once a client device 110 is assigned to a particular computational node of a corresponding computational thread of a corresponding streaming array, the client device 110 connects to the corresponding data center over a network, i.e., the client device 110 may be in communication with a data center different from the data center that performs the assignment, such as data center 3.
[0021] In particular, system 100 provides games via game cloud system 190, which, according to one embodiment of the present disclosure, are executed remotely from the client devices (e.g., thin clients) of corresponding users playing the games. System 100 can provide game control to one or more users playing one or more games via cloud gaming network or game cloud system 190 over network 150, in either single-player or multiplayer mode. In some embodiments, cloud gaming network or game cloud system 190 can include multiple virtual machines (VMs) executing on a host machine's hypervisor, with one or more virtual machines configured to execute game processor modules that utilize hardware resources available to the host's hypervisor. Network 150 can include one or more communication technologies. In some embodiments, network 150 can include fifth-generation (5G) network technology with advanced wireless communication systems.
[0022] In some embodiments, communication may be facilitated using wireless technology. Such technology may include, for example, 5G wireless communication technology. 5G is the fifth generation of cellular network technology. 5G networks are digital cellular networks in which service areas covered by providers are divided into small geographic areas called cells. Analog signals representing sound and images are digitized by the phone, converted by an analog-to-digital converter, and transmitted as a stream of bits. All 5G wireless devices within a cell communicate over the airwaves with a local antenna array and low-power automatic transceivers (transmitters and receivers) within the cell via frequency channels assigned by the transceiver from a pool of frequencies reused in other cells. The local antennas are connected to the telephone network and the Internet by high-bandwidth optical fiber or wireless backhaul connections. As with other cell networks, mobile devices moving from one cell to another are automatically transferred to the new cell. It should be understood that a 5G network is just one example type of communication network, and embodiments of the present disclosure may utilize previous generations of wireless or wired communications, as well as later generations of wired or wireless technologies following 5G.
[0023] As shown, system 100, including game cloud system 190, can provide access to multiple video games. In particular, each of the client devices may request access to a different game from the cloud gaming network. For example, game cloud system 190 may provide one or more game servers, which may be configured as one or more virtual machines running on one or more hosts to execute corresponding game applications. For example, a game server may manage virtual machines supporting game processors that instantiate instances of users' games. Thus, multiple game processors of one or more game servers associated with multiple virtual machines are configured to execute multiple instances of one or more games associated with gameplay of multiple users. In this manner, the backend server support provides streaming of gameplay media (e.g., video, audio, etc.) for multiple game applications to a corresponding number of users. That is, the game servers of the game cloud system 190 are configured to stream data (e.g., rendered images and / or frames of corresponding gameplay) back to corresponding client devices over the network 150. In this manner, computationally complex game applications can continue to run on the backend servers in response to controller inputs received and forwarded by the client devices. Each server can render images and / or frames, then encode (e.g., compress) them and stream them to a corresponding client device for display.
[0024] In one embodiment, the cloud gaming network or game cloud system 190 is a distributed game server system and / or architecture. Specifically, a distributed game engine that executes game logic is configured for each instance of a corresponding game. Generally, a distributed game engine takes each function of the game engine and distributes those functions to be performed by multiple processing entities. Individual functions may be further distributed across one or more processing entities. The processing entities may be configured in various configurations, including including physical hardware and / or as virtual components or virtual machines and / or as virtual containers, which differ from virtual machines because a container is a virtualized instance of a gaming application running on a virtualized operating system. The processing entities may utilize and / or rely on servers and their underlying hardware on one or more servers (computing nodes) of the cloud gaming network or gaming cloud system 190, which may be arranged on one or more racks. The coordination, allocation, and management of the execution of these functions across the various processing entities is performed by a distributed synchronization layer, which controls the execution of these functions to generate media (e.g., video frames, audio, etc.) for the game application in response to controller inputs by the player. The distributed synchronization layer enables critical game engine components / functions to be efficiently executed across the distributed processing entities (e.g., via load balancing) so that these functions can be distributed and restructured for more efficient processing.
[0025] 2 is a diagram of multiple rack assemblies 210 containing multiple computing nodes in a representative data center 200 of a gaming cloud system, according to one embodiment of the disclosure. For example, multiple data centers may be distributed around the world, such as in North America, Europe, and Japan.
[0026] Data center 200 includes multiple rack assemblies 220 (e.g., rack assemblies 220A through 220N). Each rack assembly includes corresponding network storage and multiple compute sleds. For example, representative rack assembly 220N includes network storage 210 and multiple compute sleds 230 (e.g., sleds 230A through 230N). Other rack assemblies may be similarly configured, with or without modifications. In particular, each of the computational threads includes one or more computational nodes that provide hardware resources (e.g., processors, CPUs, GPUs, etc.). For example, computational thread 230N in the plurality of computational threads 230 of rack assembly 220N is shown to include four computational nodes, although it is understood that a rack assembly may include one or more computational nodes. Each rack assembly is coupled to a cluster switch configured to provide communication with a management server configured for management of the corresponding data center. For example, rack assembly 220N is coupled to cluster switch 240N. The cluster switch also provides communication to an external communication network (e.g., the Internet).
[0027] Each rack assembly provides high-speed access to corresponding network storage, such as within the rack assembly. This high-speed access is provided via a PCI-express fabric that provides direct access between the compute nodes and the corresponding network storage. For example, in rack assembly 220N, the high-speed access is configured to provide data path 201 between a particular compute node of a corresponding compute thread and the corresponding network storage (e.g., storage 210). In particular, the PCIe fabric can provide network storage bandwidth (e.g., access) of over 4 gigabytes per second (GB / s) per compute node (e.g., rack assembly) at non-volatile memory express (NVMe) latency. Additionally, control path 202 is configured to communicate control and / or management information between network storage 210 and each compute node.
[0028] As illustrated, management server 210 of data center 200 communicates with assigner 191 (shown in FIG. 1 ) to allocate resources to client devices 110. In particular, management server 210 may coordinate with an instance of gaming cloud system 190′, and coordinate with the initial instance of gaming cloud system 190 (e.g., of FIG. 1 ), to allocate resources to client devices 110. In embodiments, the allocation is performed based on asset awareness, such as knowing what resources and bandwidth are needed and present in the data center. Thus, for illustrative purposes, embodiments of the present disclosure are configured to assign client devices 110 to particular compute nodes 232 of corresponding compute threads 231 of corresponding rack assembly 220B.
[0029] The streaming rack assembly is centered around compute nodes that run game applications, video games, and / or stream the audio / video of game sessions to one or more clients. Additionally, within each rack assembly, game content can be stored on storage servers that provide network storage. The network storage is equipped with large amounts of storage and a high-speed network to serve many compute nodes via Network File System (NFS)-based network storage.
[0030] 3 is a diagram of a rack assembly 300 configured to provide compute nodes with high-speed access to network storage using PCIe communications, according to one embodiment of the present disclosure. As shown, the diagram of FIG. 3 shows a high-level rack design of rack assembly 300. Rack assembly 300 may represent one or more of multiple rack assemblies 220. For example, rack assembly 300 may represent rack assembly 220N.
[0031] As mentioned above, traditional rack designs provide access to network storage using Gigabit Ethernet, which provides 40gb / s access to network storage, which is not suitable for the future of gaming.
[0032] Embodiments of the present disclosure provide access to network storage at over approximately 4 gigabytes per second (GB / s) bandwidth per compute node at NVMe-level latency, which in one embodiment is achieved through PCI Express switching technology and a rack-wide PCI Express fabric.
[0033] Each rack assembly 300 includes a network storage 310. Game content is stored or saved in the network storage 310 in each rack assembly. The network storage 310 is equipped with a large amount of storage and a high-speed network to serve many computing nodes via NFS-based network storage.
[0034] Additionally, each rack assembly 300 includes one or more streaming arrays. While rack assembly 300 is shown as having four arrays, it is understood that one or more streaming arrays may be included within rack assembly 300. More specifically, each streaming array includes a network switch, an array management server (AMS), and one or more computational threads. For example, exemplary streaming array 4 includes network switch 341, AMS 343, and one or more computational threads 345. The other streaming arrays 1-3 may be similarly configured. For illustrative purposes, the streaming arrays shown in FIG. 3 include eight computational threads per streaming array, but it is understood that a streaming array may include any number of computational threads, such that each computational thread includes one or more computational nodes.
[0035] Specifically, each streaming array is served by a corresponding PCIe switch configured as part of a PCIe fabric (e.g., Gen4) providing direct access between the compute nodes and the storage servers via the PCIe fabric. For example, representative streaming array 4 is served by PCIe switch 347. The PCIe fabric (i.e., including the PCIe switches serving each of streaming arrays 1-4) provides data path 301 (e.g., data path 201 in rack assembly 220N) that enables high-speed access to game data stored in the aforementioned network storage 310.
[0036] Additionally, each streaming array is configured with an Ethernet fabric that provides a control path 302 (eg, control path 202 within rack assembly 220N), such as for communicating control and / or management information to the streaming array.
[0037] The rack assembly 300 is also configured with shared power managed by a rack management controller (not shown). Additionally, the rack assemblies may also be configured with shared cooling (not shown).
[0038] Rack assembly 300 is designed with the requirement to provide each compute node with high-speed storage access (e.g., up to 4-5 GB / s or more). Storage is provided by network storage 310, which stores game content in RAM and NVMe drives (i.e., not a traditional bunch of disks—JBOD—storage server). In one embodiment, game content is “read-only” so it can be shared between systems. Individual compute nodes access the game content on network storage 310 via a PCIe fabric (e.g., providing data path 301) between each of the streaming arrays and network storage 310.
[0039] In particular, the PCIe fabric (e.g., Gen4) can assume that not all compute nodes simultaneously require peak performance (4-5 GB / s). Each thread has multiple (e.g., 8) lanes of PCIe (e.g., up to 16 GB / s). For example, a total of 64 lanes (for 8 threads) per streaming array can be provided to the corresponding PCIe switch, which can be configured with a multi-lane (e.g., 96-lane) PCIe switch. However, each PCIe switch can provide only 32 lanes of the corresponding array to the network storage 310, depending on the design.
[0040] Additionally, each rack assembly 300 includes a second PCIe fabric available between an array management server (AMS) and the corresponding compute threads. For example, array 4 includes a second PCIe fabric 349 that provides communication between the AMS 343 and one or more compute threads 345. This fabric has lower performance (e.g., one lane of PCIe per thread) and can be used for slower storage workloads or thread management.
[0041] Additionally, each rack assembly 300 includes a conventional Ethernet network, providing communication, for example, for control path 302. For example, each compute node has 1 x 1 Gbps Ethernet (e.g., 32 x 1 Gbps for 32 compute nodes between the compute node and the corresponding network switch), used for "audio / video streaming" and management. The AMS and network storage have faster networking for network storage and management (e.g., 40 Gbps between the corresponding AMS and the network switch, 10 Gbps between the network storage 310 and the corresponding network switch, and 100 Gbps between the network storage 310 and the cluster switch 350).
[0042] Network storage 310 (e.g., a server) may also be configured to provide network storage access to the AMS servers and compute nodes. Network storage access to the AMS servers is handled via conventional Ethernet networking (e.g., 10 Gbps between a corresponding network switch and network storage 310). However, network storage to the compute nodes is done over PCI Express (i.e., via data path 301) with a custom protocol and custom storage solution. The background to this custom storage solution is the hardware and software design of the compute nodes that utilizes PCIe switching.
[0043] In one embodiment, each compute node can request data from a location using a "command buffer" based protocol. The network storage 310 is expected to place the data. In particular, the compute node uses a direct memory access (DMA) engine to move it to its own memory during a "read operation." Data stored on the network storage 310 is stored in RAM and NVMe. Software on the network storage 310 ensures that data is cached in RAM whenever possible to avoid the need to retrieve data from NVMe. Caching is possible because it is expected that many compute nodes will access the same content.
[0044] FIG. 4 is a diagram of a streaming array 400 including multiple compute nodes arranged in a rack assembly configured to provide the compute nodes with high-speed access to network storage 410 using PCIe communications, according to one embodiment of the present disclosure. Rack assemblies configured to stream content to one or more users are divided into "streaming arrays," such as streaming arrays 1-4 in FIG. 3, that access network storage 310. In particular, the arrays are part of a rack assembly (e.g., rack assembly 300 in FIG. 3) that, as previously described, consists of a network switch, an array management server (AMS), and multiple compute threads (e.g., one or more compute threads per array, each holding one or more compute nodes). Multiple arrays 400 are configured within the rack assembly, sharing network storage but otherwise operating independently.
[0045] As shown, the Array Management Server (AMS) 403 is a server within the corresponding streaming array 400 that is responsible for managing all operations within the streaming array. It handles two broad classes of operations. First, the AMS 403 manages "configuration work," which is concerned with making sure each compute thread (e.g., threads 1-8) is functioning properly. This includes powering the threads, making sure software is up to date, configuring the network, configuring PCIe switches, etc.
[0046] A second class of AMS 403 operations is managing a cloud gaming session, which includes setting up a cloud gaming session on a corresponding compute node, providing network / internet access to one or more compute nodes, providing storage access, and monitoring the cloud gaming session.
[0047] Thus, AMS 403 is configured to manage compute nodes and compute threads, with each compute thread including one or more compute nodes. For example, AMS 403 enables power delivery to the compute nodes using general-purpose input / output (GPIO) signals to a power interposer. In one embodiment, AMS 403 is configured to control and monitor the compute nodes using universal asynchronous receiver-transmitter (UART) signals that deliver serial data (e.g., power on / off, diagnostic, and logging information). AMS 403 is configured to perform firmware updates on the compute nodes. AMS 403 is configured to perform configuration of the compute threads and corresponding PCIe switches 407.
[0048] The streaming array 400 is configured to provide storage to the compute nodes via PCI Express, as previously described. For example, a PCIe fabric provides a data path 402 between the compute nodes on the compute threads and a PCIe switch 407. In an embodiment, read-write storage access per compute node is provided at up to 500 megabytes per second (MB / s). Furthermore, in one embodiment, there is 1-2 gigabytes (GB) of storage per compute node, although other sizes of storage are supported.
[0049] Additionally, each streaming array 400 provides network / internet access to the compute nodes, as previously described. For example, network access (e.g., via network switch 411 and via paths not shown, such as Ethernet) is provided at 100 megabits per second (mb / s) per compute node.
[0050] 4, a primary function of AMS 403 is a PCI Express fabric connection to each of the compute threads. For example, PCIe fabric 420 is shown providing communication between the compute nodes on the compute threads and AMS 403. In one embodiment, the PCI Express fabric connection is implemented using a "passive PCI Express adapter" because each compute thread can be configured with a PCI Express Gen4 switch and the distance between the AMS and the compute threads must be short.
[0051] The AMS403 can consist of a central processing unit (CPU) with random access memory (RAM), may have input / output (I / O) for PCIe fabric, and has network connectivity for Ethernet.
[0052] The AMS 403 may be configured with storage (e.g., 2 x 2 terabytes of NVMe). Additionally, there may be a PCIe fabric connection to each compute sled, such as using a passive PCIe fabric adapter. There may also be a bus bar providing power (e.g., 12 volts).
[0053] 5 is a diagram of a compute sled 500 including multiple compute nodes (e.g., nodes 1-4) arranged in a rack assembly configured to provide the compute nodes with high-speed access to network storage using PCIe (e.g., Gen4) communication, according to one embodiment of the present disclosure. FIG. 5 illustrates the multiple compute nodes (e.g., nodes 1-4) and supporting hardware to support the operation of the compute nodes.
[0054] Each computational thread 500 includes one or more computational nodes. While Figure 5 illustrates a computational thread including four computational nodes (e.g., nodes 1-4), it is understood that any number of computational nodes can be provided for a computational thread including one or more computational nodes. Computational thread 500 can provide a hardware platform (e.g., a circuit board) that provides computational resources (e.g., via the computational nodes).
[0055] Compute sled 500 includes an Ethernet patch panel 510 configured to connect Ethernet cables between compute nodes (eg, nodes 1-4) and a rack-level network switch (not shown), as previously described.
[0056] The compute sled 500 includes a PCIe switch board 520 .
[0057] The computational thread 500 includes a management panel 530. For example, the management panel 530 can provide status through LEDs, buttons, and the like.
[0058] Compute sled 500 includes a power interposer board 540 configured to provide power to the compute sled.
[0059] Each computing thread includes one or more computing nodes (e.g., nodes 1-4). Each computing node disposed within the rack assembly is configured, in accordance with one embodiment of the present disclosure, to provide the computing node with high-speed access to network storage (not shown) using PCIe communication (e.g., Gen4). The computing node includes multiple I / O interfaces. For example, the computing node may include an M.2 port and multiple lanes for PCIe Gen4 (bidirectional).
[0060] A PCIe (e.g., Gen4) interface (e.g., 4 lanes) can be used to expand the system with additional devices. In particular, the PCIe interface is used to connect to a PCIe fabric, including a PCI Express switch 520 for high-speed storage. Additionally, the compute node includes an Ethernet connection (e.g., Gigabit Ethernet). The compute node also includes one or more universal asynchronous receiver-transmitter (UART) connections configured to transmit and / or receive serial data. For example, there may be one or more UART ports, intended for management purposes (e.g., connecting the compute node to a UART / GPIO controller 550). Ports can be used for remote control operations such as "power on," "power off," and diagnostics. Another UART port provides serial console functionality.
[0061] Each compute node also includes a power input connector (eg, 12 volts for designed power consumption) connected to a power interposer 540 .
[0062] FIG. 6 is a diagram of a sled-level PCIe switch 600 disposed within a rack assembly configured to provide compute nodes with high-speed access to network storage using PCIe communications, according to one embodiment of the present disclosure.
[0063] The sled PCIe switch 600 can be configured as a circuit board within the corresponding compute thread that has two roles. In one embodiment, first, the sled-level PCIe switch 600 has a "fabric role" that connects individual compute nodes (e.g., four compute nodes) to the AMS and corresponding network storage via PCIe (e.g., Gen4) bus 620 by "non-transparent bridging" (NTB). Second, the sled-level PCIe switch 600 has a "management role" in which UART and GPIO signals are provided for thread management.
[0064] In particular, PCIe (e.g., Gen4) connections are provided by external cable connectors, internal cable connectors, and PCIe edge connectors. For example, an eight-lane PCIe (e.g., Gen4) external cable connection 620 can be used to connect the compute threads to network storage for storage workloads. A second external PCIe (e.g., Gen4) connection 625 to a second PCIe fabric connects to an AMS. For example, the second PCIe connection can include one lane because it is primarily used for management functions and includes auxiliary storage functionality.
[0065] Additionally, an internal PCIe (e.g., Gen4) cable connector 610 can be used to connect the sled PCIe switch 520 to each of the compute nodes using a cable via a corresponding M.2 interface. Other connection means may be implemented. For example, instead of using an M.2 connection interface, other connectors and / or connector interfaces such as OCuLink, Slimline SAS, etc. can be used.
[0066] A management interface in the form of UART and GPIO controller 550 is used by the AMS (not shown) to communicate with and manage power for the individual compute nodes. The AMS uses multiple (e.g., two) UART interfaces per compute node for management purposes (power on / off, diagnostics, logging, etc.). The GPIO functionality is used to manage power delivery to each compute node through the power interposer board via connection 630. This also connects to a management panel (e.g., for LEDs and buttons) via connection 630, as previously described.
[0067] The thread-level PCIe switch 600 may include a PCIe (e.g., Gen4) switch 520, multiple (e.g., four) non-transparent (NT) bridging interfaces, and multiple (e.g., four) DMA (direct memory access) engines.
[0068] Additionally, a UART / GPIO controller 550 is configured and includes a PCIe interface to the PCIe switch, multiple (e.g., eight) UART channels 640, and multiple (eight) GPIO connections to the power interposer and management panel.
[0069] Additionally, there is a connector to the PCIe fabric for network storage access. For example, in one embodiment, an 8-lane external PCIe connector 620 is provided from the PCIe fabric to the network storage.
[0070] As previously mentioned, a one-lane external PCIe connector 625 to a second PCIe fabric that provides access to the AMS is also provided within the sled-level PCIe switch board 600. One or more PCIe edge connectors may also be provided.
[0071] Additionally, four multi-lane (e.g., four-lane) internal PCIe connections 610 to the compute nodes may be provided, e.g., four lanes for each compute node.
[0072] A GPIO connector 630 to the power interposer may be included. For example, four signals are needed, one for each compute node.
[0073] There may be four dual / pair UART connectors to the management panel. For example, in one embodiment, each compute node has two UART interfaces. In other embodiments, each compute node may have fewer than two UART interfaces or more than two UART interfaces.
[0074] A power interposer may be included that provides power to the sled via connection 630. A compute sled may include multiple compute nodes arranged in a rack assembly configured to provide the compute nodes with high-speed access to network storage using PCIe communication, according to one embodiment of the present disclosure. In one embodiment, the power interposer provides power to the compute sled from the rack's 12V bus bar. In other embodiments, other voltages, such as 48 volts, are used to power the rack components. For example, a higher voltage (e.g., 48 volts) may be used for power efficiency. For components requiring a specific voltage (e.g., 12 volts), the power interposer can be used to convert the power. For example, the power interposer may include conversion logic (e.g., a DC-DC converter) to convert 48 volts (or other voltages) down to 12 volts, which is used to power the compute nodes as well as supporting hardware. Power supply to the compute nodes can be controlled by GPIOs via the sled PCIe switches. Each compute node may have a dedicated signal to enable / disable power.
[0075] Additionally, a rack management control interface is provided to a rack management controller (RMC) to monitor the power interposer board, thereby providing diagnostic information such as voltage, current, temperature, etc. The rack management control interface may include voltage and / or current information, and temperature.
[0076] Power status information is delivered to the management panel using GPIO signals. This includes the power status of each compute node as well as the 12V status of the power interposer. Additionally, a bus (e.g., 12 volt) bar interface is provided.
[0077] For example, there may be hot-plug support for adding and / or removing compute threads while the power bus is powered. For example, power may be supplied at 12 volts or other levels. The voltage to auxiliary components may be lower (e.g., less than 6 volts), which can be generated from the 12 volts on the power bus.
[0078] The management panel may include a board / panel located in front of the compute sled and indicates the sled's status via LEDs. Each compute node may have two LEDs that provide control status information. The first is powered from the sled PCIe switch using a software-controllable GPIO signal. The second LED is from the power interposer board and indicates power status (e.g., voltage level). The global power status from the power interposer board indicates the overall power status of the sled.
[0079] FIG. 7 illustrates components of an exemplary device 700 that can be used to implement aspects of various embodiments of the present disclosure. For example, FIG. 7 illustrates an exemplary hardware system suitable for providing high-speed access to network storage to computational nodes of corresponding computational threads arranged in corresponding streaming arrays, such as in a rack assembly, according to embodiments of the present disclosure. The block diagram illustrates device 700, which may incorporate or be a personal computer, server computer, game console, mobile device, or other digital device, each suitable for implementing embodiments of the present invention. Device 700 includes a central processing unit (CPU) 702 for executing software applications and optionally an operating system. CPU 702 may be comprised of one or more homogeneous or heterogeneous processing cores.
[0080] According to various embodiments, CPU 702 is one or more general-purpose microprocessors having one or more processing cores. Further embodiments may be implemented using one or more CPUs with a microprocessor architecture specifically adapted for highly parallel and computationally intensive applications, such as applications configured for graphics processing during game execution, media and interactive entertainment applications, etc.
[0081] Memory 704 stores applications and data used by CPU 702 and GPU 716. Storage 706 provides non-volatile storage and other computer-readable media for applications and data and may include fixed disk drives, removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other optical storage devices, as well as signal transmission and storage media. User input device 708 communicates user input from one or more users to device 700, examples of which may include a keyboard, mouse, joystick, touchpad, touchscreen, still or video recorder / camera, and / or microphone. Network interface 709 enables device 700 to communicate with other computer systems over electronic communications networks, which may include wired or wireless communications over local area networks and wide area networks such as the Internet. The audio processor 712 is adapted to generate analog or digital audio output from instructions and / or data provided by the CPU 702, memory 704, and / or storage 706. The components of the device 700, including the CPU 702, graphics subsystem including the GPU 716, memory 704, data storage 706, user input devices 708, network interface 709, and audio processor 712, are connected via one or more data buses 722.
[0082] Graphics subsystem 714 is further connected to data bus 722 and the components of device 700. Graphics subsystem 714 includes at least one graphics processing unit (GPU) 716 and graphics memory 718. Graphics memory 718 includes display memory (e.g., a frame buffer) used to store pixel data for each pixel of an output image. Graphics memory 718 may be integrated into the same device as GPU 716, connected as a separate device from GPU 716, and / or implemented within memory 704. Pixel data may be provided directly from CPU 702 to graphics memory 718. Alternatively, CPU 702 provides data and / or instructions defining desired output images to GPU 716, which generates pixel data for one or more output images therefrom. The data and / or instructions defining desired output images may be stored in memory 704 and / or graphics memory 718. In one embodiment, GPU 716 includes 3D rendering capabilities to generate pixel data for output images from instructions and data defining geometry, lighting, shading, texturing, motion, and / or camera parameters for a scene. GPU 716 may further include one or more programmable execution units capable of executing shader programs.
[0083] Graphics subsystem 714 periodically outputs image pixel data from graphics memory 718 to be displayed on display device 710 or projected by a projection system (not shown). Display device 710 may be any device capable of displaying visual information in response to signals from device 700, including CRT, LCD, plasma, and OLED displays. Device 700 may provide analog or digital signals to display device 710, for example.
[0084] In other embodiments, graphics subsystem 714 includes multiple GPU devices combined to perform graphics processing for a single application running on a corresponding CPU. For example, multiple GPUs can perform multi-GPU rendering of an application's geometry by pre-testing the geometry against potentially interleaved screen regions before rendering objects in an image frame. In other examples, multiple GPUs can perform an alternative form of frame rendering, where, over successive frame periods, GPU1 renders the first frame, GPU2 renders the second frame, and so on, until the last GPU is reached, and the first GPU renders the next video frame (e.g., if there are only two GPUs, GPU1 renders the third frame). That is, the GPUs are cycled when rendering frames. Rendering operations may overlap, where GPU2 can begin rendering a second frame before GPU1 has finished rendering the first frame. In another embodiment, multiple GPU devices can be assigned different shader operations in the rendering and / or graphics pipeline, with a master GPU performing the main rendering and compositing. For example, in a group including three GPUs, master GPU1 may perform main rendering (e.g., a first shader operation) and compositing the output from slave GPU2 and slave GPU3, slave GPU2 may perform a second shader operation (e.g., a fluid effect such as a river), slave GPU3 may perform a third shader operation (e.g., particle smoke), and master GPU1 may composite the results from each of GPU1, GPU2, and GPU3. In this manner, different GPUs may be assigned to perform different shader operations (e.g., flag waving, wind, smoke generation, fire, etc.) to render a video frame. In yet another embodiment, each of the three GPUs may be assigned to a different object and / or portion of a scene corresponding to a video frame. In the above embodiments and implementations, these operations may be performed in the same frame cycle (concurrently in parallel) or in different frame cycles (sequentially in parallel).
[0085] Accordingly, the present disclosure describes methods and systems configured to provide high-speed access to network storage to computing nodes of corresponding computing threads configured in corresponding streaming arrays, such as in rack assemblies.
[0086] It should be understood that the various embodiments defined herein can be combined or assembled into specific implementations that use various features disclosed herein. Thus, the examples provided are only some of the possible examples and are not intended to limit the various implementations that may be defined by combining various elements. In some instances, an implementation may include fewer elements without departing from the spirit of the disclosed or equivalent implementations.
[0087] Embodiments of the present disclosure may be practiced with a variety of computer system configurations, including handheld devices, microprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc. Embodiments of the present disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a wire-based or wireless network.
[0088] With the above embodiments in mind, it should be understood that embodiments of the present disclosure may employ various computer-implemented operations involving data stored in computer systems. These operations are operations requiring physical manipulation of physical quantities. Any of the operations described herein that form part of embodiments of the present disclosure are useful machine operations. Embodiments of the disclosure also relate to devices or apparatus for performing these operations. An apparatus may be specially constructed for the required purposes. Alternatively, the apparatus may be a general-purpose computer selectively activated or configured by a computer program stored in the computer. In particular, various general-purpose machines may be used with computer programs written in accordance with the teachings herein, or it may be more convenient to construct a more specialized apparatus to perform the required operations.
[0089] The present disclosure can also be embodied as computer-readable code on a computer-readable medium. A computer-readable medium is any data storage device that can store data which can thereafter be read by a computer system. Examples of computer-readable media include hard drives, network-attached storage (NAS), read-only memory, random-access memory, CD-ROMs, CD-Rs, CD-RWs, magnetic tape, and other optical and non-optical data storage devices. The computer-readable medium can also include computer-readable tangible media distributed over network-connected computer systems so that the computer-readable code is stored and executed in a distributed fashion.
[0090] Although the method operations have been described in a particular order, it should be understood that other housekeeping operations may be performed between operations, or operations may be arranged to occur at slightly different times, or operations may be distributed within the system to allow processing operations to occur at various intervals relative to the processing, so long as the processing of the overlay operations is performed in the desired manner.
[0091] Although the foregoing disclosure has been described in some detail for clarity of understanding, it will be apparent that certain changes and modifications may be practiced within the scope of the appended claims. Accordingly, the present embodiments are to be considered as illustrative and not restrictive, and embodiments of the present disclosure are not limited to the details provided herein, but may be modified within the scope of the appended claims and their equivalents.
Claims
1. 1. A rack assembly comprising: A storage server is provided. a plurality of computational threads arranged as one or more streaming arrays; each computation thread in the plurality of computation threads includes one or more computation nodes; a PCI Express (PCIe) fabric providing direct access from each compute node in the plurality of compute threads to the storage server; and one or more PCIe switches provided for each of the one or more streaming arrays in the PCIe fabric; a PCIe switch provided for each of the one or more streaming arrays is communicatively coupled to each of the one or more computational threads of the corresponding streaming array and to the storage server; the storage server is shared by the one or more streaming arrays; The corresponding streaming array is one or more computational threads; A network switch is provided. an array management server configured to manage the one or more computational threads; and an Ethernet fabric, the network switch providing access to a remote network through the Ethernet fabric via a cluster switch; Rack assembly.
2. the network switch provides a control path for streaming management information to the one or more compute threads and to the compute nodes in the one or more streaming arrays through the Ethernet fabric. The rack assembly of claim 1 .
3. and a second PCIe fabric coupling the array management server to the one or more compute threads of the corresponding streaming array. The rack assembly of claim 1 .
4. The network switch provides access from the array management server to the storage servers through the Ethernet fabric. The rack assembly of claim 1 .
5. the array management server is configured to manage a cloud gaming session including two or more instances of a video game running on the compute nodes of the one or more streaming arrays. The rack assembly of claim 1 .
6. At least one computational node of the plurality of computational threads is configured to execute one or more instances of a plurality of video games. The rack assembly of claim 1 .
7. the storage server stores read-only game content shared by the computation nodes in the one or more streaming arrays; The rack assembly of claim 1 .
8. 1. A rack assembly comprising: A storage server is provided. a plurality of computational threads arranged as one or more streaming arrays; each computation thread in the plurality of computation threads includes one or more computation nodes; a PCI Express (PCIe) fabric providing direct access from each compute node in the plurality of compute threads to the storage server; and one or more PCIe switches provided for each of the one or more streaming arrays in the PCIe fabric; a PCIe switch provided for each of the one or more streaming arrays is communicatively coupled to each of the one or more computational threads of the corresponding streaming array and to the storage server; the storage server is shared by the one or more streaming arrays; each of the one or more computing nodes of the corresponding streaming array is communicatively coupled to a PCIe switch provided for each of the one or more streaming arrays via a dedicated line; a PCIe switch provided for each of the one or more streaming arrays is coupled to the storage server through a set of lanes having a number of lanes less than the total number of the one or more compute nodes; Rack assembly.
9. 1. A network architecture, comprising: a plurality of rack assemblies; Each rack assembly in the plurality of rack assemblies comprises: A storage server is provided. a plurality of computational threads arranged as one or more streaming arrays; each computation thread in the plurality of computation threads includes one or more computation nodes; a corresponding streaming array in each of said rack assemblies including one or more computational threads; a PCI Express (PCIe) fabric providing direct access from each compute node in the plurality of compute threads to the storage server; and one or more PCIe switches provided for each of the one or more streaming arrays in the PCIe fabric; a PCIe switch provided for each of the one or more streaming arrays is communicatively coupled to each of the one or more computational threads of the corresponding streaming array and to the storage server; the storage server in each rack assembly is shared by the one or more streaming arrays in each rack assembly; a plurality of cluster switches coupled to the plurality of rack assemblies; and a management server coupled to the plurality of cluster switches; The network architecture, wherein the management server is configured to manage the plurality of rack assemblies.
10. the management server is configured to allocate one or more computer resources in the plurality of rack assemblies to client devices. The network architecture of claim 9.
11. The corresponding streaming array in each rack assembly includes: A network switch is provided. an array management server configured to manage the one or more computational threads; and an Ethernet fabric, the network switch providing access to a remote network through the Ethernet fabric via a cluster switch; The network architecture of claim 9.
12. the network switch in each of the rack assemblies provides a control path for streaming management information over the Ethernet fabric to the one or more compute threads and to the compute nodes in the one or more streaming arrays. The network architecture of claim 11.
13. and a second PCIe fabric coupling the array management server to the one or more compute threads of the corresponding streaming array. The network architecture of claim 11.
14. the network switch in each of the rack assemblies provides access from the array management server to the storage servers over the Ethernet fabric; The network architecture of claim 11.
15. the array management server in each of the rack assemblies is configured to manage a cloud gaming session including two or more instances of a video game running on compute nodes of the one or more streaming arrays. The network architecture of claim 11.
16. at least one computing node of the plurality of computing threads in each of the rack assemblies is configured to execute one or more instances of a plurality of video games; The network architecture of claim 9.
17. the storage server in each rack assembly stores read-only game content shared by the compute nodes in the one or more streaming arrays; The network architecture of claim 9.
18. each of the one or more computing nodes of the corresponding streaming array in each of the rack assemblies is communicatively coupled to a PCIe switch provided for each of the one or more streaming arrays via a dedicated line; a PCIe switch provided for each of the one or more streaming arrays in each of the rack assemblies is coupled to the storage server through a set of lanes having a number of lanes less than the total number of the one or more compute nodes in the corresponding streaming array; The network architecture of claim 9.
Citation Information
Patent Citations
Computer system
JP2018101440A
Rack assembly structure
US20170102510A1
Ultra high-speed low-latency network storage
WO2019112710A1