Techniques for pre-configuring accelerators by predicting bitstreams

By predicting the bitstream and pre-configuring it on the accelerator, the problem of long bitstream acquisition and storage time for accelerator devices is solved, improving the working efficiency and performance of accelerator devices.

CN109426629BActive Publication Date: 2025-10-31INTEL CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201811002210.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-12-30
Filing Date
2018-08-30
Publication Date
2025-10-31
Estimated Expiration
2038-08-30

AI Technical Summary

Technical Problem

Accelerator devices take a long time to acquire and store bit streams, resulting in delays and performance losses in work acceleration.

Method used

By predicting the bitstream and pre-configuring the accelerator, the orchestrator server predicts the bitstream and pre-stores it on the accelerator, reducing bitstream acquisition and storage time.

Benefits of technology

This improved the efficiency of the accelerator equipment, reduced bitstream acquisition and storage time, and enhanced the performance of the accelerator equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109426629B_ABST
    Figure CN109426629B_ABST
Patent Text Reader

Abstract

Techniques for pre-configuring accelerators by predicting bitstreams. These techniques include communication circuitry and computing devices. The computing devices include a computing engine for determining one or more bitstreams registered on each of a plurality of accelerators. The computing engine is further used to predict the next job requested for acceleration from at least one of the plurality of computing slides, predict the bitstream from a bitstream library for executing the predicted next job to be accelerated, and determine whether the predicted bitstream is already registered on one of the accelerators. In response to determining that the predicted bitstream is not registered on one of the accelerators, the computing engine is used to select an accelerator from the plurality of accelerators that satisfies the characteristics of the predicted bitstream and register the predicted bitstream on the selected accelerator.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims the benefits of Indian Provisional Patent Application No. 201741030632, filed August 30, 2017, and U.S. Provisional Patent Application No. 62 / 584,401, filed November 10, 2017. Background Technology

[0003] The demand for accelerator devices has continued to increase as they offer significantly greater processing capacity than general-purpose accelerators for various technology areas, such as machine learning and genomics. The typical architecture of accelerator devices (such as field-programmable gate arrays (FPGAs), cryptographic accelerators, graphics accelerators, and / or compression accelerators (referred to herein as “accelerator devices,” “accelerators,” or “accelerator resources”) that are used to accelerate the execution of a set of operations in workloads (e.g., processes, applications, services, etc.) allows for the static allocation of a specified amount of shared resources (e.g., high-bandwidth memory, data storage devices, etc.) across different parts of the accelerator device’s logic (e.g., circuitry) to execute corresponding operations in one or more workloads. When an accelerator device receives a request from a computing device to accelerate a task, it can acquire a bitstream capable of executing the requested task and can register the bitstream on the accelerator device. However, the acquisition and registration of the bitstream can consume a non-negligible amount of time, thereby delaying the acceleration of the task and reducing the performance benefits associated with using an accelerator device. Attached Figure Description

[0004] The concepts described herein are illustrated in the accompanying drawings by way of example rather than limitation. For the sake of simplicity and clarity, the elements illustrated in the drawings are not necessarily drawn to scale. Where deemed appropriate, reference numerals have been repeated in the drawings to indicate corresponding or similar elements.

[0005] Figure 1 This is a simplified diagram of at least one embodiment of a data center for performing workloads using decomposed resources;

[0006] Figure 2 yes Figure 1 A simplified diagram of at least one embodiment of a data center pod;

[0007] Figure 3 It can be included Figure 2 A perspective view of at least one embodiment of the rack in the pod;

[0008] Figure 4 yes Figure 3 Side elevation view of the frame;

[0009] Figure 5 It has a sled installed inside. Figure 3 A perspective view of the rack;

[0010] Figure 6 yes Figure 5 A simplified block diagram of at least one embodiment of the top side of the skateboard;

[0011] Figure 7 yes Figure 6 A simplified block diagram of at least one embodiment of the bottom side of the skateboard;

[0012] Figure 8 Is Figure 1 A simplified block diagram of at least one embodiment of a computing skateboard that can be used in a data center;

[0013] Figure 9 yes Figure 8 Top perspective view of at least one embodiment of the computational skateboard;

[0014] Figure 10 Is Figure 1 A simplified block diagram of at least one embodiment of an accelerometer skateboard that can be used in a data center;

[0015] Figure 11 yes Figure 10 Top perspective view of at least one embodiment of the accelerator skateboard;

[0016] Figure 12 Is Figure 1 A simplified block diagram of at least one embodiment of a storage skateboard that can be used in a data center;

[0017] Figure 13 yes Figure 12 Top perspective view of at least one embodiment of a storage skateboard;

[0018] Figure 14 Is Figure 1 A simplified block diagram of at least one embodiment of a memory spool that can be used in a data center; and

[0019] Figure 15 It is possible Figure 1 A simplified block diagram of a system built within a data center to perform workloads using managed nodes composed of decomposed resources;

[0020] Figure 16 This is a simplified block diagram of at least one embodiment of a system for predicting bitstreams to preconfigure accelerators;

[0021] Figure 17 yes Figure 16 A simplified block diagram of the orchestrator server;

[0022] Figure 18 It can be made by Figure 16 and 17 A simplified block diagram of at least one embodiment of the environment established by the orchestrator server; and

[0023] Figure 19-21 It can be made by Figure 16-18 A simplified flowchart of at least one embodiment of a method performed by an orchestrator server for predicting a bitstream and pre-registering the predicted bitstream on an accelerator. Detailed Implementation

[0024] While the concepts of this disclosure allow for various modifications and alternatives, specific embodiments thereof have been illustrated by way of example in the accompanying drawings and will be described in detail herein. However, it should be understood that there is no intention to limit the concepts of this disclosure to the specific forms disclosed; rather, the invention is intended to cover all modifications, equivalents, and alternatives consistent with this disclosure and the appended claims.

[0025] References to "an embodiment," "embodiment," "illustrative embodiment," etc., in the specification indicate that the described embodiment may include a particular feature, structure, or characteristic; however, each embodiment may or may not include that particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Additionally, when a particular feature, structure, or characteristic is described in connection with an embodiment, it is assumed that implementing such a feature, structure, or characteristic in conjunction with other embodiments is within the knowledge of those skilled in the art, whether explicitly described or not. Furthermore, it should be understood that entries included in a list in the form of "at least one of A, B, and C" may mean (A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C). Similarly, entries listed in the form of "at least one of A, B, or C" may mean (A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C).

[0026] The disclosed embodiments may be implemented in some cases in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried on or stored on a transient or non-transient machine-readable (e.g., computer-readable) storage medium, which may be read and executed by one or more processors. A machine-readable storage medium may be implemented as any storage device, mechanism, or other physical structure (e.g., volatile or non-volatile memory, media disk, or other media device) for storing or transmitting information in a machine-readable form.

[0027] In the accompanying drawings, some structural or methodological features may be shown in a specific arrangement and / or order. However, it should be understood that such a specific arrangement and / or order is not required. Rather, in some embodiments, such features may be arranged in a different manner and / or order than shown in the illustrative drawings. Furthermore, the inclusion of structural or methodological features in a particular drawing does not imply that such features are required in all embodiments, and that such features may be omitted in some embodiments or may be combined with other features.

[0028] Now refer to Figure 1A data center 100, in which decomposed resources can collaboratively perform one or more workloads (e.g., applications on behalf of clients), comprises multiple pods 110, 120, 130, and 140, each comprising one or more rows of racks. As described in more detail herein, each rack houses multiple skateboards, each skateboard being implemented as a computing device primarily equipped with a specific type of resource (e.g., memory devices, data storage devices, accelerator devices, general-purpose processors), such as a server. In an illustrative embodiment, the skateboards in each pod 110, 120, 130, and 140 are connected to multiple pod switches (e.g., switches that route data communication to and from the skateboards within the pod). The pod switches are then connected to a backbone switch 150, which switches communication among the pods (e.g., pods 110, 120, 130, and 140) in the data center 100. In some embodiments, the skateboards may be connected to the infrastructure using Intel Omni-Path technology. As described in more detail herein, resources within a skateboard in data center 100 can be allocated to groups (referred to herein as “managed nodes”) that contain resources from one or more other skateboards to be shared in the execution of workloads. Workloads can be executed as if the resources belonging to a managed node were located on the same skateboard. Resources within a managed node can even belong to skateboards belonging to different racks, and even to different pods 110, 120, 130, 140. Some resources from a single skateboard can be allocated to a managed node, while other resources from the same skateboard can be allocated to different managed nodes (e.g., allocating one processor to a managed node and another processor from the same skateboard to a different managed node). By decomposing resources into skateboards that primarily consist of a single type of resource (e.g., a compute skateboard primarily consisting of compute resources, a memory skateboard primarily consisting of memory resources), and selectively allocating and decomposing the decomposed resources to form managed nodes assigned to perform workloads, Data Center 100 provides more efficient resource utilization than a typical data center (containing compute, memory, storage devices, and perhaps additional resources) that includes hyperconverged servers. Therefore, Data Center 100 can deliver greater performance (e.g., throughput, operations per second, latency, etc.) than a typical data center with the same amount of resources.

[0029] Now refer to Figure 2In the illustrative embodiment, pod 110 includes a set of rows 200, 210, 220, and 230 of racks 240. Each rack 240 may accommodate multiple slides (e.g., sixteen slides) and provide power and data connectivity to the accommodated slides, as described in more detail herein. In the illustrative embodiment, the racks in each row 200, 210, 220, and 230 are connected to multiple pod switches 250 and 260. Pod switch 250 includes a set 252 of ports to which slides of the racks of pod 110 are connected, and another set 254 of ports connecting pod 110 to a backbone switch 150 to provide connectivity to other pods in data center 100. Similarly, pod switch 260 includes a set 262 of ports to which slides of the racks of pod 110 are connected, and a set 264 of ports connecting pod 110 to the backbone switch 150. Thus, the use of pairs of switches 250 and 260 provides redundancy to pod 110. For example, if any of switches 250 and 260 fails, the skateboard in pod 110 can still maintain data communication with the rest of data center 100 (e.g., skateboards in other pods) via the other switches 250 and 260. Furthermore, in the illustrative embodiment, switches 150, 250, and 260 can be implemented as dual-mode optical switches capable of routing both Ethernet protocol communication carrying Internet Protocol (IP) packets and communication according to a second, high-performance link layer protocol (e.g., Intel's Infiniband Omni-Path architecture) via optical signaling media of optical fiber.

[0030] It should be understood that each of the other pods 120, 130, 140 (and any additional pods in data center 100) can be similarly structured as follows: Figure 2 Shown and about Figure 2 The pod 110 is described, as well as pods with similar components (e.g., each pod may have rows of racks accommodating multiple pods, as described above). Furthermore, although two pod switches 250, 260 are shown, it should be understood that in other embodiments, each pod 110, 120, 130, 140 may be connected to a different number of pod switches (e.g., providing even greater failover capacity).

[0031] Now refer to Figure 3-5 Each illustrative rack 240 of data center 100 includes two vertically arranged elongated support columns 302, 304. For example, the elongated support columns 302, 304 can extend upwards from the floor of data center 100 during deployment. Rack 240 also includes one or more horizontal pairs 310 of elongated support arms 312 configured to support the slide of data center 100. Figure 3(Identified by a dashed ellipse), as discussed below. One of the pairs of elongated support arms 312 extends outward from the elongated support rod 302, and the other elongated support arm 312 extends outward from the elongated support rod 304.

[0032] In the illustrative embodiment, each slide of the data center 100 is implemented as a chassis-less slide. That is, each slide has a chassis-less circuit board substrate on which physical resources (e.g., processors, memory, accelerators, storage devices, etc.) are mounted, as discussed in more detail below. Therefore, the rack 240 is configured to receive chassis-less slides. For example, each pair 310 of the elongated support arms 312 defines a slide slot 320 of the rack 240, which is configured to receive the corresponding chassis-less slide. To do this, each illustrative elongated support arm 312 includes a circuit board guide 330 configured to receive the chassis-less circuit board substrate of the slide. Each circuit board guide 330 is secured or otherwise mounted to the top side 332 of the corresponding elongated support arm 312. For example, in the illustrative embodiment, each circuit board guide 330 is mounted at the distal end of the corresponding elongated support arm 312 relative to the corresponding elongated support column 302, 304. For clarity of the drawings, each circuit board guide 330 may not be referenced in each figure.

[0033] Each circuit board guide 330 includes an inner wall defining a circuit board slot 380, the circuit board slot 380 being configured to receive a chassis-less circuit board substrate of a slide 400 when receiving a slide 400 in a corresponding slide slot 320 of the rack 240. To do this, as in Figure 4 As shown, the user (or robot) aligns the chassisless circuit board substrate of the illustrative chassisless slide 400 with the slide slot 320. The user or robot can then slide the chassisless circuit board substrate forward into the slide slot 320, such that each side edge 414 of the chassisless circuit board substrate is received in the corresponding circuit board slot 380 of the circuit board guide 330 of the pair 310 of the elongated support arm 312 defining the corresponding slide slot 320, as shown in Figure 4 As shown in the diagram, each type of resource can be upgraded independently of each other and at its own optimized refresh rate via a robot-accessible and robot-manipulated skateboard having decomposed resources. Furthermore, the skateboard is configured to blind-pair with power and data communication cables in each rack 240, thereby enhancing their ability to be quickly removed, upgraded, reinstalled, and / or replaced. Therefore, in some embodiments, data center 100 can operate on the data center floor without human intervention (e.g., performing workloads, undergoing maintenance and / or upgrades, etc.). In other embodiments, humans can facilitate one or more maintenance or upgrade operations in data center 100.

[0034] It should be understood that each circuit board guide 330 is double-sided. That is, each circuit board guide 330 includes an inner wall defining a circuit board slot 380 on each side of the circuit board guide 330. In this way, each circuit board guide 330 can support a chassisless circuit board substrate on either side. Therefore, a single additional elongated support post can be added to the rack 240 to make the rack 240 able to be held as in Figure 3 This represents a dual-rack solution with twice the number of slide slots 320 shown. The illustrative rack 240 includes seven pairs 310 of elongated support arms 312 defining corresponding seven slide slots 320, each configured to receive and support a corresponding slide 400 as discussed above. Of course, in other embodiments, the rack 240 may include additional pairs or fewer pairs 310 of elongated support arms 312 (i.e., additional or fewer slide slots 320). It should be understood that because the slide 400 is chassis-less, it can have a different overall height than a typical server. Therefore, in some embodiments, the height of each slide slot 320 may be shorter than the height of a typical server (e.g., shorter than a single rack unit "1U"). That is, the vertical distance between each pair 310 of the elongated support arms 312 may be less than that of a standard rack unit "1U". Furthermore, due to the relative reduction in the height of the slide slots 320, the overall height of the rack 240 may be shorter than the height of a conventional rack enclosure in some embodiments. For example, in some embodiments, each of the elongated support columns 302, 304 may have a length of six feet or less. Again, in other embodiments, rack 240 may have different dimensions. Furthermore, it should be understood that rack 240 does not include any walls, enclosures, or anything like that. Rather, rack 240 is an open-enclosure rack exposed to the local environment. Of course, in some cases, where rack 240 forms a tail rack in data center 100, end plates may be attached to one of the elongated support columns 302, 304.

[0035] In some embodiments, various interconnects can be routed upwards or downwards via elongated support pillars 302, 304. To facilitate such routing, each elongated support pillar 302, 304 includes an inner wall defining an internal cavity in which the interconnect can be located. Interconnects routed via the elongated support pillars 304, 304 can be implemented as any type of interconnect, including but not limited to data or communication interconnects for providing communication connectivity to each slide slot 320, power interconnects for providing power to each slide slot 320, and / or other types of interconnects.

[0036] In an illustrative embodiment, rack 240 includes a support platform on which corresponding optical data connectors (not shown) are mounted. Each optical data connector is associated with a corresponding slide slot 320 and configured to mate with the optical data connector of the corresponding slide 400 when a slide 400 is received in the corresponding slide slot 320. In some embodiments, optical connections between components (e.g., slides, racks, and switches) in data center 100 are made using blind-pairing optical connections. For example, a door on each cable can prevent dust from contaminating the optical fiber within the cable. During connection to the blind-pairing optical connector mechanism, the door is pushed open as the end of the cable enters the connector mechanism. Subsequently, the optical fiber within the cable enters the gel within the connector mechanism, and the optical fiber of one cable contacts the optical fiber of another cable within the gel within the connector mechanism.

[0037] The illustrative rack 240 also includes a fan array 370 coupled to the cross support arms of the rack 240. The fan array 370 includes one or more rows of cooling fans 372 aligned horizontally between elongated support columns 302, 304. In an illustrative embodiment, the fan array 370 includes one row of cooling fans 372 for each slide slot 320 of the rack 240. As discussed above, in the illustrative embodiment, each slide 400 does not include any onboard cooling system, and therefore, the fan array 370 provides cooling for each slide 400 received in the rack 240. In an illustrative embodiment, each rack 240 also includes a power source associated with each slide slot 320. Each power source is attached to one of a pair of elongated support arms 312 defining the corresponding slide slot 320. For example, the rack 240 may include a power source coupled to or attached to each elongated support arm 312 extending from the elongated support column 302. Each power source includes a power connector configured to mate with the power connector of the corresponding slide 400 when the slide 400 is received in the corresponding slide slot 320. In the illustrative embodiment, the slide 400 does not include any onboard power supply, and therefore, the power supplied in the rack 240 supplies power to the corresponding slide 400 when mounted to the rack 240.

[0038] Now refer to Figure 6 In the illustrative embodiment, the skateboard 400 is configured to be installed in a corresponding rack 240 of the data center 100 discussed above. In some embodiments, each skateboard 400 may be optimized or otherwise configured to perform a specific task, such as a computing task, an acceleration task, a data storage task, etc. For example, the skateboard 400 may be implemented as follows regarding Figure 8-9 The discussion focuses on the calculation of the skateboard 800, as follows: Figure 10-11 The discussion concerns the accelerator skateboard 1000, as follows: Figure 12-13The storage skateboard 1200 discussed, or a skateboard optimized or otherwise configured to perform other specialized tasks, such as the following regarding Figure 14 The memory slide 1400 is being discussed.

[0039] As discussed above, the illustrative slide 400 includes a chassis-less circuit board substrate 602 supporting various physical resources (e.g., electrical components) mounted thereon. It should be understood that the circuit board substrate 602 is "chassis-less" because the slide 400 does not include a housing or enclosure. Instead, the chassis-less circuit board substrate 602 is exposed to the local environment. The chassis-less circuit board substrate 602 can be formed of any material capable of supporting the various electrical components mounted thereon. For example, in the illustrative embodiment, the chassis-less circuit board substrate 602 is formed of FR-4 glass-reinforced epoxy resin laminate. Of course, in other embodiments, other materials may be used to form the chassis-less circuit board substrate 602.

[0040] As discussed in more detail below, the chassisless circuit board substrate 602 includes several features that improve the thermal cooling characteristics of various electrical components mounted on it. As discussed, the chassisless circuit board substrate 602 does not include a housing or enclosure, which improves airflow over the electrical components of the slide plate 400 by reducing those structures that might impede airflow. For example, because the chassisless circuit board substrate 602 is not positioned in a separate housing or enclosure, there is no backplate (e.g., the bottom plate of a chassis) leading to the chassisless circuit board substrate 602 that could impede airflow across the electrical components. Furthermore, the chassisless circuit board substrate 602 has a geometry configured to reduce the length of the airflow path across the electrical components to which it is mounted. For example, the illustrative chassisless circuit board 602 has a width 604 that is larger than the depth 606 of the chassisless circuit board substrate 602. In one particular embodiment, for example, compared to a typical server with a width of approximately 17 inches and a depth of approximately 39 inches, the chassis-less circuit board substrate 602 has a width of approximately 21 inches and a depth of approximately 9 inches. Therefore, the airflow path 608 extending from the front edge 610 to the rear edge 612 of the chassis-less circuit board substrate 602 has a shorter distance relative to a typical server, which can improve the thermal cooling characteristics of the slide plate 400. Furthermore, although not in... Figure 6As illustrated in the diagram, however, the various physical resources mounted to the chassisless circuit board substrate 602 are positioned such that no two substantially heat-generating electrical components shield each other, as discussed in more detail below. That is, two electrical components that do not generate considerable heat during operation (i.e., greater than the nominal heat sufficient to adversely affect the cooling of the other electrical component) are linearly aligned with each other on the chassisless circuit board substrate 602 along the direction of the airflow path 608 (i.e., along the direction extending from the front edge 610 of the chassisless circuit board substrate 602 toward the rear edge 612).

[0041] As discussed above, the illustrative slide 400 includes one or more physical resources 620 mounted to the top side 650 of the chassisless circuit board substrate 602. Although in Figure 6 Two physical resources 620 are shown, but it should be understood that in other embodiments, the skateboard 400 may include one, two, or more physical resources 620. Physical resources 620 may be implemented as any type of processor, controller, or other computing circuit capable of performing various tasks (such as computational functions) and / or controlling the functionality of the skateboard 400, for example, depending on the type or intended function of the skateboard 400. For example, as discussed in more detail below, physical resources 620 may be implemented as a high-performance processor in embodiments where the skateboard 400 is implemented as a computing skateboard, as an accelerator coprocessor or circuit in embodiments where the skateboard 400 is implemented as an accelerator skateboard, as a memory controller in embodiments where the skateboard 400 is implemented as a storage skateboard, or as a collection of memory devices in embodiments where the skateboard 400 is implemented as a memory skateboard.

[0042] The slide 400 also includes one or more additional physical resources 630 mounted to the top side 650 of the chassisless circuit board substrate 602. In the illustrative embodiment, the additional physical resources include a network interface controller (NIC) as discussed in more detail below. Of course, depending on the type and function of the slide 400, the physical resources 630 may include additional or other electrical components, circuitry, and / or devices in other embodiments.

[0043] Physical resource 630 is communicatively coupled to physical resource 630 via input / output (I / O) subsystem 622. I / O subsystem 622 may be implemented as circuitry and / or components for facilitating input / output operations using physical resources 620, physical resource 630, and / or other components of slide plate 400. For example, I / O subsystem 622 may be implemented or otherwise include a memory controller hub, input / output control hub, integrated sensor hub, firmware device, communication links (e.g., point-to-point links, bus links, lines, cables, light guides, printed circuit board traces, etc.) and / or other components and subsystems to facilitate input / output operations. In illustrative embodiments, I / O subsystem 622 is implemented or otherwise includes a dual data rate 4 (DDR4) data bus or a DDR5 data bus.

[0044] In some embodiments, the skateboard 400 may further include a resource-to-resource interconnect 624. The resource-to-resource interconnect 624 may be implemented as any type of communication interconnect capable of facilitating resource-to-resource communication. In illustrative embodiments, the resource-to-resource interconnect 624 is implemented as a high-speed point-to-point interconnect (e.g., faster than the I / O subsystem 622). For example, the resource-to-resource interconnect 624 may be implemented as a QuickPath Interconnect (QPI), an UltraPath Interconnect (UPI), or other high-speed point-to-point interconnect dedicated to resource-to-resource communication.

[0045] The slide 400 also includes a power connector 640 configured to mate with a corresponding power connector of the rack 240 when the slide 400 is mounted in the corresponding rack 240. The slide 400 receives power from the power supply of the rack 240 via the power connector 640 to supply power to its various electrical components. That is, the slide 400 does not include any local power supply (i.e., onboard power) to power its electrical components. The exclusion of local or onboard power facilitates a reduction in the overall footprint of the chassis-less circuit board substrate 602, which can increase the thermal cooling characteristics of the various electrical components mounted on the chassis-less circuit board substrate 602, as discussed above. In some embodiments, power is supplied to the processor 820 via a through-hole directly beneath the processor 820 (e.g., through the bottom side 750 of the chassis-less circuit board substrate 602), providing an increased thermal budget, additional current and / or voltage, and better voltage control compared to a typical board.

[0046] In some embodiments, the skateboard 400 may also include mounting features 642 configured to mate with a robot's mounting arm or other structure to facilitate the robot's placement of the skateboard 600 into the rack 240. Mounting features 642 may be implemented as any type of physical structure that allows the robot to grasp the skateboard 400 without damaging the chassisless circuit board substrate 602 or the electrical components mounted thereto. For example, in some embodiments, mounting features 642 may be implemented as a non-conductive pad attached to the chassisless circuit board substrate 602. In other embodiments, mounting features may be implemented as a bracket, pillar, or other similar structure attached to the chassisless circuit board substrate 602. The specific number, shape, size, and / or configuration of mounting features 642 may depend on the design of the robot configured to manage the skateboard 400.

[0047] Now refer to Figure 7 In addition to the physical resources 630 mounted on the top side 650 of the chassisless circuit board substrate 602, the slide plate 400 also includes one or more memory devices 720 mounted to the bottom side 750 of the chassisless circuit board substrate 602. That is, the chassisless circuit board substrate 602 is implemented as a double-sided circuit board. The physical resources 620 are communicatively coupled to the memory devices 720 via I / O subsystem 622. For example, the physical resources 620 and the memory devices 720 may be communicatively coupled by one or more vias extending through the chassisless circuit board substrate 602. In some embodiments, each physical resource 620 may be communicatively coupled to a different set of one or more memory devices 720. Alternatively, in other embodiments, each physical resource 620 may be communicatively coupled to each memory device 720.

[0048] The memory device 720 can be implemented as any type of memory device capable of storing data for physical resource 620 during operation of slide 400, such as any type of volatile (e.g., dynamic random access memory (DRAM) or non-volatile memory). Volatile memory can be a storage medium that requires power to maintain the state of the data stored by the medium. Non-limiting examples of volatile memory can include various types of random access memory (RAM), such as dynamic random access memory (DRAM) or static random access memory (SRAM). One particular type of DRAM that can be used in a memory module is synchronous dynamic random access memory (SDRAM). In certain embodiments, the DRAM of the memory component may conform to standards promulgated by JEDEC, such as JESD79F for DDR SDRAM, JESD79-2F for DDR2 SDRAM, JESD79-3F for DDR3 SDRAM, JESD79-4A for DDR4 SDRAM, JESD209 for low-power DDR (LPDDR), JESD209-2 for LPDDR2, JESD209-3 for LPDDR3, and JESD209-4 for LPDDR4 (these standards are available at www.jedec.org). Such standards (and similar standards) may be referred to as DDR-based standards, and the communication interface of a storage device implementing such standards may be referred to as a DDR-based interface.

[0049] In one embodiment, the memory device is a block-addressable memory device, such as those based on NAND or NOR technology. The memory device may also include next-generation non-volatile devices, such as Intel 3D XPoint. TMMemory or other byte-addressable, write-in-place non-volatile memory devices. In one embodiment, a memory device may be or may include a memory device using chalcogenide glass, multi-threshold level NAND flash memory, NOR flash memory, single-level or multi-level phase-change memory (PCM), resistive memory, nanowire memory, ferroelectric transistor random access memory (FeTRAM), antiferroelectric memory, magnetoresistive random access memory (MRAM) incorporating memristor technology, resistive memory including metal oxide substrates, oxygen vacancy substrates, and bridged random access memory (CB-RAM), or spin-transfer torque (STT)-MRAM, a device based on spintronic magnetic junction memory, a device based on magnetic tunnel junction (MTJ), a device based on domain walls (DW) and spin-orbit transfer (SOT), a thyristor-based memory device, or a combination of any of the above or other memories. A memory device may refer to the die itself and / or a packaged memory product. In some embodiments, the memory device may include a transistorless stackable cross-point architecture, wherein memory cells are located at the intersection of word lines and bit lines and are individually addressable, and wherein bit storage is based on changes in body resistance.

[0050] Now refer to Figure 8 In some embodiments, skateboard 400 can be implemented as computing skateboard 800. Computing skateboard 800 is optimized or otherwise configured to perform computational tasks. Of course, as discussed above, computing skateboard 800 may rely on other skateboards, such as accelerator skateboards and / or storage skateboards, to perform such computational tasks. Computing skateboard 800 includes various physical resources (e.g., electrical components) similar to those of skateboard 400, which already... Figure 8 They are identified using the same reference number. The above is about... Figure 6 and 7 The description of such components provided applies to the corresponding components of the calculation skateboard 800, and will not be repeated herein for the sake of clarity in the description of the calculation skateboard 800.

[0051] In the illustrative computing skateboard 800, physical resource 620 is implemented as processor 820. Although in Figure 8The illustration shows only two processors 820; however, it should be understood that in other embodiments, the computing slide 800 may include additional processors 820. Illustratively, the processors 820 are implemented as high-performance processors 820 and may be configured to operate at relatively high power ratings. Although processors 820 operating at power ratings higher than typical processors (which operate at approximately 155-230 W) generate additional heat, the enhanced thermal cooling characteristics of the chassis-less circuit board substrate 602 discussed above facilitate higher power operation. For example, in the illustrative embodiment, processors 820 are configured to operate at a power rating of at least 250 W. In some embodiments, processors 820 may be configured to operate at a power rating of at least 350 W.

[0052] In some embodiments, the computing skateboard 800 may further include a processor-to-processor interconnect 842. Similar to the resource-to-resource interconnect 624 of the skateboard 400 discussed above, the processor-to-processor interconnect 842 can be implemented as any type of communication interconnect capable of facilitating communication between the processors. In illustrative embodiments, the processor-to-processor interconnect 842 is implemented as a high-speed point-to-point interconnect (e.g., faster than the I / O subsystem 622). For example, the processor-to-processor interconnect 842 can be implemented as a Fast Path Interconnect (QPI), a Super Path Interconnect (UPI), or other high-speed point-to-point interconnect dedicated to processor-to-processor communication.

[0053] The computing board 800 also includes communication circuitry 830. The illustrative communication circuitry 830 includes a network interface controller (NIC) 832, which may also be referred to as a host fabric interface (HFI). The NIC 832 may be implemented as or otherwise include any type of integrated circuit, discrete circuitry, controller chip, chipset, interposer board, daughter card, network interface card, or other device that can be used by the computing board 800 to connect to another computing device (e.g., other boards 400). In some embodiments, the NIC 832 may be implemented as part of a system-on-a-chip (SoC) including one or more processors, or included on a multi-chip package that also includes one or more processors. In some embodiments, the NIC 832 may include a local processor (not shown) and / or local memory (not shown), both of which are local to the NIC 832. In such embodiments, the local processor of the NIC 832 may be able to perform one or more of the functions of the processor 820. Additionally or alternatively, in such embodiments, the local memory of the NIC 832 may be integrated into one or more components of the computing board at the board level, socket level, chip level, and / or other levels.

[0054] Communication circuitry 830 is communicatively coupled to optical data connector 834. Optical data connector 834 is configured to mate with a corresponding optical data connector of rack 240 when computing slide 800 is mounted in rack 240. Illustratively, optical data connector 834 includes a plurality of optical fibers extending from mating surfaces of optical data connector 834 to optical transceiver 836. Optical transceiver 836 is configured to convert incoming optical signals from rack-side optical data connectors into electrical signals and convert electrical signals into output optical signals destined for rack-side optical data connectors. Although shown as part of optical data connector 834 in the illustrative embodiment, in other embodiments, optical transceiver 836 may form part of communication circuitry 830.

[0055] In some embodiments, the computing slide 800 may further include an expansion connector 840. In such embodiments, the expansion connector 840 is configured to mate with a corresponding connector on an extended chassisless circuit board substrate to provide additional physical resources to the computing slide 800. These additional physical resources may be used, for example, by the processor 820 during operation of the computing slide 800. The extended chassisless circuit board substrate may be substantially similar to the chassisless circuit board substrate 602 discussed above and may include various electrical components mounted thereto. The specific electrical components mounted to the extended chassisless circuit board substrate may depend on the intended function of the extended chassisless circuit board substrate. For example, the extended chassisless circuit board substrate may provide additional computing resources, memory resources, and / or storage resources. Therefore, the additional physical resources of the extended chassisless circuit board substrate may include, but are not limited to, processors, memory devices, storage devices, and / or accelerator circuitry, including, for example, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), security coprocessors, graphics processing units (GPUs), machine learning circuitry, or other specialized processors, controllers, devices, and / or circuitry.

[0056] Now refer to Figure 9 An illustrative embodiment of a computing slide 800 is shown. As shown, a processor 820, communication circuitry 830, and optical data connector 834 are mounted to the top side 650 of a chassis-less circuit board substrate 602. Any suitable attachment or mounting technique can be used to mount the physical resources of the computing slide 800 to the chassis-less circuit board substrate 602. For example, various physical resources can be mounted in corresponding sockets (e.g., processor sockets), cages, or brackets. In some cases, some of the electrical components can be directly mounted to the chassis-less circuit board substrate 602 via soldering or similar techniques.

[0057] As discussed above, the processors 820 and communication circuitry 830 are mounted on the top side 650 of the chassisless circuit board substrate 602 such that no two heat-generating electrical components shield each other. In the illustrative embodiment, the processors 820 and communication circuitry 830 are mounted in corresponding positions on the top side 650 of the chassisless circuit board substrate 602 such that no two of those physical resources are linearly aligned with other physical resources along the direction of the airflow path 608. It should be understood that although the optical data connector 834 is aligned with the communication circuitry 830, the optical data connector 834 does not generate heat or nominal heat during operation.

[0058] The memory device 720 of the computing slide 800 is mounted to the bottom side 750 of the chassisless circuit board substrate 602, as discussed above with respect to slide 400. Although mounted to the bottom side 750, the memory device 720 is communicatively coupled to the processor 820 located on the top side 650 via the I / O subsystem 622. Because the chassisless circuit board substrate 602 is implemented as a double-sided circuit board, the memory device 720 and the processor 820 can be communicatively coupled by one or more through-holes, connectors, or other mechanisms extending through the chassisless circuit board substrate 602. Of course, in some embodiments, each processor 820 may be communicatively coupled to different sets of one or more memory devices 720. Alternatively, in other embodiments, each processor 820 may be communicatively coupled to each memory device 720. In some embodiments, the memory device 720 may be mounted to one or more memory interlayers on the bottom side of the chassisless circuit board substrate 602 and may be interconnected with the corresponding processor 820 via a ball grid array.

[0059] Each of the processors 820 includes a heatsink 850 attached thereto. Due to the mounting of the memory device 720 to the bottom side 750 of the chassisless circuit board substrate 602 (and the vertical spacing of the corresponding slides 400 in the rack 240), the top side 650 of the chassisless circuit board substrate 602 includes additional "free" area or space, which facilitates the use of heatsinks 850 with a larger size than conventional heatsinks used in typical servers. Furthermore, due to the improved thermal cooling characteristics of the chassisless circuit board substrate 602, no processor heatsink 850 includes a cooling fan attached thereto. That is, each of the heatsinks 850 is implemented as a fanless heatsink.

[0060] Now refer to Figure 10In some embodiments, skateboard 400 may be implemented as accelerator skateboard 1000. Accelerator skateboard 1000 is optimized or otherwise configured to perform specialized computational tasks, such as machine learning, encryption, hashing, or other computationally intensive tasks. In some embodiments, for example, computation skateboard 800 may offload tasks to accelerator skateboard 1000 during operation. Accelerator skateboard 1000 includes various components similar to those of skateboard 400 and / or computation skateboard 800, which have already... Figure 10 They are identified using the same reference number. The above is about... Figure 6 , 7 The description of such components provided in 8 applies to the corresponding components of the accelerator skateboard 1000, and will not be repeated herein for the sake of clarity in the description of the accelerator skateboard 1000.

[0061] In the illustrative accelerator slide 1000, physical resource 620 is implemented as accelerator circuit 1020. Although in Figure 10 The diagram shows only two accelerator circuits 1020; however, it should be understood that in other embodiments, the accelerator slide 1000 may include additional accelerator circuits 1020. For example, as in... Figure 11 As shown, in some embodiments, the accelerator slide 1000 may include four accelerator circuits 1020. The accelerator circuits 1020 can be implemented as any type of processor, coprocessor, computing circuit, or other device capable of performing computational or processing operations. For example, the accelerator circuits 1020 can be implemented as, for example, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a security coprocessor, a graphics processing unit (GPU), machine learning circuitry, or other specialized processors, controllers, devices, and / or circuits.

[0062] In some embodiments, the accelerator slide 1000 may further include an accelerator-to-accelerator interconnect 1042. Similar to the resource-to-resource interconnect 624 of the slide 600 discussed above, the accelerator-to-accelerator interconnect 1042 can be implemented as any type of communication interconnect capable of facilitating accelerator-to-accelerator communication. In illustrative embodiments, the accelerator-to-accelerator interconnect 1042 is implemented as a high-speed point-to-point interconnect (e.g., faster than the I / O subsystem 622). For example, the accelerator-to-accelerator interconnect 1042 can be implemented as a Fast Path Interconnect (QPI), a Super Path Interconnect (UPI), or other high-speed point-to-point interconnect dedicated to processor-to-processor communication. In some embodiments, the accelerator circuitry 1020 may be daisy-chained with a primary accelerator circuitry 1020 connected to the NIC 832 and memory 720 via the I / O subsystem 622, and the auxiliary accelerator circuitry 1020 connected to the NIC 832 and memory 720 via the primary accelerator circuitry 1020.

[0063] Now refer to Figure 11 An illustrative embodiment of an accelerator slide 1000 is shown. As discussed above, accelerator circuitry 1020, communication circuitry 830, and optical data connector 834 are mounted to the top side 650 of a chassis-less circuit board substrate 602. Again, the respective accelerator circuitry 1020 and communication circuitry 830 are mounted to the top side 650 of the chassis-less circuit board substrate 602 such that no two heat-generating electrical components shield each other, as discussed above. The memory device 720 of the accelerator slide 1000 is mounted to the bottom side 750 of the chassis-less circuit board substrate 602, as discussed above with respect to slide 600. Although mounted to the bottom side 750, the memory device 720 is communicatively coupled to the accelerator circuitry 1020 located on the top side 650 via I / O subsystem 622 (e.g., through vias). Additionally, each of the accelerator circuitry 1020 may include a heatsink 1070 larger than a conventional heatsink used in a server. As discussed above with reference to heat sink 870, heat sink 1070 can be larger than a conventional heat sink because the “free” area is provided by memory device 750 located on the bottom side 750 of the baseless circuit board substrate 602 rather than on the top side 650.

[0064] Now refer to Figure 12In some embodiments, the skateboard 400 may be implemented as a storage skateboard 1200. The storage skateboard 1200 is optimized or otherwise configured to store data in a local data storage device 1250. For example, during operation, the computing skateboard 800 or the accelerometer skateboard 1000 may store and retrieve data from the data storage device 1250 of the storage skateboard 1200. The storage skateboard 1200 includes various components similar to those of the skateboard 400 and / or the computing skateboard 800, which have already... Figure 12 They are identified using the same reference number. The above is about... Figure 6 , 7 The description of such components provided in section 8 applies to the corresponding components of storage slide 1200, and will not be repeated herein for the sake of clarity in the description of storage slide 1200.

[0065] In the illustrative storage slide 1200, the physical resource 620 is implemented as a storage controller 1220. Although in Figure 12 Only two storage controllers 1220 are shown; however, it should be understood that in other embodiments, the storage slide 1200 may include additional storage controllers 1220. The storage controller 1220 can be implemented as any type of processor, controller, or control circuit capable of controlling the storage and retrieval of data in the data storage device 1250 based on requests received via communication circuitry 830. In the illustrative embodiment, the storage controller 1220 is implemented as a relatively low-power processor or controller. For example, in some embodiments, the storage controller 1220 may be configured to operate at a rated power of approximately 75 watts.

[0066] In some embodiments, the storage slide 1200 may further include a controller-to-controller interconnect 1242. Similar to the resource-to-resource interconnect 624 of the slide 400 discussed above, the controller-to-controller interconnect 1242 can be implemented as any type of communication interconnect capable of facilitating controller-to-controller communication. In illustrative embodiments, the controller-to-controller interconnect 1242 is implemented as a high-speed point-to-point interconnect (e.g., faster than the I / O subsystem 622). For example, the controller-to-controller interconnect 1242 can be implemented as a Fast Path Interconnect (QPI), a Super Path Interconnect (UPI), or other high-speed point-to-point interconnect dedicated to processor-to-processor communication.

[0067] Now refer to Figure 13An illustrative embodiment of the storage slide 1200 is shown. In this illustrative embodiment, the data storage device 1250 is implemented as, or otherwise includes, a storage cage 1252 configured to accommodate one or more solid-state drives (SSDs) 1254. To do this, the storage cage 1252 includes a plurality of mounting slots 1256, each of which is configured to receive a corresponding solid-state drive 1254. Each of the mounting slots 1256 includes a plurality of drive guides 1258 cooperating to define an access opening 1260 of the corresponding mounting slot 1256. The storage cage 1252 is secured to a baseless circuit board substrate 602 such that the access opening faces away from (i.e., towards its front) the baseless circuit board substrate 602. Thus, the solid-state drives 1254 are accessible while the storage slide 1200 is mounted in the corresponding rack 204. For example, the solid-state drives 1254 can be removed from the rack 240 (e.g., via a robot) while the storage slide 1200 remains mounted in the corresponding rack 240.

[0068] Storage cage 1252 illustratively includes sixteen mounting slots 1256 and is capable of mounting and storing sixteen solid-state drives 1254. Of course, in other embodiments, storage cage 1252 may be configured to store additional or fewer solid-state drives 1254. Furthermore, in the illustrative embodiment, the solid-state drives are mounted vertically in storage cage 1252, but in other embodiments they may be mounted in different orientations. Each solid-state drive 1254 can be implemented as any type of data storage device capable of storing long-term data. To do so, solid-state drives 1254 may include the volatile and non-volatile memory devices discussed above.

[0069] As in Figure 13 As shown, the storage controller 1220, communication circuitry 830, and optical data connector 834 are illustratively mounted to the top side 650 of the baseless circuit board substrate 602. Again, as discussed above, any suitable attachment or mounting technique can be used to mount the electrical components of the storage slide 1200 to the baseless circuit board substrate 602, including, for example, sockets (e.g., processor sockets), retainers, brackets, soldered connections, and / or other mounting or securing techniques.

[0070] As discussed above, the respective memory controllers 1220 and communication circuits 830 are mounted on the top side 650 of the baseless circuit board substrate 602 such that no two heat-generating electrical components shield each other. For example, the memory controllers 1220 and communication circuits 830 are mounted in corresponding positions on the top side 650 of the baseless circuit board substrate 602 such that no two of those electrical components are linearly aligned with other electrical components along the direction of the airflow path 608.

[0071] The memory device 720 of the storage slide 1200 is mounted to the bottom side 750 of the baseless circuit board substrate 602, as discussed above with respect to slide 400. Although mounted to the bottom side 750, the memory device 720 is communicatively coupled to the storage controller 1220 located on the top side 650 via the I / O subsystem 622. Again, because the baseless circuit board substrate 602 is implemented as a double-sided circuit board, the memory device 720 and the storage controller 1220 can be communicatively coupled by one or more through-holes, connectors, or other mechanisms extending through the baseless circuit board substrate 602. Each of the storage controllers 1220 includes a heatsink 1270 attached thereto. As discussed above, due to the improved thermal cooling characteristics of the baseless circuit board substrate 602 of the storage slide 1200, no heatsink 1270 includes a cooling fan attached thereto. That is, each of the heatsinks 1270 is implemented as a fanless heatsink.

[0072] Now refer to Figure 14 In some embodiments, skateboard 400 may be implemented as memory skateboard 1400. Memory skateboard 1400 is optimized or otherwise configured to provide other skateboards 400 (e.g., compute skateboard 800, accelerator skateboard 1000, etc.) with access to pools of memory local to memory skateboard 1200 (e.g., in two or more sets 1430, 1432 of memory device 720). For example, during operation, compute skateboard 800 or accelerator skateboard 1000 may use a logical address space to remotely write to and / or read from one or more sets 1430, 1432 of memory skateboard 1200, which is mapped to physical addresses in memory sets 1430, 1432. Memory skateboard 1400 includes various components similar to those of skateboard 400 and / or compute skateboard 800, which have already been... Figure 14 They are identified using the same reference number. The above is about... Figure 6 , 7 The description of such components provided in 8 applies to the corresponding components of memory slide 1400, and will not be repeated herein for the sake of clarity in the description of memory slide 1400.

[0073] In the illustrative memory slide 1400, physical resource 620 is implemented as memory controller 1420. Although in Figure 14Only two memory controllers 1420 are shown; however, it should be understood that in other embodiments, memory slide 1400 may include additional memory controllers 1420. Memory controllers 1420 may be implemented as any type of processor, controller, or control circuitry capable of controlling the writing and reading of data to and from memory sets 1430, 1432 based on requests received via communication circuitry 830. In the illustrative embodiment, each memory controller 1220 is connected to the corresponding memory set 1430, 1432 to write to and read from the memory device 720 within the corresponding memory set 1430, 1432, and to implement any permissions (e.g., read, write, etc.) associated with the slide 400 that has sent a request to memory slide 1400 to perform a memory access operation (e.g., read or write).

[0074] In some embodiments, memory slide 1400 may further include a controller-to-controller interconnect 1442. Similar to the resource-to-resource interconnect 624 of slide 400 discussed above, the controller-to-controller interconnect 1442 may be implemented as any type of communication interconnect capable of facilitating controller-to-controller communication. In illustrative embodiments, the controller-to-controller interconnect 1442 is implemented as a high-speed point-to-point interconnect (e.g., faster than I / O subsystem 622). For example, the controller-to-controller interconnect 1442 may be implemented as a Fast Path Interconnect (QPI), a Super Path Interconnect (UPI), or other high-speed point-to-point interconnect dedicated to processor-to-processor communication. Thus, in some embodiments, memory controller 1420 may access memory located within a memory set 1432 associated with another memory controller 1420 via the controller-to-controller interconnect 1442. In some embodiments, scalable memory controllers are made of multiple smaller memory controllers (referred to herein as “chiplets”) on a memory slide (e.g., memory slide 1400). Chipslets may be interconnected (e.g., using EMIB (Embedded Multi-Die Interconnect Bridge)). The combined chiplet memory controller can be scaled up to a relatively large number of memory controllers and I / O ports (e.g., up to 16 memory channels). In some embodiments, memory controller 1420 can implement memory interleaving (e.g., mapping one memory address to memory set 1430, mapping the next memory address to memory set 1432, and mapping a third address to memory set 1430, etc.). Interleaving can be managed within memory controller 1420 or across network links from (e.g., computing board 800) CPU sockets to memory sets 1430, 1432, and can improve the latency associated with performing memory access operations compared to accessing adjacent memory addresses from the same memory device.

[0075] Additionally, in some embodiments, the memory slide 1400 can be connected via a waveguide connector 1480 to one or more other slides 400 (e.g., in the same rack 240 or adjacent racks 240). In an illustrative embodiment, the waveguide is a 64 mm waveguide providing 16 Rx (i.e., receive) channels and 16 Rt (i.e., transmit) channels. In an illustrative embodiment, each channel is 16 GHz or 32 GHz. In other embodiments, the frequencies may be different. Using a waveguide can provide high-throughput access to a memory pool (e.g., memory sets 1430, 1432) to another slide (e.g., a slide 400 in the same rack 240 or adjacent racks 240 as the memory slide 1400) without adding load to the optical data connector 834.

[0076] Now refer to Figure 15A system for executing one or more workloads (e.g., applications) can be implemented based on data center 100. In an illustrative embodiment, system 1510 includes orchestrator server 1520, which can be implemented as a managed node comprising computing devices (e.g., compute skateboard 800) executing management software (e.g., cloud operating environments such as OpenStack), communicatively coupled to a plurality of skateboards 400 comprising a large number of compute skateboards 1530 (e.g., each similar to compute skateboard 800), memory skateboards 1540 (e.g., each similar to memory skateboard 1400), accelerator skateboards 1550 (e.g., each similar to memory skateboard 1000), and storage skateboards 1560 (e.g., each similar to storage skateboard 1200). One or more of skateboards 1530, 1540, 1550, and 1560 can be grouped into managed nodes 1570 (e.g., via orchestrator server 1520) to collectively execute workloads (e.g., applications 1532 executed in virtual machines or containers). The managed node 1570 can be implemented as a component of physical resources 620 from the same or different skateboards 400, such as processor 820, memory resources 720, accelerator circuitry 1020, or data storage devices 1250. Additionally, the managed node can be created, defined, or "spun up" by the orchestrator server 1520 when workloads are allocated to the managed node or at any other time, and can exist regardless of whether any workloads are currently allocated to the managed node. In an illustrative embodiment, the orchestrator server 1520 can selectively allocate and / or deallocate physical resources 620 from skateboards 400 and / or add or remove one or more skateboards 400 from the managed node 1570 based on the quality of service (QoS) objectives associated with the service level protocol used for the workload (e.g., application 1532), such as performance objectives associated with throughput, latency, instructions per second, etc. In doing so, orchestrator server 1520 may receive telemetry data indicating the performance status (e.g., throughput, latency, instructions per second, etc.) of each slide 400 of managed node 1570 and compare the telemetry data with the quality objectives of the service to determine whether the quality objectives of the service are being met. If so, orchestrator server 1520 may additionally determine whether one or more physical resources can be unallocated from managed node 1570 while still meeting the QoS objectives, thereby freeing up those physical resources for use on another managed node (e.g., to perform a different workload). Alternatively, if the QoS objectives are not currently being met, orchestrator server 1520 may determine to dynamically allocate additional physical resources to assist in the execution of the workload while it is being executed (e.g., application 1532).

[0077] Furthermore, in some embodiments, orchestrator server 1520 may identify trends in resource utilization of a workload (e.g., application 1532) such as by identifying the phases of execution of a workload (e.g., time periods in which different operations are performed, each with different resource utilization characteristics) and by proactively identifying available resources in data center 100 and allocating them to managed nodes 1570 (e.g., within a predefined time period at the start of an associated phase). In some embodiments, orchestrator server 1520 may model performance based on various latency and distribution schemes to place workloads among compute skateboards and other resources (e.g., accelerator skateboards, memory skateboards, storage skateboards) in data center 100. For example, orchestrator server 1520 may utilize models that take into account the performance of resources on skateboard 400 (e.g., FPGA performance, memory access latency, etc.) and the performance of paths over the network to resources (e.g., FPGA) (e.g., congestion, latency, bandwidth). Therefore, the orchestrator server 1520 can determine which resource(s) should be used with which workloads based on the total latency associated with each potential resource available in the data center 100 (e.g., latency associated with the performance of the resource itself, in addition to the latency associated with the path through the network between the computing skateboard performing the workload and the skateboard 400 on which the resource is located).

[0078] In some embodiments, orchestrator server 1520 may use telemetry data (e.g., temperature, fan speed, etc.) reported from skateboard 400 to generate a map of heat generation in data center 100 and allocate resources to managed nodes based on the map of heat generation associated with different workloads and the predicted heat generation to maintain target temperature and heat distribution in data center 100. Additionally or alternatively, in some embodiments, orchestrator server 1520 may organize the received telemetry data into a hierarchical model indicating relationships between managed nodes (e.g., spatial relationships, such as the physical location of resources of managed nodes within data center 100, and / or functional management, such as grouping of managed nodes by clients serving them through managed nodes, the types of functions typically performed by managed nodes, managed nodes that typically share or exchange workloads with each other, etc.). Based on differences in physical location and resources among managed nodes, a given workload may exhibit different resource utilization across the resources of different managed nodes (e.g., causing different internal temperatures, using different percentages of processor or memory capacity). The orchestrator server 1520 can determine discrepancies based on telemetry data stored in the hierarchical model, and take into account predictions of future resource utilization of the workload if workloads are reallocated from one managed node to another, in order to accurately balance resource utilization in the data center 100.

[0079] To reduce the computational load on orchestrator server 1520 and the data transfer load over the network, in some embodiments, orchestrator server 1520 may send self-test information to skateboards 400 so that each skateboard 400 can locally (e.g., on skateboard 400) determine whether the telemetry data generated by skateboard 400 meets one or more conditions (e.g., available capacity meeting a predefined threshold, temperature meeting a predefined threshold, etc.). Each skateboard 400 may then report a simplified result (e.g., yes or no) back to orchestrator server 1520, which orchestrator server 1520 may utilize the simplified result when determining how to allocate resources to managed nodes.

[0080] Now refer to Figure 16 System 1600 includes an orchestrator server 1602 and an accelerator pool 1606 communicating with multiple compute skateboards 1604 (e.g., skateboard 400, compute skateboard 800, 1530 with physical resources 620) from the same or different racks (e.g., one or more of racks 240), which can be referenced above. Figure 1The described data center 100 implementation is used to predict bitstream 1622 to pre-register the predicted bitstream 1622 on an accelerator (which can execute the next job requested for acceleration). The accelerator pool 1606 includes multiple accelerator skateboards 1608A, 1608B (e.g., skateboard 400, accelerator skateboard 1000, 1550 with physical resource 620) from the same or different racks (e.g., one or more of racks 240). It should be understood that a job can be one or more tasks of an application or workload. It should be understood that in other embodiments, system 1600 may include different numbers of compute skateboards 1604, accelerator skateboards 1608, and / or other skateboards (e.g., memory skateboards or storage skateboards).

[0081] In use, as described in more detail below, the orchestrator server 1602 of system 1600 can predict the next job requested from computational skateboard 1604 to be accelerated and the bitstream 1622 capable of executing the predicted next job. It should be understood that bitstream 1622 can be executable code that can be used to implement a set of functions, and is also referred to as kernel 1680 when bitstream 1622 is registered and executed on accelerator 1670. Bitstream 1622 can be implemented as any set of instructions executable by accelerator 1670 to perform corresponding functions. For example, bitstream 1622 can be implemented as a set of instructions for performing cryptographic functions, arithmetic functions, hashing functions, and / or other functions executable by accelerator 1670. Based on the predicted bitstream 1622 and the characteristics of the next job, the orchestrator server 1602 can determine the available accelerators 1670-1, 1670-2, 1670-3, or 1670-4 capable of executing the predicted next job, and configure the determined accelerators 1670-1, 1670-2, 1670-3, or 1670-4 to pre-register the predicted bitstream 1622. By managing future bitstream registration and pre-configuring accelerators 1670, the illustrative system 1600 can reduce the overall execution time, which typically includes the time elapsed when determining an accelerator 1670 that meets the bitstream requirements and the time elapsed when registering the bitstream 1622 on the determined accelerator 1670.

[0082] The orchestrator server 1602 can be implemented as any type of computing device capable of performing the functions described herein, including predicting the next job to be requested for acceleration, predicting a bitstream 1622 capable of performing the predicted next job, and configuring the accelerator 1670 to register the predicted bitstream 1622. (As in...) Figure 16As shown, the illustrative orchestrator server 1602 includes a bitstream library 1620 that stores bitstreams 1622 currently registered on accelerators 1670-1, 1670-2, 1670-3, and 1670-4 of system 1600. Additionally, the orchestrator server 1602 tracks bitstreams 1622 currently registered on accelerators 1670-1, 1670-2, 1670-3, and 1670-4 during operation. In some embodiments, the bitstream library 1620 may further store bitstreams 1622 previously registered on accelerators 1670-1, 1670-2, 1670-3, and 1670-4. In the illustrative embodiment, the orchestrator server 1602 can select from the bitstream library 1620 the bitstream 1622 predicted to receive the predicted next job. In some embodiments, orchestrator server 1602 may support cloud operating environments such as OpenStack, and accelerator pool 1606 and compute skateboard 1604 may execute one or more applications or processes (i.e., jobs or workloads), such as in virtual machines or containers.

[0083] Compute skateboard 1604 can be implemented as any type of computing device having a central processing unit (CPU) 1640 capable of executing workloads (e.g., application 1642) and performing other functions described herein (including requesting accelerator 1670 to accelerate work via orchestrator server 1602). For example, compute skateboard 1604 can be implemented as compute skateboard 800, 1530, computer, distributed computing system, multiprocessor system, network device (e.g., physical or virtual), desktop computer, workstation, laptop computer, notebook computer, processor-based system, or network device.

[0084] In the illustrative embodiment, accelerator pool 1606 includes two accelerator slides 1608A and 1608B, and each accelerator slide 1608A and 1608B includes prefetch logic units 1660A and 1660B, bitstream caches 1662A and 1662B, and two accelerators 1670-1, 1670-2 or 1670-3, 1670-4. It should be understood that in other embodiments, accelerator slide 1608 may include other or additional components, such as those typically found in typical computing devices (e.g., various input / output devices and / or other components). Furthermore, in some embodiments, one or more of the illustrative components may be incorporated into another component or otherwise form part of another component. It should be understood that in other embodiments, accelerator pool 1606 may include different numbers of accelerator slides 1608A and 1608B, and each accelerator slide 1608A and 1608B may include different numbers of accelerators 1670. It should be understood that in other embodiments, one or more of the illustrative components may be incorporated into another component or otherwise form part of another component.

[0085] Accelerator 1670 can be implemented as a single device, such as an integrated circuit, embedded system, field-programmable array (FPGA), system-on-a-chip (SOC), application-specific integrated circuit (ASIC), reconfigurable hardware or hardware circuitry, or other specialized hardware capable of communicating with compute skateboard 1604 and orchestrator server 1602 to pre-register the predicted bitstream 1622 to execute the next task to be requested from compute skateboard 1604 for acceleration. In the illustrative embodiment, each accelerator 1670-1, 1670-2, 1670-3, 1670-4 includes two cores 1680 (e.g., each set of circuitry and / or executable code, i.e., bitstream 1622, that can be used to implement a set of functions) registered on each accelerator 1670-1, 1670-2, 1670-3, 1670-4. However, it should be understood that in other embodiments, each accelerator 1670 may include a different number of cores 1680 on the corresponding accelerator 1670.

[0086] In such a state Figure 16In the illustrative embodiment shown, accelerator slide 1608 further includes a prefetch logic unit 1660 and a bitstream cache 1662. The prefetch logic unit 1660 may be implemented as circuitry, components, or any type of device capable of prefetching bitstream 1622 from a bitstream library 1620 of orchestrator server 1602, the bitstream 1622 being predicted to perform the predicted next operation to be received from compute slide 1604. For example, if orchestrator server 1602 determines that accelerator 1670-1 of accelerator slide 1608A is configured to register predicted bitstream A 1622A, then the prefetch logic unit 1660A of accelerator slide 1608A retrieves the predicted bitstream 1622 from the bitstream library 1620 and registers the predicted bitstream 1622 on accelerator 1670-1. In some embodiments, the prefetch logic unit 1660 may be implemented as a coprocessor, embedded circuitry, ASIC, FPGA, and / or other specialized circuitry.

[0087] Bitstream cache 1662 can be implemented as any device or circuit capable of determining bitstream registered data for each accelerator 1670 on the corresponding accelerator slide 1608 and transmitting the bitstream registered data to orchestrator server 1602. For example, the bitstream registered data includes one or more bitstreams 1622 currently registered on each accelerator 1670 on the corresponding accelerator slide 1608. In some embodiments, bitstream cache 1662 can update the timestamps of bitstream registration and execution on one of the accelerators 1670 on the corresponding accelerator slide 1608 and transmit the updated timestamp data to orchestrator server 1602 to update bitstream library 1620. It should be understood that the timestamp of each bitstream 1622 can be used to determine the execution mode for each bitstream 1622 for each available application 1642, as discussed further below. Additionally, in some embodiments, bitstream cache 1662 may include one or more secure signatures (e.g., unique codes) that can be used to authenticate incoming bitstreams (e.g., to prevent the registration or execution of rogue bitstreams).

[0088] Now refer to Figure 17 The orchestrator server 1602 can be implemented as any type of computing device capable of performing the functions described herein, including determining one or more bitstreams registered on each of a plurality of accelerators, predicting the next job requested for acceleration from a computing skateboard among a plurality of computing skateboards, predicting the bitstream from a bitstream library for executing the predicted next job to be accelerated, determining whether the predicted bitstream is already registered on an accelerator, determining an accelerator that satisfies the characteristics of the predicted bitstream in response to determining that the predicted bitstream is not registered on an accelerator, and registering the predicted bitstream on the determined accelerator in response to determining the accelerator for executing the predicted next job.

[0089] As in Figure 17 As shown, the illustrative orchestrator server 1602 includes a computing engine 1702, an input / output (I / O) subsystem 1708, communication circuitry 1710, and one or more data storage devices 1714. Of course, in other embodiments, the orchestrator server 1602 may include other or additional components, such as those commonly found in computers (e.g., displays, peripherals, etc.). Furthermore, in some embodiments, one or more of the illustrative components may be incorporated into another component or otherwise form part of another component.

[0090] The computing engine 1702 can be implemented as any type of device or collection of devices capable of performing the various computing functions described below. In some embodiments, the computing engine 1702 can be implemented as a single device, such as an integrated circuit, embedded system, FPGA, system-on-a-chip (SoC), or other integrated system or device. Furthermore, in some embodiments, the computing engine 1702 includes or is implemented as a processor 1704 and a memory 1706. The processor 1704 can be implemented as any type of processor capable of performing the functions described herein. For example, the processor 1704 can be implemented as a single-core or multi-core processor, a microcontroller, or other processor or processing / control circuitry. In some embodiments, the processor 1704 can be implemented as, include, or coupled to an FPGA, ASIC, reconfigurable hardware or hardware circuitry, or other specialized hardware to facilitate the performance of the functions described herein.

[0091] Memory 1706 can be implemented as any type of volatile (e.g., dynamic random access memory (DRAM), etc.) or non-volatile memory or data storage device capable of performing the functions described herein. Volatile memory can be a storage medium that requires power to maintain the state of the data stored by the medium. Non-limiting examples of volatile memory can include various types of random access memory (RAM), such as DRAM or static random access memory (SRAM). One particular type of DRAM that can be used in a memory module is synchronous dynamic random access memory (SDRAM). In certain embodiments, the DRAM of the memory component may conform to standards promulgated by JEDEC, such as JESD79F for DDR SDRAM, JESD79-2F for DDR2 SDRAM, JESD79-3F for DDR3 SDRAM, JESD79-4A for DDR4 SDRAM, JESD209 for low-power DDR (LPDDR), JESD209-2 for LPDDR2, JESD209-3 for LPDDR3, and JESD209-4 for LPDDR4 (these standards are available at www.jedec.org). Such standards (and similar standards) may be referred to as DDR-based standards, and the communication interface of a storage device implementing such standards may be referred to as a DDR-based interface.

[0092] In one embodiment, the memory device is a block-addressable memory device, such as those based on NAND or NOR technology. The memory device may also include future-generation non-volatile devices, such as three-dimensional cross-point memory devices (e.g., Intel 3D XPoint). TM This refers to a memory device or other byte-addressable, write-in-place non-volatile memory device. In one embodiment, the memory device may be or may include a memory device using chalcogenide glass, multi-threshold level NAND flash memory, NOR flash memory, single-level or multi-level phase-change memory (PCM), resistive memory, nanowire memory, ferroelectric transistor random access memory (FeTRAM), antiferroelectric memory, magnetoresistive random access memory (MRAM) incorporating memristor technology, resistive memory including metal oxide substrates, oxygen vacancy substrates, and bridged random access memory (CB-RAM), or spin-transfer torque (STT)-MRAM, a device based on spintronic magnetic junction memory, a device based on magnetic tunnel junction (MTJ), a device based on domain walls (DW) and SOT (spin-orbit transfer), a thyristor-based memory device, or a combination of any of the above or other memories. The memory device may refer to the die itself and / or a packaged memory product.

[0093] In some embodiments, a 3D cross-point architecture (e.g., Intel 3D XPoint) TM The memory may include a transistorless, stackable cross-point architecture, wherein memory cells are located at the intersection of word lines and bit lines and are individually addressable, and wherein bit storage is based on changes in body resistance. In some embodiments, all or part of the memory 1706 may be integrated into the processor 1704. In operation, the memory 1706 may store various software and data used during operation.

[0094] The computing engine 1702 is communicatively coupled to other components of the orchestrator server 1602 via an I / O subsystem 1708, which may be implemented as circuitry and / or components for facilitating input / output operations using the computing engine 1702 (e.g., using the processor 1704 and / or memory 1706) and other components of the orchestrator server 1602. For example, the I / O subsystem 1708 may be implemented as or otherwise include a memory controller hub, an input / output control hub, an integrated sensor hub, firmware devices, communication links (e.g., point-to-point links, bus links, lines, cables, light guides, printed circuit board traces, etc.) and / or other components and subsystems to facilitate input / output operations. In some embodiments, the I / O subsystem 1708 may form part of a system-on-a-chip (SoC) and be incorporated into the computing engine 1702 along with one or more of the processor 1704, memory 1706, and other components of the orchestrator server 1602.

[0095] The communication circuit 1710 can be implemented as any communication circuit, device, or combination thereof capable of enabling communication between the orchestrator server 1602 and another computing device (e.g., computing skateboard 1604, accelerometer skateboard 1608, etc.). The communication circuit 1710 can be configured to implement such communication using any one or more communication technologies (e.g., wired or wireless communication) and associated protocols (e.g., Ethernet, Bluetooth®, Wi-Fi®, WiMAX, etc.).

[0096] The illustrative communication circuitry 1710 includes a network interface controller (NIC) 1712, which may also be referred to as a host infrastructure interface (HFI). The NIC 1712 may be implemented as one or more interposer boards, daughter cards, network interface cards, controller chips, chipsets, or other devices that can be used by the orchestrator server 1602 to connect to another computing device (e.g., compute skateboard 1604, accelerator skateboard 1608, etc.). In some embodiments, the NIC 1712 may be implemented as part of a system-on-a-chip (SoC) including one or more processors, or included in a multi-chip package that also includes one or more processors. In some embodiments, the NIC 1712 may include a local processor (not shown) and / or local memory (not shown), both of which are local to the NIC 1712. In such embodiments, the local processor of the NIC 1712 may be able to perform one or more of the functions of the computing engine 1702 described herein. Additionally or alternatively, in such embodiments, the local memory of the NIC 1712 may be integrated into one or more components of the orchestrator server 1602 at the board level, socket level, chip level, and / or other levels.

[0097] One or more illustrative data storage devices 1714 can be implemented as any type of device configured for short-term or long-term data storage, such as, for example, memory devices and circuitry, memory cards, hard disk drives (HDDs), solid-state drives (SSDs), or other data storage devices. Each data storage device 1714 may include a system partition storing data and firmware code for the data storage device 1714. Each data storage device 1714 may also include an operating system partition storing data files and executable files for the operating system. Additionally or alternatively, the orchestrator server 1602 may include one or more peripheral devices (not shown). Such peripheral devices may include any type of peripheral device typically found in computing devices, such as a monitor, speakers, mouse, keyboard, and / or other input / output devices, interface devices, and / or other peripheral devices.

[0098] The calculation skateboard 1604 and accelerator skateboard 1608 can have the same characteristics as in... Figure 17 The descriptions of those components of orchestrator server 1602 are equally applicable to the descriptions of those components of the device, and are not repeated herein for clarity of description. Furthermore, it should be understood that either computing skateboard 1604 and accelerometer skateboard 1608 may include other components (e.g., accelerometer 1670) and / or other components, sub-components, and devices typically found in computing devices, which have not been discussed above with reference to orchestrator server 1602 and are not discussed herein for clarity of description.

[0099] As described above, orchestrator server 1602 and skateboards 1604, 1608 are illustratively in communication via a network (not shown), which can be implemented as any type of wired or wireless communication network, including global networks (e.g., the Internet), local area networks (LANs) or wide area networks (WANs), cellular networks (e.g., Global System for Mobile Communications (GSM), 3G, Long Term Evolution (LTE), Global Microwave Access Interoperability (WiMAX), etc.), digital subscriber line (DSL) networks, cable networks (e.g., coaxial cables, fiber optic networks, etc.) or any combination thereof.

[0100] Now refer to Figure 18 In an illustrative embodiment, orchestrator server 1602 can establish environment 1800 during operation. In an illustrative embodiment, environment 1800 includes a bitstream library 1620, which can be implemented to indicate any data of previously and currently registered bitstreams 1622 on accelerators 1670 of accelerator pool 1606. Furthermore, illustrative environment 1800 includes a network communicator 1802, a bitstream updater 1804, a job predictor 1806, a bitstream predictor 1808, and an accelerator manager 1810. (As in...) Figure 18 As shown, the accelerometer manager 1810 further includes an accelerometer feature determiner 1812, a registered bitstream tracker 1814, and a bitstream prefetcher 1816. Each of the components of the environment 1800 can be implemented as hardware, firmware, software, or a combination thereof. Thus, in some embodiments, one or more of the components of the environment 1800 can be implemented as a collection of circuits or electrical devices (e.g., network communicator circuit 1802, bitstream updater circuit 1804, job predictor circuit 1806, bitstream predictor circuit 1808, accelerometer manager circuit 1810, accelerometer feature determiner circuit 1812, registered bitstream tracker circuit 1814, bitstream prefetcher circuit 1816, etc.).

[0101] In illustrative environment 1800, network communicator 1802 is configured to facilitate inbound and outbound network communications (e.g., network traffic, network packets, network flows, etc.) to and from orchestrator server 1602. To do this, network communicator 1802 is configured to receive and process data from remote systems or computing devices (e.g., compute skateboard 1604, accelerator skateboard 1608 of accelerator pool 1606, etc.), and to prepare and transmit data to remote systems or computing devices (e.g., compute skateboard 1604, accelerator skateboard 1608 of accelerator pool 1606, etc.). Accordingly, in some embodiments, at least a portion of the functionality of network communicator 1802 may be performed by the communication circuitry of orchestrator server 1602.

[0102] As discussed above, a bitstream updater 1804, which can be implemented as hardware, firmware, software, virtualized hardware, emulated architecture, and / or a combination thereof, is configured to update the bitstream library 1620 to keep track of which bitstream 1622 is currently registered on which accelerator 1670. As discussed above, the bitstream library 1620 stores bitstreams 1622 currently registered on accelerators 1670 of the accelerator pool 1606. To do this, the bitstream updater 1804 can receive bitstream registration data from the bitstream cache 1662 of each accelerator slide 1608, indicating which bitstreams 1622 are currently registered on each accelerator 1670 of the corresponding accelerator slide 1608, and update the bitstream library 1620 accordingly. In some embodiments, after registering a bitstream 1622 on an accelerator 1670, the bitstream updater 1804 can update the timestamp of the bitstream 1622 in the bitstream library 1620, the timestamp indicating the time of bitstream registration. Alternatively, in some embodiments, the bitstream cache 1662 of each accelerator slide 1608 may update the timestamps of bitstream registration and execution on one of the accelerators 1670 corresponding to the accelerator slide 1608, and transmit the updated timestamp data to the bitstream updater 1804 to update the bitstream library 1620. It should be understood that the timestamp of each bitstream 1622 may be used to determine the execution mode for each bitstream 1622 for each available application 1642. In some embodiments, the bitstream cache 1662 may temporarily store copies of one or more bitstreams already received from the bitstream library 1620.

[0103] As discussed above, a job predictor 1806, which can be implemented as hardware, firmware, software, virtualized hardware, emulation, and / or a combination thereof, is configured to predict the next job to be requested for acceleration from available applications 1642 on compute skateboard 1604. To do this, job predictor 1806 can predict the next job based on the available application 1642 currently executing on compute skateboard 1604. For example, the type and size of available application 1642 can infer the type of job likely to be requested for acceleration by available application 1642. Additionally or alternatively, job predictor 1806 can predict the next job to be requested based on the execution mode of bitstream 1622 for each available application 1642. To do this, in some embodiments, job predictor 1806 can determine the past execution history of each bitstream 1622 for each application 1642. For example, job predictor 1806 can analyze the timestamp of each bitstream 1622 to determine the number of times the corresponding bitstream 1622 was used to execute the job requested by each available application 1642. In other embodiments, the job predictor 1806 may utilize machine learning to predict the execution pattern for each bitstream 1622 for each application 1642.

[0104] As discussed above, a bitstream predictor 1808, which can be implemented as hardware, firmware, software, virtualized hardware, an emulated architecture, and / or a combination thereof, is configured to predict bitstream 1622 from bitstream library 1620 for performing predictions to request acceleration for the next job. To do this, bitstream predictor 1808 may predict bitstream 1622 based on available accelerators 1670 of system 1600, since each available accelerator 1670 may have different capabilities, making them differently suited to perform a given type of job. For example, accelerator 1670 may be able to perform jobs requiring encryption / decryption, compression / decompression, encoding transformations, matrix multiplication, and / or convolutional neural network operations. Additionally or alternatively, in some embodiments, bitstream predictor 1808 may predict bitstream 1622 for performing the next predicted job based on the type of the predicted next job and the type of workload that each accelerator 1670 is capable of accelerating. In other embodiments, bitstream predictor 1808 may predict bitstream 1622 based on the execution pattern of the bitstream for each available application 1642. For example, bitstream predictor 1808 may analyze the timestamp of each bitstream 1622 to determine the number of times the corresponding bitstream 1622 was used to perform work requested by each available application 1642 to determine the execution pattern.

[0105] As discussed above, the accelerator manager 1810, which can be implemented as hardware, firmware, software, virtualized hardware, emulated architecture, and / or combinations thereof, is configured to manage bitstream registers on each accelerator 1670 of the accelerator pool 1606. In an illustrative embodiment, the accelerator manager 1810 determines, based on the bitstream library 1620, whether a predicted bitstream 1622 has been registered on one of the accelerators 1670, and in response to determining that the predicted bitstream 1622 has not been registered on one of the accelerators 1670, the accelerator manager 1810 determines the accelerator 1670 capable of executing the predicted bitstream 1622 and pre-registers the predicted bitstream 1622 on the determined accelerator 1670. To do this, the accelerator manager 1810 further includes an accelerator feature determiner 1812, a registered bitstream tracker 1814, and a bitstream prefetcher 1816.

[0106] As discussed above, the accelerator characteristic determiner 1812, which can be implemented as hardware, firmware, software, virtualized hardware, emulated architectures, and / or combinations thereof, is configured to determine the characteristics of each accelerator 1670. For example, in some embodiments, the accelerator characteristic determiner 1812 may determine the availability of each accelerator 1670. To do this, the accelerator characteristic determiner 1812 may determine the current load on each accelerator 1670. In some embodiments where the accelerator 1670 is implemented as a field-programmable gate array (FPGA), the accelerator characteristic determiner 1812 may further determine the number of free slots in each FPGA and / or the number of free logic gates in each FPGA. Additionally or alternatively, the accelerator characteristic determiner 1812 may determine the type of workload that each accelerator 1670 can accelerate. For example, the accelerator characteristic determiner 1812 may determine whether each accelerator is capable of performing tasks requiring encryption / decryption, compression / decompression, encoding transformation, matrix multiplication, and / or convolutional neural network operations. Furthermore, in some embodiments, the accelerator feature determiner 1812 can determine the physical distance from the computation slide 1604 requesting acceleration to each accelerator 1670. It should be understood that the physical distance between the requesting computation slide 1604 and the accelerator 1670 can affect communication efficiency.

[0107] As discussed above, a registered bitstream tracker 1814, which can be implemented as hardware, firmware, software, virtualized hardware, emulated architecture, and / or a combination thereof, is configured to track bitstreams 1622 currently registered on accelerators 1670. The registered bitstream tracker 1814 can determine the bitstreams 1622 registered on each accelerator slide 1608 of the accelerator 1670 and store the bitstream registered data in a bitstream cache 1662. As discussed above, the bitstream cache 1662 can transmit the bitstream registered data to the orchestrator server 1602, allowing the registered bitstream tracker 1814 to monitor which bitstream 1622 is registered on which accelerator 1670 to perform work from application 1642, in order to analyze the execution patterns of the bitstream 1622 used for the corresponding application 1642.

[0108] As discussed above, a bitstream prefetcher 1816, which can be implemented as hardware, firmware, software, virtualized hardware, emulated architecture, and / or a combination thereof, is configured to prefetch a predicted bitstream 1622 that is likely to be executed on an accelerator 1670 capable of executing the predicted next job. By prefetching and registering the predicted bitstream 1622 on the accelerator 1670, the accelerator 1670 becomes pre-configured to execute the next job before receiving a corresponding job request from the application 1642 of the compute skateboard 1604. As discussed above, by managing future bitstream registration and pre-configuring the accelerator 1670, the illustrative system 1600 can reduce overall execution time, including the time spent determining which accelerator 1670 meets the bitstream requirements and registering the bitstream 1622 on the accelerator 1670.

[0109] Now refer to Figure 19-21 In use, orchestrator server 1602 can perform method 1900 to predict the bitstream 1622 to be registered on accelerator 1670 based on the predicted next job to be accelerated and to preconfigure the accelerator by registering the predicted bitstream 1622 before receiving the next job. Method 1900 begins at block 1902, where orchestrator server 1602 determines one or more bitstreams registered on each accelerator 1670. To do this, in some embodiments, orchestrator server 1602 can monitor bitstream submission and execution on each accelerator 1670 in block 1904, and can update the timestamp of each bitstream execution in block 1906, after which the timestamp can be analyzed to determine the execution mode for each bitstream for each application. In other embodiments, orchestrator server 1602 can receive bitstream registration data from bitstream cache 1662 of each accelerator slide 1608 in block 1908. As discussed above, the bitstream registered data indicates the bitstream 1622 currently registered on each of the accelerators 1670 corresponding to the accelerator slide 1608. Subsequently, in box 1910, the orchestrator server 1602 can update the bitstream library 1620 to keep track of which bitstream 1622 is currently registered on which accelerator 1670. As discussed above, the bitstream library 1620 includes the bitstream 1622 currently registered on the accelerator 1670.

[0110] In box 1912, orchestrator server 1602 predicts the next job to be requested for acceleration. An acceleration request is received from one of the applications 1642 currently executing on one of the compute spools 1604. In some embodiments, orchestrator server 1602 may predict the next job to be accelerated based on the available applications 1642 executing on compute spool 1604 in box 1914. For example, based on the type of available application 1642, orchestrator server 1602 may predict the type of data or job that is likely to be requested for acceleration. Additionally or alternatively, orchestrator server 1602 may predict the next job to be accelerated based on the execution pattern of bitstream 1622 for each available application 1642 in box 1916. To do this, orchestrator server 1602 may determine the past execution history of each bitstream 1622 for each application 1642, as indicated in box 1918, and / or may utilize machine learning to predict execution patterns, as indicated in box 1920. After determining the next predicted task to be accelerated, method 1900 proceeds to... Figure 20 Box 1922 is shown in the image.

[0111] Now refer to Figure 20 In block 1922, orchestrator server 1602 determines the characteristics of each accelerator 1670. To do this, orchestrator server 1602 may determine the availability of each accelerator 1670 in accelerator pool 1606, as indicated in block 1924. For example, in block 1926, orchestrator server 1602 may determine the current load on each accelerator 1670 to determine the availability of each accelerator 1670. As discussed above, in some embodiments, accelerator 1670 may be implemented as a field-programmable gate array (FPGA). In such an embodiment, orchestrator server 1602 may determine the number of free slots in each FPGA in block 1928 and / or the number of free logic gates in each FPGA in block 1930 to determine the availability of each FPGA.

[0112] Additionally or alternatively, in some embodiments, orchestrator server 1602 may determine the type of workload that each accelerator 1670 can accelerate, as indicated in block 1932. For example, in block 1934, orchestrator server 1602 may determine the cryptographic, compression, encoding transformation, matrix multiplication, and / or convolutional neural network computational capabilities of each accelerator 1670. In other embodiments, orchestrator server 1602 may further determine the physical distance from the requested compute skateboard 1604 to each accelerator 1670. As discussed above, the physical distance between the requested compute skateboard 1604 and the accelerator 1670 can affect communication efficiency.

[0113] In box 1938, orchestrator server 1602 predicts bitstream 1622 from bitstream library 1620 for the execution of the predicted next job. To do this, in box 1940, orchestrator server 1602 may predict bitstream 1622 based on one or more available accelerators 1670. Additionally or alternatively, orchestrator server 1602 may predict bitstream 1622 based on the type of the predicted next job to be accelerated and the type of workload that each accelerator 1670 can accelerate, as indicated in box 1942. As discussed above, each accelerator 1670 may have different characteristics that allow accelerator 1670 to execute certain types of data. Therefore, based on the type of the predicted next job, orchestrator server 1602 can determine the characteristics of the accelerators 1670 that need to execute the type of predicted next job. Additionally or alternatively, in box 1944, orchestrator server 1602 may predict bitstream 1622 based on the execution pattern of bitstream 1622 for each available application 1642 to determine which bitstream 1622 is most likely to be needed to perform the predicted next job. For example, as discussed above, orchestrator server 1602 may analyze the execution pattern based on the past execution history of each bitstream 1622 for each application 1642 and / or execution patterns predicted using machine learning. After determining the predicted bitstream for the execution of the predicted next job, method 1900 proceeds to... Figure 21 Box 1946 is shown in the image.

[0114] Now refer to Figure 21 In box 1946, orchestrator server 1602 can determine whether the predicted bitstream 1622 has already been registered on one of the available accelerators 1670 in accelerator pool 1606. If orchestrator server 1602 determines in box 1948 that the predicted bitstream 1622 has already been registered on one of the available accelerators 1670, then orchestrator server 1602 determines that registration of the predicted bitstream 1622 is not required and method 1900 jumps to the end. However, if orchestrator server 1602 determines in box 1948 that the predicted bitstream 1622 has not been registered on one of the available accelerators 1670 and needs to be registered, then method 1900 proceeds to box 1950.

[0115] In block 1950, orchestrator server 1602 determines an accelerator 1670 that satisfies the predicted bitstream characteristics. To do this, in block 1952, orchestrator server 1602 may determine accelerator 1670 based on the specific acceleration capability required by the predicted bitstream 1622. Additionally or alternatively, orchestrator server 1602 may determine accelerator 1670 based on a certain capacity on accelerator 1670 required by the predicted bitstream 1622, as indicated in block 1954. After determining accelerator 1670, method 1900 proceeds to block 1956, where orchestrator server 1602 preconfigures the determined accelerator 1670 by registering the predicted bitstream 1622 on the determined accelerator 1670. In other words, the predicted bitstream 1622 is prefetched from bitstream library 1620 and registered on the determined accelerator 1670 before receiving the predicted next job. By dynamically registering the predicted bit stream 1622 before receiving the predicted next job, the system 1600 can reduce the likelihood of delays in acquiring and registering the bit stream after receiving the next job to be accelerated.

[0116] This application provides the following technical solution:

[0117] 1. A computing device for an accelerator in a plurality of accelerators in a pre-configured system, the computing device comprising:

[0118] Communication circuits;

[0119] The computation engine is configured to: (i) determine one or more bit streams registered on each of a plurality of accelerators; (ii) predict the next job to be accelerated from at least one of a plurality of computational slides; (iii) predict the bit stream from the bit stream library to execute the predicted next job to be accelerated; (iv) determine whether the predicted bit stream has already been registered on one of the accelerators; (v) in response to determining that the predicted bit stream has not been registered on one of the accelerators, select an accelerator from the plurality of accelerators that satisfies the characteristics of the predicted bit stream; and (vi) in response to determining that an accelerator satisfies the characteristics of the predicted bit stream, register the predicted bit stream on the determined accelerator.

[0120] 2. The computing device as described in technical solution 1, wherein determining one or more bit streams registered on each accelerator includes monitoring bit stream submission and execution on each accelerator.

[0121] 3. The computing device as described in technical solution 1, wherein determining one or more bit streams registered on each accelerator includes receiving bit stream registration data from each accelerator, wherein the bit stream registration data indicates the bit stream currently registered on the corresponding accelerator.

[0122] 4. The computing device as described in technical solution 1, wherein determining one or more bit streams registered on each accelerator includes updating the bit stream library to track the bit streams currently registered on each accelerator.

[0123] 5. The computing device as described in technical solution 1, wherein predicting the next job to be requested for acceleration includes predicting the next job to be requested for acceleration based on available applications currently executing on multiple computing spools.

[0124] 6. The computing device of claim 1, wherein predicting the next job to be requested for acceleration includes predicting the next job to be requested for acceleration based on the execution mode of the bitstream for each available application currently being executed on a plurality of computing spools.

[0125] 7. The computing device as described in technical solution 1, wherein predicting bitstreams from a bitstream library includes predicting bitstreams based on available accelerators in the system.

[0126] 8. The computing device as described in technical solution 1, wherein predicting bitstreams from a bitstream library includes predicting bitstreams based on the type of the next job predicted and the type of workload that each accelerator can accelerate.

[0127] 9. The computing device as described in technical solution 1, wherein predicting bitstreams from a bitstream library includes predicting the execution mode of the bitstreams for each available application.

[0128] 10. The computing device of claim 1, wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream includes determining the accelerator based on the specific accelerator capabilities required by the predicted bitstream.

[0129] 11. The computing device as described in technical solution 1, wherein selecting an accelerator that satisfies the characteristics of the predicted bit stream includes determining the accelerator based on the capacity on the accelerator required by the predicted bit stream.

[0130] 12. The computing device as described in technical solution 1, wherein the computing engine is further used to determine the characteristics of each of the plurality of accelerators in the system.

[0131] 13. One or more machine-readable storage media, comprising a plurality of instructions stored thereon, which, when executed by a computing device, cause the computing device to:

[0132] Determine one or more bit streams registered on each of the multiple accelerators;

[0133] The prediction is that the application of at least one of a plurality of computational skateboards will be requested to accelerate the next job.

[0134] Predict the next bitstream from the bitstream library that will be accelerated in order to execute the request;

[0135] Determine whether the predicted bit stream has already been registered on one of the accelerators;

[0136] In response to determining that the predicted bit stream is not registered in one of the accelerators, an accelerator that satisfies the characteristics of the predicted bit stream is selected from a plurality of accelerators; and

[0137] In response to an accelerator that has been determined to satisfy the characteristics of the predicted bit stream, the predicted bit stream is registered on the determined accelerator.

[0138] 14. One or more machine-readable storage media as described in claim 13, wherein determining one or more bit streams registered on each accelerator includes monitoring bit stream submission and execution on each accelerator.

[0139] 15. One or more machine-readable storage media as described in claim 13, wherein determining one or more bit streams registered on each accelerator includes receiving bit stream registration data from each accelerator, wherein the bit stream registration data indicates the bit stream currently registered on the corresponding accelerator.

[0140] 16. One or more machine-readable storage media as described in claim 13, wherein determining one or more bit streams registered on each accelerator includes updating the bit stream library to track the bit streams currently registered on each accelerator.

[0141] 17. One or more machine-readable storage media as described in claim 13, wherein predicting the next job to be requested for acceleration includes predicting the next job to be requested for acceleration based on available applications currently executing on multiple computing spools.

[0142] 18. One or more machine-readable storage media as described in claim 13, wherein predicting the next job to be requested for acceleration includes predicting the next job to be requested for acceleration based on the execution mode of the bit stream for each available application currently being executed on multiple computing spools.

[0143] 19. One or more machine-readable storage media as described in technical solution 13, wherein predicting a bitstream from a bitstream library includes predicting the bitstream based on available accelerators of the system.

[0144] 20. One or more machine-readable storage media as described in technical solution 13, wherein predicting bitstreams from a bitstream library includes predicting bitstreams based on the type of the next job predicted and the type of workload that each accelerator can accelerate.

[0145] 21. One or more machine-readable storage media as described in technical solution 13, wherein predicting bitstreams from a bitstream library includes predicting the execution mode of the bitstreams for each available application.

[0146] 22. One or more machine-readable storage media as described in claim 13, wherein selecting an accelerator that satisfies the characteristics of the predicted bit stream includes determining the accelerator based on the specific accelerator capabilities required by the predicted bit stream.

[0147] 23. One or more machine-readable storage media as described in claim 13, wherein selecting an accelerator that satisfies the characteristics of the predicted bit stream includes determining the accelerator based on the capacity on the accelerator required by the predicted bit stream.

[0148] 24. One or more machine-readable storage media as described in claim 13, wherein the plurality of instructions, when executed, further cause a computing device to determine the characteristics of each of the plurality of accelerators of the system.

[0149] 25. A computing device for an accelerator in a plurality of accelerators in a pre-configured system, the computing device comprising:

[0150] Circuitry for determining one or more bit streams registered on each of a plurality of accelerators;

[0151] A device for predicting the next task that an application will request to accelerate from at least one of a plurality of computational skateboards;

[0152] A means for predicting the next bit stream from the bit stream library that is to be accelerated for the execution request;

[0153] A device for determining whether the predicted bit stream has been registered on one of the accelerators;

[0154] A means for selecting from a plurality of accelerators an accelerator that satisfies the characteristics of a predicted bit flow in response to determining that the predicted bit flow is not registered in one of the accelerators; and

[0155] A means for storing a predicted bit flow on a determined accelerator in response to a determination of the characteristics of the predicted bit flow.

[0156] 26. A method for pre-configuring an accelerator among multiple accelerators in a system, the method comprising:

[0157] The orchestrator server determines one or more bitstreams registered on each of the multiple accelerators;

[0158] The orchestrator server predicts that an application will be requested to accelerate the next job from at least one of a plurality of computational skateboards.

[0159] The orchestrator server predicts the next bitstream of the job to be accelerated from the bitstream library.

[0160] The orchestrator server determines whether the predicted bitstream has already been registered on one of the accelerators;

[0161] An accelerator server selects from multiple accelerators the accelerators that meet the characteristics of the predicted bitstream in response to determining that the predicted bitstream is not registered on one of the accelerators; and

[0162] The orchestrator server stores the predicted bitstream on the determined accelerator in response to determining the characteristics of the predicted bitstream.

[0163] 27. The method of claim 26, wherein determining one or more bit streams registered on each accelerator includes receiving bit stream registration data from each accelerator by an orchestrator server, wherein the bit stream registration data indicates the bit stream currently registered on the corresponding accelerator.

[0164] 28. The method of claim 26, wherein predicting the next job to be requested for acceleration includes the orchestrator server predicting the next job to be requested for acceleration based on available applications currently executing on multiple computing spools.

[0165] Example

[0166] The following provides illustrative examples of the techniques disclosed herein. Embodiments of the techniques may include any one or more of the examples described below, as well as any combination thereof.

[0167] Example 1 includes a computing device for accelerators in a plurality of accelerators of a pre-configured system, the computing device including communication circuitry; a computing engine configured to (i) determine one or more bit streams registered on each of the plurality of accelerators, (ii) predict an application request from at least one computing slide of the plurality of computing slides for the next job to be accelerated, (iii) predict a bit stream from a bit stream library for the predicted next job to be accelerated, (iv) determine whether the predicted bit stream has already been registered on one of the accelerators, (v) in response to determining that the predicted bit stream has not been registered on one of the accelerators, select an accelerator from the plurality of accelerators that satisfies the characteristics of the predicted bit stream, and (vi) in response to determining that an accelerator satisfies the characteristics of the predicted bit stream, register the predicted bit stream on the determined accelerator.

[0168] Example 2 includes the subject of Example 1, and wherein identifying one or more bit streams registered on each accelerator includes monitoring bit stream submissions and executions on each accelerator.

[0169] Example 3 includes the subject of either Example 1 or 2, and wherein determining one or more bit streams registered on each accelerator includes receiving bit stream registration data from each accelerator, wherein the bit stream registration data indicates the bit stream currently registered on the corresponding accelerator.

[0170] Example 4 includes the subject of any one of Examples 1-3, and wherein identifying one or more bit streams registered on each accelerator includes updating the bit stream library to track the bit streams currently registered on each accelerator.

[0171] Example 5 includes the subject of any of Examples 1-4, and the prediction of the next job to be requested for acceleration includes predicting the next job to be requested for acceleration based on the available applications currently running on multiple computing skateboards.

[0172] Example 6 includes the subject of any of Examples 1-5, and wherein predicting the next job to be requested for acceleration includes predicting the next job to be requested for acceleration based on the execution mode of the bitstream for each available application currently being executed on multiple computing skateboards.

[0173] Example 7 includes the subject of any of Examples 1-6, and wherein predicting the next work to be requested for acceleration based on the execution pattern of the bitstream for each available application includes determining the past execution history for each bitstream for each available application.

[0174] Example 8 includes the subject of any of Examples 1-7, and wherein predicting the next work to be requested for acceleration based on the execution pattern of the bitstream for each available application includes using machine learning to predict the execution pattern.

[0175] Example 9 includes the subject of any of Examples 1-8, and wherein predicting bitstreams from a bitstream library includes predicting bitstreams based on the available accelerators available in the system.

[0176] Example 10 includes the subject of any of Examples 1-9, and wherein predicting bitstreams from a bitstream library involves predicting bitstreams based on the type of the next job predicted and the type of workload that each accelerator can accelerate.

[0177] Example 11 includes the subject of any of Examples 1-10, and wherein predicting bitstreams from a bitstream library includes predicting the execution mode of the bitstreams for each available application.

[0178] Example 12 includes the subject of any one of Examples 1-11, and wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream includes determining the accelerator based on the specific accelerator capabilities required by the predicted bitstream.

[0179] Example 13 includes the subject of any one of Examples 1-12, and wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream includes determining the accelerator based on the capacity on the accelerator required by the predicted bitstream.

[0180] Example 14 includes the subject of any one of Examples 1-13, and the computational engine is further used to determine the characteristics of each of the multiple accelerators in the system.

[0181] Example 15 includes the subject of any one of Examples 1-14, and wherein determining the characteristics of each accelerator of the system includes determining the availability of each accelerator.

[0182] Example 16 includes the topic of any of Examples 1-15, and wherein determining the availability of each accelerator includes determining the current load on each accelerator.

[0183] Example 17 includes the subject of any of Examples 1-16, and wherein determining the availability of each accelerator includes determining the number of free slots in each accelerator.

[0184] Example 18 includes the subject of any of Examples 1-17, and wherein determining the availability of each accelerator includes determining the number of free logic gates in each accelerator.

[0185] Example 19 includes the subject of any one of Examples 1-18, and wherein determining the characteristics of each accelerator of the system includes determining the type of workload that each accelerator is capable of accelerating.

[0186] Example 20 includes the subject of any one of Examples 1-19, and the type of workload that each accelerator can accelerate includes determining the cryptographic, compression, encoding transformation, matrix multiplication and / or convolutional neural network computing capabilities of each accelerator.

[0187] Example 21 includes the subject of any of Examples 1-20, and wherein determining the characteristics of each accelerator of the system includes determining the physical distance from the requested computational skateboard to each accelerator.

[0188] Example 22 includes the subject of any one of Examples 1-21, and wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream includes determining the accelerator based on the characteristics of each accelerator.

[0189] Example 23 includes a method for pre-configuring accelerators in a plurality of accelerators in a system, the method comprising: determining, by an orchestrator server, one or more bitstreams registered on each of the plurality of accelerators; predicting, by the orchestrator server, the next job to be requested for acceleration from at least one of the plurality of compute skateboards; predicting, by the orchestrator server, the bitstream from a bitstream library to execute the predicted next job to be accelerated; determining, by the orchestrator server, whether the predicted bitstream is already registered on one of the accelerators; selecting, by the orchestrator server and in response to determining that the predicted bitstream is not registered on one of the accelerators, an accelerator satisfying the characteristics of the predicted bitstream from the plurality of accelerators; and registering the predicted bitstream on the determined accelerator by the orchestrator server and in response to determining that the accelerator satisfies the characteristics of the predicted bitstream.

[0190] Example 24 includes the subject of Example 23, and wherein identifying one or more bitstreams hosted on each accelerator includes monitoring bitstream submissions and executions on each accelerator by an orchestrator server.

[0191] Example 25 includes the subject of any one of Examples 23 and 24, and wherein determining one or more bit streams registered on each accelerator includes receiving bit stream registration data from each accelerator by an orchestrator server, wherein the bit stream registration data indicates the bit stream currently registered on the corresponding accelerator.

[0192] Example 26 includes the subject of any of Examples 23-25, and wherein determining one or more bitstreams registered on each accelerator includes updating the bitstream library by the orchestrator server to track the bitstreams currently registered on each accelerator.

[0193] Example 27 includes the subject of any of Examples 23-26, and wherein predicting the next job to be requested for acceleration includes the orchestrator server predicting the next job to be requested for acceleration based on the available applications currently running on multiple computing skateboards.

[0194] Example 28 includes the subject of any of Examples 23-27, and wherein predicting the next job to be requested for acceleration includes the orchestrator server predicting the next job to be requested for acceleration based on the execution pattern of the bitstream for each available application currently being executed on multiple computing skateboards.

[0195] Example 29 includes the subject of any of Examples 23-28, and wherein predicting the next job to be requested for acceleration based on the execution pattern of the bitstream for each available application includes the orchestrator server determining the past execution history of each bitstream for each available application.

[0196] Example 30 includes the subject of any of Examples 23-29, and wherein predicting the next job to be requested for acceleration based on the execution pattern of the bitstream for each available application includes the orchestrator server using machine learning to predict the execution pattern.

[0197] Example 31 includes the subject of any of Examples 23-30, and wherein predicting the bitstream from the bitstream library includes the bitstream being predicted by the orchestrator server based on the available accelerators of the system.

[0198] Example 32 includes the subject of any of Examples 23-31, and wherein the prediction of bitstreams from the bitstream library includes the prediction of bitstreams by the orchestrator server based on the type of the next job predicted and the type of workload that each accelerator can accelerate.

[0199] Example 33 includes the subject of any of Examples 23-32, and wherein the prediction of bitstreams from the bitstream library includes the execution mode of the bitstreams predicted by the orchestrator server for each available application.

[0200] Example 34 includes the subject of any of Examples 23-33, and wherein the selection of an accelerator that satisfies the characteristics of the predicted bitstream includes the orchestrator server determining the accelerator based on the specific accelerator capabilities required by the predicted bitstream.

[0201] Example 35 includes the subject of any of Examples 23-34, and wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream includes the orchestrator server determining the accelerator based on the capacity on the accelerator required by the predicted bitstream.

[0202] Example 36 includes the subject of any one of Examples 23-35, and further includes the characteristics of each of the multiple accelerators in the system determined by the orchestrator server.

[0203] Example 37 includes the subject of any one of Examples 23-36, and wherein determining the characteristics of each accelerator in the system includes determining the availability of each accelerator by the orchestrator server.

[0204] Example 38 includes the topic of any of Examples 23-37, and wherein determining the availability of each accelerator includes determining the current load on each accelerator by the orchestrator server.

[0205] Example 39 includes the topic of any of Examples 23-38, and wherein determining the availability of each accelerator includes the orchestrator server determining the number of free slots in each accelerator.

[0206] Example 40 includes the subject of any of Examples 23-39, and wherein determining the availability of each accelerator includes the orchestrator server determining the number of free logic gates in each accelerator.

[0207] Example 41 includes the subject of any of Examples 23-40, and wherein determining the characteristics of each accelerator of the system includes determining the type of workload that each accelerator can accelerate by the orchestrator server.

[0208] Example 42 includes the subject of any one of Examples 23-41, and the type of workload that each accelerator can accelerate includes the cryptographic, compression, encoding transformation, matrix multiplication and / or convolutional neural network computing capabilities of each accelerator determined by the orchestrator server.

[0209] Example 43 includes the subject of any of Examples 23-42, and wherein determining the characteristics of each accelerator in the system includes determining the physical distance from the requested computational skateboard to each accelerator by the orchestrator server.

[0210] Example 44 includes the subject of any one of Examples 23-43, and wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream includes the orchestrator server determining the accelerator based on the characteristics of each accelerator.

[0211] Example 45 includes one or more machine-readable storage media containing a plurality of instructions stored thereon, the plurality of instructions being executed to cause a computing device to perform any of the methods in Examples 23-44.

[0212] Example 46 includes a computing device comprising means for performing the method of any of Examples 23-44.

[0213] Example 47 includes a computing device for accelerators in a plurality of accelerators of a pre-configured system, the computing device including bitstream updater circuitry for determining one or more bitstreams registered on each of the plurality of accelerators; job predictor circuitry for predicting the next job to be requested for acceleration from at least one of the plurality of computing slides; bitstream predictor circuitry for predicting the bitstream from a bitstream library for executing the predicted next job to be accelerated; and accelerator manager circuitry for (i) determining whether the predicted bitstream has been registered on one of the accelerators, (ii) selecting from the plurality of accelerators an accelerator that satisfies the characteristics of the predicted bitstream in response to determining that the predicted bitstream has not been registered on one of the accelerators, and (iii) registering the predicted bitstream on the determined accelerator in response to determining that the accelerator satisfies the characteristics of the predicted bitstream.

[0214] Example 48 includes the subject of Example 47, and wherein identifying one or more bit streams registered on each accelerator includes monitoring bit stream submissions and executions on each accelerator.

[0215] Example 49 includes the subject of any one of Examples 47 and 48, and wherein determining one or more bit streams registered on each accelerator includes receiving bit stream registration data from each accelerator, wherein the bit stream registration data indicates the bit stream currently registered on the corresponding accelerator.

[0216] Example 50 includes the subject of any of Examples 47-49, and wherein identifying one or more bit streams registered on each accelerator includes updating the bit stream library to track the bit streams currently registered on each accelerator.

[0217] Example 51 includes the subject of any of Examples 47-50, and the prediction of the next job to be requested for acceleration includes predicting the next job to be requested for acceleration based on the available applications currently running on multiple computing skateboards.

[0218] Example 52 includes the subject of any of Examples 47-51, and wherein predicting the next job to be requested for acceleration includes predicting the next job to be requested for acceleration based on the execution mode of the bitstream for each available application currently being executed on multiple computing skateboards.

[0219] Example 53 includes the subject of any of Examples 47-52, and wherein predicting the next work to be requested for acceleration based on the execution pattern of the bitstream for each available application includes determining the past execution history for each bitstream for each available application.

[0220] Example 54 includes the subject of any of Examples 47-53, and wherein predicting the next job to be requested for acceleration based on the execution pattern of the bitstream for each available application includes using machine learning to predict the execution pattern.

[0221] Example 55 includes the subject of any of Examples 47-54, and wherein predicting bitstreams from a bitstream library includes predicting bitstreams based on the available accelerators of the system.

[0222] Example 56 includes the subject of any of Examples 47-55, and wherein predicting the bitstream from the bitstream library includes predicting the bitstream based on the type of the next job predicted and the type of workload that each accelerator can accelerate.

[0223] Example 57 includes the subject of any of Examples 47-56, and wherein predicting bitstreams from a bitstream library includes predicting the execution mode of the bitstreams for each available application.

[0224] Example 58 includes the subject of any of Examples 47-57, and wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream includes determining the accelerator based on the specific accelerator capabilities required by the predicted bitstream.

[0225] Example 59 includes the subject of any of Examples 47-58, and wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream includes determining the accelerator based on the capacity on the accelerator required by the predicted bitstream.

[0226] Example 60 includes the subject of any of Examples 47-59, and in which the computational engine is further used to determine the characteristics of each of the multiple accelerators in the system.

[0227] Example 61 includes the subject of any of Examples 47-60, and wherein determining the characteristics of each accelerator of the system includes determining the availability of each accelerator.

[0228] Example 62 includes the subject of any of Examples 47-61, and wherein determining the availability of each accelerator includes determining the current load on each accelerator.

[0229] Example 63 includes the subject of any of Examples 47-62, and wherein determining the availability of each accelerator includes determining the number of free slots in each accelerator.

[0230] Example 64 includes the subject of any of Examples 47-63, and wherein determining the availability of each accelerator includes determining the number of free logic gates in each accelerator.

[0231] Example 65 includes the subject of any of Examples 47-64, and wherein determining the characteristics of each accelerator of the system includes determining the type of workload that each accelerator can accelerate.

[0232] Example 66 includes the subject of any of Examples 47-65, and wherein determining the type of workload that each accelerator can accelerate includes determining the cryptographic, compression, encoding transformation, matrix multiplication, and / or convolutional neural network computational capabilities of each accelerator.

[0233] Example 67 includes the subject of any of Examples 47-66, and wherein determining the characteristics of each accelerator of the system includes determining the physical distance from the requested computational skateboard to each accelerator.

[0234] Example 68 includes the subject of any of Examples 47-67, and wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream includes determining the accelerator based on the characteristics of each accelerator.

[0235] Example 69 includes a computing device for an accelerator in a plurality of accelerators of a pre-configured system, the computing device including circuitry for determining one or more bit streams registered on each of the plurality of accelerators; means for predicting the next job to be accelerated from an application request for acceleration from at least one of the plurality of computing slides; means for predicting the predicted bit stream from a bit stream library to execute the request to be accelerated for the next job; means for determining whether the predicted bit stream has already been registered on one of the accelerators; means for selecting an accelerator from the plurality of accelerators that satisfies the characteristics of the predicted bit stream in response to determining that the predicted bit stream has not been registered on one of the accelerators; and means for registering the predicted bit stream on the determined accelerator in response to determining that the accelerator satisfies the characteristics of the predicted bit stream.

[0236] Example 70 includes the subject of Example 69, and wherein the circuitry for determining one or more bit streams registered on each accelerator includes circuitry for monitoring bit stream submissions and executions on each accelerator.

[0237] Example 71 includes the subject of any one of Examples 69 and 70, and wherein the means for determining one or more bit streams registered on each accelerator includes means for receiving bit stream registration data from each accelerator, wherein the bit stream registration data indicates the bit stream currently registered on the corresponding accelerator.

[0238] Example 72 includes the subject of any of Examples 69-71, and wherein the means for determining one or more bit streams registered on each accelerator includes means for updating the bit stream library to track the bit streams currently registered on each accelerator.

[0239] Example 73 includes the subject of any of Examples 69-72, and wherein the means for predicting the next job to be requested for acceleration includes means for predicting the next job to be requested for acceleration based on available applications currently executing on multiple computing spools.

[0240] Example 74 includes the subject of any of Examples 69-73, and wherein the means for predicting the next job to be requested for acceleration includes means for predicting the next job to be requested for acceleration based on the execution mode of the bitstream for each available application currently being executed on multiple computing spools.

[0241] Example 75 includes the subject of any of Examples 69-74, and wherein the means for predicting the next job to be requested for acceleration based on the execution mode of the bit stream for each available application includes means for determining the past execution history of each bit stream for each available application.

[0242] Example 76 includes the subject of any of Examples 69-75, and wherein the means for predicting the next job to be requested for acceleration based on the execution pattern of the bitstream for each available application includes means for using machine learning to predict the execution pattern.

[0243] Example 77 includes the subject of any of Examples 69-76, and wherein the means for predicting bitstreams from a bitstream library includes means for predicting bitstreams based on available accelerators of the system.

[0244] Example 78 includes the subject of any of Examples 69-77, and wherein the means for predicting bitstreams from a bitstream library includes means for predicting bitstreams based on the type of the next job predicted and the type of workload that each accelerator can accelerate.

[0245] Example 79 includes the subject of any of Examples 69-78, and wherein the means for predicting bitstreams from a bitstream library includes means for predicting the execution mode of the bitstream for each available application.

[0246] Example 80 includes the subject of any of Examples 69-79, and wherein the means for selecting an accelerator that satisfies the characteristics of the predicted bitstream includes means for determining the accelerator based on the specific accelerator capabilities required by the predicted bitstream.

[0247] Example 81 includes the subject of any of Examples 69-80, and wherein the means for selecting an accelerator that satisfies the characteristics of the predicted bit stream includes means for determining the accelerator based on the capacity on the accelerator required by the predicted bit stream.

[0248] Example 82 includes the subject matter of any one of Examples 69-81, and further includes means for determining the characteristics of each of the plurality of accelerators in the system.

[0249] Example 83 includes the subject of any one of Examples 69-82, and wherein the means for determining the characteristics of each accelerator of the system includes means for determining the availability of each accelerator.

[0250] Example 84 includes the subject of any one of Examples 69-83, and wherein the means for determining the availability of each accelerator includes means for determining the current load on each accelerator.

[0251] Example 85 includes the subject of any one of Examples 69-84, and wherein the means for determining the availability of each accelerator includes means for determining the number of free slots in each accelerator.

[0252] Example 86 includes the subject of any one of Examples 69-85, and wherein the means for determining the availability of each accelerator includes means for determining the number of free logic gates in each accelerator.

[0253] Example 87 includes the subject of any one of Examples 69-86, and wherein the means for determining the characteristics of each accelerator of the system includes means for determining the type of workload that each accelerator is capable of accelerating.

[0254] Example 88 includes the subject of any one of Examples 69-87, and wherein the means for determining the type of workload that each accelerator can accelerate includes means for determining the cryptographic, compression, encoding transformation, matrix multiplication and / or convolutional neural network computing capabilities of each accelerator.

[0255] Example 89 includes the subject of any one of Examples 69-88, and wherein the means for determining the characteristics of each accelerator of the system includes means for determining the physical distance from the requested computational slide to each accelerator.

[0256] Example 90 includes the subject of any one of Examples 69-89, and wherein the means for selecting an accelerator that satisfies the characteristics of the predicted bitstream includes means for determining an accelerator based on the characteristics of each accelerator.

Claims

1. A computing device for an accelerator in a plurality of accelerators in a pre-configured system, the computing device comprising: Communication circuits; A computing engine is configured to: (i) determine one or more bit streams registered on each of a plurality of accelerators; (ii) predict the next job to be accelerated from at least one of a plurality of computing slides; (iii) predict the bit stream from a bit stream library to execute the predicted next job to be accelerated; (iv) determine whether the predicted bit stream has already been registered on one of the accelerators; (v) in response to determining that the predicted bit stream has not been registered on one of the accelerators, select an accelerator from the plurality of accelerators that satisfies the characteristics of the predicted bit stream; and (vi) in response to determining that an accelerator satisfies the characteristics of the predicted bit stream, register the predicted bit stream on the determined accelerator.

2. The computing device of claim 1, wherein determining one or more bit streams registered on each accelerator includes monitoring bit stream submission and execution on each accelerator.

3. The computing device of claim 1, wherein determining one or more bit streams registered on each accelerator includes receiving bit stream registration data from each accelerator, wherein the bit stream registration data indicates the bit stream currently registered on the corresponding accelerator.

4. The computing device of claim 1, wherein determining one or more bit streams registered on each accelerator includes updating the bit stream library to track the bit streams currently registered on each accelerator.

5. The computing device of claim 1, wherein predicting the next job to be requested for acceleration comprises predicting the next job to be requested for acceleration based on available applications currently executing on a plurality of computing spools.

6. The computing device of claim 1, wherein predicting the next job to be requested for acceleration comprises predicting the next job to be requested for acceleration based on the execution mode of the bitstream for each available application currently executing on a plurality of computing spools.

7. The computing device of claim 1, wherein predicting a bitstream from a bitstream library comprises predicting the bitstream based on the available accelerators of the system.

8. The computing device of claim 1, wherein predicting bitstreams from the bitstream library comprises predicting bitstreams based on the type of the next job predicted and the type of workload that each accelerator can accelerate.

9. The computing device of claim 1, wherein predicting bitstreams from a bitstream library includes predicting the execution mode of the bitstreams for each available application.

10. The computing device of claim 1, wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream comprises determining the accelerator based on the accelerator capabilities required by the predicted bitstream.

11. The computing device of claim 1, wherein selecting an accelerator that satisfies the characteristics of the predicted bitstream comprises determining the accelerator based on the capacity on the accelerator required by the predicted bitstream.

12. The computing device of claim 1, wherein the computing engine is further used to determine the characteristics of each of the plurality of accelerators in the system.

13. A computing device for an accelerator in a plurality of accelerators in a pre-configured system, the computing device comprising: Circuitry for determining one or more bit streams registered on each of a plurality of accelerators; A device for predicting the next task that an application will request to accelerate from at least one of a plurality of computational skateboards; A means for predicting the next bit stream from the bit stream library that is to be accelerated for the execution request; A device for determining whether the predicted bit stream has been registered on one of the accelerators; A means for selecting from a plurality of accelerators an accelerator that satisfies the characteristics of the predicted bit flow in response to determining that the predicted bit flow is not registered in one of the accelerators. as well as A means for storing a predicted bit flow on a determined accelerator in response to a determination of the characteristics of the predicted bit flow.

14. The computing device of claim 13, wherein the circuitry for determining one or more bit streams registered on each accelerator includes circuitry for monitoring bit stream submissions and executions on each accelerator.

15. The computing device of claim 13, wherein the means for determining one or more bit streams registered on each accelerator includes means for receiving bit stream registration data from each accelerator, wherein the bit stream registration data indicates a bit stream currently registered on the corresponding accelerator.

16. The computing device of claim 13, wherein the means for determining one or more bit streams registered on each accelerator includes means for updating the bit stream library to track the bit streams currently registered on each accelerator.

17. The computing device of claim 13, wherein the means for predicting the next job to be requested for acceleration includes means for predicting the next job to be requested for acceleration based on available applications currently executing on a plurality of computing spools.

18. The computing device of claim 13, wherein the means for predicting the next job to be requested for acceleration includes means for predicting the next job to be requested for acceleration based on the execution mode of the bitstream for each available application currently being executed on a plurality of computing spools.

19. The computing device of claim 13, wherein the means for predicting bitstreams from a bitstream library includes means for predicting bitstreams based on available accelerators of the system.

20. The computing device of claim 13, wherein the means for predicting bitstreams from the bitstream library includes means for predicting bitstreams based on the type of the next job predicted and the type of workload that each accelerator can accelerate.

21. The computing device of claim 13, wherein the means for predicting bitstreams from a bitstream library includes means for predicting an execution mode for a bitstream for each available application.

22. The computing device of claim 13, wherein the means for selecting an accelerator that satisfies the characteristics of the predicted bitstream includes means for determining the accelerator based on the accelerator capability required by the predicted bitstream.

23. The computing device of claim 13, wherein the means for selecting an accelerator that satisfies the characteristics of the predicted bit flow includes means for determining the accelerator based on the capacity on the accelerator required by the predicted bit flow.

24. The computing device of claim 13, further comprising means for determining the characteristics of each of the plurality of accelerators in the system.

25. A method for pre-configuring an accelerator among a plurality of accelerators in a system, the method comprising: The orchestrator server determines one or more bitstreams registered on each of the multiple accelerators; The orchestrator server predicts that an application will be requested to accelerate the next job from at least one of a plurality of computational skateboards. The orchestrator server predicts the next bitstream of the job to be accelerated from the bitstream library. The orchestrator server determines whether the predicted bitstream has already been registered on one of the accelerators; The orchestrator server selects an accelerator from a plurality of accelerators that meets the characteristics of the predicted bitstream in response to determining that the predicted bitstream is not registered on one of the accelerators. as well as The orchestrator server stores the predicted bitstream on the determined accelerator in response to determining the characteristics of the predicted bitstream.

26. The method of claim 25, wherein determining one or more bit streams registered on each accelerator comprises receiving bit stream registration data from each accelerator by an orchestrator server, wherein the bit stream registration data indicates a bit stream currently registered on the corresponding accelerator.

27. The method of claim 25, wherein predicting the next job to be requested for acceleration comprises the orchestrator server predicting the next job to be requested for acceleration based on available applications currently executing on multiple computing skateboards.

28. A computer-readable medium having instructions stored thereon, which, when executed, cause a computing device to perform the method according to any one of claims 25-27.

Citation Information

Patent Citations

  • Techniques for preconfiguring accelerator by predicting bitstream

    CN114090478A

  • Resource allocation in job scheduling environment

    US20150277987A1

  • Parallel processing in hardware accelerators communicably coupled with a processor

    US20160132329A1