Hardware accelerator service aggregation
The system aggregates acceleration services from local and remote hardware accelerator cards using a standardized identifier system, addressing the limitations of existing systems in discovering and utilizing hardware accelerator card capabilities, enabling efficient workload distribution and processing.
Patent Information
- Application Number
- JP2025034672
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-11-12
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-23
AI Technical Summary
Existing systems face challenges in discovering and utilizing the full capabilities of hardware accelerator cards due to limited connections via communication interfaces like PCIe, and operating systems' inability to distinguish new functionalities, leading to restricted access to acceleration services.
A system and method for service aggregation that collects and publishes acceleration services from both local and remote hardware accelerator cards, using a standardized listing of identifiers to enable processors to discover and utilize these services across multiple cards.
Enables effective utilization of all available acceleration services from multiple hardware accelerator cards, overcoming connection and recognition limitations, and facilitating efficient workload distribution and processing.
Smart Images

Figure 2025108409000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application is a continuation of U.S. Patent Application No. 17 / 525,300, filed on November 12, 2021, the disclosure of which is incorporated herein by reference.
Background Art
[0002] Background In most systems, it is difficult for software running on a computing device, including components of the computing device and / or an operating system, to discover the functionality and capabilities provided by a hardware accelerator card connected to the computing device via a communication interface such as a PCIe bus. To avoid these problems, a processor can be hard - coded with software such as a driver to communicate with a specific hardware accelerator card. However, hard - coding a processor with the necessary software to communicate with a specific hardware accelerator card limits the processor to only that specific hardware accelerator card. Thus, the processor cannot utilize the functionality and capabilities of other hardware accelerator cards or hardware accelerator cards developed after the processor was manufactured.
[0003] Furthermore, some hardware accelerator cards may expose the functionality and capabilities of these cards as separate devices within the operating system of a computing device. In this regard, when a hardware accelerator card is connected to a computing device via a communication interface such as a PCIe bus, the operating system detects the connection or is otherwise notified of the connection and can create a list of each functionality and capability of the hardware accelerator card as an individual device within the operating system according to predefined classes and subclasses. Based on the devices listed within the operating system, the computing device can utilize the capabilities and functionality of the hardware accelerator card.
[0004] As the capabilities and functionality of hardware accelerator cards have improved and become more specialized, these new capabilities and functionality are not clearly distinguishable by the classes and subclasses provided by current operating systems. Thus, some operating systems can indicate the capabilities and functionality provided by a hardware accelerator card, but may not be able to identify all of the capabilities and functionality of the hardware accelerator card. Additionally, some of the capabilities and functionality of a hardware accelerator card may not be recognized and / or clearly distinguishable within the operating system. Accordingly, a computing device may not be able to utilize all of the features and capabilities of an available hardware accelerator card and may not even be able to recognize them.
[0005] Systems typically have a limited number of connections to communication interfaces. For example, a system that includes a PCIe bus can have only a few PCIe slots to connect hardware accelerator cards to the PCIe bus. The limited number of connections can be due to cost constraints. In this regard, each additional connection added to the system can increase the cost of physical hardware and potentially increase the overall manufacturing cost. Additionally, technical limitations such as power availability can also limit the number of devices that can be connected to the system. For example, a system can include five connections for hardware accelerator cards, but the power supply may only be able to supply power to two hardware accelerator cards simultaneously. Thus, due to the limited number of hardware accelerator cards that can be connected to the system, the system's ability to access the acceleration services provided by the hardware accelerator cards can be limited. SUMMARY OF THE INVENTION
[0006] Summary The technology described in this specification relates to a system and method for service aggregation that collects and publishes acceleration services provided by accelerators of a hardware accelerator card. Using service aggregation, a hardware accelerator card can communicate with other hardware accelerator cards and collect and publish acceleration services provided by the accelerators of these other hardware accelerator cards that can be locally or remotely connected. The collected and published acceleration services can also include the acceleration services provided by the accelerator of the hardware accelerator card by executing service aggregation. Subsequently, the hardware accelerator card, as well as the acceleration services provided by other locally or remotely connected hardware accelerator cards, can be utilized by the system.
[0007] One aspect of the present disclosure relates to a method. The method includes providing, by one or more processors of a local hardware (HW) accelerator card, via a communication interface, a listing of acceleration services from the local HW accelerator card, the listing of acceleration services including a first set of acceleration services provided by one or more accelerators of the local HW accelerator card and a second set of acceleration services provided by one or more accelerators of a remote HW accelerator card; receiving, by the one or more processors, a workload instruction from a processor of a computing device, the workload instruction defining a workload for processing by at least one of the acceleration services of the second set of acceleration services; and transferring, by the one or more processors, the workload instruction to the remote HW accelerator card.
[0008] Another aspect of the present disclosure relates to a system comprising a communication interface and a local hardware (HW) accelerator card including one or more processors and one or more accelerators. The one or more processors are configured to receive, via the communication interface, a listing of acceleration services from the local HW accelerator card, the listing of acceleration services including a first set of acceleration services provided by one or more accelerators of the local HW accelerator card and a second set of acceleration services provided by one or more accelerators of a remote HW accelerator card; receive a workload instruction from a processor of a computing device, the workload instruction defining a workload for processing by at least one of the acceleration services of the second set of acceleration services; and transfer the workload instruction to the remote HW accelerator card.
[0009] Another aspect of the present disclosure is a non-transitory tangible computer-readable medium storing computer-readable instructions of a program A computer-readable storage medium, the instructions, when executed by one or more computing devices, cause the one or more computing devices to perform a method. The method includes providing, by one or more processors of a local hardware (HW) accelerator card, via a communication interface, a listing of acceleration services from the local HW accelerator card, the listing of acceleration services including a first set of acceleration services provided by one or more accelerators of the local HW accelerator card and a second set of acceleration services provided by one or more accelerators of a remote HW accelerator card; receiving, by the one or more processors, a workload instruction from a processor of the computing device, the workload instruction defining a workload for processing by at least one of the acceleration services of the second set of acceleration services; and transferring, by the one or more processors, the workload instruction to the remote HW accelerator card.
[0010] In some examples, a processed workload may be received from the remote HW accelerator card, the processed workload being the workload after being processed by at least one of the acceleration services of the second set of acceleration services.
[0011] Optionally, the processed workload may be transferred to the processor of the computing device. In some examples, the listing of acceleration services is generated by an accelerated service manager (ASM) running on one or more processors.
[0012] Optionally, the ASM running on one or more processors communicates with another ASM running on the remote HW accelerator card.
[0013] In some cases, the ASM transfers the workload instruction to another ASM. In some cases, the ASM requests a listing of a second set of acceleration services from another ASM.
[0014] In some cases, the ASM identifies and removes an unsound acceleration service from the listing of acceleration services.
[0015] In some examples, identifying an unsound acceleration service includes the ASM determining that at least one of the acceleration services of the second set of acceleration services fails to process a workload instruction.
[0016] In some examples, removing an unsound acceleration service includes marking at least one of the acceleration services of the second set of acceleration services as unsound or removing at least one of the acceleration services of the second set of acceleration services from the listing of acceleration services.
[0017] In some cases, after determining that a workload instruction fails to be processed, the ASM sends the updated workload instruction to a different HW accelerator card for processing by at least one of the acceleration services of a different HW accelerator.
[0018] In some cases, the workload instruction is further defined to be processed by at least one acceleration service of at least one other remote HW accelerator card.
Brief Description of the Drawings
[0019]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
DETAILED DESCRIPTION OF THE INVENTION
[0020] Detailed Description The technology described herein relates to a system and method for service aggregation that collects and publishes acceleration services provided by accelerators of a hardware (HW) accelerator card. Using service aggregation, an HW accelerator card installed in a computing device can supply and publish acceleration services from local and remotely connected HW accelerator cards to a computing device. As further described herein, software executing on a computing core of an HW accelerator card in a computing device can discover or use a preconfiguration file to communicate with other HW accelerator cards locally connected to the computing device via a communication interface and / or other HW accelerator cards that are network-connected remotely to the computing device. The computing device can then communicate with other HW accelerator cards through the HW accelerator card installed in the computing device. Thus, the computing device can effectively utilize the acceleration services of any number of locally or remotely connected HW accelerator cards.
[0021] To overcome the lack of discovery of acceleration services, the technology described herein uses a standardized listing of identifiers corresponding to acceleration services that can be provided by accelerators on an HW accelerator card. In this regard, each HW accelerator card can store a listing of identifiers corresponding to the acceleration services provided by the accelerators on the card. Since the identifiers can provide a finer granularity than the currently used device classes and subclasses, a processor that retrieves the listing from the HW accelerator card can determine and utilize more accelerator services provided by the accelerators on the HW accelerator card.
[0022] As used herein, the term "acceleration service" refers to the capabilities and functionality provided by the accelerators of HW accelerator cards. The HW reference to the "acceleration service" of an accelerator card refers to the acceleration service of the accelerator on that HW accelerator card. An acceleration service may include capabilities and functionality that an accelerator can utilize to control the processing of data, which is referred to herein as the control plane acceleration service. An acceleration service may also include capabilities and functionality that an accelerator can utilize to process data, which is referred to herein as the data plane acceleration service. For example, an accelerator can assist in providing an acceleration service that provides control and / or policies for sharing memory between the memory on a host (computing device) and the accelerator. This control plane acceleration service can be identified and communicated as an acceleration service.
[0023] Since each HW accelerator card can have multiple accelerators, each HW accelerator can provide multiple acceleration services with the same capabilities and functionality and / or different capabilities and functionality. Further, each accelerator can include multiple functions and capabilities.
[0024] Example of a System FIG. 1 shows an example of the architecture of a computing device 110 in which the features described herein may be implemented. This example should not be considered as limiting the scope of the disclosure or usefulness of the features described herein. The computing device 110 can be a server, a personal computer, or other such system. The architecture of the computing device 110 includes a processor 112, a memory 114, and a hardware accelerator card 118.
[0025] Processor 112 can include one or more general-purpose processors, such as a central processing unit (CPU), and / or one or more dedicated processors, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). Processor 112 can be of any type that includes one or more microprocessors (uP), one or more microcontrollers (uC), one or more digital signal processors, or any combination thereof, without being limited thereto. The processor can include one or more levels of caching, one or more processor cores, and one or more registers. Each processor core can include an arithmetic logic unit (ALU), a floating-point arithmetic unit (FPU), a digital signal processing core (DSP core), or any combination thereof. Processor 112 can be configured to execute computer-readable program instructions that can be included in a data storage device, such as instructions 117 stored in memory 114 and / or other instructions described herein.
[0026] Memory 114 can store information accessible by processor 112, including instructions 117 that can be executed by processor 112. The memory can also include data 116 that can be obtained, manipulated, or stored by processor 112. Memory 114 can be a type of non-transitory computer-readable medium that can store information accessible by processor 112, such as a hard drive, a solid-state drive, a tape drive, an optical storage device, a memory card, a ROM, a RAM, a DVD, a CD-ROM, a writable memory, a read-only memory, etc.
[0027] Instruction 117 can be a set of instructions such as machine code directly executed by processor 112, or a set of instructions such as a script indirectly executed by processor 112. In this regard, the terms "instruction", "step", and "program" can be used interchangeably in this specification. Instruction 117 can be in object code form for direct processing by processor 112, or in another type of computer language including a script or collection of independent source code modules that are interpreted on demand or pre-compiled in advance.
[0028] Data 116 can be obtained, stored, or modified by processor 112 in accordance with instruction 117 or other such instructions. For example, although the system and method are not limited by a particular data structure, data 116 can be stored in computer registers within a distributed storage system, in a structure having a plurality of different fields and records, or documents, or as a buffer. Data 116 can be formatted in a computer-readable form such as, but not limited to, binary, ASCII, or Unicode. Further, data 116 can include information sufficient to identify numbers, descriptive text, unique codes, pointers, references to data stored in other memories including other network locations, or other related information, or information used by functions to calculate related data.
[0029] The computing device can further include a hardware (HW) accelerator card 118. The hardware accelerator card 118 can be any device configured to efficiently process a particular type of task. Some examples of HW accelerator cards include a network accelerator card, a video transcoding accelerator card, a security function accelerator card, a cryptographic accelerator card, an acoustic processing accelerator card, an artificial intelligence accelerator card, and the like. Each of these HW accelerator cards can be configured to provide specific acceleration services such as compression, encryption, code conversion, hash generation, graphic processing, simulation, and the like. Some HW accelerator cards can be configured to provide multiple acceleration services such as compression and encryption, or any other combination of acceleration services.
[0030] The computing device 110 can also include a network interface card 119. The network interface card can be any device capable of communicating directly and indirectly with other nodes of a network such as the network 470 described herein with reference to FIG. 4.
[0031] FIG. 1 functionally illustrates the processor 112, the memory 114, the HW accelerator card 118, and the network interface 119 as being within the same block, but the processor 112, the memory 114, the HW accelerator card 118, and the network interface 119 may or may not be housed within the same physical housing. For example, some of the instructions 117 and data 116 may be stored on a removable CD-ROM and elsewhere in a read-only DRAM chip. Some or all of the instructions and data may be stored at a location physically remote from the processor 112 but still accessible by the processor 112. Further, FIG. 1 illustrates the computing device 110 as including only one processor 112, memory 114, network interface 119, and HW accelerator card 118, but the computing device 110 can include any number of processors, memories, network interfaces, and HW accelerator cards. Similarly, the processor 120 can actually include a collection of processors, and these processors may or may not operate in parallel.
[0032] Referring to FIG. 2, the HW accelerator card 118 can include a compute complex 212, a memory 214, and accelerators 228a, 228b, and 228c. The compute complex can include one or more computing units 213. The compute complex can control the general operation of the other components of the hardware accelerator, for example, by distributing processing tasks to the accelerators 228a-228c and communicating with other devices within the computing device 110 such as the processor 112. Optionally, the compute complex can coordinate service aggregation or otherwise assist as described herein.
[0033] One or more computing units of the computing complex 212 can include one or more general-purpose processors and / or dedicated processors. Typically, the computing unit 213 of the hardware accelerator card can be one or more dedicated processors such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs) that can execute an ARM or MIPS-based instruction set, although other instruction sets may be used. Optionally, the computing unit 213 can be a commercially available processor.
[0034] Each of the accelerators 228a - 228c can be composed of one or more processors that can provide a specific acceleration service. For example, each accelerator can be configured to provide a specific acceleration service such as compression, encryption, code conversion, hash generation, graphics processing, simulation, etc. Some HW accelerator cards can be configured to provide multiple acceleration services such as compression and encryption, or any other combination of acceleration services. One or more processors of the accelerator can be one or more dedicated processors used, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), dedicated processors, etc. Three accelerators including accelerators 228a - 228c are shown in FIG. 2, although the HW accelerator card can include any number of accelerators. As described above, each individual accelerator can be configured to provide multiple acceleration services (e.g., multiple functions and / or capabilities).
[0035] Referring back to FIG. 2, the HW accelerator card includes a memory 214. The memory 214 can be any type of non-transitory computer-readable medium that can store information accessible by the processor 120, such as a hard drive, solid-state drive, tape drive, optical storage device, memory card, ROM, RAM, DVD, CD-ROM, writable memory, read-only memory, etc., and in this regard, can be compared with the memory 114. The memory 214 can store information accessible by the compute complex 212 and / or the accelerator 228a - 228c, including instructions 217 that can be executed by the compute unit 213 of the compute complex 212 and / or the accelerator 228a - 228c. Although not shown, each accelerator 228 - 228c can have its own memory and / or a pool of shared memory for storing data and instructions for executing tasks assigned by the compute complex 212.
[0036] The instructions 217 can include an Accelerated Service Manager (ASM) program 219. As further described herein, the ASM 219 can be executed by one or more compute units 213 of the compute complex to control or otherwise assist in the service aggregation of the HW accelerator card 118.
[0037] The data 216 in the memory 214 can be retrieved, stored, or modified by the compute complex 212 and / or the accelerator 228a - 228c according to the instructions 217 or other such instructions. As further shown in FIG. 2, the data 216 can include one or more acceleration service listings 218. Acceleration The acceleration service listing 218 can include a list of acceleration services provided by each of the accelerators 228a to 228c. The acceleration service listing 218 can take a standardized form. In this regard, a specific unique identifier can be assigned to each specific acceleration service. All accelerators having a certain acceleration service will include the unique identifier associated with that certain acceleration service within the listing of acceleration services.
[0038] Figure 3 shows listing examples 328a to 328c as stored in the memory 214, and the listing examples 328a to 328c respectively correspond to the accelerators 228a to 228c. In this regard, the memory 214 includes listings for each accelerator on the HW accelerator card 118. As shown, the listing 328a identifies the acceleration service provided by the accelerator 228a as identified by the unique identifier including function 1, capability 1, function 3, and function 5. Similarly, the accelerator 228b can provide three acceleration services, and each acceleration service is identified in the listing 328b by the unique identifier of the listing 328b including function 1, function 5, and capability 9. The accelerator 228c can provide two acceleration services. These two acceleration services are respectively identified in the listing 328c by the unique identifier including capability 9 and function 5.
[0039] As further shown in FIG. 3, accelerators that provide a common acceleration service can be associated with the same unique identifier within each listing of the accelerators. For example, Function 1 is a unique identifier associated with a specific function that accelerators 228a and 228b can execute. Thus, listings 328a and 328b include Function 1 with the same unique identifier. Similarly, Capability 9 is a unique identifier associated with a specific capability of accelerators 228b and 228c. Thus, listings 328b and 32c include the same unique identifier of Capability 9. The unique identifiers in FIG. 3 are merely examples of possible identifiers. The identifier can include any value or other such indicator including numbers, characters, symbols, etc.
[0040] Listings 328a - 328c are examples of possible forms for creating a list of unique identifiers associated with the accelerators of HW accelerator card 118. In some examples, the listing of accelerators can be stored in a composite listing such as a spreadsheet or database. For example, the composite listing can identify each accelerator and the unique identifier associated with the acceleration service provided by that accelerator. Similarly, the listing can be classified according to the accelerators. For example, the first listing can include a composite listing for a first set of accelerators, and the second listing can include a composite listing for a second set of accelerators. Other data may be included in the listing. FIGS. 2 and 3 show the listing as being stored in memory 216, but the listing can be stored in the memory of one or more accelerators.
[0041] Although not shown, an administrator can maintain a repository of acceleration services and associated unique identifiers for the acceleration services. The administrator can be an individual(s), a company, a collection of companies, standard organization(s), etc. In addition to maintaining the repository, the administrator can assign a unique identifier to each acceleration service and add additional acceleration services and corresponding unique identifiers when they are developed, received, or otherwise requested. Providing a repository of acceleration services and associated unique identifiers allows the identifiers used to indicate acceleration services to be consistent across multiple HW accelerator cards, even when the HW accelerator cards are manufactured by different suppliers.
[0042] Referring again to FIG. 2, the processor 112 can communicate directly with the hardware accelerator card 118 using a communication interface and protocol. For example, the processor(s) 112 can communicate with the hardware accelerator card(s) using a PCIe interface 260. Although FIG. 2 shows a PCIe interface 260, other communication interfaces and protocols may be used. For example, the processor(s) 112 can communicate with the HW accelerator card(s) 118 using one or more of a CAN interface and protocol, an SPI interface and protocol, a USB interface and protocol, an eSPI interface and protocol, an Ethernet interface and protocol, an IDE interface and protocol, or other such interfaces and protocols.
[0043] Communication between devices via a communication interface, such as communication between the processor 112 and the HW accelerator card 118 via the PCIe interface 260, can be controlled via an operating system running on the computing device 110. In this regard, the operating system can set up handles to provide communication channels between the devices attached to the PCIe interface 260. Optionally, the operating system can also close the communication channels between different devices connected to the PCIe interface 260.
[0044] Although not shown in FIG. 1 or FIG. 2, the computing device 110 can include a monitor having another device, such as a display device, other components commonly found in a personal computer and / or a server, e.g., a screen, a projector, a touch screen, a small LCD screen, a television, or an electrical device operable to display information processed by the processor 112. The computing device 110 can also include speakers. The computing device 110 can also include one or more user input devices, such as a mouse, a keyboard, a touch screen, a microphone, etc. The computing device 110 can also include hardware for connecting some or all of the aforementioned components together with each other.
[0045] FIG. 4 shows a networked computing system 400 that includes computing devices 410, 420, 430, and 440 connected to each other via a network 370. The computing devices 310 - 340 can each be compared to the computing device 110. For clarity, FIG. 4 shows the computing devices 410 - 440 as each including an HW accelerator card 418, 428, 438, 448 and a network interface card 419, 429, 439, 449, respectively. However, each of the computing devices 410 - 440 can include the same or different components as the computing device 110, such as one or more processors, memory, HW accelerator cards, and / or network interfaces, etc., as well as other components that are typically found within a computer or otherwise connected to a computer in other ways.
[0046] Network 470 can include various protocols and systems such that the network can be part of the Internet, World Wide Web, a particular intranet, wide area network, or local network. The network can utilize standard communication protocols such as Ethernet, WiFi, and HTTP, protocols owned by one or more companies, as well as various combinations of the aforementioned protocols. Although several advantages are obtained when information is transmitted or received as described above, other aspects of the subject matter described herein are not limited to any particular manner of transmitting information. Each of the computing devices 410 - 440 can communicate with other computing devices via the network 470. Optionally, the network 470 can be configured such that communication is possible only between a subset of the computing devices. For example, the computing device 410 can communicate with the computing devices 420 and 430 via the network 470, but may not be able to communicate with the computing device 440. Further, although FIG. 4 shows only four computing devices connected via the network 470, the networked computing system can include any number of computing devices and networks.
[0047] FIG. 5 shows acceleration service listings 510 - 540 of acceleration services respectively provided by HW accelerators in computing devices 410 - 440. Similar to listings 328a - 328c in FIG. 3 that identify the acceleration services provided by the accelerator of computing device 110, acceleration service listings 510 - 540 include the capabilities and functionality respectively provided by the accelerators of HW accelerator cards 418 - 448 in computing devices 410 - 440. Each acceleration service provided by the accelerator of an HW accelerator card mounted on the communication interface of a computing device is identified with a "(local)" label. For example, the HW accelerator card 418 in computing device 410 can provide acceleration services for networking and compression. Similarly, the HW accelerator card 428 in computing device 420 can provide acceleration services for networking and encoding, the HW accelerator card 438 in computing device 430 can provide acceleration services for networking and hashing, and the HW accelerator card 448 in computing device 440 can provide services for networking and encryption.
[0048] Listings 510 through 540 also include acceleration services of other computing devices on network 470 identified via service aggregation. These remotely available acceleration services are identified in FIG. 5 with the "(remote)" label. For example, Listing 510 identifies the encoding provided by the HW accelerator card in computing device 320, the hashing provided by the HW accelerator card in computing device 330, and the encryption provided by the HW accelerator card in computing device 340 as aggregated acceleration services available to computing device 310. FIG. 5 shows each computing device as being able to use all acceleration services provided by all connected computing devices, although in some cases, acceleration services provided by a remote device may be blocked from remote access. For example, computing device 340 can prevent the local encryption acceleration service from being made available by service aggregation.
[0049] The acceleration service listing for each computing device can be generated by ASM software running on one or more HW accelerator cards. For example, during initialization of a HW accelerator card such as HW accelerator card 418 of computing device 410, the compute complex of the HW accelerator card can execute ASM. ASM can prepare a listing of acceleration services that can be provided through the HW accelerator card, including local acceleration services and remote acceleration services. This acceleration service listing can be provided by the operating system running on computing device 410 or from a configuration file stored in memory on HW accelerator card 418, computer device 410, or elsewhere.
[0050] In some cases, the acceleration service listing can be discovered dynamically. For example, the ASM running on the HW accelerator card 418 can communicate with the HW accelerator card locally attached to the computing device 410 and / or other HW accelerator cards of computing devices connected to the network 470. During communication, the HW accelerator card 418 can request a listing of the acceleration services provided by other local or remotely connected HW accelerators. The HW accelerator card 418 can collect these other acceleration services into an acceleration listing 510. In some cases, the HW accelerator card can maintain separate acceleration listings for each other local or remotely connected HW accelerator card. For example, although FIG. 5 shows only a single listing 510 for the computing device 410, the computing device 410 can have a listing for the acceleration services provided by the HW accelerator card 418 and a separate listing for the acceleration services provided by other remotely connected HW accelerator cards. The foregoing example describes discovering and collecting acceleration services by an ASM program running on the HW accelerator card 418, but other HW accelerator cards, such as accelerator cards 428, 438, and / or 448, can also prepare listings of the acceleration services available to their respective computing devices 420 - 440.
[0051] The ASM running on the HW accelerator card 418 can manage the acceleration service listing. In this regard, the ASM can determine which acceleration services are healthy (e.g., operable, available for processing a workload, have sufficient processing capacity for a workload, etc.). The ASM can monitor the operation of the acceleration services in the acceleration service listing to determine the state of the acceleration services (e.g., healthy, unhealthy / busy, unavailable, etc.). For example, the ASM can request status updates from other ASMs to determine the state of acceleration services provided by other HW accelerator cards. The ASM can remove the acceleration service listing to remove acceleration services identified as unavailable (e.g., unreachable) by other ASMs. In another example, the ASM can mark acceleration services identified as unhealthy (e.g., inefficient / slow in operation) or busy (e.g., processing other workloads, reserved for other workloads, etc.) so that the ASM does not send workloads to these busy / unhealthy acceleration services.
[0052] In another example, the ASM can monitor the workload or response indicator sent for completion. If no response is received or the workload completion is not identified by the ASM, the ASM can determine that the workload was not received by the remote HW accelerator and / or not processed by the expected acceleration service provided by the remote HW accelerator. In such a case, the ASM can request that the workload be executed by another acceleration service.
[0053] In some cases, ASM can distribute the workload across multiple acceleration services for load balancing, fault handling, performance, or other such considerations. For example, the workload can be large, and ASM can utilize the acceleration services of many HW accelerator cards to efficiently process the workload. In another example, ASM can prioritize several remote HW accelerator cards. For example, ASM can direct the workload to be processed by a first remote HW accelerator. If the first remote HW accelerator cannot process the workload, a second remote HW accelerator can process the workload. This process can be repeated until the workload is processed, with additional fallback HW accelerators being directed to process the workload.
[0054] As part of the initialization of the HW accelerator card, ASM running on the compute complex can prepare services for sending and receiving service calls both locally and remotely. In this regard, ASM can initialize and / or confirm that the end service code is enabled to process local service calls (e.g., calls to be processed by the accelerator of the HW accelerator card on which ASM is running).
[0055] ASM can also initialize and / or confirm that its proxy code for remote services is enabled. The proxy code can be used when the medium for invoking the remote service (e.g., a call sent from one ASM to another ASM for processing by an accelerator on another HW accelerator card) is enabled. In this regard, the proxy code refers to the code that enables communication from one ASM to another ASM. The proxy code may also refer to the code that enables the ASM on the local HW accelerator card to call an acceleration service from an accelerator on another remote HW accelerator card where there may be no ASM running for it. That is, the ASM of the local HW accelerator card by the proxy code can be configured to define the path of a request for an acceleration service from the local computing device to the remote HW accelerator card when the remote HW accelerator card does not include an ASM.
[0056] ASM can publish an acceleration service listing to a processor of a computing device. For example, the ASM running on the HW accelerator card 418 can provide the listing 510 to one or more processors of the computing device 410. Optionally, an operating system running on the computing device can request the listing from the ASM. For example, the operating system running on the computing device 410 can request the listing 510 of acceleration services from the HW accelerator card 418.
[0057] Example of method FIG. 6 is a flowchart showing a process of discovering an acceleration service provided by an HW accelerator card such as HW accelerator card 118 connected to a processor such as processor 112 via a communication interface such as PCIe bus 260. Processor 112 can request to communicate with HW accelerator card 118 (shown by the dashed line) via a PCIe interface. An operating system running on the computing device can provide a communication channel via the PCIe bus between the HW accelerator card and processor 112.
[0058] Using the communication channel, processor 112 can send a request for a listing of acceleration services provided by the accelerators on HW accelerator card 118, as shown at line 623. In response to receiving the request from processor 118, compute complex 212 of HW accelerator card 118 can query and receive from memory 214 of the HW accelerator card (or the memory of the accelerator), as indicated by arrows 625 and 627 respectively, a listing of acceleration ser vices. In this regard, the HW accelerator card can gather the acceleration services of all accelerators. In some cases, HW accelerator card 118 can query only some of the accelerators.
[0059] Optionally, HW accelerator card 118 can gather the acceleration services of the accelerators in a hierarchical manner. In this regard, the acceleration services can be hierarchical in that one acceleration service can rely on or be dependent on another acceleration service. This hierarchical relationship between the acceleration services can be identified and stored in this listing. Optionally, each level in the hierarchical relationship can identify the capabilities and functionality of the lower level.
[0060] The compute complex 212 can provide a listing of acceleration services to the processor 112 via the PCIe bus 260 as shown at line 629. When the processor receives the listing of acceleration services, the communication channel can be closed.
[0061] If the processor can utilize one or more acceleration services, the processor 112 can request the HW accelerator card to complete one or more tasks using one of the provided acceleration services provided by the accelerator on the HW accelerator card 118. FIG. 7 shows the processor 112 requesting information about the acceleration services of the HW accelerator card 118 connected via the PCIe bus 260. In this regard, steps 723-729 correspond to steps 623-629 described above.
[0062] As indicated by arrow 729, the HW accelerator indicates that the HW accelerator can provide a compression service. Upon receiving an acceleration service, the processor 112 can provide, as indicated by arrow 731, a workload instruction including an indication of a location to store data and instructions to the HW accelerator card 118 to compress the data. Next, the compute complex 212 of the HW accelerator card can, as indicated by arrow 733, confirm the instruction and provide an ID that the processor 112 can communicate with to obtain status update information about the compression by the HW accelerator card 118. Next, the processor 112 can, as indicated by arrows 735 and 737 respectively, request and receive the status of the compression. When a polling request indicates that the compression is complete, communication between the processor 112 and the HW accelerator card 118 can end, or additional tasks can be sent from the processor 112 to the HW accelerator card. Although FIG. 7 shows a compression service, the processing performed by the HW accelerator can be any kind of operation or combination of operations.
[0063] FIG. 8 shows a flowchart of a service aggregation operation. Software running on the processor 812 of a computing device can request one or more acceleration services from an HW acceleration card 818 to process a workload. The request indicated at line 823 can be sent by the processor 812 to the local HW accelerator card 818 via a communication link established by a communication interface such as the PCE interface 860. For illustration purposes, the acceleration service being requested by the processor 812 is a "compression" service for processing a workload. However, any other acceleration service may be requested.
[0064] Before sending request 823, the software executing on processor 812 can be provided with a listing of acceleration services available locally (i.e., by the accelerators of HW accelerator card 818) or remotely (i.e., by other remote accelerators of HW accelerators connected via a network) from HW accelerator card 818. For example, this listing can be provided during the initialization of HW accelerator card 818. In another example, software such as an operating system or another program executing on the computing device can request the listing. In yet another example, the listing of acceleration services can be received by the software from a configuration system without communication with HW accelerator card 118. The configuration system can be a centralized listing, and the software executing on the computing device can obtain the listing therefrom. For example, the configuration system can be located remotely, together with a storage device accessible via a network or other connection.
[0065] Request 823 can include the name (e.g., unique identifier) of the requested acceleration service, in-line parameters, and the pinned memory location(s) where the input(s) (e.g., data for processing the workload) for the requested acceleration service(s) can be found.
[0066] After receiving the request, the ASM of the HW accelerator card 818 can determine whether the requested acceleration service is local, as shown at line 825. If the acceleration service is local, the ASM can provide service to the request using in-line parameters and access the memory address passed using DMA (Direct Memory Access). Then, a local accelerator that can execute the requested acceleration service can process the input (e.g., data). Referring again to FIG. 8, the requested compression acceleration service can be performed by the accelerator of the HW accelerator card 818. Then, the result (e.g., output) of the completed compression request can be passed from the HW accelerator card 818 to the processor 812, as shown at line 827.
[0067] If the requested acceleration service is not locally available, the ASM can act as a client of the remote service HW accelerator card for the requester. In this regard, the ASM of the HW accelerator card can pass the request to a remote HW accelerator card that provides the requested acceleration service. The HW accelerator card can also pass the pinned memory to the remote HW accelerator card via Remote Direct Memory Access (RDMA).
[0068] FIG. 8 further shows the flow of the remote service aggregation operation. For example, the processor 812 can request the encryption acceleration service provided by the remote HW accelerator card 838 to the local HW accelerator card 818, as shown at line 829. After receiving the request, the local HW accelerator card 818 can determine that the encryption is not an acceleration service provided by this card, as shown at line 831.
[0069] When the local HW accelerator card is unable to complete the acceleration service, the ASM of the local HW accelerator card 818 can determine which remote HW accelerator card can provide the service. Referring to FIG. 8, the ASM of the local HW accelerator 818 determines that the remote HW accelerator card 838 can provide the encryption acceleration service and can transfer the request as shown at line 833.
[0070] Optionally, the request can include the identifier of the HW accelerator card that can provide the acceleration service. In such a case, the local HW accelerator card 818 can skip determining whether this card can perform the requested acceleration service and continue to transfer the request to the identified remote HW accelerator card.
[0071] Next, the remote HW accelerator card 838 can perform the encryption acceleration service shown at line 835 in FIG. 8. Then, the result of the completed encryption (e.g., the output) can be passed from the remote HW accelerator card 838 to the local HW accelerator card 818 as shown at line 837 and from the local HW accelerator card 818 to the processor 812 as shown at line 839.
[0072] Communication between HW accelerator cards can be performed by an ASM program executed on each HW accelerator card. In this regard, each ASM can communicate directly. Alternatively, the ASM can encapsulate communication data (e.g., requests, outputs, etc.) into data packets within an outer header that can define a route to a remote router table destination for the packets, and the address of the remote router table destination can be discovered via regular routing. When a packet reaches the intended ASM, the ASM can perform decapsulation of the packet and direct the packet to the correct accelerator on that HW accelerator card as identified within the packet header.
[0073] Unless otherwise specified, the foregoing alternatives are not mutually exclusive and rather can be implemented in various combinations to achieve their respective advantages. Since these and other variations and combinations of the above-described features can be utilized without departing from the subject matter defined by the claims, the foregoing description of the embodiments should be received as illustrative rather than as a limitation of the subject matter defined by the claims. Additionally, the examples provided herein, as well as clauses expressed as "such as," "including," etc., should not be construed as limiting the subject matter of the claims to specific examples, but rather the examples are for the purpose of illustrating only one of many possible embodiments. Further, the same reference numbers within different drawings can identify the same or similar elements.
Claims
1. Providing, by one or more processors of a local hardware (HW) accelerator card, a listing of acceleration services via a communication interface from the local HW accelerator card, wherein the listing of acceleration services includes a first set of acceleration services provided by one or more accelerators of the local HW accelerator card and a second set of acceleration services provided by one or more accelerators of a remote HW accelerator card; Receiving, by the one or more processors, a workload instruction from a processor of a computing device, wherein the workload instruction defines a workload to be processed by at least one of the acceleration services of the second set of acceleration services; Transferring, by the one or more processors, the workload instruction to the remote HW accelerator card A method comprising the above steps.
2. Receiving, by the one or more processors, a processed workload from the remote HW accelerator card, wherein the processed workload is the workload after being processed by at least one of the acceleration services of the second set of acceleration services; The method according to claim 1, further comprising the above step.
3. Transferring, by the one or more processors, the processed workload to the processor of the computing device The method according to claim 2, further comprising the above step.
4. The method according to claim 1, wherein the listing of acceleration services is generated by an accelerated service manager (ASM) running on the one or more processors.
5. The method according to claim 4, wherein the ASM running on the one or more processors communicates with another ASM running on the remote HW accelerator card.
6. The method according to claim 5, wherein the communication includes the ASM transferring the workload instruction to the other ASM.
7. The method of claim 5, wherein the communication includes the ASM requesting a listing of the second set of acceleration services from the other ASM.
8. The method of claim 5, wherein the ASM identifies and removes an unsound acceleration service from the listing of acceleration services.
9. The method of claim 8, wherein identifying the unsound acceleration service includes the ASM determining a failure in processing of the workload instruction by at least one of the acceleration services of the second set of acceleration services.
10. Removing the unsound acceleration service includes marking at least one of the acceleration services of the second set of acceleration services as unsound, or removing at least one of the acceleration services of the second set of acceleration services from the listing of acceleration services. The method of claim 9.
11. After determining the failure in processing of the workload instruction, a step of the ASM transmitting an updated workload instruction to a different HW accelerator card for processing by at least one of the acceleration services of the different HW accelerators The method of claim 10, further comprising.
12. The method of claim 1, further defining that the workload instruction is processed by at least one acceleration service of at least one other remote HW accelerator card.
13. A communication interface and A local hardware (HW) accelerator card including one or more processors and one or more accelerators, wherein the one or more processors Receiving, via the communication interface, a listing of acceleration services from the local HW accelerator card, the listing of acceleration services including a first set of acceleration services provided by one or more accelerators of the local HW accelerator card and a second set of acceleration services provided by one or more accelerators of a remote HW accelerator card; Receiving a workload instruction from a processor of a computing device, the workload instruction defining a workload to be processed by at least one of the acceleration services of the second set of acceleration services; Transferring the workload instruction to the remote HW accelerator card; A local hardware (HW) accelerator card configured to perform; A system comprising.
14. The one or more processors are further configured to: Receive a processed workload from the remote HW accelerator card, the processed workload being the workload after being processed by the at least one of the acceleration services of the second set of acceleration services; The system according to claim 13, further configured to perform.
15. The one or more processors are further configured to transfer the processed workload to the processor of the computing device; The system according to claim 14, further configured to perform.
16. The system according to claim 13, wherein the listing of acceleration services is generated by an accelerated service manager (ASM) running on the one or more processors.
17. The system according to claim 16, wherein the ASM running on the one or more processors communicates with another ASM running on the remote HW accelerator card.
18. The system according to claim 17, wherein the communication includes the ASM transferring the workload instruction to the other ASM.
19. The system of claim 17, wherein the communication includes the ASM requesting a listing of the second set of acceleration services from the other ASM. **Claim 20** The system of claim 18, wherein the ASM identifies and removes an unsound acceleration service from the listing of acceleration services.
Citation Information
Patent Citations
Method and device for processing display video card fault in virtual cloud environment
CN113157476A
Acceleration resource processing method and device
JP2019522293A
Method and Apparatus for Configuring Accelerator
US20180307499A1
Automatic localization of acceleration in edge computing environments
US20200026575A1
Technologies for accelerator fabric protocol multipathing
US20200218684A1
Cited By
Control device, industrial machine, and control method
US12377507B2