Data processing system with a hardware-accelerated plane and a software plane
By designing a data processing system that integrates software-driven host components and hardware acceleration components, using a common physical network to realize communication between host components and hardware acceleration components, the challenges that are difficult to integrate in the prior art are solved and efficient and flexible data processing is achieved.
Patent Information
- Application Number
- CN202110844740.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-05-20
- Filing Date
- 2016-04-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2036-04-07
AI Technical Summary
The prior art is difficult to effectively integrate software-driven computing devices and hardware acceleration components, resulting in challenges in meeting the design requirements of specific data processing environments.
A data processing system is designed that includes two or more software-driven host components and two or more hardware acceleration components, allowing host components and hardware acceleration components to communicate with each other through a common physical network, and enabling transparent communication between hardware acceleration components in the hardware acceleration plane.
It realizes a solution to effectively integrate software-driven computing devices and hardware acceleration components, improves the efficiency and flexibility of the data processing system, and can meet various design requirements in different data processing environments.
Smart Images

Figure CN113553185B_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application for an invention named "Data Processing System with Hardware Acceleration Plane and Software Plane", with the international application date of April 7, 2016, entering the Chinese national phase on December 15, 2017, and the Chinese national application number 201680035401.8. Technical Field
[0002] The present disclosure relates to the field of computers, and more particularly to a data processing system with a hardware acceleration plane and a software plane. Background Art
[0003] The computer industry faces increasing challenges in its efforts to improve the speed and efficiency of software-driven computing devices, for example, due to power limitations and other factors. Software-driven computing devices employ one or more central processing units (CPUs) that process machine-readable instructions in a conventional timing manner. To address this issue, the computing industry has proposed using hardware acceleration components (such as field-programmable gate arrays (FPGAs)) to supplement the processing performed by software-driven computing devices. However, software-driven computing devices and hardware acceleration components are different types of devices with fundamentally different architectures, performance characteristics, power requirements, program configuration paradigms, interface characteristics, etc. Therefore, integrating these two types of devices in a manner that meets the various design requirements of a specific data processing environment is a challenging task. Summary of the Invention
[0004] A data processing system is described herein that includes two or more software-driven host components. The two or more host components together provide a software plane. The data processing system also includes two or more hardware acceleration components (such as FPGA devices) that together provide a hardware acceleration plane. In one implementation, a common physical network allows the host components to communicate with each other and also allows the hardware acceleration components to communicate with each other. Further, the hardware acceleration components in the hardware acceleration plane include functions that enable them to communicate with each other in a transparent manner without assistance functions from the software plane. Generally speaking, the data processing system can be considered to support two logical networks that share a common physical network substrate. The logical networks can interact with each other but operate in an independent manner.
[0005] The functions summarized above can be manifested in various types of systems, devices, components, methods, computer-readable storage media, data structures, graphical user interface presentations, articles of manufacture, etc.
[0006] The present invention content is provided to introduce a selection of concepts in a simplified form; these concepts are further described in the following detailed description. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 An overview of a data processing system including a software plane and a hardware acceleration plane is shown.
[0008] Figure 2 Shows Figure 1 a first example of the operation of a data processing system.
[0009] Figure 3 Shows Figure 1 a second example of the operation of a data processing system.
[0010] Figure 4 Shows an implementation corresponding to a data center Figure 1 of a data processing system.
[0011] Figure 5 Is Figure 4 a more inclusive depiction of a data center implementation.
[0012] Figure 6 Shows an alternative way of implementing server unit components compared to the way shown in Figure 4
[0013] Figure 7 Shows another way of implementing server unit components compared to the way shown in Figure 4
[0014] Figure 8 Shows an alternative data processing system compared to the data processing system shown in Figure 1 e.g., compared to the network infrastructure shown in Figure 1 which uses a different network infrastructure.
[0015] Figure 9 Is a flowchart showing Figure 1 an operating mode of a data processing system.
[0016] Figure 10 Shows an overview of an implementation of a management function for managing Figure 1 a data processing system.
[0017] Figure 11 Provides an overview of a request-driven operating mode of a service mapping component (SMC) as a component of the management function of Figure 10
[0018] Figures 12 to 15 Shows different corresponding options for processing service requests made by tenant function instances residing on a host component.
[0019] Figure 16 Provides Figure 10 An overview of another background-related mode of operation of the SMC.
[0020] Figures 17 to 20 Shows different corresponding architectures for physically implementing Figure 10 The management function.
[0021] Figures 21 to 24 Shows various corresponding strategies for configuring a hardware acceleration component in a Figure 1 Data processing system.
[0022] Figure 25 Shows one way of implementing Figure 1 The hardware acceleration component.
[0023] Figure 26 Shows a hardware acceleration component including separate configurable domains.
[0024] Figure 27 Shows a function for performing data transfer between a local host component and an associated local hardware acceleration component.
[0025] Figure 28 Shows an implementation of a router introduced in Figure 25 Something.
[0026] Figure 29 Shows an implementation of a transport component introduced in Figure 25 Something.
[0027] Figure 30 Shows an implementation of a 3-port switch introduced in Figure 25 Something.
[0028] Figure 31 Shows Figure 1 An implementation of the host component shown.
[0029] Throughout the disclosure and the figures, the same reference numerals are used to indicate the same components and features. 100-series numbers refer to features first found in Figure 1 Something, 200-series numbers refer to features first found in Figure 2 Something, 300-series numbers refer to features first found in Figure 3 Something, and so on. Detailed Description
[0030] The present disclosure is organized as follows. Section A describes an illustrative data processing system that includes a hardware acceleration plane and a software plane. Section B describes management functions for managing the data processing system of Section A. Section C elaborates on one implementation of illustrative hardware acceleration components in the hardware acceleration plane.
[0031] As a preliminary matter, some of the figures in the drawings describe concepts in the context of one or more structural components, which are variously referred to as functions, modules, features, elements, etc. The various components shown in the drawings can be implemented in any manner by any physical and tangible mechanism, such as, by software running on a computing device, hardware (e.g., chip-implemented logic functions), etc., and / or any combination thereof. In one case, the various components illustrated in the drawings are divided into different units, which may reflect corresponding different physical and tangible components used in the actual implementation. Alternatively or additionally, any single component illustrated in the drawings can be implemented by multiple actual physical components. Alternatively or additionally, the depiction of any two or more separate components in the drawings may reflect different functions performed by a single actual physical component.
[0032] Other figures describe concepts in the form of flowcharts. In this form, certain operations are described as constituting different boxes that are executed in a certain order. These embodiments are illustrative and not restrictive. Certain boxes described herein can be combined together and executed in a single operation, certain boxes can be broken down into multiple component boxes, and certain boxes can be executed in an order different from that illustrated herein (including executing boxes in parallel). The boxes shown in the flowcharts can be implemented in any manner by any physical and tangible mechanism, such as, by software running on a computing device, hardware (e.g., chip-implemented logic functions), etc., and / or any combination thereof.
[0033] Regarding terminology, the phrase "configured to" encompasses any manner by which any kind of physical and tangible functionality can be constructed to perform the identified operation. The functionality can be configured to perform the operation using, for example, software running on a computing device, hardware (e.g., chip-implemented logic functions), etc., and / or any combination thereof.
[0034] The term "logic" encompasses any physical and tangible functionality for performing a task. For example, each operation illustrated in a flowchart corresponds to a logic component for performing that operation. The operations can be performed using, for example, software running on a computing device, hardware (e.g., chip-implemented logic functions), etc., and / or any combination thereof. When implemented by a computing device, the logic components represent electrical components that are physical parts of the computing system, however implemented.
[0035] Any storage resource or any combination of storage resources described herein may be considered a computer-readable medium. In many instances, a computer-readable medium represents some form of physical and tangible entity. The term computer-readable medium also encompasses propagated signals, such as those transmitted or received via physical conduits and / or air or other wireless media. However, the specific terms “computer-readable storage medium” and “computer-readable medium device” expressly exclude the propagated signals themselves, while including all other forms of computer-readable media.
[0036] The following explanations may identify one or more features as “optional.” Statements of this type should not be construed as an exhaustive indication of features that may be considered optional; that is, other features may be considered optional, even though not explicitly identified in the text. Further, any description of a single entity is not intended to exclude the use of multiple such entities; similarly, a description of multiple entities is not intended to exclude the use of a single entity. Further, although the description may interpret certain features as alternative ways of performing the identified function or implementing the identified mechanism, the features may also be combined together in any combination. Finally, the term “exemplary” or “illustrative” refers to one implementation among potentially many implementations.
[0037] A. Overview
[0038] Figure 1 An overview of a data processing system 102 including a software plane 104 and a hardware acceleration plane 106 is shown. The software plane 104 includes a collection of software driver components (each represented by the Figure 1 symbol “S” in ) while the hardware plane includes a collection of hardware acceleration components (each represented by the Figure 1 symbol “H” in ). For example, each host component may correspond to a server computer that executes machine-readable instructions using one or more central processing units (CPUs). In turn, each CPU may execute instructions on one or more hardware threads. On the other hand, each hardware acceleration component may correspond to hardware logic for implementing functions such as field programmable gate array (FPGA) devices, massively parallel processor array (MPPA) devices, graphics processing units (GPUs), application specific integrated circuits (ASICs), multi-processor system-on-chip (MPSoCs), etc.
[0039] The term "hardware" acceleration component is also intended to broadly cover different ways of using hardware devices to perform functions, including, for example, at least: (a) cases where at least some tasks are implemented in the form of hard ASIC logic, etc.; (b) cases where at least some tasks are implemented in the form of soft (configurable) FPGA logic, etc.; (c) cases where at least some tasks run as software on top of an FPGA software processor overlay; (d) cases where at least some tasks run on an MPPA of a soft processor, etc.; (e) cases where at least some tasks run as software on a hard ASIC processor, etc., or any combination thereof. Similarly, the data processing system 102 can accommodate different manifestations of software-driven devices in the software plane 104.
[0040] To simplify repeated references to the hardware acceleration components, the following explanations will refer to these devices simply as "acceleration components". Further, the following explanations will present the main example where the acceleration components correspond to FPGA devices, although as noted, the data processing system 102 can be constructed using other types of acceleration components. Further, the hardware acceleration plane 106 can be constructed using a heterogeneous collection of acceleration components, including different types of FPGA devices with different corresponding processing capabilities and architectures, a mixture of FPGA devices and other devices, and so on.
[0041] Host components typically use a sequential execution paradigm to perform operations, for example, by using each CPU hardware thread in its CPU hardware threads to execute machine-readable instructions one by one. In contrast, acceleration components can use a spatial paradigm to perform operations, for example, by using a large number of parallel logic elements to execute computational tasks. Therefore, compared with software-driven host components, acceleration components can perform some operations in a shorter time. In the context of the data processing system 102, the "acceleration" qualifier associated with the term "acceleration component" reflects its potential to accelerate functions performed by host components.
[0042] In one example, the data processing system 102 corresponds to a data center environment including multiple computer servers. The computer servers correspond to Figure 1 the host components in the software plane 104 shown. In other cases, the data processing system 102 corresponds to an enterprise system. In other cases, the data processing system 102 corresponds to a user device or appliance using at least one host component that can access two or more acceleration components, etc. These examples are cited by way of illustration rather than limitation; other applications are possible.
[0043] In one implementation, each host component in the data processing system 102 is coupled to at least one acceleration component via a local link. The basic unit of the processing device is referred to herein as a "server unit component" because the devices can be grouped together and maintained as a single serviceable unit (although not necessarily so) within the data processing system 102. The host components in the server unit component are referred to as "local" host components to distinguish them from other host components associated with other server unit components. Similarly, the acceleration components of the server unit component are referred to as "local" acceleration components to distinguish them from other acceleration components associated with other server unit components.
[0044] For example, Figure 1 An illustrative local host component 108 is shown, which is coupled to a local acceleration component 110 via a local link 112 (e.g., a Peripheral Component Interconnect Express (PCIe) link as will be described below). The pairing of the local host component 108 and the local acceleration component 110 forms at least a part of a single server unit component. More generally, Figure 1 The software plane 104 is shown coupled to the hardware acceleration plane via a number of individual local links, Figure 1 collectively referred to as the local H-to-local S coupling 114.
[0045] The local host component 108 can also communicate indirectly with any other remote acceleration component in the hardware acceleration plane 106. For example, the local host component 108 can access a remote acceleration component 116 via the local acceleration component 110. More specifically, the local acceleration component 110 communicates with the remote acceleration component 116 via a link 118.
[0046] In one implementation, a common network 120 is used to couple host components in the software plane 104 to other host components and to couple acceleration components in the hardware acceleration plane 106 to other acceleration components. That is, two host components can communicate with each other using the same network 120 in the same way as two acceleration components. As another feature, the interaction between host components in the software plane 104 is independent of the interaction between acceleration components in the hardware acceleration plane 106. This means that, for example, from the perspective of host components in the software plane 104, two or more acceleration components can communicate with each other in a transparent manner, outside the direction of the host components, and the host components do not "know" the specific interactions occurring in the hardware acceleration plane 106. However, the host components can initiate interactions occurring in the hardware acceleration plane 106 by issuing service requests targeted at services hosted by the hardware acceleration plane 106.
[0047] According to a non - restrictive implementation, the data processing system 102 uses the Ethernet protocol to transmit IP packets over the public network 120. In one implementation, each local host component in the server unit component is given a single physical IP address. The local acceleration components within the same server unit component can adopt the same IP address. The server unit component can determine whether an incoming packet is destined for a local host component rather than a local acceleration component in different ways. For example, a packet destined for a local acceleration component can be expressed as a User Datagram Protocol (UDP) packet specifying a particular port; on the other hand, a packet destined for a host is not expressed in this way. In another case, packets belonging to the acceleration plane 106 can be distinguished from packets belonging to the software plane 104 based on the value of a status flag in each packet (e.g., in the header or body of the packet).
[0048] Given the above characteristics, the data processing system 102 can be conceptually thought of as forming two logical networks that share the same physical communication link. Packets associated with the two logical networks can be distinguished from each other by their respective traffic classes in the manner described above. However, in other implementations (e.g., as described below with respect to Figure 8 ), the data processing system 102 can use two different physical networks to handle host - to - host traffic and hardware - to - hardware traffic respectively. Further, in implementations that do use the public network 120, the host - to - host network infrastructure does not need to be identical to the hardware - to - hardware network infrastructure; that is, the two infrastructures are common in the sense that most of their network resources are shared but not necessarily all of them are shared.
[0049] Finally, the management function 122 is used to manage the operation of the data processing system 102. As will be elaborated in more detail in Section B (below), the management function 122 can be physically implemented using different control architectures. For example, in one control architecture, the management function 122 can include multiple local management components coupled to one or more global management components.
[0050] Through the introduction of Section B, the management function 122 can include several sub - components that perform different corresponding logical functions (which can be physically implemented in different ways). For example, the location determination component 124 identifies the current location of services within the data processing system 102 based on the current allocation information stored in the data storage device 126. As used herein, a service refers to any function performed by the data processing system 102. For example, a service can correspond to an encryption function. Another service can correspond to a document sorting function. Another service can correspond to a data compression function, and so on.
[0051] In operation, the location determination component 124 can receive a request for a service. In response, the location determination component 124 returns the address associated with the service, if that address exists in the data storage device 126. The address can identify the specific acceleration component hosting the requested service.
[0052] The service mapping component (SMC) 128 maps services to specific acceleration components. The SMC 128 can operate in at least two modes based on the type of trigger event it receives that invokes its operation. In the first case, the SMC 128 processes requests for services made by tenant function instances. A tenant function instance can correspond to a software program running on a specific local host component, or more specifically, to a program executing on a virtual machine, which in turn is associated with a specific local host component. The software program can request services during its execution. The SMC 128 processes the request by determining the appropriate component(s) in the data processing system 102 to provide the service. The possible components considered include local acceleration components (associated with the local host component initiating the request); remote acceleration components; and / or the local host component itself (the local host component would thus implement the service in software). The SMC 128 makes its determination based on one or more mapping considerations, such as whether the requested service is related to a line rate service.
[0053] In another mode of operation, the SMC 128 generally operates in a background and global mode, thereby allocating services to acceleration components based on global conditions in the data processing system 102 (rather than processing individual requests from tenant function instances, or in addition to that). For example, the SMC 128 can invoke its allocation function in response to a change in demand that affects one or more services. In this mode, the SMC 128 again makes its determination based on one or more mapping considerations, such as historical demand associated with the service.
[0054] In performing its functions, the SMC 128 can interact with the location determination component 124. For example, when the SMC 128 attempts to determine the address of an already allocated service provided by an acceleration component, the SMC 128 can consult the data storage device 126. When the data storage device 126 maps a service to one or more acceleration components, it can also update the data storage device 126, for example, by storing the addresses of those acceleration components associated with the service.
[0055] Although not shown in Figure 1 the sub-components of the SMC 128 also manage multi-component services. A multi-component service is a service that consists of multiple parts. Multiple corresponding acceleration components execute the corresponding parts.
[0056] It should be noted that for convenience,Figure 1 Management functionality 122 is illustrated as being separate from components in software plane 104 and hardware plane 106. However, as will be explained in Section B, any aspect of management functionality 122 may be implemented using resources of software plane 104 and / or hardware plane 106. When implemented by hardware plane 106, management functionality may be accelerated like any service.
[0057] Figure 2 Shows the corresponding to a single transaction or a portion of a single transaction Figure 1 1. In operation (1), in the course of executing a single computing task, first host component 202 communicates with second host component 204. Second host component 204 then requests to use a service implemented in hardware acceleration plane 106 (although second host component 204 may not "know" where the service is implemented, except that the service can be accessed at a specified address).
[0058] In many cases, the requested service is implemented on a single acceleration component (although there may be multiple redundant such acceleration components from which to choose). Figure 2 In a specific example of , the requested service corresponds to a multi-component service dispersed across a collection (or cluster) of acceleration components, each of which performs an assigned portion of the service. The graph structure may specify the manner in which the individual acceleration components are coupled together in the collection. In some implementations, the graph structure also identifies at least one header component. The header component corresponds to a contact point through which entities in the data processing system 102 can interact with the multi-component service in the hardware acceleration plane 106. The header component may also serve as an initial processing stage in a processing pipeline defined by the graph structure.
[0059] exist Figure 2In a specific case, assume that the acceleration component 206 corresponds to a local acceleration component that is locally linked to the second host component 204, and the acceleration component 208 is the head component of a multi-component service. In operations (2) and (3), the second host component 204 is requested to access the acceleration component 208 via its local acceleration component 206. Then, the acceleration component 208 executes a part of its multi-component service to generate an intermediate output result. In operation (4), the acceleration component 208 then calls another acceleration component 210 that executes another corresponding part of the multi-component service to generate the final result. In operations (5), (6), and (7), the hardware acceleration plane 106 forwards the final result back to the requesting second host component 204 sequentially through the same chain of components as described above, but in the reverse direction. It should be noted that the data flow operations described above, including the process operations that define the return path, are cited only by way of example and not restrictively; other multi-component services may use other graphical structures that specify any other flow paths. For example, the acceleration component 210 may directly forward the final result to the local acceleration component 206.
[0060] First, it should be noted that the operations that occur in the hardware acceleration plane 106 are performed in a manner independent of the operations executed in the software plane 104. In other words, the host components in the software plane 104 do not manage the operations in the hardware acceleration plane 106. However, the host components can call the operations in the hardware acceleration plane 106 by issuing requests for services hosted by the hardware acceleration plane 106.
[0061] Second, it should be noted that the hardware acceleration plane 106 performs its transactions in a manner transparent to the requesting host components. For example, the second host component 204 may be "unaware" of how its request is processed in the hardware acceleration plane, including the fact that the service corresponds to a multi-component service.
[0062] Third, it should be noted that in this implementation, the communication in the software plane 104 (e.g., corresponding to operation (1)) uses the same common network 120 as the communication in the hardware acceleration plane 106 (e.g., corresponding to operations (3) to (6)). Operations (2) and (7) can be performed through local links, corresponding to the Figure 1 local H-to-local S coupling 114 shown.
[0063] Figure 2 The multi-component service shown is similar to a ring in that a series of acceleration components are traversed in one direction to obtain the final result; this final result is then propagated back through the same series of acceleration components in the direction opposite to the head component. However, as noted above, other multi-component services may use different sets of acceleration components with different corresponding process structures.
[0064] For example,Figure 3 shows a second example of the operation of a data processing system 102 that employs a different flow structure as compared to the example of Figure 1 . More specifically, in operation (1), a local host component (not shown) sends a request to its local acceleration component 302. In this case, it is assumed that the local acceleration component is also the head component of the service. In operation (2), the head component can then forward multiple messages to multiple corresponding acceleration components. Each acceleration component that receives a message can perform a part of the multi-component service in parallel with other acceleration components. (It should be noted that Figure 1 may represent only a part of a more complete transaction.) Figure 3
[0065] Moreover, a multi-component service does not necessarily need to employ a single head component or any head component. For example, a multi-component service can employ a cluster of acceleration components that all perform the same function. The data processing system 102 can be configured to invoke such a multi-component service by contacting any arbitrary member of the cluster. This acceleration component can be referred to as the head component because it is the first component to be accessed, but it does not have a special status. In other cases, the host component can initially distribute multiple requests to multiple members of a set of acceleration components.
[0066] Figure 4 shows a part of a data center 402 that represents an implementation of the Figure 1 data processing system 102. Specifically, Figure 4 shows a rack in the data center 402. The rack includes multiple server unit components (404, 406,..., 408), each component being coupled to a top-of-rack TOR switch 410. A top-of-rack switch is a switch that couples the components in the rack to other components in the data center. Although not shown, other racks may present a similar architecture. A rack is a physical structure for housing or otherwise grouping multiple processing components.
[0067] Figure 4 Also shown is an illustrative combination of a representative server unit component 404. It includes a local host component 412, which includes one or more central processing units (CPUs) (414, 416,...) and a local acceleration component 418. The local acceleration component 418 is directly coupled to the local host component 412 via a local link 420. The local link 420 can be implemented as a PCIe link, for example. The local acceleration component 418 is also indirectly coupled to the local host component 412 through a network interface controller (NIC) 422.
[0068] Finally, it should be noted that the local acceleration component 418 is coupled to the TOR switch 410. Thus, in this particular implementation, the local acceleration component 418 represents the only path through which the local host component 412 interacts with other components in the data center (including other host components and other acceleration components). Among other effects, Figure 4 the architecture allows the local acceleration component 418 to perform processing (e.g., by performing encryption, compression, etc.) on packets received from (and / or sent to) the TOR switch 410 without burdening the CPU-based operations performed by the local host component 412.
[0069] It should be noted that the local host component 412 can communicate with the local acceleration component 418 via the local link 420 or via the NIC 422. Different entities can utilize these two paths in different corresponding situations. For example, assume that a program running on the local host component 412 requests a service. In one implementation, assume that the local host component 412 provides local instantiations of the location determination component 124 and the data storage device 126. Alternatively, the global management component can provide the location determination component 124 and its data storage device 126. In either case, the local host component 412 can consult the data storage device 126 to determine the address of the service. Then, the local host component 412 can access the service using the identified address via the NIC 422 and the TOR switch 410.
[0070] In another implementation, assume that the local acceleration component 418 provides local instantiations of the location determination component 124 and the data storage device 126. The local host component 412 can access the local acceleration component 418 via the local link 420. The local acceleration component 418 can then consult the local data storage device 126 to determine the address of the service, and based on that address, it accesses the service via the TOR switch 410. There are other possible ways to access the service as well.
[0071] Figure 5 is Figure 4 a more inclusive depiction of the data center 402 shown. The data center 402 includes multiple racks (502 to 512,...). Each rack includes multiple server unit components. Each server unit component can in turn have the architecture described above in Figure 4 For example, the representative server unit component 514 includes local host component(s) 516, network interface controller (N) 518, and local acceleration component(s) 520.
[0072] Figure 5 The routing infrastructure shown above with reference to Figure 1corresponds to one implementation of the described common network 120. The routing infrastructure includes a plurality of top-of-rack (TOR) switches 522 and a higher-layer switching infrastructure 524. The higher-layer switching infrastructure 524 connects the TOR switches 522 together. The higher-layer switching infrastructure 524 can have any architecture and can be driven by any routing protocol. In Figure 5 the illustrated example, the higher-layer switching infrastructure 524 includes at least a set of aggregation switches 526, core switches 528, etc. Traffic routed through the illustrated infrastructure can correspond to Ethernet IP packets.
[0073] Figure 5 The illustrated data center 402 can correspond to a set of resources provided at a single geographical location, or a distributed set of resources distributed across multiple geographical locations (e.g., multiple individual contributing data centers located in different parts of the world). In a distributed context, the management function 122 can send work from a first contributing data center to a second contributing data center based on any mapping considerations, such as: (1) determining that acceleration components are available at the second contributing data center; (2) determining that the acceleration components are configured to perform the desired service or services at the second contributing data center; and / or (3) determining that the acceleration components are not only configured to perform the desired service or services, but they are immediately available (e.g., “online”) to perform those services, etc. As used herein, the term “global” generally refers to any scope that is more inclusive than the local domain associated with a single server unit component.
[0074] Generally, it should be noted that although Figure 4 and Figure 5 the focus is on using relatively broad data processing systems (corresponding to data centers), some of the principles set forth herein can be applied to smaller systems, including cases where a single local host component (or other type of component) is coupled to multiple acceleration components, and the smaller system includes local acceleration components and one or more remote acceleration components. Such smaller systems can even be implemented in user devices or appliances, etc. The user device can have the option of using local acceleration resources and / or remote acceleration resources.
[0075] Compared with Figure 4 the illustrated architecture, Figure 6 shows an alternative way of implementing the server unit component 602. Similar to Figure 4 the case of Figure 6 the server unit component 602 includes a local host component 604, which consists of one or more CPUs (606, 608,...), local acceleration components 610, and a local link 612 for coupling the local host component 604 to the local acceleration components 610. Compared withFigure 4 In a different case, the server unit component 602 implements a network interface controller (NIC) 614 as an internal component of the local acceleration component 610 rather than as a separate component.
[0076] In contrast to Figure 4 the architecture shown, Figure 7 FIG. shows another alternative way of implementing the server unit component 702. In Figure 7 this case, the server unit component 702 includes any number n of local host components (704, ..., 706) and any number m of local acceleration components (708, ..., 710) (other components of the server unit component 702 are omitted from the figure for ease of explanation). For example, the server unit component 702 may include a single host component coupled to two local acceleration components. The two acceleration components may perform different corresponding tasks. For example, one acceleration component may be used to process outgoing traffic to its local TOR switch, while the other acceleration component may be used to process incoming traffic from the TOR switch. Additionally, the server unit component 702 may load any service on any one of the local acceleration components (708, ..., 710).
[0077] It should also be noted that in the examples set forth above, the server unit component may refer to, for example, a physical grouping of components by forming a single servable unit within a rack of a data center. In other cases, the server unit component may include one or more host components and one or more acceleration components, which do not necessarily have to be housed together in a single physical unit. In such cases, the local acceleration components may be considered to be logically associated with their corresponding local host components rather than physically associated.
[0078] Alternatively or additionally, the local host component and one or more remote acceleration components may be implemented on a single physical component such as a single MPSoC-FPGA die. A network switch may also be incorporated into the single component.
[0079] In contrast to Figure 1 that shown, Figure 8 FIG. shows an alternative data processing system 802. Similar to Figure 1 the data processing system 102, the data processing system 802 includes a software plane 104 and a hardware acceleration plane 106, and a local H-to-local S coupling 114 for connecting the local host component to the corresponding local acceleration component. However, in contrast to Figure 1Different from the data processing system 102, the data processing system 802 includes a first network 804 for coupling host components together and a second network 806 for coupling hardware components together, where the first network 804 is at least partially different from the second network 806. For example, the first network 804 may correspond to the type of data center switching architecture shown in Figure 5 The second network 806 may correspond to dedicated links with any network topology for connecting acceleration components together. For example, the second network 806 may correspond to a p×r ring network. Each acceleration component in the ring network is coupled to adjacent acceleration components to the east, west, north, and south via appropriate cable links and the like. Alternatively, other types of ring networks may be used, which have any corresponding sizes and dimensions.
[0080] In other cases, a local hard CPU, and / or a soft CPU and / or acceleration logic provided by a single processing component (e.g., implemented on a single die) may be coupled to other elements of other processing components (e.g., implemented on other dies, boards, racks, etc.) via different networks. The individual services themselves may utilize one or more recursive local interconnection networks.
[0081] Further, it should be noted that the above description is constructed in the context of a service request being issued by a host component and being satisfied by an acceleration component. However, alternatively or additionally, any acceleration component may also make a service request that can be satisfied by any other component (e.g., another acceleration component and / or even a host component). The SMC 102 may resolve such requests in a manner similar to the way described above. In fact, some of the features described herein may be implemented separately on the hardware acceleration plane rather than the software plane.
[0082] More generally, some features may be implemented by any first component that requests a service, and the service may be satisfied by the first component, and / or one or more local components relative to the first component, and / or one or more remote components relative to the first component. However, for ease of explanation, the following description will mainly continue to be constructed in the context where the entity making the request corresponds to a local host component.
[0083] Finally, other implementations may adopt different strategies for coupling host components to hardware components, for example, different from Figure 14 The local H to local S coupling 114 shown.
[0084] Figure 9 Shows a representation of Figure 1Process 902 of an illustrative mode of operation of data processing system 102. In block 904, the local host component issues a request for a service. In block 906, the local host component receives a reply to the request, which may identify the address of the service. In an alternative implementation, the associated local acceleration component may perform blocks 904 and 906 after receiving the request from the local host component. In other words, either the local host component or the local acceleration component may perform the address lookup function.
[0085] In block 908, assuming that the identified address pertains to a function locally implemented by the local acceleration component, the associated local acceleration component may perform the service locally. Alternatively or additionally, in block 910, the local acceleration component routes the request to a remote acceleration component. As noted above, the local acceleration component is configured to perform the routing to the remote acceleration component without involving the local host component. Further, multiple host components communicate with each other in data processing system 102 via the same physical network as the multiple acceleration components.
[0086] In summary, for Section A, data processing system 102 has several useful features. First, data processing system 102 uses a common network 120 (except for the example of Figure 8 ), which avoids the costs associated with a custom network used to couple acceleration components together. Second, the common network 120 enables the addition of an acceleration plane to an existing data processing environment such as a data center. And after installation, the resulting data processing system 102 can be effectively maintained because it utilizes the existing physical links found in the existing data processing environment. Third, data processing system 102 integrates acceleration plane 106 without imposing large additional power requirements. For example, considering the manner described above, where the local acceleration component can be integrated with an existing server unit component. Fourth, data processing system 102 provides an effective and flexible mechanism for allowing host components to access any acceleration resources provided by hardware acceleration plane 106, for example, without strictly pairing the host component with a specific fixed acceleration resource and without burdening the host component by managing the hardware acceleration plane 106 itself. Fifth, data processing system 102 provides an effective mechanism for managing acceleration resources by intelligently distributing these resources across hardware plane 106, thereby: (a) reducing overutilization and underutilization of resources (e.g., corresponding to the "stranded capacity" problem); (b) facilitating rapid access to these services by consumers of these services; (c) accommodating increased processing requirements specified by some consumers and / or services, etc. The above effects are illustrative and not exhaustive; data processing system 102 also provides other useful effects.
[0087] B. Management Functions
[0088] Figure 10 illustrates an overview of one implementation of the management function 122 of the data processing system 102 for management. More specifically, Figure 1 depicts a logical view of the functions performed by the management function 122, which includes its main engine, the service mapping component (SMC) 128. Different sub-components correspond to different main functions performed by the management function 122. The various possible physical implementations of the logical functions described below Figure 10 are shown. Figures 17 to 20 illustrates various possible physical implementations of the logical functions.
[0089] As described in the introductory Section A, the location determination component 124 identifies the current location of a service within the data processing system 102 based on the current allocation information stored in the data storage device 126. In operation, the location determination component 124 receives a request for a service. In response, it returns the address of the service (if it exists within the data storage device 126). This address can identify the specific acceleration component that implements the service.
[0090] The data storage device 126 can maintain any type of information that maps services to addresses. In Figure 10 the small excerpt shown, the data storage device 126 maps a small number of services (service w, service x, service y, and service z) to the acceleration components that are currently configured to provide these services. For example, the data storage device 126 indicates that the configuration image for service w is currently installed on devices with addresses a1, a6, and a8. The address information can be expressed in any way. Here, for ease of explanation, the address information is represented in a high-level symbolic form.
[0091] In some implementations, data storage device 126 may also optionally store status information characterizing each current service-to-component assignment in any way. Generally, the status information for service-to-component assignments specifies how the assigned services, as implemented on the component (or components) to which it is assigned, are to be processed within data processing system 102, such as by specifying a persistence level, specifying its access rights (e.g., "ownership"), and so on. In one non-limiting implementation, for example, a service-to-component assignment may be designated as reserved or non-reserved. When performing configuration operations, SMC 128 may consider the reserved / non-reserved status information associated with the assignment when determining whether it is appropriate to change the assignment, e.g., to meet a current request for a service, a change in the demand for one or more services, and so on. For example, data storage device 126 indicates that acceleration components with addresses a1, a6, and a8 are currently configured to perform service w, but only the assignments to acceleration components a1 and a8 are considered to be reserved. Thus, compared to the other two acceleration components, SMC 128 views the assignment to acceleration component a6 as a more suitable candidate for re-assignment (re-configuration).
[0092] Additionally, or alternatively, data storage device 126 may provide information indicating whether a service-to-component assignment is to be shared by all instances of a tenant function or is dedicated to one or more specific instances of a tenant function (or consumers of some services otherwise indicated). In the former (fully shared) case, all tenant function instances compete for the same resources provided by the acceleration components. In the latter (dedicated) case, only those clients associated with the service assignment are permitted to use the assigned acceleration components. Figure 10 At a high level, it is shown that services x and y running on an acceleration component with address bit a3 are reserved for use by one or more specified tenant function instances, while any tenant function instance may use another service-to-component assignment.
[0093] SMC 128 may also interact with data storage device 1002, which provides availability information. The availability information identifies a pool of acceleration components having available capacity to implement one or more services. For example, in one usage, SMC 128 may determine that it is appropriate to assign one or more acceleration components as providers of a function. To do so, SMC 128 utilizes data storage device 1002 to find acceleration components having idle capacity to implement the function. SMC 128 then assigns the function to one or more of these idle acceleration components. Doing so changes the availability-related state of the selected acceleration components.
[0094] The SMC 128 also manages and maintains availability information in the data storage device 1002. In doing so, the SMC 128 can use different rules to determine whether an acceleration component is available or unavailable. In one approach, the SMC 128 can consider an acceleration component that is currently being used as unavailable, while considering an acceleration component that is not currently being used as available. In other cases, an acceleration component can have different configurable domains (e.g., tiles), some of which are currently being used and some of which are not currently being used. Here, the SMC 128 can specify the availability of an acceleration component by expressing a portion of its processing resources that are not currently being used. For example, Figure 10 indicates that the acceleration component with address a1 has 50% of its processing resources available. On the other hand, the acceleration component with address a2 is fully available, while the acceleration component with address a3 is fully unavailable. As will be described in more detail below, individual acceleration components can notify the SMC 128 of their relative utilization levels in different ways.
[0095] In other cases, the SMC 128 can consider pending requests for an acceleration component when registering the acceleration component as available or unavailable. For example, the SMC 128 can indicate that an acceleration component is unavailable because it is scheduled to deliver a service to one or more tenant function instances, even though it cannot engage in providing the service at the current time.
[0096] In other cases, the SMC 128 can also register the type of each available acceleration component. For example, the data processing system 102 can correspond to a heterogeneous environment that supports acceleration components with different physical characteristics. The availability information in such a case can indicate not only the identity of the available processing resources, but also the type of these resources.
[0097] In other cases, when registering an acceleration component as available or unavailable, the SMC 128 can also consider the state of service-to-component allocation. For example, assume that a specific acceleration component is currently configured to perform a certain service, and furthermore, assume that the allocation has been specified as reserved rather than non-reserved. Given its reserved state, the SMC 128 can specify the acceleration component as unavailable (or a portion of it as unavailable), regardless of whether the service is currently being actively used to perform a function at present. In fact, the reserved state of an acceleration component serves as a lock to prevent the SMC 128 from reconfiguring the acceleration component in at least some cases.
[0098] Now referring to the core mapping operation of the SMC 128 itself, the SMC 128 allocates or maps services to the acceleration components in response to a trigger event. More specifically, the SMC 128 operates in different modes according to the type of the received trigger event. In the request-driven mode, the SMC 128 processes requests from tenant functions for services. Here, each trigger event corresponds to a request made by a tenant function instance that resides at least partially on a specific local host component. In response to each request from the local host component, the SMC 128 determines the appropriate component to implement the service. For example, the SMC 128 can choose from the following: a local acceleration component (associated with the local host component that issued the request), a remote acceleration component, or the local host component itself (subsequently, the local host component will implement the service in software), or some combination thereof.
[0099] In the second background mode, the SMC 128 operates by globally allocating services to the acceleration components within the data processing system 102 to meet the overall expected demands in the data processing system 102 and / or to meet other system-wide goals and other factors (rather than just focusing on individual requests from host components). Here, each received trigger event corresponds to some condition in the data processing system 102 as a whole, which guarantees the allocation (or reallocation) of the service, such as a change in the demand for the service.
[0100] However, it should be noted that the modes described above are not mutually exclusive domains of analysis. For example, in the request-driven mode, the SMC 128 can attempt to achieve at least two goals. As a first primary goal, the SMC 128 will attempt to find the acceleration component (or components) that satisfy the request for the service to be solved, while meeting one or more performance goals related to the data processing system 102 as a whole. As a second goal, regarding the future use of the service by other tenant function instances, the SMC 128 can optionally consider the long-term impact of its allocation of the service. In other words, the second goal involves background considerations triggered exactly by the requests of specific tenant function instances.
[0101] For example, consider the following simplified scenario. A tenant function instance can make a request for a service, where the tenant function instance is associated with a local host component. The SMC 128 can execute the service by configuring the local acceleration component in response to the request. When making this decision, the SMC 128 first attempts to find an allocation that satisfies the request of the tenant function instance. However, the SMC 128 can also make its allocation based on the determination that many other host components have requested the same service, and most of these host components are located in the same rack as the tenant function instance that has generated the current request for the service. In other words, this additional discovery further supports the decision to place the service on the in-rack acceleration component.
[0102] Figure 10 Depicts SMC 128 that optionally includes multiple logic components that perform different respective analyses. As a first optional component of the analysis, SMC 128 can use state determination logic 1004 to define the state of the allocation it is making, e.g., as reserved or non-reserved, dedicated or fully shared, etc. For example, assume SMC 128 receives a request from a tenant function instance for a service. In response, SMC 128 can decide to configure a local acceleration component to provide the service, and in the process, designate the allocation as non-reserved, e.g., under the initial assumption that the request can be a "one-time" request for the service. In another scenario, assume SMC 128 makes an additional determination that the same tenant function instance has repeatedly made requests for the same service within a short period of time. In this case, SMC 128 can make the same allocation decision as described above, but this time SMC 128 can designate it as reserved. SMC 128 can also optionally designate the service as dedicated only to the requesting tenant function. By doing so, SMC 128 can enable the data processing system 102 function to more effectively meet future requests from the tenant function example for the service. In other words, when the local acceleration component is heavily used by the local host component, the reserved state may reduce the chance for SMC 128 to move the service away from the local acceleration component at a later time.
[0103] In addition, a tenant function (or local host component) instance can specifically request that it be granted reserved and dedicated use of a local acceleration component. The state determination logic 1004 can use different environment-specific rules when determining whether to honor the request. For example, the state determination logic 1004 can decide to honor the request as long as no other trigger event is received that causes the request to be overridden. The state determination logic 1004 can override the request, e.g., when it attempts to fulfill another request that is determined to be more urgent than the tenant function's request for any environment-specific reason.
[0104] In some implementations, it should be noted that a tenant function (or local host component or some other consumer of a service) instance can independently control the use of its local resources. For example, the local host component can pass utilization information to the management function 122, which indicates that its local acceleration component is unavailable or not fully available, regardless of whether the local acceleration component is actually busy at the moment. By doing so, the local host component may cause SMC 128 to "steal" its local resources. Different implementations can use different environment-specific rules to determine whether an entity is permitted to restrict access to its local resources in the manner described above, and if so, under what circumstances.
[0105] In another example, assume that the SMC 128 determines that there is a general increase in demand for a specific service. In response, the SMC 128 can find a specified number of idle acceleration components corresponding to a "pool" of acceleration components, and then designate the pool of acceleration components as a reserved (but fully shared) resource for providing the specific service. After that, the SMC 128 can detect a general decrease in demand for the specific service. In response, the SMC 128 can reduce the pool of reserved acceleration components, for example, by changing the status of one or more acceleration components previously registered as "reserved" to "not reserved".
[0106] It should be noted that the specific dimensions of the states described above (reserved vs. not reserved, dedicated vs. fully shared) are referred to by way of illustration and not limitation. Other implementations can employ any other state-related dimensions, or can simply accommodate a single state designation (thus omitting the use of the functionality of the state determination logic 1004).
[0107] As a second component of the analysis, the SMC 128 can use sizing determination logic 1006 to determine the number of acceleration components suitable for providing the service. The SMC 128 can make such a determination based on considerations of the processing requirements associated with the service and the resources available to meet those processing requirements.
[0108] As a third component of the analysis, the SMC 128 can use type determination logic 1008 to determine the type of acceleration components suitable for providing the service. For example, consider the case where the data processing system 102 has a heterogeneous collection of acceleration components with different corresponding capabilities. The type determination logic 1008 can determine one or more acceleration components of a specific type suitable for providing the service.
[0109] As a fourth component of the analysis, the SMC 128 can use layout determination logic 1010 to determine the specific acceleration component (or components) suitable for addressing a specific trigger event. This decision, in turn, can have one or more aspects. For example, as part of its analysis, the layout determination logic 1010 can determine whether it is appropriate to configure an acceleration component to perform the service, where the component is not currently configured to perform the service.
[0110] The above aspects of the analysis are referred to by way of illustration and not limitation. In other implementations, the SMC 128 can provide additional analysis phases.
[0111] Generally, the SMC 128 performs its various allocation determinations based on one or more mapping considerations. For example, one mapping consideration can involve historical demand information provided in the data storage device 1012.
[0112] However, it should be noted that SMC 128 does not need to perform a multi-factor analysis in all cases. In some cases, for example, the host component can make a request for a service associated with a single fixed location (e.g., corresponding to a local acceleration component or a remote acceleration component). In those cases, SMC 128 can simply follow the location determination component 124 to map the service request to the address of the service, rather than evaluating the costs and benefits of performing the service in a different way. In other cases, the data storage device 126 can associate multiple addresses with a single service, each address associated with an acceleration component that can perform the service. SMC 128 can use any mapping considerations, such as load balancing considerations, when allocating requests for a service to a specific address.
[0113] As a result of its operation, SMC 128 can update the data storage device 126 with information that maps services to the addresses where those services can be found (assuming the information has been changed by SMC 128). SMC 128 can also store state information related to the new service-to-component assignment.
[0114] To configure one or more acceleration components to perform a function (if not already so configured), SMC 128 can call the configuration component 1014. In one implementation, the configuration component 1014 configures the acceleration components by sending a configuration stream to the acceleration components. The configuration stream specifies the logic to be "programmed" into the receiving acceleration component. The configuration component 1014 can use different strategies to configure the acceleration components, several of which are elaborated below.
[0115] The fault monitoring component 1016 determines whether an acceleration component has failed. SMC 128 can respond to a fault notification by replacing the failed acceleration component with a spare acceleration component.
[0116] B.1. Operation of SMC in Request-Driven Mode
[0117] Figure 11 Provide an overview of an operation mode of SMC 128 when applied to tasks that process requests for tenant function instances running on a host component. In the illustrated scenario, assume that the host component 1102 implements multiple tenant functions (T 1 , T 2 ,..., T n ) instances. Each tenant function instance can correspond to a software program that is at least partially executed on the host component 1102, e.g., in a virtual machine that uses the physical resources of the host component 1102 (in addition to other possible host components). Further assume that a tenant function instance initiates by generating a request for a specific service Figure 11The transactions shown. For example, a tenant function can perform a photo editing function and can call a compression service as part of its overall operation. Or, a tenant function can perform a search algorithm and can call a ranking service as part of its overall operation.
[0118] In operation (1), the host component 1102 can send its request for a service to the SMC 128. In operation (2), among other analyses, the SMC 128 can determine at least one suitable component to implement the service. In this case, assume that the SMC 128 determines that the remote acceleration component 1104 is the most suitable component to implement the service. The SMC 128 can obtain the address of the remote acceleration component 1104 from the location determination component 124. In operation (3), the SMC 128 can communicate its answer to the host component 1102, for example, in the form of an address associated with the service. In operation (4), the host component 1102 can invoke the remote acceleration component 1104 via its local acceleration component 1106. Other ways of handling requests from tenant functions are possible. For example, the local acceleration component 1106 can query the SMC 128 instead of, or in addition to, the local host component 1102.
[0119] Path 1108 represents an example in which the representative acceleration component 1110 (and / or its associated local host component) communicates utilization information to the SMC 128. The utilization information can identify whether the acceleration component 1110 is fully or partially available or unavailable. The utilization information can also optionally specify the type of processing resources available that are owned by the acceleration component 1110. As noted above, the utilization information can also be selected to purposefully organize the SMC 128 to utilize the resources of the acceleration component 1110 later, for example, by indicating that the resources are unavailable in whole or in part.
[0120] Although not shown, any acceleration component can also make a directed request to the SMC 128 for a specific resource. For example, the host component 1102 can specifically request to use its local acceleration component 1106 as a reserved and dedicated resource. As noted above, the SMC 128 can use different environment-specific rules when determining whether to honor such a request.
[0121] Further, although not shown, components other than the host component can make requests. For example, a hardware acceleration component can run a tenant function instance that issues a request for a service that can be satisfied by itself, another hardware acceleration component (or multiple acceleration components), a host component (or multiple components), etc., or any combination thereof.
[0122] Figures 12 to 15Shows different corresponding options for processing requests for services made by tenant functions residing on a host component. Starting from Figure 12 Assume that the local host component 1202 includes at least two tenant function instances T1(1204) and T2(1206), both running simultaneously (but in reality, the local host component 1202 can host more tenant function instances). The first tenant function instance T1 requires the acceleration service A1 to perform its operations, while the second tenant function instance T2 requires the acceleration service A2 to perform its operations.
[0123] Further assume that the local acceleration component 1208 is coupled to the local host component 1202 via, for example, a PCIe local link, etc. At the current moment, the local acceleration component 1208 hosts the A1 logic 1210 for performing the acceleration service A1, and the A2 logic 1212 for performing the acceleration service A2.
[0124] According to an administrative decision, the SMC 128 assigns T1 to the A1 logic 1210 and assigns T2 to the A2 logic 1212. However, the decision made by the SMC 128 is not a fixed rule; as will be described, the SMC 128 can make its decision based on multiple factors, some of which may reflect conflicting considerations. Thus, based on other factors (not described at this time), the SMC 128 can choose to assign jobs to the acceleration logic in a way different from Figure 12 the way shown.
[0125] In Figure 13 the scenario, the host component 1302 has the same tenant function instances (1304, 1306) with the same service requirements as those described above. However, in this case, the local acceleration component 1308 only includes the A1 logic 1310 for performing the service A1. That is, it no longer hosts the A2 logic for performing the service A2.
[0126] In response to the above scenario, the SMC 128 can choose to assign T1 to the A1 logic 1310 of the acceleration component 1308. Then, the SMC 128 can assign T2 to the A2 logic 1312 of the remote acceleration component 1314, which has been configured to perform the service. Again, the illustrated assignment is presented here in the spirit of illustration rather than limitation; the SMC 128 can choose a different allocation based on another combination of input considerations. In one implementation, the local host component 1302 and the remote acceleration component 1314 can optionally compress the information they send to each other, for example, to reduce bandwidth consumption.
[0127] It should be noted that the host component 1302 accesses the A2 logic 1312 via the local acceleration component 1308. However, in another case (not shown), the host component 1302 may access the A2 logic 1312 via a local host component (not shown) associated with the acceleration component 1314.
[0128] Figure 14 Another scenario is presented in which the host component 1402 has the same tenant function instances (1404, 1406) with the same service requirements as described above. In this case, the local acceleration component 1408 includes A1 logic 1410 for performing service A1 and A3 logic 1412 for performing service A3. Further assume that the availability information in the data storage device 1002 indicates that the A3 logic 1412 is not currently being used by any tenant function instance. In response to the above scenario, the SMC 128 may use the ([ Figure 10 of) configuration component 1014 to reconfigure the acceleration component 1408 such that it includes A2 logic 1414 and does not include A3 logic 1412 (as shown at the bottom of Figure 14 ). Then, the SMC 128 may assign T2 to the A2 logic 1414. Although not shown, the SMC 128 may alternatively or additionally decide to reconfigure any remote acceleration component to perform the A2 service.
[0129] Generally, the SMC 128 may perform the configuration in a complete or partial manner to meet any requests of the tenant function instances. The SMC performs a complete configuration by reconfiguring all the application logics provided by the acceleration component. The SMC 128 may perform a partial configuration by reconfiguring a part (e.g., one or more tiles) of the application logics provided by the acceleration component, thereby keeping other parts (e.g., one or more other tiles) intact and operable during the reconfiguration. The same applies to the operation of the SMC 128 in the background operation mode described below. Further, it should be noted that additional factors may play a role in determining whether the A3 logic 1412 is a valid candidate for reconfiguration, such as whether the service is considered reserved, whether there are pending requests for the service, etc.
[0130] Figure 15Presents another scenario where the host component 1502 has the same tenant function instances (1504, 1506) with the same service requirements as those described above. In this case, the local acceleration component 1508 only includes the A1 logic 1510 for executing Service A1. In response to the above scenario, the SMC 128 can assign T1 to the A1 logic 1510. Further, assume that the SMC 128 determines that it is not feasible to execute Service A2 for any acceleration component. In response, if the logic is actually available at the host component 1502, the SMC 128 can instruct the local host component 1502 to assign T2 to the local A2 software logic 1512. The SMC 128 can make the Figure 15 decision based on various reasons. For example, the SMC 128 may conclude that hardware acceleration is not possible because there is currently no configuration image for the service. Or, there may be a configuration image, but the SMC 128 concludes that the capacity of any of the acceleration devices in the acceleration device is not sufficient to load and / or run such a configuration.
[0131] Finally, the above examples have been described in the context of tenant function instances running on host components. But as already pointed out above, tenant function instances can more generally correspond to service requesters, and those service requesters can run on any component including acceleration components. Thus, for example, a requester running on an acceleration component can generate a request for a service to be executed by one or more other acceleration components and / or by itself and / or by one or more host components. The SMC 102 can process the requester's request in any of the ways described above.
[0132] B.2. SMC Operation in Background Mode
[0133] Figure 16Provides an overview of an operating mode of the SMC 128 when operating in the background mode. In operation (1), the SMC 128 can receive a certain type of trigger event that initiates the operation of the SMC 128. For example, the trigger event can correspond to a change in demand affecting the service, etc. In operation (2), in response to the trigger event, the SMC 128 determines the allocation of one or more services to the acceleration components based on one or more mapping considerations and the availability information in the data storage device 1002. For example, by assigning the service to a set of one or more available acceleration components. In operation (3), the SMC 128 executes its allocation decision. As part of this process, the SMC 128 can call the configuration component 1014 to configure the acceleration components that have been assigned to perform the service, assuming these components have not been configured to perform the service. The SMC 128 also updates the service location information in the data storage device 126 and, if appropriate, the availability information in the data storage device 1002.
[0134] In Figure 16 a specific example of, the SMC 102 assigns the first acceleration component group 1602 to execute the first service ("Service y"), and assigns the second acceleration component group 1604 to execute the second service ("Service z"). In actual practice, the assigned acceleration component groups can have any number of members, and these members can be distributed across the hardware acceleration plane 106 in any way. However, the SMC 128 can attempt to group the acceleration components associated with the service in a specific way to achieve satisfactory bandwidth and latency performance (among other factors). The SMC 128 can apply further analysis when allocating the acceleration components associated with a single multi-component service.
[0135] The SMC 128 can also operate in the background mode to assign one or more acceleration components that implement a specific service to at least one tenant function instance, without necessarily requiring the tenant function to make a request for this specific service each time. For example, assume that the tenant function instance often uses the compression function corresponding to Figure 16 "Service z" in. The SMC 128 can proactively assign one or more dedicated second acceleration component groups 1604 to at least this tenant function instance. When the tenant function needs to use the service, it can obtain it from the available address pool associated with the second acceleration component group 1604 that has been assigned to it. The same dedicated mapping operation can be performed with respect to a group of tenant function instances (rather than a single instance).
[0136] B.3. Physical implementation manner of the management function
[0137] Figure 17 Shows Figure 10The first physical implementation of the management function 122. In this case, the management function 122 is provided on a single global management component (M G ) 1702 or on multiple global management components (1702,..., 1704). If multiple global management components (1702,..., 1704) are used, they can provide redundant logic and information to achieve the required load balancing and fault management performance. In one case, each global management component can be implemented on a computer server device that can correspond to one of the host components in the host component or a dedicated management computing device. In operation, any individual host component (S) or acceleration component (H) can interact with the global management component via Figure 1 the common network 120 shown.
[0138] Figure 18 Illustrates Figure 10 The second physical implementation of the management function 122. In this case, each server unit component (such as the representative server unit component 1802) provides at least one local management component (M L ) 1804. For example, the local host component 1806 can implement the local management component 1804 (e.g., as part of its hypervisor function), or the local acceleration component 1808 can implement the local management component 1804, or some other component within the server unit component 1802 can implement the local management component 1804 (or some combination thereof). The data processing system 102 also includes one or more global management components (1810,..., 1812). Each global management component can provide redundant logic and information in the manner described above with respect to Figure 17 . As elaborated above, the management function 122 collectively presents all the local and global management components in the data processing system 102.
[0139] Figure 18 The architecture of can implement the request-driven aspect of the SMC 128 in the following manner. The local management component 1804 can first determine whether the local acceleration component 1808 can execute the service requested by the tenant function. In the case where the local acceleration component 1808 cannot perform the task, the global management component (M G ) can make other decisions, such as identifying a remote acceleration component for performing the service. On the other hand, in Figure 17 's architecture, a single global management component can make all the decisions involved in mapping requests to acceleration components.
[0140] Further, the local management component 1804 can send utilization information to the global management component on any basis such as a periodic basis and / or an event-driven basis (e.g., in response to a utilization change). The global management component can use the utilization information to update the master record of its availability information in the data storage device 1002.
[0141] Figure 19 illustrates Figure 10 a third physical implementation of the management function 122. In this case, each server unit component stores its own dedicated local management component (M L )(which can be implemented by the local host component, the local acceleration component, some other local component, or some combination thereof as part of its hypervisor function). For example, the server unit component 1902 provides the local management component 1904, as well as the local host component 1906 and the local acceleration component 1908. Similarly, the server unit component 1910 provides the local management component 1912, as well as the local host component 1914 and the local acceleration component 1916. Each instance of the local management component stores redundant logic and information about other instances of the same component. Known distributed system tools can be used to ensure that all distributed versions of the component contain the same logic and information, such as the ZOOKEEPER tool provided by the Apache Software Foundation of Forest Hill, Maryland. (In addition, it should be noted that the same technology can be used to maintain redundant logic and information in other examples described in this subsection.) As elaborated above, the management function 122 centrally presents all the local management components in the data processing system 102. That is, there is no central global management component in this implementation.
[0142] Figure 20 illustrates Figure 10 a fourth physical implementation of the management function 122. In this case, the management function 122 implements a hierarchy of individual management components. For example, in one merely representative structure, each server unit component includes a lower-level local management component (M L3)(which may be implemented by a local host component, a local acceleration component, some other local components, or some combination thereof). For example, server unit component 2002 provides a lower-level local management component 2004, as well as a local host component 2006 and a local acceleration component 2008. Similarly, server unit component 2010 provides a lower-level local management component 2012, as well as a local host component 2014 and an acceleration component 2016. The next management level of this structure includes at least a middle-level management component 2018 and a middle-level management component 2020. The top level of this structure includes a single global management component 2022 (or multiple redundant such global management components). Thus, the illustrated control architecture forms a structure with three levels, but the architecture may have any number of levels.
[0143] In operation, the low-level management components (2004, 2012, ...) handle certain low-level management decisions that directly affect resources associated with individual server unit components. The middle-level management components (2018, 2020) may make decisions that affect relevant parts of the data processing system 102, such as individual racks or groups of racks. The top-level management component (2022) may make global decisions that are widely applicable to the entire data processing system 102.
[0144] B.4. Configuration Components
[0145] Figures 21 to 24 Illustrates different corresponding strategies for configuring the acceleration component, which correspond to different ways of implementing Figure 10 the configuration component 1014. Starting from Figure 21 the global management component 2102 accesses the data storage device 2104, which provides one or more configuration images. Each configuration image contains the logic that can be used to implement the corresponding service. The global management component 2102 can configure the acceleration component by forwarding a configuration stream (corresponding to the configuration image) to the acceleration component. For example, in one method, the global management component 2102 can send the configuration stream to the local management component 2106 associated with a specific server unit component 2108. The local management component 2106 can then coordinate the configuration of the local acceleration component 2110 based on the received configuration stream. Alternatively, the local host component 2112 can perform the operations described above instead of, or in addition to, the local management component 2106.
[0146] Figure 22Another strategy for configuring the acceleration component is shown. In this case, the global management component 2202 sends instructions to the local management component 2204 of the server unit component 2206. In response, the local management component 2204 accesses the configuration image in the local data storage device 2208 and then uses it to configure the local acceleration component 2210. Alternatively, the local host component 2212 can perform the operations described above in place of, or in addition to, the local management component 2204.
[0147] Figure 23 Another technique for configuring the local acceleration component 2302 is shown. In this approach, it is assumed that the acceleration component 2302 includes application logic 2304, which in turn is governed by the current model 2306 (where the model corresponds to the logic that performs functions in a specific manner). It is further assumed that the acceleration component 2302 can access the local storage device 2308. The local storage device 2308 stores configuration images associated with one or more other models (Model 1,..., Model n). When triggered, the local model loading component 2310 can replace the configuration associated with the current model 2306 with a configuration associated with another model in the local storage device 2308. The model loading component 2310 can be implemented by the acceleration component 2302 itself, the host component, the local management component, etc., or a combination thereof. In one implementation, Figure 23 the configuration operation shown can be performed in less time than a total reconfiguration of the overall application logic 2304, because it requires replacing some of the logic used by the application logic 2304 rather than replacing the entire application logic 2304 in a large-scale manner.
[0148] Finally, Figure 24 an acceleration component with application logic 2402 that supports partial configuration is shown. The management function 122 can utilize this function by configuring Application 1 (2404) separately from Application 2 (2406), and vice versa.
[0149] C. Illustrative Implementations of Hardware Acceleration Components
[0150] Figure 25 An implementation of the acceleration component 2502 in the Figure 1 data processing system shown, which can be physically implemented as an FPGA device. It should be noted that the details presented below are elaborated in the spirit of illustration rather than limitation; compared with the Figure 25 acceleration component shown, other data processing systems can use acceleration components with architectures that vary in one or more ways. Further, other data processing systems can adopt heterogeneous designs including acceleration components of different types.
[0151] From a high-level perspective, the acceleration component 2502 can be implemented as a hierarchy with different functional layers. At the lowest level, the acceleration component 2502 provides a "shell" that provides components related to a basic interface that generally remains the same across most application scenarios. The core component 2504 located inside the shell can include an "inner shell" and application logic 2506. The inner shell corresponds to all resources in the core component 2504 except for the application logic 2506 and represents a second level of resources that remains the same in a certain set of application scenarios. The application logic 2506 itself represents the highest level of resources that are most vulnerable to change. However, it should be noted that any component of the acceleration component 2502 can be reconfigured technically.
[0152] In operation, the application logic 2506 interacts with the shell resources and the inner shell resources in a manner similar to how a software-implemented application interacts with its underlying operating system resources. From the perspective of application development, using common shell resources and inner shell resources can save developers from having to recreate these common components for each application they create. This strategy also reduces the risk that developers may change the inner shell or shell functionality in a way that causes problems within the overall data processing system 102.
[0153] First referring to the shell, the acceleration component 2502 includes a bridge 2508 that is used to couple the acceleration component 2502 to a network interface controller (via the NIC interface 2510) and a local top-of-rack switch (via the TOR interface 2512). The bridge 2508 supports two modes. In the first mode, the bridge 2508 provides a data path that allows traffic from the NIC or TOR to flow into the acceleration component 2502 and traffic from the acceleration component 2502 to flow out to the NIC or TOR. The acceleration component 2502 can perform any processing on the traffic it "intercepts", such as compression, encryption, etc. In the second mode, the bridge 2508 supports a data path that allows traffic to flow between the NIC and the TOR without further processing by the acceleration component 2502. Internally, the bridge can consist of various FIFOs (2514, 2516) that buffer the received packets and various selectors and arbitration logics that route the packets to their desired destinations. The bypass control component 2518 controls whether the bridge 2508 operates in the first mode or the second mode.
[0154] The memory controller 2520 governs the interaction between the acceleration component 2502 and local memory 2522 (such as DRAM memory). The memory controller 2520 can perform error correction as part of its services.
[0155] The host interface 2524 provides a means for the acceleration component to communicate with local host components ( Figure 25The functions of the interaction (not shown in the figure) are not shown. In one implementation, the host interface 2524 can use the Peripheral Component Interconnect Express (PCIe) in combination with Direct Memory Access (DMA) to exchange information with local host components.
[0156] Finally, the shell can include various other features 2526, such as a clock signal generator, status LEDs, error correction functions, and so on.
[0157] In one implementation, the inner shell can include a router 2528 that routes messages between various internal components of the acceleration component 2502 and between the acceleration component 2502 and external entities (via the transport component 2530). Each such endpoint is associated with a corresponding port. For example, the router 2528 is coupled to the memory controller 2520, the host interface 1120, the application logic 2506, and the transport component 2530.
[0158] The transport component 2530 schedules packets for transmission to a remote entity (such as a remote acceleration component) and receives packets from a remote acceleration component (such as a remote acceleration component).
[0159] When activated, the 3-port switch 2532 takes over the function of the bridge 2508 by routing packets between the NIC and the TOR, and between the NIC or the TOR and the local ports associated with the acceleration component 2502 itself.
[0160] Finally, the optional diagnostic recorder 2534 stores transaction information about the operations performed by the router 2528, the transport component 2530, and the 3-port switch 2532 in a circular buffer. For example, the transaction information can include data about the source and destination IP addresses of the packets, host-specific data, timestamps, and so on. Technicians can study the log of the transaction information to try to diagnose the causes of faults or suboptimal performance in the acceleration component 2502.
[0161] Figure 26 The acceleration component 3202 is shown, which includes separate configurable domains (2604, 2606,...). The configuration component (e.g., Figure 10 the configuration component 1014) can configure each configurable domain without affecting the other configurable domains. Therefore, the configuration component 1014 can configure one or more configurable domains while other configurable domains are performing operations based on their respective configurations, and these configurations are not disturbed.
[0162] In some implementations, Figure 1The data processing system 102 can dynamically reconfigure its acceleration components to address any mapping considerations. This reconfiguration can be performed on a partial and / or full service basis and can be performed on a periodic and / or event-driven basis. In fact, in some cases, the data processing system 102 may appear to be continuous in the process of adapting itself to changing conditions in the data processing system 102 by reconfiguring its acceleration logic.
[0163] C.1. Local Link
[0164] Figure 27 Illustrates the function by which the local host component 2702 can forward information to its local acceleration component 2704 via Figure 25 the host interface 2524 shown (e.g., using PCIe in conjunction with DMA memory transfers). In one non-limiting protocol, in operation (1), the host logic 2706 places the data to be processed into a kernel-pinned input buffer 2708 in the main memory associated with the host logic 2706. In operation (2), the host logic 2706 instructs the acceleration component 2704 to retrieve the data and start processing it. The thread of the host logic is then placed in a sleep state until it receives a notification event from the acceleration component 2704 or it continues to process other data asynchronously. In operation (3), the acceleration component 2704 transfers data from the host logic's memory and places it in the acceleration component input buffer 2710.
[0165] In operations (4) and (5), the application logic 2712 retrieves data from the input buffer 2710, processes it to generate an output result, and places the output result in the output buffer 2714. In operation (6), the acceleration component 2704 copies the contents of the output buffer 2714 to an output buffer in the host logic's memory. In operation (7), the acceleration component notifies the host logic 2706 that the data is ready to be retrieved. In operation (8), the host logic thread wakes up and consumes the data in the output buffer 2716. Then, the host logic 2706 can discard the contents of the output buffer 2716, which allows the acceleration component 2704 to reuse it in the next transaction.
[0166] C.2. Router
[0167] Figure 28 Illustrates in Figure 25One implementation of router 2528 introduced in []. The router includes any number of input units (here four, 2802, 2804, 2806, 2808) for receiving messages from corresponding ports and any number of output units (here four, 2810, 2812, 2814, 2814) for forwarding messages to corresponding ports. As described above, the endpoints associated with the ports include memory controller 2520, host interface 2524, application logic 2506, and transport component 2530. The crossbar component 2818 forwards messages from input ports to output ports based on the address information associated with the messages. More specifically, a message consists of a plurality of "flits", and router 2528 sends messages on a flit-by-flit basis.
[0168] In one non-limiting implementation, router 2528 supports a number of virtual channels (such as eight) for carrying different classes of traffic on the same physical link. That is, for those scenarios where multiple services are implemented by application logic 2506, router 2528 can support multiple traffic classes, and these traffic classes need to communicate on separate classes of traffic.
[0169] Router 2528 can use credit-based flow control techniques to manage access to the resources of the router (e.g., its available buffer space). In this technique, the input units (2802 to 2808) provide credits to upstream entities corresponding to the exact number of flits available in their buffers. This credit authorizes the upstream entity to transmit its data to the input units (2802 to 2808). More specifically, in one implementation, router 2528 supports "elastic" input buffers that can be shared among multiple virtual channels. The output units (2810 to 2816) are responsible for tracking the available credits in their downstream receivers and providing authorization to any input unit (2802 to 2808) that requests to send a flit to a given output port.
[0170] C.3. Transport Component
[0171] Figure 29 Shows one implementation of transport component 2530 introduced in Figure 25 []. The transport component 2530 can provide a register interface to establish connections between nodes. That is, each such connection is unidirectional and links a send queue on a source component to a receive queue on a destination component. Before the transport component 2530 can transmit or receive data, a software process can establish the connection by statically allocating the connection. The data storage device 2902 stores two tables that control the connection state, a send connection table and a receive connection table.
[0172] The packet processing component 2904 processes messages arriving from the router 2528 and destined for a remote endpoint (e.g., another acceleration component). It does so by buffering and packetizing the messages. The packet processing component 2904 also processes packets received from some remote endpoints and destined for the router 2528.
[0173] For messages arriving from the router 2528, the packet processing component 2904 matches each message request to a transmit connection table entry in the transmit connection table, e.g., using the header information and virtual channel (VC) information associated with the message as provided by the router 2528 as query items. The packet processing component 2904 uses the information retrieved from the transmit connection table entry (such as sequence numbers, address information, etc.) to construct the packets it sends to the remote entity.
[0174] More specifically, in one non-limiting approach, the packet processing component 2904 encapsulates the packets in UDP / IP Ethernet frames and sends them to the remote acceleration component. In one implementation, the packet can include an Ethernet header, followed by an IPv4 header, followed by a UDP header, followed by a transport header (specifically associated with the transport component 2530), and then followed by the payload.
[0175] For packets arriving from the network (e.g., received on the local port of the 3-port switch 2532), the packet processing component 2904 matches each packet to a receive connectable table entry provided in the packet header. If there is a match, the packet processing component retrieves the virtual channel field of the entry and uses this information to forward the received message to the router 2528 (in accordance with the credit flow technology used by the router 2528).
[0176] The fault handling component 2906 buffers all sent packets until it receives an acknowledgement (ACK) from the receiving node (e.g., the remote acceleration component). If the ACK for the connection does not arrive within the specified timeout period, the fault handling component 2906 can retransmit the packet. The fault handling component 2906 will repeat such retransmissions a specified number of times (e.g., 128 times). If the packet remains unacknowledged after all these attempts, the fault handling component 2906 can discard it and release its buffer.
[0177] C.4. 3-port Switch
[0178] Figure 30 An implementation of the 3-port switch 2532 is shown. The 3-port switch 2532 operates to securely insert (and remove) network packets generated by the acceleration component onto the data center network without compromising host to TOR network traffic.
[0179] The 3-port switch 2532 is connected to the NIC interface 2510 (corresponding to the host interface), the TOR interface 2512, and the local interface associated with the local acceleration component 2502 itself. The 3-port switch 2532 can be conceptually thought of as including receive interfaces (3002, 3004, 3006) that are used to receive packets from the host component and the TOR switch respectively, and to receive packets at the local acceleration component. The 3-port switch 2532 also includes transmit interfaces (3008, 3010, 3012) that are used to provide packets to the TOR switch and the host component respectively, and to receive packets transmitted by the local acceleration component.
[0180] The packet classifier (3014, 3016) determines the packet class received from the host component or the TOR switch based on, for example, the status information specified by the packet. In one implementation, each packet is classified as belonging to a lossless flow (e.g., Remote Direct Memory Access (RDMA) traffic) or a lossy flow (e.g., Transmission Control Protocol / Internet Protocol (TCP / IP) traffic). Traffic belonging to a lossless flow does not tolerate packet loss, while traffic belonging to a lossy flow can tolerate some packet loss.
[0181] The packet buffers (3018, 3020) store incoming packets in different respective buffers according to the class of the traffic to which they belong. If there is no available space in the buffer, the packet will be discarded. (In one implementation, since the application logic 2506 can regulate the flow of packets by using "back pressuring", the 3-port switch 2532 does not provide packet buffering for the packets provided by the local acceleration component (via the local port).) The arbitration logic 3022 selects and transmits the selected packets among the available packets.
[0182] As described above, the traffic destined for the local acceleration component is encapsulated in UDP / IP packets on a fixed port number. The 3-port switch 2532 examines the incoming packets (e.g., as received from the TOR) to determine whether they are UDP packets on the correct port number. If so, the 3-port switch 2532 outputs the packets on the local RX port interface 3006. In one implementation, all traffic arriving on the local TX port interface 3012 is sent out from the TOR TX port interface 3008, but can also be sent to the host TX port interface 3010. Further, it should be noted that Figure 30 the acceleration component 2502 is indicated to intercept traffic from the TOR but not from the host component, but can also be configured to intercept traffic from the host component.
[0183] The PFC processing logic 3024 allows the 3-port switch 2532 to insert priority flow control frames into the traffic flow being transmitted to the TOR or host components. That is, for the lossless traffic class, if the packet buffer fills up, the PFC processing logic 3024 sends a PFC message to the link partner to request suspension of traffic on that class. If a PFC control frame for the lossless traffic class is received on the host RX port interface 3002 or the TOR RX port interface 3004, the 3-port switch 2532 will stop sending packets on the port where the control message was received.
[0184] C.5. Illustrative Host Component
[0185] Figure 31 Shows an implementation of the host component 3102 corresponding to any one of the host components in the Figure 1 illustrated host component(s). The host component 3102 may include one or more processing devices 3104, such as one or more central processing units (CPUs), and each processing unit may implement one or more hardware threads. The host component 3102 may also include any storage resources 3106 for storing any kind of information such as code, settings, data, etc. Non-limiting examples include any of the following: any type of RAM, any type of ROM, flash devices, hard disks, optical discs, etc. More generally, any storage resource may use any technology to store information. Further, any storage resource may provide volatile or non-volatile retention of information. Further, any storage resource may represent a fixed or removable component of the host component 3102. In one case, when the processing device 3104 executes the associated instructions stored in any storage resource or combination of storage resources, the host component 3102 may perform any of the operations associated with local tenant functions. The host component 3102 also includes one or more drive mechanisms 3108 for interacting with any storage resources, such as hard disk drive mechanisms, optical disc drive mechanisms, etc.
[0186] The host component 3102 also includes an input / output module 3110 for receiving various inputs (via input device 3112) and for providing various outputs (via output device 3114)). A particular output mechanism may include a presentation device 3116 and an associated graphical user interface (GUI) 3118. The host component 3102 may also include one or more network interfaces 3120 for exchanging data with other devices via one or more communication conduits 3122. One or more communication buses 3124 communicatively couple the components described above together.
[0187] The communication pipeline 3122 can be implemented in any way, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication pipeline 3722 can include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.
[0188] The following summary provides a non-exhaustive list of illustrative aspects of the techniques set forth herein.
[0189] According to a first aspect, a data processing system is described that includes two or more host components, each host component using one or more central processing units to execute machine-readable instructions, and the two or more host components jointly provide a software plane. The data processing system further includes two or more hardware acceleration components that jointly provide a hardware acceleration plane. The data processing system also includes a common network for allowing the host components to communicate with each other and for allowing the hardware acceleration components to communicate with each other. Further, the hardware acceleration components in the hardware acceleration plane have functions that enable the hardware acceleration components to communicate with each other in a transparent manner without assistance from the software acceleration plane.
[0190] According to a second aspect, the two or more hardware acceleration components in the hardware acceleration plane correspond to field programmable gate array (FPGA) devices.
[0191] According to a third aspect, the two or more host components in the software plane exchange packets through the common network via a first logical network, and the two or more hardware acceleration components in the hardware acceleration plane exchange packets through the common network via a second logical network. The first logical network and the second logical network share the physical link of the common network and are distinguished from each other based on the traffic class to which their respective packets belong.
[0192] According to a fourth aspect, the packets sent through the second logical network use a specified protocol on the identified ports, which constitutes a characteristic that differentiates the packets to be sent through the second logical network from the packets sent through the first logical network.
[0193] According to a fifth aspect, the data processing system further includes a plurality of server unit components. Each server unit component includes: a local host component; a local hardware acceleration component; and a local link for coupling the local host component to the local hardware acceleration component. The local hardware acceleration component is coupled to the common network and serves as a pipeline through which the local host component communicates with the common network.
[0194] According to a sixth aspect, at least one server unit component includes a plurality of local host components and / or a plurality of local hardware acceleration components.
[0195] According to a seventh aspect, a local hardware acceleration component is coupled to a top-of-rack switch in a data center.
[0196] According to an eighth aspect, the local hardware acceleration component is further coupled to a network interface controller, and the network interface controller is coupled to a local host component.
[0197] According to a ninth aspect, the local host component or the local hardware acceleration component is configured to issue a request for a service; and receive a reply to the request that identifies an address of the service. The local hardware acceleration component is configured to: locally execute the service when the identified address pertains to a function locally implemented by the local hardware acceleration component; and route the request to a specific remote hardware acceleration component via a common network when the identified address pertains to a function remotely implemented by a remote hardware acceleration component. Further, the local hardware acceleration component is configured to perform the routing without involving the local host component.
[0198] According to a tenth aspect, the data processing system further includes a management function for identifying an address in response to the request.
[0199] According to an eleventh aspect, a method for performing functions in a data processing environment is described. The method includes: performing the following operations in a local host component that executes machine-readable instructions using one or more central processing units, or in a local hardware acceleration component coupled to the local host component: (a) issuing a request for a service; and (b) receiving a reply to the request that identifies an address of the service. The method further includes: performing the following operations in the local hardware acceleration component: (a) locally executing the service when the identified address pertains to a function locally implemented by the local hardware acceleration component; and (b) routing the request to a remote hardware acceleration component when the identified address pertains to a function remotely implemented by a remote hardware acceleration component. Further, the local hardware acceleration component is configured to perform the routing to the remote hardware acceleration component without involving the local host component. Additionally, multiple host components communicate with each other in the data processing environment, and multiple hardware acceleration components communicate with each other in the data processing environment via a common network.
[0200] According to a twelfth aspect, each hardware acceleration component in the method described above corresponds to a field-programmable gate array (FPGA) device.
[0201] According to a thirteenth aspect, the common network in the method described above supports a first logical network and a second logical network that share a physical link of the common network. Host components in the data processing environment use the first logical network to exchange packets with each other, and hardware acceleration components in the data processing environment use the second logical network to exchange packets with each other. The first logical network and the second logical network are differentiated from each other based on the traffic class to which their respective packets belong.
[0202] According to a fourteenth aspect, packets sent through a second network in the method described above use a specified protocol on the identified ports, which constitutes a characteristic that differentiates the packets to be sent through the second logical network from those sent through the first logical network.
[0203] According to a fifteenth aspect, the local hardware acceleration component in the method described above is coupled to a common network, and the local host component interacts with the common network via the local hardware acceleration component.
[0204] According to a sixteenth aspect, the local hardware acceleration component in the method described above is coupled to a top-of-rack switch in a data center.
[0205] According to a seventeenth aspect, a server unit component in a data center is described. The server unit component includes: a local host component that executes machine-readable instructions using one or more central processing units; a local hardware acceleration component; and a local link for coupling the local host component and the local hardware acceleration component. The local hardware acceleration component is coupled to a common network and serves as a pipeline through which the local host component and the common network communicate. More generally, a data center includes multiple host components and multiple hardware acceleration components, which are provided in other corresponding server unit components, where the common network serves as a shared pipeline through which multiple host components communicate with each other and multiple hardware acceleration components communicate with each other. Further, the local hardware acceleration component is configured to interact with remote hardware acceleration components of other corresponding server unit components without involving the local host component.
[0206] According to an eighteenth aspect, the server unit component includes multiple local host components and / or multiple local hardware acceleration components.
[0207] According to a nineteenth aspect, the local hardware acceleration component is coupled to a top-of-rack switch in a data center.
[0208] According to a twentieth aspect, the local hardware acceleration component is configured to receive service requests from the local host component, and the local hardware acceleration component is configured to: (a) locally execute the service when the address associated with the service involves a function locally implemented by the local hardware acceleration component; and (b) route the request to a remote hardware acceleration component via the common network when the address involves a function remotely implemented by a remote hardware acceleration component. The local hardware acceleration component is configured to perform the routing to the remote hardware acceleration component without involving the local host component.
[0209] A twenty-first aspect corresponds to any combination (e.g., any permutation or subset) of the first to twentieth aspects described above.
[0210] The twenty-second aspect corresponds to any method counterpart, device counterpart, system counterpart, component counterpart, computer-readable storage medium counterpart, data structure counterpart, article of manufacture counterpart, graphical user interface presentation counterpart, etc. associated with the first aspect to the twenty-first aspect.
[0211] Finally, although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are disclosed as example forms for implementing the claims.
Claims
1. A data processing system, comprising: two or more host components, each having a respective central processing unit configured to execute machine-readable instructions; two or more hardware acceleration components; a first link that couples a local host component from the two or more host components to a local hardware acceleration component from the two or more hardware acceleration components; and a second link that couples the local hardware acceleration component to at least one other hardware acceleration component from the two or more hardware acceleration components, wherein the local hardware acceleration component has a function that enables the local hardware acceleration component to communicate with the at least one other hardware acceleration component without assistance from the local host component, wherein each of the two or more hardware acceleration components is configured to implement a shell function via an interface that is shared across different application scenarios, wherein the shell function of each respective hardware acceleration component includes: a transport component configured to schedule packets for transmission by the respective hardware acceleration component; a router configured to route messages between internal components of the respective hardware acceleration component, the router having a first port coupled to the application logic of the respective hardware acceleration component and a second port coupled to the transport component; and a bridge configured to control network traffic flow by directing traffic received by the respective hardware acceleration component for processing in a first mode and to allow other traffic to flow through the bridge without being processed by the respective hardware acceleration component in a second mode.
2. The data processing system according to claim 1, wherein the first link comprises a Peripheral Component Interconnect Express (PCIe) link.
3. The data processing system according to claim 2, wherein at least one of the host components is configured via respective machine-readable instructions to: obtain a configuration image having logic to implement a service; and communicate the configuration image to the local hardware acceleration component, the local hardware acceleration component being configured by the configuration image to implement the service using the interface provided by the shell function.
4. The data processing system according to claim 1, wherein the interface provided by the shell function comprises: a memory controller configured to govern interactions between the respective hardware acceleration component and a respective memory, and a host interface configured to enable services on the respective hardware acceleration component to interact with a respective host component.
5. The data processing system according to claim 4, wherein the local hardware acceleration component further comprises a multi-port switch having other ports connected to: a network interface controller, a top-of-rack switch, and an interface associated with the local hardware acceleration component.
6. The data processing system according to claim 1, wherein the received traffic is directed between the network interface controller and the top-of-rack switch for compression or encryption processing by the corresponding hardware acceleration component in the first mode, and the other traffic is directed between the network interface controller and the top-of-rack switch while bypassing the compression or encryption processing by the corresponding hardware acceleration component in the second mode.
7. The data processing system according to claim 6, wherein the local hardware acceleration component is configured to receive the result of the service performed by the at least one other hardware acceleration component, and the local hardware acceleration component is configured to obtain the result from the at least one other hardware acceleration component via the second link; and the local hardware acceleration component is configured to provide the result to the local host component via the first link.
8. The data processing system according to claim 7, wherein the local host component is configured to: provide the result to the tenant function executed on the local host component.
9. A data processing method comprising: providing two or more host components, the two or more host components having respective central processing units configured to execute machine-readable instructions; providing two or more hardware acceleration components; configuring a local host component from the two or more host components to communicate with a local hardware acceleration component from the two or more hardware acceleration components on a first link, the first link directly connecting the local hardware acceleration component to the local host component; and configuring the local hardware acceleration component to communicate with at least one other hardware acceleration component from the two or more hardware acceleration components without the assistance of the local host component; and configuring each of the hardware acceleration components to implement a shell function via an interface shared across different application scenarios, wherein the shell function of each corresponding hardware acceleration component includes: a transport component configured to schedule packets for transmission by the corresponding hardware acceleration component; a router configured to route messages between internal components of the corresponding hardware acceleration component, the router having a first port coupled to the application logic of the corresponding hardware acceleration component and a second port coupled to the transport component; and a bridge configured to control network traffic flow by directing received traffic for processing by the corresponding hardware acceleration component in a first mode and allowing other traffic to flow through the bridge without being processed by the corresponding hardware acceleration component in a second mode.
10. The method according to claim 9, the method further comprising: configuring parallel logic elements of the local hardware acceleration component to perform the requested service.
11. The method according to claim 10, further comprising: receiving a request from a tenant function on the local host component to perform the requested service; and In response to the request from the tenant function, configure the parallel logic elements of the local hardware acceleration component.
12. The method according to claim 9, further comprising: Configuring the local hardware acceleration component to communicate with the at least one other hardware acceleration component using a second link.
13. The method according to claim 12, wherein the first link comprises a Peripheral Component Interconnect Express (PCIe) link, and the second link comprises a packetized network.
14. A data processing system, comprising: Two or more host components, each having a corresponding central processing unit configured to execute machine-readable instructions; Two or more hardware acceleration components; A link that directly connects a local host component from the two or more host components to a local hardware acceleration component from the two or more hardware acceleration components; and A network that connects the local hardware acceleration component to at least one other hardware acceleration component from the two or more hardware acceleration components, wherein the local hardware acceleration component has a function that enables the local hardware acceleration component to communicate with the at least one other hardware acceleration component on the network without assistance from the local host component, each hardware acceleration component is configured to implement a shell function via an interface shared across different application scenarios, the shell function of each corresponding hardware acceleration component includes: A transport component configured to schedule packets for transmission on the network by the corresponding hardware acceleration component; A router configured to route messages between internal components of the corresponding hardware acceleration component, the router having a first port coupled to the application logic of the corresponding hardware acceleration component and a second port coupled to the transport component; and A bridge configured to control network traffic flow by directing traffic received by the corresponding hardware acceleration component for processing in a first mode, and to allow other traffic to flow through the bridge without being processed by the corresponding hardware acceleration component in a second mode.
15. The data processing system according to claim 14, wherein the local host component is configured to communicate on the network with another host component from the two or more host components.
16. The data processing system according to claim 15, wherein the local host component and the local hardware acceleration component use the same network interface controller to communicate on the network.
17. The data processing system according to claim 15, wherein the local host component and the other host component are configured to exchange other packets on the network using a special category of network traffic, and the local hardware acceleration component and the at least one other hardware acceleration component are configured to exchange the packets on the network using another category of network traffic.
18. The data processing system according to claim 17, wherein the packets use a protocol on an identified port to distinguish the packets from the other packets.
19. The data processing system according to claim 18, wherein the packet comprises a User Datagram Protocol (UDP) packet.
20. The data processing system according to claim 14, wherein the router has: a third port connected to a memory controller, the memory controller being configured to govern the interaction of the corresponding hardware acceleration component with the memory of the corresponding hardware acceleration component; and a fourth port connected to a host interface, the host interface being configured to provide a function to the corresponding hardware acceleration component to interact with a corresponding host component.
Citation Information
Patent Citations
Device capable of conducting business hardware acceleration and method thereof
CN102769574A
Acceleration for Virtual Bridged Hosts
US20130152075A1