Multi-Criteria Power Management Scheme for Pooling-Based Accelerator Architectures

Through the computing nodes and accelerators connected to the switch, the power supply of the accelerator is dynamically adjusted using telemetry meters and power management controllers, solving the intelligent load balancing problem of accelerator power management in the data center, reducing the total cost of ownership and improving resource utilization efficiency.

CN111052039BActive Publication Date: 2025-07-04INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201880056533.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-09-28
Filing Date
2018-09-28
Publication Date
2025-07-04
Estimated Expiration
2038-09-28

AI Technical Summary

Technical Problem

In existing data centers, the power management solution of pooled accelerators cannot be intelligently load balancing and optimization based on the accelerator's performance requirements, resulting in waste of resources and an increase in total cost of ownership.

Method used

Connect multiple computing nodes and accelerators through switches, and use telemetry meters and power management controllers to monitor and adjust the power supply of each accelerator in real time, adjusting power distribution dynamically according to its current performance level and performance targets.

Benefits of technology

It realizes intelligent power load balancing based on performance requirements, reduces the total cost of ownership of the computing environment, and improves resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111052039B_ABST
    Figure CN111052039B_ABST
Patent Text Reader

Abstract

A computing device, method, and system for controlling power. The computing device is configured to be used as part of a network architecture including a plurality of nodes and a plurality of pooled accelerators communicatively coupled to the nodes. The computing device includes: a memory storing instructions; and a processing circuit configured to execute the instructions. The processing circuit is used to receive corresponding requests from respective nodes among the plurality of nodes, the requests being addressed to a plurality of corresponding accelerators, each of the corresponding requests including information about a kernel to be executed by the corresponding accelerator, about the corresponding accelerator, and about a performance target for the execution of the kernel. The processing circuit is further used to control the power supply to the corresponding accelerator based on the information in each of the corresponding requests.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments described herein generally relate to power management, which involves data centers that use aggregated subsystems. Background Art

[0002] As data centers have evolved, such architectural concepts have moved from rack implementations, which use shared power, shared cooling, and rack-level management, to more decentralized implementations that involve subsystem aggregation and the use of pooled computing resources, pooled storage and storage machines, and / or shared boot. For newly emerging data center and network fabric architectures, changes to power management schemes are needed. Brief Description of the Drawings

[0003] For simplicity and clarity of illustration, the elements shown in the drawings are not necessarily drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity of presentation. Additionally, reference numerals may be repeated among the drawings to indicate corresponding or analogous elements. The drawings are listed below.

[0004] Figure 1 is a schematic illustration of a computing environment according to some illustrative embodiments, the computing environment including a plurality of nodes communicatively coupled via a switch to a plurality of pooled accelerators;

[0005] Figure 2 is a telemetry / power meter according to some illustrative embodiments;

[0006] Figure 3 is a flowchart of a first method according to some illustrative embodiments; and

[0007] Figure 4 is a flowchart of a second method according to some illustrative embodiments. Detailed Description

[0008] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of some embodiments. However, one of ordinary skill in the art will understand that some embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, units, and / or circuits have not been described in detail so as not to obscure the discussion.

[0009] For simplicity and clarity of illustration, the figures show a general construction and the description and details of well-known features and techniques may be omitted so as not to unnecessarily obscure the discussion of the described embodiments of the present invention. Additionally, the elements in the figures are not necessarily drawn to scale. For example, the dimensions of some of the elements in the figures may be enlarged relative to other elements to help improve the understanding of the disclosed embodiments. The same reference numerals in different figures represent the same elements, while like reference numerals may but do not necessarily represent like elements.

[0010] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims are used to distinguish similar elements and not necessarily to describe a particular order or temporal sequence. It should be understood that such terms are interchangeable under appropriate circumstances, e.g., such that the embodiments of the present invention described herein can be operated in an order different from those illustrated or otherwise described herein. Similarly, if a method is described herein as including a series of acts, the order of the acts presented herein is not necessarily the only order in which such acts can be performed, and certain of the described acts may be omitted and / or certain other acts not described herein may be added to the method. Further, the terms "comprising", "including", "having" and any variations thereof are intended to cover non-exclusively, such that a process, method, article or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to such process, method, article or apparatus.

[0011] The term "coupled" or "communicatively coupled" as used herein is defined as connected directly or indirectly in either an electrical or non-electrical manner. As used herein, a "processing circuit" may refer to a single instance of a processor block in different physical locations, or to multiple processor blocks in various physical locations within a platform, and can be implemented in hardware, firmware, software, or a combination thereof. As used herein, a "memory" may refer to a single instance of a memory block in different physical locations, or to multiple memory blocks in various locations within a platform, and can likewise be implemented in hardware, firmware, software, or a combination thereof.

[0012] Although the following embodiments are described with reference to energy conservation and energy efficiency in specific computing environments (such as those including a computing platform or a processor), the similar techniques and teachings of the embodiments described herein can be applied to other types of circuits or semiconductor devices that can also benefit from better energy efficiency and better energy conservation. For example, the disclosed embodiments are not limited to any specific type of computing environment, such as a rack-scale architecture data center. That is, the disclosed embodiments can be used in many different system types, including server computers (e.g., tower, rack-mounted, blade, microserver, etc.), communication systems, storage systems, desktop computers configured in any way, laptops, notebooks, and tablet computers (including 2:1 tablets, phablets, etc.), and the disclosed embodiments can also be used in other devices, such as handheld devices, wearable devices, IoT devices, to name a few.

[0013] The embodiments are not limited to physical computing devices but can also relate to software optimizations for energy conservation and energy efficiency. As will become apparent in the following description, the embodiments of the methods, apparatuses, and systems described herein (whether referring to hardware, firmware, software, or a combination thereof) are indispensable for future 'green technologies' such as power savings and power efficiency in products that cover a large portion of the U.S. economy.

[0014] According to the prior art, it is known that pooled accelerators in a data center use a fixed amount of power depending on their respective power requirements. In the context of an embodiment, by "pooled accelerators" is meant a group of two or more accelerators connected by a switch and "pooled" together, within the same rack / enclosure or on different racks and / or at different locations in a data center, where pooling is with respect to the nodes used to issue kernels to the accelerators for execution. However, with a fixed power supply, the data center architecture cannot achieve optimal load balancing and intelligent power distribution based on the performance requirements of the workload on a given accelerator. As an example, according to the prior art, if application A issues instructions to accelerator 1 to run an accelerated kernel a, then considering that accelerator 1 is under a fixed power supply, kernel a may take 100 seconds to run on accelerator 1. However, the service level agreement (SLA) associated with application A may have a time requirement of 200 seconds. Assuming there is a relationship between power and performance, kernel a can be executed at half the power supplied to accelerator 1 and still meet the SLA timing requirement. The remaining power can then be provided to other accelerators that are running accelerated kernels and have a more stringent SLA requirement than accelerator 1. As will be appreciated by a person skilled in the art, the SLA can be based, for example, on time requirements and / or on power performance requirements.

[0015] Embodiments present a mechanism for more efficiently utilizing power in pooled accelerators and heterogeneous workload architectures with different performance requirements. An objective of embodiments is to reduce the total cost of ownership (TCO) of a computing environment while providing a flexible architecture that allows for an intelligent power load balancing scheme based on the performance requirements of each component.

[0016] Some embodiments include a switch configured to connect multiple computer nodes (nodes) to multiple accelerators. According to some illustrative embodiments, a node can access any and all accelerators to which it is connected via the switch. As an example, the switch can include a fabric interconnect switch such as for Transmission Control Protocol (TCP) or User Datagram Protocol (UDP) communications and / or Remote Direct Memory Access (RDMA). As an example, the switch can include a coherence switch that includes one or more interconnections for a communication protocol that provides cache coherence transactions. The switch can include, for example, a memory switch. The switch can include, for example, a (Peripheral Component Interconnect Express) PCIe switch or any suitable high-speed serial computer expansion bus. The switch can include a memory device for storing instructions and processing circuitry coupled to the memory. The switch can be configured to process requests from nodes among the multiple nodes and to use the instructions in the memory and based on the request to determine the cores to be issued to an accelerator among the multiple accelerators and the performance goals of the cores. The switch can further be configured to control the power supply to be delivered to the accelerator based on the performance goals of the cores. The switch can be configured to perform the above functions for each accelerator among the multiple accelerators.

[0017] The switch can include a telemetry table that includes data mapping each accelerator among the multiple accelerators to their current performance level, their current power supply level, and their threshold power supply level. The switch can control the amount of power supplied to each accelerator based on the current performance level of the accelerator, based on the current power supply level of the accelerator, and based on the threshold power supply level of the accelerator. For example, the switch can be configured to reduce the amount of power supplied to each accelerator in response to determining that the current performance level of the accelerator is higher than the performance goal. Similarly, the switch can be configured to reduce the amount of power supplied to each accelerator in response to determining that the current performance level of the accelerator is lower than the performance goal. The switch can be configured to adjust the power to each accelerator based on the threshold power level of the accelerator. The switch can further be configured to monitor the current performance level of each accelerator and to adjust the power supplied to the accelerator based on the monitored and updated current performance level of the accelerator.

[0018] The switch can be further configured to redirect power from one of the accelerators to another one of the accelerators based on the corresponding current power levels and corresponding performance goals of the accelerators.

[0019] Exemplary embodiments will now be described in further detail below with respect to Figures 1-4 Exemplary embodiments will now be described in further detail below with respect to

[0020] Figure 1 is a schematic illustration of a computing environment 100, which is part of a network architecture including computing devices such as nodes, switches, and pooled accelerators. The nodes in the illustrated environment include Node 1, Node 2, and Node 3. According to some exemplary embodiments, the nodes are shown communicatively coupled via a switch 110 to a plurality of pooled accelerators - Accelerator 1, Accelerator 2, and Accelerator 3. Although only three nodes - Node 1, Node 2, and Node 3, only three accelerators - Accelerator 1, Accelerator 2, and Accelerator 3, and only one switch 110 are shown in Figure 1 , the computing environment 100 can include any suitable number of computing nodes, accelerators, and switches coupled to each other via a network architecture such as a low-latency network architecture connection 109. Each node as shown herein can include a central processing unit (CPU) 102, a cache memory 104, a main memory 106, and a node network architecture 108. As will be appreciated by those skilled in the art, each node can include any suitable number of processors and cores within each CPU, a field-programmable gate array (FPGA), a controller, a memory, and / or other components. The computing environment can represent any suitable computing environment, such as, a high-performance computing environment, a data center, a communication service provider infrastructure (e.g., one or more parts of an evolved packet core), a memory-in-compute environment, another computing environment, or a combination thereof.

[0021] The switch 110 can include, for example, an ingress interface 112 and an egress interface 113. The ingress interface 112 is used to receive, for example, node-by-node, payloads from the associated nodes and queue the payloads. The egress node 113 is used to queue, for example, accelerator-by-accelerator, the processed payloads for transmission from the switch. The ingress interface 112 and the egress interface 113 can include, for example, PCIe interfaces as shown, and can further include interfaces for the transmission of power or voltage signals to each of the accelerators - Accelerator 1, Accelerator 2, and Accelerator 3.

[0022] Switch 110 may further include a switch processing circuit 114 and a switch memory 118, and the switch memory 118 stores instructions for execution by the switch processing circuit 114. Additionally, switch 110 may include a power meter or a telemetry meter 116. As will be appreciated by those skilled in the art, the telemetry meter 116 may include information that maps each of the plurality of pooled accelerators - accelerator 1, accelerator 2, and accelerator 3 to their current performance levels, their current power levels (current power supply levels), and their threshold power levels, where the threshold power levels are predetermined based on, for example, accelerator parameters and capacities. The telemetry meter 116 may further include other telemetry regarding each accelerator. Although not shown, as will be appreciated by those skilled in the art, switch 110 may include other components thereon. The components of switch 110 may be communicatively coupled to each other on switch 110 and / or directly communicatively coupled to each other, such as via bus 120 as shown, such as the components of the switch, interface 112, interface 113, switch processing circuit 114, switch memory 118, and telemetry meter 116.

[0023] Now referring to each accelerator, accelerator 1, accelerator 2, and accelerator 3 may each include an accelerator network interface (NI) 124, a processor 126, an accelerator memory 128 (including volatile and non-volatile memory 128), and a direct memory access component (DMA) 132, all of which may be interconnected, such as via bus 147 and / or directly. As will be recognized by those skilled in the art, the DMA may be used for "memory-to-memory" copying or for moving data within memory. Each accelerator may further include an accelerator unit 149, which, in one embodiment, may be an accelerator unit SoC, such as a field programmable gate array-based accelerator unit. The accelerator unit 149 may include a network interface (NI) 134, an input buffer 136, and an output buffer 138, a memory controller 142, a programmable processor 140, and a series of switches 144, all of which are connected to each other via bus 146. The accelerator unit may be connected to the other components of the accelerator via NI 134 and further via bus 147 and / or directly. Additionally, according to one embodiment, each accelerator includes a power management controller (PMC) therein (PMC 1211, 1212, and 1213 for accelerator 1, accelerator 2, and accelerator 3, respectively), and the PMC is configured to control the amount of power, for example, through a voltage input terminal (VI) 30 on each accelerator to each accelerator. The VI 30 may include one or more voltage input pins. According to one embodiment, the PMC may be incorporated within the processor 126 of each accelerator. According to another embodiment, the processor 126 and the programmable processor 140 of each accelerator may be collectively referred to as the processing circuit of the accelerator.

[0024] Embodiments include other components or different components on each accelerator, accelerator units not based on FPGAs, and pooled accelerators that are different from each other within their scope. Note that, by way of example only, the structures shown for each of Accelerator 1, Accelerator 2, and Accelerator 3 in the illustrated embodiments are shown to be identical. Additionally, an accelerator according to one embodiment may be configured to execute one kernel at a time, or it may be configured to execute more than one kernel at a time. Moreover, embodiments include accelerators that do not include an on-board PMC, where power management functions for the accelerator may be performed externally, such as, by way of example, within a switch such as Switch 110, or within a PMC (not shown) external to and communicatively coupled to Switch 110, or within a PMC in a suitable accelerator. Due to the many possible locations of the PMC, the PMC is shown as a dashed line in Figure 1 the figure.

[0025] Now referring to Figure 2 , an example of a telemetry table 216 is shown, such as telemetry table 116 that is part of a switch 110 such as Figure 1 . The telemetry table 216 may include information for each of Accelerator 1, Accelerator 2, and Accelerator 3. Specifically, for each accelerator, the telemetry table 216 may include its current performance level represented by the amount of floating point operations per second (FLOPS) A, B, C for the respective accelerators in Accelerator 1, Accelerator 2, and Accelerator 3. Additionally, for each accelerator, the table 216 may include its current performance level represented by the amount of X, Y, or Z watts for the respective accelerators in Accelerator 1, Accelerator 2, and Accelerator 3. Moreover, for each accelerator, the table 216 may include its threshold power level represented by the amount of X', Y', or Z' watts for the respective accelerators in Accelerator 1, Accelerator 2, and Accelerator 3.

[0026] Now a power management scheme according to some embodiments will be described with respect to Figure 1 and Figure 2 .

[0027] Now referring to Figure 1 and Figure 2, the switch processing circuit 114 can be coupled to the switch memory to retrieve instructions from the switch memory and execute these instructions to perform operations, which include: determining the cores to be issued to the accelerator and the performance goals of the cores, and controlling the amount of power to be delivered to the accelerator based on the performance goals of the cores. According to one embodiment, the switch 110 can be configured to adjust the amount of power supplied to each of the accelerators - accelerator 1, accelerator 2, and accelerator 3 based on the current performance level of each accelerator, based on the current power supply to each accelerator, and based on the threshold power supply to each accelerator. For example, according to one embodiment, the switching processing circuit 114 can retrieve such information by accessing the information about the current performance levels of each of the accelerators - accelerator 1, accelerator 2, and accelerator 3 in the telemetry table 116 / 216 on the switch. For example, the switch 110 can be configured to reduce the amount of power supplied to each of the accelerators - accelerator 1, accelerator 2, and accelerator 3 in response to determining that the current performance level of the accelerator is higher than the performance goal. Similarly, the switch 110 can be configured to increase the amount of power supplied to each of the accelerators - accelerator 1, accelerator 2, and accelerator 3 in response to determining that the current performance level of the accelerator is lower than the performance goal.

[0028] Still referring to Figure 1 and Figure 2 , the switch 110 can be configured to process requests from nodes in multiple nodes (such as, for example, from node 1). The request can be transmitted to the switch 110 via the ingress interface 112 and the network fabric connection 109, and the request can be a request from node 1 to the switch 110 to issue a core such as core a to one of the accelerators (such as, accelerator 1) for execution by the accelerator. For example, the switch 110 can issue the core to accelerator 1 via the egress interface 113 and the network fabric connection 109. The switch processing circuit 114 can be configured to use the instructions in the switch memory 118 and based on the request from node 1 to determine: (1) that core a is to be issued to accelerator 1, and further (2) the performance level associated with the SLA of core a. The switch 110 can be further configured to control the amount of power to be delivered to accelerator 1 based on the performance goal of core a, for example, by the switch processing circuit 114 and using the instructions in the switch memory 118.

[0029] Still referring to Figure 1 and Figure 2, for example, by the switch processing circuitry 114 and using the instructions within the switch memory 118 and the information within the telemetry tables 116 / 216, the switch 110 can be configured to reduce the amount of power supplied to accelerator 1 in response to determining that the current performance level of accelerator 1 is higher than the performance target of core a. Similarly, for example, by the switch processing circuitry 114 and using the instructions within the switch memory 118 and the information within the telemetry tables 116 / 216, the switch 110 can be configured to increase the amount of power supplied to accelerator 1 in response to determining that the current performance level of accelerator 1 is lower than the performance target of core a. For example, by using the switch processing circuitry 1114 and using the instructions within the switch memory 118 and the information within the telemetry tables 116 / 216, the switch 110 can further be configured to adjust the power to the accelerator based on a threshold power level of accelerator 1. For example, by the switch processing circuitry 114 and using the instructions within the switch memory 18 and the information within the telemetry tables 116 / 216, the switch 110 can further be configured to monitor the current performance level of accelerator 1 to update the current performance level of accelerator 1 in the telemetry tables 116 / 216, and to adjust the power supplied to accelerator 1 based on the updated current performance level of accelerator 1 determined based on the monitoring.

[0030] The switch 110 can control the power to accelerator 1 in the manner described above while maintaining the amount of power supplied to accelerator 1 within the threshold power levels set forth in the telemetry tables 116 / 216. The switch 110 can include, for example, mechanisms therein for initially setting the telemetry tables 116 / 216 including the threshold power levels for each of accelerators 1, 2, and 3 via, for example, the switch processing circuitry 114.

[0031] Although the above two paragraphs provide examples of embodiments in which one node among the nodes sends a request to the switch to issue one core to an accelerator (i.e., node 1 sends a request to the switch to cause accelerator 1 to execute core a) and in which the switch then controls the power to one accelerator (i.e., accelerator 1), according to an embodiment: (1) any number of nodes can transmit any number of requests to the switch; (2) each request can include a request for one or more cores to be executed by one or more accelerators; (3) each accelerator can be configured to execute one or more cores based on instructions from the switch; (4) the switch can be configured to control the power to any number of accelerators based on the respective performance levels of the cores to be executed by the accelerators; and (5) the switch can be configured to monitor the telemetry of any number of accelerators and to update its telemetry tables based on such monitoring.

[0032] According to one embodiment, the switch 110 can be configured to supply the amount of power to be supplied to accelerators 1, 2, and 3, for example, by processing circuitry 114 of the switch and using instructions within switch memory 118, such as by controlling the power through one or more of PMCs 1211, 1212, and 1213 on each respective accelerator. Each of PMCs 1211, 1212, and 1213 can be configured to receive corresponding instructions from the switch 110 to adjust the power to its particular accelerator. The instructions to each PMC can be routed to each PMC, for example, through bus 147 of each of accelerators 1, 2, or 3 as appropriate. Each of PMCs 1211, 1212, and 1213 can then be configured to execute the instructions to the PMC to control the amount of power arriving at its corresponding accelerator through a power input connection, such as one or more voltage input pins VI 130 of the respective accelerator among accelerators 1, 2, and 3. Additionally or alternatively, one or more PMCs can exist external to accelerators 1, 2, and 3 that can be configured to regulate the amount of power arriving at each of accelerators 1, 2, and 3 through the respective VI 130. According to one embodiment, one or more PMCs can reside within the appropriate switch processing circuitry 114. According to another embodiment, one or more PMCs can reside within the respective accelerator among the accelerators. According to yet another embodiment, the switch 110 can control the power to each accelerator by sending instructions to each accelerator to execute one or more kernels at a certain clock frequency level, thereby indirectly controlling the power drawn from each accelerator. In the latter case, one or both of the processor 126 and the programmable processor 140 of each accelerator can control the clock frequency at which the accelerator executes one or more kernels.

[0033] The processing circuits 114 and 126 or any other processing circuit or controller in a computing environment according to an embodiment may include any suitable processing circuit (note that in this specification and the associated drawings, processing circuit and processor may be used interchangeably), such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, a handheld processor, an application processor, a coprocessor, a system on a chip (SOC), or other devices for executing code (i.e., software instructions). The processing circuits 114 and 126 or any other processing circuit in a computing environment according to an embodiment may include multiple processing cores or a single core, and the multiple processing cores or the single processing core may include asymmetric processing elements or symmetric processing elements. However, the processor or processing circuit mentioned herein may include any number of processing elements, which may be symmetric or asymmetric. A processing element may refer to hardware or logic for supporting software threads. Examples of hardware processing elements include: thread units, thread slots, threads, process units, contexts, context units, logical processors, hardware threads, cores, and / or any other element capable of maintaining the state of a processor (such as an execution state or an architectural state). In other words, in one embodiment, a processing element refers to any hardware capable of independently associating with code (such as software threads, operating systems, applications, or other code). A physical processor (or processor socket) typically refers to an integrated circuit that potentially includes any number of other processing elements such as cores or hardware threads. A processing element may also include one or more arithmetic logic units (ALUs), floating point units (FPUs), caches, instruction pipelines, interrupt handling hardware, registers, or other hardware for facilitating the operation of the processing element.

[0034] According to an embodiment, an FPGA-based accelerator unit, such as accelerator unit 149, may include any number of semiconductor devices, which may include configurable / reprogrammable logic circuitry in the form of a programmable processor 140. The FPGA-based accelerator unit may be configured via a data structure (e.g., a bitstream) having any suitable format that defines how the logic will be configured. The FPGA-based accelerator unit may be reprogrammed any number of times after the FPGA-based accelerator unit has been manufactured. The configurable logic of the FPGA-based accelerator unit may be programmed to implement one or more kernels. A kernel may include the configured logic of the FPGA-based accelerator unit, which may receive a set of one or more inputs, process the set of inputs using the configured logic, and provide a set of one or more outputs. A kernel may perform any suitable type of processing. In embodiments, a kernel may include a video processor, an image processor, a waveform generator, a pattern recognition module, a packet processor, an encryptor, a decryptor, an encoder, a decoder, a compression device, a processor operable to perform any number of operations each specified by a different instruction sequence, or any other suitable processing function. Some FPGA-based accelerator units may be limited to executing a single kernel at a time, while other FPGA-based accelerator units may be capable of executing multiple kernels simultaneously.

[0035] Any suitable entity of a compute node may be configured to instruct an accelerator to implement one or more kernels (i.e., one or more kernels may be registered at the accelerator) and / or execute one or more kernels (i.e., provide one or more input parameters to the accelerator to perform the functions of the kernels based on these input parameters).

[0036] Memories such as switch memory 118, telemetry meters 116 / 216, memory 128, and memories within each of the nodes - Node 1, Node 2, and Node 3 - can store any suitable data, such as data used by processors communicatively coupled thereto to provide the functionality of computing environment 100. For example, data associated with a program executed or a file accessed by switch processing circuitry 114 can be stored in switch memory 118. Thus, a memory device according to some embodiments can include a system memory that stores data used by processing circuitry and / or a sequence of instructions executed by processing circuitry. In embodiments, a memory device according to some embodiments can store persistent data (e.g., a user's files or sequence of instructions), which remains stored even after power to the memory device according to the embodiments is removed. A memory device according to an embodiment can be dedicated to a particular processing circuitry or shared with other devices of computing environment 100. In embodiments, a memory device according to an embodiment can include a memory that includes any number of memory modules, a memory device controller, and other support logic. The memory modules can include a plurality of memory cells each operable to store one or more bits. The cells of the memory modules are arranged in any suitable form, such as in columns and rows or a three-dimensional structure. The cells can be logically grouped into banks, blocks, pages (where a page is a subset of a block), sub-blocks, frames, word lines, bit lines, bytes, or other suitable groups. The memory modules can include non-volatile memory and / or volatile memory.

[0037] Memory controller 142 can be an integrated memory controller (i.e., it is integrated on the same die or integrated circuit as programmable processor 140) that includes logic for controlling the data flow to or from volatile and non-volatile memory 128. Memory controller 142 can include logic operable to read from, write to, or request other operations from a memory device according to an embodiment. During operation, memory controller 142 can issue commands that include one or more addresses of memory 128 according to an embodiment to read data from or write data to the memory (or perform other operations). In some embodiments, memory controller 142 can be implemented on a die or memory different from the die or integrated memory on which processor 106 is implemented.

[0038] For using inter-connected inter-component communication and intra-component communication, such as, for communication between nodes, switches, and accelerators or for communication within a node, within a switch, or within an accelerator, the protocol for communicating via the interconnection may have any suitable features of the Intel Ultra Path Interconnect (UPI), the Intel Quick Path Interconnect (QPI), or other known communication protocols. The network fabric connection 109 may include any suitable network fabric, such as, an Ethernet fabric, an Intel Omni-Path fabric, an Intel True Scale fabric, an InfiniBand-based fabric (e.g., an InfiniBand enhanced data rate fabric), a RapidIO fabric, or other suitable network fabrics. In other embodiments, the network fabric connection 109 may include any other suitable board-to-board, socket-to-socket interconnection.

[0039] Although not depicted, the computer environment 100 may use one or more batteries, renewable energy converters (e.g., solar or motion-based energy), and / or a power supply socket connector and associated systems for receiving power, a display for outputting data provided by one or more processing circuits, or a network interface that allows the processing circuits to communicate over a network. In embodiments, the battery, the power supply socket connector, the display, and / or the network interface may be communicatively coupled to the processing circuits.

[0040] Figure 3 is a flowchart of a first method 300 according to some illustrative embodiments. At operation 302, the method includes: processing respective requests from respective nodes among a plurality of nodes within a network fabric, the respective requests being addressed to a plurality of corresponding accelerators among a plurality of pooled accelerators within the network fabric, each respective request among the respective requests including information about a kernel to be executed by a corresponding accelerator among the plurality of corresponding accelerators, about the corresponding accelerator, and about a performance target for the kernel execution. At operation 304, the method includes: controlling the power supply to the corresponding accelerators based on the information in each respective request among the respective requests.

[0041] Figure 4 is a flowchart of a second method 400 according to some illustrative embodiments. At operation 402, the method includes: executing a kernel posted from a node among a plurality of nodes to a computing device, the kernel having a performance target associated therewith, the computing device being configured as part of a group of pooled accelerators within a network fabric. At operation 404, the method includes: processing an instruction for controlling the power supply to the processing circuits of the computing device based on the performance target of the kernel.

[0042] Examples described herein may include logic or several components, modules, or mechanisms, or may operate on logic or several components, modules, or mechanisms. A module is a tangible entity (e.g., hardware) capable of performing specified operations when operating. A module includes hardware. In an example, the hardware may be specifically configured to perform particular operations (e.g., hardwired). In another example, the hardware may include configurable execution units (e.g., transistors, circuits, etc.) and a computer-readable medium that contains instructions, where the instructions configure the execution units to perform particular operations when in operation. The configuration may occur under the guidance of the execution unit or a loading mechanism. Accordingly, when the device is operating, the execution units are communicatively coupled to the computer-readable medium. In this example, the execution units may be components of more than one module. For example, when operating, the execution units may be configured by a first set of instructions to implement a first module at a first point in time and reconfigured by a second set of instructions to implement a second module at a second point in time.

[0043] For example, returning to Figure 1 , a storage unit or memory (such as the switch memory 18 of computing environment 100, or other memory or combination of memories) may include a machine-readable medium on which is stored one or more sets of data structures or instructions (e.g., software) embodied or utilized by one or more of the techniques or functions described herein. The instructions may also reside, wholly or at least partially, in the main memory, in the static memory, or within the processing circuitry when executed by the machine. In an example, one or any combination of the processing circuitry, the main memory, the static memory, or other storage devices may constitute a machine-readable medium.

[0044] Some illustrative embodiments may be implemented wholly or in part in software and / or firmware. The software and / or firmware may take the form of instructions contained in or on a non-transitory computer-readable storage medium. Those instructions may then be read and executed by one or more processors to effect performance of the operations described herein. Those instructions may then be read and executed by one or more processors to cause a switch or accelerator to perform the methods and / or operations described herein. The instructions may be in any suitable form, such as but not limited to, source code, compiled code, interpreted code, executable code, static code, dynamic code, etc. Such computer-readable media may include any tangible non-transitory medium for storing information in a form readable by one or more computers, such as but not limited to: read only memory (ROM); random access memory (RAM); magnetic disk storage media; optical storage media; flash memory, etc.

[0045] The functions, operations, components, and / or features described herein with reference to one or more embodiments may be combined with or utilized in combination with one or more other functions, operations, components, and / or features described herein with reference to one or more embodiments, or vice versa.

[0046] Example:

[0047] The following examples relate to further embodiments.

[0048] Example 1 includes a computing device configured to be used as part of a network architecture that includes a plurality of nodes and a plurality of pooled accelerators communicatively coupled to the nodes. The computing device includes: a memory storing instructions; and a processing circuit coupled to the memory, the processing circuit for executing the instructions to: receive respective requests from respective ones of the plurality of nodes, the respective requests being addressed to respective ones of the plurality of pooled accelerators, each of the respective requests including information about a kernel to be executed by a corresponding one of the plurality of pooled accelerators, about the corresponding accelerator, and about a performance target for the execution of the kernel; and control the power supply to the corresponding accelerator based on the information in each of the respective requests.

[0049] Example 2 includes the subject matter of Example 1, and optionally, wherein the processing circuit is further for executing the instructions to issue the kernel to the corresponding accelerator for execution by the corresponding accelerator.

[0050] Example 3 includes the subject matter of Example 1, and optionally, wherein the processing circuit is further for executing the instructions to: implement monitoring of the current performance level of the corresponding accelerator during execution of the kernel; and control the power supply to the corresponding accelerator during execution of the kernel based on an updated version of the current performance level obtained from the monitoring.

[0051] Example 4 includes the subject matter of Example 1, and optionally, wherein: the device is for storing a telemetry table that includes data mapping each of the plurality of corresponding accelerators to the current performance level, the current power supply level, and a threshold power supply level of each of the plurality of corresponding accelerators; and the processing circuit is for executing the instructions to control the power supply by controlling the power supply to each of the plurality of corresponding accelerators based on determining the current performance level, the current power supply level, and the threshold supply level of each of the plurality of corresponding accelerators from the telemetry table.

[0052] Example 5 includes the subject matter as described in Example 4, and optionally, wherein the processing circuitry is configured to execute instructions to control power supply by reducing the power supply to each of the plurality of corresponding accelerators in response to determining that the current performance level of each of the plurality of corresponding accelerators is higher than the performance target of the core being executed by each of the plurality of corresponding accelerators.

[0053] Example 6 includes the subject matter as described in Example 4, and optionally, wherein the processing circuitry is configured to execute instructions to control power supply by increasing the power supply to each of the plurality of corresponding accelerators in response to determining that the current performance level of each of the plurality of corresponding accelerators is lower than the performance target of the core being executed by each of the plurality of corresponding accelerators.

[0054] Example 7 includes the subject matter as described in Example 4, and optionally, wherein the processing circuitry is configured to execute instructions to redirect power supply from a first corresponding accelerator among the plurality of corresponding accelerators to a second corresponding accelerator among the plurality of corresponding accelerators based on the current performance level, the current power supply level, and a threshold power supply level of each of a first corresponding accelerator and a second corresponding accelerator among the plurality of corresponding accelerators and further based on the respective performance targets of the cores being executed by each of the first corresponding accelerator and the second corresponding accelerator among the plurality of corresponding accelerators.

[0055] Example 8 includes the subject matter as described in Example 4, and optionally, wherein the processing circuitry is configured to execute instructions to implement an initial setting of data within a telemetry meter.

[0056] Example 9 includes the subject matter as described in Example 1, and optionally, further includes at least one of a coherence switch, a memory switch, or a peripheral component interconnect express (PCIe) switch including an ingress interface for receiving a corresponding request from a node and an egress interface for controlling power supply.

[0057] Example 10 includes a product that includes one or more tangible computer-readable non-transitory storage media including computer-executable instructions that, when executed by at least one computer processor, enable the at least one computer processor to implement operations at a computing device, the operations including: processing respective requests from respective nodes of a plurality of nodes within a network fabric, the respective requests being addressed to respective ones of a plurality of pooled accelerators within the network fabric, each of the respective requests including information regarding a kernel to be executed by a corresponding one of the plurality of pooled accelerators, regarding the corresponding accelerator, and regarding performance goals for execution of the kernel; and controlling power supply to the corresponding accelerator based on the information in each of the respective requests.

[0058] Example 11 includes the subject matter of Example 10 and, optionally, wherein the operations further include: issuing the kernel to the corresponding accelerator for execution by the corresponding accelerator.

[0059] Example 12 includes the subject matter of Example 10 and, optionally, wherein the operations further include: implementing monitoring of a current performance level of the corresponding accelerator during execution of the kernel; and controlling power supply to the corresponding accelerator during execution of the kernel based on an updated version of the current performance level obtained from the monitoring.

[0060] Example 13 includes the subject matter of Example 10 and, optionally, wherein the operations further include: controlling power supply by controlling power supply to each of the plurality of corresponding accelerators based on a current performance level, a current power supply level, and a threshold power supply level of each of the plurality of corresponding accelerators.

[0061] Example 14 includes the subject matter of Example 13 and, optionally, wherein the operations further include: controlling power supply by reducing power supply to each of the plurality of corresponding accelerators in response to determining that a current performance level of each of the plurality of corresponding accelerators is higher than a performance goal of a kernel being executed by each of the corresponding accelerators.

[0062] Example 15 includes the subject matter of Example 13 and, optionally, wherein the operations further include: controlling power supply by increasing power supply to each of the plurality of corresponding accelerators in response to determining that a current performance level of each of the plurality of corresponding accelerators is lower than a performance goal of a kernel being executed by each of the corresponding accelerators.

[0063] Example 16 includes the subject matter as described in Example 13, and optionally, wherein the operation further includes: redirecting power supply from a first corresponding accelerator among a plurality of corresponding accelerators to a second corresponding accelerator among the plurality of corresponding accelerators based on the current performance level, the current power supply level, and a threshold power supply level of each of the first corresponding accelerator and the second corresponding accelerator among the plurality of corresponding accelerators and further based on the respective performance goals of the kernels being executed by each of the first corresponding accelerator and the second corresponding accelerator among the plurality of corresponding accelerators.

[0064] Example 17 includes the subject matter as described in Example 13, and optionally, wherein the operation further includes: implementing an initial setting of data within a telemetry table.

[0065] Example 18 includes a method of operating a computing device, the method including: processing respective requests from respective nodes among a plurality of nodes within a network fabric, the respective requests being addressed to a plurality of corresponding accelerators among a plurality of pooled accelerators within the network fabric, each of the respective requests including information regarding a kernel to be executed by a corresponding accelerator among the plurality of corresponding accelerators, regarding the corresponding accelerator, and regarding a performance goal of the execution of the kernel; and controlling power supply to the corresponding accelerators based on the information within each of the respective requests.

[0066] Example 19 includes the subject matter as described in Example 18, and optionally, further includes: issuing a kernel to a corresponding accelerator for execution by the corresponding accelerator.

[0067] Example 20 includes the subject matter as described in Example 18, and optionally, further includes: implementing monitoring of the current performance level of a corresponding accelerator during execution of a kernel; and controlling power supply to the corresponding accelerator during execution of the kernel based on an updated version of the current performance level obtained from the monitoring.

[0068] Example 21 includes the subject matter as described in Example 18, and optionally, further includes: controlling power supply by controlling power supply to each of the plurality of corresponding accelerators based on the current performance level, the current power supply level, and a threshold power supply level of each of the plurality of corresponding accelerators.

[0069] Example 22 includes the subject matter as described in Example 21 and, optionally, further includes: controlling the power supply by respectively reducing the power supply to each of the plurality of corresponding accelerators or increasing the power supply to each of the plurality of corresponding accelerators in response to determining that the current performance level of each of the plurality of corresponding accelerators is higher than the performance target of the kernel being executed by each of the corresponding accelerators and determining that the current performance level of each of the plurality of corresponding accelerators is lower than the performance target of the kernel being executed by each of the corresponding accelerators.

[0070] Example 23 includes the subject matter as described in Example 21 and, optionally, further includes: redirecting the power supply from a first corresponding accelerator among the plurality of corresponding accelerators to a second corresponding accelerator among the plurality of corresponding accelerators based on the current performance level, the current power supply level, and a threshold power supply level of each of a first corresponding accelerator and a second corresponding accelerator among the plurality of corresponding accelerators and further based on the respective performance targets of the kernels being executed by each of the first corresponding accelerator and the second corresponding accelerator among the plurality of corresponding accelerators.

[0071] Example 24 includes a computing device configured to be used as part of a network architecture that includes a plurality of nodes and a plurality of pooled accelerators communicatively coupled to the nodes. The computing device includes: means for processing respective requests from respective nodes among the plurality of nodes, the respective requests being addressed to a plurality of corresponding accelerators among the plurality of pooled accelerators, each of the respective requests including information about a kernel to be executed by the corresponding accelerator among the plurality of corresponding accelerators, about the corresponding accelerator, and about the performance target of the execution of the kernel; and means for controlling the power supply to the corresponding accelerator based on the information in each of the respective requests.

[0072] Example 25 includes the subject matter as described in Example 24 and, optionally, further includes means for issuing a kernel to a corresponding accelerator for execution by the corresponding accelerator.

[0073] Example 26 includes a computing device configured to be used as part of a set of pooled accelerators communicatively coupled to a plurality of nodes via a switch within a network fabric. The computing device includes: a network interface configured to communicatively couple to the switch; a processing circuit communicatively coupled to the network interface to receive a kernel from the network interface. The processing circuit is further configured to: execute the kernel, which is issued to the network interface from a node among the plurality of nodes via the switch, and the kernel further has a performance target associated therewith; and process instructions from the switch to control power supply to the processing circuit based on the performance target of the kernel.

[0074] Example 27 includes the subject matter as described in Example 26, and optionally, wherein the processing circuit is further configured to: implement monitoring of the current performance level of the computing device during execution of the kernel; transmit data regarding the updated current performance level of the computing device during execution of the kernel to the switch; and control power supply to the processing circuit during execution of the kernel based on the updated current performance level obtained from the monitoring and based on a threshold power supply level of the computing device.

[0075] Example 28 includes the subject matter as described in Example 27, and optionally, wherein the processing circuit is further configured to: implement monitoring of the current power supply level of the computing device during execution of the kernel; transmit data regarding the updated current power supply level of the computing device during execution of the kernel to the switch; and control power supply to the processing circuit during execution of the kernel based on the updated current power supply level obtained from the monitoring.

[0076] Example 29 includes the subject matter as described in Example 27, and optionally, wherein the processing circuit is configured to transmit a threshold power supply level to the switch to effect an initial setting of data in a telemetry table of the switch.

[0077] Example 30 includes the subject matter as described in Example 26, and optionally, wherein the processing circuit includes a power management controller (PMC) configured to process instructions from the switch to control power supply to the processing circuit based on the performance target of the kernel.

[0078] Example 31 includes the subject matter as described in Example 26, and optionally, further includes an accelerator, which includes a field programmable gate array (FPGA)-based accelerator unit.

[0079] Example 32 includes the subject matter as described in Example 31, and optionally, wherein the FPGA-based accelerator unit includes a network interface (NI) and an input buffer, an output buffer, a programmable processor, and a memory controller communicatively coupled to the NI.

[0080] Example 33 includes a product that includes one or more tangible computer-readable non-transitory storage media including computer-executable instructions that, when executed by at least one computer processor, enable the at least one computer processor to implement operations at a computing device configured to be used as part of a set of pooled accelerators communicatively coupled within a network fabric to a plurality of nodes, the operations including: executing a kernel posted to the computing device from a node among the plurality of nodes, the kernel having a performance goal associated therewith; and processing instructions to control a power supply to a processing circuit of the computing device based on the performance goal of the kernel.

[0081] Example 34 includes the subject matter of Example 33 and optionally, wherein the computing device is used to communicatively couple to the plurality of nodes via a switch, and the operations further include: implementing monitoring of a current performance level of the computing device during execution of the kernel; transmitting data regarding the updated current performance level of the computing device during execution of the kernel to the switch; and controlling the power supply to the processing circuit during execution of the kernel based on the updated current performance level obtained from the monitoring and based on a threshold power supply level of the computing device.

[0082] Example 35 includes the subject matter of Example 34 and optionally, wherein the operations further include: implementing monitoring of a current power supply level of the computing device during execution of the kernel; transmitting data regarding the updated current power supply level of the computing device during execution of the kernel to the switch; and controlling the power supply to the processing circuit during execution of the kernel based on the updated current power supply level obtained from the monitoring.

[0083] Example 36 includes the subject matter of Example 34 and optionally, the operations further include: transmitting a threshold power supply level to the switch to effect an initial setting of data within a telemetry table of the switch.

[0084] Example 37 includes a method for execution at a computing device configured to be used as part of a set of pooled accelerators communicatively coupled within a network fabric to a plurality of nodes, the method including: executing a kernel posted to the computing device from a node among the plurality of nodes, the kernel having a performance goal associated therewith; and processing instructions to control a power supply to a processing circuit of the computing device based on the performance goal of the kernel.

[0085] Example 38 includes the subject matter as described in Example 37, and optionally, wherein the computing device is communicatively coupled to a plurality of nodes via a switch, and the method further includes: implementing monitoring of the current performance level of the computing device during execution of the kernel; transmitting data regarding the updated current performance level of the computing device during execution of the kernel to the switch; and controlling the power supply to the processing circuitry during execution of the kernel based on the updated current performance level obtained from the monitoring and based on a threshold power supply level of the computing device.

[0086] Example 39 includes the subject matter as described in Example 38, and optionally, further includes: implementing monitoring of the current power supply level of the computing device during execution of the kernel; transmitting data regarding the updated current power supply level of the computing device during execution of the kernel to the switch; and controlling the power supply to the processing circuitry during execution of the kernel based on the updated current power supply level obtained from the monitoring.

[0087] Example 40 includes the subject matter as described in Example 38, and optionally, further includes: transmitting a threshold power supply level to the switch to effect an initial setting of data within a telemetry table of the switch.

[0088] Example 41 includes a device configured to be used as part of a set of pooled accelerators communicatively coupled to a plurality of nodes via a switch within a network fabric, the computing device including: executing a kernel posted to the computing device from a node among the plurality of nodes, the kernel having a performance target associated therewith; and processing instructions to control the power supply to the processing circuitry of the computing device based on the performance target of the kernel.

[0089] Example 42 includes the subject matter as described in Example 41, and optionally, wherein the computing device is communicatively coupled to a plurality of nodes via a switch, and the computing device further includes: means for implementing monitoring of the current performance level of the computing device during execution of the kernel; means for transmitting data regarding the updated current performance level of the computing device during execution of the kernel to the switch; and means for controlling the power supply to the processing circuitry during execution of the kernel based on the updated current performance level obtained from the monitoring and based on a threshold power supply level of the computing device.

[0090] Although certain features have been illustrated and described herein, many modifications, substitutions, changes, and equivalents may occur to those skilled in the art. Accordingly, it is to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the present disclosure.

Claims

1. A computing device configured to be used as part of a network architecture, the network architecture including a plurality of nodes and a plurality of pooled accelerators communicatively coupled to the nodes, the computing device comprising: a memory storing instructions; and a processing circuit coupled to the memory, the processing circuit for executing the instructions to: receive a respective request from a respective node of the plurality of nodes, the respective request being addressed to a plurality of corresponding accelerators of the plurality of pooled accelerators, each respective request of the respective requests including information about a kernel to be executed by a corresponding accelerator of the plurality of corresponding accelerators, about the corresponding accelerator, and about a performance target of the execution of the kernel; and control a power supply to the corresponding accelerator based on the information in each respective request of the respective requests.

2. The computing device according to claim 1, wherein, The processing circuit is further for executing the instructions to issue the kernel to the corresponding accelerator for execution by the corresponding accelerator.

3. The computing device according to claim 1, wherein, The processing circuit is further for executing the instructions to: monitor a current performance level of the corresponding accelerator during execution of the kernel; and control the power supply to the corresponding accelerator during execution of the kernel based on an updated version of the current performance level obtained from the monitoring.

4. The computing device of claim 1, wherein: the device is for storing a telemetry table including data mapping each corresponding accelerator of the plurality of corresponding accelerators to a current performance level, a current power supply level, and a threshold power supply level of each corresponding accelerator of the plurality of corresponding accelerators; and the processing circuit is for executing the instructions to control the power supply by controlling the power supply to each corresponding accelerator of the plurality of corresponding accelerators based on determining the current performance level, the current power supply level, and the threshold supply level of each corresponding accelerator of the plurality of corresponding accelerators from the telemetry table.

5. The computing device according to claim 4, wherein, The processing circuit is for executing the instructions to control the power supply by reducing the power supply to each corresponding accelerator of the plurality of corresponding accelerators in response to determining that the current performance level of each corresponding accelerator of the plurality of corresponding accelerators is higher than a performance target of a kernel being executed by each corresponding accelerator of the plurality of corresponding accelerators.

6. The computing device according to claim 4, wherein, The processing circuit is for executing the instructions to control the power supply by increasing the power supply to each corresponding accelerator of the plurality of corresponding accelerators in response to determining that the current performance level of each corresponding accelerator of the plurality of corresponding accelerators is lower than a performance target of a kernel being executed by each corresponding accelerator of the plurality of corresponding accelerators.

7. The computing device according to claim 4, wherein The processing circuit is configured to execute the instruction to redirect power supply from a first corresponding accelerator among the plurality of corresponding accelerators to a second corresponding accelerator among the plurality of corresponding accelerators based on the current performance level, the current power supply level, and the threshold power supply level of each of the first corresponding accelerator and the second corresponding accelerator among the plurality of corresponding accelerators and further based on the respective performance objectives of the kernels being executed by each of the first corresponding accelerator and the second corresponding accelerator among the plurality of corresponding accelerators.

8. The computing device according to claim 4, wherein, The processing circuit is configured to execute the instruction to implement an initial setting of data in the telemetry meter.

9. The computing device according to any one of claims 1-8, further comprising at least one of a coherence switch, a memory switch, or a Peripheral Component Interconnect Express (PCIe) switch that includes an ingress interface for receiving the respective request from the node and an egress interface for controlling the power supply.

10. A method of operating a computing device, the method comprising: processing respective requests from respective nodes among a plurality of nodes within a network fabric, the respective requests being addressed to a plurality of corresponding accelerators among a plurality of pooled accelerators within the network fabric, each respective request among the respective requests including information about a kernel to be executed by a corresponding accelerator among the plurality of corresponding accelerators, about the corresponding accelerator, and about a performance objective of the execution of the kernel; and controlling power supply to the corresponding accelerators based on the information in each respective request among the respective requests.

11. The method according to claim 10, further comprising: Issuing the kernel to the corresponding accelerator for execution by the corresponding accelerator.

12. The method according to claim 10, further comprising: implementing monitoring of the current performance level of the corresponding accelerator during execution of the kernel; and controlling the power supply to the corresponding accelerator during execution of the kernel based on an updated version of the current performance level obtained from the monitoring.

13. The method according to claim 10, further comprising: Controlling the power supply by controlling the power supply to each of the plurality of corresponding accelerators based on the current performance level, the current power supply level, and the threshold power supply level of each of the plurality of corresponding accelerators.

14. An apparatus comprising means for performing the method according to any one of claims 10-13.

15. A machine-readable storage including machine-readable instructions that, when executed, implement the method according to any one of claims 10-13 or implement the computing device according to any one of claims 1-9.

Citation Information

Patent Citations

  • Method and system for controlling power consumption of pipeline processor

    CN101464721A

  • Delegating network processor operations to star topology serial bus interfaces

    CN101878475A