Reconfiguring a computing system using a circuit switch
By using circuit switches to quickly route data signals in the computing system, the problem of long temporary configuration reconfiguration time is solved, efficient execution of workloads with short duration is achieved, and the performance of computing system is improved.
Patent Information
- Application Number
- CN202310851703.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-19
- Filing Date
- 2023-07-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-07-12
AI Technical Summary
In some computing systems, reconfiguring temporary configurations can involve a longer period of time, making it unfeasible for short-term work.
By introducing circuit switches into the computing system, electronic or optical circuits are used to quickly route data signals between computing resources, allowing temporary configurations to be initiated with shorter lag times. The controller analyzes the upcoming workload, divides it into parts at multiple granularity levels, and decides whether to initiate temporary configurations to perform these parts.
Implementing optimized workload execution at multiple granularity levels, providing faster and more efficient workload execution, improving computing system performance, and allowing increased number of workloads executed simultaneously.
Smart Images

Figure CN117909057B_ABST
Abstract
Description
Background Art
[0001] Modern computer systems can include any number of components, such as a central processing unit (CPU), memory, chipset, and / or many other devices coupled together via an interconnect (e.g., a computer bus, network, etc.). The interconnect can transfer data between devices or components within a computer and between computers. For example, the interconnect can be used to read a data element from memory and provide that data element to a processor. Brief Description of the Drawings
[0002] Some embodiments are described with reference to the following drawings.
[0003] Figure 1 is a schematic diagram of an example computing system according to some embodiments.
[0004] Figure 2 is an illustration of an example operation according to some embodiments.
[0005] Figures 3A to 3D is an illustration of an example configuration according to some embodiments.
[0006] Figure 4 is an illustration of an example software interface according to some embodiments.
[0007] Figure 5 is an illustration of an example process according to some embodiments.
[0008] Figure 6 is a schematic diagram of an example computing system according to some embodiments.
[0009] Figure 7 is an illustration of an example process according to some embodiments.
[0010] Figure 8 is a diagram of an example machine-readable medium storing instructions according to some embodiments.
[0011] Throughout the drawings, like reference numerals represent similar but not necessarily identical elements. The drawings are not necessarily to scale, and the dimensions of some parts may be exaggerated to more clearly show the examples illustrated. Additionally, the drawings provide examples and / or embodiments consistent with the description; however, the description is not limited to the examples and / or embodiments provided in the drawings. Detailed Description
[0012] In the present disclosure, unless the context clearly indicates otherwise, the use of the terms "a", "an", or "the" is also intended to include the plural forms. Additionally, the terms "includes", "including", "comprises", "comprising", "have", or "having", when used in the present disclosure, specify the presence of the stated elements, but do not preclude the presence or addition of other elements.
[0013] In some examples, a computing system may include a plurality of computing resources connected via a network architecture. For example, such computing resources may include processing devices, memory devices, accelerator devices, storage devices, etc. In some examples, a subset of the computing resources (referred to herein as a "temporary configuration") may be allocated to perform different computing tasks. Additionally, after completing its computing task, the temporary configuration may be deleted, and its computing resources may be made available for reconfiguration into a new (or multiple) temporary configuration. However, in some computing systems, reconfiguring a temporary configuration may involve a relatively long time (referred to herein as "latency"). For example, in a computing system using a packet-switching networking architecture, it may take a significant amount of time to route data in packets (e.g., form, send, cache, and decode data packets). Thus, in such a computing system, using a temporary configuration may not be feasible for tasks of relatively short duration (e.g., tasks that include a relatively small number of instructions).
[0014] According to some embodiments of the present disclosure, a computing system may include a circuit switch to form a temporary configuration of computing resources. The circuit switch may use electrical circuitry or optical circuitry to quickly route data signals between computing resources and may thus allow a temporary configuration to be initiated with a relatively short latency (e.g., one percent of the latency when using packet switching). In some embodiments, a controller may analyze the upcoming workload of the computing system and may partition the workload into parts at multiple granularity levels. For example, the controller may partition the upcoming workload into a series of parts, where the parts correspond to tasks, microservices, or instructions. For each workload part, the controller may determine whether using a temporary configuration will improve the performance of the computing system, and if so, may initiate a temporary configuration to perform the workload part.
[0015] In some embodiments, since the circuit switch can initiate a temporary configuration with a relatively short latency, it is feasible to use the temporary configuration to execute a workload portion with a relatively short duration (e.g., for a relatively small number of instructions). Thus, the execution of the workload can be optimized at multiple granularity levels. In this way, some embodiments can provide improved execution of the workload. Additionally, since computing resources can be dynamically reallocated among workloads during execution, some embodiments can allow a computing system to increase the number of concurrently executed workloads. Details regarding the use of the temporary configuration are described below with reference to Figures 1 to 8 describe various details of using the temporary configuration.
[0016] Figure 1 - Example computing system
[0017] Figure 1 FIG. 100 illustrates an example computing system 100 according to some embodiments. The computing system 100 may include one or more circuit switches 110 to provide interconnection between any number of computing resources. Such computing resources may include general-purpose processors (GPPs) 120, special-purpose processors (SPPs) 130, memory devices 140, storage devices 150, and other computing resources (not shown). For example, the general-purpose processor 120 may include various types of central processing units (CPUs), system-on-chips (SoCs), processing cores, etc. The special-purpose processor 130 may include various types of special-purpose processing devices, such as graphics processing units (GPUs), digital signal processors (DSPs), math processors, cryptographic processors, network processors, etc. The memory device 140 may include various types of memories, such as dynamic random access memory (DRAM), static random access memory (SRAM), etc. The storage device 150 may include one or more non-transitory storage media, such as hard disk drives (HDDs), solid-state drives (SSDs), optical discs, etc., or combinations thereof. In some embodiments, the computing system 100 and the included components may operate according to the Compute Express Link (CXL) protocol or specification (e.g., the CXL 1.1 specification).
[0018] In some embodiments, the circuit switch 110 may be an electronic circuit switch that routes signals within the electrical domain (e.g., as electrical signals). In other embodiments, the circuit switch 110 may be an optical circuit switch that routes signals within the optical domain (e.g., as photon signals) without converting the signal to the electrical domain. The circuit switch 110 may be controlled by a controller 105. In some examples, the controller 105 may be implemented via hardware (e.g., an electronic circuit) or a combination of hardware and programming (e.g., including at least one processor and instructions executable by the at least one processor and stored on at least one machine-readable storage medium).
[0019] In some embodiments, the controller 105 may analyze an upcoming workload of the computing system 100 and may partition the workload into portions at multiple granularity levels. For example, the controller may partition an upcoming workload into a series of portions, where each portion corresponds to one or more jobs, one or more microservices, or one or more instructions. As used herein, the term "microservice" may refer to a software component that performs a single function, includes multiple instructions, and executes independently of other microservices. Additionally, as used herein, the term "job" may refer to a software application or module that includes multiple microservices.
[0020] For each workload portion, the controller 105 may determine whether to initiate a temporary configuration to execute the workload portion. If so, the controller 105 may allocate a subset of resources to form a temporary configuration. The temporary configuration may be used to execute the corresponding workload portion. An example operation using a temporary configuration is described below with reference to Figure 2 Describe an example operation using a temporary configuration.
[0021] In some embodiments, the computing system 100 may process different workloads at the same time. For example, the (multiple) circuit switches 110 may be configured such that the resources of the computing system 100 are partitioned into a first resource group allocated to a first workload and a second resource group allocated to a second workload. In such an embodiment, different workloads may be executed independently (e.g., in parallel) in different resource groups. Each workload may be associated with a different user entity (e.g., a customer, an application, an organization, etc.). In some embodiments, the controller 105 may initiate one or more temporary configurations within each resource group (i.e., for each individual workload). For example, within a resource group, the available resources may be allocated to two or more temporary configurations (e.g., to execute two workload portions in parallel). In another example, within a resource group, the available resources may initially be allocated to a first temporary configuration to execute a first workload portion, released after the first workload portion is completed, and then allocated to a second temporary configuration to execute a second workload portion.
[0022] Figure 2 - Example operation
[0023] Figure 2 A diagram illustrating an example operation 200 in accordance with some embodiments is shown. Operation 200 may be performed by the computing system 100 (shown in Figure 1 ). As Figure 2As shown, block 210 may include the source code of the compilation workload 201, thereby generating the compiled code 220 to be executed. The workload 201 consists of any number of jobs 202 (e.g., programs, applications, modules, etc.). Each job 202 consists of any number of microservices 203. In addition, each microservice 203 consists of any number of instructions 204.
[0024] Block 230 may include analyzing the source code to determine a series of workload portions at multiple granularity levels (e.g., jobs, microservices, or instructions). Block 240 may include determining a temporary configuration for each workload portion. Block 250 may include controlling a circuit switch to provide the temporary configuration for executing the workload portion. Block 260 may include executing the workload portion using the corresponding temporary configuration and thereby generating a workload output 270.
[0025] In some embodiments, each temporary configuration may be a subset of the computing resources included in the computing system (e.g., Figure 1 a subset of the GPP 120, SPP 130, memory device 140, and / or storage device 150 as shown), and this subset of computing resources is selected to execute the corresponding workload portion. For example, for a workload portion that includes performing data encryption, the temporary configuration may include at least one encryption processor to perform the encryption process, and sufficient memory to store the variables and other data used during the encryption process. In addition, to initiate the temporary configuration, a circuit switch (e.g., Figure 1 the circuit switch 110 as shown) may be controlled to provide an appropriate interconnection between the computing resources included in the temporary configuration. Some examples of temporary configurations are described below with reference to Figures 3A to 3D
[0026] In some embodiments, each workload portion may be defined as a job portion, a microservice portion, or an instruction-level portion. As used herein, the term "job portion" refers to program code for executing one or more jobs. In addition, the term "microservice portion" refers to program code for executing one or more microservices. In addition, the term "instruction-level portion" refers to program code for executing one or more iterations of a specific instruction. For example, an instruction-level portion may include a computational instruction and associated (multiple) loop instructions to repeat the computational instruction a given number of times. In this example, the instruction-level portion may be referred to as a "loop instruction" that is executed in multiple iterations. In another example, the instruction-level portion may include only a single computational instruction that is executed once.
[0027] In some embodiments, each workload portion may be determined as a contiguous portion for which execution can be accelerated or otherwise improved by using a particular temporary configuration of the computing system. For example, if the workload includes two adjacent instructions and if these two instructions execute most quickly when using different temporary configurations for each of the two instructions, then these two instructions can be identified as two separate instruction-level portions. In another example, if the workload includes two adjacent microservices and if these two microservices execute most quickly by using the same temporary configuration for the two microservices, then these two microservices can be identified together as a single microservice portion.
[0028] Figures 3A to 3D — Example temporary configuration
[0029] Figures 3A to 3D An example temporary configuration according to some embodiments is shown. For example, these example temporary configurations may be implemented in the computing system 100 (as Figure 1 shown).
[0030] Now referring to Figure 3A , an example sequence 310 of workload portions is shown. In the example, the first workload portion (instruction-level portion "Instruction 1") in the sequence 310 is analyzed to determine a first temporary configuration 320. As shown, the first temporary configuration 320 includes a single GPP 120 and a single memory device 140 connected by a circuit switch 110. For example, the first workload portion may be a relatively simple instruction-level portion that does not require any specialized processing or a large amount of memory.
[0031] Now referring to Figure 3B , the second workload portion ("Microservice 1") in the sequence 310 is analyzed to determine a second temporary configuration 330. As shown, the second temporary configuration 330 includes a single GPP 120, a single memory device 140, and two SPPs 130 connected by a circuit switch 110. For example, the second workload portion may be a microservice portion that includes video signal processing and thus may require multiple SPPs 130 as a graphics accelerator.
[0032] Now referring to Figure 3C , the third workload portion ("Microservice 2") in the sequence 310 is analyzed to determine a third temporary configuration 340. As shown, the third temporary configuration 340 includes a single GPP 120, two memory devices 140, and three SPPs 130 connected by a circuit switch 110. For example, the third workload portion may be a microservice portion that includes complex encryption calculations and thus may require multiple SPPs 130 as an encryption accelerator.
[0033] Now referring toFigure 3D , the fourth workload portion (instruction-level portion "Instruction 2") in Sequence 310 is analyzed to determine the fourth temporary configuration 350. As shown, the fourth temporary configuration 350 includes a single GPP 120 and three memory devices 140 connected by a circuit switch 110. For example, the fourth workload portion can be a loop instruction that performs multiple iterations to modify a very large data matrix in memory, and thus may require multiple memory devices 140 to maintain the data matrix in memory.
[0034] Figure 4 - Example software interface
[0035] Figure 4 A diagram illustrating an example software interface according to some embodiments is shown. Circuit switch control logic 470 can analyze the upcoming workload of a computing system and can determine a series of workload portions at multiple granularity levels. Circuit switch control logic 470 can determine whether to initiate a temporary configuration for a workload portion, and if so, can cause the circuit switch to provide a temporary configuration for executing the workload portion. Circuit switch control logic 470 can be implemented in a hardware controller (e.g., Figure 1 the controller 105 shown), in software executed by the controller, or a combination thereof. In some embodiments, circuit switch control logic 470 can interact with multiple software layers of a computing system to determine whether to initiate each temporary configuration. For example, as Figure 4 shown, circuit switch control logic 470 can interact with an application layer 410, a compiler layer 420, an operating system layer 430, an instruction set architecture (ISA) hardware layer 440, an input / output (I / O) hardware layer 450, and a hardware system management layer 460.
[0036] Figure 5 – Example process for initiating a temporary configuration
[0037] Now referring to Figure 5 , an example process 500 for initiating a temporary configuration according to some embodiments is shown. Process 500 can be executed by a controller that executes instructions (e.g., Figure 1 the controller 105 shown). Process 500 can be implemented in hardware or a combination of hardware and programming (e.g., machine-readable instructions executable by one or more processors). The machine-readable instructions can be stored in a non-transitory computer-readable medium, such as an optical, semiconductor, or magnetic storage device. The machine-readable instructions can be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc. For illustrative purposes, the details of process 500 are described below with reference to Figures 1 to 4 , Figures 1 to 4 showing an example according to some embodiments. However, other embodiments are also possible.
[0038] Block 510 may include analyzing program code to be executed by a computing system. Block 520 may include determining a series of workload portions having multiple granularity levels. For example, referring to Figures 1 to 3A , controller 105 analyzes the source code of upcoming workload 201 and identifies a sequence 310 of workload portions at multiple granularity levels (e.g., jobs, microservices, instructions).
[0039] At block 530, a loop (defined by blocks 530 to 590) may be entered to process each workload portion. Block 540 may include determining a temporary configuration for executing the current workload portion. Decision block 550 may include determining whether the temporary configuration is a change from the current system configuration. If not ("no"), then process 500 may return to block 530 (e.g., to process another workload portion). For example, referring to Figures 1 to 3A , controller 105 determines a first temporary configuration 320 to execute a first workload portion (instruction-level portion "instruction 1") and determines whether the first temporary configuration 320 is different from the current existing configuration of the computing system being used to execute workload 210.
[0040] Referring again to Figure 5 , if at decision block 550 it is determined that the temporary configuration is a change from the current system configuration ("yes"), then process 500 may continue at block 560, which includes determining the performance benefit of the new temporary configuration. Block 570 may include determining the performance cost of the new temporary configuration. Decision block 580 may include determining whether the performance benefit exceeds the performance cost. If not ("no"), then process 500 may return to block 530 (e.g., to process another workload portion). For example, referring to Figures 1 to 3A , controller 105 determines that the first temporary configuration 320 is different from the current existing configuration of computing system 100, and in response, controller 105 calculates an estimated performance improvement (e.g., reduction in execution time, energy savings, etc.) when using the temporary configuration to execute the workload portion. Further, controller 105 calculates an estimated performance cost for initiating the temporary configuration (e.g., loss of processing time when the computing system is reconfigured, energy cost, etc.). Controller 105 may determine whether to initiate the temporary configuration based on a comparison of the performance improvement with the performance cost. In some embodiments, the performance benefit may be determined, at least in part, based on the number of iterations in which the workload portion is repeated. For example, for an instruction-level portion that is a loop instruction, the performance benefit may increase as the number of iterations of the loop instruction increases.
[0041] In some embodiments, the decision of whether to initiate a temporary configuration can also be based on the relative priority or importance of different temporary configurations that can be allocated the same computing resources. For example, if two different temporary configurations both require a particular processing device, then the processing device can be allocated to the temporary configuration with the higher priority (e.g., based on a service level agreement (SLA)).
[0042] Referring again to Figure 5 , if it is determined at decision block 580 that the performance gain exceeds the performance cost (“yes”), then process 500 can continue at block 590, which includes initiating a new temporary configuration to execute a workload portion. After block 590, process 500 can return to block 530 (e.g., to process another workload portion). After all workload portions are completed at block 530, process 500 can be completed. For example, referring to Figures 1 to 3A , the controller 105 determines that the performance gain from using the new temporary configuration exceeds the performance cost of initiating the new temporary configuration. In response to this determination, the controller 105 deallocates the computing resources used by the previous temporary configuration (if any) and initiates the new temporary configuration. When initiating the new temporary configuration, the controller 105 prepares the memory device 140 for executing the workload portion, controls the circuit switch 110 to provide a data path for the temporary configuration, and verifies the changes to the circuit switch 110. In addition, the controller 105 grants the processing devices (e.g., GPPs 120 and / or SPP 130) access to the prepared memory device 140 via the circuit switch 110, and the processing devices execute the workload portion.
[0043] Figure 6 - Example computing system
[0044] Figure 6 FIG. shows a schematic diagram of an example computing system 600. In some examples, the computing system 600 can generally correspond to part or all of the computing system 100 (as Figure 1 shown). As shown, the computing system 600 can include a processing device 602, at least one circuit switch 606, a memory device 604, and a controller 605. In some embodiments, the controller 605 can be a hardware processor that executes instructions 610 to 650. The instructions 610 to 650 can be stored in a non-transitory machine-readable memory.
[0045] Instruction 610 may be executed to identify a first instruction - level portion and a second instruction - level portion to be continuously executed by a computing system. Instruction 620 may be executed to determine a first subset of processing devices and a first subset of memory devices to be used for executing the first instruction - level portion. Instruction 630 may be executed to control a circuit switch to interconnect the first subset of processing devices and the first subset of memory devices during the execution of the first instruction - level portion. For example, referring to Figures 1 to 3A , the controller 105 analyzes the source code of the upcoming workload 201 and identifies a sequence 310 of workload portions at multiple granularity levels (e.g., tasks, microservices, instructions). The controller 105 determines a first temporary configuration to execute the first instruction - level portion (e.g., a loop instruction) in the workload sequence and controls the circuit switch 110 and other system components to provide the first temporary configuration. The first instruction - level portion is executed using the first temporary configuration of the computing system 100.
[0046] Referring again to Figure 6 , instruction 640 may be executed to determine a second subset of processing devices and a second subset of memory devices to be used for executing the second instruction - level portion. Instruction 650 may be executed to control a circuit switch to interconnect the second subset of processing devices and the second subset of memory devices during the execution of the second instruction - level portion. For example, referring to Figures 1 to 3A , the controller 105 determines a second temporary configuration to execute the second instruction - level portion that follows the first instruction - level portion and controls the circuit switch 110 and other system components to provide the second temporary configuration. Then the second instruction - level portion is executed using the second temporary configuration of the computing system 100.
[0047] Figure 7 - Example process
[0048] Now referring to Figure 7 , an example process 700 according to some embodiments is shown. The process 700 may be executed by a controller that executes instructions (e.g., the controller 105 shown in Figure 1 ). The process 700 may be implemented in hardware or a combination of hardware and programming (e.g., machine - readable instructions executable by one or more processors). The machine - readable instructions may be stored in a non - transitory computer - readable medium, such as an optical, semiconductor, or magnetic storage device. The machine - readable instructions may be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc.
[0049] Block 710 may include a first instruction-level portion and a second instruction-level portion to be continuously executed by a controller of a computing system, the computing system including a plurality of processing devices, a plurality of memory devices, and a circuit switch. Block 720 may include a first subset of processing devices and a first subset of memory devices to be used for executing the first instruction-level portion, determined by the controller. Block 730 may include the controller controlling the circuit switch to interconnect the first subset of processing devices and the first subset of memory devices during execution of the first instruction-level portion.
[0050] Block 740 may include a second subset of processing devices and a second subset of memory devices to be used for executing the second instruction-level portion, determined by the controller. Block 750 may include the controller controlling the circuit switch to interconnect the second subset of processing devices and the second subset of memory devices during execution of the second instruction-level portion. After block 750, process 700 may be completed.
[0051] Figure 8 - Exemplary machine-readable medium
[0052] Figure 8 A machine-readable medium 800 storing instructions 810 to 850 is shown in accordance with some embodiments. Instructions 810 to 850 may be executed by a single processor, multiple processors, a single processing engine, multiple processing engines, etc. The machine-readable medium 800 may be a non-transitory storage medium, such as an optical, semiconductor, or magnetic storage medium.
[0053] Executable instruction 810 to identify a first instruction-level portion and a second instruction-level portion to be continuously executed by a computing system, the computing system including a plurality of processing devices, a plurality of memory devices, and a circuit switch. Executable instruction 820 to determine a first subset of processing devices and a first subset of memory devices to be used for executing the first instruction-level portion. Executable instruction 830 to control the circuit switch to interconnect the first subset of processing devices and the first subset of memory devices during execution of the first instruction-level portion.
[0054] Executable instruction 840 to determine a second subset of processing devices and a second subset of memory devices to be used for executing the second instruction-level portion. Executable instruction 850 to control the circuit switch to interconnect the second subset of processing devices and the second subset of memory devices during execution of the second instruction-level portion.
[0055] According to some embodiments of the present disclosure, a computing system may include a circuit switch to form a temporary configuration of computing resources. The circuit switch may use electronic or optical circuits to rapidly route data signals between computing resources, thereby allowing the temporary configuration to be initiated with a relatively short latency. A controller may analyze an upcoming workload to identify workload portions at multiple granularity levels and may determine a temporary configuration of the computing system to execute the workload portions. In this way, some embodiments may provide faster and / or more efficient execution of workloads and may thereby improve the performance of the computing system.
[0056] Note that while Figures 1 to 8 various examples are shown, embodiments are not limited in this regard. For example, referring to Figure 1 , it may be envisioned that the computing system 100 may include additional devices and / or components, fewer components, different components, different arrangements, etc. In another example, it may be envisioned that the functionality of the controller 105 described above may be included in any other engine or software of the computing system 100. Other combinations and / or variations are also possible.
[0057] Data and instructions are stored in respective storage devices, which are implemented as one or more computer-readable or machine-readable storage media. The storage media include different forms of non-transitory memory, including semiconductor memory devices such as dynamic or static random access memory (DRAM or SRAM), erasable and programmable read-only memory (EPROM), electrically erasable and programmable read-only memory (EEPROM), and flash memory; magnetic disks such as fixed disks, floppy disks, and removable disks; other magnetic media including magnetic tape; optical media such as compact disks (CDs) or digital video disks (DVDs); or other types of storage devices.
[0058] Note that the instructions discussed above may be provided on a computer-readable or machine-readable storage medium, or alternatively, may be provided on multiple computer-readable or machine-readable storage media distributed in a large system that may have multiple nodes. Such a computer-readable or machine-readable storage medium or media is considered to be part of an article (or manufacture). An article or manufacture may refer to any single component or multiple components that are manufactured. The storage medium or media may be located in a machine that runs the machine-readable instructions or at a remote site from which the machine-readable instructions may be downloaded over a network for execution.
[0059] In the foregoing description, numerous details are set forth to provide an understanding of the subject matter disclosed herein. However, implementations may be practiced without some of these details. Other embodiments may include modifications and variations to the above details. The appended claims are intended to cover such modifications and variations.
Claims
1. A computing system, the computing system comprising: a plurality of processing devices; a plurality of memory devices; a circuit switch; and a controller, the controller being configured to: identify a first instruction-level portion and a second instruction-level portion to be consecutively executed by the computing system; determine a first subset of the plurality of processing devices and a first subset of the plurality of memory devices to be used for executing the first instruction-level portion; control the circuit switch to interconnect the first subset of the plurality of processing devices and the first subset of the plurality of memory devices during the execution of the first instruction-level portion; determine a second subset of the plurality of processing devices and a second subset of the plurality of memory devices to be used for executing the second instruction-level portion; and control the circuit switch to interconnect the second subset of the plurality of processing devices and the second subset of the plurality of memory devices during the execution of the second instruction-level portion, wherein a first temporary configuration includes the first subset of the plurality of processing devices and the first subset of the plurality of memory devices, wherein a second temporary configuration includes the second subset of the plurality of processing devices and the second subset of the plurality of memory devices; determine a performance benefit of the second temporary configuration; determine a performance cost of the second temporary configuration; determine whether the performance benefit exceeds the performance cost; and in response to determining that the performance benefit exceeds the performance cost, initiate the second temporary configuration in the computing system.
2. The computing system according to claim 1, wherein the performance benefit is an estimated reduction in execution time for executing the second instruction-level portion using the second temporary configuration; and the performance cost is an estimated loss of processing time for initiating the second temporary configuration.
3. The computing system according to claim 1, wherein the controller is further configured to determine whether the second temporary configuration is different from the first temporary configuration; and in response to determining that the second temporary configuration is different from the first temporary configuration, initiate the second temporary configuration.
4. The computing system according to claim 1, wherein the circuit switch is an optical path switch for routing signals in the optical domain.
5. The computing system according to claim 1, wherein the circuit switch is a circuit switch for routing electrical signals, and wherein the circuit switch does not process data packets.
6. The computing system according to claim 1, the controller being configured to: analyze an upcoming workload of the computing system; and partition the upcoming workload into portions at multiple granularity levels including tasks, microservices, and instructions.
7. The computing system according to claim 1, wherein the first instruction-level portion includes a loop instruction to be repeated for a plurality of iterations, and wherein the plurality of processing devices comprising: a plurality of general-purpose processors; and a plurality of dedicated processors.
8. A method, the method comprising: identifying, by a controller of a computing system, a first instruction-level portion and a second instruction-level portion to be consecutively executed by the computing system, the computing system including a plurality of processing devices, a plurality of memory devices, and a circuit switch; The controller determines a first subset of the plurality of processing devices and a first subset of the plurality of memory devices to be used for executing the first instruction-level portion; The controller controls the circuit switch to interconnect the first subset of the plurality of processing devices and the first subset of the plurality of memory devices during the execution of the first instruction-level portion; The controller determines a second subset of the plurality of processing devices and a second subset of the plurality of memory devices to be used for executing the second instruction-level portion; and The controller controls the circuit switch to interconnect the second subset of the plurality of processing devices and the second subset of the plurality of memory devices during the execution of the second instruction-level portion, wherein the first temporary configuration includes the first subset of the plurality of processing devices and the first subset of the plurality of memory devices, wherein the second temporary configuration includes the second subset of the plurality of processing devices and the second subset of the plurality of memory devices; Determine the performance benefit of the second temporary configuration; Determine the performance cost of the second temporary configuration; Determine whether the performance benefit exceeds the performance cost; and In response to determining that the performance benefit exceeds the performance cost, initiate the second temporary configuration in the computing system.
9. The method according to claim 8, wherein: The performance benefit is an estimated reduction in the execution time of executing the second instruction-level portion using the second temporary configuration; and The performance cost is an estimated loss of processing time for initiating the second temporary configuration.
10. The method according to claim 8, further comprising: Determine whether the second temporary configuration is different from the first temporary configuration; and In response to determining that the second temporary configuration is different from the first temporary configuration, initiate the second temporary configuration.
11. The method according to claim 8, further comprising Using the first subset of the plurality of processing devices and the first subset of the plurality of memory devices to execute the first instruction-level portion, including executing multiple iterations of a loop instruction.
12. A non-transitory machine-readable medium storing instructions that, when executed, cause a processor to: Identify a first instruction-level portion and a second instruction-level portion to be continuously executed by a computing system, the computing system including a plurality of processing devices, a plurality of memory devices, and a circuit switch; Determine a first subset of the plurality of processing devices and a first subset of the plurality of memory devices to be used for executing the first instruction-level portion; Control the circuit switch to interconnect the first subset of the plurality of processing devices and the first subset of the plurality of memory devices during the execution of the first instruction-level portion; Determine a second subset of the plurality of processing devices and a second subset of the plurality of memory devices to be used for executing the second instruction-level portion; and Control the circuit switch to interconnect the second subset of the plurality of processing devices and the second subset of the plurality of memory devices during the execution of the second instruction-level portion, Wherein, the first temporary configuration includes the first subset of the plurality of processing devices and the first subset of the plurality of memory devices, Wherein, the second temporary configuration includes the second subset of the plurality of processing devices and the second subset of the plurality of memory devices; Determine the performance benefit of the second temporary configuration; Determine the performance cost of the second temporary configuration; Determine whether the performance benefit exceeds the performance cost; and In response to determining that the performance benefit exceeds the performance cost, initiate the second temporary configuration in the computing system.
13. The non-transitory machine-readable medium according to claim 12, Wherein: The performance benefit is an estimated reduction in the execution time of executing the second instruction-level portion using the second temporary configuration; and The performance cost is an estimated loss of processing time for initiating the second temporary configuration.
14. The non-transitory machine-readable medium according to claim 12, wherein the circuit switch is an optical path switch that routes signals within the optical domain.
Citation Information
Patent Citations
Call center and method thereof for processing a large number of speech path requests
CN107968895A
Apparatus and method for memory management in a graphics processing environment
US20180293183A1