Power consumption management method, multi-processing unit system and power consumption management module
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA DAMO (HANGZHOU) TECH CO LTD
- Filing Date
- 2021-11-05
- Publication Date
- 2026-08-07
AI Technical Summary
其对于各运算单元自身功耗利用效率较优,但对于多运算单元系统整体,并不有利于总体功耗利用效率
[0032] In the embodiments of this disclosure, the allocation of local power budgets for each computing unit takes into account the power management parameters of each computing unit, rather than distributing them equally. This makes the local power budgets more closely match the actual needs of each computing unit, thereby improving overall power utilization efficiency. Furthermore, the allocation of local power budgets for each computing unit also considers the constraints of the global power budget, preventing individual computing unit power efficiency optimization from resulting in low overall power efficiency, further improving the overall power utilization efficiency of the multi-computing-unit system.
Smart Images

Figure CN116088662B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a power management method, a multiprocessor unit system, and a power management module. Background Technology
[0002] In power management of multi-processor multi-unit systems, power consumption and performance are closely related. When the power requirements of each unit cannot be fully met, the performance of some units will degrade, leading to phenomena such as dark silicon. Conversely, when the power requirements of each unit are fully met, the overall power consumption of the multi-unit system becomes very high, resulting in resource waste and poor energy efficiency for both.
[0003] Generally, power management in multi-computing unit systems can be categorized into centralized power management and distributed power management. In centralized power management, each computing unit shares the same power budget and manages its own power consumption. However, the actual power consumption required by each computing unit varies significantly, and a shared functional budget is detrimental to achieving optimal overall power efficiency. In distributed power management, different computing units perform their own power budget allocation and management. While this offers better power efficiency for each individual computing unit, it is not beneficial for the overall power efficiency of the multi-computing unit system.
[0004] Therefore, there is room for improvement in the overall power utilization efficiency of both centralized power management and distributed power management. Summary of the Invention
[0005] In view of this, embodiments of the present disclosure provide a power management method, a multi-processor unit system, and a power management module to improve the overall power utilization efficiency of a multi-processor unit system.
[0006] According to a first aspect of the present disclosure, a power management method for a multi-computing unit system is provided. The multi-computing unit system includes a plurality of local power management units and a global power management unit, each local power management unit corresponding to a computing unit in the multi-computing unit system. The method includes: using the global power management unit to obtain a global power budget for the multi-computing unit system; using the global power management unit to allocate a local power budget for each computing unit according to the global power budget and power management parameters of each computing unit; using the local power management unit to manage the local power resources of the corresponding computing unit according to the allocated local power budget; and using the local power management unit to report the power management parameters of the computing unit to the global power management unit.
[0007] In another implementation of this disclosure, the power management parameters include at least one of energy consumption parameters, performance parameters, and workload parameters.
[0008] In another implementation of this disclosure, the energy consumption parameter includes at least one of power consumption value, power consumption status, and local budget utilization rate.
[0009] In another implementation of this disclosure, obtaining the global power budget of the multi-computing unit system includes: determining the global power budget based on the previous power management parameters of each computing unit and the previous global power budget.
[0010] In another implementation of this disclosure, determining the global power budget based on the previous power management parameters and the previous global power budget of each computing unit includes: using a first machine learning model or a first rule to determine the global power budget based on the previous power management parameters and the previous global power budget of each computing unit, so that the predicted power utilization of the global power budget is not lower than the actual power utilization of the previous global power budget.
[0011] In another implementation of this disclosure, determining the global power budget based on the previous power management parameters and the previous global power budget of each computing unit includes: determining the global power budget of the multi-computing unit system in the second power management period based on the power management parameters of each computing unit in the first power management period and the global power budget of the multi-computing unit system in the first power management period, wherein the second power management period follows the first power management period, and the first power management period and the second power management period form a continuous power management period.
[0012] In another implementation of this disclosure, the first power management period and the second power management period are set according to one of the following: the first power management period and the second power management period are two consecutive periods of equal duration; the first power management period and the second power management period are respectively used to execute two consecutive power resource management tasks.
[0013] In another implementation of this disclosure, the step of allocating the local power budget of each computing unit according to the global power budget and the power management parameters of each computing unit includes: allocating the local power budget of each computing unit in the second power management period according to the global power budget of the multi-computing unit system in the second power management period and the power management parameters of each computing unit in the first power management period.
[0014] In another implementation of this disclosure, the step of allocating the local power budget of each computing unit according to the global power budget and the power management parameters of each computing unit includes: determining the total power budget of each computing unit in the global power budget and the power budget of the global power management unit; and allocating the local power budget of each computing unit according to the total power budget of each computing unit and the power management parameters of each computing unit.
[0015] In another implementation of this disclosure, the step of allocating the local power budget of each computing unit based on the total power budget of each computing unit and the power management parameters of each computing unit includes: using a second rule to allocate the local power budget of each computing unit based on the total power budget of each computing unit and the power management parameters of each computing unit; or, obtaining the local power budget of each computing unit based on the total power budget of each computing unit and the power management parameters of each computing unit as input to a second machine learning model.
[0016] In another implementation of this disclosure, the step of obtaining the local power budget of each computing unit based on the total power budget of each computing unit and the power management parameters of each computing unit as input to the second machine learning model includes: determining the current power budget reference parameters of each computing unit according to the previous power management parameters of each computing unit; and inputting the current power budget reference parameters of each computing unit and the current total power budget of each computing unit into the second machine learning model to obtain the local power budget of each computing unit.
[0017] In another implementation of this disclosure, the second machine learning model is a reinforcement learning model, wherein determining the current power budget reference parameter of each computing unit based on the previous power management parameters of each computing unit includes: calculating the current power budget reference parameter of each computing unit based on the previous power management parameters of each computing unit, as the state and reward of the reinforcement learning model.
[0018] In another implementation of this disclosure, the multi-computing unit system further includes shared resources shared by the multi-computing units, wherein allocating the local power budget of each computing unit according to the global power budget and the power management parameters of each computing unit includes: allocating the local power budget of each computing unit according to the global power budget, the power management parameters of each computing unit and the power management parameters of the shared resources.
[0019] In another implementation of this disclosure, the step of managing the local power resources of the corresponding computing unit according to the allocated local power budget includes: determining and executing a local power resource management strategy based on the allocated local power budget and the power management parameters of the local power management unit.
[0020] In another implementation of this disclosure, the local power resource management strategy includes at least one of dynamic voltage frequency scaling, power gating, clock gating, setting clock frequency, and adjusting the order of multiple tasks to be executed.
[0021] In another implementation of this disclosure, determining the local power resource management strategy based on the allocated local power budget and the power management parameters of the local power management unit includes: using a third machine learning model or a third rule to determine the local power resource management strategy based on the allocated local power budget and the power management parameters of the local power management unit.
[0022] In another implementation of this disclosure, the global power management unit performs a first process when the power management parameters of each arithmetic unit meet predetermined conditions during the first power management period.
[0023] In another implementation of this disclosure, the power management parameters include power consumption values, and the predetermined conditions include: the power consumption values of each computing unit during a first power management period and the global power budget exceeding the first power management period.
[0024] In another implementation of this disclosure, the first process includes one of the following: reducing the power consumption value of each computing unit during the first power management period, so that the sum of the power consumption values of each computing unit during the first power management period does not exceed the global power budget of the first power management period; determining the amount by which the sum of the power consumption values of the computing units during the first power management period exceeds the global power budget, and using the amount to offset the global power budget of the second power management period.
[0025] In another implementation of this disclosure, the global power management unit is located within one of the multiple computing units, or outside the multiple computing units.
[0026] According to a second aspect of the present disclosure, a multi-processor unit system is provided, including multiple processing units, multiple local power management units, and a global power management unit. Each local power management unit corresponds to one of the multiple processing units. The global power management unit is used to obtain the global power budget of the multi-processor unit system, and allocate the local power budget of each processing unit according to the global power budget and the power management parameters reported by each processing unit. The local power management unit is used to manage the local power resources of the corresponding processing unit according to the allocated local power budget to execute tasks, and report the power management parameters when executing tasks to the global power management unit.
[0027] In another implementation of this disclosure, the global power management unit includes: a global storage unit for storing the global power budget and the power management parameters of each processing unit; and a global management unit for allocating a local power budget for each processing unit according to the global power budget and the power management parameters of each processing unit.
[0028] In another implementation of this disclosure, the local power management unit includes: a local storage unit for storing the local power budget and power management parameters of the corresponding processing unit; and a local management unit for determining a local power resource management strategy based on the local power budget and power management parameters of the processing unit.
[0029] In another implementation of this disclosure, the multi-processing unit system is a distributed processing system, and the processing unit is a processing unit within the distributed processing system.
[0030] In another implementation of this disclosure, the multi-processing unit system is a multi-core processing unit, and the processing unit is a core in the multi-core processing unit.
[0031] According to a third aspect of the present disclosure, a power management module for a multi-processor unit system is provided, including multiple local power management units and a global power management unit. Each local power management unit corresponds to a processing unit in the multi-processor unit system. The global power management unit is used to obtain the global power budget of the multi-processor unit system and allocate a local power budget for each processing unit according to the global power budget and the power management parameters reported by each processing unit. The local power management units are used to manage the local power resources of the corresponding processing unit according to the allocated local power budget to execute tasks and report the power management parameters when executing tasks to the global power management unit.
[0032] In the embodiments of this disclosure, the allocation of local power budgets for each computing unit takes into account the power management parameters of each computing unit, rather than distributing them equally. This makes the local power budgets more closely match the actual needs of each computing unit, thereby improving overall power utilization efficiency. Furthermore, the allocation of local power budgets for each computing unit also considers the constraints of the global power budget, preventing individual computing unit power efficiency optimization from resulting in low overall power efficiency, further improving the overall power utilization efficiency of the multi-computing-unit system. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings.
[0034] Figure 1 This is a schematic architecture diagram of a multi-processor system according to an embodiment of the present disclosure;
[0035] Figure 2 This is a flowchart of the steps of a power management method according to another embodiment of the present disclosure;
[0036] Figure 3A As another embodiment of this disclosure Figure 1 A detailed architecture diagram of a processor (CPU) for another example of a multi-processor unit system;
[0037] Figure 3B In accordance with this disclosure Figure 3A The example shows the processing logic block diagram of the global power management unit.
[0038] Figure 3C In accordance with this disclosure Figure 3A The example corresponds to the processing logic block diagram of the local power management unit;
[0039] Figure 4 This is a schematic structural diagram of a multi-processor system according to another embodiment of the present disclosure;
[0040] Figure 5 In accordance with this disclosure Figure 4 A schematic structural diagram of the global power management unit in an embodiment;
[0041] Figure 6 In accordance with this disclosure Figure 4 A schematic structural diagram of the local power management unit in an embodiment. Detailed Implementation
[0042] To enable those skilled in the art to better understand the technical solutions in the embodiments of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art should fall within the protection scope of this disclosure.
[0043] The specific implementation of the embodiments of this disclosure will be further described below with reference to the accompanying drawings.
[0044] Figure 1 This is a schematic architecture diagram of a multi-computation unit system according to an embodiment of the present disclosure. Figure 1 The multi-processor unit system 100 includes a global power management unit 110, a local power management unit 120, and a processing unit 130. Generally, the multi-processor unit system 100 is a data processing system integrating multiple processing units, including but not limited to multi-core processors and multi-processor systems. Multi-core processors include, but are not limited to, multi-core CPUs (Central Processing Units) and GPUs (Graphics Processing Units). When the multi-processor unit system 100 is a multi-core CPU, the processing unit 130 can act as a CPU core; when the multi-processor unit system 100 is a GPU, the processing unit 130 can be an SM (Streaming Multiprocessor) or SP (Streaming Processor). In addition to the processing units, the multi-processor unit system 100 may also include shared resources such as shared memory resources and / or shared instruction scheduling resources. Shared memory resources can be implemented as a level-one or multi-level cache composed of Static Random-Access Memory (SRAM) or Dynamic Random-Access Memory (DRAM), and shared instruction scheduling resources can be implemented as scheduling units.
[0045] It should also be understood that the global power management unit 110 refers to the software or hardware configuration for power management of the multi-computing unit system 100. The multi-computing unit system 100 is used to manage the power consumption of each local power management unit 120, the global power management unit 110 itself, and one or more computing units 130 included in the multi-computing unit system 100. The local power management unit 120 corresponds to one or more computing units 130 and is a software or hardware configuration for managing the power consumption of one or more computing units 130 and the local power management unit 120 itself. It should also be understood that either the global power management unit 110 or the local power management unit 120 can be implemented as a software configuration or a hardware configuration, and the embodiments of this disclosure do not limit the location of the above-mentioned software configuration or hardware configuration. In one example, the global power management unit 110 can be configured in hardware or software outside of each computing unit 130, or configured in computing units 130. In another example, the local power management unit 120 can be configured in hardware or software within or outside the computing units 130 it manages.
[0046] There can be multiple local power management units 120 and arithmetic units 130, for example, Figure 1The diagram illustrates arithmetic units 1-K and local power management units 1-K. Each local power management unit 120 and each arithmetic unit 130 can have a corresponding relationship, which can be implemented as a one-to-one correspondence, or one local power management unit 120 corresponding to multiple arithmetic units 130. In this embodiment, for ease of explanation, the local power management unit 120 and the arithmetic unit 130 are shown in a one-to-one correspondence. The multi-arithmetic unit system 100 may further include shared storage resources 140 and main memory 150, wherein the arithmetic unit 130 can obtain instructions and data for data processing from the main memory 150. The shared storage resource 140 can serve as a cache between the arithmetic unit 130 and the main memory 150, for example, a last-level cache (LLC).
[0047] Specifically, the local power management unit 120 performs local power management for the corresponding arithmetic unit 130. For example, the local power management unit 120 calculates the power usage based on its respective local power budget.
[0048] The global power management unit 110 performs global power management. For example, the global power management unit 110 can allocate power budgets to each computing unit. Specifically, when the global power management unit is configured inside a computing unit, its own power budget can be included in the local power budget of that computing unit; when the global power management unit is configured outside each computing unit, it can also allocate its own power budget. Similarly, when the local power management unit 120 is configured inside a computing unit, its power budget can be included in the local power budget of that computing unit; when the local power management unit 120 is configured outside each computing unit, the global power management unit 110 can also allocate its own power budget.
[0049] The local power management unit 120 and the global power management unit 110 can communicate with each other, and the global power management unit 110 can also communicate with the shared storage resource 140 to realize the above-mentioned global power management.
[0050] It should also be understood that Figure 1 The multiple processing units are merely exemplary architectures, and other architectural designs are also applicable to the embodiments disclosed herein.
[0051] Figure 2 This is a flowchart of the steps of a power management method according to another embodiment of the present disclosure. Figure 2 The power management method applied to multi-processor systems can be applied to Figure 1 Multi-computation unit systems, but not limited to Figure 1 A multi-computation unit system. Figure 2 The multi-computing unit system includes multiple local power management units and a global power management unit. Each local power management unit corresponds to one computing unit in the multi-computing unit system. Specifically, the method includes:
[0052] S210: Utilizes the global power management unit to obtain the global power budget of the multi-processor system.
[0053] S220: Using the global power management unit, the local power budget of each computing unit is allocated according to the global power budget and the power management parameters of each computing unit.
[0054] S230: Utilizes the local power management unit to manage the local power resources of the corresponding computing unit according to the allocated local power budget.
[0055] S240: Using the local power management unit, report the power management parameters of the computing unit to the global power management unit.
[0056] The following will combine Figure 1 A schematic architecture diagram of a multi-computing unit system Figure 2 The power consumption management method will be explained in general.
[0057] The global power management unit 110 obtains the global power budget of the multi-computing unit system 100. Based on the global power budget and the power management parameters of each computing unit 130, the global power management unit 100 allocates a local power budget to each computing unit 130. The local power management unit 120 manages the local power resources of the corresponding computing unit 130 according to the allocated local power budget and reports the power management parameters of the computing unit 130 to the global power management unit 110.
[0058] It should be understood that the global power budget is the overall power budget required for the multi-computing unit system 100 to perform data processing, and can cover the power consumption of each part or unit in the multi-computing unit system 100.
[0059] In one example, the local power management unit 120 is configured in the corresponding arithmetic unit 130. The local power budget refers to the power budget allocated to the corresponding arithmetic unit 130, while the global power budget includes the overall power budget allocated to all arithmetic units 130. The global power budget and the local power budget can be calculated using power values or in units of power counters corresponding to those power values. The power counter can be a device configured within the arithmetic unit itself, or such a device can be configured in software or hardware within the local power management unit 120 and the global power management unit 110.
[0060] In addition, local power consumption resources refer to the hardware or software configurations that consume power. Hardware configurations can be circuit components, etc., while software configurations include tasks such as instruction processing, process processing, and data processing based on CPUs or GPUs.
[0061] It should be understood that power management parameters refer to information or data related to power management, such as data related to power usage or performance parameters of the computing unit related to power management. The power management parameters of the computing unit 130 may include, but are not limited to, at least one of energy consumption parameters, performance parameters, and workload parameters.
[0062] The power consumption parameters of the arithmetic unit 130 refer to data or information characterizing the power usage of the arithmetic unit 130 when executing processing tasks, and may include at least one of power consumption value, power consumption status, and local budget utilization. The task load parameters of the arithmetic unit 130 refer to parameters characterizing the amount of data processed in the processing tasks executed by the arithmetic unit 130, and may include, but are not limited to, at least one of the number of instructions, the number of addresses, and the main memory access latency information of the processing task. The performance parameters of the arithmetic unit 130 refer to parameters characterizing the performance of the software or hardware configuration of the arithmetic unit 130, and may include, but are not limited to, at least one of clock frequency, number of cores, base frequency, and front-side bus frequency.
[0063] In the embodiments of this disclosure, the allocation of local power budgets for each computing unit takes into account the power management parameters of each computing unit, rather than distributing them equally. This makes the local power budgets more closely match the actual needs of each computing unit, thereby improving overall power utilization efficiency. Furthermore, the allocation of local power budgets for each computing unit also considers the constraints of the global power budget, preventing individual computing unit power efficiency optimization from resulting in low overall power efficiency, further improving the overall power utilization efficiency of the multi-computing-unit system.
[0064] The following will combine Figures 3A-3C The examples of the processing units and the processing logic of the power management method are described in detail. It should be understood that each processing logic can be understood as a sub-step included in the power management method, which can be implemented through software configuration or hardware configuration. Figure 3A As another embodiment of this disclosure Figure 1 A detailed block diagram of a processor (CPU) for another example of a multi-processor unit system.
[0065] In some embodiments, the CPU 1300 may include one or more CPU cores 130 (examples of arithmetic units) for processing instructions, the processing and execution of which can be controlled by a user (e.g., through an application) and / or the system platform. In some embodiments, each CPU core 130 may be used to process a specific instruction set. In some embodiments, the instruction set may support Complex Instruction Set Computing (CISC), Reduced Instruction Set Computing (RISC), or computation based on Very Long Instruction Words (VLIW). Different CPU cores 130 may each process different or the same instruction sets. In some embodiments, the CPU core 130 may also include other processing modules, such as a Digital Signal Processor (DSP). As an example, Figure 3A The diagram shows the processing units cores 1 to m, where m is a non-zero natural number.
[0066] In some embodiments, the cache memory 18 may be wholly or partially integrated into the CPU 1300. And depending on the architecture, Figure 3A Level 3 cache memories L1 to L3 are shown. L3 140 may be an internal cache memory located outside each processing unit core 101, while L1 181 and L2 181 are internal cache memories inside each processing unit core 101. L1 to L3 may include instruction-oriented instruction caches and data-oriented data caches. In some embodiments, the components in CPU 1300 may share at least a portion of the cache memory, and processing unit cores 1 to m may, for example, share the third-level cache memory L3. CPU 1300 may also include an external cache (not shown), and other cache structures may also serve as external caches for CPU 1300.
[0067] In some embodiments, the CPU 1300 may include a register file 139, which may include multiple registers for storing different types of data and / or instructions. These registers may be of different types. For example, the register file 139 may include integer registers, floating-point registers, status registers, instruction registers, and pointer registers. The registers in the register file 139 may be implemented using general-purpose registers, or a specific design may be adopted according to the actual needs of the CPU 1300.
[0068] CPU 1300 may include a Memory Management Unit (MMU) 133 for translating virtual addresses to physical addresses. The MMU 133 caches a portion of page table entries and can also retrieve uncached entries from memory. Each CPU core 130 may have one or more MMUs 133. The MMUs 120 in different CPU cores 130 can be synchronized with those in other processing units or processing unit cores, allowing each processing unit or processing unit core to share a unified virtual memory system.
[0069] The CPU 1300 is used to execute instruction sequences (i.e., programs). The process of the CPU 1300 executing each instruction includes: fetching the instruction from the memory where the instructions are stored, decoding the fetched instruction, executing the decoded instruction, and saving the instruction execution result. This process is repeated until all instructions in the instruction sequence have been executed or a halt instruction is encountered.
[0070] To implement the above process, the CPU 1300 may include an instruction fetch unit 137, an instruction decode unit 135, an instruction dispatch unit (not shown), an instruction execution unit 131, and an instruction de-initialization unit (not shown), etc.
[0071] The instruction fetch unit 137 serves as the boot engine of the CPU 1300, used to move instructions from main memory 150 to the instruction register (which can be...). Figure 3A The instruction is stored in one of the registers in the register file 26 shown, and the next fetch address is received or calculated according to the fetch algorithm, which may include, for example, incrementing or decrementing the address based on the instruction length.
[0072] After the instruction is fetched, the CPU 1300 enters the instruction decoding stage. The instruction decoding unit 135 decodes the fetched instruction according to a predetermined instruction format to obtain the operand fetch information required by the fetched instruction, thereby preparing for the operation of the instruction execution unit 131. Operand fetch information may include pointers to immediate values, registers, or other software / hardware that can provide source operands.
[0073] The instruction dispatch unit, typically located within a high-performance CPU 1300, lies between the instruction decoding unit 135 and the instruction execution unit 131. It is used for instruction scheduling and control, efficiently allocating instructions to different instruction execution units 131, thus enabling parallel operations of multiple instructions. After an instruction is fetched, decoded, and scheduled to the corresponding instruction execution unit 131, that unit begins executing the instruction, performing the operation indicated by the instruction and implementing the corresponding function.
[0074] The instruction write-back unit (or instruction write-back unit) is mainly responsible for writing the execution results generated by the instruction execution unit 131 back to the corresponding storage location (e.g., a register inside the CPU 1300) so that subsequent instructions can quickly obtain the corresponding execution results from that storage location.
[0075] Different instruction execution units 131 can be set up in the CPU 1300 for different types of instructions. The instruction execution unit 131 can be an arithmetic unit (e.g., containing an arithmetic logic unit, vector operation unit, etc., used to perform operations based on operands and output the results), a memory execution unit (e.g., used to access memory according to instructions to read data from memory or write specified data to memory), or a coprocessing unit, etc. In the CPU 1300, each instruction execution unit 131 can run in parallel and output corresponding execution results.
[0076] When executing a certain type of instruction (such as a memory access instruction), the instruction execution unit 131 needs to access the main memory 150 to obtain information stored in the main memory 150 or to provide data that needs to be written to the main memory 150.
[0077] It should also be understood that both the global power management unit 110 and the local power management unit 120 can be configured externally or internally to the CPU core. (See reference...) Figure 3B The global power management unit 110 may include a parameter acquisition logic unit 111, a budget acquisition logic unit 112, a budget allocation logic unit 113, and a distribution logic unit 114. The parameter acquisition logic unit 111 acquires power management parameters from the local power management unit 120. At least some of these power management parameters can be obtained by the local power management unit 120 monitoring the managed CPU cores 130, and some can also be obtained by monitoring shared resources such as L3 140. The budget acquisition logic unit 112 acquires the global power budget, which can be a preset value or obtained by processing the power management parameters acquired by the parameter acquisition logic unit 111. The budget allocation logic unit 113 can allocate the local power budget of each CPU core 130 according to the global power budget obtained by the budget acquisition logic unit 112 and the power management parameters obtained by the local power management unit 120, and send the allocation result to the local power management unit 120 via the distribution logic unit 114. The local power management unit 120 performs power management on the CPU core 130 based on the local power budget as the allocation result.
[0078] Reference Figure 3CThe local power management unit 120 may include a data reporting logic unit 121, a data acquisition logic unit 122, a data monitoring logic unit 123, a policy analysis logic unit 124, and a policy execution logic unit 125. The data monitoring logic unit 123 monitors the real-time power management parameters of the CPU core 130 and sends these parameters to the policy analysis logic unit 124. The real-time power management parameters include at least one of the following: CPU core 130's energy consumption parameters, performance parameters, and workload parameters.
[0079] The data reporting logic unit 121 is used to report the monitored real-time power management parameters of the CPU core 130 to the global power management unit 110, for example, to the parameter acquisition logic unit 111 of the global power management unit 110. It should be understood that the data reporting logic unit 121 can report real-time power management parameters directly or indirectly according to the power management period set for the global power management unit. Direct reporting means forwarding the monitored real-time power management parameters directly to the global power management unit 110 without any processing. Indirect reporting means processing the monitored real-time power management parameters, such as statistical processing or feature extraction, before reporting them to the global power management unit 110.
[0080] The data acquisition logic unit 122 is used at least to acquire the allocated local power budget from the global power management unit 110, for example, from the distribution logic unit 114 of the global power management unit 110; it can also be used to acquire configuration information from sources such as user interfaces, including but not limited to the correspondence between the power status of resource management tasks and power management parameters. The data acquisition logic unit 122 can also be used to send the local power budget, configuration information, etc., to the policy analysis logic unit 124.
[0081] The strategy analysis logic unit 124 is used to determine the power consumption status of the resource management task based on the local power consumption budget, configuration information and power consumption management parameters, and send the power consumption status of the resource management task to the strategy execution logic unit 125.
[0082] The strategy execution logic unit 125 is used to execute the resource management task in this power consumption state and generate control instructions, which are then sent to the arithmetic unit 130. It should be understood that after the arithmetic unit 130 executes the task according to the control instructions, the power management parameters of the CPU core 130 monitored by the data monitoring logic unit 123 will change. In one example, the data monitoring logic unit 123 monitors the power management parameters of the CPU core 130 for a monitoring period shorter than the power management period. Preferably, the power management period is an integer multiple of the monitoring period.
[0083] As an example, in accordance with Figure 3BIn the global power management unit 110 that performs configuration, step S210 can be executed by the budget acquisition logic unit 112, and step S220 can be executed by the budget allocation logic unit 113. Figure 3C The configured local power management unit 120 can be executed in step S230 through the data acquisition logic unit 122, the policy analysis logic unit 124, and the policy execution logic unit 125, and in step S240 through the data monitoring logic unit 123 and the data reporting logic unit 121.
[0084] However, it should be understood that Figure 3B and Figure 3C The configurations of the global power management unit 110 and the local power management unit 120 shown are merely exemplary descriptions, and the configurations of the global power management unit 110 and the local power management unit 120 in the embodiments of this disclosure are not limited to the above examples.
[0085] The following will Figure 2 The various possible implementations are described and illustrated in detail. It should be understood that the following examples can be applied to the multi-operation unit system of the above embodiments, but are not limited to the multi-operation unit system of the above embodiments. In other words, the global power management unit can be the global power management unit 110 in the above embodiments; the local power management unit can be the local power management unit 120 in the above embodiments.
[0086] Specifically, step S210 can be implemented as follows: determining the global power budget based on the previous power management parameters of each computing unit and the previous global power budget. The previous power management parameters are power management parameters obtained by managing the resources of each computing unit using the previous global power budget, that is, the result of power management based on the previous local power budget. In other examples, the determination of the global power budget may be based not only on the previous power management parameters of each computing unit, but also on the power management parameters after management using the current global power budget. Previous data reflects the results of power management based on the previously allocated local power budget, and has high power management reference value for the current local power budget allocation. Therefore, it helps to determine a global power budget that better meets the requirements, which is beneficial for achieving dynamic power management.
[0087] In some examples, step S210 can be implemented as follows: using a first machine learning model, a global power budget is determined based on the previous power management parameters and the previous global power budget of each computing unit. The first machine learning model can improve the efficiency of determining the global power budget and can learn the inherent relationship between previous data and current data, thereby improving the accuracy of the global power budget. It should be understood that in this example, the global power budget determined in this step is the current global power budget, and the current local power budget allocation for each computing unit can be based on the current global power budget. It should be understood that each local power management unit can perform power management on the corresponding computing unit according to the current local power budget allocation to obtain the current power management parameters.
[0088] The management objective of the global power budget can be that the predicted power utilization rate of the global power budget is not lower than the actual power utilization rate of the previous global power budget, thereby improving the power utilization rate of the global power budget to a certain extent and contributing to the overall energy efficiency of the multi-computing unit system. The management objective of the global power budget can also be to allow the overall power consumption to exceed a preset threshold (or percentage) of the global power budget within a preset time period. If the overall power consumption based on the current global management budget exceeds the preset threshold, after determining the initial value of the next global power budget, the initial value can be reduced to obtain the actual value of the next global power budget. For example, the initial value can be multiplied by a coefficient less than 1 to obtain the actual value. If the overall power consumption based on the current global management budget does not exceed the preset threshold, the initial value is directly used as the actual value of the next global power budget.
[0089] In addition, the power utilization rate or actual power utilization value of the global power budget can be calculated, and the first machine learning model can be further trained based on the power utilization rate or actual power utilization value and the corresponding power management parameters as samples to update the first machine learning model.
[0090] It should be understood that the first machine learning model includes, but is not limited to, feedforward neural networks, convolutional neural networks, recurrent neural networks, etc., and its training methods include, but are not limited to, supervised learning, unsupervised learning, and reinforcement learning. The global power management unit may include an interface for deploying and updating the first machine model, which can communicate with main memory.
[0091] Alternatively, step S210 can also be implemented as follows: using a first rule, based on the previous power management parameters of each computing unit and the previous global power budget, a global power budget is determined. The first rule can indicate the adjustment rules for the global power budget according to the previous power management parameters, so that the predicted power utilization rate of the current global power budget is not lower than the actual power utilization rate of the previous global power budget. On the one hand, the first rule can correspond to the aforementioned global power budget management objective; in other words, the first rule can be set so that the determined global power budget can achieve the predetermined management objective. On the other hand, the first rule can be implemented in software, improving the efficiency of user configuration. Especially when power management has different strategies due to different scenarios or different tasks performed by multi-computing unit systems, the set first rule improves the flexibility of power management.
[0092] In other examples, step S210 can also be implemented as follows: based on the power management parameters of each computing unit in the first power management period and the global power budget of the multi-computing unit system in the first power management period, determine the global power budget of the multi-computing unit system in the second power management period, thereby improving the real-time performance and reliability of global power management. Regarding the aforementioned first power management period and second power management period, the second power management period can be after the first power management period. Furthermore, the first power management period and the second power management period can form a continuous power management period or a non-continuous power management period.
[0093] Furthermore, the duration of the first power management period and the second power management period can each be set arbitrarily. For example, the durations of the first power management period and the second power management period can be set independently. Preferably, the first power management period and the second power management period can be two consecutive periods of equal duration, which helps to improve the reliability of power management.
[0094] Furthermore, the first power management period and the second power management period can indicate periods for executing different tasks, or they can indicate different periods for executing the same task. These tasks can be power resource management tasks, including clock frequency management and operating voltage management. Preferably, the first power management period and the second power management period can be configured to execute two consecutive power resource management tasks.
[0095] Figure 2The scheme may further include: when the power management parameters of each computing unit meet predetermined conditions during the first power management period, the global power management unit executes a first process. It should be understood that the first process can be power management between different power management periods by the global power management unit. Different power management periods can be continuous or non-continuous power management periods. The first process realizes power management between different power management periods according to predetermined conditions, further improving dynamic management efficiency and reliability.
[0096] Correspondingly, in another implementation of step S210, the global power management unit can also determine the global power budget for the second power management period based on the result of the first processing, wherein the first processing indicates the management of the global power budget between different power management periods.
[0097] The following are examples of the first processing corresponding to different predetermined conditions:
[0098] For example, a predetermined condition could be that the sum of the power consumption values of each computing unit during the first power management period exceeds the global power budget for the first power management period. Accordingly, the global power management unit can reduce the local power budget of each computing unit during the first power management period, ensuring that the sum of the power consumption values of each computing unit during the first power management period does not exceed the global power budget for the first power management period. This allows the global power management unit to allocate the budget for the second power management period without considering the previous power consumption values. Furthermore, the global power management unit can also accumulate any unused remaining budget from the global power budget of the first power management period into the budget allocation for the second power management period, thereby achieving better global dynamic power management.
[0099] Furthermore, the first process, executed under predetermined conditions, is used to ensure that the operating power of each computing unit during the second power management period conforms to the local power budget of the second power management period. The first process can also be used to ensure that the operating power of each computing unit during the second power management period conforms to the product of the local power budget of the second power management period and a predetermined ratio, wherein the predetermined ratio is less than 1.
[0100] For example, the predetermined condition could be: the sum of the power consumption values of each computing unit during the first power management period does not exceed the global power budget for the first power management period. Correspondingly, the global power management unit can determine the difference between the sum of the power consumption values of the computing units during the first power management period and the global power budget, and use this difference to reduce the global power budget for the second power management period. This better achieves global dynamic power management and avoids device malfunctions caused by insufficient real-time global power budget. More specifically, the global power management unit can determine the difference between the sum of the running power reported by each computing unit and the local power budget for the second power management period, and then, when obtaining the global power budget for the computing units in the third power management period, use this difference to subtract the global power budget. The third power management period is after the second power management period.
[0101] Furthermore, when the global power budget is determined based on the first power management period and the second power management period, step S220 can be implemented as follows: according to the global power budget of the multi-operation unit system in the second power management period and the power management parameters of each operation unit in the first power management period, allocate the local power budget of each operation unit in the second power management period. Thus, based on the first power management period and the second power management period, not only is the global power budget managed, but also the local power budget is allocated, improving the uniformity and consistency of the management periods in different processes of dynamic power management.
[0102] When the first power management period and the second power management period form a continuous power management period, power management can be performed in more real-time. In other words, each moment falls within a power management period, and power management parameters at any given time can be used to determine the global power budget and allocate the local power budget. It should be understood that a continuous power management period means that the end of the first power management period is the start of the second power management period.
[0103] When the first and second power management periods form a non-contiguous power management period, only the power management parameters within these two periods can be used to determine the global power budget. In other words, power budget management can be performed without considering the time outside of these two periods. It should be understood that a non-contiguous power management period refers to a period with a duration greater than zero between the end of the first power management period and the start of the second. Therefore, the power management parameters within the first and second power management periods can be considered as valuable sampled data among all power management parameters. Using this sampled data to determine the global power budget improves data processing efficiency.
[0104] For step S220, when allocating the local power budget for each computing unit, it can be specifically implemented as follows: the second rule can be used to allocate the local power budget for each computing unit during the second power management period. Several examples of the second rule will be illustrated below:
[0105] For example, the total power budget allocated to each computing unit in the global power budget can be determined first. Then, based on the power management parameters of each computing unit in the first power management period, the power consumption ratio of each computing unit in the second power management period can be determined. Finally, based on the power consumption ratio of each computing unit, the local power budget of each computing unit in the second power management period can be determined from the total power budget. To determine the power consumption ratio of each computing unit in the second power management period, the target task volume of the processing tasks executed by each computing unit in the second power management period, as indicated by the task volume parameters, can be obtained. Then, the predicted power consumption value required for each computing unit to complete the target task volume, based on the hardware and software performance indicated by the performance parameters, can be calculated. The power consumption ratio can be the ratio between the predicted power consumption value of each computing unit and the sum of the predicted power consumption values. It should be understood that power management parameters can indicate actual power consumption usage. For each computing unit, the more actual power consumption is used in the first power management period, the more local power budget is allocated in the second power management period.
[0106] For example, when determining the total power budget allocated to each computing unit in the global power budget, the power budget of the global power management unit can be excluded from the global power budget to obtain the total power budget for each computing unit. It should also be understood that the power budget of the global power management unit can be set to a predetermined value. Alternatively, statistical values (e.g., average values) of the power consumption of the global power management unit can be collected during the budget statistics period. Because the data processing volume and power consumption of the global power management unit itself are more stable, its own power budget is included in the allocation of the global power budget, thus efficiently and accurately estimating its own power budget. It should be understood that the duration of the aforementioned budget statistics period can be longer than the power management period. For example, when the first power management period and the second power management period are equal, the budget statistics period can be set to an integer multiple of the power management period.
[0107] For example, the previous power management parameters of the global power management unit (GW) can be combined with the power management parameters of each computing unit during the previous power management period to allocate the power budget. For instance, the previous power consumption ratio between the GW and each computing unit can be calculated, and the current global power budget can be allocated based on this ratio. This achieves a unified dynamic allocation of the GW's own power budget and the local power budgets of each computing unit. It should be understood that when the GW allocates the power budget across power management periods, the current global power budget may include unused remaining budget from the previous global power budget, or it may exclude any excess consumption based on the previous global power budget.
[0108] Alternatively, step S220 can be specifically implemented as follows: power budget allocation is performed using a second machine learning model, which can be obtained by pre-training a neural network. The input of the second machine learning model can be determined based on the global power budget of the multi-operation unit system during the second power management period and the power management parameters of each operation unit during the first power management period. Then, the local power budget of each operation unit during the second power management period is output from the second machine learning model.
[0109] The input to the second machine learning model can be determined in several ways: First, by directly extracting features from both the global power budget of the multi-operation unit system during the second power management period and the power management parameters of each operation unit during the first power management period. Second, by first determining the total power budget of each operation unit in the global power budget of the second power management period, and then extracting features based on the total power budget and the power management parameters of each operation unit during the first power management period to obtain the input to the second machine learning model. Third, by determining the input to the second machine learning model solely based on the power management parameters of each operation unit during the first power management period, outputting the power consumption ratio of each operation unit during the first power management period, and then allocating the local power budget of each operation unit during the second power management period from the global power budget of the first power management period or the total power budget of each operation unit during the first power management period according to the power consumption ratio.
[0110] In one specific example, the global power budget of the multi-computing unit system during the second power management period is the same as the global power budget of the multi-computing unit system during the first power management period. In other words, given the global power budget of the multi-computing unit system, the local power budget allocation for each computing unit is performed. In another specific example, the total power budget of each computing unit during the first power management period is the same as the total power budget during the second power management period. In other words, with the total power budget of each computing unit remaining unchanged, the local power budget allocation for each computing unit is performed. In these two specific examples, the second rule is simplified, or the second machine learning model is made more efficient, thus improving the efficiency of local power budget allocation.
[0111] In addition, step S220 can be implemented as follows: determining the total power budget of each computing unit and the power budget of the global power management unit in the global power budget, and allocating the local power budget of each computing unit according to the total power budget of each computing unit and the power management parameters of each computing unit, thereby improving the allocation efficiency of the local power budget of each computing unit.
[0112] In one example, a second rule can be used to allocate the local power budget of each computing unit based on the total power budget of each computing unit and the power management parameters of each computing unit. The second rule can be implemented in software, which improves the efficiency of local power budget allocation. In particular, when power management has different power allocation strategies due to different scenarios or different tasks performed by multi-computing unit systems, using the set second rule helps to improve the flexibility of power management.
[0113] In another example, the total power budget of each computing unit and the power management parameters of each computing unit can be used as inputs to a second machine learning model to obtain the local power budget of each computing unit. The second machine learning model can improve the allocation efficiency of the local power budget and learn the intrinsic relationship between the power management parameters of each computing unit and the local power budget, thereby improving the allocation accuracy of the local power budget.
[0114] Specifically, the current power budget reference parameters of each computing unit can be determined based on the previous power management parameters of each computing unit, and the current power budget reference parameters and the current total power budget of each computing unit can be input into the second machine learning model to obtain the local power budget of each computing unit.
[0115] For example, the second machine learning model can be a model pre-trained via supervised or unsupervised training. The direct or indirect inputs of the second machine model can include the current power budget reference parameters of each computing unit. These current power budget reference parameters can be data determined based on previous power management parameters, such as power feature data of each computing unit obtained by feature extraction of previous power management parameters. The output of the second machine model can be the local power budget of each computing unit, or data indicating the local power budget of each computing unit. It should be understood that the power budget reference parameters can be determined for each computing unit; in other words, the current power budget reference parameters of that computing unit can be determined based on its previous power management parameters. Alternatively, multiple computing units can be grouped, and the power budget reference parameters can be determined for all computing units within each group; in other words, the current overall power budget reference parameters of each computing unit within that group can be determined based on its previous overall power management parameters. It should be understood that the computing units can be grouped based on the correlation between the processing tasks performed by each computing unit. Furthermore, grouping processing can be performed in the global power management unit. For example, the global power management unit performs clustering processing based on the power management parameters of each computing unit, identifies each computing unit in the cluster as a related computing unit, and then determines the current power budget reference parameters of each computing unit in the cluster based on the previous power management parameters of each computing unit in the cluster as a whole. This grouping method can improve the processing efficiency of the second machine learning model when there are a large number of computing units, while ensuring the accuracy of power budget allocation. In particular, when the computing units are SMs or SPs in the GPU, it improves the energy efficiency of the GPU.
[0116] For example, the second machine learning model can be a reinforcement learning model. The current power budget reference parameters for each computing unit can serve as the state and reward of the reinforcement learning model. Based on the state and reward, the reinforcement learning model can predict the current execution strategy, for example, determining the current local power budget for each computing unit. The reward of the reinforcement learning model can indicate at least one of the following: budget utilization rate, budget allocation error, and budget allocation error rate.
[0117] It should be understood that the second machine learning model includes, but is not limited to, feedforward neural networks, convolutional neural networks, recurrent neural networks, etc., and its training methods include, but are not limited to, supervised learning, unsupervised learning, and reinforcement learning. The global power management unit may include an interface for deploying and updating the second machine model, which can communicate with main memory.
[0118] In a specific example, power budget reference parameters for each computing unit in the second power management period can be calculated based on the power management parameters of each computing unit in the first power management period. These parameters serve as the state and reward for the reinforcement learning model corresponding to the first power management period. For example, the state and reward for the reinforcement learning model corresponding to the first power management period can be calculated at the end of the first power management period. Alternatively, the power management parameters of each computing unit in the first power management period can be obtained at the end of the first power management period, and the state and reward for the first power management period can be calculated at the beginning of the second power management period. It should be understood that the definition and determination method of the first and second power management periods can be as described above, and will not be repeated here. It should also be understood that when calculating the state and reward for the first power management period, calculations can be performed based on the power management parameters of each computing unit and other data. For example, other data may include the local power budget and / or total power budget of each computing unit in the first power management period. In this way, the calculated state and reward can more accurately reflect the feedback results of the aforementioned local power budget allocation.
[0119] Furthermore, step S220 can also be implemented as follows: allocating the local power budget for each computing unit based on the global power budget, the power management parameters of each computing unit, and other data. Other data includes the power management parameters of shared resources. Shared resources may include shared storage resources such as the last-level cache mentioned above; considering the power management parameters of shared resources can further improve the accuracy of local power budget allocation.
[0120] Furthermore, step S230 can be implemented as follows: determining and executing a local power resource management strategy based on the allocated local power budget and the power management parameters of the local power management unit. It should be understood that the local power resource management strategy includes at least one of Dynamic Voltage and Frequency Scaling (DVFS), PowerGating, Clock Gating, setting the clock frequency, and adjusting the order of multiple tasks to be executed. It should be understood that Dynamic Voltage and Frequency Scaling is a method that can dynamically change the voltage and frequency of the arithmetic unit during program execution, effectively improving the energy efficiency of the arithmetic unit. PowerGating can be used to turn off the power supply to currently inactive parts of the circuitry in the arithmetic unit to save power. Clock Gating is an effective means of reducing the power consumption of the arithmetic unit; for example, it can manage the dynamic power consumption caused by register flipping corresponding to the arithmetic unit.
[0121] In one example, step S230 can be specifically implemented as follows: using a third machine learning model, a local power resource management strategy is determined based on the allocated local power budget and the power management parameters of the local power management unit. The third machine learning model can be a pre-trained neural network model. The allocated local power budget and the power management parameters of the local power management unit can be used as direct or indirect inputs to the third machine learning model, and the local power resource management strategy can be used as its direct or indirect output. The third machine learning model can improve the analysis efficiency of the local power resource management strategy and learn the intrinsic relationship between the power management parameters of each computing unit and the local power budget and local power resource management strategy, thereby improving the reliability and efficiency of local power management. More specifically, the third machine learning model can learn the intrinsic relationship between at least one of dynamic voltage frequency scaling, power selection, clock selection, setting the clock frequency, and adjusting the order of multiple tasks to be executed, and the power management parameters and local power budget, achieving accurate power management of the computing units.
[0122] It should be understood that the third machine learning model includes, but is not limited to, feedforward neural networks, convolutional neural networks, recurrent neural networks, etc., and its training methods include, but are not limited to, supervised learning, unsupervised learning, and reinforcement learning. The local power management unit may include an interface for deploying and updating the third machine model, which may be connected to main memory or to the global power management unit, through which the third machine learning model is deployed or updated.
[0123] In another example, step S230 can be specifically implemented as follows: using a third rule, a local power resource management strategy is determined based on the allocated local power budget and the power management parameters of the local power management unit. The third rule can be implemented in software or in hardware, such as digital circuits, improving the analysis efficiency of the local power resource management strategy. Especially since different power management scenarios or tasks performed by the computing unit result in different power resource management strategies, utilizing the established third rule helps to improve the flexibility of the power resource management strategy.
[0124] In one example, the third rule indicates the correspondence between power state and clock frequency and / or voltage in the local power management strategy. For instance, the power state can be determined based on the power management parameters of the local power management unit, then the third rule can be used to determine the voltage and / or clock frequency with that power state, and then power management instructions for the local power management strategy can be generated based on that voltage and / or clock frequency.
[0125] In another example, the third rule indicates the relationship between a power adjustment strategy and the clock frequency and / or voltage, whereby the power adjustment strategy indicates increasing or decreasing power in predetermined steps. The power management parameters of the local power management unit can be determined periodically, and the current voltage and / or clock frequency can be determined. Then, the third rule is periodically consulted to determine the power adjustment strategy corresponding to the current voltage and / or clock frequency, and a power management instruction indicating the power adjustment strategy is generated based on that voltage and / or clock frequency. Furthermore, the third rule may also include the relationship between a predetermined step size and the clock frequency and / or voltage; accordingly, the aforementioned power management instruction can also be generated considering the current predetermined step size.
[0126] The following will be combined again. Figure 3C Describe the possible implementations of the local power resource management strategy for the arithmetic unit, and the implementations of the management strategy for each arithmetic unit may be the same or different.
[0127] The strategy analysis logic unit 124 can generate control instructions based on at least one of dynamic voltage frequency scaling, power selection, clock selection, setting clock frequency, and adjusting the order of multiple tasks to be executed. These control instructions correspond to the control circuit that executes the strategy. This control circuit can be a portion of the digital circuitry in the CPU core 130, or a digital or analog circuitry corresponding to the CPU core 130. For example, for dynamic voltage frequency scaling, the control instruction instructs the control circuit to perform dynamic voltage frequency scaling on the arithmetic unit; for power selection or clock selection, the control instruction instructs to turn off currently inactive circuitry, or to adjust the toggle frequency of registers in the arithmetic unit; for setting the clock frequency, the control instruction instructs the control circuit to adjust the clock frequency; for adjusting the order of multiple tasks to be executed, the control instruction instructs the control circuit to adjust the order of the multiple tasks to be executed. It should be understood that the above control circuitry can be implemented through software configuration or hardware configuration, and this embodiment does not limit this implementation.
[0128] Figure 4 This is a schematic structural diagram of a multi-processor unit system according to another embodiment of the present disclosure. For example, the multi-processor unit system of this embodiment can be a distributed processing system, where the processing unit is a processing unit within the distributed processing system. Another example is a multi-core processing unit, where the processing unit is a core within the multi-core processing unit. Specifically, a multi-processor system includes, but is not limited to, multiple CPU systems, multiple GPU systems, a heterogeneous system consisting of at least one CPU and at least one GPU, and a system consisting of at least one CPU or at least one GPU and other processing units. Figure 4The multi-processor unit system includes multiple processing units 430, multiple local power management units 420, and a global power management unit 410. Each local power management unit 420 corresponds to one processing unit 430.
[0129] In addition, the global power management unit 410 is used to obtain the global power budget of the multi-processor system, and allocate the local power budget of each processing unit 430 according to the global power budget and the power management parameters reported by each processing unit 430.
[0130] In addition, the local power management unit 420 manages the local power resources of the corresponding processing unit 430 according to the allocated local power budget to perform tasks, and reports the power management parameters when performing tasks to the global power management unit 410.
[0131] Generally, there are multiple local power management units 420 and processing units 430, for example, Figure 4 The processing unit 1-K and the local power management unit 1-K are shown. Each local power management unit 420 and each processing unit 430 can have a one-to-one correspondence. The global power management unit 410 and the local power management unit 420 can be formed into a power management module to manage the power consumption of multiple processing units 430 and the power management module itself.
[0132] In the embodiments of this disclosure, the allocation of local power budgets for each computing unit takes into account the power management parameters of each computing unit, rather than distributing them equally. This makes the local power budgets more closely match the actual needs of each computing unit, thereby improving overall power utilization efficiency. Furthermore, the allocation of local power budgets for each computing unit also considers the constraints of the global power budget, preventing individual computing unit power efficiency optimization from resulting in low overall power efficiency, further improving the overall power utilization efficiency of the multi-computing-unit system.
[0133] Figure 5 In accordance with this disclosure Figure 4 A schematic structural diagram of the global power management unit in an embodiment. Figure 5 The global power management unit 410 includes a global storage unit 411 and a global management unit 412.
[0134] The global storage unit 411 is used to store the global power consumption budget and the power consumption management parameters of each processing unit. The global management unit 412 is used to allocate the local power consumption budget of each processing unit according to the global power consumption budget and the power consumption management parameters of each processing unit.
[0135] Figure 6 In accordance with this disclosure Figure 4 A schematic structural diagram of the local power management unit in an embodiment. Figure 6The local power management unit 420 includes a local storage unit 421 and a local management unit 422.
[0136] The local storage unit 421 is used to store the local power consumption budget and power consumption management parameters of the corresponding processing unit. The local management unit 422 is used to determine the local power consumption resource management strategy based on the local power consumption budget and power consumption management parameters of the processing unit.
[0137] It should be understood that the various units and their operations in this embodiment can be referred to the description and explanation of the above embodiments. The processing unit can perform similar or the same operations and functions as the arithmetic unit, which will not be repeated here.
[0138] The following still combines Figure 4 The power management module of a multi-processor unit system is described below. The power management module includes multiple local power management units 420 and a global power management unit 410. Each local power management unit 420 corresponds to one processing unit 430. The global power management unit 410 is used to obtain the global power budget of the multi-processor unit system and allocate local power budgets for each processing unit 430 based on the global power budget and the power management parameters reported by each processing unit 430. Furthermore, the local power management units 420 are used to manage the local power resources of the corresponding processing unit 430 to execute tasks according to the allocated local power budget, and report the power management parameters during task execution to the global power management unit 410.
[0139] In the embodiments of this disclosure, the allocation of local power budgets for each computing unit takes into account the power management parameters of each computing unit, rather than distributing them equally. This makes the local power budgets more closely match the actual needs of each computing unit, thereby improving overall power utilization efficiency. Furthermore, the allocation of local power budgets for each computing unit also considers the constraints of the global power budget, preventing individual computing unit power efficiency optimization from resulting in low overall power efficiency, further improving the overall power utilization efficiency of the multi-computing-unit system.
[0140] It should be understood that the various units and their operations in this embodiment can be referred to the description and explanation of the above embodiments. The processing unit can perform similar or the same operations and functions as the arithmetic unit, which will not be repeated here.
[0141] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this disclosure can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this disclosure.
[0142] The methods described above according to embodiments of this disclosure can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded over a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.
[0143] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments disclosed herein.
[0144] The above embodiments are only used to illustrate the embodiments of this disclosure, and are not intended to limit the embodiments of this disclosure. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this disclosure. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this disclosure, and the patent protection scope of the embodiments of this disclosure should be defined by the claims.
Claims
1. A power management method for a multi-processor unit system, wherein, The multi-computing unit system includes multiple local power management units and a global power management unit. The system also includes a cache and main memory. Each local power management unit corresponds to one of the computing units in the system and is configured within that unit. The computing unit retrieves instructions and data for data processing from the main memory. The cache is located between the computing unit and the main memory. The global power management unit communicates with the cache. The method includes: Using the global power management unit, the global power budget is determined based on the previous power management parameters of each computing unit and the previous global power budget; Using the global power management unit, the local power budget of each computing unit is allocated according to the global power budget and the power management parameters of each computing unit. The allocation of the local power budget of each computing unit according to the global power budget and the power management parameters of each computing unit includes: allocating the local power budget of each computing unit according to the global power budget, the power management parameters of each computing unit and the power management parameters of the cache. The local power management unit manages the local power resources of the corresponding computing unit according to the allocated local power budget. The local power management unit reports the power management parameters of the computing unit to the global power management unit.
2. The method according to claim 1, wherein, The power management parameters include at least one of energy consumption parameters, performance parameters, and workload parameters; wherein, the energy consumption parameters include at least one of power consumption value, power consumption status, and local budget utilization rate.
3. The method according to claim 1, wherein, The determination of the global power budget based on the previous power management parameters of each computing unit and the previous global power budget includes: Using a first machine learning model or a first rule, the global power budget is determined based on the previous power management parameters of each computing unit and the previous global power budget, so that the predicted power utilization of the global power budget is not lower than the actual power utilization of the previous global power budget.
4. The method according to claim 1, wherein, The determination of the global power budget based on the previous power management parameters of each computing unit and the previous global power budget includes: Based on the power management parameters of each computing unit during the first power management period and the global power budget of the multi-computing unit system during the first power management period, the global power budget of the multi-computing unit system during the second power management period is determined, wherein the second power management period follows the first power management period, and the first power management period and the second power management period form a continuous power management period; and The step of allocating a local power budget for each computing unit based on the global power budget and the power management parameters of each computing unit includes: Based on the global power budget of the multi-computing unit system during the second power management period and the power management parameters of each computing unit during the first power management period, the local power budget of each computing unit during the second power management period is allocated.
5. The method according to claim 1, wherein, The step of allocating a local power budget for each computing unit based on the global power budget and the power management parameters of each computing unit includes: Determine the total power consumption budget for each computing unit in the global power consumption budget; Based on the total power consumption budget of each computing unit and the power consumption management parameters of each computing unit, the local power consumption budget of each computing unit is allocated.
6. The method according to claim 5, wherein, The step of allocating a local power budget for each computing unit based on the total power budget of each computing unit and the power management parameters of each computing unit includes: Using the second rule, the local power budget of each computing unit is allocated according to the total power budget of each computing unit and the power management parameters of each computing unit; Alternatively, the local power budget of each computing unit can be obtained by using the total power budget of each computing unit and the power management parameters of each computing unit as input to the second machine learning model; wherein, obtaining the local power budget of each computing unit by using the total power budget of each computing unit and the power management parameters of each computing unit as input to the second machine learning model includes: determining the current power budget reference parameters of each computing unit according to the previous power management parameters of each computing unit; and inputting the current power budget reference parameters of each computing unit and the current total power budget of each computing unit into the second machine learning model to obtain the local power budget of each computing unit.
7. The method according to claim 1, wherein, The management of local power resources of the corresponding computing units according to the allocated local power budget includes: Based on the allocated local power consumption budget and the power consumption management parameters of the local power consumption management unit, determine and execute the local power consumption resource management strategy; The local power resource management strategy includes at least one of dynamic voltage frequency scaling, power gating, clock gating, setting clock frequency, and adjusting the order of multiple tasks to be executed.
8. The method according to claim 4, wherein, The global power management unit performs the first process when the power management parameters of each arithmetic unit meet predetermined conditions during the first power management period. The power management parameters include power consumption values, and the predetermined conditions include: the power consumption values of each computing unit during the first power management period and the global power budget exceeding the first power management period. The first process includes one of the following: Reduce the power consumption of each computing unit during the first power management period so that the sum of the power consumption of each computing unit during the first power management period does not exceed the global power budget of the first power management period. The sum of the power consumption values of the computing unit during the first power management period and the amount exceeding the global power budget are determined, and the amount is used to offset the global power budget during the second power management period.
9. The method according to claim 1, wherein, The global power management unit is located within one of the multiple computing units, or outside the multiple computing units.
10. A multi-processing unit system, comprising multiple processing units, multiple local power management units, and a global power management unit, wherein each local power management unit corresponds to one of the multiple processing units and is configured inside the processing unit, the processing unit retrieves instructions and data for data processing from main memory, and the global power management unit communicates with a cache between the processing unit and the main memory. The global power management unit is used to determine the global power budget based on the previous power management parameters and the previous global power budget of each processing unit, and to allocate the local power budget of each processing unit according to the global power budget, the power management parameters reported by each processing unit and the cached power management parameters. The local power management unit is used to manage the local power resources of the corresponding processing unit according to the allocated local power budget to execute tasks, and to report the power management parameters when executing tasks to the global power management unit.
11. The multi-processor unit system according to claim 10, wherein, The global power management unit includes: A global storage unit is used to store the global power consumption budget and the power consumption management parameters of each processing unit; A global management unit is configured to allocate a local power budget for each processing unit based on the global power budget and the power management parameters of each processing unit; and The local power consumption management unit includes: Local storage units are used to store the local power budget and power management parameters of the corresponding processing units; The local management unit is used to determine the local power resource management strategy based on the local power budget and power management parameters of the processing unit.
12. The multi-processor unit system according to claim 10, wherein, The multi-processing unit system is a distributed processing system, and the processing unit is a processing unit in the distributed processing system; or, the multi-processing unit system is a multi-core processing unit, and the processing unit is a core in the multi-core processing unit.
13. A power management module for a multiprocessor unit system, comprising a plurality of local power management units and a global power management unit, each local power management unit corresponding to a processing unit in the multiprocessor unit system and configured inside the processing unit, wherein the processing unit retrieves instructions and data for data processing from main memory, and the global power management unit communicates with a cache between the processing unit and the main memory. The global power management unit is used to determine the global power budget based on the previous power management parameters and the previous global power budget of each processing unit, and to allocate the local power budget of each processing unit according to the global power budget, the power management parameters reported by each processing unit and the cached power management parameters. The local power management unit is used to manage the local power resources of the corresponding processing unit according to the allocated local power budget to execute tasks, and to report the power management parameters when executing tasks to the global power management unit.
Citation Information
Patent Citations
Dynamic power budget allocation in multi-processor system
US20190377395A1
Managing power consumption of multiple computing node clusters in a computing rack system
US20200042068A1