A GPGPU Streaming Multiprocessor Power Saving Scheduling Method, Device and Medium
By identifying and suspending memory access-intensive loads on the SM core, prioritizing the processing of computation-intensive loads, and entering the SM core with the stochastic state after the calculation is completed, the problem of low energy efficiency utilization of intensive loads in GPGPU scheduling execution is solved, achieving higher energy efficiency ratio and power consumption savings.
Patent Information
- Application Number
- CN202510096500.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-22
AI Technical Summary
During the existing GPGPU scheduling and execution, the energy efficiency of intensive loads is low, which is easy to form memory contention, resulting in waste of resources and energy consumption, which is not conducive to saving power consumption.
By identifying and suspending memory access-intensive loads on the SM core, priority is given to processing computation-intensive loads. When the computation-intensive load is completed, the thread block of the memory access-intensive load is activated for calculation, and the SM core that completes the computation-intensive load is powered or clocked to enable power consumption.
It improves the energy efficiency ratio of GPGPU, reduces memory contention, reduces the impact of computing-intensive tasks on memory-intensive tasks, and saves power consumption.
Smart Images

Figure CN119537038B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of GPGPU scheduling and execution, and particularly to a power-saving scheduling method, device, and medium for GPGPU streaming multiprocessors. Background Art
[0002] Due to their significant computing power, accelerators such as GPGPUs are widely used for different workloads. These architectures allow thousands of threads to be executed in parallel through programming models such as CUDA (Compute Unified Device Architecture, a parallel computing platform and programming model) or OpenCL (Open Computing Language). The programming model results in a hierarchical structure of threads, where a group of threads forms a warp (a basic scheduling unit consisting of 32 threads) or a wavefront (a set of threads executing in parallel), and a group of warps forms a thread block, called a CTA (Cooperative Thread Array) or a workgroup. Inside the GPGPU, there are two levels of scheduling: a thread block or CTA scheduler assigns CTAs to each core, and a warp scheduler determines which warp to execute within the core. To improve the performance of these parallel architectures, improved warp schedulers or CTA schedulers have been proposed to increase resource utilization and program execution parallelism.
[0003] However, for memory-intensive workloads, maximizing the utilization of all cores does not necessarily result in the highest energy efficiency. At the same time, it may cause pipeline stalls during operation due to memory contention and insufficient resources of the memory access controller, resulting in waste of resources and energy consumption. Summary of the Invention
[0004] Embodiments of this application provide a power-saving scheduling method, device, and medium for GPGPU streaming multiprocessors to solve the following technical problems: During the existing GPGPU scheduling and execution process, the energy efficiency utilization rate of intensive loads is low, memory contention is easily formed, resulting in waste of resources and energy consumption, which is not conducive to power saving.
[0005] Embodiments of this application adopt the following technical solutions:
[0006] On the one hand, an embodiment of the present application provides a GPGPU stream multiprocessor resource saving scheduling method, comprising: activating all stream multiprocessor units connected to a thread block scheduling unit; transmitting the task load assigned by the thread block scheduling unit to the stream multiprocessor unit, and judging the category of the task load as an intensive load through a monitoring cycle and an SM core in the stream multiprocessor unit, and determining a memory access SM core corresponding to a memory access intensive load; wherein the intensive load comprises: a computationally intensive load and the memory access intensive load; suspending the stream multiprocessor unit corresponding to the memory access SM core; and performing computational control on the computational SM cores in the remaining stream multiprocessor units to obtain a computational execution state of the computational SM core; wherein the computational execution state comprises: an activation state and a resource saving state; according to the computational execution state of the computational SM core, activating and controlling the stream multiprocessor unit corresponding to the memory access SM core to obtain a task execution state of the current task load.
[0007] The embodiment of the present application can identify and suspend memory access intensive loads on SM (streaming multiprocessor, which is an important computing unit in NVIDIA Volta and later processor architectures), give priority to computing compute-intensive loads, and when all compute-intensive loads are completed, activate the thread blocks of the suspended memory access intensive loads for computing, and at the same time, power gate or clock gate enable the SM core that completes the compute-intensive load, thereby reducing static power consumption and dynamic power consumption, thereby improving energy efficiency. In this way, compute-intensive tasks and memory-access intensive tasks are executed separately, reducing the impact of memory-access intensive tasks on compute-intensive tasks, reducing memory contention, and at the same time, the streaming multiprocessor enters a power-saving state after the calculation is completed, thereby saving power consumption.
[0008] In a feasible implementation, all stream multiprocessor units connected to the thread block scheduling unit are activated, specifically including: transmitting the received thread task load to the thread block scheduling unit through the Host server; wherein each thread block in the thread block scheduling unit can only be executed on one of the stream multiprocessor units, and one of the stream multiprocessors can execute multiple thread blocks simultaneously; before allocating the task load to the stream multiprocessor units through the thread block scheduling unit, controlling all SM cores in the stream multiprocessor units to remain activated.
[0009] In a feasible implementation manner, before determining, through the monitoring period and the SM core in the streaming multiprocessor unit, the category of the task load related to intensive loads and determining the memory access SM core corresponding to the memory access intensive load, the method further includes: configuring the monitoring period for each SM core in each streaming multiprocessor unit to obtain a clock monitoring period; wherein the unit quantity of the clock monitoring period is associated with the task attributes of the thread task load; through the clock monitoring period, performing cycle monitoring on the memory access intensive load for memory access stalls to obtain a memory access stall cycle; wherein the memory access stall cycle is the total number of cycles for monitoring the stalls of the SM core due to memory access.
[0010] In a feasible implementation manner, determining, through the monitoring period and the SM core in the streaming multiprocessor unit, the category of the task load related to intensive loads and determining the memory access SM core corresponding to the memory access intensive load specifically includes: when the memory access stall cycle is greater than one-third of the clock monitoring period, determining the task load running under the current SM core as the memory access intensive load; identifying and collecting the number of running SM cores corresponding to the memory access intensive load; when the number of running SM cores is greater than one-half of the number of SM cores in the active state, taking the running SM cores corresponding to the number of running SM cores as the memory access SM cores; wherein the memory access SM cores only contain the memory access intensive load; determining the streaming multiprocessor unit corresponding to the memory access SM cores as the memory access streaming multiprocessor unit for processing the memory access intensive load.
[0011] In a feasible implementation manner, performing a suspension process on the streaming multiprocessor unit corresponding to the memory access SM cores specifically includes: the streaming multiprocessor unit includes an improved Warp scheduler and a pipelining module; when the streaming multiprocessor unit is determined as the memory access streaming multiprocessor unit for processing the memory access intensive load, performing execution identification processing on the Warp scheduler in the memory access streaming multiprocessor unit for memory access instructions; when a memory access instruction is identified, pausing the Warp scheduler related to the memory access instruction; based on the Warp scheduler in the memory access streaming multiprocessor unit in the paused state, controlling the memory access streaming multiprocessor unit to enter the suspended state.
[0012] In a feasible implementation, the computing SM cores in the remaining stream multiprocessor units are operationally controlled to obtain the computing execution status of the computing SM cores, which specifically includes: determining the remaining stream multiprocessor units as computing stream multiprocessor units; where the computing stream multiprocessor units only contain the compute-intensive loads; through the computing stream multiprocessor units, operationally controlling the compute-intensive loads; if the computing SM cores in the computing stream multiprocessor units are in continuous operation of the compute-intensive loads, then determining the computing SM cores as the active state; if the computing SM cores in the computing stream multiprocessor units have completed the operation of the compute-intensive loads, then determining the computing SM cores as the power-saving state; continuously monitoring the computing execution status of the computing SM cores in all the computing stream multiprocessor units.
[0013] In a feasible implementation, according to the computing execution status of the computing SM cores, activation control is performed on the stream multiprocessor units corresponding to the memory access SM cores to obtain the task execution status of the current task load, which specifically includes: if the activation state exists in the computing execution status of the computing SM cores, then no activation control is performed on the stream multiprocessor units corresponding to the memory access SM cores; if the computing execution status of the computing SM cores are all in the power-saving state, then activation control is performed on the memory access stream multiprocessor units corresponding to the memory access SM cores; when the memory access SM cores in the memory access stream multiprocessor units are in the active state, task execution processing is performed on the memory access-intensive loads in the memory access stream multiprocessor units until all the memory access-intensive loads in the memory access stream multiprocessor units are completed, obtaining the memory access completion status; where the task execution status includes: the power-saving state of the compute-intensive loads and the memory access completion status of the memory access-intensive loads.
[0014] In a feasible implementation, after performing activation control on the stream multiprocessor units corresponding to the memory access SM cores to obtain the task execution status of the current task load, the method further includes: if the task execution status all includes the power-saving state and the memory access completion status, then controlling all the stream multiprocessor units to be readjusted to the active state; based on the active state, controlling the stream multiprocessor units to wait to receive the next round of thread task loads allocated by the thread block scheduler.
[0015] Second aspect, an embodiment of the present application further provides a power-saving scheduling device for a GPGPU streaming multiprocessor. The device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, so that the at least one processor can execute a power-saving scheduling method for a GPGPU streaming multiprocessor according to any of the above embodiments.
[0016] Third aspect, an embodiment of the present application further provides a non-volatile computer storage medium, characterized in that the storage medium is a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores at least one program, and each program includes instructions, and when the instructions are executed by a terminal, the terminal executes a power-saving scheduling method for a GPGPU streaming multiprocessor according to any of the above embodiments.
[0017] The present application provides a power-saving scheduling method, device and medium for a GPGPU streaming multiprocessor. Compared with the prior art, the embodiments of the present application have the following beneficial technical effects:
[0018] 1. Resource optimization: By activating processing and suspending processing, this method can dynamically adjust the use of streaming multiprocessor units according to the task type, thereby optimizing the usage efficiency of hardware resources.
[0019] 2. Power consumption reduction: Through power-saving control (placing some SM cores in a low-power state), this method helps to reduce the overall power consumption of the GPU, which is of great significance for improving energy efficiency and extending the device life.
[0020] 3. Performance improvement: By judging the category of the task load and specifically allocating memory access-intensive loads to specific SM cores, the efficiency of task processing and the overall performance of the GPU can be improved.
[0021] 4. Load balancing: By monitoring the cycle and classifying the task load by SM cores, balanced distribution of different types of loads can be achieved, thereby avoiding overloading of some units while other units are idle.
[0022] 5. Dynamic response: This method can dynamically adjust the state of the memory access SM core according to the current computing execution state of the computing SM core, enabling the GPU to flexibly respond to different computing requirements.
[0023] 6. Latency reduction: By optimizing the scheduling strategy, task execution latency caused by resource conflicts or resource shortages can be reduced.
[0024] 7. Enhanced scalability: This method can adapt to GPGPU systems of different scales, provides a general resource management method, and enhances the scalability of the system.
[0025] 8. Simplified programming model: This scheduling method may provide a more transparent and efficient programming model for developers because the tasks of task scheduling and resource management are automatically handled.
[0026] 9. Improved stability: By precisely controlling the state of the SM cores, the system instability phenomenon caused by resource competition can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings:
[0028] Figure 1 is a flowchart of a power-saving scheduling method for a GPGPU stream multi-processor provided by an embodiment of the present application;
[0029] Figure 2 is a schematic diagram of a low-power architecture provided by an embodiment of the present application;
[0030] Figure 3 is a schematic diagram of a suspension power-saving process provided by an embodiment of the present application;
[0031] Figure 4 is a schematic diagram of the operation effect of a power-saving scheduling method provided by an embodiment of the present application;
[0032] Figure 5 is a schematic diagram of the structure of a power-saving scheduling device for a GPGPU stream multi-processor provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0034] An embodiment of the present application provides a power-saving scheduling method for a GPGPU stream multi-processor. As Figure 1 shown, the power-saving scheduling method for a GPGPU stream multi-processor specifically includes steps S101-S104:
[0035] It should be noted thatFigure 2 A schematic diagram of a low-power architecture provided by an embodiment of the present application. As Figure 2 shown, a power-saving scheduling system for a GPGPU stream multiprocessor section. Among them, the low-power architecture includes a thread block scheduling unit, a stream multiprocessor unit (SM), a dynamic monitoring unit, a running reduction unit, and a gating enable unit.
[0036] The thread block scheduling unit can distribute the threads received from the Host to the stream multiprocessor unit. Each thread block can only be executed on one stream multiprocessor unit. One stream multiprocessor can execute multiple thread blocks simultaneously. For a suspended stream multiprocessor, even if there are still remaining computing resources, no more thread blocks will be allocated. Only after a round of suspension-power-saving process is completed will threads be reallocated.
[0037] The stream multiprocessor unit contains an improved Warp scheduler and a pipelined part. Once a stream multiprocessor enters the suspended state, the Warp scheduler will continue to issue Warp executions until a Warp needs to execute a memory access instruction. These Warps will be paused to limit requests to the memory system. Once all Warps are paused, the stream multiprocessor core is truly in the suspended state.
[0038] The dynamic monitoring unit can dynamically monitor the load situation in the current stream multiprocessor, determine whether it is a memory access-intensive load or a compute-intensive load, and send the judgment result to the running reduction unit.
[0039] The running reduction unit will suspend the memory access-intensive load and give priority to running the compute-intensive load. After the compute-intensive load is calculated, it will enable the stream multiprocessor that executes the compute-intensive load to enter the power-saving state and turn off the power gating or clock gating as needed.
[0040] The gating enable unit has clock gating and power gating enable, and can dynamically start and stop the clock and power of the stream multiprocessor.
[0041] As a feasible implementation, the thread block scheduling unit can distribute the threads received from the Host (server) to the streaming multiprocessor units. Each thread block can only be executed on one streaming multiprocessor unit, and one streaming multiprocessor can execute multiple thread blocks simultaneously. For the suspended streaming multiprocessor, even if there are still remaining computing resources, no more thread blocks will be allocated. Only after a round of suspension - power saving process is completed, will threads be re - allocated. Among them, the dynamic monitoring unit can dynamically monitor the load situation in the current streaming multiprocessor, determine whether it is a memory - access - intensive load or a compute - intensive load, and send the judgment result to the operation reduction unit. Among them, the operation reduction unit will suspend the memory - access - intensive load and give priority to running the compute - intensive load. After the compute - intensive load is calculated, it will enable the streaming multiprocessor that executes the compute - intensive load to enter the power - saving state and turn off the power gating or clock gating as needed. Among them, the gating enable unit has clock gating and power gating enable, and can dynamically turn on and off the clock and power of the streaming multiprocessor, enabling the streaming multiprocessor to enter the power - saving state.
[0042] S101. Activate all the streaming multiprocessor units connected to the thread block scheduling unit.
[0043] Specifically, it is necessary to first transport the received thread task load to the thread block scheduling unit through the Host server. Among them, each thread block in the thread block scheduling unit can only be executed on one streaming multiprocessor unit, and one streaming multiprocessor can execute multiple thread blocks simultaneously.
[0044] Furthermore, before the task load is allocated to the streaming multiprocessor units through the thread block scheduling unit, control all the SM cores in the streaming multiprocessor units to remain in the active state.
[0045] In one embodiment, Figure 3 This is a schematic diagram of a suspension and power - saving process provided by an embodiment of the present application. As Figure 3 shown, taking 6 streaming multiprocessors (SMs) as an example, at step 0, all SM cores are in the active state so that the task load can be allocated to the streaming multiprocessor units through the thread block scheduling unit in the subsequent process.
[0046] S102. Transmit the task load allocated by the thread block scheduling unit to the streaming multiprocessor units, and through the monitoring period and SM cores in the streaming multiprocessor units, judge the category of the task load for intensive load, and determine the memory - access SM cores corresponding to the memory - access - intensive load. Among them, the intensive load includes: compute - intensive load and memory - access - intensive load.
[0047] Specifically, first configure the monitoring period for each SM core in each streaming multi-processor unit to obtain a clock monitoring period. Among them, the unit quantity of the clock monitoring period is associated with the task attributes of the thread task load.
[0048] Further, through the clock monitoring period, perform cycle monitoring on the memory access intensive load regarding memory access stalls to obtain the memory access stall cycle. Among them, the memory access stall cycle is the total number of cycles for monitoring the SM core to stall due to memory access.
[0049] Further, when the memory access stall cycle is greater than one-third of the clock monitoring period, determine the task load running under the current SM core as a memory access intensive load.
[0050] Further, identify and collect the number of running SM cores corresponding to the memory access intensive load.
[0051] Further, when the number of running SM cores is greater than one-half of the number of SM cores in the active state, the running SM cores corresponding to the number of running SM cores are memory access SM cores. Among them, the memory access SM cores only contain memory access intensive loads.
[0052] Further, determine the streaming multi-processor unit corresponding to the memory access SM core as the memory access streaming multi-processor unit for processing the memory access intensive load.
[0053] In one embodiment, as Figure 2 shown, at step 1, the loads running on three SM cores are determined as memory access intensive loads and enter the suspended state, pausing execution. Among them, the determination algorithm process for the SM to be determined as a memory access intensive load is: each SM core will have a monitoring period, set to 15,000 clock cycles. Within these 15,000 clock cycles, monitor the total number of cycles for the SM core to stall due to memory access. When the number of stalled clock cycles (memory access stall cycle) is greater than one-third of the monitoring period (clock monitoring period), the load running on this core is determined as a memory access intensive load. When the number of SM cores determined as memory access intensive loads (number of running SM cores) is greater than or equal to one-half of the number of cores of the already activated SM, suspend the SM cores determined as memory access intensive loads (memory access SM cores: SM4, SM5).
[0054] In one embodiment, it includes the following core code:
[0055] Function MONITOR
[0056] for each (Monitor cycle) do (here do represents a for loop)
[0057] for (i = 1; i <= Counter_active_state; i++) do (Here, "do" represents a for loop)
[0058] if stall cycle(i) > mon cycle / 3 then
[0059] Counter_mem++.
[0060] stall cycle(i) = 0.
[0061] end if
[0062] end for
[0063] if Counter_mem >= Counter_active_state / 2 then
[0064] select a COREj that is in active_state.
[0065] COREj switch from active_state to suspended state.
[0066] Counter_suspended_state++.
[0067] Counter_active_state --.
[0068] Counter_mem --.
[0069] end if
[0070] end for
[0071] end function
[0072] Among them, the explanations of keywords in the code:
[0073] Monitor cycle: The monitoring cycle, which is set according to the running cycle of a thread block executed by the current SM core (the number of cycles from when the thread block is dispatched to the SM core until it is executed, and this value is an estimated value).
[0074] Stall cycle: The total number of cycles that the SM core pauses due to memory access operations.
[0075] Counter_active_state: Represents the number of active SM cores. The activation status flag indicates that the compute cores can operate normally.
[0076] Counter_ suspended state: Represents the number of SM cores that enter the suspended state.
[0077] Counter_mem: Represents the number of SM cores determined to be memory access intensive workloads.
[0078] Through the code implementation of the above dynamic monitoring unit algorithm, the following operations can be cycled in each detection period (one clock cycle in hardware implementation): For each activated compute core, judge the memory access stall cycles. When the total memory access stall cycles of a certain compute core are greater than 1 / 3 of the monitoring period, it is determined that the current compute core is a memory access intensive workload. At this time, the memory access intensive workload counter (Counter_mem) is incremented by 1. When it is judged that the memory access intensive workload counter is greater than or equal to 1 / 2 of the number of activated compute cores, an activated compute core is selected to enter the suspended state, the corresponding suspended state counter is incremented by 1, the activation state counter is decremented by 1, and the memory intensive workload counter is decremented by 1.
[0079] S103. Suspend the streaming multiprocessor unit corresponding to the memory access SM core. And perform operation control on the compute SM cores in the remaining streaming multiprocessor units to obtain the compute execution status of the compute SM cores. Among them, the compute execution status includes: activation status and power saving status.
[0080] Specifically, the streaming multiprocessor unit includes an improved Warp scheduler and a pipelining module. When the streaming multiprocessor unit is determined to be a memory access streaming multiprocessor unit for processing memory access intensive workloads, perform execution recognition processing on the Warp scheduler in the memory access streaming multiprocessor unit for memory access instructions. When a memory access instruction is recognized, suspend the Warp scheduler related to the memory access instruction.
[0081] Furthermore, based on the Warp scheduler in the memory access streaming multiprocessor unit that is in the suspended state, control the memory access streaming multiprocessor unit to enter the suspended state.
[0082] Furthermore, determine the remaining streaming multiprocessor units as compute streaming multiprocessor units. Among them, the compute streaming multiprocessor unit only contains compute intensive workloads.
[0083] Further, it is also necessary to perform operation control on compute-intensive loads through the compute stream multi-processor unit. If the compute SM core in the compute stream multi-processor unit is in continuous operation on a compute-intensive load, the compute SM core is determined to be in an active state. If the compute SM core in the compute stream multi-processor unit has completed the operation on the compute-intensive load, the compute SM core is determined to be in a power-saving state.
[0084] Further, continuously monitor the compute execution status of the compute SM cores in all compute stream multi-processor units.
[0085] In one embodiment, as Figure 2 shown, at step 2, the compute SM cores (SM2, SM3) of two compute-intensive loads have completed their computations and are enabled to enter the power-saving state (have completed the operations on the compute-intensive loads). Two compute SM cores (SM0, SM1) are still in continuous operation on the compute-intensive loads, so the compute SM cores are determined to be in an active state, and the remaining two memory access stream multi-processor units (SM4, SM5) are still in a suspended state.
[0086] In one embodiment, the following core code is included:
[0087] Function REDUCE
[0088] while Counter_suspended_state > 0 do
[0089] if COREi finished then
[0090] power gate COREi
[0091] select a COREj that is in suspended state.
[0092] COREj switch from suspended to active state.
[0093] Counter_suspended_state --.
[0094] end if
[0095] end while
[0096] end Function
[0097] Among them, the explanations of the keywords in the code are:
[0098] COREi: represents a certain SM core in the active state.
[0099] COREj: represents a certain SM core in the suspended state.
[0100] Through the above code implementation of the running reduction unit algorithm, when the value of the suspended state counter > 0, it indicates that there is already a suspended memory access-intensive load at this time. Under this judgment condition, when any one of the compute-intensive cores finishes execution, the power gating enable of its corresponding clock core is turned off, enabling it to enter the power-saving mode, and a suspended access-intensive compute core is selected to enter the active state and continue running. In this way, the access-intensive loads can be executed separately, which not only reduces the impact on the compute-intensive loads but also reduces the competition among the access-intensive loads.
[0101] S104. According to the compute execution status of the compute SM cores, perform activation control on the streaming multiprocessor units corresponding to the memory access SM cores to obtain the task execution status of the current task load.
[0102] Specifically, if there is an active state in the compute execution status of the compute SM cores, no activation control is performed on the streaming multiprocessor units corresponding to the memory access SM cores.
[0103] If the compute execution status of all compute SM cores is in the power-saving state, activation control is performed on the memory access streaming multiprocessor units corresponding to the memory access SM cores.
[0104] When the memory access SM cores in the memory access streaming multiprocessor units are in the active state, task execution processing is performed on the memory access-intensive loads in the memory access streaming multiprocessor units until all the memory access-intensive loads in the memory access streaming multiprocessor units complete execution, obtaining the memory access completion status.
[0105] Among them, the task execution status includes: the power-saving state of the compute-intensive load and the memory access completion status of the memory access-intensive load.
[0106] In one embodiment, as Figure 2 shown, at step 3, all the compute SM cores (SM0, SM1, SM2, SM3) of the compute-intensive task loads have completed their operations and become in the power-saving state. Then, the previously suspended memory access-intensive loads are activated one by one for execution, that is, activation control is performed on the streaming multiprocessor units corresponding to the memory access SM cores (SM4, SM5), all of which become in the active state. Then, task execution processing is performed on the memory access-intensive loads until all the memory access-intensive loads in the memory access streaming multiprocessor units complete execution, obtaining the memory access completion status, which means the completion of the processing of this round of thread task loads.
[0107] Further, if the task execution statuses all include the power-saving state and the memory access completion state, then control all streaming multi-processor units to be readjusted to the active state. Finally, based on the active state, control the streaming multi-processor units to wait for receiving the next round of thread task loads allocated by the thread block scheduling unit.
[0108] In one embodiment, Figure 4 is a schematic diagram of the operation effect of a power-saving scheduling method provided by an embodiment of the present application. As Figure 4 shown, during the execution process, after judging the number of cycles per square inch of the SM computing core, based on the preset operating cycle value, using the proportional relationship between the operating cycles, the computing-intensive load or memory access-intensive load to which different SM computing cores belong is determined, and then corresponding activation operations or suspension waits are performed. That is, power gating or clock gating is enabled for the SM cores that have completed the computing-intensive load to reduce the consumption of static power and dynamic power, thereby improving the energy efficiency ratio. In this way, the computing-intensive tasks and memory access-intensive tasks are executed separately, reducing the influence of the memory access-intensive tasks on the computing-intensive tasks, reducing memory contention, and at the same time, after the calculation is completed, the streaming multi-processor enters the power-saving state, saving power consumption.
[0109] In addition, an embodiment of the present application also provides a GPGPU streaming multi-processor power-saving scheduling device. As Figure 5 shown, the GPGPU streaming multi-processor power-saving scheduling device 500 specifically includes:
[0110] At least one processor 501. And a memory 502 communicatively connected to the at least one processor 501. Among them, the memory 502 stores instructions that can be executed by the at least one processor 501, so that the at least one processor 501 can execute:
[0111] Perform all activation processing on the streaming multi-processor units connected to the thread block scheduling unit;
[0112] Transmit the task loads allocated by the thread block scheduling unit to the streaming multi-processor units, and through the monitoring cycles and SM cores in the streaming multi-processor units, judge the category of the intensive load of the task loads to determine the memory access SM cores corresponding to the memory access-intensive load; among them, the intensive load includes: computing-intensive load and memory access-intensive load;
[0113] Perform suspension processing on the streaming multi-processor units corresponding to the memory access SM cores; and perform operation control on the computing SM cores in the remaining streaming multi-processor units to obtain the computing execution status of the computing SM cores; among them, the computing execution status includes: active state and power-saving state;
[0114] According to the computing execution status of the SM core, activation control is performed on the streaming multiprocessor unit corresponding to the memory access SM core to obtain the task execution status of the current task load.
[0115] Embodiments of the present application can identify and suspend memory access-intensive loads on the SM (streaming multiprocessor, an important computing unit in NVIDIA Volta and subsequent processor architectures), prioritize the operation of compute-intensive loads, and after all compute-intensive loads are completed, activate the thread blocks of the suspended memory access-intensive loads for operation. At the same time, power gating or clock gating is enabled for the SM cores that have completed the compute-intensive loads to reduce static and dynamic power consumption, thereby improving the energy efficiency ratio. By separating the execution of compute-intensive tasks and memory access-intensive tasks in this way, the impact of memory access-intensive tasks on compute-intensive tasks is reduced, memory contention is reduced, and after the computation is completed, the streaming multiprocessor enters a power-saving state, saving power consumption.
[0116] The various embodiments in the present application are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0117] The devices and media provided by the embodiments of the present application correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0118] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0119] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks, and the functions specified in the one block or multiple blocks.
[0120] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks, and the functions specified in the one block or multiple blocks.
[0121] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks, and the functions specified in the one block or multiple blocks.
[0122] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0123] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0124] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0125] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0126] The above is only the embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the specification of the present application.
Claims
1. A GPGPU streaming multiprocessor resource saving scheduling method, characterized in that: The method comprises: Activate all the stream multiprocessor units connected to the thread block scheduling unit; The task load assigned by the thread block scheduling unit is transmitted to the stream multiprocessor unit, and the task load is subjected to a category judgment of a relevant intensive load through a monitoring cycle and an SM core in the stream multiprocessor unit, so as to determine a memory access SM core corresponding to a memory access intensive load; wherein the intensive load includes: a computationally intensive load and the memory access intensive load; Suspending the stream multiprocessor unit corresponding to the memory access SM core; and performing computation control on the computation SM cores in the remaining stream multiprocessor units to obtain the computation execution state of the computation SM core; wherein the computation execution state includes: an activation state and a resource saving state; According to the computing execution state of the computing SM core, the stream multiprocessor unit corresponding to the memory access SM core is activated and controlled to obtain the task execution state of the current task load, specifically including: If the computing execution state of the computing SM core has the activation state, then activation control is not performed on the stream multiprocessor unit corresponding to the memory access SM core; If the computing execution states of the computing SM cores are all in the resource-saving state, activating and controlling the memory access stream multiprocessor unit corresponding to the memory access SM core; When the memory access SM core in the memory access stream multiprocessor unit is in an activated state, performing task execution processing on the memory access intensive load in the memory access stream multiprocessor unit until all the memory access intensive loads in the memory access stream multiprocessor unit are executed, and a memory access completion state is obtained; The task execution status includes: a resource saving status of the computationally intensive load and a memory access completion status of the memory access intensive load.
2. A GPGPU streaming multiprocessor resource saving scheduling method according to claim 1, characterized in that: The streaming multiprocessor units connected to the thread block scheduling units are fully activated, including: The received thread task load is transmitted to the thread block scheduling unit through the host server; wherein each thread block in the thread block scheduling unit can only be executed on one stream multiprocessor unit, and one stream multiprocessor can execute multiple thread blocks at the same time; Before distributing the task load to the stream multiprocessor unit through the thread block scheduling unit, all SM cores in the stream multiprocessor unit are controlled to remain in an activated state.
3. A GPGPU streaming multiprocessor resource saving scheduling method according to claim 1, characterized in that: Before determining the category of the task load related to the intensive load through the monitoring cycle and the SM core in the stream multiprocessor unit and determining the memory access SM core corresponding to the memory access intensive load, the method further includes: Performing a monitoring cycle configuration on the SM core in each of the stream multiprocessor units to obtain a clock monitoring cycle; wherein the unit number of the clock monitoring cycle is associated with a task attribute of a thread task load; The memory access intensive load is monitored for cycles related to memory access pauses through the clock monitoring cycle to obtain a memory access pause cycle; wherein the memory access pause cycle is the total number of cycles during which the SM core is monitored to pause due to memory access.
4. A GPGPU streaming multiprocessor resource saving scheduling method according to claim 3, characterized in that: By using the monitoring cycle and the SM core in the stream multiprocessor unit, the task load is subjected to a category judgment related to the intensive load, and a memory access SM core corresponding to the memory access intensive load is determined, specifically including: When the memory access pause period is greater than one third of the clock monitoring period, determining the task load currently being run by the SM core as the memory access intensive load; Identify and collect the number of running SM cores corresponding to the memory access intensive load; When the number of the running SM cores is greater than half of the number of the SM cores in the activated state, the running SM cores corresponding to the number of the running SM cores are used as the memory access SM cores; wherein the memory access SM cores only include the memory access intensive load; A streaming multiprocessor unit corresponding to the memory access SM core is determined as a memory access streaming multiprocessor unit for processing the memory access intensive load.
5. A GPGPU streaming multiprocessor resource saving scheduling method according to claim 4, characterized in that: The stream multiprocessor unit corresponding to the memory access SM core is suspended, specifically comprising: The stream multiprocessor unit includes an improved Warp scheduler and a pipeline module; When the stream multiprocessor unit is determined to be a memory access stream multiprocessor unit for processing the memory access intensive load, performing execution identification processing on the Warp scheduler in the memory access stream multiprocessor unit regarding the memory access instruction; When the memory access instruction is identified, pausing the Warp scheduler related to the memory access instruction; Based on the Warp scheduler in the memory access stream multiprocessor unit being in a paused state, the memory access stream multiprocessor unit is controlled to enter a suspended state.
6. A GPGPU streaming multiprocessor resource saving scheduling method according to claim 1, characterized in that: The calculation SM core in the remaining stream multiprocessor unit is operated and controlled to obtain the calculation execution state of the calculation SM core, specifically including: Determining the remaining stream multiprocessor units as computational stream multiprocessor units; wherein the computational stream multiprocessor units only contain the computationally intensive load; Performing computational control on the computationally intensive load by means of the computational stream multiprocessor unit; If the computing SM core in the computing stream multiprocessor unit is in the continuous operation of the computing intensive load, determining the computing SM core as an activated state; If the computing SM core in the computing stream multiprocessor unit has completed the operation of the computing intensive load, the computing SM core is determined to be in a power saving state; The computing execution status of the computing SM cores in all the computing stream multiprocessor units is continuously monitored.
7. A GPGPU streaming multiprocessor resource saving scheduling method according to claim 1, characterized in that: After activating and controlling the stream multiprocessor unit corresponding to the memory access SM core to obtain the task execution status of the current task load, the method further includes: If the task execution status includes the resource saving status and the memory access completion status, controlling all the stream multiprocessor units to readjust to the activation status; Based on the activation state, the streaming multiprocessor unit is controlled to wait for receiving a next round of thread task loads allocated by the thread block scheduling unit.
8. A GPGPU streaming multiprocessor resource-saving scheduling device, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, so that the at least one processor can execute the GPGPU streaming multi-processor resource scheduling method according to any one of claims 1-7.
9. A non-volatile computer storage medium, characterized in that: The storage medium is a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores at least one program, each of which includes instructions, and when the instructions are executed by a terminal, the terminal executes a GPGPU stream multi-processor resource scheduling method according to any one of claims 1-7.
Citation Information
Patent Citations
Multi-core apparatus and job scheduling method thereof
CN104216774A
Load balancing scheduling method and computing device
CN115391031A