Information processing device and information processing method
The information processing device optimizes power consumption and performance in real-time by using a performance-power optimization unit with roofline model data and application tables to adjust arithmetic cores and main memory, addressing the challenge of heterogeneous core configurations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-06-08
- Publication Date
- 2026-03-30
AI Technical Summary
Existing technologies fail to optimize power consumption and performance in real-time for information processing devices with multiple computing cores, especially in heterogeneous configurations, without hindering processing performance.
An information processing device equipped with a performance-power optimization unit that utilizes roofline model data and application performance definition tables to dynamically adjust the computational performance and power consumption of arithmetic cores and main memory based on application scheduling, supporting heterogeneous core configurations.
Enables highly accurate, real-time performance-power optimization that adapts to the computational intensity of applications running on multiple processing cores, ensuring optimal processing performance without excessive power consumption.
Smart Images

Figure 0007837409000001 
Figure 0007837409000002 
Figure 0007837409000003
Abstract
Description
Technical Field
[0001] This application relates to an information processing apparatus and an information processing method for controlling power consumption according to arithmetic processing performance.
Background Art
[0002] An automatic control system is generally a system in which multiple functions cooperate and integrate to perform recognition, judgment, and control. For example, an automatic driving system includes an automatic driving control unit that generates optimal control parameters from surrounding situations, and an engine control unit, a brake control unit, and a steering control unit that respectively realize engine control, brake control, and steering control of a vehicle. As the autonomy level (e.g., the automatic driving level) increases, more computing performance is required for the automatic control system.
[0003] On the other hand, in order to meet the processing performance required by the system, there is a system configuration that mounts a system-on-a-chip (SoC) with high performance, multiple arithmetic cores, and a large-capacity main memory device. On the other hand, as the performance of arithmetic cores increases and they become multi-core / single-core, and the capacity of the main memory device increases, the power consumption and heat generation of the system increase. In response to this, for example, there is an arithmetic processing device equipped with dynamic voltage and frequency scaling (DVFS). The DVFS function is a power-saving mechanism that changes the operating frequency and operating voltage of the arithmetic core to reduce power consumption. However, it is not easy to optimize and control power consumption in real time without affecting the arithmetic performance of arithmetic processing such as applications. Furthermore, optimization control considering applications executed in parallel in a system configuration including multiple arithmetic cores or applications within a plurality of container environments is required.
[0004] To address these challenges, Patent Document 1 discloses a method for determining VM / container and data placement in a hyperconverged infrastructure (HCI) environment. Furthermore, Patent Document 2 discloses a device that reduces power consumption by changing the processor frequency and instruction width based on memory access information obtainable within the processor. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2020-52730 [Patent Document 2] International Publication No. 2008 / 120274 [Overview of the project] [Problems that the invention aims to solve]
[0006] Patent Document 1 describes a method for managing the utilization status of shared computing resources and determining the destination node for new virtual machines, containers, and storage volumes based on that utilization status, so as not to exceed the upper limit of the computing resources of the destination node. It does not target the optimization of performance power of an information processing device adapted to the computational processing of an application. Patent Document 2 describes a device that reduces processor power based on memory access information within the processor, and does not target high-precision performance power optimization based on information related to the computational processing of an application. Furthermore, it does not target applications that perform parallel processing with multiple computing cores. In addition, it does not include main memory as a target for performance power optimization.
[0007] Even when combining the technologies disclosed in Patent Documents 1 and 2, there was a problem in that it was not possible to optimize the performance power of a real-time information processing device in a way that is suitable for the computational processing of applications that are processed in parallel with multiple computing cores, without hindering processing performance.
[0008] This invention was made in view of these problems and aims to enable real-time performance-power optimization control of an information processing device that is adapted to the computational processing of applications and does not hinder processing performance. Furthermore, this invention aims to enable performance-power optimization control of an information processing device adapted to the execution of applications that perform parallel processing with multiple computing cores. The multiple computing cores are not limited to computing cores of the same type, but also support heterogeneous information processing device configurations that have multiple computing cores with different processing methods. [Means for solving the problem]
[0009] The information processing device disclosed in this application is Equipped with a power-saving mechanism multiple Includes arithmetic core and main memory Computer hardware and System software that runs on computer hardware, System software and applications that run in a container execution environment for system software, In an information processing device including, It includes a performance-power optimization unit that performs optimization processing for the computational performance and power consumption of the information processing device. The performance-power optimization unit includes a performance-power optimization program, roofline model data indicating the computational intensity and computational performance per unit time of the computer hardware, and an application performance definition table including application computational intensity information. The performance-power optimization program includes a scheduling information acquisition unit that obtains application scheduling information for the computing cores of the computer hardware, and an optimization calculation of the computer hardware's roofline. When multiple computing cores are executing their respective threads at the same time, the program calculates the bandwidth performance that each thread requests from main memory from the roofline model data for each computing core, and then calculates the bandwidth performance to be allocated to each thread based on the ratio of each individual bandwidth performance to the sum of the calculated bandwidth performances. It includes a roofline optimization calculation unit. [Effects of the Invention]
[0010] The present invention provides an information processing device that enables highly accurate performance-power optimization of the information processing device (processing cores and main memory) adapted to the computational intensity of an application processed in parallel by multiple processing cores. It provides an information processing device that works in conjunction with the application's scheduling information to provide real-time performance-power optimization that does not hinder the application's processing performance. Furthermore, the device supports a heterogeneous configuration where the multiple processing cores are not all of the same type, but rather multiple processing cores with different processing methods. [Brief explanation of the drawing]
[0011] [Figure 1] This is a block diagram showing the configuration of the information processing device according to Embodiment 1. [Figure 2] The figure which shows the example of the roof line model data in the information processing apparatus which concerns on Embodiment 1. [Figure 3] The table figure which shows the example of the content of the application performance definition table in the information processing apparatus which concerns on Embodiment 1. [Figure 4A] The figure which shows the example of the information acquired from the scheduling information acquisition part in the information processing apparatus which concerns on Embodiment 1. [Figure 4B] The figure which shows the example of the information acquired from the scheduling information acquisition part in the information processing apparatus which concerns on Embodiment 1. [Figure 5] The figure which shows the operation flow of the performance power optimization part in the information processing apparatus which concerns on Embodiment 1. [Figure 6] The figure which shows the operation flow of the roof line optimization calculation part in the information processing apparatus which concerns on Embodiment 1. [Figure 7A] The figure which shows the example of the roof line optimization calculation when the arithmetic intensity of the application by the roof line optimization calculation part in the information processing apparatus which concerns on Embodiment 1 has an intersection point with the gradient part of the roof line model data. [Figure 7B] The figure which shows the example of the roof line optimization calculation when the arithmetic intensity of the application by the roof line optimization calculation part in the information processing apparatus which concerns on Embodiment 1 has an intersection point with the gradient part of the roof line model data. [Figure 7C] The figure which shows the example of the roof line optimization calculation when the arithmetic intensity of the application by the roof line optimization calculation part in the information processing apparatus which concerns on Embodiment 1 has an intersection point with the gradient part of the roof line model data. [Figure 8A] The figure which shows the example of the roof line optimization calculation when the arithmetic intensity of the application by the roof line optimization calculation part in the information processing apparatus which concerns on Embodiment 1 has an intersection point with the roof part of the roof line model data. [Figure 8B]A diagram showing an example of loop line optimization calculation when the operation intensity of an application by the loop line optimization calculation unit in the information processing apparatus according to Embodiment 1 has an intersection with the loop part of the loop line model data. [Figure 8C] A diagram showing an example of loop line optimization calculation when the operation intensity of an application by the loop line optimization calculation unit in the information processing apparatus according to Embodiment 1 has an intersection with the loop part of the loop line model data. [Figure 9] A diagram showing the operation flow of the loop line control unit in the information processing apparatus according to Embodiment 1. [Figure 10] A diagram showing an example of information acquired from the scheduling information acquisition unit in the information processing apparatus according to Embodiment 2. [Figure 11] A diagram showing the operation flow of the loop line optimization calculation unit in the information processing apparatus according to Embodiment 2. [Figure 12] A diagram showing the operation flow of the control performance calculation process included in the loop line optimization calculation unit in the information processing apparatus according to Embodiment 2. [Figure 13A] A diagram showing an example of loop line optimization calculation by the loop line optimization calculation unit in the information processing apparatus according to Embodiment 2. [Figure 13B] A diagram showing an example of loop line optimization calculation by the loop line optimization calculation unit in the information processing apparatus according to Embodiment 2. [Figure 13C] A diagram showing an example of loop line optimization calculation by the loop line optimization calculation unit in the information processing apparatus according to Embodiment 2. [Figure 13D] A diagram showing an example of loop line optimization calculation by the loop line optimization calculation unit in the information processing apparatus according to Embodiment 2. [Figure 14] A diagram showing the operation flow of the execution thread acquisition process in the loop line control unit of the information processing apparatus according to Embodiment 2.
Embodiments for Carrying Out the Invention
[0012] Embodiment 1. Figure 1 is a block diagram showing the configuration of the information processing device 1000 according to Embodiment 1. The information processing device 1000 includes at least a performance-power optimization unit 1100, system software 1200, computer hardware 1300, and an application 1400.
[0013] The performance-power optimization unit 1100 performs calculations to optimize and control the performance-power of the computer hardware 1300 to suit the computational processing of the application 1400. The system software 1200 at least assigns and executes application 1400 on computer hardware 1300, and also obtains the execution status of computer hardware 1300 and controls the performance and power of computer hardware 1300.
[0014] Application 1400 runs using the resources of the computer hardware 1300 allocated by the system software 1200. Alternatively, application 1400 may be application 1610 running within the container execution environment 1600 provided by the container runtime 1500 together with the system software 1200.
[0015] The performance-power optimization unit 1100 includes roofline model data 1101, an application performance definition table 1102, and a performance-power optimization program 1110.
[0016] The roofline model data 1101 represents the relationship between computational intensity and computational performance per unit time in the computer hardware 1300. Computational intensity represents the amount of computation per data size, e.g., 1 byte. Computational performance represents the amount of computation per unit time, e.g., 1 second.
[0017] The application performance definition table 1102 describes the information necessary to perform optimization calculations for the performance power control of the computer hardware 1300 when running application 1400 and application 1610, which runs within the container execution environment 1600, on the computer hardware 1300.
[0018] The performance power optimization program 1110 includes a scheduling information acquisition unit 1111, a roofline optimization calculation unit 1112, and a roofline control setting unit 1113.
[0019] The scheduling information acquisition unit 1111 acquires allocation information for the computer hardware 1300 for application 1400 and application 1610 running within the container execution environment 1600.
[0020] The roofline optimization calculation unit 1112 performs optimization calculations for the performance and power control of the computer hardware 1300 when executing application 1400 and application 1610, which runs within the container execution environment 1600, on the computer hardware 1300.
[0021] The roofline control setting unit 1113, based on the optimization calculation of performance power control performed by the roofline optimization calculation unit 1112, sets the performance power control settings for the computer hardware 1300 to the system software 1200.
[0022] The system software 1200 includes a scheduler 1201, a roofline control unit 1202, a computing core performance power control unit 1203, and a memory bandwidth power control unit 1204.
[0023] The scheduler 1201 controls the allocation and execution of application 1400 and application 1610, which runs within the container execution environment 1600, on the computer hardware 1300.
[0024] The roofline control unit 1202 controls the roofline model in the computer hardware 1300, which represents the relationship between computational intensity and computational performance per unit time.
[0025] The arithmetic core performance power control unit 1203 controls the processing performance and power consumption of the arithmetic cores provided in the computer hardware 1300.
[0026] The memory bandwidth power control unit 1204 controls the bandwidth performance and power consumption of the main memory provided in the computer hardware 1300.
[0027] The computer hardware 1300 includes at least one or more arithmetic cores 1310 and main memory 1320. The arithmetic core 1310 includes at least a power-saving mechanism that controls the power consumption of the arithmetic machine and the arithmetic core 1310. The arithmetic core 1310 may be configured not only as a single type of arithmetic core, but also as a heterogeneous information processing device having multiple arithmetic cores with different processing methods. The main memory 1320 stores data for arithmetic processing, and the arithmetic core 1310 loads and stores data into the main memory 1320. The controller (not shown) of the main memory 1320 includes a power saving mechanism that controls the power consumption of the main memory 1320. The main memory 1320 may also include a cache memory (not shown) between it and the arithmetic core 1310.
[0028] Figure 2 shows an example of roofline model data 1101 in Embodiment 1.
[0029] The roofline model data 1101 consists of at least roofline model data 2000 for each computing core 1310 provided in the computer hardware 1300. For example, the example in Figure 2 shows the roofline model data 2000 represented as a graph with computational intensity on the horizontal axis and computational performance on the vertical axis. Furthermore, it includes information on the maximum and minimum values A and B of the controllable computational performance of the computing core 1310, and information on the maximum and minimum values D of the controllable bandwidth performance of the main memory 1320.
[0030] The roofline model data 2000 defines upper limits on computational performance relative to computational intensity for the controllable computational performance of the computing core 1310 and the controllable bandwidth performance of the main memory 1320. Further details on the roofline model can be found, for example, in "Samuel Williams, Andrew Waterman and David Patterson, "Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore, (2009)"".
[0031] Figure 3 is a table diagram showing an example of the contents of the application performance definition table 1102 in Embodiment 1.
[0032] The application performance definition table 1102 includes a container identifier 3000, an application identifier 3001, an execution thread identifier 3002 (hereinafter also referred to as an execution thread) indicating an execution thread, computational intensity 3003, and memory transfer execution efficiency 3004.
[0033] The container identifier 3000 contains an identifier indicating the container execution environment 1600. If application 1400 is not an application running within the container execution environment 1600, the container identifier 3000 will be an invalid identifier.
[0034] Application identifier 3001 contains an identifier for each container identifier 3000 that indicates an application 1610 running within the container execution environment 1600. If application 1400 is not an application running within the container execution environment 1600, an identifier is listed for each application 1400. For example, Figure 3 illustrates that application 1610, which runs within container execution environment 1600 and is specified by identifier C1 of container identifier 3000, has two identifiers, App1 and App2, derived from application identifier 3001. It also illustrates that the identifier App3 in container identifier 3000 is invalid and is not an application running within container execution environment 1600. Execution thread identifier 3002 contains an identifier that indicates the execution thread for each application. An application's execution threads typically contain one or more. The computational intensity of 3003 indicates the computational intensity for each execution thread of the application. Memory transfer execution efficiency 3004 describes the data transfer execution efficiency between the computing core 1310 and the main memory 1320 for each execution thread of the application.
[0035] Figures 4A and 4B show examples of scheduling information acquired from the scheduling information acquisition unit 1111 in Embodiment 1. The scheduling information includes, at a minimum, for each of the computing cores 0 and 1 that make up the computing core 1310, the timings T1, T2, and T3 for assigning application execution threads to the computing core 1310, the assignment time intervals T1a, T2a, and T3a, the scheduling period P, and time T information.
[0036] For example, Figure 4A illustrates a case where application execution threads do not run simultaneously on multiple computing cores 0 and 1. It shows the scheduling information for application execution threads on computing cores 0 and 1, ensuring that threads running on each of them do not overlap at any given time.
[0037] Furthermore, Figure 4B shows an example where the arithmetic core 1310 is a single arithmetic core 0. For example, it is also possible that only one arithmetic core 0 is operating while the other arithmetic cores are in a low-power sleep state. In this case as well, only one thread will be executed on the arithmetic core at any given time.
[0038] Figure 5 shows the operation flow of the performance power optimization unit 1100 in Embodiment 1.
[0039] In step S5000, the performance power optimization unit 1100 starts the performance power optimization process. In step S5001, the performance power optimization unit 1100 obtains scheduling information for the application's execution threads using the scheduling information acquisition unit 1111. Examples of scheduling information are shown in Figures 4A and 4B mentioned above.
[0040] In step S5002, the performance-power optimization unit 1100, using the roofline optimization calculation unit 1112, performs an optimization calculation of the performance-power of the computer hardware 1300 that is adapted to the computational processing of the execution threads of the application 1400.
[0041] In step S5003, the performance power optimization unit 1100, using the roofline control setting unit 1113, performs performance power control settings for the computer hardware 1300 based on the optimization calculation in step S5002. In step S5004, the performance power optimization unit 1100 completes the performance power optimization process.
[0042] Figure 6 shows the operation flow of the roofline optimization calculation unit 1112 in Embodiment 1.
[0043] In step S6000, the roofline optimization calculation unit 1112 starts the optimization calculation process. In step S6001, the roofline optimization calculation unit 1112 repeats steps S6001 to S6013 for each application execution thread. The application execution thread is the application execution thread described in the scheduling information acquired by the scheduling information acquisition unit 1111 shown in step S5001 of Figure 5.
[0044] In step S6002, the roofline optimization calculation unit 1112 obtains the computation intensity 3003 corresponding to the execution thread identifier 3002 in Figure 3 from the application performance definition table 1102.
[0045] In step S6003, the roofline optimization calculation unit 1112 calculates the intersection point (first intersection point) between the roofline model data 2000 (see Figure 2), which is composed of the maximum controllable computational performance A of the computational core 1310 and the maximum controllable bandwidth performance C of the main memory 1320, and the computational intensity 3003 obtained in step S6002, for the computational core 1310 on which the execution thread operates. If the first intersection point, intersection A1, is in the slope portion of the roofline model data 2000, the process proceeds to step S6004. If intersection B1 is in the roof portion of the roofline model data 2000, the process proceeds to step S6009.
[0046] In step S6004, the roofline optimization calculation unit 1112 calculates a roof that represents the calculation performance value P1 of the intersection A1 in the sloped section of the roofline model data 2000.
[0047] In step S6005, the roofline optimization calculation unit 1112 obtains the memory transfer execution efficiency 3004 corresponding to the execution thread identifier 3002 from the application performance definition table 1102.
[0048] In step S6006, the roofline optimization calculation unit 1112 calculates the bandwidth performance by multiplying the maximum value C of the controllable bandwidth performance of the main memory 1320 by the memory transfer execution efficiency 3004 obtained in step S6005. Note that the example graph of roofline model data 2000 is a log-log graph, and the bandwidth performance is represented by the slope portion of the roofline model data 2000. The intersection point A2 is calculated between the new slope portion based on the calculated bandwidth performance and the calculation intensity 3003 obtained in step S6002.
[0049] In step S6007, the roofline optimization calculation unit 1112 calculates the roof that represents the calculation performance value P2 of intersection A2 obtained in step S6006.
[0050] In step S6008, the roofline optimization calculation unit 1112 sets the control bandwidth performance of the main memory 1320 to its maximum value C, the theoretical value of the control performance of the arithmetic core 1310 to the roof calculation performance value P1 calculated in step S6004, and the lower limit to the roof calculation performance value P2 calculated in step S6007. After the processing steps of step S6008, the process proceeds to step S6013.
[0051] In step S6009, the roofline optimization calculation unit 1112 translates the gradient portion of the example roofline model data 2000 and calculates a new gradient passing through intersection B1 on the roof portion of the roofline model data 2000, and the gradient value of that gradient (which is the theoretical bandwidth value). In step S6010, the roofline optimization calculation unit 1112 obtains the memory transfer execution efficiency 3004 corresponding to the execution thread identifier 3002 from the application performance definition table 1102, similar to step S6005.
[0052] In step S6011, the roofline optimization calculation unit 1112 calculates a bandwidth performance value (referred to as the actual bandwidth value) by multiplying the theoretical bandwidth value calculated in step S6009 by the reciprocal of the memory transfer execution efficiency 3004 obtained in step S6010. However, if the calculated actual bandwidth value exceeds the maximum value C of the controllable bandwidth performance of the main memory 1320, the calculated actual bandwidth value may be set to the maximum value C.
[0053] In step S6012, the roofline optimization calculation unit 1112 sets the control bandwidth performance of the main memory 1320 to the execution bandwidth value calculated in step S6011, and sets the control performance of the arithmetic core 1310 to the maximum controllable arithmetic performance A.
[0054] In step S6013, the roofline optimization calculation unit 1112 determines whether the application's execution thread has terminated. If it has terminated, the process proceeds to step S6014. If it has not terminated, the process returns to step S6001. In step S6014, the roofline optimization calculation unit 1112 terminates the optimization calculation process.
[0055] Figures 7A, 7B, and 7C show examples of roofline optimization calculations when the calculation intensity of the application by the roofline optimization calculation unit 1112 of Embodiment 1 intersects with the slope portion of the roofline model data 2000 in Figure 2.
[0056] The roofline model 7000 shown in Figure 7A is an example where, in step S6003 shown in Figure 6, the intersection point between the roofline model data, which consists of the maximum controllable computational performance A of the computational core 1310 and the maximum controllable bandwidth performance C of the main memory 1320, and the computational intensity 3003 obtained in step S6002 is calculated for the computational core 1310 on which the execution thread operates, and the case where the intersection point A1 is located on the roof portion of the roofline model data.
[0057] The roofline model 7010 shown in Figure 7B is illustrated in step S6004 shown in Figure 6, including the intersection A1 and the dashed arrow 7011 which corresponds to the calculation process including the calculation performance value P1 of intersection A1.
[0058] The roofline model 7020 shown in Figure 7C is illustrated by including the dashed arrow 7021, which corresponds to the process of calculating bandwidth performance and intersection A2 in step S6006 shown in Figure 6, and the arrow 7022, which corresponds to the process of calculating the computational performance value P2 of intersection A2.
[0059] Figures 8A, 8B, and 8C show examples of roofline optimization calculations when the application's computational intensity by the roofline optimization calculation unit 1112 of Embodiment 1 intersects with the roof portion of the roofline model data 2000 in Figure 2.
[0060] The roofline model 8000 shown in Figure 8A illustrates the case where, in step S6003 shown in Figure 6, the intersection point between the roofline model data, which consists of the maximum controllable computational performance A of the computational core 1310 on which the execution thread operates and the maximum controllable bandwidth performance C of the main memory 1320, and the computational intensity 3003 obtained in step S6002 is calculated, and the intersection point B1 is located on the roof portion of the roofline model data.
[0061] The roofline model 8010 shown in Figure 8B illustrates the state after calculating the theoretical bandwidth value in step S6009 shown in Figure 6. The dashed arrow 8011, which represents the theoretical bandwidth value calculation process, is also included in the example.
[0062] The roofline model 8020 shown in Figure 8C illustrates the state after the effective bandwidth value has been calculated in step S6011 shown in Figure 6. The arrow 8021, which corresponds to the calculation process of the effective bandwidth value, is also included in the example.
[0063] Figure 9 shows the operation flow of the roofline control unit 1202 in Embodiment 1. The roofline control unit 1202 may be called and executed immediately before the timing T1 in which the scheduler 1201 assigns the execution thread to the computing core and executes it, as illustrated in Figures 4A and 4B. Alternatively, it may be executed a little earlier, taking into account the processing time of the roofline control unit 1202.
[0064] In step S9000, the roofline control unit 1202 starts the roofline control process. In step S9001, the roofline control unit 1202 obtains the execution thread of the application to be executed from the scheduler 1201.
[0065] In step S9002, the roofline control unit 1202 calculates the control bandwidth performance of the main memory 1320, which is calculated by the roofline optimization calculation unit 1112 corresponding to the execution thread and set by the roofline control setting unit 1113, and then executes the corresponding power control of the main memory 1320 via the memory bandwidth power control unit 1204.
[0066] In step S9003, the roofline control unit 1202 proceeds to step S9004 if there is a control performance value calculated by the roofline optimization calculation unit 1112 and set by the roofline control setting unit 1113 for the arithmetic core 1310 on which the execution thread operates. If the control performance value calculated by the roofline optimization calculation unit 1112 and set by the roofline control setting unit 1113 for the arithmetic core 1310 on which the execution thread operates is set to a lower limit or theoretical value for control performance, the process proceeds to step S9005.
[0067] In step S9004, the roofline control unit 1202 sets the control performance of the arithmetic core 1310 on which the execution thread operates to the control performance value obtained in step S9003, and performs power control of the arithmetic core 1310 according to that control performance value via the arithmetic core performance power control unit 1203. For example, dynamic voltage and frequency scaling (DVFS) may be used for the performance power control of the arithmetic core 1310. After the processing step S9004, the process proceeds to step S9006.
[0068] In step S9005, the roofline control unit 1202 sets the control performance of the arithmetic core 1310 on which the execution thread operates to the theoretical value or lower limit of the control performance obtained in step S9003, and performs power control of the arithmetic core 1310 according to that control performance value via the arithmetic core performance power control unit 1203. If there is a lower limit of control performance in step S9003, the lower limit may be used.
[0069] After processing in step S9004 or step S9005, the process proceeds to step S9006, in which step S9006 the roofline control unit 1202 terminates the roofline control process.
[0070] Embodiment 2. Figure 10 shows an example of information acquired from the scheduling information acquisition unit 1111 in Embodiment 2. The difference from Figure 4 is that the application's execution threads run concurrently on each computing core. For example, Figure 10 illustrates the scheduling information for application execution threads on computing core 0 and computing core 1, where execution threads running on computing core 0 and computing core 1 overlap at the same time.
[0071] Figure 11 is a diagram showing the operation flow of the roofline optimization calculation unit in Embodiment 2. The difference from Figure 6 is that the application's execution threads run concurrently on each computing core. In step S11000, the roofline optimization calculation unit 1112 starts the optimization calculation process. In step S11001, the roofline optimization calculation unit 1112 repeats steps S11001 to S11012 for each set of application execution threads that are running simultaneously on each calculation core 1310.
[0072] In step S11002, the roofline optimization calculation unit 1112 obtains the calculation intensity 3003 for each pair of execution threads from the application performance definition table 1102.
[0073] In step S11003, the roofline optimization calculation unit 1112 calculates the bandwidth performance of the main memory 1320 required by each execution thread. The method for calculating the bandwidth performance required by an execution thread is, for example, to calculate the intersection point between the roofline model data 2000, which is composed of the maximum controllable computational performance A of the computational core 1310 on which the execution thread operates and the maximum controllable bandwidth performance C of the main memory 1320, and the computational intensity 3003 obtained in step S11002. If the intersection point is in the slope portion of the roofline model data 2000, that slope value may be used as the bandwidth performance value required by the execution thread. If the intersection point is located on the roof portion of the roofline model data 2000, the slope portion of the roofline model data 2000 is translated, a new slope passing through the intersection point on the roof portion of the roofline model data 2000 is calculated, and this slope value is multiplied by the reciprocal of the memory transfer execution efficiency 3004 of the execution thread to calculate a new slope value, which may be used as the bandwidth performance value requested by the execution thread. However, if this value exceeds the maximum controllable bandwidth performance C of the main memory 1320, the bandwidth performance value requested by the execution thread may be set to the maximum value C.
[0074] In step S11004, the roofline optimization calculation unit 1112 determines whether the sum of the bandwidth performance of the main memory 1320 requested by each execution thread, calculated in step S11003, is greater than the maximum controllable bandwidth performance C of the main memory 1320. If it is greater, the unit proceeds to step S11005. If it is less than the sum of the bandwidth performance C, the unit proceeds to step S11006.
[0075] In step S11005, the roofline optimization calculation unit 1112 calculates the bandwidth ratio for each execution thread relative to the total bandwidth performance of the main memory 1320 requested by each execution thread, which was calculated in step S11003. For example, the method for calculating the bandwidth ratio for each execution thread may be to calculate the ratio of the bandwidth performance requested by each execution thread to the total bandwidth performance requested by each execution thread, which was calculated in step S11003.
[0076] In step S11006, the roofline optimization calculation unit 1112 repeats steps S11006 to S11008 for each execution thread. In step S11007, the roofline optimization calculation unit 1112 calculates the control performance of the arithmetic core 1310 on which the execution thread operates.
[0077] In step S11008, the roofline optimization calculation unit 1112 determines whether the execution thread has terminated or not. If it has terminated, it proceeds to step S11009. If it has not terminated, it returns to step S11006.
[0078] In step S11009, the same determination as in step S11004 is made. If the sum of the bandwidth performance of the main memory 1320 requested by each execution thread, calculated in step S11003, is greater than the maximum controllable bandwidth performance C of the main memory 1320, the process proceeds to step S11010. Otherwise, the process proceeds to step S11011.
[0079] In step S11010, the roofline optimization calculation unit 1112 sets the control bandwidth performance of the main memory 1320 to its maximum value C. In step S11011, the roofline optimization calculation unit 1112 sets the control bandwidth performance of the main memory 1320 to the sum of the bandwidth performance of the main memory 1320 requested by each execution thread, which was calculated in step S11003.
[0080] In step S11012, the roofline optimization calculation unit 1112 determines whether a pair of application execution threads running simultaneously on each processing core 1310 has finished. If it has finished, the process proceeds to step S11013. If it has not finished, the process returns to step S11001. In step S11013, the roofline optimization calculation unit 1112 terminates the optimization calculation process.
[0081] Figure 12 is a diagram showing the operation flow of the control performance calculation process included in the roofline optimization calculation unit of Embodiment 2. In step S12000, the roofline optimization calculation unit 1112 starts the control performance calculation process for the arithmetic core 1310. In step S12001, the roofline optimization calculation unit 1112 obtains the memory transfer execution efficiency 3004 corresponding to the execution thread identifier 3002 from the application performance definition table 1102.
[0082] In step S12002, the roofline optimization calculation unit 1112 calculates the intersection point between the roofline model data 2000, which is composed of the maximum controllable computational performance A of the computational core 1310 on which the execution thread operates, and the maximum controllable bandwidth performance C of the main memory 1320, and the computational intensity 3003 obtained in step S11002. If intersection point A1a is in the slope portion of the roofline model data 2000, the process proceeds to step S12003. If intersection point B1a is in the roof portion of the roofline model data 2000, the process proceeds to step S12008.
[0083] In step S12003, the roofline optimization calculation unit 1112 calculates a roof that represents the calculation performance value P1a of the intersection A1a in the sloped section of the roofline model data 2000.
[0084] In step S12004, the roofline optimization calculation unit 1112 calculates a bandwidth performance value by multiplying the maximum controllable bandwidth performance C of the main memory 1320 by the bandwidth ratio of the execution thread calculated in step S11005 shown in Figure 11 and the memory transfer execution efficiency 3004 corresponding to the execution thread obtained in step S12001. The unit calculates the intersection point A2a between the gradient portion of the roofline model data 2000, which is the calculated bandwidth performance value, and the calculation intensity 3003 obtained in step S11002.
[0085] In step S12005, the roofline optimization calculation unit 1112 calculates the roof that represents the calculation performance value P2a of the intersection A2a obtained in step S12004. In step S12006, the roofline optimization calculation unit 1112 sets the theoretical value of the control performance of the calculation core 1310 to the calculation performance value P1a for the roof calculated in step S12003.
[0086] In step S12007, the roofline optimization calculation unit 1112 sets the lower limit of the control performance of the calculation core 1310 to the calculation performance value P2a for the roof calculated in step S12005. After the processing step S12007, the process proceeds to step S12014.
[0087] In step S12008, the roofline optimization calculation unit 1112 determines, as in step S11004 illustrated in Figure 11, whether the total bandwidth performance of the main memory 1320 requested by the execution threads is greater than the maximum controllable bandwidth performance C of the main memory 1320. If it is greater, the unit proceeds to step S12009. If it is less than the value of the main memory 1320, the unit proceeds to step S12013.
[0088] In step S12009, the roofline optimization calculation unit 1112 calculates a bandwidth performance value by multiplying the maximum controllable bandwidth performance C of the main memory 1320 by the bandwidth ratio of the execution thread calculated in step S11005 shown in Figure 11. The unit then calculates the intersection point B2a between the gradient portion of the roofline model data 2000, which is the calculated bandwidth performance value, and the calculation intensity 3003 obtained in step S11002.
[0089] In step S12010, the roofline optimization calculation unit 1112 calculates a roof that represents the calculation performance value Q2a of the intersection B2a calculated in step S12009. In step S12011, the roofline optimization calculation unit 1112 determines whether the calculation performance value Q2a calculated in step S12010 is less than the maximum controllable calculation performance A of the calculation core 1310. If it is less, the process proceeds to step S12012. If it is greater, the process proceeds to step S12013.
[0090] In step S12012, the roofline optimization calculation unit 1112 sets the lower limit of the control performance of the calculation core 1310 to the calculation performance value Q2a for the roof calculated in step S12010. After the processing steps of step S12012, the process proceeds to step S12014.
[0091] In step S12013, the roofline optimization calculation unit 1112 sets the control performance value of the calculation core 1310 to the maximum value A of the controllable calculation performance of the calculation core 1310, and sets the calculation performance value to Q1a. After the processing steps of step S12013, the process proceeds to step S12014. In step S12014, the roofline optimization calculation unit 1112 completes the control performance calculation process of the arithmetic core 1310.
[0092] Figures 13A, 13B, 13C, and 13D show examples of roofline optimization calculations performed by the roofline optimization calculation unit 1112 of Embodiment 2.
[0093] The roofline model 13000 shown in Figure 13A exemplifies a case in step S11003 shown in Figure 11 where, in the process of calculating the bandwidth performance value of the main memory 1320 requested by the execution thread, the intersection point A1a, which is the second intersection point between the roofline model data 2000, which is composed of the maximum controllable computational performance A of the computational core 1310 and the maximum controllable bandwidth performance C of the main memory 1320, and the computational intensity 3003 obtained in step S11002, is located in the slope portion of the roofline model data 2000.
[0094] The roofline model 13010 shown in Figure 13B illustrates a case in step S11003 shown in Figure 11 where, in the process of calculating the bandwidth performance value of the main memory 1320 requested by the execution thread, the intersection point B1a of the roofline model data 2000, which is composed of the maximum controllable computational performance A of the computational core 1310 and the maximum controllable bandwidth performance C of the main memory 1320, and the computational intensity 3003 obtained in step S11002, is located on the roof portion of the roofline model data 2000. The dashed arrow 13011, which corresponds to the process of calculating the bandwidth performance value requested by the execution thread, is also included as an example.
[0095] The roofline model 13020 shown in Figure 13C is illustrated by including the dashed arrow 13021, which corresponds to the calculation process of the roof showing the calculation performance value P1a at intersection A1a in step S12003 shown in Figure 12; the dashed arrow 13022, which corresponds to the calculation process of intersection A2a in step S12004; and the dashed arrow 13023, which corresponds to the calculation process of the roof showing the calculation performance value P2a at intersection A2a in step S12005.
[0096] The roofline model 13030 shown in Figure 13D includes the dashed arrow 13031, which corresponds to the calculation process of intersection B2a in step S12009 shown in Figure 12.
[0097] Figure 14 shows the operation flow of the execution thread acquisition process in the roofline control unit 1202 of Embodiment 2. The difference from Figure 9 is that the process of acquiring the execution thread in step S9001 corresponds to the fact that the application's execution threads are executed simultaneously on each processing core 1310. The operation flow of the roofline control unit 1202 is the same as step S9001 in Figure 9, but replaced with steps S14000 to S14006 shown in Figure 14.
[0098] In step S14000, the roofline control unit 1202 starts the process of acquiring execution threads.
[0099] In step S14001, the roofline control unit 1202 determines whether or not it is operating in response to a roofline control interrupt. If it is operating in response to an interrupt, the process proceeds to step S14005. If it is not operating in response to an interrupt, the process proceeds to step S14002. The roofline control interrupt is issued in step S14004.
[0100] In step S14002, the roofline control unit 1202 obtains the execution thread of the application to be assigned to its own computing core 1310 from the scheduler 1201.
[0101] In step S14003, the roofline control unit 1202 determines from the scheduler 1201 whether there are any execution threads running on other arithmetic cores 1310 besides itself. If there are execution threads running on other arithmetic cores 1310, the unit proceeds to step S14004. Otherwise, the unit proceeds to step S14006.
[0102] In step S14004, the roofline control unit 1202 issues an interrupt to the other computing cores 1310 to execute the roofline control unit.
[0103] In step S14005, the roofline control unit 1202 obtains the set of execution threads currently running on each processing core 1310. According to the process illustrated in step S14005, each processing core 1310 obtains the set of execution threads for applications running at the same time. In step S14006, the roofline control unit 1202 terminates the execution thread acquisition process.
[0104] After completing the processing step S14006, the processing steps S9002 to S9006 in Figure 9 are performed. The processing steps S9002 to S9006 in Figure 9 are performed based on the control bandwidth performance of the main memory 1320 and the control performance of the arithmetic core 1310 on which each execution thread operates, calculated for each set of execution threads of an application running simultaneously on each arithmetic core 1310 as illustrated in Figure 11. The control bandwidth performance and power control of the main memory 1320 are performed accordingly, as well as the control performance and power control of the arithmetic core 1310 on which each execution thread operates. The control bandwidth performance and power control of the main memory 1320 may be performed on any of the arithmetic cores 1310. Furthermore, the control performance and power control of the arithmetic core 1310 on which each execution thread operates may be performed on each individual arithmetic core 1310.
[0105] Although this application describes various exemplary embodiments and examples, the various features, aspects, and functions described in one or more embodiments are not limited to the application of a particular embodiment, but can be applied individually or in various combinations to the embodiments. Accordingly, countless variations not illustrated are conceivable within the scope of the technology disclosed herein. These include, for example, modifications, additions, or omissions of at least one component, as well as the extraction of at least one component and its combination with components of other embodiments. [Explanation of Symbols]
[0106] 1000 Information processing device, 1100 Performance-power optimization unit, 1101 Roofline model data, 1102 Application performance definition table, 1110 Performance-power optimization program, 1111 Scheduling information acquisition unit, 1112 Roofline optimization calculation unit, 1113 Roofline control setting unit, 1200 System software, 1201 Scheduler, 1202 Roofline control unit, 1203 Computation core performance-power control unit, 1204 Memory bandwidth-power control unit, 1300 Computer hardware, 1310 Computation core, 1320 Main memory, 1400 Application
Claims
1. Computer hardware including multiple computing cores and main memory equipped with power saving mechanisms, System software running on the aforementioned computer hardware, The system software and an application that runs in the container execution environment of the system software, In an information processing device including, The information processing device is equipped with a performance-power optimization unit that performs optimization processing of the computational performance and power consumption of the information processing device, The performance-power optimization unit includes a performance-power optimization program, roofline model data indicating the computational intensity and computational performance per unit time of the computer hardware, and an application performance definition table including computational intensity information of the application. The performance-power optimization program includes a scheduling information acquisition unit that acquires scheduling information for applications to the computing cores of the computer hardware, and a roofline optimization calculation unit that performs optimization calculations of the roofline of the computer hardware, calculates the bandwidth performance that each thread requests from the main memory from the roofline model data for each computing core when there are execution threads running simultaneously in the multiple computing cores, and calculates the bandwidth performance to be allocated to each thread based on the ratio of each bandwidth performance to the sum of the calculated bandwidth performances. An information processing device characterized by comprising:
2. The information processing apparatus according to claim 1, characterized in that the performance power optimization program includes a roofline control setting unit that performs performance power control settings for the computer hardware to the system software based on the optimization calculation of the roofline.
3. The aforementioned system software is A scheduler that assigns and runs one or more of the aforementioned applications on one or more computing cores, Based on the setting information for the roofline control settings, a roofline control unit controls the roofline of the computer hardware, A computing core performance power control unit that controls the performance and power of the computing core, A memory bandwidth-power control unit that controls the bandwidth and power of the main memory, The information processing apparatus according to claim 1, characterized by including the following:
4. The information processing apparatus according to claim 1, characterized in that the roofline model data includes roofline model data for each processing core.
5. The roofline model data is, in the roofline model data for each computing core, The maximum computing performance is the roofline that indicates the highest computing performance, The minimum computational performance value is the roofline that shows the minimum computational performance, The maximum memory bandwidth is the slope that indicates the maximum memory bandwidth, The minimum memory bandwidth is the slope that indicates the minimum memory bandwidth, and The information processing apparatus according to claim 1, characterized by comprising:
6. The aforementioned application performance definition table is: A container identifier indicating the container execution environment, An application identifier indicating an application running in the aforementioned container execution environment, An execution thread identifier indicating the execution thread of the aforementioned application, The computational intensity of the execution thread of the aforementioned application, The memory transfer execution efficiency of the execution thread of the aforementioned application, The information processing apparatus according to claim 1, characterized by including the following:
7. A scheduling information acquisition step for acquiring scheduling information for application execution threads to a computing core of computer hardware including a computing core equipped with a power saving mechanism and main memory, A roofline optimization calculation step that performs an optimization calculation of the roofline of the aforementioned computer hardware, calculates the bandwidth performance that each thread requests from the main memory from the roofline model data for each computing core when multiple computing cores are executing their respective execution threads at the same time, and calculates the bandwidth performance to be allocated to each thread based on the ratio of each bandwidth performance to the sum of the calculated bandwidth performances, An information processing method characterized by including a roofline control setting step for setting control settings for the performance power of the computer hardware with respect to system software running on the computer hardware.
8. The aforementioned scheduling information acquisition step is: The information processing method according to claim 7, characterized by including the step of obtaining information including the activation status of all of the computing cores of the computer hardware and whether or not the execution threads of the application are running simultaneously on multiple computing cores.
9. The roofline optimization calculation step is: When the execution threads of the aforementioned application are not running simultaneously on multiple computing cores, For each execution thread of the aforementioned application, the step of obtaining the computational intensity from an application performance definition table which includes computational intensity information of the aforementioned application, A step of finding a first intersection point between roofline model data, which is composed of the maximum controllable computational performance of the computation core and the maximum controllable bandwidth performance of the main memory, and the computational intensity. When the first intersection point is the slope portion of the roofline model data, the first intersection point is defined as intersection point A1, and the roof is calculated to show the calculation performance value P1 of intersection point A1. The steps include obtaining the memory transfer execution efficiency of the application's execution threads from the application performance definition table, The steps include: calculating the bandwidth performance by multiplying the maximum controllable bandwidth performance of the main memory by the memory transfer execution efficiency, and calculating the intersection point A2 of the slope portion of the roofline model data and the calculation intensity based on the bandwidth performance; The steps include: calculating the roof that represents the calculation performance value P2 of the intersection A2; The step includes setting the control bandwidth performance of the main memory to its maximum value, the theoretical value of the control performance of the arithmetic core to the arithmetic performance value P1, and the lower limit of the control performance of the arithmetic core to the arithmetic performance value P2, When the first intersection point corresponds to the roof of the roofline model data, the first intersection point is defined as intersection point B1, and in the log-log graph of the roofline model data, the gradient portion indicating the maximum controllable bandwidth performance of the main memory is shifted parallel to obtain a new gradient portion passing through intersection point B1 and a gradient value that is the theoretical bandwidth value of that gradient. The gradient value obtained by multiplying the calculated gradient value by the reciprocal of the memory transfer execution efficiency is the effective bandwidth value. The steps to calculate, The steps include setting the control bandwidth performance of the main memory to the effective bandwidth value and the control performance of the computing core to the maximum value of the controllable computing performance, The information processing method according to claim 7, characterized by including the following:
10. The roofline control setting step is, When the execution threads of the aforementioned application are not running simultaneously on multiple computing cores, The steps include obtaining an execution thread for the application to be assigned and executed on the aforementioned computing core, The steps include controlling the control bandwidth performance of the main memory corresponding to the execution thread, and controlling the power of the main memory according to the control bandwidth performance, Includes, If there is a control performance value for the processing core, the process includes the steps of controlling the control performance of the processing core to the control performance value and controlling the power of the processing core to a level corresponding to the control performance. If there is a theoretical value for the control performance of the aforementioned computing core, the steps include: controlling the control performance of the computing core to the theoretical value, and controlling the power of the computing core to a level corresponding to the control performance. or, The information processing method according to claim 7, characterized in that, if there is a lower limit to the control performance of the processing core, the method includes the steps of controlling the control performance of the processing core to the lower limit and controlling the power of the processing core to a level corresponding to the control performance.
11. The roofline optimization calculation step is as follows: When the execution threads of the application are executed simultaneously on multiple computing cores, the method includes the steps of obtaining the computational intensity of each set of execution threads of the application from an application performance definition table containing computational intensity information of the application, for each set of execution threads of the application that are executed simultaneously on multiple computing cores, and calculating the bandwidth performance that each execution thread requests from the main memory. If the calculated total bandwidth performance requested by each execution thread from the main memory is greater than the maximum controllable bandwidth performance of the main memory, the process includes the steps of calculating the bandwidth ratio of each execution thread of the application to the total bandwidth performance, and setting the controllable bandwidth performance of the main memory to its maximum value. The process includes the step of setting the controllable bandwidth performance of the main memory to the sum of the bandwidth performance requested by each of the execution threads that have calculated the main memory, if the sum of the bandwidth performance requested by each of the execution threads that have calculated the main memory is not greater than the maximum controllable bandwidth performance of the main memory. The information processing method according to claim 7, comprising a calculation step of calculating the control performance of the arithmetic core for each execution thread of the application.
12. In the calculation process for calculating the control performance of the arithmetic core when the execution threads of the application are executed simultaneously on multiple arithmetic cores, For each execution thread of the application, the steps include: obtaining the memory transfer execution efficiency of the execution thread of the application from an application performance definition table including computational intensity information of the application; and finding a second intersection point between roofline model data, which consists of the maximum controllable computational performance of the computation core on which the execution thread operates and the maximum controllable bandwidth performance of the main memory, and the computational intensity of the execution thread. The steps include: when the second intersection point is the slope portion of the roofline model data, the second intersection point is defined as intersection point A1a, and a calculation of the roof showing the computational performance value P1a of the intersection point A1a; a calculation of a bandwidth performance value obtained by multiplying the bandwidth performance of the main memory requested by the execution thread by the calculated bandwidth ratio of the execution thread and the acquired memory transfer execution efficiency; a calculation of the intersection point A2a between the slope portion of the roofline model data showing the calculated bandwidth performance value and the acquired computational strength of the execution thread; a calculation of the roof showing the computational performance value P2a of the intersection point A2a; and setting the theoretical value of the control performance of the computation core as the computational performance value P1a and the lower limit of the control performance of the computation core as the computational performance value P2a. The steps include: when the second intersection point is the roof of the roofline model data, the second intersection point is set as intersection point B1a, and if the calculated total bandwidth performance of the main memory requested by the execution thread is greater than the maximum controllable bandwidth performance of the main memory, a bandwidth performance value is calculated by multiplying the calculated bandwidth performance of the main memory requested by the execution thread by the calculated bandwidth ratio of the execution thread; a step of calculating the intersection point B2a between the gradient portion of the roofline model data showing the calculated bandwidth performance value and the computational intensity of the acquired execution thread; a step of calculating the roof showing the computational performance value Q2a at intersection point B2a; and if the calculated computational performance value Q2a is less than the maximum controllable computational performance of the computation core, the lower limit of the controllable performance of the computation core is set as the computational performance value Q2a. The information processing method according to claim 7, comprising the step of setting the controllable performance of the arithmetic core to the maximum controllable performance if the calculated total bandwidth performance of the main memory requested by the execution thread is less than the maximum controllable bandwidth performance of the main memory, or if the calculated computation performance value Q2a is greater than the maximum controllable computation performance of the arithmetic core.
13. In the process of obtaining information on the execution threads of the roofline control process, if the execution threads of the application are running simultaneously on multiple computing cores, the process includes the steps of determining whether or not they are operating triggered by a roofline control interrupt, and, if they are operating triggered by a roofline control interrupt, obtaining a set of execution threads running on each of the computing cores. If the operation is not triggered by the roofline control interrupt, the process includes the steps of obtaining the execution thread of the application to be assigned to and executed on the own computing core, and determining whether or not there are any other execution threads running on other computing cores. The information processing method according to claim 7, characterized in that, if there are execution threads running on other arithmetic cores, it includes the steps of issuing an interrupt to other arithmetic cores other than itself to perform the roofline control, and obtaining a set of execution threads running on each of the arithmetic cores.
Citation Information
Patent Citations
System and method for controlling central processing unit power with guaranteed transient deadlines
JP2016511880A
Method to determine disposition for VM / container and volume in HCI environment, and storage system
JP2020052730A
Processor for controlling arithmetic capacity
WO2008120274A1
Information processing system and information processing system control method
WO2021250737A1