vRAN system and accelerator offloading method
Patent Information
- Application Number
- JP2024528071
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2022-06-17
- Filing Date
- 2022-06-17
- Publication Date
- 2025-09-30
AI Technical Summary
Existing methods for processing high-load calculations using CPUs result in performance bottlenecks and excessive power consumption due to inefficient processing by the CPU alone or continuous operation of accelerators, even when data processing is low.
An accelerator offload device dynamically switches processing between a CPU and an accelerator based on data processing amounts, using a data processing amount acquisition unit, calculation destination determination unit, arithmetic processing offload unit, and ACC power supply control unit to optimize processing efficiency and power usage without modifying application programs.
This solution minimizes power consumption and resolves performance bottlenecks by dynamically switching processing between the CPU and accelerator, ensuring efficient use of resources and maintaining application program integrity.
Abstract
Description
Accelerator offload device, accelerator offload method, and program
[0001] The present invention relates to an accelerator offload device, an accelerator offload method, and a program.
[0002] Depending on the type of processor, different workloads are suited to different tasks (high processing power). While general-purpose central processing units (CPUs) are capable of quickly and efficiently processing highly parallel workloads that CPUs are not good at (low processing power), accelerators (hereinafter referred to as ACCs) such as field programmable gate arrays (FPGAs), graphics processing units (GPUs), and application-specific integrated circuits (ASICs) are available. By combining these heterogeneous processors and offloading workloads that CPUs are not good at (low processing power) to the ACCs, offloading technology is being increasingly utilized to improve overall processing time and efficiency.
[0003] Typical examples of workloads for which ACC offloading is performed include encoding / decoding processing (FEC: Forward Error Correction processing) in a virtual Radio Access Network (vRAN), audio and video media processing, and encryption / decryption processing.
[0004] 25 is a schematic diagram of a process of offloading part of a process performed by a CPU to an accelerator (ACC). As shown in FIG. 25, the accelerator system includes hardware (HW) 10, an OS (Operating System) 20, and an application (APL) 1.
[0005] The hardware 10 includes a CPU 11 and an accelerator (ACC) 12. The ACC 12 is a calculation unit hardware that performs specific calculations at high speed based on input from the CPU 11. Specifically, the accelerator 12 is a PLD (Programmable Logic Device) such as a GPU or FPGA.
[0006] The CPU 11 offloads part of the processing of the APL 1 (a workload that the CPU 11 is not good at) to the ACC 12, thereby achieving performance and power efficiency that cannot be achieved by software (CPU processing) alone.
[0007] Patent document 1 describes a control device that includes a communication unit that receives packets from a network, a plurality of first control units that function as a plurality of virtual control units, a distribution circuit that distributes the received packets to a plurality of groups, and a plurality of second control units that distribute the packets distributed by the distribution circuit to the plurality of virtual control units.
[0008] Non-Patent Document 1 describes the CUDA Toolkit (registered trademark), a development environment for creating high-performance GPU-accelerated applications. The CUDA Toolkit enables developers to develop, optimize, and deploy applications on GPU-accelerated embedded systems, desktop workstations, enterprise data centers, cloud-based platforms, and HPC supercomputers.
[0009] Japanese Patent Application Laid-Open No. 2019-153019
[0010] CUDA, [online], [Retrieved May 11, 2022], Internet <https: / / developer.nvidia.com / cuda-toolkit>
[0011] When performing processing including highly parallel operations or high-load operations, there are two methods that can be considered in the prior art: (1) a method in which the operations are performed by the CPU alone, and (2) a method in which part of the processing is offloaded to the ACC 12. In the (1) method in which the operations are performed by the CPU alone, all processing is performed by the CPU alone, eliminating the need to use the ACC 12 and reducing power consumption throughout the system. On the other hand, when the amount of data processing is large, processing that is inefficiently performed by the CPU 11 becomes a bottleneck, limiting processing performance. In the (2) method in which part of the processing is offloaded to the ACC 12, the ACC 12 performs processing that is inefficiently performed by the CPU 11, thereby eliminating bottlenecks and improving performance. On the other hand, there is power consumption overhead due to the ACC 12 always operating regardless of the amount of data processing.
[0012] Figure 26 is a diagram illustrating the problems with the method (1) of performing calculations by a single CPU. As shown in Figure 26, the CPU 11 performs calculations for process 1 in step S1, process 2 in step S2, the offload target process in step S10, and process 3 in step S3 by itself without using the ACC. Because the CPU 11 does not use the ACC, no power consumption overhead occurs in the ACC. However, because the CPU 11 performs offload target processes, which it is not good at, by itself, a performance bottleneck occurs when the amount of data processing is large (see symbol a in Figure 26).
[0013] FIG. 27 is a diagram illustrating the issues with the method (2) of offloading part of processing to the ACC 12. As shown in FIG. 27 , after processing 1 in step S1 and processing 2 in step S2, the CPU 11 offloads the offload target processing (step S10) to the ACC 12. The CPU 11 receives the processing result from the ACC 12 and performs the calculation of processing 3 in step S3. The high-speed processing by the ACC 12 can eliminate the performance bottleneck of the CPU 11 (see symbol b in FIG. 27 ). However, the ACC 12 always operates even when there is no data processing or when the processing volume is low, generating power consumption overhead (see symbol c in FIG. 27 ). Thus, in processing including high-load calculations, the issue is to minimize the power consumption overhead while suppressing the performance bottleneck caused by the CPU.
[0014] The present invention has been made in light of this background, and its objective is to minimize power consumption overhead while suppressing performance bottlenecks caused by the CPU without modifying application programs.
[0015] In order to solve the above-mentioned problems, an accelerator offload device that offloads specific processing of an application program to an accelerator is provided, comprising: a data processing amount acquisition unit that acquires the current data processing amount and / or the predicted data processing amount of the processing to be offloaded; an operation destination determination unit that determines whether to perform the operation in a CPU or offload the operation to the accelerator based on the data processing amount acquired by the data processing amount acquisition unit, and changes the offload destination if a change is necessary; and an operation processing offload unit that has a shared memory between the CPU or the accelerator that executes parallel operations, receives an operation processing request from the application program, stores the data to be processed and a processing command in the shared memory, and causes the CPU or the accelerator to execute the corresponding processing.
[0016] According to the present invention, it is possible to minimize power consumption overhead while suppressing performance bottlenecks caused by the CPU without modifying the application program.
[0017] FIG. 1 is a schematic configuration diagram of an accelerator offload system according to a first embodiment of the present invention. FIG. 2 is a diagram illustrating features and an operational image of the accelerator offload device according to the first embodiment of the present invention. FIG. 3 is an operational explanatory diagram for explaining dynamic switching between ACC and CPU in the accelerator offload device according to the first embodiment of the present invention. FIG. 4 is a flowchart illustrating data processing operations during APL operation in the accelerator offload device according to the first embodiment of the present invention. FIG. 5 is an operational explanatory diagram of the accelerator offload system from when a data processing amount acquisition unit in the accelerator offload device according to the first embodiment of the present invention acquires a processing amount and passes it to a computation destination determination unit. FIG. 6 is a flowchart illustrating operations from when a data processing amount acquisition unit in the accelerator offload device according to the first embodiment of the present invention acquires a processing amount and passes it to a computation destination determination unit. FIG. 7 is an operational explanatory diagram of operations until an offload target processing execution location is switched (ACC → CPU) in the accelerator offload device according to the first embodiment of the present invention. FIG. 8 is a flowchart illustrating operations until an offload target processing execution location is switched (ACC → CPU) in the accelerator offload device according to the first embodiment of the present invention. FIG. 9 is an operational explanatory diagram of operations until an offload target processing execution location is switched (CPU → ACC) in the accelerator offload device according to the first embodiment of the present invention. FIG. 1 is a flowchart showing the operation of the accelerator offload device according to a first embodiment of the present invention until the offload target processing execution location is switched (CPU → ACC). FIG. 2 is a schematic configuration diagram of an accelerator offload system according to a second embodiment of the present invention. FIG. 3 is an explanatory diagram of the operation until the data processing amount acquisition unit of the accelerator offload device according to the second embodiment of the present invention acquires information and passes it to the computation destination determination unit. FIG. 4 is a flowchart showing the operation until the data processing amount acquisition unit of the accelerator offload device according to the second embodiment of the present invention acquires information and passes it to the computation destination determination unit 120. FIG. 5 is an explanatory diagram of the operation until the offload target processing execution location of the accelerator offload device according to the second embodiment of the present invention is switched (ACC → CPU). FIG. 6 is a flowchart showing the operation until the offload target processing execution location of the accelerator offload device according to the second embodiment of the present invention is switched (ACC → CPU).1 is an explanatory diagram of the operation of an accelerator offload device according to a second embodiment of the present invention until the offload target processing execution location is switched (CPU → ACC). FIG. 2 is a flowchart showing the operation of an accelerator offload device according to the second embodiment of the present invention until the offload target processing execution location is switched (CPU → ACC). FIG. 3 is a schematic configuration diagram of an accelerator offload system according to a third embodiment of the present invention. FIG. 4 is an explanatory diagram of the operation of an accelerator offload device according to a third embodiment of the present invention when the number of execution processing cores is changed (while the offload target processing is being executed by the CPU). FIG. 5 is an explanatory diagram of the operation of an accelerator offload device according to the third embodiment of the present invention until the offload target processing execution location is switched (CPU → ACC). FIG. 6 is a flowchart showing the operation of an accelerator offload device according to the third embodiment of the present invention until the offload target processing execution location is switched (ACC → CPU). FIG. 7 is an explanatory diagram of the operation of an accelerator offload device according to the third embodiment of the present invention until the offload target processing execution location is switched (ACC → CPU). FIG. 8 is a flowchart showing the operation of an accelerator offload device according to the third embodiment of the present invention <until the offload target processing execution location is switched (ACC → CPU)>. FIG. 9 is a hardware configuration diagram showing an example of a computer that realizes the functions of the accelerator offload device in the accelerator offload system according to each embodiment of the present invention. It is a schematic diagram of a process in which a part of a process performed by a CPU is offloaded to an accelerator (ACC). It is a diagram explaining a problem of a method in which a calculation is performed by a CPU alone. It is a diagram explaining a problem of a method in which a part of a process is offloaded to an ACC.
[0018] An accelerator offload system and the like in an embodiment of the present invention (hereinafter referred to as "the present embodiment") will be described below with reference to the drawings. (First Embodiment) [Overall Configuration] Fig. 1 is a schematic configuration diagram of an accelerator offload system according to a first embodiment of the present invention. Components that are the same as those in Fig. 21 are assigned the same reference numerals. As shown in Fig. 1, an accelerator offload system 1000 includes hardware (HW) 10, an OS 20, a high-speed data communication unit 30 that is high-speed data transfer middleware, an accelerator offload device 100, and an APL 1.
[0019] The hardware 10 includes a CPU 11, an accelerator (ACC) 12, and a ring buffer 13. The ACC 12 is a calculation unit hardware that performs specific calculations at high speed based on input from the CPU 11. Specifically, the accelerator 12 is a PLD such as a GPU or FPGA. The ring buffer 13 is provided in the hardware 10 and copies target workloads. A calculation processing offload unit 130 (described below) of the accelerator offload device 100 exchanges data with the accelerator via the ring buffer 13.
[0020] [High-Speed Data Communication Unit 30] The high-speed data communication unit 30 is a high-speed data communication layer including CUDA, OpenCL BBDEV API, and the like. For example, the high-speed data communication unit 30 is the CUDA Toolkit (registered trademark) for using an NVIDIA (registered trademark) GPU, or OpenCL (registered trademark) for calculations using a heterogeneous processor. The BBDEV API (registered trademark) also provides accelerator I / O functions for processing wireless access signals as a Development Kit (library). The high-speed data communication unit 30 can provide the accelerator I / O functions for processing wireless access signals to the APL1 by incorporating the accelerator I / O functions provided as libraries by the CUDA, OpenCL BBDEV API, and the like into the APL1.
[0021] [Accelerator Offload Device 100 ] The accelerator offload device 100 includes a data processing amount acquisition unit 110 , a computation destination determination unit 120 , a computation offload unit 130 , and an ACC power supply control unit 140 .
[0022] The accelerator offload device 100 has the following features: it dynamically switches the parallel computation processing execution unit (a general term for a processing unit that executes computations using either the CPU or the ACC) between the ACC 12 and the CPU 11 (feature <1>); when switching the parallel computation processing execution unit, there is no modification of the APL 1 or interruption of the APL 1 that is currently running (feature <2>); when the amount of data processing is small, processing is performed by the CPU 11, reducing the power consumption of the ACC 12 to minimize power overhead; and when the ACC 12 is not operating, the power consumption of the ACC 12 is minimized (feature <3>).
[0023] Feature <1> eliminates the performance bottleneck of the CPU 11 through high-speed processing by the ACC 12. Feature <2> is that switching is transparent, does not require modification of APL1, and does not interrupt APL1 during switching. Feature <3> is that when the amount of data processing is low, processing is performed by the CPU 11, reducing the power consumption of the ACC 12 and minimizing power overhead.
[0024] <Data processing amount acquisition unit 110> As a function of feature <1>, the data processing amount acquisition unit 110 acquires the current data processing amount and / or the predicted data processing amount of the offload target processing, and notifies the calculation destination determination unit 120 (<Data processing amount notification>).
[0025] Examples of the information to be acquired include (1) the amount of parallel computation data per second [Byte] and (2) the CPU utilization rate (when parallel computation is performed by the CPU 11). Examples of the acquisition method include (1) acquiring the amount of processed data from the computation processing offload unit 130 and (2) acquiring the CPU utilization rate from the OS.
[0026] <Calculation Destination Determination Unit 120> As a function of feature <1>, the calculation destination determination unit 120 determines, based on the data processing volume of APL1, whether to perform calculations in the CPU 11 or offload the calculations to the ACC 12, and changes the offload destination of the calculation processing offload unit 130 if a change is necessary (<Offload Destination Change Instruction>).
[0027] In one example of the determination method, the operator determines a threshold value in advance for the amount of processed data, CPU utilization rate, or traffic volume, and changes the processing destination when the threshold value is exceeded.
[0028] As a function of feature <3>, when the processing destination is changed, the processing destination determination unit 120 notifies the ACC power supply control unit 140 of that information (<offload destination information>). Furthermore, when the processing destination is changed from the CPU 11 to the ACC 12, the processing destination determination unit 120 receives power supply information about the ACC 12 (<power supply information>), and instructs the processing offload unit 130 to make the change when the ACC 12 becomes electrically capable of executing the processing.
[0029] <Computational processing offload unit 130> As a function of feature <1>, the computational processing offload unit 130 has a shared memory (not shown) with the CPU 11 or ACC 12 that executes parallel computations, receives a computational processing request from the APL 1, stores the data to be processed and processing commands in the shared memory, and has the CPU 11 or ACC 12 execute the corresponding processing (<computational processing offload>) (see symbol aa in Figure 1).
[0030] The computational processing offload unit 130 receives a processing request from the APL 1 through the same interface, stores the data to be processed and a processing command in the shared memory, and causes the CPU 11 or the ACC 12 to execute the corresponding processing.
[0031] When the computation offload unit 130 receives a change from the computation destination determination unit 120, it changes the execution location of the computation from the ACC 12 to the CPU 11 or from the CPU 11 to the ACC 12.
[0032] As a function of feature <2>, the computation processing offload unit 130 receives a parallel computation processing request from APL1 and causes the CPU 11 or ACC 12 to execute the corresponding processing. A shared memory is provided between the CPU 11 or ACC 12 that executes the parallel computation, and the data to be processed and processing instructions are stored in the shared memory to execute the processing. Regardless of the execution destination, the processing request is received from APL1 via the same interface, so there is no interruption when APL1 is modified or switched.
[0033] <ACC power supply control unit 140> As a function of feature <3>, the ACC power supply control unit 140 reduces the power of the ACC 12 or increases the power of the ACC 12 to return it to a processing-enabled state based on offload destination information acquired from the processing destination determination unit 120 (<ACC power operation>). When the processing destination is changed from the ACC 12 to the CPU 11, the ACC power supply control unit 140 performs processing to reduce the power of the ACC 12. When the processing destination is changed from the CPU 11 to the ACC 12, the ACC power supply control unit 140 increases the power of the ACC 12 to return it to a processing-enabled state. The ACC power supply control unit 140 provides information to the processing destination determination unit 120 that processing is now possible.
[0034] The operation of the accelerator offload device 100 of the accelerator offload system 1000 configured as described above will be described below. (Features and Operational Overview) Figure 2 is a diagram illustrating the features and operational overview of the accelerator offload device 100. The same processing portions as in Figure 27 are assigned the same step numbers. As indicated by arrow aa1 in Figure 2, high-speed processing by the ACC 12 eliminates performance bottlenecks in the CPU 11 (Feature <1>).
[0035] As shown by the symbol aa2 in FIG. 2, the switching is transparent, there is no need to modify APL1, and no interruption of APL1 occurs during the switching (characteristic <2>).
[0036] As shown by symbol aa1 in FIG. 2, when the amount of data processing is small, the CPU 11 performs the processing, and the power consumption of the ACC 12 is reduced, thereby minimizing the power overhead (feature <3>).
[0037] [Dynamic Switching Between ACC 12 and CPU 11 (Feature <1>)] The dynamic switching between ACC 12 and CPU 11 (Feature <1>) will be described. Feature <1> dynamically switches whether offload target processing is performed by the CPU 11 or the ACC 12 depending on the amount of data processing. As a result, when the amount of data processing is large, the processing can be offloaded to the ACC 12, thereby eliminating the performance bottleneck of the CPU 11.
[0038] 3 is an explanatory diagram of the operation of the accelerator offload system to explain dynamic switching between the ACC 12 and the CPU 11 (feature <1>). The same components as in FIG. 1 are assigned the same reference numerals. In the following explanation, the functional units of the relevant operations are outlined in bold. The computation offload unit 130 receives a computation request from the APL 1 and causes the CPU 11 (see symbol cc in FIG. 3) or the ACC 12 to execute the requested processing (see symbol bb in FIG. 3). When a change is received from the computation destination determination unit 120, the computation execution location is changed from the ACC 12 to the CPU 11, or from the CPU 11 to the ACC 12.
[0039] <Data Processing Flow During APL1 Operation> Fig. 4 is a flowchart showing the data processing operation during the operation of APL1. The processing shown in Fig. 4 is performed continuously while the application is running.
[0040] In step S11, APL1 executes processing other than the offload target processing on CPU 11. In step S12, when the offload target processing becomes necessary, APL1 starts the computation offload unit 130. The computation offload unit 130 checks the "offload target processing execution location" that was set at the time of startup or that was dynamically changed by the computation destination determination unit 120.
[0041] In step S13, the computation offload unit 130 determines whether the "offload target process execution location" is ACC 12. If the "offload target process execution location" is ACC 12 (S13: Yes), in step S14, the computation offload unit 130 writes the offload process content and the data to be processed to the memory shared with ACC 12.
[0042] In step S15, the ACC 12 reads the information written in the shared memory and executes the offload target process. After that, the ACC 12 stores the processing result data in the memory shared with the arithmetic processing offload unit 130.
[0043] In step S16, the computation offload unit 130 reads the data and passes it to the application (APL1). The application continues processing after the offload target processing, and then ends the processing of this flow.
[0044] If the "offload target process execution location" is not ACC12 in step S13 (No in S13), the computation offload unit 130 determines in step S17 whether or not to execute the offload target process on the same core.
[0045] If the offload target process is to be executed on the same core (S17: Yes), the computation offload unit 130 executes the offload target process in step S18. In step S19, the computation offload unit 130 passes the processing result data to the application. The application continues processing after the offload target process and ends the processing of this flow.
[0046] If the offload target process is not to be executed on the same core in step S17 (No in S17), in step S20 the computational processing offload unit 130 writes the process content and the data to be processed to the memory shared with the CPU 11 core that executes the offload target process. In step S19, the corresponding core of the CPU 11 reads the information written in the shared memory and executes the offload target process. Thereafter, the corresponding core of the CPU 11 stores the processing result data in the memory shared with the computational processing offload unit 130 and proceeds to step S19.
[0047] <From when the data processing amount acquisition unit 110 acquires the processing amount and passes it to the computation destination determination unit 120> Figure 5 is a diagram illustrating the operation of the accelerator offload system from when the data processing amount acquisition unit 110 acquires the processing amount and passes it to the computation destination determination unit 120. The same components as in Figure 1 are assigned the same reference numerals. As shown in Figure 5, the data processing amount acquisition unit 110 acquires the data processing amount from APL1 and the computation processing offload unit 130 (see reference numeral dd in Figure 5).
[0048] First, the processing amount received by the data processing amount acquisition unit 110 and example methods for doing so will be described. Example method (1): The data processing amount acquisition unit 110 receives the processing amount from APL1. APL1 periodically transfers the processing amount to the data processing amount acquisition unit 110. This is done via shared memory, but requires modification on the APL1 side.
[0049] Method example (2): The data processing amount acquisition unit 110 receives the processing amount from the arithmetic processing offload unit 130. The arithmetic processing offload unit 130 periodically passes the offloaded processing amount to the data processing amount acquisition unit 110. This passing is performed via a shared memory.
[0050] Method example (3): Method of referring to CPU usage rate The data processing amount acquisition unit 110 receives the usage rate of the corresponding core of the CPU 11 performing offload processing or the corresponding core of the CPU 11 running APL 1. This is done by the data processing amount acquisition unit 110 periodically referring to CPU usage rate information held by the OS (OS etc. 20).
[0051] Method example (4): Method of referring to ACC 12 usage rate When the ACC 12 is performing offload processing, the data processing amount acquisition unit 110 receives information on the operating level of the ACC 12. This information is received via a shared memory that is provided in advance between the data processing amount acquisition unit 110 and the ACC 12.
[0052] Next, the frequency with which the data processing amount acquisition unit 110 receives data and passes it on to the computation destination determination unit 120 will be described. The frequency with which the data processing amount is received is set in advance by the operator. For example, the CPU utilization rate may be referenced once every 10 μsec (in the case of the above-mentioned method example (3)), or information from the memory shared with the computation processing offload unit 130 may be read once every 20 μsec (in the case of the above-mentioned method example (2)). The read information may be shared with the computation destination determination unit 120 immediately, or may be shared in batches. For example, the average value of 10 CPU utilization rates read once every 10 μsec may be shared once every 100 μsec.
[0053] 6 is a flowchart showing the operation of the data processing amount acquisition unit 110 in FIG. 5 from acquiring the processing amount to transferring it to the computation destination determination unit 120. The processing shown in FIG. 6 is performed continuously while the application is running. In step S31, the data processing amount acquisition unit 110 acquires the data processing amount from APL1 or the computation processing offload unit 130. In step S32, the data processing amount acquisition unit 110 transfers the acquired data processing amount to the computation destination determination unit 120. This transfer is performed via shared memory. In step S33, the computation destination determination unit 120 reads the relevant data from the shared memory and ends the processing of this flow.
[0054] <Until the Offload Target Process Execution Location is Switched (ACC 12 → CPU 11)> Figure 7 is an explanatory diagram of the operation of the accelerator offload system until the offload target process execution location is switched (ACC 12 → CPU 11). Components identical to those in Figure 1 are designated by the same reference numerals. As shown in Figure 7, the computation destination determination unit 120 reads the data processing amount from the memory shared with the data processing amount acquisition unit 110 (see reference numeral dd in Figure 7). If the computation destination determination unit 120 needs to change the execution location from ACC 12 to CPU 11, it issues an instruction to the computation offload unit 130 to change the execution location (see reference numeral ee in Figure 7) and an instruction to the ACC power supply control unit 140 to put ACC 12 into power-saving mode (see reference numeral ff in Figure 7). The ACC power supply control unit 140 puts ACC 12 into power-saving mode (see reference numeral gg in Figure 7).
[0055] The following describes the logic for determining whether to change the offload processing execution location (ACC 12 → CPU 11) until the offload processing execution location is switched. The logic for determining whether to change the offload processing location may be statically specified, or may be dynamically changed by acquiring statistical information such as delay time and throughput from the computation offload unit 130.
[0056] Examples of static designation The following are triggers for changing the processing execution location from ACC 12 to CPU 11: (1) The usage rate of ACC 12 falls below 50%. (2) The amount of processing offloaded by the computational processing offload unit 130 falls below a predetermined value (for example, 1000 requests per second).
[0057] Examples of dynamic designation The following are triggers for changing the processing execution location from ACC 12 to CPU 11: (1) If the delay requirements are sufficiently met, the processing execution location is changed to CPU 11. (1-1) The delay requirements that the offload target processing must meet are received in advance from APL 1. (1-2) The computation processing offload unit 130 periodically measures the delay time required for processing. (1-3) The measured delay time is passed to the computation destination determination unit 120 via shared memory. (1-4) The computation destination determination unit 120 compares the measured delay time T with the delay requirement R in (1-1) above, and if T is within half of R, the processing execution location is changed from ACC 12 to CPU 11.
[0058] Next, a method for changing the offload processing execution location will be described. The calculation destination determination unit 120 writes to the memory shared with the calculation processing offload unit 130 when determining the offload processing execution location. For example, if [ACC12:1, CPU11:2], the value written in the shared memory is rewritten from 1 to 2 (1 → 2) when changing the location. Whenever the calculation processing offload unit 130 offloads processing, it checks this written value and offloads to the execution location that corresponds to the value.
[0059] Next, a method for putting the accelerator into a power saving mode will be described. The ACC power supply control unit 140 receives an instruction from the computation destination determination unit 120 to save power for the ACC 12. The power saving for the ACC 12 can be achieved in the following ways: Power saving example (1): Turning off the power supply to the accelerator. Power saving example (2): Rewriting the FPGA to a circuit that consumes less power. Power saving example (3): Power saving the bus connecting to the ACC 12, such as PCIe (Peripheral Component Interconnect Express). Power saving example (4): Degenerating the number of execution cores of the accelerator.
[0060] 8 is a flowchart showing the operation up to switching the offload target processing execution location (ACC 12 → CPU 11) in FIG. 7. The processing shown in FIG. 8 is performed continuously while the application is running. In step S41, the processing destination determination unit 120 reads the data processing amount from the memory shared with the data processing amount acquisition unit 110. In step S42, the processing destination determination unit 120 determines whether it is necessary to change the offload processing execution location based on the data processing amount. In step S43, the processing destination determination unit 120 determines whether it is necessary to change the execution location from ACC 12 to CPU 11. If it is not necessary to change the execution location from ACC 12 to CPU 11 (No in S43), the processing of this flow ends.
[0061] If it is necessary to change the execution location from the ACC 12 to the CPU 11 (Yes in S43), in step S44 the computation destination determination unit 120 instructs the computation offload unit 130 to change the execution location. In step S45, the computation destination determination unit 120 instructs the ACC power supply control unit 140 to put the ACC 12 into power saving mode. In step S46, the ACC power supply control unit 140 puts the ACC 12 into power saving mode, and the processing of this flow ends.
[0062] <Until the Offload Target Processing Execution Location is Switched (CPU 11 → ACC 12)> Figure 9 is a diagram illustrating the operation of the accelerator offload system until the offload target processing execution location is switched (CPU 11 → ACC 12). The same components as in Figure 1 are designated by the same reference numerals. As shown in Figure 9, the computation destination determination unit 120 instructs the ACC power supply control unit 140 to restore the ACC 12 from power-saving mode (see reference numeral hh in Figure 9). The ACC power supply control unit 140 restores the ACC 12 from power-saving mode (see reference numeral ii in Figure 9) and notifies the computation destination determination unit 120 that the restoration is complete (see reference numeral jj in Figure 9). The computation destination determination unit 120 instructs the computation offload unit 130 to change the processing execution location (see reference numeral kk in Figure 9).
[0063] The logic for determining whether to change the offload processing execution location (CPU 11 → ACC 12) until the offload processing execution location is switched will be described. The logic for determining whether to change the offload processing execution location may be statically specified, or may be dynamically changed by acquiring statistical information such as delay time and throughput from the computation offload unit 130.
[0064] Examples of static designation The following are triggers for changing the processing execution location from the CPU 11 to the ACC 12: (1) The CPU usage rate used for offload processing exceeds a predetermined value (e.g., 80%). (2) The amount of processing offloaded by the computation processing offload unit 130 falls below a predetermined value (e.g., 1,000 requests per second).
[0065] Examples of dynamic designation The following are triggers for changing the processing execution location from CPU 11 to ACC 12: (1) If the delay requirements are sufficiently met, the processing execution location is changed to CPU 11. (1-1) The delay requirements that the offload target processing must meet are received in advance from APL 1. (1-2) The computation processing offload unit 130 periodically measures the delay time required for processing. (1-3) The measured delay time is passed to the computation destination determination unit 120 via shared memory. (1-4) The computation destination determination unit 120 compares the measured delay time T with the delay requirement R in (1-1), and if T exceeds R, the processing execution location is changed from CPU 11 to ACC 12.
[0066] Next, a method for returning the accelerator from power saving mode will be described. The ACC power supply control unit 140 receives an instruction from the computation destination determination unit 120 and returns the ACC 12 to an active state. There are the following ways to return the ACC 12. Return example (1): Power on the accelerator. Return example (2): Write a circuit in the FPGA that executes the offload target processing. Return example (3): Cancel power saving on the bus connecting to the ACC 12, such as PCIe. Return example (4): Increase the number of execution cores of the accelerator.
[0067] Next, we will explain why the order in which the calculation destination determination unit 120 issues instructions is different when changing from CPU 11 to ACC 12 and when changing from ACC 12 to CPU 11. If ACC 12 is made power-saving, processing by ACC 12 becomes impossible. Therefore, when changing the execution location from ACC 12 to CPU 11, the processing execution location is switched first and then power saving is made on ACC 12, and when changing the execution location from CPU 11 to ACC 12, ACC 12 is first restored and then the processing execution location is changed.
[0068] 10 is a flowchart showing the operation up to switching the offload target processing execution location (CPU 11 → ACC 12) in FIG. 9. The processing shown in FIG. 10 is performed continuously while the application is running. In step S51, the processing destination determination unit 120 reads the data processing amount from the memory shared with the data processing amount acquisition unit 110. In step S52, the processing destination determination unit 120 determines whether it is necessary to change the offload processing execution location based on the data processing amount. In step S53, the processing destination determination unit 120 determines whether it is necessary to change the execution location from CPU 11 to ACC 12. If it is not necessary to change the execution location from CPU 11 to ACC 12 (S53: No), the processing of this flow ends.
[0069] If it is necessary to change the execution location from the CPU 11 to the ACC 12 (S53: Yes), in step S54 the calculation destination determination unit 120 instructs the ACC power supply control unit 140 to restore the ACC 12 from power saving mode. In step S55, the ACC power supply control unit 140 restores the ACC 12 from power saving mode and notifies the calculation destination determination unit 120 that the restoration is complete. In step S56, the calculation destination determination unit 120 instructs the calculation processing offload unit 130 to change the processing execution location, and the processing of this flow ends.
[0070] Second Embodiment [Overall Configuration] Figure 11 is a schematic configuration diagram of an accelerator offload system according to a second embodiment of the present invention. Components that are the same as those in Figure 1 are assigned the same reference numerals, and overlapping descriptions will be omitted. As shown in Figure 11, an accelerator offload system 1000A includes hardware (HW) 10, an OS 20, a high-speed data communication unit 30, a traffic prediction information provider 210 provided outside a server 150, an accelerator offload device 200, and an APL 1.
[0071] In the accelerator offload device 200 , the data processing amount acquiring unit 110 does not acquire the predicted data processing amount from within the server 150 , but receives it from a traffic prediction information providing unit 210 outside the server 150 .
[0072] [Accelerator offload device 200] The accelerator offload device 200 includes a traffic prediction information providing unit 210 provided outside the server 150, a data processing amount obtaining unit 110, a computation destination determining unit 120, a computation processing offload unit 130, and an ACC power supply control unit 140.
[0073] <Traffic Prediction Information Providing Unit 210> The traffic prediction information providing unit 210 provides prediction information of traffic volume increase / decrease to the data processing volume acquiring unit 110. The traffic prediction information providing unit 210 is configured using, for example, a RAN Intelligent Controller (RIC) in a vRAN (virtualized distributed unit). The traffic prediction information providing unit 210 provides information such as a prediction of traffic volume increase / decrease (hereinafter referred to as traffic volume increase / decrease prediction information) to a server 150 (for example, a vDU server in a vRAN) equipped with the accelerator offload device 200. An example of the traffic volume increase / decrease prediction information is that the traffic volume will increase fivefold from 12:34 due to the occurrence of a planned event such as a fireworks display.
[0074] <Data Processing Amount Acquisition Unit 110> The data processing amount acquisition unit 110 receives predicted increases and decreases in traffic volume and their time information from the traffic prediction information providing unit 210. The data processing amount acquisition unit 110 passes the received information to the computation destination determination unit 120.
[0075] The accelerator offload device 200 is mounted on a vDU server (server 150) in the vRAN. In this case, the data processing amount acquisition unit 110 in the vDU server receives traffic volume increase / decrease prediction information from the RIC, which is the traffic prediction information provider 210. As an example of an acquisition method, the data processing amount acquisition unit 110 receives traffic volume increase / decrease predictions and their time information from a network interface connected to the traffic prediction information provider 210.
[0076] <Computation Destination Determination Unit 120> The computation destination determination unit 120 determines the execution location of the offload target process based on the “traffic volume prediction information” received from the data processing volume acquisition unit 110. Note that cooperation with the ACC power supply control unit 140 and the computation processing offload unit 130 is the same as in the first embodiment.
[0077] The operation of the accelerator offload device 200 of the accelerator offload system 1000A configured as described above will now be described.
[0078] <From when the data processing amount acquisition unit 110 acquires information to when it passes it to the computation destination determination unit 120> Figure 12 is a diagram illustrating the operation of the accelerator offload system from when the data processing amount acquisition unit 110 acquires information to when it passes it to the computation destination determination unit 120. The same components as in Figure 11 are assigned the same reference numerals. As shown in Figure 12, the traffic prediction information provider 210 transmits information such as predicted increases or decreases in traffic volume to the data processing amount acquisition unit 110 (see reference numeral 11 in Figure 12). The data processing amount acquisition unit 110 passes the acquired traffic prediction information to the computation destination determination unit 120 (see reference numeral mm in Figure 12). This transfer is performed via shared memory.
[0079] The following describes the processing volume received by the data processing volume acquisition unit 110 and examples of the method for receiving it. Method Example (1): Receiving Traffic Increase / Decrease Volume Information via Network The data processing volume acquisition unit 110 receives prediction information from the traffic prediction information provider 210, such as that the traffic volume will increase fivefold or reach 6 Gbps from 12:34. The data processing volume acquisition unit 110 receives the prediction information from a network interface connected to the traffic prediction information provider 210.
[0080] Method example (2): Periodically receiving information on current traffic volume via a network The data processing volume acquisition unit 110 periodically receives information on current traffic volume from the traffic prediction information provision unit 210. The data processing volume acquisition unit 110 receives the traffic volume information from a network interface connected to the traffic prediction information provision unit 210.
[0081] Method example (3): Receiving a traffic increase / decrease event via a network The data processing amount acquisition unit 110 receives event information indicating that the traffic volume will increase or decrease significantly from, for example, 12:34 p.m. from the traffic prediction information providing unit 210. The data processing amount acquisition unit 110 receives the event information from a network interface connected to the traffic prediction information providing unit 210.
[0082] 13 is a flowchart showing the operation of the data processing amount obtaining unit 110 in FIG. 12 from obtaining information to passing the information to the calculation destination determining unit 120. In step S61, the traffic prediction information providing unit 210 transmits information such as a predicted increase or decrease in traffic volume to the data processing amount obtaining unit 110. In step S62, the data processing amount obtaining unit 110 receives information such as a predicted increase or decrease in traffic volume from the traffic prediction information providing unit 210. This is achieved by accessing an interface connected to the traffic prediction information providing unit 210 via a network.
[0083] In step S63, the data processing amount acquiring unit 110 passes the acquired traffic prediction information to the computation destination determining unit 120. This passing is performed via shared memory. In step S64, the computation destination determining unit 120 reads the relevant data from the shared memory and ends the processing of this flow.
[0084] <Until the Offload Target Processing Execution Location is Switched (ACC 12 → CPU 11)> Figure 14 is an explanatory diagram of the operation of the accelerator offload system until the offload target processing execution location is switched (ACC 12 → CPU 11). The same components as in Figure 11 are assigned the same reference numerals. As shown in Figure 14, the computation destination determination unit 120 reads traffic prediction information from a memory shared with the data processing amount acquisition unit 110 (see symbol mm in Figure 14). Based on the traffic prediction information, the computation destination determination unit 120 determines whether or not it is necessary to change the offload processing execution location. If it is necessary to change the execution location from ACC 12 to CPU 11, it issues an instruction to the computation offload unit 130 to change the execution location (see symbol nn in Figure 14). The computation destination determination unit 120 then issues an instruction to the ACC power supply control unit 140 to place ACC 12 in power-saving mode (see symbol oo in Figure 14). The ACC power supply control unit 140 puts the ACC 12 into a power saving mode (see symbol pp in FIG. 14).
[0085] The logic for determining whether to change the offload processing location will now be described: The computation destination determination unit 120 determines whether or not it is necessary to change the offload processing location based on the traffic volume (traffic prediction information).
[0086] Examples of Determining a Change in Execution Location The following are triggers for changing the processing execution location from ACC 12 to CPU 11: (1) The received traffic volume per second falls below 1 Gbps. (2) Based on traffic reduction information, it is predicted that the traffic volume per second will fall below 1 Gbps from a certain time. In this case, the computation destination determination unit 120 instructs ACC 12 to change the processing execution location to CPU 11 when the relevant time arrives. (3) When event information indicating a significant decrease in traffic is received, the computation destination determination unit 120 instructs ACC 12 to change the processing execution location to CPU 11.
[0087] 15 is a flowchart showing the operation of FIG. 14 until the offload target processing execution location is switched (ACC 12 → CPU 11). This process is performed continuously while the application is running. In step S71, the processing destination determination unit 120 reads traffic prediction information from a memory shared with the data processing amount acquisition unit 110. In step S72, the processing destination determination unit 120 determines whether or not it is necessary to change the offload processing execution location based on the traffic prediction information. In step S73, the processing destination determination unit 120 determines whether or not it is necessary to change the execution location from ACC 12 to CPU 11. If it is not necessary to change the execution location from ACC 12 to CPU 11 (S73: No), the processing of this flow ends.
[0088] If it is necessary to change the execution location from the ACC 12 to the CPU 11 (S73: Yes), in step S74 the computation destination determination unit 120 instructs the computation offload unit 130 to change the execution location. In step S75, the computation destination determination unit 120 instructs the ACC power supply control unit 140 to put the ACC 12 into power saving mode. In step S76, the ACC power supply control unit 140 puts the ACC 12 into power saving mode, and the processing of this flow ends.
[0089] <Until the Offload Target Processing Execution Location is Switched (CPU 11 → ACC 12)> Figure 16 is an explanatory diagram of the operation of the accelerator offload system until the offload target processing execution location is switched (CPU 11 → ACC 12). The same components as in Figure 11 are assigned the same reference numerals. As shown in Figure 16, the computation destination determination unit 120 reads traffic prediction information from a memory shared with the data processing amount acquisition unit 110 (see symbol mm in Figure 16). Based on the traffic prediction information, the computation destination determination unit 120 determines whether or not it is necessary to change the offload processing execution location. If it is necessary to change the execution location from CPU 11 to ACC 12, the computation destination determination unit 120 instructs the ACC power supply control unit 140 to restore the ACC 12 from power-saving mode (see symbol qq in Figure 16). The ACC power supply control unit 140 restores the ACC 12 from power-saving mode (see symbol rr in Figure 16) and notifies the computation destination determination unit 120 that the restoration is complete (see symbol ss in Figure 16). The computation destination determination unit 120 instructs the computation offload unit 130 to change the processing execution location (see symbol tt in FIG. 16).
[0090] The logic for determining whether to change the offload processing execution location will be described below. The computation destination determination unit 120 determines whether a change in the offload processing execution location is necessary based on the traffic volume (traffic prediction information). Examples of execution location change determination Triggers for changing the processing execution location from the CPU 11 to the ACC 12 include the following: (1) The received traffic volume per second exceeds 1 Gbps. (2) Based on traffic increase volume information, it is predicted that the traffic volume per second will exceed 1 Gbps from a certain time. In this case, the computation destination determination unit 120 issues a processing execution location change instruction from the CPU 11 to the ACC 12 when the relevant time arrives. (3) The computation destination determination unit 120 issues a processing execution location change instruction from the CPU 11 to the ACC 12 when event information indicating a significant increase in traffic is received.
[0091] 17 is a flowchart showing the operation of FIG. 16 until the offload target processing execution location is switched (CPU 11 → ACC 12). This process is performed continuously while the application is running. In step S81, the processing destination determination unit 120 reads traffic prediction information from a memory shared with the data processing amount acquisition unit 110. In step S82, the processing destination determination unit 120 determines whether or not it is necessary to change the offload processing execution location based on the traffic prediction information. In step S83, the processing destination determination unit 120 determines whether or not it is necessary to change the execution location from CPU 11 to ACC 12. If it is not necessary to change the execution location from CPU 11 to ACC 12 (S83: No), the processing of this flow ends.
[0092] If it is necessary to change the execution location from the CPU 11 to the ACC 12 (S83: Yes), in step S84 the computation destination determination unit 120 instructs the ACC power supply control unit 140 to restore the ACC 12 from power saving mode. In step S85, the ACC power supply control unit 140 restores the ACC 12 from power saving mode and notifies the computation destination determination unit 120 that the restoration is complete. In step S86, the computation destination determination unit 120 instructs the computation processing offload unit 130 to change the processing execution location, and the processing of this flow ends.
[0093] (Third Embodiment) [Overall Configuration] Figure 18 is a schematic configuration diagram of an accelerator offload system according to a third embodiment of the present invention. Components that are the same as those in Figure 1 are assigned the same reference numerals, and overlapping descriptions will be omitted. As shown in Figure 18, an accelerator offload system 1000B has hardware (HW) 10, an OS etc. 20, a high-speed data communication unit 30, an accelerator offload device 300, and an APL 1. A CPU 11 has multiple CPU cores (core #a) 11a, (core #b) 11b, (core #c) 11c, ...
[0094] [Accelerator Offload Device 300 ] The accelerator offload device 300 includes a data processing amount acquisition unit 110 , a computation destination determination unit 120 , a computation offload unit 130 , an ACC power supply control unit 140 , and a CPU performance control unit 310 .
[0095] The accelerator offload device 300 includes a CPU performance control unit 310, which controls CPU performance by scaling in or out the number of processing execution cores when the CPU 11 executes the offload target processing.
[0096] When the offload target process is offloaded to the ACC 12, the accelerator offload device 300 controls the frequency of the CPU 11 to reduce power consumption.
[0097] The CPU performance control unit 310 includes a CPU core scaler 311 and a CPU frequency control unit 312 .
[0098] When the CPU 11 executes the offload target process, the CPU core scaling unit 311 scales out or scales in the number of execution cores to satisfy the performance requirements.
[0099] The CPU frequency control unit 312 saves power by lowering the CPU frequency while no processing is being performed on the CPU core, such as while the offload target processing is being performed by the ACC 12. When processing is started again, the CPU frequency control unit 312 returns the frequency of the CPU 11 to its original value.
[0100] The operation of the accelerator offload device 300 of the accelerator offload system 1000B configured as described above will be described below. The operation of the accelerator offload device 300 will be described separately for an operation for changing the number of execution processing cores (while the CPU 11 is executing the offload target process), an operation until the offload target process execution location is switched (CPU 11 → ACC 12), and an operation until the offload target process execution location is switched (ACC 12 → CPU 11).
[0101] <Change in the Number of Execution Processing Cores (CPU 11 Executing Offload Target Processing)> Figure 19 is a diagram illustrating the operation of the accelerator offload system when the number of execution processing cores is changed (CPU 11 Executing Offload Target Processing). The same components as in Figure 18 are assigned the same reference numerals. The computation offload unit 130 is currently executing offload target processing on the CPU core (core #a) 11a of the CPU 11 (see reference numeral uu in Figure 19).
[0102] The computation destination determination unit 120 reads the data processing amount from the shared memory with the data processing amount acquisition unit 110 (see symbol vv in FIG. 19 ). Based on the data processing amount, the computation destination determination unit 120 determines whether the offload processing execution location needs to be changed, and if so, notifies the CPU performance control unit 310 of this (see symbol ww in FIG. 19 ). When the offload processing target is performed by the CPU 11, the CPU core scaling unit 311 performs core allocation by scaling out / in the number of execution cores to meet performance requirements (see symbol xx in FIG. 19 ). The CPU frequency control unit 312 lowers the CPU frequency when no processing is being performed on the CPU core, such as when the offload processing target is being performed by the ACC 12 (see symbol yy in FIG. 19 ).
[0103] The following describes the determination and notification of the number of cores required for offload target processing. While the CPU 11 is executing the offload target processing, the computation destination determination unit 120 determines the number of cores to be assigned according to the processing volume, and notifies the CPU performance control unit 310. The following are examples of how the number of cores to be assigned can be determined. Determination example (1): Determine the increase or decrease based on the average CPU 11 usage rate of the assigned cores. Average usage rate is 80% or higher: Increase the number of CPU 11 cores by one. Average usage rate is 30% or lower: Decrease the number of CPU 11 cores by one. Determination example (2): Determine the required number based on the number of requests per second from the computation processing offload unit 130. Number of requests per second: Less than 500: 1 CPU 11 core; 500 to 750: 2 CPU 11 cores; 750 to 1000: 3 CPU 11 cores, etc.
[0104] The operation of changing the number of execution processing cores (while the CPU 11 is executing the offload target process) in FIG. 19 corresponds to steps S51, S52, S53, and S94 to S96 in the flowchart in FIG. 21, which will be described later.
[0105] <Until the Offload Target Processing Execution Location is Switched (CPU 11 → ACC 12)> Figure 20 is an explanatory diagram of the operation of the accelerator offload system until the offload target processing execution location is switched (CPU 11 → ACC 12). The same components as in Figure 18 are assigned the same reference numerals. As shown in Figure 20, when it is necessary to change the execution location from CPU 11 to ACC 12, the computation destination determination unit 120 instructs the computation offload unit 130 to change the processing execution location (see symbol zz in Figure 20). The computation destination determination unit 120 notifies the CPU performance control unit 310 that the processing execution location has been changed from CPU 11 to ACC 12 (see symbol ww in Figure 20). When it is necessary to change the number of execution cores of CPU 11, the computation destination determination unit 120 notifies the CPU performance control unit 310 of the required number of cores. The CPU core scaling unit 311 allocates cores appropriate for the required number of cores (see symbol xx in Figure 19). The CPU frequency control unit 312 reduces the frequency of the CPU core used to execute the offload target process (see symbol yy in FIG. 19).
[0106] The following describes scaling in when switching the offload target process from CPU 11 to ACC 12. If multiple CPU 11 cores have been assigned to the offload target process, the CPU core scaling unit 311 scales in the cores when the process execution location is shifted from CPU 11 to ACC 12. That is, the CPU core scaling unit 311 releases the CPU cores reserved for the offload target process and reallocates them to application processing or the like.
[0107] FIG. 21 is a flowchart showing the operation up to switching the offload target process execution location (CPU 11 → ACC 12) in FIG. 20 . This flow includes the operation of changing the number of execution processing cores (while the offload target process is being executed by the CPU 11) in FIG. 19 . The process shown in FIG. 21 is continuously performed while the application is running. Note that steps that perform the same processes as those in the flow of FIG. 10 are assigned the same step numbers. In step S51, the computation destination determination unit 120 reads the data processing amount from the memory shared with the data processing amount acquisition unit 110. In step S52, the computation destination determination unit 120 determines whether or not it is necessary to change the offload process execution location based on the data processing amount. In step S53, the computation destination determination unit 120 determines whether or not it is necessary to change the execution location from the CPU 11 to the ACC 12. If it is not necessary to change the execution location from the CPU 11 to the ACC 12 (S53: No), the process proceeds to step S94.
[0108] If it is necessary to change the execution location from CPU 11 to ACC 12 in step S53 (S53: Yes), in step S91 the computation destination determination unit 120 instructs the computation processing offload unit 130 to change the processing execution location. In step S92, the computation destination determination unit 120 notifies the CPU performance control unit 310 that the processing execution location has been changed from CPU 11 to ACC 12. In step S93, the CPU frequency control unit 312 lowers the frequency of the CPU core used to execute the offload target processing, thereby saving power and ending the processing of this flow.
[0109] If it is determined in step S53 that there is no need to change the execution location from CPU 11 to ACC 12 (S53: No), then in step S94 the calculation destination determination unit 120 determines, based on the amount of data processing, whether or not it is necessary to change the number of execution cores of the CPU 11. If it is not necessary to change the number of execution cores of the CPU 11 (S94: No), the processing of this flow ends.
[0110] If the number of execution cores of the CPU 11 needs to be changed (S94: Yes), in step S95 the computation destination determination unit 120 notifies the CPU performance control unit 310 of the required number of cores. In step S96, the CPU core scaling unit 311 allocates cores appropriate for the required number of cores, and then ends the processing of this flow.
[0111] <Until the Offload Target Processing Execution Location is Switched (ACC 12 → CPU 11)> Figure 22 is an explanatory diagram of the operation of the accelerator offload system until the offload target processing execution location is switched (ACC 12 → CPU 11). The same components as in Figure 18 are assigned the same reference numerals. As shown in Figure 22, the computation destination determination unit 120 notifies the CPU performance control unit 310 that the processing execution location has been changed from ACC 12 to CPU 11 and the number of CPU 11 cores required for the offload target processing (see reference symbol aaa in Figure 22). The CPU core scaling unit 311 allocates the notified required number of cores for the offload target processing (see reference symbol xx in Figure 22). The CPU frequency control unit 312 increases the frequency of the allocated core to its maximum (see reference symbol yy in Figure 22). The CPU frequency control unit 312 notifies the computation destination determination unit 120 of the end of core allocation and frequency variation (see reference symbol bbb in Figure 22). The computation destination determination unit 120 issues an instruction to the computation offload unit 130 to change the execution location (see symbol zz in FIG. 22).
[0112] Notification of the number of cores required for offload target processing will be described. When changing the execution location of the offload target processing from ACC 12 to CPU 11, the computation destination determination unit 120 determines the number of cores to be allocated according to the amount of processing, and notifies the CPU performance control unit 310. The following are examples of how the number of cores to be allocated can be determined.
[0113] Determination example (1): Determined based on the usage rate of ACC12. If the usage rate of ACC12 is less than 20%, the number of CPU cores is 1. If the usage rate is between 20% and 50%, the number of CPU cores is 2.
[0114] Determination example (2): Determined based on the number of requests per second of the computational processing offload unit 130. The number of requests per second is as follows: Less than 500: 1 CPU core; 500 or more but less than 750: 2 CPU cores; 750 or more but less than 1000: 3 CPU cores, etc.
[0115] 23 is a flowchart showing the operation of FIG. 22 <Up until switching the offload target processing execution location (ACC 12 → CPU 11)>. Note that the same step numbers are used for steps that perform the same processing as in the flow of FIG. 8. In step S41, the processing destination determination unit 120 reads the data processing amount from the memory shared with the data processing amount acquisition unit 110. In step S42, the processing destination determination unit 120 determines whether it is necessary to change the offload processing execution location based on the data processing amount. In step S43, the processing destination determination unit 120 determines whether it is necessary to change the execution location from ACC 12 to CPU 11. If it is not necessary to change the execution location from ACC 12 to CPU 11 (No in S43), the processing of this flow ends.
[0116] If it is necessary to change the execution location from ACC 12 to CPU 11 (Yes in S43), in step S101, the computation destination determination unit 120 notifies the CPU performance control unit 310 that the processing execution location has been changed from ACC 12 to CPU 11 and the number of CPU 11 cores required for the offloaded processing. In step S102, the CPU core scaling unit 311 allocates the notified required number of cores for the offloaded processing. In step S103, the CPU frequency control unit 312 increases the frequency of the allocated core to its maximum. In step S104, the CPU frequency control unit 312 notifies the computation destination determination unit 120 of the completion of core allocation and frequency variation.
[0117] In step S44, the computation destination determination unit 120 instructs the computation offload unit 130 to change the execution location. In step S45, the computation destination determination unit 120 instructs the ACC power supply control unit 140 to set the ACC 12 to the power saving mode. In step S46, the ACC power supply control unit 140 sets the ACC 12 to the power saving mode, and the processing of this flow ends.
[0118] [Hardware Configuration] The accelerator offload devices 100, 200, and 300 according to the above embodiments are realized by, for example, a computer 900 configured as shown in FIG. 24. FIG. 24 is a hardware configuration diagram showing an example of the computer 900 that realizes the functions of the accelerator offload devices 100, 200, and 300. The computer 900 includes a CPU 901, a RAM 902, a ROM 903, a HDD 904, an accelerator 905, an input / output interface (I / F) 906, a media interface (I / F) 907, and a communication interface (I / F) 908. The accelerator 905 corresponds to the accelerator offload devices 100, 200, and 300 shown in FIGS. 1, 11, and 18.
[0119] The accelerator 905 is an accelerator (device) 11 (FIGS. 1 and 5) that processes at least one of data from the communication I / F 908 and data from the RAM 902 at high speed. Note that the accelerator 905 may be of a type (look-aside type) that executes processing from the CPU 901 or RAM 902 and then returns the execution result to the CPU 901 or RAM 902. On the other hand, the accelerator 905 may be of a type (in-line type) that performs processing between the communication I / F 908 and the CPU 901 or RAM 902.
[0120] The accelerator 905 is connected to an external device 915 via a communication I / F 908. The input / output I / F 906 is connected to an input / output device 916. The media I / F 907 reads and writes data from and to a recording medium 917.
[0121] The CPU 901 operates based on a program stored in the ROM 903 or HDD 904, and controls each unit of the accelerator offload devices 100, 200, and 300 shown in Figures 1, 11, and 18 by executing a program (also called an application or an app for short) loaded into the RAM 902. This program can also be distributed via a communication line or recorded on a recording medium 917 such as a CD-ROM. The ROM 903 stores a boot program executed by the CPU 901 when the computer 900 starts up, programs that depend on the hardware of the computer 900, and the like.
[0122] The CPU 901 controls an input / output device 916, which is made up of input units such as a mouse and a keyboard, and output units such as a display and a printer, via an input / output I / F 906. The CPU 901 acquires data from the input / output device 916 via the input / output I / F 906, and outputs generated data to the input / output device 916. Note that a GPU (Graphics Processing Unit) or the like may be used as a processor together with the CPU 901.
[0123] The HDD 904 stores programs executed by the CPU 901 and data used by the programs. The communication I / F 908 receives data from other devices via a communication network (e.g., NW (Network)) and outputs the data to the CPU 901, and also transmits data generated by the CPU 901 to other devices via the communication network.
[0124] The media I / F 907 reads a program or data stored in a recording medium 917 and outputs it to the CPU 901 via the RAM 902. The CPU 901 loads a program related to a target process from the recording medium 917 onto the RAM 902 via the media I / F 907, and executes the loaded program. The recording medium 917 is an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto Optical Disc), a magnetic recording medium, a conductive memory tape medium, a semiconductor memory, or the like.
[0125] For example, when a computer 900 functions as the accelerator offload devices 100, 200, and 300 configured as one device according to this embodiment, a CPU 901 of the computer 900 executes a program loaded onto a RAM 902 to realize the functions of the accelerator offload devices 100, 200, and 300. Furthermore, data stored in the RAM 902 is stored in the HDD 904. The CPU 901 reads and executes a program related to a target process from a recording medium 917. Alternatively, the CPU 901 may read a program related to a target process from another device via a communications network.
[0126] [Effects] As described above, the accelerator offload device 100 offloads specific processing of an application program (APL1) to an accelerator (ACC12), and includes: a data processing amount acquisition unit 110 that acquires the current data processing amount and / or the predicted data processing amount of the offload target processing being performed; a calculation destination determination unit 120 that determines whether to perform calculations in the CPU 11 or to offload the calculations to the ACC 12 based on the data processing amount of the APL1, and changes the offload destination of the calculation processing offload unit 130 if a change is necessary; and the calculation processing offload unit 130 that has a shared memory between itself and the CPU 11 or the ACC 12 that executes parallel calculations, receives a calculation processing request from the APL1, stores the data to be processed and a processing command in the shared memory, and causes the CPU 11 or the ACC 12 to execute the corresponding processing.
[0127] This allows for minimizing power consumption overhead while suppressing CPU-related performance bottlenecks without modifying application programs. Specifically, by dynamically switching between the CPU 11 and the ACC 12 for offloaded processing depending on the amount of data processing, when the amount of data processing is large, the processing can be offloaded to the ACC 12, eliminating the performance bottleneck of the CPU 11 (Feature <1>). Dynamic switching of the arithmetic processing unit is transparent to the APL 1, so no modification of the APL 1 is required, and no interruption of the APL 1 occurs during switching (Feature <2>). When parallel processing is performed by the CPU 11, reducing the power consumption of the ACC 12 minimizes the power consumption overhead caused by the inclusion of the ACC 12 (Feature <3>).
[0128] In the accelerator offload device 100 (FIG. 1), the computational processing offload unit 130 receives processing requests from the APL 1 via the same interface, stores the data to be processed and processing commands in shared memory, and causes the CPU 11 or the ACC 12 to execute the corresponding processing.
[0129] By doing this, the arithmetic processing offload unit 130 stores the data to be processed and processing instructions in shared memory on the hardware 10, and can execute specific processing of the application program (APL1) at high speed without modifying the application program, regardless of whether the parallel arithmetic processing execution unit is assigned to ACC12 or CPU11.
[0130] In the accelerator offload device 100 (FIG. 1), when the calculation processing offload unit 130 receives a change from the calculation destination determination unit 120, it changes the execution location of the calculation processing from the ACC 12 to the CPU 11 or from the CPU 11 to the ACC 12.
[0131] In this way, the arithmetic processing offload unit 130 can dynamically switch the parallel arithmetic processing execution unit between the ACC 12 and the CPU 11. Dynamic switching of the arithmetic processing unit is performed transparently to the APL1, so there is no need to modify the APL1, and no interruption of the APL1 occurs during switching.
[0132] The accelerator offload device 100 (FIG. 1) is characterized by the following: an ACC power control unit 140 that performs ACC power control to reduce the power of the ACC 12 or increase the power of the ACC 12 to return it to a processing-enabled state based on offload destination information acquired from the computation destination determination unit 120;
[0133] In this way, the arithmetic processing offload unit 130 can minimize the power consumption overhead caused by the inclusion of the ACC 12 by reducing the power consumption of the ACC 12 when parallel calculations are being performed by the CPU 11. Also, when the ACC 12 is not operating, the power consumption of the ACC 12 can be minimized.
[0134] The accelerator offload device 200 (FIG. 14) is characterized by including a traffic prediction information providing unit 210 that provides the data processing amount acquiring unit 110 with prediction information on increases and decreases in traffic volume.
[0135] In this way, the computation offload unit 130 can perform computation offloading using not only the current data processing volume but also traffic volume increase / decrease prediction information. Also, the data processing volume acquisition unit 110 can acquire the predicted data processing volume from the traffic prediction information providing unit 210 outside the server 150, rather than from within the server 150. For example, the traffic prediction information providing unit 210 can use the RIC in the vRAN and transmit traffic volume increase / decrease prediction information to the vDU server on which the present system is installed.
[0136] In the accelerator offload device 300 (FIG. 18), the CPU has multiple CPU cores, and is characterized by including a CPU performance control unit 310 that controls CPU performance by scaling in or out the number of processing execution cores when offload target processing is performed by the CPU.
[0137] By doing this, when the offload target processing is performed by CPU 11, the number of processing execution cores can be scaled in or out to vary performance, and when the offload target processing is offloaded to ACC 12, power consumption can be reduced by controlling the frequency of CPU 11.
[0138] Note that, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. Furthermore, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. Furthermore, the components of each device shown in the drawings are functionally conceptual and do not necessarily have to be physically configured as shown. In other words, the specific form of distribution and integration of each device is not limited to that shown in the drawings, and all or part of the devices can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0139] Furthermore, the above-described configurations, functions, processing units, processing means, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented by software that causes a processor to interpret and execute programs that implement the respective functions. Information such as programs, tables, and files that implement the respective functions may be stored in a memory, a recording device such as a hard disk or a solid-state drive (SSD), or a recording medium such as an integrated circuit (IC) card, a secure digital (SD) card, or an optical disk.
[0140] 1 APL (application program) 10 Hardware 11 CPU 11a, 11b, 11c CPU core (core#) 12 ACC (accelerator) 13 Ring buffer 30 High-speed data communication unit 100, 200, 300 Accelerator offload device 110 Data processing amount acquisition unit 120 Computation destination determination unit 130 Computation processing offload unit 140 ACC power supply control unit 210 Traffic prediction information providing unit 1000 to 1000B Accelerator offload system
Claims
1. A vRAN (virtualized Radio Access Network) system having an accelerator and a CPU, in which part or all of the data processing performed by the CPU is offloaded to the accelerator, a traffic prediction information providing unit that provides traffic volume increase / decrease prediction information as traffic prediction information; a computation destination determination unit that determines whether or not it is necessary to change the offload processing execution location based on the current data processing volume of the offload target processing and traffic prediction information; a computation offload unit that, when it is necessary to change the execution location from the accelerator to the CPU, causes the CPU to execute the process to be offloaded to the accelerator; a power supply control unit that sets the accelerator to a power saving state when it is necessary to change the execution location from the accelerator to the CPU. A vRAN system characterized by:
2. The traffic prediction information providing unit is a RAN Intelligent Controller (RIC). The vRAN system of claim 1 .
3. The power supply control unit reduces the power of the accelerator or turns off the power when the accelerator is set to power saving. The vRAN system of claim 1 .
4. The CPU has multiple CPU cores, The calculation destination determination unit When the offload target process is performed by the CPU, the CPU performance is controlled by scaling in or out the number of processing execution cores. The vRAN system of claim 1 .
5. A method for managing a traffic forecast using a vRAN, comprising: The vDU server includes the computation destination determination unit, the computation offload unit, and the power supply control unit. The vRAN system of claim 1 .
6. An accelerator offloading method for a vRAN (virtualized Radio Access Network) system having an accelerator and a CPU, offloading part or all of data processing performed by the CPU to the accelerator, The vRAN system includes a traffic prediction information providing unit, a computation destination determining unit, a computation processing offload unit, and a power supply control unit, a step in which the traffic prediction information providing unit provides prediction information of an increase or decrease in traffic volume as traffic prediction information; a step in which the computation destination determination unit determines whether or not it is necessary to change the offload processing execution location based on the current data processing volume of the offload target processing and traffic prediction information; a step in which the arithmetic processing offload unit causes the CPU to execute the processing to be offloaded to the accelerator when it is necessary to change the execution location from the accelerator to the CPU; the power supply control unit, when it is necessary to change the execution location from the accelerator to the CPU, performs a power saving setting on the accelerator. An accelerator offload method comprising: