Server internal data transfer device, server internal data transfer method, and program
The server internal data transfer device addresses high power consumption in polling models by using a packet arrival monitoring unit and virtual queues to optimize CPU usage and data transfer efficiency in multi-application environments.
Patent Information
- Application Number
- JP2023565786
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2041-12-08
AI Technical Summary
Existing data transfer technologies using polling models, such as DPDK and KBP, result in high power consumption due to constant CPU usage for monitoring packet arrival, even during intermittent data reception.
A server internal data transfer device that includes a packet arrival monitoring unit and a polling control unit to monitor data arrival and wake up application threads only when necessary, using a virtual queue system to replace physical queues and dynamically adjust connections between virtual and physical queues.
Reduces CPU power consumption by minimizing unnecessary polling, while maintaining efficient data transfer and reducing delays, allowing for flexible resource management and power savings in multi-application environments.
Smart Images

Figure 0007709645000001 
Figure 0007709645000002 
Figure 0007709645000003
Abstract
Description
Technical Field
[0001] The present invention relates to an in-server data transfer device, an in-server data transfer method, and a program.
Background Art
[0002] Against the background of the progress of virtualization technology such as NFV (Network Functions Virtualization), systems are being constructed and operated for each service. Further, from the form of constructing a system for each service, service functions are divided into reusable module units and operated on an independent virtual machine (VM: Virtual Machine or container, etc.) environment.
[0003] As a technology for constructing a virtual machine, a hypervisor environment composed of Linux (registered trademark) and KVM (kernel-based virtual machine) is known. In this environment, a Host OS (the OS installed on a physical server is called a Host OS) incorporating a KVM module operates in a memory area different from a user space called a kernel space as a hypervisor. In this environment, a virtual machine operates in the user space, and a Guest OS (the OS installed on the virtual machine is called a Guest OS) operates in the virtual machine.
[0004] Unlike the physical server on which the Host OS operates, the virtual machine in which the Guest OS operates includes all HW (hardware) including a network device (represented by an Ethernet (registered trademark) card device, etc.), and becomes register control necessary for interrupt processing from the HW to the Guest OS and writing from the Guest OS to the hardware. In such register control, since notifications and processes that should originally be executed by physical hardware are emulated by software in a pseudo manner, the performance is generally lower than that of the Host OS environment.
[0005] In this performance degradation, there is a technology that reduces the emulation of HW and improves the communication performance and versatility with a fast and unified interface, especially for the Host OS and external processes existing outside the self-virtual machine from the Guest OS. As this technology, a device abstraction technology called virtio, that is, a para-virtualization technology, has been developed and has already been incorporated into many general-purpose OSs such as Linux (registered trademark) and FreeBSD (registered trademark) and is currently in use (see Patent Document 1).
[0006] In virtio, regarding data input / output such as console, file input / output, and network communication, data exchange by a queue designed as a ring buffer is defined by queue operations as a transport for single-direction transfer of transfer data. Then, by preparing the number and size of queues suitable for each device using the virtio queue specification at the time of Guest OS startup, communication between the Guest OS and the outside of the self-virtual machine can be realized only by queue operations without executing hardware emulation.
[0007] As data transfer technologies within a server, there are New API (NAPI), DPDK (Data Plane Development Kit), and KBP (Kernel Busy Poll).
[0008] New API (NAPI) performs packet processing by a software interrupt request after a hardware interrupt request when a packet arrives.
[0009] DPDK realizes packet processing functions in the user space where applications operate and performs immediate harvesting upon packet arrival in a polling model from the user space (see Non-Patent Document 1). Specifically, DPDK is a framework for performing the control of a NIC (Network Interface Card) that was conventionally carried out by the Linux kernel (registered trademark) in the user space. The biggest difference from the processing in the Linux kernel is that it has a polling-based receiving mechanism called PMD (Pull Mode Driver). Normally, in the Linux kernel, when data arrives at the NIC, an interrupt occurs, and based on this, the receiving process is executed. On the other hand, PMD continuously performs data arrival confirmation and receiving processing with a dedicated thread. By eliminating overheads such as context switches and interrupts, high-speed packet processing can be performed. DPDK significantly improves the performance and throughput of packet processing and enables securing a lot of time for data plane application processing. However, DPDK uses computer resources such as the CPU (Central Processing Unit) and NIC in an exclusive manner.
[0010] Non-Patent Document 2 describes a server-internal network delay control device (KBP: Kernel Busy Poll). KBP constantly monitors packet arrival by a polling model within the kernel. Thereby, softIRQ is suppressed, and low-latency packet processing is realized.
[0011] Next, the DPDK system will be described. [DPDK System Configuration] Figure 23 is a diagram showing the configuration of a DPDK system that controls HW10 equipped with an accelerator 11. The DPDK system has HW10, a packet processing API (Application Programming Interface) 14, and an application (APL) 20.
[0012] APL20 is packet processing that is performed prior to the execution of APL. Here, APL20 is APL1 and APL2.
[0013] The packet processing API 14 is an API for offloading packet processing to a NIC or an accelerator. The packet processing API 14 is data transfer middleware placed in the user space, which is DPDK. DPDK realizes packet processing functions in the user space where APL20 operates, and enables reduction of packet transfer delay by performing immediate harvesting at the arrival of a packet in the polling model from the user space. That is, DPDK performs harvesting of packets (refer to the content of the packets accumulated in the buffer, and delete the corresponding queue entry from the buffer in consideration of the next process to be performed) by polling (busy polling of the queue by the CPU), so there is no waiting and the delay is small.
[0014] HW10 conducts data transmission and reception communication with APL1 and APL2. In the following description, the data flow in which APL1 and APL2 receive packets from HW10 is referred to as Rx-side reception, and the data flow in which APL1 and APL2 transmit packets to HW10 is referred to as Tx-side transmission.
[0015] HW10 includes an accelerator 11. Also, HW10 may include a NIC (physical NIC) for connecting to a communication network. The accelerator 11 is a computing unit hardware that performs specific operations at high speed based on inputs from the CPU. Specifically, the accelerator 11 is a PLD (Programmable Logic Device) such as a GPU (Graphics Processing Unit) or an FPGA (Field Programmable Gate Array). In FIG. 23, the accelerator 11 includes a plurality of IP cores (Intellectual Property Core) 12 and a physical queue 13 consisting of an Rx queue (queue: waiting queue) and a Tx queue that holds data in a first-in-first-out list structure. The IP core 12 is design information of a reusable circuit component that constitutes a semiconductor such as an FPGA, an IC, or an LSI, and is sometimes called a device core (Core processor).
[0016] Offload part of the processing of APL1 and APL2 to the accelerator 11 to achieve performance and power efficiency that cannot be achieved by software (CPU processing) alone. In large-scale server clusters such as data centers that constitute NFV (Network Functions Virtualization) and SDN (Software Defined Network), cases where the above-described accelerator 11 is applied are assumed.
[0017] Existing applications (APL20, APL1, APL2) that transfer data to the accelerator in poll mode, such as DPDK, operate with a physical queue 13 used by the application fixedly associated at initialization (see the dashed box in FIG. 23). The thread of the application (hereinafter referred to as the application thread) performs transmission and reception processing via a Ring Buffer 16 corresponding to the accelerator 11 (see FIG. 24). The application thread 15 is a PollingThread here.
[0018] FIG. 24 is a diagram for explaining the reception processing by polling of the DPDK system in FIG. 23. The application thread (PollingThread) 15 and the packet processing API 14 are arranged on the user space 30. In the reception process by polling, when there is data to be transferred, the pointer of the corresponding data is stored in the Ring Buffer 16 (see reference symbol a in FIG. 24). In order to suppress the delay from the accelerator 11, the application thread 15 polls the Ring Buffer 16, obtains the pointer of the corresponding data when there is data to be transferred, and performs the reception process (see reference symbol b in FIG. 24). At this time, the application thread 15 has a 100% CPU usage rate due to polling, and the power consumption increases. DPDK has little delay because the application thread 15 polls for packet arrival, but the power consumption increases. As shown in FIG. 23, when there are multiple applications APL1 and APL2, the impact of the increase in power consumption is very large.
Prior Art Documents
Patent Documents
[0019]
Patent Document 1
Non-Patent Documents
[0020]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0021] However, packet transfer using the polling model has the following problems. In DPDK, since the kernel thread performs polling (busy polling the queue with the CPU), it occupies a CPU core. Therefore, for example, even in the case of intermittent packet reception, in DPDK, regardless of whether a packet arrives or not, the CPU is always used at 100%, resulting in a problem of high power consumption. KBP also has the same problem as the above DPDK. That is, KBP can suppress softIRQ and achieve low-latency packet processing by constantly monitoring packet arrival in the kernel using the polling model. However, since the kernel thread that constantly monitors packet arrival occupies a CPU core and always uses CPU time, there is a problem of high power consumption.
[0022] In view of such a background, the present invention has been made. When using accelerator resources mounted on a physical server in a plurality of applications, the present invention aims to reduce the power consumption of the CPU used for data polling while suppressing the data transfer delay from the accelerator to the application.
Means for Solving the Problems
[0023] In order to solve the above-mentioned problems, when an application uses a device including an accelerator, a server internal data transfer device that performs data transfer from the device to the application, A data arrival monitoring unit that monitors communication between the device and an app thread corresponding to the application and measures the data arrival timing, and a polling control unit that wakes up the app thread at the time of data arrival when the data arrival monitoring unit detects the arrival of data to receive the data and puts it to sleep while there is no arrival of the data. characterized by comprising: a device virtual part that connects the device to the application by causing data processing to be performed on an app thread corresponding to the device and the application using a virtual queue that substitutes for the physical queue of the device; and a proxy part that dynamically changes the connection between the virtual queue and the physical queue.
Advantages of the Invention
[0024] According to the present invention, when using the accelerator resources installed in a physical server in a plurality of applications, it is possible to reduce the power consumption of the CPU used for data polling while suppressing the data transfer delay from the accelerator to the application.
Brief Description of the Drawings
[0025]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Embodiments for Carrying Out the Invention
[0026] Hereinafter, a server internal data transfer system and the like in an embodiment for carrying out the present invention (hereinafter referred to as "this embodiment") will be described with reference to the drawings. (Embodiment) [Overall Configuration] FIG. 1 is a schematic configuration diagram of a server internal data transfer system according to an embodiment of the present invention. The same components as those in FIG. 23 are denoted by the same reference numerals. As shown in FIG. 1, the server internal data transfer system 1000 includes HW10, a packet processing API 14, a controller (CTRL) (server internal data transfer device) 100, and an application (APL) 20.
[0027] The controller 100 is arranged between the application 20 and the HW10. For this reason, the packet processing API 14 is arranged between the application 20 and the controller 100. However, from the perspective of the application 20, the existence of the controller 100 is not visible, and it is an API for offloading packet processing to the NIC or the accelerator.
[0028] The controller 100 is a server internal data transfer device that transfers data from the device to the application 20 when the application 20 uses a device including the accelerator 11. The controller 100 manages the association between the application thread 15 and the IP cores 12 of the plurality of accelerators 11, and communicates with the accelerator 11 on behalf of the application 20.
[0029] The controller 100 includes a packet arrival monitoring unit 110, a polling control unit 120, a device emulation unit 130, and a proxy unit 140.
[0030] The packet arrival monitoring unit 110 monitors the communication between the accelerator 11 (device) and the application thread 15 and measures the packet arrival timing. That is, after the processing between the accelerator 11 and the application thread 15 is completed, the packet comes up for reception in the application. However, the packet arrival monitoring unit 110 monitors this packet arrival and measures the packet arrival timing from the accelerator 11.
[0031] The polling control unit 120 wakes up the application thread 15 at the time of packet arrival when the packet arrival monitoring unit 110 detects packet arrival to prompt packet processing, and puts it to sleep when there is no packet. Previously, the device was always awake and performing polling processing. Now, when it is not necessary, the packet processing is stopped (the polling is stopped; put to sleep). The polling control unit 120 wakes up the application thread 15 only when a packet arrives.
[0032] The device emulation unit 130 connects to the application 20 with an interface equivalent to the existing one and provides a virtual queue 200 (see Fig. 5) instead of the physical queue 13 of the device. The device emulation unit 130 corresponds to the physical queue of the device and causes the application thread 15 to perform packet processing using the virtual queue 200 instead of the physical queue, thereby connecting the device to the application 20. The virtual queue 200 is a queue shown to the application thread 15 instead of the physical queue 13.
[0033] From the application 20, it is desired to make it seem that the accelerator 11 is always being used (always communicating with the accelerator 11) (in other words, to hide the controller 100 from the accelerator 11). The device emulation unit 130 connects to the application 20 with an interface equivalent to the existing one and provides a virtual queue 200 instead of the physical queue 13 of the device. Thereby, the application 20 appears to be always communicating with the accelerator 11.
[0034] The proxy unit 140 dynamically changes the connection between the virtual queue 200 and the physical queue 13. When a packet arrives from the application thread 15 via the virtual queue 200, For example it is necessary to deliver the packet to the accelerator 11. In this case, the proxy unit 140 connects the virtual queue 200 and the physical queue 13.
[0035] The proxy unit 140 changes the association of the physical queue 13 with respect to the virtual queue 200 and simultaneously instructs the increase or decrease of the IP core 12 and the application thread 15. The proxy unit 140 further associates a plurality of physical queues (1:N) with one application thread 15 or one physical queue (M:1) with a plurality of application threads 15 (N and M are arbitrary natural numbers).
[0036] [In-Server Data Transfer System] FIG. 2 is a block configuration diagram of an in-server data transfer system 1000 including a controller 100. The same components as those in FIG. 1 are denoted by the same reference numerals. The in-server data transfer system 1000 includes an application thread 15 arranged on the user space 30, an accelerator 11 having a plurality of physical queues 13 and an IP core 12, and a controller 100.
[0037] In addition, an external controller 50 such as a RIC (RAN Intelligent Controller) is connected to the in-server data transfer system 1000. The RIC is a component defined by software of a RAN (Radio Access Network) architecture, and realizes the control, optimization, and intelligentization of RAN functions.
[0038] The white arrows in FIG. 2 indicate the data flow. From the perspective of the data transmission and reception processing flow, the controller 100 is configured to be interposed between the application thread 15 and the accelerator 11.
[0039] The controller 100 includes a packet arrival monitoring unit 110, a polling control unit 120, a device virtualization unit 130, a proxy unit 140, an integrated control unit 150, an application control unit 160, a device control unit 170, and an external controller IF 180. The integrated control unit 150 performs integrated control of each unit. The integrated control unit 150 makes a scale-in determination in response to a request from the external controller IF 180 (see FIG. 19 described later). Further, the integrated control unit 150 issues a "discussion object mapping table update request" to the proxy unit 140 (see FIG. 19 described later). The integrated control unit 150 issues a scale-in request to the device control unit 170 (FIG. 19 described later).
[0040] The application control unit 160 controls the application thread 15 according to an instruction from the integrated control unit 150. The device control unit 170 controls the accelerator 11 according to an instruction from the integrated control unit 150. The external controller IF 180 receives an instruction from the external controller 50.
[0041] Hereinafter, the operation of the in-server data transfer system configured as described above will be described. [Sleep Control Operation of Application Thread] FIG. 3 is a diagram for explaining the Sleep control operation of the application thread. The same components as those in FIGS. 1 and 24 are denoted by the same reference numerals. The application thread 15, the packet processing API 14, and the controller 100 are arranged on the user space 30.
[0042] As described above, when there are a plurality of applications APL1, APL2 (see FIG. 1), the influence of increased power consumption is significant. Therefore, the Sleep control operation of the application thread when there are a plurality of applications APL1, APL2 will be described. For each of the plurality of applications APL1, APL2, there are a plurality of application threads 15. The polling control unit 120 (see FIG. 1) of the controller 100 wakes up the application thread 15 when a packet arrives to prompt packet processing, and puts it to sleep when there is no packet.
[0043] The packet arrival monitoring unit 110 (see FIG. 1) of the controller 100 monitors the communication with the accelerator 11 and measures the packet arrival timing. In the reception process by polling, the packet arrival monitoring unit 110 (see FIG. 1) stores the pointer of the data to be transferred in the Ring Buffer 17 (Ring Buffer <1>) using the physical queue 41 (see reference symbol a in FIG. 3).
[0044] Incidentally, conventionally, as shown in FIG. 24, the application thread 15 polled the Ring Buffer 16 and retrieved packets.
[0045] In contrast, in this embodiment, the packet arrival monitoring unit 110 (see FIG. 1) polls the Ring Buffer 17 using the virtual queue 200 (see FIG. 5) instead of the application thread 15, and when there is data to be transferred, acquires the pointer of the corresponding data and performs reception processing (see reference symbol b in FIG. 3).
[0046] The proxy unit 140 stores the pointer of the data to be transferred in the Ring Buffer 18 (Ring Buffer <2>) (= virtual queue 200) instead of the application thread 15 (see reference symbol c in FIG. 3).
[0047] The polling control unit 120 notifies the application thread 15 of an event (see reference symbol d in FIG. 3).
[0048] The application thread 15 acquires the pointer of the corresponding data using the virtual queue 200 when there is data to be transferred (see 3 reference symbol e in the figure).
[0049] Here, the packet arrival monitoring unit 110 polls the physical queue 13 on behalf of all the application threads 15. As a result, even when there are a plurality of application threads 15, the physical queue 13 can be monitored by only one packet arrival monitoring unit 110. For example, in the case of FIG. 23, there are a plurality of (six) application threads 15 processing in parallel for a plurality of applications APL1, APL2, and in the conventional example of FIG. 23, it was necessary to poll the six application threads 15. In contrast, in the present embodiment, for the plurality of application threads 15 existing in parallel, all the physical queues 13 can be monitored only by polling all the physical queues 13 by one packet arrival monitoring unit 110. Therefore, by simple calculation, the power consumption can be suppressed to 1 / 6. The problem can be solved even if not all polling can be eliminated.
[0050] As described above, instead of the application thread 15, the controller 100 monitors the packet arrival of the accelerator 11 on behalf and polls the Ring Buffer 17. The controller 100 notifies the application thread 15 of the packet arrival by a method such as event notification.
[0051] FIG. 4 is a control sequence diagram of the sleep control operation of the application thread in FIG. 3. The application thread 15 notifies the accelerator execution request to the Ring Buffer 17 (S1), and the Ring Buffer 17 notifies this execution request to the Ring Buffer 18 (S2). The Ring Buffer 18 notifies this execution request to the accelerator 11 (S3).
[0052] Here, when the application thread 15 notifies the accelerator execution request to the Ring Buffer 17, the controller 100 puts the application thread 15 to sleep (see reference symbol f in FIG. 4).
[0053] The accelerator 11 executes the request from the accelerator 11 (S4) and transmits the processing result to the Ring Buffer 18 (S5). The packet arrival monitoring unit 110 of the controller 100 monitors the communication with the accelerator 11, and the controller 100 polls the Ring Buffer 18 instead of the application thread 15 between the controller 100 and the Ring Buffer 18 (see reference symbol g in FIG. 4).
[0054] The controller 100 notifies the application thread 15 of the packet arrival by a method such as event notification (see reference symbol h in FIG. 4).
[0055] The Ring Buffer 18 transmits the processing result to the Ring Buffer 17 (S6), and when there is a data transmission request from the application thread 15 that has woken up upon receiving the event notification, the Ring Buffer 17 transmits the processing result of the accelerator 11 to the application thread 15 (S7).
[0056] In this way, the controller 100 can typically monitor the packet arrival from the accelerator 11 by polling. Therefore, it is not necessary for each application thread 15 to monitor the packet arrival by polling, and the power consumption can be reduced.
[0057] Since the controller 100 can monitor the communication of all application threads 15, the polling of all application threads 15 can be made more efficient, and the power consumption of the entire server can be reduced.
[0058] The in-server data transfer system 1000 can suppress the delay by methods such as performing event notification from the controller 100 to the application thread 15 when a packet arrives, and performing event notification based on the timing information obtained from the statistical data of the packet transmission and reception intervals.
[0059] [Communication Proxy with Accelerator] FIG. 5 is a diagram for explaining the communication proxy with the accelerator. The same reference symbols are assigned to the same components as in FIG. 1. So that the controller 100 can communicate with the application 20 as a pseudo-device having an interface equivalent to the existing one, the controller 100 communicates with the application 20 as a pseudo-device having an interface equivalent to the existing one. At this time, a virtual queue 200 is provided to the application 20 instead of the physical queue 13. The device emulation unit 130 (FIGS. 1 and 2) connects to the application with an interface equivalent to the existing one and provides a virtual queue 200 instead of the physical queue 13 of the device.
[0060] FIG. 6 is a diagram showing a correspondence table 210 between the thread ID of the proxy unit 140 and the physical queue 13. A physical queue 13 corresponding to each thread ID "1001..." is assigned. For example, when there are physical queues (0 to 5) in order from the left of the IP core 12 of the accelerator 11 shown in FIG. 5, the thread ID "1001" is assigned to the leftmost physical queue (0) of the IP core 12 of the accelerator 11 shown in FIG. 5. It is also possible to assign a plurality of physical queues to a thread ID. For example, the thread ID "1003" is assigned to the physical queues (2, 3) of the IP core 12 of the accelerator 11 shown in FIG. 5. This example shows that the connection branches to the physical queues (2, 3) when viewed from the virtual queue. Conversely, the thread IDs "2001" and "2002" are both assigned to the physical queue (4). This example shows that the connection converges to the physical queue (4). Note that conventionally, the thread ID and the physical queue correspond one-to-one, and there is no such branching / aggregation as described above.
[0061] FIG. 7 is a diagram showing the virtual queue table 220 of the proxy unit 140. The virtual queue table 220 describes which virtual queue the thread ID corresponds to. The thread ID and the virtual queue have a one-to-one correspondence. For example, thread ID "1001" corresponds to virtual queue (0), and thread ID "1002" corresponds to virtual queue (1). Note that Proc_type (process type) defines what type of processing is to be performed on the accelerator. For example, Proc_type "1" can be connected to IP core 12 corresponding to processing of type 1, and Proc_type "2" can be connected to IP core 12 corresponding to processing of type 2, and so on. APL describes the application type when a thread corresponds to multiple applications. "Status" describes whether the thread is active (Used) or inactive (blank).
[0062] Figure 8 is a diagram showing the argument mapping table 230 of the proxy unit 140. The argument mapping table 230 assigns physical queues (0 to 3) to virtual queues (0 to 4). For example, physical queue 0 is assigned to virtual queue (0), and physical queue 1 is assigned to virtual queue (1). Note that it is also possible to assign multiple physical queues (physical queues 2, 3) to virtual queue (3).
[0063] Figure 9 is a diagram showing the physical queue table 240 of the proxy unit 140. The IP core ID is the ID of the core that actually performs processing. For each IP core ID, a physical queue, Proc_type, device ID, and "status" are associated. Proc_type describes that connections cannot be made unless the processing is of the same type. The device ID indicates the type of accelerator processed by the accelerator. For example, if the device ID is "1", it describes that accelerator 1 is corresponding.
[0064] [Polling Control] <Wake-up Judgment Using Packet Arrival Information> Figure 10 is a flowchart showing the wake-up judgment process using packet arrival information in polling control. In step S11, the polling control unit 120 determines whether the packet arrival monitoring unit 110 has confirmed the arrival of a packet. If the arrival of the packet cannot be confirmed (S11: No), the process returns to step S11 to wait for the arrival of the packet.
[0065] If the arrival of the packet can be confirmed (S11: Yes), in step S12, the polling control unit 120 sends a wake-up instruction to the application thread 15 corresponding to the virtual queue 200 (in the following description, since the application thread 15 is a thread in polling control, it is denoted as the polling thread 15).
[0066] In step S13, the polling control unit 120 determines whether the polling thread 15 has woken up and harvested the packet. If the packet has not been harvested (S13: No), the process returns to step S13 to wait for packet harvesting.
[0067] If the packet has been harvested (S13: Yes), in step S14, the polling control unit 120 instructs the polling thread 15 corresponding to the virtual queue 200 to Sleep until the next harvest and ends the processing of this flow.
[0068] <Wake-up determination using timing information (statistics)> FIG. 11 is a flowchart showing a wake-up determination process using timing information (statistics) in polling control. In step S21, the polling control unit 120 selects the virtual queue 200. In step S22, the polling control unit 120 checks the largest one among the packet timers 250 of the virtual queue 200 (see FIG. 12).
[0069] FIG. 12 is a diagram showing the packet timer 250 of the virtual queue 200 in a table. The virtual queue and timer are set for each packet ID. By setting a timer for each virtual queue 200, it is possible to set in advance the wake-up timing when a packet arrives, and the effectiveness of Sleep control can be achieved.
[0070] Returning to FIG. 11, in step S23, the polling control unit 120 determines whether the packet timer 250 exceeds the threshold value of the virtual queue 200.
[0071] FIG. 13 is a diagram showing the threshold table 260 of the virtual queue 200. Set a threshold value for whether to perform pruning for each virtual queue. For example, it is set that the virtual queue (0) does not perform pruning until the threshold value "200" is reached.
[0072] Returning to FIG. 11, when the packet timer 250 does not exceed the threshold value of the virtual queue 200 (S23: No), the process returns to step S21 and steps S21 and S22 are repeated until the packet timer 250 exceeds the threshold value of the virtual queue 200. The threshold value of the packet timer 250 is set separately from the statistical information from packet transmission to reception.
[0073] When the packet timer 250 exceeds the threshold value of the virtual queue 200 (S23: Yes), in step S24, the polling control unit 120 sends a wake-up instruction to the polling thread corresponding to the virtual queue 200.
[0074] In step S25, the polling control unit 120 determines whether the thread has woken up and pruned the packet. If the packet has not been pruned (S25: No), the process returns to step S25 and waits for packet pruning.
[0075] When the thread has woken up and pruned the packet (S25: Yes), in step S26, the polling control unit 120 instructs the polling thread 15 corresponding to the virtual queue 200 to sleep until the next pruning and ends the processing of this flow.
[0076] [Packet Transmission] Figure 14 is a control sequence diagram for packet transmission. This packet transmission has a <polling pattern> and a <non-polling pattern>. Note that this polling is the polling between the packet arrival monitoring unit 110 and the virtual queue 200.
[0077] <polling pattern> The application thread 15 transmits the pointer information of the packet to the virtual queue 200 (S101). On the other hand, the packet arrival monitoring unit 110 confirms the packet arrival by polling (S151). The virtual queue 200 notifies the packet arrival monitoring unit 110 of "packet arrival" and "packet pointer information" (S152).
[0078] The packet arrival monitoring unit 110 receives "packet arrival" and "packet pointer information" from the virtual queue 200 and transmits the "packet pointer information" to the proxy unit 140 (S105).
[0079] The proxy unit 140 receives the "packet pointer information" and checks the physical queue 13 corresponding to the virtual queue 200 (S106). The proxy unit 140 transmits the "packet pointer information" to the physical queue 13 (S107).
[0080] This packet transmission according to the <polling pattern> (polling present pattern) does not notify the controller 100 like the following <non-polling pattern>.
[0081] <non-polling pattern> The same step numbers are assigned to the same procedures as the above <polling pattern>. The application thread 15 transmits the pointer information of the packet to the virtual queue 200 (S101) and notifies the packet arrival monitoring unit 110 of "packet arrival" (S102). Upon receiving the "Packet Arrival" notification from the application thread 15, the packet arrival monitoring unit 110 performs a packet arrival confirmation for the virtual queue 200 (S103). The virtual queue 200 transmits "Packet Arrival" and "Packet Pointer Information" to the packet arrival monitoring unit 110 (S104).
[0082] Upon receiving "Packet Arrival" and "Packet Pointer Information" from the virtual queue 200, the packet arrival monitoring unit 110 transmits the "Packet Pointer Information" to the proxy unit 140 (S105).
[0083] Upon receiving the "Packet Pointer Information", the proxy unit 140 checks the physical queue 13 corresponding to the virtual queue 200 (S106). The proxy unit 140 transmits the "Packet Pointer Information" to the physical queue 13 (S107).
[0084] [Packet Reception and Polling Control] Figure 15 is a control sequence diagram of packet reception and polling. The packet arrival monitoring unit 110 checks for packet arrival at the physical queue 13 by polling (S111). The physical queue 13 transmits "Packet Arrival" and "Packet Pointer Information" to the packet arrival monitoring unit 110 (S112).
[0085] Upon receiving "Packet Arrival" and "Packet Pointer Information" from the physical queue 13, the packet arrival monitoring unit 110 transmits the "Packet Pointer Information" to the proxy unit 140 (S113).
[0086] Upon receiving the "Packet Pointer Information", the proxy unit 140 checks the virtual queue 200 corresponding to the physical queue 13 (S114). The proxy unit 140 transmits the "Packet Pointer Information" to the virtual queue 200 (S115), and transmits the "Corresponding Virtual Queue Information" to the packet arrival monitoring unit 110 (S116).
[0087] Upon receiving the "corresponding virtual queue information" from the proxy unit 140, the packet arrival monitoring unit 110 transmits "packet arrival detected" to the polling control unit 120 (S117). Upon receiving "packet arrival detected" from the packet arrival monitoring unit 110, the polling control unit 120 issues an "event notification" to the application thread 15 (S118).
[0088] The application thread 15 performs a "wake-up process" (S119). The application thread 15 requests "packet pointer information" from the virtual queue 200 (S120). In response to this request, the virtual queue 200 transmits "packet pointer information" to the virtual queue 200 (S121).
[0089] [IP Core Scale-in] Figure 16 is a control sequence diagram of IP core scale-in. The external controller IF 180 issues a "scale-in determination request" to the integrated control unit 150 (S131). The integrated control unit 150 performs a scale-in determination (S132). The integrated control unit 150 issues a "logical object mapping table update request" to the proxy unit 140 (S133). The proxy unit 140 updates the logical object mapping table 230 (see Figure 8) (refer to the logical object mapping table update in Figure 18) and notifies the integrated control unit 150 to that effect (S135).
[0090] In the IP core scale-in process, take the case where IP core #2 among the IP cores 12 is scaled in as an example (see Figure 17). The integrated control unit 150 issues an "IP core #2 scale-in request" to the device control unit 170 (S136). The device control unit 170 performs a power-off operation on IP core #2 (S137) and notifies the integrated control unit 150 of "IP core #2 scale-in completed" (S138).
[0091] FIG. 17 is a diagram for explaining the IP core scale-in of FIG. 16. The same components as those in FIG. 5 are denoted by the same reference numerals. By the IP core scale-in process of FIG. 16, IP core #2 is scale-in (dashed arrow in FIG. 17 i reference).
[0092] FIG. 18 is a flowchart showing the argument mapping table update process. In step S31, the proxy unit 140 determines whether there is a virtual queue 200 corresponding to the physical queue to be scale-in. If there is no corresponding virtual queue 200 (S31: No), the argument mapping table update process according to this flow ends normally.
[0093] If there is a corresponding virtual queue 200 (S 31 : Yes), in step S32, the proxy unit 140 selects one virtual queue 200.
[0094] In step S33, the proxy unit 140 determines whether there is a physical queue 13 with the same proc_type as that of the virtual queue 200. If there is no physical queue 13 with the same proc_type as that of the virtual queue 200 (S33: No), the argument mapping table update process according to this flow ends with NG.
[0095] If there is a physical queue 13 with the same proc_type as that of the virtual queue 200 in step S33 above (S 33 : Yes), judging that re-association of the virtual queue 200 and the physical queue 13 is to be performed, in step S34, the proxy unit 140 selects a device (however, according to the policy) if there are multiple devices. This policy is, for example, to concentrate on the device with the highest load as much as possible (aggregation), or to disperse to the device with the lowest load (load distribution).
[0096] In step S35, when there are multiple physical queues 13, the proxy unit 140 selects (according to a policy) a physical queue 13. This policy is, for example, to concentrate (aggregate) as much as possible on the physical queue with the highest load or to distribute (load balance) to the physical queue with the lowest load.
[0097] In step S36, the proxy unit 140 changes the value of the cell of the physical queue 13 associated with the virtual queue 200 to 1 from the argument mapping table 230 (see FIG. 8).
[0098] In step S37, the proxy unit 140 changes the value of the cell of the virtual queue 200 and the scale-in physical queue 13 to 0 from the argument mapping table 230 and returns to step S31 above.
[0099] [Thread scale-in] FIG. 19 is a control sequence diagram of thread scale-in. The external controller IF180 makes a "scale-in determination request" to the integrated control unit 150 (S141). The integrated control unit 150 makes a scale-in determination (S142). The integrated control unit 150 makes an "argument mapping table update request" to the proxy unit 140 (S143). The proxy unit 140 updates the argument mapping table 230 (see FIG. 8) (S144) (see the argument mapping table update in FIG. 21) and notifies the integrated control unit 150 to that effect (S145).
[0100] In the thread scale-in process, take as an example the case where thread #2 among the application threads 15 is scale-in (see FIG. 20). The integrated control unit 150 makes a "thread #2 scale-in request" to the device control unit 170 (S146). The device control unit 170 performs a polling stop operation for thread #2 (S147) and notifies the integrated control unit 150 of "thread #2 scale-in completion" (S148).
[0101] FIG. 20 is a diagram for explaining thread scale-in. The same components as those in FIGS. 5 and 17 are denoted by the same reference numerals. By the thread scale-in process of FIG. 19, the thread #2 shown in FIG. 20 is scaled in (see the dashed arrow j in FIG. 20).
[0102] FIG. 21 is a flowchart showing the processing for updating the object mapping table. In step S41, the proxy unit 140 determines whether there is a physical queue 13 corresponding to the virtual queue to be scaled in. If there is no corresponding physical queue 13 (S41: No), the processing for updating the object mapping table by this flow ends normally.
[0103] If there is a corresponding physical queue 13 (S41: Yes), in step S42, the proxy unit 140 selects one physical queue 13.
[0104] In step S43, the proxy unit 140 determines whether there is a virtual queue 200 having the same proc_type as that of the physical queue 13. If there is no virtual queue 200 having the same proc_type as that of the physical queue 13 (S43: No), the process proceeds to step S47.
[0105] If there is a virtual queue 200 having the same proc_type as that of the physical queue 13 in step S43 (S43: Yes), the proxy unit 140 determines to re-associate the virtual queue 200 and the physical queue 13, and in step S44, selects an application if there are multiple applications (however, according to a policy). This policy is, for example, to concentrate (aggregate) on the APL with the highest load as much as possible, or to distribute (load distribution) to the APL with a small load.
[0106] In step S45, the proxy unit 140 selects a virtual queue 200 if there are multiple virtual queues 200 (however, according to a policy). This policy is, for example, to concentrate (aggregate) on the virtual queue with the highest load as much as possible, or to distribute (load distribution) to the virtual queue with a small load.
[0107] In step S46, the proxy unit 140 changes the cell value of the virtual queue 200 associated with the physical queue 13 to 1 from the argument mapping table 230 (see FIG. 8).
[0108] In step S47, the proxy unit 140 changes the cell values of the physical queue 13 and the scale-in virtual queue 200 to 0 from the argument mapping table 230 and returns to the above step S 41 to return.
[0109] [Hardware Configuration] The controller (in-server data transfer device) 100 according to the above embodiment is realized by a computer 900 configured as shown in FIG. 22, for example. FIG. 22 is a hardware configuration diagram showing an example of a computer 900 that realizes the functions of the controller 100. The computer 900 includes a CPU 901, a RAM 902, a ROM 903, an HDD 904, an accelerator 905, an input / output interface (I / F) 906, a media interface (I / F) 907, and a communication interface (I / F: Interface) 908. The accelerator 905 corresponds to the accelerator 11 in FIGS. 1, 3, and 5.
[0110] The accelerator 905 is an accelerator (device) 11 (FIGS. 1 and 5) that processes at high speed at least one of the data from the communication I / F 908 and the data from the RAM 902. Note that, as the accelerator 905, a type (look-aside type) that returns the execution result to the CPU 901 or the RAM 902 after executing the processing from the CPU 901 or the RAM 902 may be used. On the other hand, as the accelerator 905, a type (in-line type) that enters between the communication I / F 908 and the CPU 901 or the RAM 902 and performs processing may be used.
[0111] The accelerator 905 is connected to an external device 915 via the communication I / F 908. The input / output I / F 906 is connected to an input / output device 916. The media I / F 907 reads and writes data from and to a recording medium 917.
[0112] The CPU 901 operates based on programs stored in the ROM 903 or the HDD 904, and controls each part of the controller 100 shown in FIGS. 1, 3, and 5 by executing the programs (also called applications or their abbreviated apps) read into the RAM 902. And this program can also be distributed via a communication line or recorded on a recording medium 917 such as a CD-ROM and distributed. The ROM 903 stores a boot program executed by the CPU 901 when the computer 900 is started up, programs dependent on the hardware of the computer 900, and the like.
[0113] The CPU 901 controls an input / output device 916 including an input part such as a mouse and a keyboard and an output part such as a display and a printer via the input / output I / F 906. The CPU 901 acquires data from the input / output device 916 via the input / output I / F 906 and also outputs the generated data to the input / output device 916. Note that, together with the CPU 901, a GPU (Graphics Processing Unit) or the like may be used as a processor.
[0114] The HDD 904 stores programs executed by the CPU 901 and data used by the programs. The communication I / F 908 receives data from other devices via a communication network (for example, NW (Network)) and outputs it to the CPU 901, and also transmits data generated by the CPU 901 to other devices via the communication network.
[0115] The media I / F 907 reads the program or data stored in the recording medium 917 and outputs it to the CPU 901 via the RAM 902. The CPU 901 loads the program related to the target process from the recording medium 917 onto the RAM 902 via the media I / F 907 and executes the loaded program. The recording medium 917 is an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto Optical disk), a magnetic recording medium, a conductor memory tape medium, or a semiconductor memory, etc.
[0116] For example, when the computer 900 functions as the in-server data transfer device 100 configured as one device according to this embodiment, the CPU 901 of the computer 900 realizes the functions of the controller (in-server data transfer device) 100 by executing the program loaded onto the RAM 902. Also, the data in the RAM 902 is stored in the HDD 904. The CPU 901 reads and executes the program related to the target process from the recording medium 917. Additionally, the CPU 901 may read the program related to the target process from another device via the communication network.
[0117] [Effect] As described above, the in-server data transfer device according to this embodiment is an in-server data transfer device (controller 100) that transfers data from a device to the application 20 when using a device including the accelerator 11 in the application 20, and includes a packet arrival monitoring unit 110 that monitors the communication between the device and the application thread 15 corresponding to the application 20 and measures the packet arrival timing, and a polling control unit 120 that wakes up the application thread 15 at the arrival of the packet detected by the packet arrival monitoring unit 110 to perform packet processing and puts it to sleep while there is no packet arrival.
[0118] By doing so, a controller (in-server data transfer device) 100 equipped with a packet arrival monitoring unit 110 and a polling control unit 120 is configured to mediate between an application 20 and an accelerator 11.
[0119] Since the controller 100 can typically monitor the arrival of packets from the accelerator 11 by polling, it is not necessary for each application thread 15 to monitor the arrival of packets by polling, and power consumption can be reduced.
[0120] Since the controller 100 can monitor the communications of all application threads 15, it is possible to improve the polling efficiency of all application threads 15 and reduce the power consumption of the entire server.
[0121] Also, the controller 100 can suppress delays by, for example, performing event notifications to the application thread 15 when a packet arrives, or performing event notifications based on timing information from statistical data on the packet transmission / reception interval.
[0122] Furthermore, when multiple network applications that accelerate data processing using an accelerator, such as vRAN (virtual Radio Access Network), are installed on the same server, applying this system can achieve the following effects.
[0123] Power saving: Polling in each application thread 15 becomes unnecessary, and the power consumption of the CPU core can be significantly reduced.
[0124] Efficiency (dynamic aspect): In cases where demand changes significantly, such as between day and night, power consumption can be reduced by stopping unnecessary resources through IP core scale-in or thread scale-in.
[0125] Efficiency (static aspect): When there is spare capacity in the accelerator 11, the number of users that can be accommodated by the accelerator can be increased by associating the virtual queues of multiple applications with one physical queue.
[0126] The in-server data transfer device (controller) 100 according to the present embodiment corresponds to the physical queue of the device, and by using the virtual queue 200 instead of the physical queue and causing the application thread 15 to perform packet processing, it includes a device virtualization unit 130 that connects the device to the application 20, and a proxy unit 140 that dynamically changes the connection between the virtual queue 200 and the physical queue 13.
[0127] In this way, it includes the device virtualization unit 130 and the proxy unit 140, and the device virtualization unit 130 provides the virtual queue 200 instead of the physical queue 13 to the application. Since it appears to be communicating with the actual accelerator 11 from the application 20, no modification is required.
[0128] The controller 100 manages the association between the application thread 15 and the IP cores 12 of the multiple accelerators 11, and communicates with the accelerator 11 instead of the application 20. The controller 100 can perform dynamic resource changes such as increasing (scaling out) or decreasing (scaling in) the IP cores 12 and application threads 15 according to demand by changing the association between the physical queue 13 and the virtual queue 200 and simultaneously instructing the increase or decrease of the IP cores 12 and application threads 15.
[0129] Furthermore, by associating multiple physical queues (1:N) with one application thread 15 or one physical queue (M:1) with multiple application threads, more flexible resource management and resource aggregation by overlay become possible.
[0130] In addition, among the respective processes described in the above embodiments, all or part of the processes described as being automatically performed can also be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, regarding the process procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified. Moreover, each component of each illustrated device is conceptually functional, and it is not necessarily physically configured as illustrated. That is, the specific form of the distribution and integration of each device is not limited to that shown, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads, usage situations, etc.
[0131] In addition, each of the above configurations, functions, processing units, processing means, etc. may be realized in hardware by designing part or all of them, for example, with an integrated circuit. Also, each of the above configurations, functions, etc. may be realized by software for a processor to interpret and execute a program for realizing each function. Information such as programs, tables, and files for realizing each function can be held in a memory, a recording device such as a hard disk or an SSD (Solid State Drive), or a recording medium such as an IC (Integrated Circuit) card, an SD (Secure Digital) card, or an optical disk.
Explanation of Reference Numerals
[0132] 10 HW 11,905 Accelerator (Device) 12 IP Core (Device Core) 14 Packet Processing API 17,18 Ring Buffer 15 App Thread (Polling Thread) 20 Application (APL) 100 Controller (In-Server Data Transfer Device) 110 Packet Arrival Monitoring Unit 120 Polling Control Unit 130 Device Emulation Unit 140 Proxy Unit 150 Integrated Control Unit 160 Application Control Unit 170 Device Control Unit 180 External Controller IF 220 Virtual Queue Table 230 Argument Mapping Table 240 Physical Queue Table 250 Packet Timer 260 Threshold Table 1000 In-Server Data Transfer System APL1, APL2 Applications
Claims
1. A server internal data transfer device that performs data transfer from the device to the application when using a device including an accelerator in an application, comprising: a data arrival monitoring unit that monitors communication between the device and an application thread corresponding to the application and measures data arrival timing; a polling control unit that wakes up the application thread to receive data when data arrives as detected by the data arrival monitoring unit, and puts it to sleep while no data arrives; a device emulation unit that connects the device to the application by causing the device and an application thread corresponding to the application to perform data processing using a virtual queue that substitutes for the physical queue of the device; and a proxy unit that dynamically changes the connection between the virtual queue and the physical queue. A server internal data transfer device characterized by the above.
2. A server internal data transfer device that performs data transfer from the device to the application when using a device including an accelerator in an application, comprising: a device emulation unit that connects the device to the application by causing the device and an application thread corresponding to the application to perform data processing using a virtual queue that substitutes for the physical queue of the device; and a proxy unit that dynamically changes the connection between the virtual queue and the physical queue, and correspondence information that associates the application thread with one or more of the physical queues, description information that describes which virtual queue the application thread corresponds to, and object mapping information that assigns a physical queue to the virtual queue, and wherein the proxy unit when a decrease in the application thread is instructed, refers to the object mapping information to determine whether there is a virtual queue corresponding to the physical queue to be scaled in, and if there is a corresponding virtual queue, dynamically decreases the number of virtual queues connected to the physical queue based on the correspondence information and the description information, and stops the application thread using the virtual queue. A server internal data transfer device characterized by the above.
3. A server internal data transfer device that transfers data from a device including an accelerator to an application when the device is used by the application, comprising: A device virtualization unit that connects the device to the application by causing data processing to be performed on an application thread corresponding to the device and the application using a virtual queue that substitutes for the physical queue of the device; A proxy unit that dynamically changes the connection between the virtual queue and the physical queue, and Association information that associates the application thread with one or more of the physical queues, Description information that describes which virtual queue the application thread corresponds to, Object mapping information that assigns the physical queue to the virtual queue, and having The proxy unit When an increase in the application thread is instructed, determines whether there is a virtual queue corresponding to the physical queue to be scaled out with reference to the object mapping information, and if there is a corresponding virtual queue, dynamically increases the number of virtual queues connected to the physical queue based on the association information and the description information, and resumes the application thread using the virtual queue A server internal data transfer device characterized by the above.
4. A server internal data transfer device that transfers data from a device including an accelerator to an application when the device is used by the application, comprising: A device virtualization unit that connects the device to the application by causing data processing to be performed on an application thread corresponding to the device and the application using a virtual queue that substitutes for the physical queue of the device; A proxy unit that dynamically changes the connection between the virtual queue and the physical queue, and Association information that associates the application thread with one or more of the physical queues, Description information that describes which virtual queue the application thread corresponds to, Object mapping information that assigns the physical queue to the virtual queue, and Physical queue information in which a physical queue, Proc_type, and device ID are associated for each IP core of the device, and having The proxy unit When a decrease in the IP core is instructed, referring to the object mapping information, it is determined whether there is a virtual queue corresponding to the physical queue to be scaled in. If there is a corresponding virtual queue, based on the correspondence information, the description information, and the physical queue information, the number of physical queues connected to the virtual queue is dynamically decreased, and the IP core of the device using the physical queue is stopped. A data transfer device within a server, characterized in that.
5. A data transfer device within a server that transfers data from a device to an application when using a device including an accelerator in an application, using a virtual queue that substitutes for the physical queue of the device, and causing data processing to be performed on an application thread corresponding to the device and the application, a device virtualization unit that connects the device to the application; a proxy unit that dynamically changes the connection between the virtual queue and the physical queue, and correspondence information that associates the application thread with one or more of the physical queues; description information that describes which virtual queue the application thread corresponds to; object mapping information that assigns a physical queue to the virtual queue; for each IP core of the device, physical queue information in which a physical queue, Proc_type, and device ID are associated; The proxy unit is When an increase in the IP core is instructed, referring to the object mapping information, it is determined whether there is a virtual queue corresponding to the physical queue to be scaled out. If there is a corresponding virtual queue, based on the correspondence information, the description information, and the physical queue information, the number of physical queues connected to the virtual queue is dynamically increased, and the IP core of the device using the physical queue is restarted. A data transfer device within a server, characterized in that.
6. A data transfer method within a server of a data transfer device within a server that transfers data from a device to an application when using a device including an accelerator in an application, The data transfer device within the server is a step of monitoring communication between the device and an application thread corresponding to the application and measuring the data arrival timing; When data arrival is detected, waking up the application thread to receive data and putting it to sleep while there is no data arrival; Using a virtual queue that substitutes for the physical queue of the device to cause the device and the application thread corresponding to the application to perform data processing, thereby connecting the device to the application; Dynamically changing the connection between the virtual queue and the physical queue; and Executing the above steps, a method for transferring data within a server, characterized in that;
7. When using a device including an accelerator in an application, on a computer as a data transfer device within a server that transfers data from the device to the application, Data arrival monitoring means for monitoring communication between the device and the application thread corresponding to the application and measuring data arrival timing; Polling control means for waking up the application thread to receive data when data arrival is detected by the data arrival monitoring means and putting it to sleep while there is no data arrival; Device emulation means for using a virtual queue that substitutes for the physical queue of the device to cause the device and the application thread corresponding to the application to perform data processing, thereby connecting the device to the application; Proxy means for dynamically changing the connection between the virtual queue and the physical queue; and A program for causing the above to be executed.
Citation Information
Patent Citations
Connection control system of virtual machine and connection control method of virtual machine
JP2018032156A
Techniques for Received Packet Processing and Associated Power Management in Network Devices
JP2018507457A
Intra-server delay control device, intra-server delay control method, and program
WO2021130828A1