Intra-server data transfer apparatus, intra-server data transfer method, and program

The server internal data transfer device with a sleep control management unit addresses high CPU usage and power consumption in existing data transfer methods by implementing a polling model that reduces CPU usage and maintains low latency through scheduled sleep and wake-up mechanisms.

JP2025100826AActive Publication Date: 2025-07-03NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2025070463
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-03
Estimated Expiration
2041-07-19

AI Technical Summary

Technical Problem

Existing data transfer methods in server environments, such as interrupt and polling models, suffer from high CPU usage and power consumption due to constant monitoring for data arrival, leading to increased latency and inefficiency.

Method used

A server internal data transfer device with a sleep control management unit that manages data arrival schedules and performs sleep control based on data arrival timing, using a polling model to reduce CPU usage and power consumption while maintaining low latency.

Benefits of technology

The solution reduces CPU usage and achieves power saving while maintaining low latency by putting the data transfer thread to sleep when no data is arriving and waking it up at the scheduled time, thereby optimizing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025100826000001_ABST
    Figure 2025100826000001_ABST
Patent Text Reader

Abstract

To enable power saving by reducing CPU usage rate while maintaining low latency performance.SOLUTION: An intra-server data transfer device 200 includes: a sleep control management unit 210 that manages a data arrival schedule and performs sleep control in accordance with data arrival timing, where the sleep control management unit 210 sleeps a thread that monitors arrival of data using a polling model based on the data arrival timing in a data flow having information on specific data arrival timing, and performs operation of releasing sleep of the thread based on the data arrival timing.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an in-server data transfer device, an in-server data transfer method, and a program.

Background Art

[0002] Against the backdrop of the advancement of virtualization technologies such as NFV (Network Functions Virtualization), systems are being constructed and operated for each service. Further, from the form of constructing a system for each service, the service functions are divided into reusable module units and operated on an independent virtual machine (VM: Virtual Machine or container, etc.) environment, thereby enhancing the operability by using them as needed like parts. This form is called SFC (Service Function Chaining) and is becoming the mainstream.

[0003] As a technology for constructing a virtual machine, a hypervisor environment composed of Linux (registered trademark) and KVM (kernel-based virtual machine) is known. In this environment, a Host OS (the OS installed on a physical server is called a Host OS) incorporating a KVM module operates in a memory area different from the user space called the kernel space as a hypervisor. In this environment, a virtual machine operates in the user space, and a Guest OS (the OS installed on the virtual machine is called a Guest OS) operates within the virtual machine.

[0004] Unlike the physical server on which the Host OS runs, in the virtual machine where the Guest OS runs, all HW (hardware), including network devices (represented by Ethernet (registered trademark) card devices, etc.), becomes register control necessary for interrupt processing from the HW to the Guest OS and writing from the Guest OS to the hardware. In such register control, since notifications and processes that the physical hardware should originally execute are emulated pseudo - software, the performance is generally lower than that of the Host OS environment.

[0005] In this performance degradation, there is a technology that reduces the emulation of HW and improves the communication performance and versatility with a high - speed and unified interface, especially for the Guest OS to the Host OS and external processes existing outside the self - virtual machine. As this technology, a device abstraction technology called virtio, that is, para - virtualization technology, has been developed and has already been incorporated into many general - purpose OSs such as Linux (registered trademark) and FreeBSD (registered trademark) and is currently in use (see Patent Documents 1 and 2).

[0006] In virtio, regarding data input / output such as console, file input / output, and network communication, data exchange by a queue designed as a ring buffer is defined by queue operations as a transport for single - direction transfer of transfer data. And by preparing the number and size of queues suitable for each device using the virtio queue specification at the time of Guest OS startup, communication between the Guest OS and the outside of the self - virtual machine can be realized only by queue operations without executing hardware emulation.

[0007] [Packet Transfer by Interrupt Model (Example of General - Purpose VM Configuration)] Figure 19 is a diagram for explaining packet transfer by the interrupt model in a server virtualization environment of a general - purpose Linux kernel (registered trademark) and VM configuration. HW10 has a NIC (Network Interface Card) 11 (physical NIC) (interface section) and communicates data with the data processing APL (Application) 1 on the user space 60 via a virtual communication path constructed by the Host OS 20, the hypervisor KVM 30 for constructing virtual machines, the virtual machines (VM1, VM2) 40, and the Guest OS 50. In the following description, as shown by the thick arrows in Fig. 19, the data flow in which the data processing APL 1 receives packets from HW10 is referred to as Rx-side reception, and the data flow in which the data processing APL 1 transmits packets to HW10 is referred to as Tx-side transmission.

[0008] The Host OS 20 has a kernel 21, a Ring Buffer 22, and a Driver 23. The kernel 21 has a vhost-net module 221A which is a kernel thread, a tap device 222A, and a virtual switch (br) 223A.

[0009] The tap device 222A is a kernel device of the virtual network and is supported by software. The virtual machine (VM1) 40 can communicate between the Guest OS 50 and the Host OS 20 via the virtual switch (br) 223A created in the virtual bridge. The tap device 222A is a device connected to the virtual NIC (vNIC) of the Guest OS 50 created in this virtual bridge.

[0010] The host OS 20 copies the configuration information (such as the size of the shared buffer queue, the number of queues, identifiers, and start address information for accessing the ring buffer) built within the virtual machine of the guest OS 50 to the vhost-net module 221A, and constructs the information of the endpoints on the virtual machine side inside the host OS 20. This vhost-net module 221A is a kernel-level backend for virtio networking, and can reduce the virtualization overhead by moving the virtio packet processing task from the user area (user space) to the vhost-net module 221A of the kernel 21.

[0011] The guest OS 50 has the guest OS (Guest1) installed on the virtual machine (VM1) and the guest OS (Guest2) installed on the virtual machine (VM2), and the guest OS 50 (Guest1, Guest2) operates within the virtual machines (VM1, VM2) 40. Taking Guest1 as an example of the guest OS 50, the guest OS 50 (Guest1) has a kernel 51, a ring buffer 52, and a driver 53, and the driver 53 includes a virtio-driver 531.

[0012] Specifically, as a PCI (Peripheral Component Interconnect) device, there are virtio devices for the console, file input / output, and network communication respectively within the virtual machine (the console is a virtio-console, file input / output is a virtio-blk, and the network is a device called virtio-net, and the driver of the corresponding OS is defined by a virtio queue). When the guest OS starts up, two data transfer endpoints (transmission and reception endpoints) are created between the guest OS and the other side, and a parent-child relationship for data transmission and reception is constructed. In most cases, the parent-child relationship is composed of the virtual machine side (child side) and the guest OS (parent side).

[0013] The child side exists as the configuration information of the devices within the virtual machine, and requests from the parent side the sizes of each data area, the number of combinations of required endpoints, and the types of devices. The parent side allocates and secures the memory for the shared buffer queue for storing and transferring the necessary data according to the requests of the child side, and returns the address to the child side so that the child side can access it. Regarding the operations of the shared buffer queue required for data transfer, all are common in virtio and are executed with the consent of both the parent side and the child side. Furthermore, the size of the shared buffer queue is also agreed upon by both sides (that is, it is determined for each device). Thereby, it becomes possible for both the parent side and the child side to operate the queue shared between them just by transmitting the address to the child side.

[0014] Since the shared buffer queue prepared in virtio is for single - direction use, for example, a virtual network device called a virtio - net device is composed of three Ring Buffers 52 for transmission, reception, and control. The communication between the parent and the child is realized by writing to the shared buffer queue and notifying of buffer updates. After writing to the Ring Buffer 52, it notifies the other side. When the other side receives the notification, it uses the common operations of virtio to confirm which shared buffer queue contains how much new data and retrieves the new buffer area. Thereby, the transfer of data from the parent to the child or from the child to the parent is established.

[0015] As described above, by sharing the Ring Buffer 52 for data exchange between the parent and the child and the operation methods (common in virtio) for each ring buffer, communication between the Guest OS 50 and the outside without the need for hardware emulation is realized. Thereby, compared with conventional hardware emulation, it is possible to realize the high - speed transmission and reception of data between the Guest OS 50 and the outside.

[0016] When the Guest OS 50 in the virtual machine communicates with the outside, the child side needs to connect to the outside and send and receive data as an intermediary between the child side and the parent side. For example, the communication between the Guest OS 50 and the Host OS 20 is one such example. Here, when the outside is the Host OS 20, there are two existing communication methods.

[0017] The first method (hereinafter referred to as external communication method 1) is to construct a child-side endpoint within the virtual machine and connect the communication between the Guest OS 50 and the virtual machine and the communication endpoint provided by the Host OS 20 (usually called a tap / tun device) within the virtual machine. Through this connection, the following connection is constructed to realize the communication from the Guest OS 50 to the Host OS 20.

[0018] At this time, the Guest OS 50 operates in a memory area of the user space with different permissions from the memory area of the kernel space where the tap driver or the Host OS 20 operates. Therefore, at least one memory copy will occur for the communication from the Guest OS 50 to the Host OS 20.

[0019] The second method (hereinafter referred to as external communication method 2) is a technology called vhost-net as a means to solve this problem. In vhost-net, the configuration information of the parent side (such as the size of the shared buffer queue, the number of queues, the identifier, the head address information for accessing the ring buffer, etc.) once constructed within the virtual machine is copied to the vhost-net module 221A inside the Host OS 20, and the information of the child-side endpoint is constructed inside the host. This construction enables the operation of the shared buffer queue to be directly carried out between the Guest OS 50 and the Host OS 20. As a result, the number of copies is substantially zero, and compared with virtio-net, the number of copies is one less than that of external communication method 1, so data transfer can be realized faster compared with external communication method 1.

[0020] In this way, in the Host OS 20 and Guest OS 50 connected by virtio, the packet transfer process can be accelerated by reducing the number of memory copies related to virtio-net.

[0021] Note that since kernel v4.10 (2017.2~), the specification of the tap interface has changed, and the packets inserted from the tap device are completed within the same context as the process of copying the packets to the tap device. As a result, the generation of software interrupts (softIRQ) has disappeared.

[0022] [Packet Transfer by Polling Model (Example of DPDK)] The method of connecting and coordinating multiple virtual machines is called Inter-VM Communication. In large-scale environments such as data centers, virtual switches have been standardly used for connections between VMs. However, since it is a method with a large communication delay, a faster method has been newly proposed. For example, a method using special hardware called SR-IOV (Single Root I / O Virtualization), or a software method using Intel DPDK (Intel Data Plane Development Kit) (hereinafter referred to as DPDK), which is a high-speed packet processing library, has been proposed (see Non-Patent Document 1).

[0023] DPDK is a framework for performing the control of NIC (Network Interface Card) that was conventionally carried out by the Linux kernel (registered trademark) in the user space. The biggest difference from the processing in the Linux kernel is that it has a polling-based receiver mechanism called PMD (Pull Mode Driver). Usually, in the Linux kernel, when data arrives at the NIC, an interrupt occurs, and based on this, the reception processing is executed. On the other hand, PMD continuously performs data arrival confirmation and reception processing with a dedicated thread. By eliminating overheads such as context switches and interrupts, high-speed packet processing can be performed. DPDK significantly improves the performance and throughput of packet processing, making it possible to secure a lot of time for data plane application processing.

[0024] DPDK uses computer resources such as the CPU (Central Processing Unit) and NIC in an exclusive manner. Therefore, it is difficult to apply to uses that can be flexibly reconnected in module units like SFC. There is an application called SPP (Soft Patch Panel) to mitigate this. SPP prepares shared memory between VMs so that each VM can directly reference the same memory space, thereby omitting packet copying in the virtualization layer. Also, for the packet exchange between the physical NIC and the shared memory, high-speed operation is realized using DPDK. SPP can change the packet input destination and output destination software-wise by controlling the reference destination of the memory exchange of each VM. Through this processing, SPP realizes dynamic connection switching between VMs and between a VM and the physical NIC (see Non-Patent Document 2).

[0025] Figure 20 is a diagram for explaining packet transfer by the polling model in the configuration of OvS-DPDK (Open vSwitch with DPDK). The same components as in Figure 19 are labeled with the same reference numerals, and the description of overlapping parts is omitted. As shown in FIG. 20, the Host OS 20 includes OvS-DPDK 70 which is software for packet processing. OvS-DPDK 70 has a vhost-user 71 which is a functional unit for connecting to a virtual machine (here, VM1), and a dpdk (PMD) 72 which is a functional unit for connecting to a NIC (DPDK) 11 (physical NIC). Also, the data processing APL 1A includes a dpdk (PMD) 2 which is a functional unit for performing polling in the Guest OS 50 section. That is, the data processing APL 1A is an APL obtained by equipping the data processing APL 1 in FIG. 19 with the dpdk (PMD) 2 and modifying the data processing APL 1.

[0026] Packet transfer by the polling model enables path operation through a GUI in the SPP that performs high-speed packet copying between the Host OS 20 and the Guest OS 50 and between the Guest OS 50s without copying via shared memory as an extension of DPDK.

[0027] [Rx-side Packet Processing by New API (NAPI)] FIG. 21 is a schematic diagram of Rx-side packet processing by the New API (NAPI) implemented from Linux kernel 2.5 / 2.6 (see Non-Patent Document 1). The same components as those in FIG. 19 are denoted by the same reference numerals. As shown in FIG. 21, the New API (NAPI) executes the data processing APL 1 arranged in the user space 60 that can be used by the user on a server including the OS 70 (for example, the Host OS), and performs packet transfer between the NIC 11 of the HW 10 connected to the OS 70 and the data processing APL 1.

[0028] The OS 70 has a kernel 71, a Ring Buffer 72, and a Driver 73, and the kernel 71 has a protocol processing unit 74. Kernel71 is a core function of the OS70 (e.g., Host OS), which monitors hardware and manages the execution status of programs in terms of processes. Here, Kernel71 responds to requests from the data processing APL1 and conveys requests from the HW10 to the data processing APL1. Kernel71 processes requests from the data processing APL1 through system calls (where a "user program operating in non-privileged mode" requests processing from a "kernel operating in privileged mode"). Kernel71 transmits packets to the data processing APL1 via Socket75. Kernel71 receives packets from the data processing APL1 via Socket75.

[0029] Ring Buffer72 is managed by Kernel71 and is in the memory space of the server. Ring Buffer72 is a buffer of a fixed size that stores messages output by Kernel71 as logs, and when the upper limit size is exceeded, it is overwritten from the beginning.

[0030] Driver73 is a device driver for Kernel71 to monitor hardware. Note that Driver73 depends on Kernel71, and if the created (built) kernel source changes, it becomes a different thing. In this case, the corresponding driver source needs to be obtained, rebuilt on the OS using the driver, and the driver needs to be created.

[0031] The protocol processing unit 74 performs protocol processing for L2 (Data Link Layer) / L3 (Network Layer) / L4 (Transport Layer) defined by the OSI (Open Systems Interconnection) reference model.

[0032] Socket75 is an interface for kernel71 to perform inter-process communication. Socket75 has a socket buffer and does not frequently cause data copy processing. The flow until communication establishment via Socket75 is as follows. 1. The server side creates a socket file to accept clients. 2. Name the acceptance socket file. 3. Create a socket queue. 4. Accept the first connection from the client in the socket queue. 5. On the client side, create a socket file. 6. Send a connection request from the client side to the server. 7. On the server side, create a connection socket file separately from the acceptance socket file. As a result of communication establishment, data processing APL1 can call system calls such as read() and write() on kernel71.

[0033] In the above configuration, Kernel71 receives the notification of packet arrival from NIC11 by hardware interrupt (hardIRQ) and schedules software interrupt (softIRQ) for packet processing. The New API (NAPI) implemented from Linux kernel 2.5 / 2.6 performs packet processing by software interrupt (softIRQ) after hardware interrupt (hardIRQ) when a packet arrives. As shown in Figure 21, packet transfer by the interrupt model performs packet transfer by interrupt processing (see reference c in Figure 21), so waiting for interrupt processing occurs and the delay of packet transfer increases.

[0034] The following describes the outline of NAPI Rx-side packet processing. [Rx-side Packet Processing Configuration by New API (NAPI)] Figure 22 is a diagram for explaining the outline of Rx-side packet processing by New API (NAPI) at the location surrounded by the dashed line in Figure 21. <Device driver> As shown in FIG. 22, in the device driver, there are arranged a NIC 11 (physical NIC) which is a network interface card, a hard IRQ 81 which is a handler that executes a process (hardware interrupt) called and requested by the occurrence of a processing request of the NIC 11, and a netif_rx 82 which is a processing functional unit for software interrupts.

[0035] <Networking layer> In the Networking layer, there are arranged a soft IRQ 83 which is a handler that executes a process (software interrupt) called and requested by the occurrence of a processing request of the netif_rx 82, and a do_softirq 84 which is a control functional unit that performs the entity of the software interrupt (soft IRQ). Also, there are arranged a net_rx_action 85 which is a packet processing functional unit that executes upon receiving a software interrupt (soft IRQ), a poll_list 86 which registers information of a net device indicating which device the hardware interrupt from the NIC 11 is from, a netif_receive_skb 87 which creates a sk_buff structure (a structure for Kernel 71 to perceive the state of a packet), and a Ring Buffer 72.

[0036] <Protocol layer> In the Protocol layer, there are arranged an ip_rcv 88, an arp_rcv 89, etc., which are packet processing functional units.

[0037] The above netif_rx 82, do_softirq 84, net_rx_action 85, netif_receive_skb 87, ip_rcv 88, and arp_rcv 89 are parts (function names) of programs used for packet processing in Kernel 71.

[0038] [Rx-side Packet Processing Operation by New API (NAPI)] The arrows (symbols) d to o in FIG. 22 indicate the flow of Rx-side packet processing. When the hardware function unit 11a of NIC11 (hereinafter referred to as NIC11) receives a packet (or frame) in a frame from a counterpart device, it copies the arrived packet to Ring Buffer 72 without using the CPU by DMA (Direct Memory Access) transfer (see reference symbol d in FIG. 22). This Ring Buffer 72 is a memory space in the server and is managed by Kernel 71 (see FIG. 21).

[0039] However, if NIC11 only copies the arrived packet to Ring Buffer 72, Kernel 71 cannot recognize the packet. Therefore, when a packet arrives, NIC11 raises a hardware interrupt (hardIRQ) to hardIRQ81 (see reference symbol e in FIG. 22), and Kernel 71 recognizes the packet by netif_rx82 executing the following processing. Note that hardIRQ81 shown surrounded by an ellipse in FIG. 22 represents a handler, not a functional unit.

[0040] netif_rx82 is a function that actually performs processing. When hardIRQ81 (the handler) is raised (see reference symbol f in FIG. 22), it stores in poll_list86 the information of the net_device that indicates which device the hardware interrupt from NIC11, which is one of the information of the content of the hardware interrupt (hardIRQ), belongs to, and registers queue pruning (deleting the corresponding queue entry from the buffer in consideration of the next process to be performed by referring to the content of the packet accumulated in the buffer) (see reference symbol g in FIG. 22). Specifically, upon receiving that the packet has been stuffed into Ring Buffer 72, netif_rx82 uses the driver of NIC11 to register subsequent queue pruning in poll_list86 (see reference symbol g in FIG. 22). As a result, queue pruning information due to the packet being stuffed into Ring Buffer 72 is registered in poll_list86.

[0041] Thus, in the <Device driver> of Figure 22, when NIC11 receives a packet, it copies the packet that has arrived at Ring Buffer72 through DMA transfer. Also, NIC11 raises hardIRQ81 (handler), and netif_rx82 registers the net_device in poll_list86 and schedules a software interrupt (softIRQ). So far, the processing of the hardware interrupt in the <Device driver> of Figure 22 stops.

[0042] After that, netif_rx82 raises it to softIRQ83 (handler) with a software interrupt (softIRQ) to harvest the data stored in Ring Buffer72 using the information (specifically, pointers) in the queue stored in poll_list86 (see reference h in Figure 22), and notifies do_softirq84, which is the control function part of the software interrupt (see reference i in Figure 22).

[0043] do_softirq84 is the software interrupt control function part that defines each function of the software interrupt (there are various packet processes, and interrupt processing is one of them. Define the interrupt processing). Based on this definition, do_softirq84 notifies net_rx_action85, which actually performs the software interrupt processing, of the request for the current (corresponding) software interrupt (see reference j in Figure 22).

[0044] When the turn of the softIRQ comes, net_rx_action85 calls a polling routine to harvest packets from Ring Buffer72 based on the net_device registered in poll_list86 (see reference k in Figure 22) and harvests the packets (see reference l in Figure 22). At this time, net_rx_action85 continues harvesting until poll_list86 becomes empty. After that, net_rx_action85 notifies netif_receive_skb87 (see reference m in Figure 22).

[0045] netif_receive_skb87 creates an sk_buff structure, analyzes the content of the packet, and transfers the processing to the subsequent protocol processing unit 74 (see Figure 21) for each type. That is, when netif_receive_skb87 analyzes the content of the packet and processes it according to the content of the packet, it transfers the processing to ip_rcv88 of <Protocol layer> (symbol n in Figure 22), and also, for example, if it is L2, it transfers the processing to arp_rcv89 (symbol o in Figure 22).

[0046] Non-Patent Document 3 describes a server internal network delay control device (KBP: Kernel Busy Poll). KBP constantly monitors packet arrival by the polling model within the kernel. Thereby, it suppresses softIRQ and realizes low-latency packet processing.

[0047] Figure 23 is an example of data transfer of video (30 FPS). The workload shown in Figure 23 performs intermittent data transfer at a transfer rate of 350 Mbps every 30 ms.

[0048] Figure 24 is a diagram showing the CPU usage rate used by the busy poll thread in KBP described in Non-Patent Document 3. As shown in Figure 24, in KBP, the kernel thread exclusively occupies the CPU core to perform busy poll. Even for the intermittent packet reception shown in Figure 23, in KBP, since it always uses the CPU regardless of the presence or absence of packet arrival, there is a problem that power consumption increases.

[0049] Next, the DPDK system will be described. [DPDK System Configuration] Figure 25 is a diagram showing the configuration of a DPDK system that controls HW110 equipped with an accelerator 120. The DPDK system has DPDK 150, a data high-speed transfer middleware placed on HW 110, OS 140, and user space 160, and a data processing APL 1. The data processing APL 1 is packet processing performed prior to the execution of the APL. HW 110 communicates data transmission and reception with the data processing APL 1. In the following description, as shown in FIG. 25, the data flow in which the data processing APL 1 receives a packet from HW 110 is referred to as Rx-side reception, and the data flow in which the data processing APL 1 transmits a packet to HW 110 is referred to as Tx-side transmission.

[0050] HW 110 includes an accelerator 120 and a NIC 130 (physical NIC) for connecting to a communication network. The accelerator 120 is a computational unit hardware that performs specific operations at high speed based on an input from the CPU. Specifically, the accelerator 120 is a PLD (Programmable Logic Device) such as a GPU (Graphics Processing Unit) or an FPGA (Field Programmable Gate Array). In FIG. 25, the accelerator 120 includes a plurality of cores (core processors) 121, an Rx queue (queue: waiting queue) 122 for holding data in a first-in first-out list structure, and a Tx queue 133.

[0051] Offload a part of the processing of the data processing APL 1 to the accelerator 120 to achieve performance and power efficiency that cannot be achieved by software (CPU processing) alone. In a large-scale server cluster such as a data center that constitutes NFV (Network Functions Virtualization) or SDN (Software Defined Network), a case where the above-described accelerator 120 is applied is assumed.

[0052] NIC130 is NIC hardware that implements a NW interface, and includes an Rx queue 131 and a Tx queue 132 that hold data in a first-in, first-out list structure. NIC130 is connected to a counterpart device 170 via, for example, a communication network, and performs packet transmission and reception. Note that NIC130 may be, for example, a SmartNIC which is a NIC with an accelerator. A SmartNIC is a NIC that can offload processing-intensive tasks such as IP packet processing that cause a drop in processing power, thereby reducing the load on the CPU.

[0053] DPDK150 is a framework for controlling a NIC in user space 160, and specifically consists of data high-speed transfer middleware. DPDK150 has a PMD (Poll Mode Driver) 151 which is a polling-based reception mechanism (a driver that can select data arrival in polling mode or interrupt mode). PMD151 continuously performs data arrival confirmation and reception processing with a dedicated thread.

[0054] DPDK150 realizes a packet processing function in user space 160 where APL operates, and enables reduction of packet transfer delay by immediately harvesting packets in polling model when packets arrive from user space 160. That is, since DPDK150 performs packet harvesting by polling (busy polling the queue with the CPU), there is no waiting and the delay is small.

Prior Art Documents

Patent Documents

[0055]

Patent Document 1

Patent Document 2

Non-Patent Documents

[0056]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0057] However, both packet transfer using the interrupt model and the polling model have the following problems. In the interrupt model, the kernel that receives an event (hardware interrupt) from the HW performs packet transfer by means of software interrupt processing for packet processing. Therefore, since the interrupt model performs packet transfer by interrupt (software interrupt) processing, there are problems such as competition with other interrupts and waiting if the interrupt destination CPU is being used by a higher-priority process, resulting in a large delay in packet transfer. In this case, if the interrupt processing becomes congested, the waiting delay becomes even larger. For example, as shown in FIG. 19, packet transfer by the interrupt model performs packet transfer by interrupt processing (see reference numerals a and b in FIG. 19), so waiting for interrupt processing occurs and the delay in packet transfer becomes large.

[0058] Supplement the mechanism that causes delays in the interrupt model. In a general kernel, packet transfer processing is transmitted by software interrupt processing after hardware interrupt processing. When a software interrupt for packet transfer processing occurs, under the following conditions (1) to (3), the software interrupt processing cannot be executed immediately. For this reason, it is mediated by a scheduler such as ksoftirqd (a kernel thread for each CPU, which is executed when the load of software interrupts becomes high), and the interrupt processing is scheduled, resulting in a waiting time on the order of milliseconds. (1) When competing with other hardware interrupt processing (2) When competing with other software interrupt processing (3) When other processes or kernel threads (such as migration threads) with high priority and the interrupt destination CPU are in use Under the above conditions, the software interrupt processing cannot be executed immediately.

[0059] Similarly, for packet processing by New API (NAPI), as shown in the dashed box p in Figure 22, due to the competition of interrupt processing (softIRQ), a NW delay on the order of milliseconds occurs.

[0060] <Problems of KBP> As described above, KBP can suppress softIRQ and achieve low-latency packet processing by constantly monitoring packet arrival using the polling model within the kernel. However, since the kernel thread that constantly monitors packet arrival exclusively occupies the CPU core and always uses CPU time, there is a problem of high power consumption. With reference to Figures 23 and 24, the relationship between the workload and CPU usage will be described. As shown in FIG. 24, in KBP, the kernel thread exclusively occupies the CPU core to perform busy polling. Even for the intermittent packet reception shown in FIG. 23, in KBP, since the CPU is always used regardless of the arrival of packets, there is a problem of increased power consumption.

[0061] DPDK also has the same problem as the above KBP. <Problems of DPDK> In DPDK, the kernel thread exclusively occupies the CPU core to perform polling (busy polling the queue with the CPU). Therefore, even for the intermittent packet reception shown in FIG. 23, in DPDK, regardless of the arrival of packets, the CPU is always used at 100%, so there is a problem of increased power consumption.

[0062] In this way, since DPDK realizes the polling model in user space and no softIRQ conflict occurs, and KBP realizes the polling model in the kernel and no softIRQ conflict occurs, low-latency packet transfer is possible. However, both DPDK and KBP have the problem of wasting CPU resources for packet arrival monitoring at all times regardless of the arrival of packets, resulting in increased power consumption.

[0063] In view of such a background, the present invention has been made, and an object of the present invention is to reduce the CPU usage rate and enable power saving while maintaining low latency.

Means for Solving the Problems

[0064] To solve the above problems, a data transfer device within a server is provided with a sleep control management unit that manages a data arrival schedule and performs sleep control in accordance with the data arrival timing. The sleep control management unit causes a thread that monitors the arrival of data using a polling model based on the data arrival timing to sleep in a data flow having information regarding a specific data arrival timing, and performs an operation to wake up the thread based on the data arrival timing. The data transfer device within a server is characterized by this.

Advantages of the Invention

[0065] According to the present invention, it is possible to reduce the CPU usage rate while maintaining low latency, thereby achieving power saving.

Brief Description of the Drawings

[0066]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Embodiments for Carrying Out the Invention

[0067] Hereinafter, a server internal data transfer system and the like in an embodiment for carrying out the present invention (hereinafter referred to as "the present embodiment") will be described with reference to the drawings. (First Embodiment) [Overall Configuration] FIG. 1 is a schematic configuration diagram of a server internal data transfer system according to the first embodiment of the present invention. The same components as those in FIG. 25 are denoted by the same reference numerals. As shown in FIG. 1, the server internal data transfer system 1000 includes an HW 110, an OS 140, and a server internal data transfer device 200 which is a data high-speed transfer middleware arranged on a user space 160. In user space 160, there are further arranged a data processing APL 1 and a data flow time slot management scheduler 2. The data processing APL 1 is a program executed in user space 160. The data flow time slot management scheduler 2 transmits schedule information to the data processing APL 1 (see reference symbol q in FIG. 1). Further, the data flow time slot management scheduler 2 transmits data arrival schedule information to a sleep control management unit 210 (described later) (see reference symbol r in FIG. 1).

[0068] HW 110 performs data transmission and reception communication with the data processing APL 1. The data flow in which the data processing APL 1 receives a packet from HW 110 is referred to as Rx-side reception, and the data flow in which the data processing APL 1 transmits a packet to HW 110 is referred to as Tx-side transmission. HW 110 includes an accelerator 120 and a NIC 130 (physical NIC) for connecting to a communication network.

[0069] The accelerator 120 is computing unit hardware such as a GPU or an FPGA. The accelerator 120 includes a plurality of cores (core processors) 121, an Rx queue 122 for holding data in a first-in first-out list structure, and a Tx queue 123. A part of the processing of the data processing APL 1 is offloaded to the accelerator 120 to achieve performance and power efficiency that cannot be achieved by software (CPU processing) alone.

[0070] The NIC 130 is NIC hardware that realizes a NW interface, and includes an Rx queue 131 and a Tx queue 132 for holding data in a first-in first-out list structure. The NIC 130 is connected to a counterpart device 170 via, for example, a communication network and performs packet transmission and reception.

[0071] The OS 140 is, for example, Linux (registered trademark). The OS 140 includes a high-resolution timer 141 that performs timer management in more detail than the kernel timer. The high-resolution timer 141 uses, for example, the hrtimer of Linux (registered trademark). In hrtimer, the time at which a callback occurs can be specified using the unit of ktime_t. The high-resolution timer 141 notifies the sleep control unit 221 of the data transfer unit 220 described later of the data arrival timing at the specified time (see reference u in FIG. 1).

[0072] [Server internal data transfer device 200] The server internal data transfer device 200 is DPDK for controlling the NIC in the user space 160, and specifically consists of data high-speed transfer middleware. The server internal data transfer device 200 includes a sleep control management unit 210 and a data transfer unit 220. The server internal data transfer device 200 has a PMD 151 (driver capable of selecting data arrival in polling mode or interrupt mode) (see FIG. 25) similar to the DPDK arranged on the user space 160. The PMD 151 is a driver capable of selecting data arrival in polling mode or interrupt mode, and a dedicated thread continuously performs confirmation of data arrival and reception processing.

[0073] <sleep control management unit 210> The sleep control management unit 210 manages the data arrival schedule and performs sleep control of the data transfer unit 220 in accordance with the data arrival timing. The sleep control management unit 210 collectively controls the Sleep / start timing of each data transfer unit 220 (see reference t in FIG. 1).

[0074] The sleep control management unit 210 manages the data arrival schedule information, distributes the data arrival schedule information to the data transfer unit 220, and performs sleep control of the data transfer unit 220. The sleep control management unit 210 includes a data transfer unit management unit 211, a data arrival schedule management unit 212, and a data arrival schedule distribution unit 213.

[0075] The data transfer unit management unit 211 holds information such as the number of data transfer units 220 and process IDs (PIDs: Process Identification) in a list. The data transfer unit management unit 211 transmits information such as the number of data transfer units 220 and process IDs to the data transfer unit 220 in response to a request from the data arrival schedule distribution unit 213.

[0076] The data arrival schedule management unit 212 manages the data arrival schedule. The data arrival schedule management unit 212 acquires data arrival schedule information from the data flow time slot management scheduler 2 (see reference symbol r in FIG. 1). When there is a change in the data arrival schedule information, the data arrival schedule management unit 212 receives a change notification of the data arrival schedule information from the data flow time slot management scheduler 2 and detects the change in the data arrival schedule information. Alternatively, the data arrival schedule management unit 212 detects it by snooping the data including the data arrival schedule information (see FIGS. 4 and 5). The data arrival schedule management unit 212 transmits the data arrival schedule information to the data arrival schedule distribution unit 213 (see reference symbol s in FIG. 1).

[0077] The data arrival schedule distribution unit 213 acquires information such as the number of data transfer units 220 and process IDs from the data transfer unit management unit 211. The data arrival schedule distribution unit 213 distributes the data arrival schedule information to each data transfer unit 220 (see reference symbol t in FIG. 1).

[0078] <Data transfer unit 220> The data transfer unit 220 starts a polling thread that monitors packet arrival using the polling model. Based on the data arrival schedule information distributed from the sleep control management unit 210, the data transfer unit 220 puts the thread to sleep and activates a timer immediately before data arrival to wake up the thread for sleep release. Here, in case the data transfer unit 220 receives a packet at an unintended timing by the timer, when releasing the sleep, the corresponding thread is woken up by a hardware interrupt. Sleep / release will be described later in [Sleep / Release].

[0079] The data transfer unit 220 includes a sleep control unit 221, a data arrival monitoring unit 222, an Rx data transfer unit 223 (packet harvesting unit), and a Tx data transfer unit 224. The data arrival monitoring unit 222 and the Rx data transfer unit 223 are functional units on the Rx side, and the Tx data transfer unit 224 is a functional unit on the Tx side.

[0080] <sleep control unit 221> Based on the data arrival schedule information from the sleep control management unit 210, the sleep control unit 221 performs sleep control to stop data arrival monitoring and go to sleep when there is no data arrival. The sleep control unit 221 holds the data arrival schedule information received from the data arrival schedule distribution unit 213.

[0081] The sleep control unit 221 sets a timer for the data arrival timing for the data arrival monitoring unit 222 (see reference v in Figure 1). That is, the sleep control unit 221 sets a timer so that the data arrival monitoring unit 222 can start polling immediately before data arrival. Here, the sleep control unit 221 may use hrtimers or the like, which are high-resolution timers 141 held by the Linux kernel, and activate the data arrival monitoring unit 222 by a hardware interrupt trigger when the timer is activated by the hardware clock.

[0082] Figure 2 is a diagram showing an operation example of the polling thread of the in-server data transfer device 200. The vertical axis represents the CPU usage rate [%] of the CPU core used by the polling thread, and the horizontal axis represents time. Note that Figure 3 shows an operation example of the polling thread due to packet arrival corresponding to the data transfer example of the video (30 FPS) in which packets are received intermittently shown in Figure 13. As shown in Figure 2, the data transfer unit 220 puts the thread (polling thread) to sleep (see reference w in Figure 2) based on the data arrival schedule information received from the sleep control management unit 210, and when the sleep is released, the sleep is released by a hardware interrupt (hardIRQ) (see reference x in Figure 2). Note that reference y in Figure 2 is a fluctuation in the wiring voltage due to congestion use of the core CPU (Core processor), etc.

[0083] <Rx side> The data arrival monitoring unit 222 is activated immediately before data arrives according to the data arrival schedule information managed by the sleep control unit 221. The data arrival monitoring unit 222 monitors the Rx queues 122 and 131 of the accelerator 120 or the NIC 130 to check for the arrival of data.

[0084] Regardless of whether data arrives or not, the data arrival monitoring unit 222 monopolizes the CPU core and monitors the arrival of data by polling. Incidentally, if this is made an interrupt model, the delay described in the prior art of Figure 22 (that is, when a softIRQ competes with other softIRQs, a waiting occurs regarding the execution of the softIRQ, and a NW delay on the order of ms due to this waiting) occurs. In this embodiment, it is a feature that the polling model sleep control is used on the Rx side.

[0085] When there is data arrival in the Rx queues 122 and 131, the Data Arrival Monitoring Unit 222 deletes the entries of the corresponding queues from the Rx queues 122 and 131 (refer to the content of the packets stored in the buffer, and delete the corresponding queue entries from the buffer considering the next processing to be performed for the processing of the packet), and transfers them to the Rx Data Transfer Unit 223.

[0086] The Rx Data Transfer Unit 223 transfers the received data to the Data Processing APL1. Similar to the Tx Data Transfer Unit 224, since it operates only when data arrives, it does not waste the CPU.

[0087] <Tx side> The Tx Data Transfer Unit 224 stores the received data in the Tx queues 123 and 132 of the accelerator 120 or the NIC 130. The Tx Data Transfer Unit 224 is activated by inter - process communication when the Data Processing APL1 sends data, and returns to CPU idle when the data transfer is completed. Therefore, unlike the Data Arrival Monitoring Unit 222, it does not waste the CPU.

[0088] [Sleep / Release] Based on the data arrival schedule information received from the sleep control unit 221, the Data Transfer Unit 220 puts the thread to sleep and wakes it up when triggered by a timer. <Normal time> Based on the scheduling information of the data arrival timing (data arrival schedule information), the Data Transfer Unit 220 activates a timer immediately before the data arrival to wake up the data arrival monitoring unit thread of the Data Transfer Unit 220. For example, using the hr_timer function standardly installed in the Linux kernel, when the timer expiration occurs, it activates the hardware interrupt of the timer, and the Data Arrival Monitoring Unit 222 wakes up the thread.

[0089] <Unexpected (when data arrives outside the schedule)> When data arrives outside the scheduled timing, the thread of the data arrival monitoring unit 222 is in a sleeping state. Also, there is no timer activation for normal operation. Therefore, a hardware interrupt that notifies packet arrival is triggered when a packet arrives. As described above, during normal operation, since packets are constantly monitored in polling mode, a hardware interrupt is not necessary, and the function of the hardware interrupt is disabled in the driver (PMD). However, when putting the polling thread to sleep, assume that data may arrive outside the scheduling, and change the mode so that a hardware interrupt is raised when a packet arrives. By doing so, when a packet arrives, the hardware interrupt is raised, and in this hardware interrupt handler, the data arrival monitoring unit 222 can wake up the thread.

[0090] [Example of obtaining data arrival schedule information] An example of obtaining data arrival schedule information in the server internal data transfer system according to this embodiment will be described. As an example of a data flow where the data arrival schedule is determined, signal processing in the RAN (Radio Access Network) can be cited. The signal processing in the RAN manages the data arrival timing by time division multiplexing by the MAC scheduler of MAC4 (described later).

[0091] For signal processing of vRAN (virtual RAN) and vDU (virtual Distributed Unit), DPDK is often used for high-speed data transfer. By applying the inventive method, sleep control of the data transfer unit (such as DPDK PMD) is performed according to the data arrival timing managed by the MAC scheduler.

[0092] As a method for obtaining the data arrival timing managed by the MAC scheduler, there are <obtaining data arrival schedule information from the MAC scheduler> (directly obtained from the MAC Scheduler) (see Figure 3), <obtaining data arrival schedule information by snooping FAPI P7> (obtained by snooping the FAPI P7 IF) (see Figure 4), and <obtaining data arrival schedule information by snooping CTI> (obtained by snooping the O-RAN CTI) (see Figure 5). These will be described in order below.

[0093] <obtaining data arrival schedule information from the MAC scheduler> Figure 3 is a schematic configuration diagram of the in-server data transfer system of Acquisition Example 1. Acquisition Example 1 is an example applied to the vDU system. The same components as those in Figure 1 are labeled with the same reference numerals, and the description of overlapping parts is omitted. As shown in Figure 3, in the in-server data transfer system 1000A of Acquisition Example 1, in the user space 160, further, a PHY (High) (Physical) 3, a MAC (Medium Access Control) 4, and an RLC (Radio Link Control) 5 are arranged. As a countermeasure device connected to the NIC 130, an RU (Radio Unit) 171 is connected to the receiving side of the NIC 130, and a vCU 172 is connected to the transmitting side of the NIC 130.

[0094] The sleep control management unit 210 of the in-server data transfer system 1000A modifies the MAC scheduler of the MAC 4 to obtain data arrival schedule information from the MAC 4 (see the reference numeral z in Figure 3). Although an example applied to the vDU system has been described, it may be applied not only to the vDU but also to vRAN systems such as the vCU.

[0095] <obtaining data arrival schedule information by snooping CTI> Figure 4 is a schematic configuration diagram of the in-server data transfer system of Acquisition Example 2. Acquisition Example 2 is an example applied to the vCU system. The same components as those in Figure 3 are denoted by the same reference numerals, and the description of overlapping parts is omitted. As shown in Figure 4, in the in-server data transfer system 1000B of Acquisition Example 2, in the user space 160, further, an FAPI (FAPI P7) 6 is arranged between the PHY (High) 3 and the MAC 4. Note that although the FAPI 6 is drawn inside the in-server data transfer device 200 for the sake of notation, the FAPI 6 is arranged outside the in-server data transfer device 200. The FAPI 6 is an IF (interface) for exchanging data scheduling information and the like for connecting the PHY (High) 3 and the MAC 4 defined in the SCF (Small Cell Forum) (see reference numeral aa in Figure 4).

[0096] The sleep control management unit 210 of the in-server data transfer system 1000B snoops on the FAPI 6 and then acquires data arrival schedule information (see reference numeral bb in Figure 4).

[0097] <Snoop on CTI7 and acquire data arrival schedule information> Figure 5 is a schematic configuration diagram of the in-server data transfer system of Acquisition Example 3. Acquisition Example 3 is an example applied to the vCU system. The same components as those in Figure 3 are denoted by the same reference numerals, and the description of overlapping parts is omitted. As shown in Figure 5, in the in-server data transfer system 1000C of Acquisition Example 3, a transmission device 173 is arranged outside the user space 160. The transmission device 173 is a transmission device defined in the O-RAN community. The MAC 4 in the user space 160 and the transmission device 173 are connected via a CTI (Collaborative Transport Interface) 7. The CTI 7 is an IF for exchanging data scheduling information and the like with the transmission device defined in the O-RAN community (see reference numeral cc in Figure 5).

[0098] The sleep control management unit 210 of the in-server data transfer system 1000C snoops on the CTI7 and then acquires data arrival schedule information (refer to the reference symbol dd in Fig. 5).

[0099] The operation of the in-server data transfer system will be described below. Since the basic operations of the in-server data transfer systems 1000 (refer to Fig. 1), 1000A (refer to Fig. 3), 1000B (refer to Fig. 4), and 1000C (refer to Fig. 5) are the same, the in-server data transfer system 1000 (refer to Fig. 1) will be described.

[0100] [Operation of the sleep control management unit 210] <When there is a change in the data arrival schedule information> Fig. 6 is a flowchart showing the operation of the sleep control management unit 210 when there is a change in the data arrival schedule information. Step S10 shown by the dashed line in Fig. 6 represents an external factor for the start of the operation of the sleep control management unit 210 (hereinafter, in this specification, the dashed line in the flowchart represents an external factor for the start of the operation). In step S10 [external factor], when there is a change in the data arrival schedule information, the data flow time slot management scheduler 2 (refer to Fig. 1) notifies the data arrival schedule management unit 212 of the sleep control management unit 210 that there has been a change (refer to the reference symbol r in Fig. 1). Or, as shown in Figs. 4 and 5, the data arrival schedule management unit 212 (refer to Fig. 1) of the sleep control management unit 210 detects it by snooping on the data containing the data arrival schedule information.

[0101] In step S11, the data arrival schedule management unit 212 (refer to Fig. 1) of the sleep control management unit 210 acquires the data arrival schedule information from the data flow time slot management scheduler 2 (refer to Fig. 1).

[0102] In step S12, the data arrival schedule management unit 212 transmits data arrival schedule information to the data arrival schedule distribution unit 213 (see FIG. 1).

[0103] In step S13, the data arrival schedule distribution unit 213 of the sleep control management unit 210 acquires information such as the number and process ID of the data transfer unit 220 (see FIG. 1) from the data transfer unit management unit 211 (see FIG. 1).

[0104] In step S14, the data arrival schedule distribution unit 213 distributes the data arrival schedule information to each data transfer unit 220 (see FIG. 1) to complete the processing of this flow.

[0105] <When the addition or deletion of the data transfer unit 220 occurs> FIG. 7 is a flowchart showing the operation of the sleep control management unit 210 when the addition or deletion of the data transfer unit 220 occurs. In step S20 [external factor], when the addition or deletion of the data transfer unit 220 (see FIG. 1) occurs, the operation system of this system, maintenance operator, etc. set information such as the number and process ID of the data transfer unit 220 for the data transfer unit management unit 211 (see FIG. 1) of the sleep control management unit 210.

[0106] In step S21, the data transfer unit management unit 211 of the sleep control management unit 210 holds information such as the number and process ID of the data transfer unit 220 as a list.

[0107] In step S22, the data transfer unit management unit 211 transmits information such as the number and process ID of the data transfer unit 220 in response to a request from the data arrival schedule distribution unit 213 to complete the processing of this flow. The operation of the sleep control management unit 210 has been described above. Next, the operation of the data transfer unit 220 will be described.

[0108] [Operation of the data transfer unit 220] <Sleep control> FIG. 8 is a flowchart showing the operation of the sleep control unit 221 of the data transfer unit 220. In step S31, the sleep control unit 221 (see FIG. 1) of the data transfer unit 220 holds the data arrival schedule information received from the data arrival schedule distribution unit 213 (see FIG. 1) of the sleep control management unit 210.

[0109] Here, due to reasons such as not being time-synchronized with the opposing device 170 (see FIG. 1), there may be a constant difference between the data arrival timing managed by the sleep control management unit 210 (see FIG. 1) and the actual data arrival timing. In this case, the data transfer unit 220 stores the difference from the data arrival timing, and if this difference data is constant, the sleep control management unit 210 may correct it by a certain difference time to handle it (details will be described later with reference to FIGS. 11 and 12).

[0110] In step S32, the sleep control unit 221 (see FIG. 1) of the data transfer unit 220 sets a timer for the data arrival timing for the data arrival monitoring unit 222 (see FIG. 1). That is, the sleep control unit 221 sets a timer so that the data arrival monitoring unit 222 can start polling immediately before data arrival.

[0111] At this time, a high-resolution timer 141 (see FIG. 1) such as hrtimers (registered trademark) held by the Linux kernel (registered trademark) may be used to activate the data arrival monitoring unit 222 by means of a hardware interrupt opportunity when the timer is activated by the hardware clock. The operation of the sleep control unit 221 has been described above. Next, the operations of the <Rx side> and <Tx side> of the data transfer unit 220 will be described. The present invention is characterized in that the operations are different between the <Rx side> and the <Tx side>.

[0112] <Rx side> FIG. 9 is a flowchart showing the operation of the data arrival monitoring unit 222 of the data transfer unit 220. In step S41, the data arrival monitoring unit 222 (see FIG. 1) of the data transfer unit 220 is activated immediately before data arrives in accordance with the data arrival schedule information managed by the sleep control unit 221 (see FIG. 1).

[0113] Here, when data is received from the accelerator 120 or the NIC 130 (see FIG. 1) while the data arrival monitoring unit 222 is sleeping, a hardware interrupt is activated at the time of data reception, and the data arrival monitoring unit 222 may be activated within this hardware interrupt handler. This method is effective for handling the case where data arrives at a timing that deviates from the data arrival schedule managed by the sleep control management unit 210.

[0114] In step S42, the data arrival monitoring unit 222 monitors the Rx queues 122, 131 (see FIG. 1) of the accelerator 120 or the NIC 130 to confirm the presence or absence of data arrival. At this time, regardless of the presence or absence of data arrival, the CPU core is exclusively occupied to monitor the presence or absence of data arrival by polling. If this is made an interrupt model, the delay described in the prior art of FIG. 22 (that is, when a softIRQ competes with other softIRQs, a wait occurs regarding the execution of the softIRQ, and an NW delay on the order of ms due to this wait) occurs. The feature of this embodiment is that a polling model sleep control is used on the Rx side.

[0115] In step S43, the data arrival monitoring unit 222 determines whether or not data has arrived in the Rx queues 122, 131.

[0116] If data has arrived in the Rx queues 122, 131 (S43: Yes), in step S44, the data arrival monitoring unit 222 culls the data (queue) stored in the Rx queues 122, 131 (refer to the content of the packets accumulated in the buffer, and delete the corresponding queue entry from the buffer in consideration of the next process to be performed for the processing of that packet), and transfers it to the Rx data transfer unit 223 (see FIG. 1). If there is no data arrival at the Rx queues 122 and 131 (S43: No), the process returns to step S42.

[0117] In step S45, the Rx data transfer unit 223 transfers the received data to the data processing APL1 (see FIG. 1). Similar to the Tx data transfer unit 224 (see FIG. 1) described later, the Rx data transfer unit 223 operates only when data arrives, so it does not waste the CPU.

[0118] In step S46, when no data arrives even after the elapse of a certain period specified by the operator, the sleep control management unit 210 (see FIG. 1) puts the data arrival monitoring unit 222 (see FIG. 1) to sleep and ends the processing of this flow.

[0119] <Tx side> FIG. 10 is a flowchart showing the operation of the Tx data transfer unit 224 of the data transfer unit 220. In step S50 [external factor], the data processing APL1 (see FIG. 1) transfers data to the data transfer unit 220 of the in-server data transfer device 200 (see FIG. 1).

[0120] In step S51, the Tx data transfer unit 224 of the data transfer unit 220 stores the received data in the Tx queues 123 and 132 (see FIG. 1) of the accelerator 120 or the NIC 130 (see FIG. 1) and ends the processing of this flow. The Tx data transfer unit 224 is activated by inter-process communication when the data processing APL1 sends data, and returns to CPU idle when the data transfer is completed. Therefore, unlike the data arrival monitoring unit 222 on the <Rx side>, it does not waste the CPU. The operation of the data transfer unit 220 has been described above.

[0121] [Example of handling when there is a difference in the data arrival schedule] Next, the handling when there is a certain time difference between the data arrival schedule grasped by the sleep control management unit 210 and the actually arriving data arrival schedule will be described. This is a supplementary explanation of step S31 in FIG. 8. In the present embodiment, a use case where the data arrival schedule of the RAN or the like is predetermined is assumed. Data arrival with a non-constant time difference is not allowed by the RAN system (APL side) and is thus excluded from consideration.

[0122] <When the schedule of the data transfer unit 220 is ahead of the actual data arrival: Case1> FIG. 11 is a flowchart showing the operation of the data transfer unit 220 when there is a difference in the data arrival schedule. In step S61, the data arrival monitoring unit 222 (see FIG. 1) of the data transfer unit 220 monitors the Rx queues 122 and 131 (see FIG. 1) of the accelerator 120 or the NIC 130, and records the time difference Δ (the symbol representing the difference is denoted as Δ) T from the data arrival schedule to the actual data arrival in a memory (not shown).

[0123] In step S62, when there are consecutive data arrival differences of ΔT a plurality of times, the data arrival monitoring unit 222 (see FIG. 1) transmits to the sleep control unit 221 (see FIG. 1) that the data arrival schedule has advanced by ΔT. The consecutive plurality of times here are arbitrarily set by the operator of this system.

[0124] In step S63, the sleep control unit 221 (see FIG. 1) of the data transfer unit 220 receives the transmission that the data arrival schedule has advanced by ΔT, delays the data arrival schedule by ΔT, and ends the processing of this flow. Thereby, it becomes possible to correct the schedule when the data arrival schedule is early for a certain period of time.

[0125] <When the schedule of the data transfer unit 220 is behind the actual data arrival: Case2> FIG. 12 is a flowchart showing the operation of the data transfer unit 220 when there are differences in the data arrival schedule. In step S71, the data arrival monitoring unit 222 (see FIG. 1) of the data transfer unit 220 monitors the Rx queues 122 and 131 (see FIG. 1) of the accelerator 120 or the NIC 130, and if data has already arrived in the first polling when the data arrival monitoring is started, it records this in a memory (not shown). To supplement the explanation, the data arrival monitoring unit 222 is activated immediately before the data arrives (see the process of step S32 in FIG. 8). However, even though it is "immediately before", there is a time interval of Δt for "immediately before", and it is assumed that there will be some cycles of idle polling. Therefore, if data has already arrived when the polling is started, it can be determined that there is a high possibility that the schedule of the data transfer unit 220 is delayed.

[0126] In step S72, when the data arrival monitoring unit 222 has data arrival at the start of polling continuously for a plurality of times, it transmits to the sleep control unit 221 (see FIG. 1) to delay the data arrival schedule by a micro time ΔS. Here, since it is not possible to grasp exactly how much the data arrival schedule is shifted, by repeatedly delaying the micro time ΔS set arbitrarily by the operator, the schedule is adjusted little by little.

[0127] In step S73, the sleep control unit 221 receives the transmission indicating that the data arrival schedule should be advanced by ΔS, advances the data arrival schedule by ΔS, and ends the process of this flow. By repeatedly performing this time correction of ΔS, it becomes possible to correct the schedule when there is a delay in the data arrival schedule for a certain period of time.

[0128] As described above, in the in-server data transfer system 1000, the in-server data transfer device 200 is arranged on the user space 160. Therefore, like DPDK, the data transfer unit 220 of the in-server data transfer device 200 can bypass the kernel and refer to a ring-structured buffer (a ring-structured buffer created in the memory space managed by DPDK when a packet arrives at the accelerator 120 or the NIC 130 through DMA (Direct Memory Access)). That is, the in-server data transfer device 200 does not use the ring buffer (Ring Buffer72) (see FIG. 22) or the pole list (Ring Buffer72) (see FIG. 22) in the kernel. The data transfer unit 220 can instantly detect the arrival of a packet by constantly monitoring the ring-structured buffer (mbuf; a ring-structured buffer to which the PMD151 copies data by DMA) created in the memory space managed by this DPDK (that is, it is a polling model rather than an interrupt model).

[0129] In addition to the feature of being arranged on the user space 160 described above, the in-server data transfer device 200 has the following features regarding the wake-up method of the polling thread. That is, for a workload with a determined data arrival timing, the in-server data transfer device 200 wakes up the polling thread by a timer based on the scheduling information (data arrival schedule information) of the data arrival timing. Note that the in-server data transfer device 200B (see FIG. 17) of the third embodiment described later provides a polling thread in the kernel and wakes up the polling thread by a hardware interrupt opportunity from the NIC 11.

[0130] A supplementary explanation of the operation of the in-server data transfer device 200 will be given. <Normal operation: Polling mode> The in-server data transfer device 200 has the polling thread in user space 160 monitor the ring buffer expanded from the accelerator 120 or the NIC 130 (see FIG. 1) in the memory space. Specifically, the PMD 151 (see FIG. 25) of the in-server data transfer device 200 is a driver that can select data arrival in polling mode or interrupt mode. When data arrives at the accelerator 120 or the NIC 130, since the buffer of the ring structure called mbuf is in the memory space, the PMD 151 copies the data to the ring structure buffer mbuf by DMA. The polling thread in user space 160 monitors this ring structure buffer mbuf. Therefore, the in-server data transfer device 200 does not use the poll_list prepared by the kernel. The normal operation (polling mode) has been described above. Next, the operation in the unexpected interrupt mode will be described.

[0131] <Unexpected operation: Interrupt mode> When data arrives while the polling thread of the in-server data transfer device 200 is sleeping, the mode of the driver (PMD 151) is changed so that a hardware interrupt (hardIRQ) can be raised from the accelerator 120 or the NIC 130 (see FIG. 1). When data arrives at the accelerator 120 or the NIC 130, a hardware interrupt is triggered to wake up the polling thread. In this way, the driver (PMD 151) of the in-server data transfer device 200 has two modes: polling mode and interrupt mode.

[0132] (Second Embodiment) FIG. 13 is a schematic configuration diagram of an in-server data transfer system according to the second embodiment of the present invention. The same components as those in FIG. 1 are denoted by the same reference numerals, and the description of overlapping parts is omitted. As shown in FIG. 13, the in-server data transfer system 1000D includes an HW 110, an OS 140, and an in-server data transfer device 200A which is data high-speed transfer middleware arranged on a user space 160. Similar to the in-server data transfer device 200 in FIG. 1, the in-server data transfer device 200A consists of data high-speed transfer middleware. The in-server data transfer device 200A includes a sleep control management unit 210 and a data transfer unit 220A.

[0133] The data transfer unit 220A further includes a CPU frequency / CPU idle control unit 225 (CPU frequency control unit, CPU idle control unit) in addition to the configuration of the data transfer unit 220 in FIG. 13. The CPU frequency / CPU idle control unit 225 performs control to vary the CPU operation frequency and the CPU idle setting. Specifically, the CPU frequency / CPU idle control unit 225 of the polling thread (in-server data transfer device 200A) started by the hardware interrupt handler sets the CPU operation frequency of the CPU core used by the polling thread to be lower than that during normal use.

[0134] Here, the kernel can change the operation frequency of the CPU core by the governor setting, and the CPU frequency / CPU idle control unit 225 can use the governor setting or the like to set the CPU operation frequency to be lower than that during normal use. However, the CPU idle setting depends on the CPU model. When the CPU core enables the CPU idle setting, it can also be disabled.

[0135] The operation of the in-server data transfer system 1000D will be described below. <Rx side> FIG. 14 is a flowchart showing the operation of the data arrival monitoring unit 222 of the data transfer unit 220A. The same step numbers are assigned to the parts that perform the same processing as the flowchart shown in FIG. 9, and the description of the overlapping parts is omitted. When the data arrival monitoring unit 222 (see FIG. 13) starts immediately before data arrives in step S41, the CPU frequency / CPU idle control unit 225 (see FIG. 13) in step S81 then returns the operating frequency of the CPU core used by the data transfer unit 220A to its original value (increases the CPU operating frequency of the CPU core). Further, the CPU frequency / CPU idle control unit 225 returns the CPU idle state (depending on the CPU architecture such as C-State) setting and proceeds to step S42.

[0136] When the sleep control management unit 210 (see FIG. 13) puts the data arrival monitoring unit 222 (see FIG. 13) to sleep in step S46, the CPU frequency / CPU idle control unit 225 in step S82 sets the operating frequency of the CPU core used by the data transfer unit 220A to a lower value. Further, the CPU frequency / CPU idle control unit 225 inputs the CPU idle state (depending on the CPU architecture such as C-State) setting and ends the processing of this flow with the corresponding CPU core set to the CPU idle setting.

[0137] In this way, the in-server data transfer device 200A can also achieve further power savings by having the data transfer unit 220A include the CPU frequency / CPU idle control unit 225 and setting the CPU frequency / CPU idle state in conjunction with the sleep control of the data arrival monitoring unit 222. Note that the process of lowering the CPU frequency setting and the process of putting it into this sleep state may be executed simultaneously. Also, it may sleep after confirming that the packet transfer process has been completed.

[0138] [Application Example] The in-server data transfer devices 200 and 200A may be in-server data transfer devices that start a thread for monitoring packet arrival using the polling model in the Kernel, and the OS is not limited. Also, it is not limited to being in a server virtualization environment. Therefore, the in-server data transfer systems 1000 to 1000D can be applied to each configuration shown in FIGS. 15 and 16.

[0139] <Example of Application to VM Configuration> FIG. 15 is a diagram showing an example in which the in-server data transfer system 1000E is applied to an interrupt model in a server virtualization environment of a general-purpose Linux kernel (registered trademark) and a VM configuration. The same components as those in FIGS. 1, 13, and 19 are denoted by the same reference numerals. As shown in FIG. 15, the in-server data transfer system 1000E includes an HW10, a Host OS20, in-server data transfer devices 200 and 200A which are data high-speed transfer middleware arranged on a user space 160, a virtual switch 184, and a Guest OS70.

[0140] Specifically, the server includes a Host OS20 on which a virtual machine and an external process formed outside the virtual machine can operate, and a Guest OS70 that operates within the virtual machine. The Host OS20 includes a Kernel91, a Ring Buffer22 (see FIG. 19) managed by the Kernel91 in a memory space in the server including the Host OS20, a poll_list86 (see FIG. 22) for registering information of a network device indicating which device a hardware interrupt (hardIRQ) from the NIC11 belongs to, a vhost-net module 221A (see FIG. 19) which is a kernel thread, a tap device 222A (see FIG. 19) which is a virtual interface created by the Kernel91, and a virtual switch (br) 223A (see FIG. 19).

[0141] On the other hand, the Guest OS70 includes a Kernel181, a Driver73, a Ring Buffer52 (see FIG. 19) managed by the Kernel181 in a memory space in the server including the Guest OS70, and a poll_list86 (see FIG. 22) for registering information of a network device indicating which device a hardware interrupt (hardIRQ) from the NIC11 belongs to.

[0142] In the server internal data transfer system 1000E, the server internal data transfer devices 200 and 200A are arranged on the user space 160. Therefore, like DPDK, the data transfer unit 220 of the server internal data transfer devices 200 and 200A can bypass the kernel and refer to the ring-structured buffer. That is, the server internal data transfer devices 200 and 200A do not use the ring buffer (Ring Buffer72) (see Fig. 22) or the pole list (Ring Buffer72) (see Fig. 22) in the kernel. The data transfer unit 220 can bypass the kernel and refer to the ring-structured buffer (Ring Buffer72) (mbuf; a ring-structured buffer where PMD151 copies data by DMA), and can instantaneously grasp the arrival of packets (that is, it is a polling model instead of an interrupt model).

[0143] By doing so, in the system with the virtual server configuration of the VM, in either the Host OS20 or the Guest OS70, when there is data arrival, the kernel is bypassed in the polling mode and packet transfer is performed with low latency to achieve low latency. Also, when there is no data arrival, power consumption is reduced by stopping data arrival monitoring and sleeping. As a result, by performing sleep control with timer control considering the data arrival timing, both low latency and power saving can be achieved. Also, without modifying the APL, the delay in the server can be reduced and packet transfer can be performed.

[0144] <Example of application to container configuration> Fig. 16 is a diagram showing an example in which the server internal data transfer system 1000B is applied to the interrupt model in the server virtualization environment with the container configuration. The same components as in Fig. 15 are labeled with the same reference numerals. As shown in FIG. 16, the in-server data transfer system 1000F includes a Guest OS 180 and has a container configuration with the OS replaced by a Container 210A. The Container 210A has a vNIC (virtual NIC) 211A. The in-server data transfer devices 200 and 200A are arranged on a user space 160.

[0145] In a system with a virtual server configuration such as a container, by performing sleep control through timer control considering the data arrival timing, it is possible to achieve both low latency and power saving. Also, packet transfer can be performed with reduced latency within the server without modifying the APL.

[0146] <Example of application to a pair metal configuration (non-virtualized configuration)> The present invention can be applied to a system with a non-virtualized configuration such as a pair metal configuration. In a system with a non-virtualized configuration, by performing sleep control through timer control considering the data arrival timing, it is possible to achieve both low latency and power saving. Also, packet transfer can be performed with reduced latency within the server without modifying the APL.

[0147] <Expansion technology> When the number of traffic flows increases, the present invention can scale out against network loads by increasing the number of CPUs assigned to the packet arrival monitoring thread in cooperation with RSS (Receive-Side Scaling) that can process inbound network traffic with multiple CPUs.

[0148] <Example of application to a network system with a determined data arrival schedule> As an example of a network system with a determined data arrival schedule, it can also be applied to a high-speed packet transfer processing function unit in a network system that must guarantee data arrival timing, such as a TAS (Time Aware Shaper) in a TSN (Time Sensitive Network). In a network system with a determined data arrival schedule, it is possible to achieve both low latency and power saving.

[0149] (Third Embodiment) In the first and second embodiments, the in-server data transfer devices 200 and 200A are arranged on the user space 160. In the third embodiment, instead of the in-server data transfer devices 200 and 200A arranged on the user space 160, the in-server data transfer device 200B that deploys a polling thread in the kernel and performs sleep control is provided in the kernel.

[0150] FIG. 17 is a schematic configuration diagram of an in-server data transfer system according to the third embodiment of the present invention. The same components as those in FIGS. 1, 13, and 21 are denoted by the same reference numerals, and the description of overlapping parts is omitted. This embodiment is an example applied to packet processing by a New API (NAPI) implemented from Linux kernel 2.5 / 2.6. Note that when a polling thread is installed inside the kernel, it is necessary to consider the kernel version when using the NAPI base.

[0151] As shown in FIG. 17, the in-server data transfer system 1000G includes an HW10, an OS70, and an in-server data transfer device 200B disposed in the Kernel71 of the OS70. More specifically, the data transfer unit 220 of the in-server data transfer device 200B exists only inside the kernel71, and the sleep control management unit 210 of the in-server data transfer device 200B may exist in either the user space 160 or inside the kernel71 (the sleep control management unit 210 may be disposed in either the user space 160 or inside the kernel71). FIG. 17 shows an example in which the data transfer unit 220 and the sleep control management unit 210 (i.e., the in-server data transfer device 200B) are disposed inside the kernel71.

[0152] Here, if a configuration is adopted in which the in-server data transfer device 200B that performs sleep control is disposed inside the kernel71, the in-server data transfer devices 200 and 200A disposed on the space160 become unnecessary (in this case, considering general operation, the in-server data transfer devices 200 and 200A are disposed in the in-server data transfer system, and a mode in which the in-server data transfer devices 200 and 200A are adaptively not used is also included). The reason why the in-server data transfer devices 200 and 200A become unnecessary will be described. That is, software interrupts that cause delay problems occur only inside the kernel71 when DPDK is not used, and when DPDK is not used, data is transferred to the data processing APL1 without interrupts using the socket75. Therefore, even if the in-server data transfer devices 200 and 200A are not present on the user space 160, data can be transferred to the data processing APL1 at high speed.

[0153] OS70 includes a Kernel71, a Ring Buffer22 (see FIG. 19) managed by the Kernel71 in the memory space of a server equipped with the OS70, a poll_list86 (see FIG. 22) for registering information of a network device indicating which device a hardware interrupt (hardIRQ) from the NIC11 belongs to, a vhost-net module 221A (see FIG. 19) which is a kernel thread, a tap device 222A (see FIG. 19) which is a virtual interface created by the Kernel91, and a virtual switch (br) 223A (see FIG. 19). As described above, in the server internal data transfer device 200B, at least the data transfer unit 220 (see FIG. 1) is arranged in the Kernel71 of the OS70.

[0154] The data transfer unit 220 of the server internal data transfer device 200B has a data arrival monitoring unit 222 (see FIG. 1) for monitoring data arrival from the interface unit (NIC11). When data arrives from the interface unit, the interface unit copies the arriving data to the memory space without using the CPU by DMA (Direct Memory Access) and arranges this data using a buffer in a ring configuration. The data arrival monitoring unit 222 starts a thread for monitoring packet arrival using a polling model and detects the arrival of data by monitoring the buffer in the ring configuration.

[0155] Specifically, the data transfer unit 220 of the server internal data transfer device 200B has a Kernel (Kernel71) and a Ring Buffer72 managed by the Kernel in the memory space of a server equipped with the OS (OS70), and a poll list (poll_list86) (see FIG. 22) for registering information of a network device indicating which device a hardware interrupt (hardIRQ) from the interface unit (NIC11) belongs to, and starts a thread for monitoring packet arrival using a polling model inside the Kernel.

[0156] In this way, the data transfer unit 220 of the in-server data transfer device 200B includes a data arrival monitoring unit 222 that monitors (polls) a pole list, and when a packet has arrived, refers to the packets held in the ring buffer and executes pruning to delete the corresponding queue entry from the ring buffer based on the next process to be performed. The Rx data transfer unit (packet pruning unit) 223, and a sleep control unit 221 that puts a thread (polling thread) to sleep when no packet arrives for a predetermined period and wakes up the thread by a hardware interrupt (hardIRQ) of this thread when a packet arrives.

[0157] By doing so, the in-server data transfer device 200B stops the software interrupt (softIRQ) of packet processing, which is the main cause of NW delay. The data arrival monitoring unit 222 of the in-server data transfer device 200B executes a thread that monitors packet arrival, and the Rx data transfer unit (packet pruning unit) 223 performs packet processing by a polling model (without softIRQ) when a packet arrives. Then, when no packet arrives for a predetermined period, the sleep control unit 221 puts the thread (polling thread) to sleep, so the thread sleeps when no packet has arrived. The sleep control unit 221 wakes up the thread by a hardware interrupt (hardIRQ) when a packet arrives.

[0158] As described above, the in-server data transfer system 1000G includes an in-server data transfer device 200B that provides a polling thread in the kernel. The data transfer unit 220 of the in-server data transfer device 200B wakes up the polling thread by a hardware interrupt trigger from the NIC11. In particular, when providing a polling thread in the kernel, the data transfer unit 220 is characterized by being woken up by a timer. Thereby, the in-server delay control device 200B can achieve both low latency and power saving by performing sleep management of the polling thread that performs packet transfer processing.

[0159] [Hardware Configuration] The in-server data transfer devices 200, 200A, and 200B according to the above embodiments are realized by a computer 900 having a configuration as shown in FIG. 18, for example. FIG. 18 is a hardware configuration diagram showing an example of a computer 900 that realizes the functions of the in-server data transfer devices 200 and 200A. The computer 900 includes a CPU 901, a ROM 902, a RAM 903, an HDD 904, a communication interface (I / F) 906, an input / output interface (I / F) 905, and a media interface (I / F) 907.

[0160] The CPU 901 operates based on a program stored in the ROM 902 or the HDD 904, and controls each part of the in-server data transfer devices 200, 200A, and 200B shown in FIGS. 1 and 13. The ROM 902 stores a boot program executed by the CPU 901 when the computer 900 is started up, a program dependent on the hardware of the computer 900, and the like.

[0161] The CPU 901 controls an input device 910 such as a mouse or a keyboard and an output device 911 such as a display via the input / output I / F 905. The CPU 901 acquires data from the input device 910 via the input / output I / F 905 and outputs the generated data to the output device 911. Note that, as the processor, a GPU (Graphics Processing Unit) or the like may be used together with the CPU 901.

[0162] The HDD 904 stores programs executed by the CPU 901, data used by the programs, and the like. The communication I / F 906 receives data from other devices via a communication network (for example, NW (Network) 920) and outputs the data to the CPU 901, and transmits the data generated by the CPU 901 to other devices via the communication network.

[0163] The media I / F 907 reads a program or data stored in the recording medium 912 and outputs the program or data to the CPU 901 via the RAM 903. The CPU 901 loads a program related to a target process from the recording medium 912 onto the RAM 903 via the media I / F 907 and executes the loaded program. The recording medium 912 is an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto Optical disk), a magnetic recording medium, a conductor memory tape medium, or a semiconductor memory or the like.

[0164] For example, when the computer 900 functions as the in-server data transfer devices 200, 200A, and 200B configured as one device according to the present embodiment, the CPU 901 of the computer 900 realizes the functions of the in-server data transfer device 100 by executing the program loaded on the RAM 903. Also, the data in the RAM 903 is stored in the HDD 904. The CPU 901 reads and executes the program related to the target process from the recording medium 912. Additionally, the CPU 901 may read the program related to the target process from another device via the communication network (NW920).

[0165] [Effect] As described above, an in-server data transfer device 200 that performs data transfer control of the interface unit (accelerator 120, NIC 130) in user space, where the OS (OS 70) includes a kernel (Kernel 171), a ring buffer (mbuf; a buffer with a ring structure where PMD 151 copies data by DMA) in the memory space of the server equipped with the OS, and a driver (PMD 151) that can select data arrival from the interface unit (accelerator 120, NIC 130) in polling mode or interrupt mode, and a data transfer unit 220 that starts a thread (polling thread) to monitor packet arrival using the polling model, and a sleep control management unit (sleep control management unit 210) that manages data arrival schedule information and distributes the data arrival schedule information to the data transfer unit 220 to perform sleep control of the data transfer unit 220. The data transfer unit 220 puts the thread to sleep based on the data arrival schedule information distributed from the sleep control management unit 210, and activates a timer immediately before data arrival to wake up the thread for sleep release.

[0166] By doing so, the sleep control management unit 210 collectively controls the Sleep / activation timing of each data transfer unit 220 in order to perform sleep control of a plurality of data transfer units in accordance with the data arrival timing. When there is data arrival, the kernel is bypassed in the polling mode and packet transfer is performed with low latency, thereby achieving low latency. Also, when there is no data arrival, data arrival monitoring is stopped and the device goes to sleep, thereby achieving power saving. As a result, by performing sleep control through timer control considering the data arrival timing, it is possible to achieve both low latency and power saving.

[0167] The in-server data transfer device 200 can achieve low latency by implementing data transfer delay within the server in the polling model instead of the interrupt model. That is, in the in-server data transfer device 200, similar to DPDK, the data transfer unit 220 arranged in the user space 160 can bypass the kernel and refer to the buffer with a ring structure. Then, by constantly monitoring this ring-structured buffer by the polling thread, it is possible to instantly grasp packet arrival (it is the polling model, not the interrupt model).

[0168] Also, for a data flow such as time-division multiplexed data flow where the data arrival timing is fixedly determined, like in signal processing in vRAN, by performing sleep control of the data transfer unit 220 in consideration of the data arrival schedule, it is possible to reduce the CPU usage rate while maintaining low latency, and achieve power saving. That is, by performing sleep control through timer control considering the data arrival timing for the problem of wasteful use of CPU resources in the polling model, it is possible to achieve both low latency and power saving.

[0169] In addition, a Guest OS (Guest OS70) operating within a virtual machine includes a kernel (Kernel171), a ring buffer (mbuf; a buffer having a ring structure for PMD151 to copy data by DMA) in the memory space of a server equipped with the Guest OS, a driver (PMD151) capable of selecting data arrival from an interface unit (accelerator 120, NIC130) in either polling mode or interrupt mode, and a protocol processing unit 74 that performs protocol processing on packets for which pruning has been executed. It also includes a data transfer unit 220 that starts a thread (polling thread) to monitor packet arrival using the polling model, and a sleep control management unit (sleep control management unit 210) that manages data arrival schedule information and distributes the data arrival schedule information to the data transfer unit 220 to perform sleep control on the data transfer unit 220. The data transfer unit 220 is characterized by putting the thread to sleep based on the data arrival schedule information distributed by the sleep control management unit 210 and activating a timer immediately before data arrival to perform sleep release to wake up the thread.

[0170] By doing so, in a system with a virtual server configuration of a VM, for a server equipped with a Guest OS (Guest OS70), it is possible to reduce the CPU usage rate while maintaining low latency, and achieve power savings.

[0171] Also, a Host OS (Host OS20) on which a virtual machine and an external process formed outside the virtual machine can operate includes a kernel (Kernel91), a ring buffer (mbuf; a buffer having a ring structure for PMD151 to copy data by DMA) in the memory space of the server equipped with the Host OS, a driver (PMD151) capable of selecting data arrival from the interface unit (accelerator 120, NIC130) in polling mode or interrupt mode, and a tap device 222A which is a virtual interface created by the kernel (Kernel91). A data transfer unit 220 that starts a thread (polling thread) for monitoring packet arrival using the polling model, and a sleep control management unit (sleep control management unit 210) that manages data arrival schedule information and distributes the data arrival schedule information to the data transfer unit 220 to perform sleep control on the data transfer unit 220. The data transfer unit 220 is characterized by putting the thread to sleep based on the data arrival schedule information distributed from the sleep control management unit 210 and activating a timer immediately before data arrival to perform wake-up of the thread.

[0172] By doing so, in a system with a virtual server configuration of a VM, for a server equipped with a kernel (Kernel191) and a Host OS (Host OS20), it is possible to reduce the CPU usage rate while maintaining low latency, and it is possible to achieve power saving.

[0173] Also, there is a server internal data transfer device 200B, where the OS (OS 70) includes a kernel (Kernel 171), a ring buffer (Ring Buffer 72) managed by the kernel in the memory space of the server equipped with the OS, a poll list (poll_list 86) for registering network device information indicating which device the hardware interrupt (hardIRQ) from the interface unit (NIC 11) belongs to, a data transfer unit 220 that starts a thread in the kernel to monitor packet arrival using the polling model, a sleep control management unit (sleep control management unit 210) that manages the data arrival schedule, manages the data arrival schedule information, and distributes the data arrival schedule information to the data transfer unit 220 to perform sleep control on the data transfer unit 220. The data transfer unit 220 includes a data arrival monitoring unit 222 that monitors (polls) the poll list, a packet pruning unit (Rx data transfer unit 223) that, when a packet has arrived, refers to the packet held in the ring buffer and executes pruning to delete the corresponding queue entry from the ring buffer based on the next process to be performed, and a sleep control unit (sleep control unit 221) that puts the thread (polling thread) to sleep based on the data arrival schedule information received from the sleep control management unit 210 and wakes up from the sleep by a hardware interrupt (hardIRQ) when the sleep is released.

[0174] By doing so, the in-server data transfer device 200B can achieve low latency by implementing data transfer delay within the server in a polling model instead of an interrupt model. In particular, for data flows with fixed data arrival timings such as time-division multiplexed data flows in vRAN, by performing sleep control on the data transfer unit 220 in consideration of the data arrival schedule, it is possible to reduce the CPU usage rate while maintaining low latency, thereby achieving power saving. That is, by performing sleep control on the problem of wasteful use of CPU resources in the polling model through timer control considering the data arrival timing, it is possible to achieve both low latency and power saving.

[0175] Based on the data arrival schedule information received from the sleep control management unit 210, the data transfer unit 220 puts a thread (polling thread) to sleep and wakes up from the sleep by a hardware interrupt (hardIRQ) when the sleep is released. As a result, in addition to the above effects, the effects of (1) to (2) are further achieved.

[0176] (1) Stop the software interrupt (softIRQ) at the time of packet arrival, which is the cause of delay, and implement the polling model within the kernel (Kernel171). That is, the in-server data transfer system 1000G implements a polling model instead of an interrupt model, which is the main cause of NW delay, unlike the existing NAPI technology. Since packets are immediately harvested without waiting at the time of arrival, low-latency packet processing can be achieved.

[0177] (2) The polling thread in the in-server data transfer device 200 operates as a kernel thread and monitors packet arrivals in polling mode. The kernel thread (polling thread) that monitors packet arrivals sleeps while there are no packet arrivals. When there are no packet arrivals, since it does not use the CPU due to sleeping, the effect of power saving can be obtained.

[0178] And when a packet arrives, the polling thread that is sleeping is woken up (the sleep is released) by the hard IRQ handler at the time of packet arrival. By releasing the sleep by the hard IRQ handler, the polling thread can be immediately started while avoiding soft IRQ contention. Here, the sleep release is characterized in that it is not caused by a timer but by the hard IRQ handler. In addition, when the traffic load is known in advance, for example, when a 30ms sleep is known as in the workload transfer rate shown in FIG. 23, it may be caused by the hard IRQ handler in accordance with this timing.

[0179] In this way, the in-server data transfer device 200B can achieve both low latency and power saving by performing sleep management of the polling thread that performs packet transfer processing.

[0180] The in-server data transfer device 200A is characterized in that it includes a CPU frequency setting unit (CPU frequency / CPU idle control unit 225) that sets the CPU operation frequency of the CPU core used by the thread to be low during sleep.

[0181] In this way, the in-server data transfer device 200A dynamically varies the CPU operation frequency according to the traffic, that is, if the CPU is not used due to sleeping, the CPU operation frequency during sleep is set low, so that the effect of power saving can be enhanced.

[0182] The in-server data transfer device 200A is characterized by including a CPU idle setting unit (CPU frequency / CPU idle control unit 225) that sets the CPU idle state of the CPU core used by a thread to a power-saving mode during sleep.

[0183] By doing so, the in-server data transfer device 200A can enhance the power-saving effect by dynamically varying the CPU idle state (power-saving function according to the CPU model such as changing the operating voltage) according to the traffic.

[0184] Among the processes described in each of the above embodiments, all or part of the processes described as being automatically performed can also be performed manually, or all or part of the processes described as being performed manually can also be automatically performed by a known method. In addition, regarding the processing procedures, control procedures, specific names, information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified. Moreover, each component of each illustrated device is conceptually functional and does not necessarily need to be physically configured as shown in the figure. That is, the specific form of the distribution and integration of each device is not limited to that shown in the figure, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads and usage situations.

[0185] In addition, each of the above configurations, functions, processing units, processing means, etc. may be realized in hardware by designing part or all of them, for example, with an integrated circuit. Also, each of the above configurations, functions, etc. may be realized by software for a processor to interpret and execute programs that realize their respective functions. Information such as programs, tables, files, etc. that realize each function can be held in a memory, a recording device such as a hard disk, an SSD (Solid State Drive), or a recording medium such as an IC (Integrated Circuit) card, an SD (Secure Digital) card, or an optical disk.

Description of Symbols

[0186] 1 Data Processing APL (Application) 2 Data Flow Time Slot Management Scheduler 3 PHY (High) 4 MAC 5 RLC 6 FAPI (FAPI P7) 20,70 Host OS (OS) 50 Guest OS (OS) 86 poll_list (Poll List) 72 Ring Buffer (Ring Buffer) 91,171,181 Kernel (Kernel) 110 HW 120 Accelerator (Interface Section) 121 Core (Core Processor) 122,131 Rx Queue 123,132 Tx Queue 130 NIC (Physical NIC) (Interface Section) 140 OS 151 PMD (Driver Capable of Selecting Data Arrival in Polling Mode or Interrupt Mode) 160 user space (User Space) 200,200A,200B Server Internal Data Transfer Device 210 sleep Control Management Section 210A Container 211 Data Transfer Section Management Section 212 Data Arrival Schedule Management Section 213 Data Arrival Schedule Distribution Section 220 Data Transfer Section 221 sleep Control Section 222 Data Arrival Monitoring Section 223 Rx Data Transfer Section (Packet Harvesting Section) 224 Tx Data Transfer Section 225 CPU Frequency / CPU Idle Control Unit (CPU Frequency Control Unit, CPU Idle Control Unit) 1000, 1000A, 1000B, 1000C, 1000D, 1000E, 1000F, 1000G Server Internal Data Transfer System Ring-Structured Buffer Where Mbuf PMD Copies Data via DMA

Claims

1. A data transfer device within a server, comprising a sleep control management unit that manages a data arrival schedule and performs sleep control in accordance with the data arrival timing, wherein the sleep control management unit, in a data flow having information regarding a specific data arrival timing, puts to sleep a thread that monitors the arrival of data using a polling model based on the data arrival timing, and performs an operation to wake up the thread based on the data arrival timing characterizing the data transfer device within the server.

2. The data flow having information regarding a specific data arrival timing in Claim 1 is signal processing in a RAN (Radio Access Network) characterizing the data transfer device within the server.

3. The specific data arrival timing in Claim 2 is obtained from MAC (Medium Access Control) scheduler information that manages the data transmission and reception of radio signals in a RAN characterizing the data transfer device within the server.

4. The MAC scheduler information in Claim 3 is obtained via an IF that exchanges data schedule information connecting a PHY (Physical) layer and a MAC layer characterizing the data transfer device within the server.

5. The specific data arrival timing in Claim 2 is obtained from a message flowing through an IF (Interface) between a transmission device facing a RAN device characterizing the data transfer device within the server.

6. The data arrival timing in Claim 2 is obtained from an IF that exchanges data schedule information between a transmission device in an O-RAN characterizing the data transfer device within the server.

7. A data transfer method within a server for a data transfer device within a server that manages a data arrival schedule and performs sleep control in accordance with the data arrival timing, wherein the data transfer device within the server in a data flow having information regarding a specific data arrival timing, executes a step of putting to sleep a thread that monitors the arrival of data using a polling model based on the data arrival timing, and a step of performing an operation to wake up the thread based on the data arrival timing characterizing the data transfer method within the server.

8. A program for causing a computer to function as the in-server data transfer device according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Low-power-consumption adaptive polling

    JP2004199683A

  • Portable telephone set

    JP2005110050A

  • Radio base station and communication control method

    JP2017188834A

  • Techniques for Received Packet Processing and Associated Power Management in Network Devices

    JP2018507457A

  • Variable polling interval based on historical timing results

    US20090089784A1