Intra-server data transfer device, intra-server data transfer method and program

By setting up a data transmission device in the user space of the server, the device uses a sleep control mechanism to manage CPU activities based on the data arrival timeline, solving the problems of high CPU usage and large power consumption in the prior art, and realizing low latency and low power consumption data transmission.

JP7673808B2Active Publication Date: 2025-05-09NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023536248
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-19
Publication Date
2025-05-09
Estimated Expiration
2041-07-19

AI Technical Summary

Technical Problem

When transmitting data in servers, the prior art faces the problems of high CPU usage and high power consumption, and it is difficult to maintain low-latency data transmission.

Method used

Using a new model of data transmission in a server, by setting up a data transmission device in the user space, the device includes a thread monitoring data arrival and a unit managing sleep control. The device performs sleep control based on the schedule of data arrival, and only wakes up the thread when expected data arrives, reducing unnecessary CPU activity.

Benefits of technology

Effectively reduces CPU usage and power consumption, while maintaining low latency data transmission performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007673808000001
    Figure 0007673808000001
  • Figure 0007673808000002
    Figure 0007673808000002
  • Figure 0007673808000003
    Figure 0007673808000003
Patent Text Reader

Abstract

Provided is a server internal data transfer device (200) for performing data transfer control of an interface unit in a user space, the server internal data transfer device comprising a data transfer unit (220) that activates a thread for monitoring arrival of a packet by using a polling model, and a sleep control management unit (210) that manages data arrival schedule information and performs sleep control of the data transfer unit (220) by distributing the data arrival schedule information to the data transfer unit (220), wherein the data transfer unit (220) causes a thread to sleep and performs sleep cancellation of initiating a timer immediately before arrival of data to activate the thread on the basis of the data arrival schedule information distributed from the sleep control management unit (210).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an intra-server data transfer device, an intra-server data transfer method, and a program. [Background technology]

[0002] With the advancement of virtualization technology such as NFV (Network Functions Virtualization), systems are being built and operated for each service.In addition, a form called Service Function Chaining (SFC) is becoming mainstream, in which service functions are divided into reusable modules and run on independent virtual machine (VM: Virtual Machine, container, etc.) environments, allowing them to be used as components when needed and improving operability, as opposed to the form in which systems are built for each service.

[0003] A known technology for configuring virtual machines is a hypervisor environment consisting of Linux (registered trademark) and KVM (kernel-based virtual machine). In this environment, a Host OS (an OS installed on a physical server is called a Host OS) with a built-in KVM module operates as a hypervisor in a memory area called kernel space, which is different from the user space. In this environment, a virtual machine operates in the user space, and a Guest OS (an OS installed on a virtual machine is called a Guest OS) operates within the virtual machine.

[0004] A virtual machine running a Guest OS is different from a physical server running a Host OS in that all HW (hardware), including network devices (typified by Ethernet (registered trademark) card devices, is subject to register control required for interrupt processing from the HW to the Guest OS and writing from the Guest OS to the hardware. In this type of register control, notifications and processing that should be executed by physical hardware are simulated by software, so performance is generally lower than in a Host OS environment.

[0005] To address this performance degradation, there is a technology that reduces HW emulation, particularly from the Guest OS to the Host OS and external processes that exist outside the virtual machine, and improves communication performance and versatility through a high-speed, unified interface. One such technology is a device abstraction technology called Virtio, or paravirtualization technology, which has already been incorporated into many general-purpose operating systems, including Linux (registered trademark) and FreeBSD (registered trademark), and is currently in use (see Patent Documents 1 and 2).

[0006] Virtio defines data exchange by queues designed with ring buffers as a one-way transport for data input / output such as console, file input / output, and network communication, by queue operations. By using the Virtio queue specifications to prepare the number and size of queues appropriate for each device when the Guest OS starts, communication between the Guest OS and the outside of the virtual machine can be achieved by queue operations alone, without performing hardware emulation.

[0007] [Packet forwarding using the interrupt model (example of a generic VM configuration)] FIG. 19 is a diagram for explaining packet transfer according to an interrupt model in a server virtualization environment having a generic Linux kernel (registered trademark) and a VM configuration. The HW10 has a NIC (Network Interface Card) 11 (physical NIC) (interface unit), and transmits and receives data to and from a data processing APL (Application) 1 on a user space 60 via a virtual communication path constructed by a Host OS 20, a KVM 30 which is a hypervisor that constructs virtual machines, virtual machines (VM1, VM2) 40, and a Guest OS 50. In the following description, as shown by the thick arrows in Fig. 19, the data flow in which the data processing APL1 receives packets from the HW10 is referred to as Rx side reception, and the data flow in which the data processing APL1 transmits packets to the HW10 is referred to as Tx side transmission.

[0008] The host OS 20 includes a kernel 21, a ring buffer 22, and a driver 23. The kernel 21 includes a vhost-net module 221A, which is a kernel thread, a tap device 222A, and a virtual switch (br) 223A.

[0009] The tap device 222A is a kernel device of a virtual network and is supported by software. The virtual machine (VM1) 40 allows communication between the Guest OS 50 and the Host OS 20 via a virtual switch (br) 223A created in a virtual bridge. The tap device 222A is a device connected to a virtual NIC (vNIC) of the Guest OS 50 created in this virtual bridge.

[0010] The Host OS 20 copies the configuration information (size of the shared buffer queue, number of queues, identifier, top address information for accessing the ring buffer, etc.) constructed within the virtual machine of the Guest OS 50 to the vhost-net module 221A, and constructs the endpoint information on the virtual machine side within the Host OS 20. This vhost-net module 221A is a kernel-level backend for virtio networking, and can reduce the virtualization overhead by transferring the virtio packet processing task from the user area (user space) to the vhost-net module 221A of the kernel 21.

[0011] The Guest OS 50 includes a Guest OS (Guest 1) installed on a virtual machine (VM1) and a Guest OS (Guest 2) installed on a virtual machine (VM2), and the Guest OS 50 (Guest 1, Guest 2) run in the virtual machines (VM1, VM2) 40. Taking Guest 1 as an example of the Guest OS 50, the Guest OS 50 (Guest 1) includes a kernel 51, a ring buffer 52, and a driver 53, and the driver 53 includes a virtio-driver 531.

[0012] Specifically, within the virtual machine there are virtio devices for the console, file I / O, and network communication as PCI (Peripheral Component Interconnect) devices (the console is virtio-console, file I / O is virtio-blk, and the network is virtio-net, and the corresponding drivers that the OS has are defined in the virtio queue), and when the Guest OS starts up, two data transfer endpoints (sending and receiving endpoints) are created between the Guest OS and the other side, establishing a parent-child relationship for sending and receiving data. In many cases, the parent-child relationship is configured between the virtual machine (child side) and the Guest OS (parent side).

[0013] The child side exists as device configuration information within the virtual machine, and requests the parent side for the size of each data area, the number of endpoint combinations required, and the device type. In accordance with the child side's request, the parent side allocates and secures memory for a shared buffer queue to store and transfer the required amount of data, and returns the address to the child side so that the child can access it. In Virtio, all shared buffer queue operations required for data transfer are common, and are executed with both the parent and child sides having agreed upon them. Furthermore, the size of the shared buffer queue is also agreed upon by both sides (i.e. it is determined for each device). This makes it possible for both the parent and child sides to operate the shared queue simply by communicating the address to the child side.

[0014] The shared buffer queues prepared in virtio are prepared for single direction use, so for example, a virtual network device called a virtio-net device consists of three Ring Buffers 52: one for sending, one for receiving, and one for control. Communication between parent and child is achieved by writing to the shared buffer queue and notifying the buffer update, and after writing to the Ring Buffer 52, the other side is notified. When the other side receives the notification, it uses virtio's common operations to check how much new data is in which shared buffer queue, and extracts new buffer space. This completes the transfer of data from parent to child or from child to parent.

[0015] As described above, the parent and child share the Ring Buffer 52 for data exchange with each other and the operation method for each ring buffer (common to virtio), thereby realizing communication between the Guest OS 50 and the outside world without the need for hardware emulation. This makes it possible to send and receive data between the Guest OS 50 and the outside world at high speeds compared to conventional hardware emulation.

[0016] When Guest OS50 in a virtual machine communicates with the outside, the child side must connect to the outside and act as a relay between the outside and the parent side to send and receive data. One example is communication between Guest OS50 and Host OS20. Here, if the outside is the Host OS20, there are two existing communication methods.

[0017] The first method (hereinafter referred to as external communication method 1) establishes a child endpoint within the virtual machine, and connects within the virtual machine the communication between Guest OS 50 and the virtual machine and the communication endpoint (usually called a tap / tun device) provided by Host OS 20. This connection establishes the following connection, realizing communication from Guest OS 50 to Host OS 20.

[0018] At this time, the Guest OS 50 runs in a memory area that is a user space with different privileges from the memory area called the kernel space in which the tap driver and the Host OS 20 run. For this reason, communication from the Guest OS 50 to the Host OS 20 requires at least one memory copy.

[0019] The second method (hereinafter referred to as external communication method 2) is a technology called vhost-net that solves this problem. In vhost-net, the parent configuration information (size of the shared buffer queue, number of queues, identifier, top address information for accessing the ring buffer, etc.) once constructed within the virtual machine is copied to the vhost-net module 221A inside the Host OS 20, and the child end point information is constructed within the host. This construction makes it possible to directly operate the shared buffer queue between the Guest OS 50 and the Host OS 20. This effectively eliminates the need for copying, and since there is one less copy time compared to virtio-net, faster data transfer can be achieved compared to external communication method 1.

[0020] In this way, by reducing the number of memory copies related to virtio-net in the Host OS 20 and the Guest OS 50 connected by virtio, the packet transfer process can be accelerated.

[0021] In addition, since kernel v4.10 (2017.2~), the specifications of the tap interface have been changed so that packets inserted from the tap device are completed within the same context as the process that copied the packet to the tap device. This has eliminated the occurrence of software interrupts (softIRQs).

[0022] [Packet forwarding using the polling model (DPDK example)] The method of connecting and coordinating multiple virtual machines is called Inter-VM Communication, and in large-scale environments such as data centers, virtual switches have been used as a standard method for connecting VMs. However, this method has a large communication delay, so new faster methods have been proposed. For example, a method that uses special hardware called SR-IOV (Single Root I / O Virtualization) and a software method that uses Intel DPDK (Intel Data Plane Development Kit) (hereinafter referred to as DPDK), a high-speed packet processing library, have been proposed (see Non-Patent Document 1).

[0023] DPDK is a framework for controlling NICs (Network Interface Cards), which was previously handled by the Linux kernel (registered trademark), in user space. The biggest difference with the processing in the Linux kernel is that it has a polling-based reception mechanism called PMD (Pull Mode Driver). Normally, in the Linux kernel, an interrupt occurs when data arrives at the NIC, which triggers the reception process. On the other hand, in PMD, a dedicated thread continuously checks whether data has arrived and performs the reception process. By eliminating overhead such as context switches and interrupts, high-speed packet processing can be achieved. DPDK significantly improves the performance and throughput of packet processing, making it possible to secure more time for data plane application processing.

[0024] DPDK exclusively uses computer resources such as CPU (Central Processing Unit) and NIC. For this reason, it is difficult to apply to applications that flexibly change connections on a module basis like SFC. There is an application called SPP (Soft Patch Panel) that alleviates this problem. SPP provides shared memory between VMs and configures each VM to directly reference the same memory space, thereby omitting packet copying in the virtualization layer. In addition, DPDK is used to speed up packet exchange between physical NICs and shared memory. SPP can change the input and output destinations of packets in software by controlling the reference destination of memory exchange for each VM. Through this process, SPP realizes dynamic connection switching between VMs and between VMs and physical NICs (see Non-Patent Document 2).

[0025] Fig. 20 is a diagram for explaining packet forwarding by a polling model in the configuration of OvS-DPDK (Open vSwitch with DPDK). The same components as those in Fig. 19 are given the same reference numerals and the description of the overlapping parts is omitted. As shown in FIG. 20, the host OS 20 includes an OvS-DPDK 70, which is software for packet processing. The OvS-DPDK 70 includes a vhost-user 71, which is a functional unit for connecting to a virtual machine (here, VM1), and a dpdk (PMD) 72, which is a functional unit for connecting to a NIC (DPDK) 11 (physical NIC). 19. Moreover, the data processing APL1A includes a dpdk(PMD)2 which is a functional unit that performs polling in the Guest OS 50 section. That is, the data processing APL1A is an APL obtained by modifying the data processing APL1 in FIG.

[0026] Packet forwarding using the polling model is an extension of DPDK, and enables route operations using a GUI in SPP, which performs high-speed packet copying between Host OS 20 and Guest OS 50, and between Guest OS 50, with zero copy via shared memory.

[0027] [Rx side packet processing using New API (NAPI)] 21 is a schematic diagram of Rx-side packet processing by New API (NAPI) implemented in Linux kernel 2.5 / 2.6 (see Non-Patent Document 1). The same components as those in FIG. 19 are denoted by the same reference numerals. As shown in FIG. 21, the New API (NAPI) executes a data processing APL1 placed in a user space 60 available to a user on a server having an OS70 (e.g., a Host OS), and transfers packets between the NIC 11 of the HW 10 connected to the OS70 and the data processing APL1.

[0028] The OS 70 includes a kernel 71 , a ring buffer 72 , and a driver 73 , and the kernel 71 includes a protocol processing unit 74 . Kernel 71 is a core function of OS 70 (for example, a host OS), and monitors hardware and manages the execution status of programs on a process-by-process basis. Here, kernel 71 responds to requests from data processing APL1, and also conveys requests from HW 10 to data processing APL1. Kernel 71 processes requests from data processing APL1 via a system call (a "user program running in non-privileged mode" requests processing from the "kernel running in privileged mode"). The Kernel 71 transmits a packet to the data processing APL1 via the Socket 75. The Kernel 71 receives a packet from the data processing APL1 via the Socket 75.

[0029] The ring buffer 72 is located in the memory space of the server and is managed by the Kernel 71. The ring buffer 72 is a buffer of a certain size that stores messages output by the Kernel 71 as logs, and when the upper limit size is exceeded, the messages are overwritten from the top.

[0030] Driver73 is a device driver that monitors hardware with kernel71. Driver73 depends on kernel71, and if the created (built) kernel source changes, it will become a different driver. In this case, you will need to obtain the driver source, rebuild it on the OS that uses the driver, and create the driver.

[0031] The protocol processing unit 74 performs protocol processing of L2 (data link layer), L3 (network layer), and L4 (transport layer) defined by the OSI (Open Systems Interconnection) reference model.

[0032] Socket75 is an interface that kernel71 uses for inter-process communication. Socket75 has a socket buffer and does not require frequent data copying. The process for establishing communication via Socket75 is as follows: 1. The server creates a socket file that accepts clients. 2. The acceptance socket file is named. 3. A socket queue is created. 4. The first connection from a client in the socket queue is accepted. 5. A socket file is created on the client side. 6. The client issues a connection request to the server. 7. On the server side, a connection socket file is created in addition to the acceptance socket file. As a result of establishing communication, data processing APL1 becomes able to call system calls such as read() and write() on kernel71.

[0033] In the above configuration, the Kernel 71 receives notification of the arrival of a packet from the NIC 11 via a hardware interrupt (hardIRQ), and schedules a software interrupt (softIRQ) for packet processing. The New API (NAPI) implemented in Linux kernel 2.5 / 2.6 above processes packets by a hardware interrupt (hardIRQ) and then a software interrupt (softIRQ) when a packet arrives. As shown in Figure 21, packet transfer using the interrupt model transfers packets using interrupt processing (see symbol c in Figure 21), which creates a wait for the interrupt processing and increases the delay in packet transfer.

[0034] The following provides an overview of packet processing on the NAPI Rx side. [Rx side packet processing configuration using New API (NAPI)] FIG. 22 is a diagram for explaining an outline of Rx-side packet processing by New API (NAPI) in the area surrounded by the dashed line in FIG. <Device driver> As shown in Figure 22, the device driver includes NIC11 (physical NIC), which is a network interface card, hardIRQ81, which is a handler that is called when a processing request is made to NIC11 and executes the requested processing (hardware interrupt), and netif_rx82, which is a processing function for software interrupts.

[0035] <Networking layer> In the networking layer, there are disposed softIRQ 83, which is a handler that is called when a processing request from netif_rx 82 occurs and executes the requested processing (software interrupt), and do_softirq 84, which is a control function unit that executes the software interrupt (softIRQ).In addition, there are disposed net_rx_action 85, which is a packet processing function unit that receives and executes a software interrupt (softIRQ), poll_list 86, which registers information on the net device (net_device) that indicates which device the hardware interrupt from the NIC 11 is from, netif_receive_skb 87, which creates an sk_buff structure (a structure that enables the Kernel 71 to recognize the state of the packet), and the Ring Buffer 72.

[0036] <Protocol layer> In the protocol layer, packet processing functional units such as ip_rcv88 and arp_rcv89 are arranged.

[0037] The above netif_rx82, do_softirq84, net_rx_action85, netif_receive_skb87, ip_rcv88, and arp_rcv89 are program components (function names) used for packet processing in Kernel71.

[0038] [Rx side packet processing operation using New API (NAPI)] The arrows (symbols) d to o in FIG. 22 indicate the flow of packet processing on the Rx side. When the hardware function unit 11a of the NIC 11 (hereinafter referred to as NIC 11) receives a packet (or a frame) in a frame from the other device, it copies the arriving packet to the Ring Buffer 72 by DMA (Direct Memory Access) transfer without using the CPU (see symbol d in FIG. 22). This Ring Buffer 72 is a memory space in the server, and is managed by the Kernel 71 (see FIG. 21).

[0039] However, if the NIC 11 simply copies an arriving packet to the Ring Buffer 72, the Kernel 71 cannot recognize the packet. Therefore, when a packet arrives, the NIC 11 raises a hardware interrupt (hardIRQ) to hardIRQ 81 (see symbol e in FIG. 22), and netif_rx 82 executes the following process, which allows the Kernel 71 to recognize the packet. Note that hardIRQ 81, shown in an oval in FIG. 22, indicates a handler rather than a functional unit.

[0040] Netif_rx82 is a function that actually performs processing. When hardIRQ81 (handler) is started (see symbol f in FIG. 22), it saves information on the net device (net_device) that indicates which device the hardware interrupt from NIC11 is from, which is one of the pieces of information on the contents of the hardware interrupt (hardIRQ), in poll_list86, and registers queue pruning (referring to the contents of the packets stored in the buffer, deleting the corresponding queue entry from the buffer in consideration of the processing of the packets to be performed next) (see symbol g in FIG. 22). Specifically, when packets are stuffed into the Ring Buffer72, netif_rx82 uses the driver of NIC11 to register future queue pruning in poll_list86 (see symbol g in FIG. 22). As a result, queue pruning information resulting from the packing of packets into the Ring Buffer72 is registered in poll_list86.

[0041] In this way, in Fig. 22<Device driver> In the above, when the NIC 11 receives a packet, it uses DMA transfer to copy the packet that has arrived to the Ring Buffer 72. In addition, the NIC 11 raises the hardIRQ 81 (handler), and the netif_rx 82 registers the net_device in the poll_list 86 and schedules a software interrupt (softIRQ). So far, we have completed the steps in Figure 22.<Device driver> Processing of hardware interrupts in stops.

[0042] Thereafter, netif_rx82 uses the information (specifically, the pointer) in the queue stored in poll_list86 to notify softIRQ83 (handler) of the reaping of the data stored in Ring Buffer72 via a software interrupt (softIRQ) (see symbol h in Figure 22), and notifies do_softirq84, which is the software interrupt control function unit (see symbol i in Figure 22).

[0043] do_softirq84 is a software interrupt control function section, which defines each function of software interrupt (there are various types of packet processing, and interrupt processing is one of them. This defines interrupt processing). Based on this definition, do_softirq84 notifies net_rx_action85, which actually processes the software interrupt, of the current (corresponding) software interrupt request (see symbol j in Figure 22).

[0044] When the turn of the softIRQ comes, net_rx_action 85 calls a polling routine for reaping packets from Ring Buffer 72 based on net_device registered in poll_list 86 (see symbol k in FIG. 22), and reaps the packets (see symbol l in FIG. 22). At this time, net_rx_action 85 continues reaping until poll_list 86 becomes empty. Thereafter, net_rx_action 85 notifies netif_receive_skb 87 (see symbol m in FIG. 22).

[0045] The netif_receive_skb 87 creates an sk_buff structure, analyzes the contents of the packet, and sends the packet to the downstream protocol processor 74 (see FIG. 21) for processing according to the packet type. That is, the netif_receive_skb 87 analyzes the contents of the packet and, when processing is to be performed according to the contents of the packet,<Protocol layer> If it is L2, the process is passed to arp_rcv 89 (symbol o in FIG. 22).

[0046] Non-Patent Document 3 describes a server network delay control device (KBP: Kernel Busy Poll). KBP constantly monitors packet arrivals using a polling model within the kernel. This suppresses softIRQ and achieves low-latency packet processing.

[0047] Fig. 23 shows an example of video (30 FPS) data transfer. The workload shown in Fig. 23 has a transfer rate of 350 Mbps, and data is transferred intermittently every 30 ms.

[0048] FIG. 24 is a diagram showing the CPU utilization rate used by the busy poll thread in the KBP described in Non-Patent Document 3. In FIG. As shown in Fig. 24, in KBP, the kernel thread occupies a CPU core to perform busy poll. Even in the case of intermittent packet reception as shown in Fig. 23, KBP has a problem of high power consumption because the CPU is always used regardless of whether packets arrive or not.

[0049] Next, we will explain the DPDK system. [DPDK system configuration] FIG. 25 is a diagram showing a configuration of a DPDK system that controls the HW 110 including the accelerator 120. The DPDK system includes HW 110, OS 140, DPDK 150, which is high-speed data transfer middleware arranged on user space 160, and data processing APL 1. The data processing APL1 is a packet processing carried out prior to the execution of the APL. The HW 110 communicates with the data processing APL 1 to transmit and receive data. In the following description, as shown in Fig. 25, a data flow in which the data processing APL 1 receives packets from the HW 110 is referred to as Rx side reception, and a data flow in which the data processing APL 1 transmits packets to the HW 110 is referred to as Tx side transmission.

[0050] The HW 110 includes an accelerator 120 and a NIC 130 (physical NIC) for connecting to a communication network. The accelerator 120 is a calculation unit hardware that performs a specific calculation at high speed based on an input from a CPU. Specifically, the accelerator 120 is a PLD (Programmable Logic Device) such as a GPU (Graphics Processing Unit) or an FPGA (Field Programmable Gate Array). In Fig. 25, the accelerator 120 includes a plurality of Cores (Core processors) 121, an Rx queue (queue) 122 that holds data in a first-in-first-out list structure, and a Tx queue 133.

[0051] A part of the processing of the data processing APL1 is offloaded to the accelerator 120, thereby achieving performance and power efficiency that cannot be achieved by software (CPU processing) alone. It is assumed that the accelerator 120 described above will be applied to a large-scale server cluster such as a data center that implements NFV (Network Functions Virtualization) or SDN (Software Defined Network).

[0052] The NIC 130 is NIC hardware that realizes a NW interface, and includes an Rx queue 131 and a Tx queue 132 that hold data in a first-in, first-out list structure. The NIC 130 is connected to an opposing device 170 via, for example, a communication network, and transmits and receives packets. The NIC 130 may be, for example, a SmartNIC, which is a NIC with an accelerator. The SmartNIC is a NIC that can offload load-intensive processing, such as IP packet processing that reduces processing power, to reduce the load on the CPU.

[0053] DPDK150 is a framework for controlling NIC in user space160, and specifically consists of high-speed data transfer middleware. DPDK150 has PMD (Poll Mode Driver) 151, a polling-based reception mechanism (a driver that can select polling mode or interrupt mode for data arrival). PMD151 has a dedicated thread that continuously checks for data arrival and performs reception processing.

[0054] DPDK150 realizes packet processing function in user space160 where APL runs, and makes it possible to reduce packet transfer delay by immediately reaping packets when they arrive from user space160 using a polling model. In other words, DPDK150 harvests packets by polling (busy polling the queue in the CPU), so there is no waiting and delay is small. [Prior art documents] [Patent documents]

[0055] [Patent Document 1] JP 2015-197874 A [Patent Document 2] JP 2018-32156 A [Non-patent literature]

[0056] [Non-Patent Document 1] New API Intel, [online], [Retrieved July 5, 2021], Internet 〈http: / / lwn.net / 2002 / 0321 / a / napi-howto.php3〉 [Non-Patent Document 2] “Resource Configuration (NIC) ~Introduction to DPDK Part 6~,” NTT Technocross, [online], [Retrieved July 5, 2021], Internet 〈https: / / www.ntt-tx.co.jp / column / dpdk_blog / 190610 / 〉 [Non-Patent Document 3] Kei Fujimoto, Kenichi Matsui, Masayuki Akutsu, “KBP: Kernel Enhancements for Low-Latency Networking without Application Customization in Virtual Server”, IEEE CCNC 2021. Summary of the Invention [Problem to be solved by the invention]

[0057] However, both packet transfer using the interrupt model and the polling model have the following problems. In the interrupt model, packets are transferred by software interrupt processing, which allows the kernel to process packets after receiving an event (hardware interrupt) from HW. Because the interrupt model transfers packets by interrupt (software interrupt) processing, there are issues with conflicts with other interrupts and waiting when the interrupt destination CPU is used by a higher priority process, resulting in large delays in packet transfer. In this case, if the interrupt processing becomes congested, the waiting delays will become even larger. For example, as shown in FIG. 19, packet transfer according to the interrupt model transfers packets by interrupt processing (see symbols a and b in FIG. 19), which causes waiting for the interrupt processing, resulting in large delays in packet transfer.

[0058] Supplement the mechanism that causes delay in the interrupt model. In a general kernel, packet transfer processing is transmitted by software interrupt processing after hardware interrupt processing. When a software interrupt for packet transfer processing occurs, under the following conditions (1) to (3), the software interrupt processing cannot be executed immediately. For this reason, it is mediated by a scheduler such as ksoftirqd (a kernel thread for each CPU, which is executed when the load of software interrupts becomes high), and the interrupt processing is scheduled, resulting in a waiting time on the order of milliseconds. (1) When competing with other hardware interrupt processing (2) When competing with other software interrupt processing (3) When other processes or kernel threads (such as migration threads) with high priority and the interrupt destination CPU are in use Under the above conditions, the software interrupt processing cannot be executed immediately.

[0059] Similarly, for packet processing by New API (NAPI), as shown in the dashed box p in Fig. 22, due to the competition of interrupt processing (softIRQ), a NW delay on the order of milliseconds occurs.

[0060] <Problems of KBP> As described above, KBP can suppress softIRQ and achieve low-latency packet processing by constantly monitoring packet arrival using the polling model within the kernel. However, since the kernel thread that constantly monitors packet arrival monopolizes the CPU core and always uses CPU time, there is a problem of high power consumption. With reference to Figs. 23 and 24, the relationship between the workload and CPU utilization will be described. As shown in FIG. 24, in KBP, the kernel thread exclusively occupies a CPU core to perform busy polling. Even for the intermittent packet reception shown in FIG. 23, in KBP, since the CPU is always used regardless of the arrival of packets, there is a problem of increased power consumption.

[0061] DPDK also has the same problem as the above-mentioned KBP. <Problems of DPDK> In DPDK, the kernel thread exclusively occupies a CPU core to perform polling (busy polling of the queue by the CPU). Therefore, even for the intermittent packet reception shown in FIG. 23, in DPDK, since the CPU is always used 100% regardless of the arrival of packets, there is a problem of increased power consumption.

[0062] In this way, since DPDK realizes the polling model in user space and no softIRQ conflict occurs, and KBP realizes the polling model in the kernel and no softIRQ conflict occurs, low-latency packet transfer is possible. However, both DPDK and KBP have the problem of wasting CPU resources for packet arrival monitoring at all times regardless of the arrival of packets, resulting in increased power consumption.

[0063] In view of such a background, the present invention has been made, and an object of the present invention is to reduce the CPU usage rate and enable power saving while maintaining low latency.

Means for Solving the Problems

[0064] To solve the above-mentioned problems, The OS has a kernel, a ring-structured buffer in a memory space in a server having the OS, and a driver capable of selecting a polling mode or an interrupt mode for data arrival from an interface unit, a server internal data transfer device that performs data transfer control of an interface unit on user space, P using a ring model dataThe data transfer device within a server comprises a data transfer unit that starts a thread that monitors arrivals, and a sleep control management unit that manages data arrival schedule information and distributes the data arrival schedule information to the data transfer unit to perform sleep control of the data transfer unit, wherein the data transfer unit puts the thread to sleep based on the data arrival schedule information distributed from the sleep control management unit, and starts a timer just before data arrives to perform a sleep release that wakes up the thread. Effect of the Invention

[0065] According to the present invention, it is possible to reduce CPU utilization and achieve power saving while maintaining low latency. [Brief description of the drawings]

[0066] [Figure 1] 1 is a schematic configuration diagram of an intra-server data transfer system according to a first embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram illustrating an example of the operation of a polling thread in the server data transfer system according to the first embodiment of the present invention. [Diagram 3] FIG. 2 is a schematic configuration diagram of an intra-server data transfer system of an acquisition example 1 of the intra-server data transfer system according to the first embodiment of the present invention. [Figure 4] FIG. 2 is a schematic configuration diagram of an intra-server data transfer system of an acquisition example 2 of the intra-server data transfer system according to the first embodiment of the present invention. [Diagram 5] FIG. 11 is a schematic configuration diagram of an intra-server data transfer system of an acquisition example 3 of the intra-server data transfer system according to the first embodiment of the present invention. [Figure 6] 10 is a flowchart showing the operation of a sleep control management unit when there is a change in data arrival schedule information of the intra-server data transfer system according to the first embodiment of the present invention. [Figure 7] 5 is a flowchart showing the operation of a sleep control management unit when an expansion / reduction of a data transfer unit occurs in the intra-server data transfer system according to the first embodiment of the present invention. [Figure 8] 5 is a flowchart showing the operation of a sleep control unit of the data transfer unit of the intra-server data transfer system according to the first embodiment of the present invention. [Figure 9] 5 is a flowchart showing the operation of a data arrival monitoring unit of a data transfer unit of the intra-server data transfer system according to the first embodiment of the present invention. [Figure 10] 5 is a flowchart showing a Tx data transfer unit operation of the data transfer unit of the server data transfer system according to the first embodiment of the present invention. [Figure 11] 10 is a flowchart showing an operation of a data transfer unit when there is a difference in data arrival schedules in the intra-server data transfer system according to the first embodiment of the present invention. [Figure 12] 10 is a flowchart showing an operation of a data transfer unit when there is a difference in data arrival schedules in the intra-server data transfer system according to the first embodiment of the present invention. [Figure 13] FIG. 11 is a schematic configuration diagram of an intra-server data transfer system according to a second embodiment of the present invention. [Figure 14] 10 is a flowchart showing the operation of a data arrival monitoring unit of a data transfer unit of an intra-server data transfer system according to a second embodiment of the present invention. [Figure 15] FIG. 11 is a diagram illustrating an example in which an intra-server data transfer system is applied to an interrupt model in a server virtualization environment having a generic Linux kernel and a VM configuration. [Figure 16] FIG. 11 is a diagram illustrating an example in which an intra-server data transfer system is applied to an interrupt model in a server virtualization environment having a container configuration. [Figure 17] FIG. 11 is a schematic configuration diagram of an intra-server data transfer system according to a third embodiment of the present invention. [Figure 18] 2 is a hardware configuration diagram showing an example of a computer that realizes the functions of an intra-server data transfer device of the intra-server data transfer system according to the embodiment of the present invention. FIG. [Figure 19]1 is a diagram for explaining packet forwarding based on an interrupt model in a server virtualization environment having a generic Linux kernel and a VM configuration. [Figure 20] FIG. 1 is a diagram illustrating packet forwarding according to a polling model in the configuration of OvS-DPDK. [Figure 21] This is a schematic diagram of Rx-side packet processing using the New API (NAPI) implemented in the Linux kernel 2.5 / 2.6. [Figure 22] 22 is a diagram for explaining an outline of Rx-side packet processing by New API (NAPI) in the area surrounded by a dashed line in FIG. 21. FIG. [Figure 23] FIG. 1 is a diagram showing an example of video (30 FPS) data transfer. [Figure 24] FIG. 13 is a diagram showing the CPU utilization rate used by a busy poll thread in the KBP described in Non-Patent Document 3. [Diagram 25] FIG. 1 is a diagram illustrating the configuration of a DPDK system that controls HW equipped with an accelerator. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0067] Hereinafter, an intra-server data transfer system and the like in an embodiment for carrying out the present invention (hereinafter, referred to as "the present embodiment") will be described with reference to the drawings. (First embodiment) [Overall configuration] Fig. 1 is a schematic diagram of an internal server data transfer system according to a first embodiment of the present invention. The same components as those in Fig. 25 are denoted by the same reference numerals. As shown in FIG. 1, the server data transfer system 1000 includes HW 110, OS 140, and a server data transfer device 200 that is high-speed data transfer middleware arranged on a user space 160. In the user space 160, a data processing APL1 and a data flow timeslot management scheduler 2 are further arranged. The data processing APL1 is a program executed in the user space 160. The data flow timeslot management scheduler 2 transmits schedule information to the data processing APL1 (see symbol q in FIG. 1). In addition, the data flow timeslot management scheduler 2 transmits data arrival schedule information to a sleep control management unit 210 (described later) (see symbol r in FIG. 1).

[0068] The HW 110 communicates with the data processing APL 1 for sending and receiving data. A data flow in which the data processing APL 1 receives packets from the HW 110 is referred to as Rx side reception, and a data flow in which the data processing APL 1 transmits packets to the HW 110 is referred to as Tx side transmission. The HW 110 includes an accelerator 120 and a NIC 130 (physical NIC) for connecting to a communication network.

[0069] The accelerator 120 is a computation unit hardware such as a GPU, an FPGA, etc. The accelerator 120 includes a plurality of Cores (core processors) 121, an Rx queue 122 that holds data in a first-in, first-out list structure, and a Tx queue 123. A part of the processing of the data processing APL1 is offloaded to the accelerator 120, thereby achieving performance and power efficiency that cannot be achieved by software (CPU processing) alone.

[0070] The NIC 130 is NIC hardware that realizes a NW interface, and includes an Rx queue 131 and a Tx queue 132 that hold data in a first-in, first-out list structure. The NIC 130 is connected to an opposing device 170 via, for example, a communication network, and transmits and receives packets.

[0071] The OS 140 is, for example, Linux (registered trademark). The OS 140 includes a high-resolution timer 141 that performs timer management in more detail than the kernel timer. The high-resolution timer 141 uses, for example, the hrtimer of Linux (registered trademark). In the hrtimer, the time at which a callback occurs can be specified using the unit of ktime_t. The high-resolution timer 141 notifies the sleep control unit 221 of the data transfer unit 220 described later of the data arrival timing at the specified time (see reference symbol u in FIG. 1).

[0072] [Server internal data transfer device 200] The server internal data transfer device 200 is DPDK for controlling the NIC in the user space 160, and specifically consists of data high-speed transfer middleware. The server internal data transfer device 200 includes a sleep control management unit 210 and a data transfer unit 220. The server internal data transfer device 200 has a PMD 151 (a driver capable of selecting data arrival in polling mode or interrupt mode) (see FIG. 25) similar to the DPDK arranged on the user space 160. The PMD 151 is a driver capable of selecting data arrival in polling mode or interrupt mode, and a dedicated thread continuously performs confirmation of data arrival and reception processing.

[0073] <Sleep control management unit 210> The sleep control management unit 210 manages the data arrival schedule and performs sleep control of the data transfer unit 220 in accordance with the data arrival timing. The sleep control management unit 210 collectively performs timing control of Sleep / startup of each data transfer unit 220 (see reference symbol t in FIG. 1).

[0074] The sleep control management unit 210 manages the data arrival schedule information, distributes the data arrival schedule information to the data transfer unit 220, and performs sleep control of the data transfer unit 220. The sleep control management unit 210 includes a data transfer unit management unit 211 , a data arrival schedule management unit 212 , and a data arrival schedule distribution unit 213 .

[0075] The data transfer unit management unit 211 holds information such as the number of data transfer units 220 and process IDs (PIDs) as a list. In response to a request from the data arrival schedule distribution unit 213, the data transfer unit management unit 211 transmits information such as the number of data transfer units 220 and process IDs to the data transfer units 220.

[0076] The data arrival schedule management unit 212 manages the data arrival schedule. The data arrival schedule management unit 212 acquires data arrival schedule information from the data flow time slot management scheduler 2 (see symbol r in FIG. 1). When there is a change in the data arrival schedule information, the data arrival schedule management unit 212 detects the change in the data arrival schedule information by receiving a change notification of the data arrival schedule information from the data flow time slot management scheduler 2. Alternatively, the data arrival schedule management unit 212 detects the change by snooping data including the data arrival schedule information (see Figs. 4 and 5). The data arrival schedule management unit 212 transmits data arrival schedule information to the data arrival schedule distribution unit 213 (see symbol s in FIG. 1).

[0077] The data arrival schedule distribution unit 213 acquires information such as the number of data transfer units 220 and process IDs from the data transfer unit management unit 211 . The data arrival schedule distribution unit 213 distributes data arrival schedule information to each data transfer unit 220 (see symbol t in FIG. 1).

[0078] <Data transfer unit 220> The data transfer unit 220 starts a polling thread that monitors packet arrival using the polling model. Based on the data arrival schedule information distributed from the sleep control management unit 210, the data transfer unit 220 puts the thread to sleep and activates a timer immediately before data arrival to wake up the thread for sleep release. Here, when the data transfer unit 220 receives a packet at an unintended timing with the timer, upon sleep release, it wakes up the corresponding thread by means of a hardware interrupt. Sleep / release will be described later in [Sleep / Release].

[0079] The data transfer unit 220 includes a sleep control unit 221, a data arrival monitoring unit 222, an Rx data transfer unit 223 (packet extraction unit), and a Tx data transfer unit 224. The data arrival monitoring unit 222 and the Rx data transfer unit 223 are Rx-side functional units, and the Tx data transfer unit 224 is a Tx-side functional unit.

[0080] <sleep control unit 221> Based on the data arrival schedule information from the sleep control management unit 210, the sleep control unit 221 performs sleep control to stop data arrival monitoring and go to sleep when there is no data arrival. The sleep control unit 221 holds the data arrival schedule information received from the data arrival schedule distribution unit 213.

[0081] The sleep control unit 221 sets a timer for the data arrival timing for the data arrival monitoring unit 222 (see reference v in Fig. 1). That is, the sleep control unit 221 sets a timer so that the data arrival monitoring unit 222 can start polling immediately before data arrival. Here, the sleep control unit 221 may use hrtimers or the like, which are high-resolution timers 141 held by the Linux kernel, and activate the data arrival monitoring unit 222 at the hardware interrupt trigger when the timer is activated by the hardware clock.

[0082] FIG. 2 is a diagram showing an operation example of the polling thread of the in-server data transfer device 200. The vertical axis represents the CPU usage rate [%] of the CPU core used by the polling thread, and the horizontal axis represents time. Note that FIG. 3 shows an operation example of the polling thread due to packet arrival corresponding to the data transfer example of the video (30 FPS) in which packets are received intermittently shown in FIG. 13. As shown in FIG. 2, the data transfer unit 220 puts the thread (polling thread) to sleep (see reference w in FIG. 3) based on the data arrival schedule information received from the sleep control management unit 210, and when the sleep is released, the sleep is released by a hardware interrupt (hardIRQ) (see reference w in FIG. 3). Note that reference y in FIG. 3 is a fluctuation in the wiring voltage due to congestion use of the core CPU (Core processor), etc.

[0083] <Rx side> The data arrival monitoring unit 222 is activated immediately before the data arrives according to the data arrival schedule information managed by the sleep control unit 221. The data arrival monitoring unit 222 monitors the Rx queues 122 and 131 of the accelerator 120 or the NIC 130 to check for the arrival of data.

[0084] Regardless of whether data has arrived or not, the data arrival monitoring unit 222 monopolizes the CPU core and monitors the arrival of data by polling. Incidentally, if this is made an interrupt model, the delay described in the prior art of FIG. 22 (that is, when the softIRQ competes with other softIRQs, a wait occurs regarding the execution of the softIRQ, and an NW delay on the order of ms due to this wait) occurs. In this embodiment, it is a feature that the polling model sleep control is used on the Rx side.

[0085] When there is data arrival in the Rx queues 122 and 131, the Data Arrival Monitoring Unit 222 deletes the entries of the corresponding queues from the buffer (by referring to the content of the packets stored in the buffer and deleting the corresponding queue entries from the buffer considering the next processing to be performed for the processing of the packet), and transfers them to the Rx Data Transfer Unit 223.

[0086] The Rx Data Transfer Unit 223 transfers the received data to the Data Processing APL1. Similar to the Tx Data Transfer Unit 224, since it operates only when data arrives, it does not waste the CPU.

[0087] <Tx side> The Tx Data Transfer Unit 224 stores the received data in the Tx queues 123 and 132 of the accelerator 120 or the NIC 130. The Tx Data Transfer Unit 224 is activated by inter - process communication when the Data Processing APL1 sends data, and returns to CPU idle when the data transfer is completed. Different from the Data Arrival Monitoring Unit 222, it does not waste the CPU.

[0088] [Sleep / Release] Based on the data arrival schedule information received from the sleep control unit 221, the Data Transfer Unit 220 puts the thread to sleep and wakes it up at the timer trigger. <Normal time> Based on the scheduling information of the data arrival timing (data arrival schedule information), the Data Transfer Unit 220 activates the timer immediately before the data arrival to wake up the data arrival monitoring unit thread of the Data Transfer Unit 220. For example, using the hr_timer function standardly installed in the Linux kernel, when the timer expiration time comes, it activates the hardware interrupt of the timer, and the Data Arrival Monitoring Unit 222 wakes up the thread.

[0089] <Unexpected (when data arrives outside the schedule)> If data arrives outside the scheduled timing, the thread of the data arrival monitor 222 goes to sleep. Also, the normal timer is not activated. For this reason, a hardware interrupt is activated to notify the arrival of a packet when the packet arrives. As described above, under normal circumstances, packets are constantly being monitored in polling mode, so a hardware interrupt is not necessary and the hardware interrupt function is disabled by the driver (PMD). However, when putting the polling thread to sleep, in case data arrives outside of the schedule, the mode is changed so that a hardware interrupt is raised when a packet arrives. By doing so, when a packet arrives, a hardware interrupt is raised, and the data arrival monitor 222 can wake up the thread with this hardware interrupt handler.

[0090] [Example of obtaining data arrival schedule information] An example of acquiring data arrival schedule information in the intra-server data transfer system according to this embodiment will be described. An example of a data flow with a fixed data arrival schedule is signal processing in a Radio Access Network (RAN), where the MAC scheduler of MAC4 (described later) manages the data arrival timing by time division multiplexing.

[0091] Signal processing in vRAN (virtual RAN) and vDU (virtual Distributed Unit) often uses DPDK for high-speed data transfer. By applying the invented method, the data transfer unit (DPDK PMD, etc.) is controlled to sleep according to the data arrival timing managed by the MAC scheduler.

[0092] As methods for obtaining the data arrival timing managed by the MAC scheduler, there are <obtaining data arrival schedule information from the MAC scheduler> (directly obtained from the MAC Scheduler) (see Figure 3), <obtaining data arrival schedule information by snooping FAPI P7> (obtained by snooping the FAPI P7 IF) (see Figure 4), and <obtaining data arrival schedule information by snooping CTI> (obtained by snooping the O-RAN CTI) (see Figure 5). These will be described in order below.

[0093] <obtaining data arrival schedule information from the MAC scheduler> Figure 3 is a schematic configuration diagram of the in-server data transfer system of Acquisition Example 1. Acquisition Example 1 is an example applied to the vDU system. The same components as those in Figure 1 are labeled with the same reference numerals, and the description of overlapping parts is omitted. As shown in Figure 3, in the in-server data transfer system 1000A of Acquisition Example 1, in the user space 160, further, a PHY (High) (Physical) 3, a MAC (Medium Access Control) 4, and an RLC (Radio Link Control) 5 are arranged. As a countermeasure device connected to the NIC 130, an RU (Radio Unit) 171 is connected to the receiving side of the NIC 130, and a vCU 172 is connected to the transmitting side of the NIC 130to.

[0094] The sleep control management unit 210 of the in-server data transfer system 1000A modifies the MAC scheduler of the MAC 4 to obtain data arrival schedule information from the MAC 4 (see reference symbol z in Figure 3). Although an example applied to the vDU system has been described, it may be applied not only to the vDU but also to vRAN systems such as the vCU.

[0095] <obtaining data arrival schedule information by snooping CTI> FIG. 4 is a schematic configuration diagram of the in-server data transfer system of Acquisition Example 2. Acquisition Example 2 is an example applied to the vCU system. The same components as those in FIG. 3 are denoted by the same reference numerals, and the description of overlapping parts is omitted. As shown in FIG. 4, in the in-server data transfer system 1000B of Acquisition Example 2, in the user space 160, further, an FAPI (FAPI P7) 6 is arranged between the PHY (High) 3 and the MAC 4. Note that although the FAPI 6 is drawn inside the in-server data transfer device 200 for the sake of notation, the FAPI 6 is arranged outside the in-server data transfer device 200. The FAPI 6 is an IF (interface) for exchanging data scheduling information and the like for connecting the PHY (High) 3 and the MAC 4 defined in the SCF (Small Cell Forum) (see reference numeral aa in FIG. 4).

[0096] The sleep control management unit 210 of the in-server data transfer system 1000B snoops the FAPI 6 and then acquires the data arrival schedule information (see reference numeral bb in FIG. 4).

[0097] <Snoop <CTI7> and acquire data arrival schedule information> FIG. 5 is a schematic configuration diagram of the in-server data transfer system of Acquisition Example 3. Acquisition Example 3 is an example applied to the vCU system. The same components as those in FIG. 3 are denoted by the same reference numerals, and the description of overlapping parts is omitted. As shown in FIG. 5, in the in-server data transfer system 1000C of Acquisition Example 3, a transmission device 173 is arranged outside the user space 160. The transmission device 173 is a transmission device defined in the O-RAN community. The MAC 4 in the user space 160 and the transmission device 173 are connected via a CTI (Collaborative Transport Interface) 7. The CTI 7 is an IF for exchanging data scheduling information and the like with the transmission device defined in the O-RAN community (see reference numeral cc in FIG. 5).

[0098] The sleep control management unit 210 of the intra-server data transfer system 1000C snoops the CTI7 and then acquires the data arrival schedule information (see symbol dd in FIG. 5).

[0099] The operation of the intra-server data transfer system will now be described. The basic operations of the intra-server data transfer systems 1000 (see FIG. 1), 1000A (see FIG. 3), 1000B (see FIG. 4), and 1000C (see FIG. 5) are the same, so only the intra-server data transfer system 1000 (see FIG. 1) will be described.

[0100] [Operation of the sleep control management unit 210] <If there is a change in the data arrival schedule information> FIG. 6 is a flowchart showing the operation of the sleep control management unit 210 when the data arrival schedule information is changed. Step S10 enclosed by a dashed line in FIG. 6 represents an external factor that causes the sleep control management unit 210 to start operating (hereinafter, in this specification, a dashed line in a flowchart represents an external factor that causes the operation to start). In step S10 [external factor], if there is a change in the data arrival schedule information, the data flow time slot management scheduler 2 (see FIG. 1) notifies the data arrival schedule management unit 212 of the sleep control management unit 210 of the change (see symbol r in FIG. 1). Alternatively, as shown in FIG. 4 and FIG. 5, the data arrival schedule management unit 212 of the sleep control management unit 210 (see FIG. 1) detects the change by snooping data including the data arrival schedule information.

[0101] In step S11, the data arrival schedule management section 212 (see FIG. 1) of the sleep control management section 210 acquires data arrival schedule information from the data flow time slot management scheduler 2 (see FIG. 1).

[0102] In step S12, the data arrival schedule management unit 212 transmits the data arrival schedule information to the data arrival schedule distribution unit 213 (see FIG. 1).

[0103] In step S13, the data arrival schedule distribution unit 213 of the sleep control management unit 210 acquires information such as the number and process ID of the data transfer unit 220 (see FIG. 1) from the data transfer unit management unit 211 (see FIG. 1).

[0104] In step S14, the data arrival schedule distribution unit 213 distributes the data arrival schedule information to each data transfer unit 220 (see FIG. 1) to complete the processing of this flow.

[0105] <When the addition or deletion of the data transfer unit 220 occurs> FIG. 7 is a flowchart showing the operation of the sleep control management unit 210 when the addition or deletion of the data transfer unit 220 occurs. In step S20 [external factor], when the addition or deletion of the data transfer unit 220 (see FIG. 1) occurs, the operation system of this system, maintenance operator, etc. set information such as the number and process ID of the data transfer unit 220 for the data transfer unit management unit 211 (see FIG. 1) of the sleep control management unit 210.

[0106] In step S21, the data transfer unit management unit 211 of the sleep control management unit 210 holds information such as the number and process ID of the data transfer unit 220 as a list.

[0107] In step S22, the data transfer unit management unit 211 transmits information such as the number and process ID of the data transfer unit 220 in response to a request from the data arrival schedule distribution unit 213 to complete the processing of this flow. The operation of the sleep control management unit 210 has been described above. Next, the operation of the data transfer unit 220 will be described.

[0108] [Operation of the data transfer unit 220] <Sleep control> FIG. 8 is a flowchart showing the operation of the sleep control unit 221 of the data transfer unit 220. In step S31, the sleep control unit 221 (see FIG. 1) of the data transfer unit 220 holds the data arrival schedule information received from the data arrival schedule distribution unit 213 (see FIG. 1) of the sleep control management unit 210.

[0109] Here, due to reasons such as not being time-synchronized with the counterpart device 170 (see FIG. 1), there may be a constant difference between the data arrival timing managed by the sleep control management unit 210 (see FIG. 1) and the actual data arrival timing. In this case, the data transfer unit 220 stores the difference from the data arrival timing, and if this difference data is constant, the sleep control management unit 210 may correct it by a certain difference time to handle it accordingly (details will be described later with reference to FIGS. 11 and 12).

[0110] In step S32, the sleep control unit 221 (see FIG. 1) of the data transfer unit 220 sets a timer for the data arrival timing for the data arrival monitoring unit 222 (see FIG. 1). That is, the sleep control unit 221 sets a timer so that the data arrival monitoring unit 222 can start polling immediately before data arrival.

[0111] At this time, a high-resolution timer 141 (see FIG. 1) such as hrtimers (registered trademark) held by the Linux kernel (registered trademark) may be used to activate the data arrival monitoring unit 222 by means of a hardware interrupt opportunity when the timer is triggered by the hardware clock. The operation of the sleep control unit 221 has been described above. Next, the operations of the <Rx side> and <Tx side> of the data transfer unit 220 will be described. The present invention is characterized in that the operations of the <Rx side> and <Tx side> are different.

[0112] <Rx side> FIG. 9 is a flowchart showing the operation of the data arrival monitoring unit 222 of the data transfer unit 220. In step S41, the data arrival monitor 222 (see FIG. 1) of the data transfer unit 220 starts up immediately before data arrives, in accordance with the data arrival schedule information managed by the sleep control unit 221 (see FIG. 1).

[0113] Here, when data is received from the accelerator 120 or the NIC 130 (see FIG. 1) while the data arrival monitor 222 is sleeping, a hardware interrupt may be activated when the data is received, and the data arrival monitor 222 may be activated within this hardware interrupt handler. This method is effective in dealing with the case where data arrives at a timing that deviates from the data arrival schedule managed by the sleep control manager 210.

[0114] In step S42, the data arrival monitor 222 monitors the Rx queues 122, 131 (see FIG. 1) of the accelerator 120 or the NIC 130 to check whether data has arrived. At this time, regardless of whether data has arrived, the CPU core is exclusively used to monitor whether data has arrived by polling. If this is an interrupt model, the delay described in the prior art of FIG. 22 will occur (i.e., when a softIRQ competes with another softIRQ, a wait will occur for the execution of the softIRQ, and a network delay of the order of ms will occur due to this wait). This embodiment is characterized in that the polling model sleep control is used on the Rx side.

[0115] In step S43, the data arrival monitor 222 determines whether or not data has arrived in the Rx queues 122 and 131.

[0116] If data has arrived in the Rx queues 122, 131 (S43: Yes), in step S44 the data arrival monitoring unit 222 prunes the data (queue) stored in the Rx queues 122, 131 (referring to the contents of the packets stored in the buffer, and deleting the corresponding queue entry from the buffer while taking into account the next processing to be performed on the packet), and transfers it to the Rx data transfer unit 223 (see FIG. 1). If there is no data arrival at the Rx queues 122 and 131 (S43: No), the process returns to step S42.

[0117] In step S45, the Rx data transfer unit 223 transfers the received data to the data processing APL1 (see Fig. 1). Similar to the Tx data transfer unit 224 (see Fig. 1) described later, the Rx data transfer unit 223 operates only when data arrives, so it does not waste the CPU.

[0118] In step S46, when no data arrives even after the elapse of a certain period specified by the operator, the sleep control management unit 210 (see Fig. 1) puts the data arrival monitoring unit 222 (see Fig. 1) to sleep and ends the processing of this flow.

[0119] <Tx side> Fig. 10 is a flowchart showing the operation of the Tx data transfer unit 224 of the data transfer unit 220. In step S50 [external factor], the data processing APL1 (see Fig. 1) transfers data to the data transfer unit 220 of the in-server data transfer device 200 (see Fig. 1).

[0120] In step S51, the Tx data transfer unit 224 of the data transfer unit 220 stores the received data in the Tx queues 123 and 132 (see Fig. 1) of the accelerator 120 or NIC 130 (see Fig. 1) and ends the processing of this flow. The Tx data transfer unit 224 is activated by inter-process communication when the data processing APL1 sends data, and returns to CPU idle when the data transfer is completed. Therefore, unlike the data arrival monitoring unit 222 on the <Rx side>, it does not waste the CPU. The operation of the data transfer unit 220 has been described above.

[0121] [Example of handling when there is a difference in the data arrival schedule] Next, a description will be given of a case where there is a certain time difference between the data arrival schedule grasped by the sleep control management unit 210 and the data arrival schedule that actually arrives, which is a supplementary explanation of step S31 in FIG. In this embodiment, a use case is assumed in which the data arrival schedule of the RAN, etc. is predetermined. Data arrivals with non-constant time differences are not accepted by the RAN system (APL side), and are therefore excluded.

[0122] <When the schedule of the data transfer unit 220 is ahead of the actual data arrival: Case 1> FIG. 11 is a flowchart showing the operation of the data transfer unit 220 when there is a difference in the data arrival schedule. In step S61, the data arrival monitoring unit 222 (see FIG. 1) of the data transfer unit 220 monitors the Rx queues 122, 131 (see FIG. 1) of the accelerator 120 or the NIC 130, and records the time difference △ (the symbol representing the difference is written as △) T from the data arrival schedule to the actual arrival of the data in a memory not shown.

[0123] In step S62, when there is a data arrival difference of ΔT several times in succession, the data arrival monitor unit 222 (see FIG. 1) notifies the sleep control unit 221 (see FIG. 1) that the data arrival schedule has advanced by ΔT. The number of consecutive times is set arbitrarily by the system operator.

[0124] In step S63, the sleep control unit 221 (see FIG. 1) of the data transfer unit 220 receives the notification that the data arrival schedule is advanced by ΔT, delays the data arrival schedule by ΔT, and ends the processing of this flow. This makes it possible to correct the schedule when the data arrival schedule is ahead for a certain period of time.

[0125] <When the schedule of the data transfer unit 220 is delayed compared to the actual arrival of data: Case 2> FIG. 12 is a flowchart showing the operation of the data transfer unit 220 when there is a difference in the data arrival schedule. In step S71, the data arrival monitor 222 (see FIG. 1) of the data transfer unit 220 monitors the Rx queues 122, 131 (see FIG. 1) of the accelerator 120 or the NIC 130, and if data has already arrived in the first polling after data arrival monitoring has started, records this in a memory (not shown). A supplementary explanation: The data arrival monitor 222 starts up just before data arrives (see the process of step S32 in FIG. 8). However, even just before, there is a time interval of just before = Δt, and it is expected that empty polling will be performed for several cycles. Therefore, if data has already arrived after polling has started, it can be determined that there is a high possibility that the schedule of the data transfer unit 220 is behind schedule.

[0126] In step S72, if data has already arrived at the start of polling several times in succession, the data arrival monitor 222 notifies the sleep control unit 221 (see FIG. 1) to delay the data arrival schedule by an infinitesimal time ΔS. Since it is not possible to know how much the data arrival schedule is actually out of sync, the schedule is gradually adjusted by repeatedly delaying it by an infinitesimal time of ΔS set by the operator.

[0127] In step S73, the sleep control unit 221 receives a notification that the data arrival schedule should be advanced by ΔS, advances the data arrival schedule by ΔS, and ends the processing of this flow. By repeating this time correction by ΔS, it is possible to correct the schedule when there is a delay in the data arrival schedule for a certain period of time.

[0128] As described above, in the server data transfer system 1000, the server data transfer device 200 is arranged on the user space 160. For this reason, like DPDK, the data transfer unit 220 of the server data transfer device 200 can bypass the kernel and refer to a ring-structured buffer (a ring-structured buffer created in a memory space managed by DPDK by DMA (Direct Memory Access) when a packet arrives at the accelerator 120 or NIC 130). That is, the server data transfer device 200 does not use a ring buffer (Ring Buffer 72) (see FIG. 22) or a poll list (Ring Buffer 72) (see FIG. 22) in the kernel. The data transfer unit 220 is able to instantly grasp the arrival of packets by having a polling thread constantly monitor a ring-structured buffer (mbuf; a ring-structured buffer to which PMD 151 copies data using DMA) created in the memory space managed by this DPDK (i.e., it is a polling model rather than an interrupt model).

[0129] In addition to the above-mentioned feature of being placed on the user space 160, the intra-server data transfer device 200 has the following feature regarding the method of waking up the polling thread. That is, for a workload with a fixed data arrival timing, the intra-server data transfer device 200 wakes up a polling thread using a timer based on scheduling information of the data arrival timing (data arrival schedule information). Note that an intra-server data transfer device 200B (see FIG. 17) of a third embodiment described later provides a polling thread in the kernel and wakes up the polling thread when triggered by a hardware interrupt from the NIC 11.

[0130] The operation of the intra-server data transfer device 200 will now be described in more detail. <Normal operation: Polling mode> In the intra-server data transfer device 200, the polling thread of the user space 160 monitors the ring buffer deployed in the memory space from the accelerator 120 or the NIC 130 (see FIG. 1). Specifically, the PMD 151 (see FIG. 25) of the intra-server data transfer device 200 is a driver capable of selecting the polling mode or the interrupt mode for data arrival, and when data arrives at the accelerator 120 or the NIC 130, the PMD 151 copies the data to the ring-structured buffer mbuf by DMA since a ring-structured buffer called mbuf exists in the memory space. The polling thread of the user space 160 monitors this ring-structured buffer mbuf. For this reason, the intra-server data transfer device 200 does not use the poll_list prepared by the kernel. The normal operation (polling mode) has been explained above. Next, the unexpected interrupt mode operation will be described.

[0131] <Unexpected operation: interrupt mode> The intra-server data transfer device 200 changes the mode of the driver (PMD 151) so that a hardware interrupt (hardIRQ) can be raised from the accelerator 120 or NIC 130 (see FIG. 1) if data arrives while the polling thread is sleeping, and when data arrives at the accelerator 120 or NIC 130, a hardware interrupt is triggered to wake up the polling thread. In this way, the driver (PMD 151) of the intra-server data transfer device 200 has two modes: polling mode and interrupt mode.

[0132] Second embodiment Fig. 13 is a schematic diagram of an internal server data transfer system according to a second embodiment of the present invention. The same components as those in Fig. 1 are given the same reference numerals and the description of the overlapping parts will be omitted. As shown in FIG. 13, the in-server data transfer system 1000D includes an HW 110, an OS 140, and an in-server data transfer device 200A which is data high-speed transfer middleware disposed on a user space 160. Similar to the in-server data transfer device 200 shown in FIG. 1, the in-server data transfer device 200A consists of data high-speed transfer middleware. The in-server data transfer device 200A includes a sleep control management unit 210 and a data transfer unit 220A.

[0133] The data transfer unit 220A further includes a CPU frequency / CPU idle control unit 225 (CPU frequency control unit, CPU idle control unit) in addition to the configuration of the data transfer unit 220 shown in FIG. 13. The CPU frequency / CPU idle control unit 225 performs control to vary the CPU operating frequency and CPU idle settings. Specifically, the CPU frequency / CPU idle control unit 225 of the polling thread (in-server data transfer device 200A) started by the hardware interrupt handler sets the CPU operating frequency of the CPU core used by the polling thread lower compared to normal use.

[0134] Here, the kernel can change the operating frequency of the CPU core by the governor setting, and the CPU frequency / CPU idle control unit 225 can use the governor setting or the like to set the CPU operating frequency lower compared to normal use. However, the CPU idle setting depends on the CPU model. When the CPU core has enabled the CPU idle setting, it is also possible to disable it.

[0135] The operation of the in-server data transfer system 1000D will be described below. <Rx side> FIG. 14 is a flowchart showing the operation of the data arrival monitoring unit 222 of the data transfer unit 220A. Parts that perform the same processing as the flowchart shown in FIG. 9 are assigned the same step numbers, and the description of overlapping parts is omitted. When the data arrival monitor 222 (see FIG. 13) is started up just before data arrives in step S41, the CPU frequency / CPU idle control unit 225 (see FIG. 13) returns the operating frequency of the CPU core used by the data transfer unit 220A to its original frequency (increases the CPU operating frequency of the CPU core) in step S81. The CPU frequency / CPU idle control unit 225 also returns the CPU idle state (depending on the CPU architecture, such as C-State) setting to its original frequency and proceeds to step S42.

[0136] When the sleep control management unit 210 (see FIG. 13) puts the data arrival monitoring unit 222 (see FIG. 13) into sleep in step S46, the CPU frequency / CPU idle control unit 225 sets the operating frequency of the CPU core used by the data transfer unit 220A to a low value in step S82. The CPU frequency / CPU idle control unit 225 also sets the CPU idle state (depending on the CPU architecture, such as C-State) setting, sets the corresponding CPU core to CPU idle setting, and ends the processing of this flow.

[0137] In this way, the data transfer unit 220A of the server data transfer device 200A is equipped with a CPU frequency / CPU idle control unit 225, and by setting the CPU frequency / CPU idle state in conjunction with the sleep control of the data arrival monitoring unit 222, it is possible to achieve further power savings. The process of lowering the CPU frequency setting and the process of entering the sleep state may be executed simultaneously. Also, the process may go to sleep after confirming that the packet transfer process is completed.

[0138] [Example of application] The intra-server data transfer devices 200 and 200A may be any intra-server data transfer device that starts a thread in the kernel for monitoring packet arrival using a polling model, and the OS is not limited. In addition, the intra-server data transfer devices 200 and 200A are not limited to a server virtualization environment. Therefore, the intra-server data transfer systems 1000 to 1000D can be applied to the configurations shown in Figs. 15 and 16.

[0139] <Example of Application to VM Configuration> FIG. 15 is a diagram showing an example in which the in-server data transfer system 1000E is applied to an interrupt model in a server virtualization environment of a general-purpose Linux kernel (registered trademark) and a VM configuration. The same components as those in FIGS. 1, 13, and 19 are denoted by the same reference numerals. As shown in FIG. 15, the in-server data transfer system 1000E includes an HW10, a Host OS20, in-server data transfer devices 200 and 200A which are data high-speed transfer middleware arranged on a user space 160, a virtual switch 184, and a Guest OS70.

[0140] Specifically, the server includes a Host OS20 on which a virtual machine and an external process formed outside the virtual machine can operate, and a Guest OS70 that operates inside the virtual machine. The Host OS20 includes a Kernel91, a Ring Buffer22 (see FIG. 19) managed by the Kernel91 in a memory space in the server including the Host OS20, a poll_list86 (see FIG. 22) for registering information on a network device indicating which device a hardware interrupt (hardIRQ) from the NIC11 belongs to, a vhost-net module 221A (see FIG. 19) which is a kernel thread, a tap device 222A (see FIG. 19) which is a virtual interface created by the Kernel91, and a virtual switch (br) 223A (see FIG. 19).

[0141] On the other hand, the Guest OS70 includes a Kernel181, a Driver73, a Ring Buffer52 (see FIG. 19) managed by the Kernel181 in a memory space in the server including the Guest OS70, and a poll_list86 (see FIG. 22) for registering information on a network device indicating which device a hardware interrupt (hardIRQ) from the NIC11 belongs to.

[0142] In the server intra-data transfer system 1000E, the server intra-data transfer devices 200, 200A are arranged on the user space 160. For this reason, like DPDK, the data transfer unit 220 of the server intra-data transfer devices 200, 200A can bypass the kernel and refer to the ring-structured buffer. In other words, the server intra-data transfer devices 200, 200A do not use the ring buffer (Ring Buffer 72) (see FIG. 22) or the poll list (Ring Buffer 72) (see FIG. 22) in the kernel. The data transfer unit 220 can bypass the kernel and refer to a ring-structured buffer (Ring Buffer 72) (mbuf; a ring-structured buffer to which PMD 151 copies data using DMA), making it possible to instantly grasp the arrival of a packet (i.e., it is a polling model rather than an interrupt model).

[0143] By doing this, in a system with a VM virtual server configuration, in both the Host OS 20 and the Guest OS 70, when data arrives, the kernel is bypassed in polling mode and packets are transferred with low latency, thereby achieving low latency. Also, when no data arrives, data arrival monitoring is stopped and the system goes to sleep, thereby achieving power saving. As a result, by controlling sleep through timer control that takes into account the timing of data arrival, it is possible to achieve both low latency and power saving. Also, packets can be transferred with low latency within the server without modifying the APL.

[0144] <Example of application to container configuration> 16 is a diagram showing an example in which an intra-server data transfer system 1000B is applied to an interrupt model in a server virtualization environment with a container configuration. The same components as those in FIG. 15 are given the same reference numerals. 16, the intra-server data transfer system 1000F includes a Guest OS 180 and a container configuration in which the OS is replaced with a Container 210A. The Container 210A includes a vNIC (virtual NIC) 211A. The intra-server data transfer devices 200 and 200A are arranged on a user space 160.

[0145] In a system with a virtual server configuration such as containers, low latency and power saving can be achieved by controlling sleep with a timer that takes into account the timing of data arrival. In addition, packet transfer can be performed with reduced latency within the server without modifying the APL.

[0146] <Example of application to paired metal configuration (non-virtualized configuration)> This invention can be applied to systems with a non-virtualized configuration, such as a paired metal configuration. In a non-virtualized system, it is possible to achieve both low latency and power saving by controlling sleep using a timer that takes into account the timing of data arrival. In addition, it is possible to transfer packets with reduced latency within the server without modifying the APL.

[0147] <Extension Technology> When the number of traffic flows increases, the present invention works in conjunction with RSS (Receive-Side Scaling), which can process inbound network traffic using multiple CPUs, to increase the number of CPUs assigned to the packet arrival monitoring thread, making it possible to scale out in response to network load.

[0148] <Example of application to a network system with a fixed data arrival schedule> As an example of a network system with a fixed data arrival schedule, such as TAS (Time Aware Shaper) in TSN (Time Sensitive Network), it is also possible to apply this technology to the high-speed packet forwarding processing function unit in a network system in which data arrival timing must be guaranteed. In a network system with a fixed data arrival schedule, it is possible to achieve both low latency and power saving.

[0149] Third embodiment In the first and second embodiments, the intra-server data transfer devices 200, 200A are placed in a user space 160. In the third embodiment, instead of the intra-server data transfer devices 200, 200A placed in the user space 160, an intra-server data transfer device 200B that provides a polling thread in the kernel and performs sleep control is provided in the kernel.

[0150] Fig. 17 is a schematic diagram of an intra-server data transfer system according to a third embodiment of the present invention. The same components as those in Figs. 1, 13, and 21 are given the same reference numerals, and explanations of overlapping parts are omitted. This embodiment is an example applied to packet processing using New API (NAPI) implemented in Linux kernel 2.5 / 2.6. Note that when a polling thread is installed inside the kernel, if it is based on NAPI, it is necessary to take the kernel version into consideration.

[0151] As shown in Fig. 17, the intra-server data transfer system 1000G includes HW 10, OS 70, and intra-server data transfer device 200B arranged in kernel 71 of OS 70. More specifically, the data transfer unit 220 of the intra-server data transfer device 200B exists only inside the kernel 71, and the sleep control management unit 210 of the intra-server data transfer device 200B may exist either in the user space 160 or inside the kernel 71 (the sleep control management unit 210 may be arranged either in the user space 160 or inside the kernel 71). Fig. 17 shows an example in which the data transfer unit 220 and the sleep control management unit 210 (i.e., the intra-server data transfer device 200B) are arranged inside the kernel 71.

[0152] Here, if a configuration is adopted in which the intra-server data transfer device 200B that performs sleep control is placed inside the kernel 71, the intra-server data transfer devices 200, 200A placed on the space 160 become unnecessary (in this case, in consideration of general-purpose operation, the intra-server data transfer devices 200, 200A are placed in the intra-server data transfer system, and the intra-server data transfer devices 200, 200A may be adaptively not used). The reason why the intra-server data transfer devices 200, 200A become unnecessary will be explained. That is, when DPDK is not used, software interrupts that cause delays occur only inside the kernel 71, and when DPDK is not used, data is transferred to the data processing APL1 without interrupts using the socket 75. For this reason, even if the intra-server data transfer devices 200, 200A are not placed on the user space 160, data can be transferred to the data processing APL1 at high speed.

[0153] OS70 comprises a Kernel 71, a Ring Buffer 22 (see FIG. 19) managed by the Kernel 71 in the memory space in the server having the OS70, a poll_list 86 (see FIG. 22) that registers information about network devices indicating which device the hardware interrupt (hardIRQ) from the NIC 11 belongs to, a vhost-net module 221A (see FIG. 19) which is a kernel thread, a tap device 222A (see FIG. 19) which is a virtual interface created by the Kernel 91, and a virtual switch (br) 223A (see FIG. 19). As described above, in the intra-server data transfer device 200B, at least the data transfer unit 220 (see FIG. 1) is arranged within the Kernel 71 of the OS 70.

[0154] The data transfer unit 220 of the intra-server data transfer device 200B has a data arrival monitor 222 (see FIG. 1) for monitoring the arrival of data from the interface unit (NIC 11), and when data arrives from the interface unit, the interface unit copies the arriving data into memory space by DMA (Direct Memory Access) without using the CPU, and arranges this data in a ring-configured buffer. The data arrival monitor 222 starts a thread that monitors packet arrivals using a polling model, and detects the arrival of data by monitoring the ring-configured buffer.

[0155] Specifically, the data transfer unit 220 of the intra-server data transfer device 200B has an OS (OS70) which has a kernel (Kernel71), a ring buffer (Ring Buffer72) managed by the kernel in the memory space of the server having the OS, and a poll list (poll_list86) (see Figure 22) which registers information on network devices indicating which device a hardware interrupt (hardIRQ) from the interface unit (NIC11) belongs to, and launches a thread within the kernel which monitors for arriving packets using a polling model.

[0156] In this way, the data transfer unit 220 of the intra-server data transfer device 200B comprises a data arrival monitoring unit 222 that monitors (polls) the poll list, an Rx data transfer unit (packet reaping unit) 223 that, if a packet has arrived, refers to the packet held in the ring buffer and performs reaping to delete the corresponding queue entry from the ring buffer based on the next processing to be performed, and a sleep control unit 221 that puts a thread (polling thread) to sleep if no packet arrives for a specified period of time, and wakes up the thread (polling thread) from sleep mode using a hardware interrupt (hardIRQ) when a packet arrives.

[0157] In this way, the intra-server data transfer device 200B stops the software interrupt (softIRQ) of packet processing, which is the main cause of NW delays, and the data arrival monitor 222 of the intra-server data transfer device 200B executes a thread that monitors packet arrival, and the Rx data transfer unit (packet harvester) 223 processes the packet by a polling model (without softIRQ) when a packet arrives. Then, the sleep control unit 221 puts the thread (polling thread) to sleep if no packet arrives for a predetermined period, so that the thread (polling thread) sleeps when no packet arrives. The sleep control unit 221 releases the sleep state by a hardware interrupt (hardIRQ) when a packet arrives.

[0158] As described above, the intra-server data transfer system 1000G includes an intra-server data transfer device 200B that provides a polling thread in the kernel, and the data transfer unit 220 of the intra-server data transfer device 200B wakes up the polling thread when triggered by a hardware interrupt from the NIC 11. In particular, the data transfer unit 220 is characterized in that, when the polling thread is provided in the kernel, it wakes up the thread using a timer. This allows the intra-server delay control device 200B to achieve both low latency and power saving by performing sleep management of the polling thread that performs packet transfer processing.

[0159] [Hardware configuration] The intra-server data transfer devices 200, 200A, 200B according to the above-described embodiments are realized by a computer 900 having a configuration as shown in FIG. 18, for example. FIG. 18 is a hardware configuration diagram showing an example of a computer 900 that realizes the functions of the intra-server data transfer devices 200 and 200A. The computer 900 has a CPU 901 , a ROM 902 , a RAM 903 , a HDD 904 , a communication interface (I / F) 906 , an input / output interface (I / F) 905 , and a media interface (I / F) 907 .

[0160] The CPU 901 operates based on a program stored in the ROM 902 or the HDD 904, and controls each part of the intra-server data transfer devices 200, 200A, and 200B shown in Figures 1 and 13. The ROM 902 stores a boot program executed by the CPU 901 when the computer 900 is started up, programs that depend on the hardware of the computer 900, and the like.

[0161] The CPU 901 controls an input device 910 such as a mouse and a keyboard, and an output device 911 such as a display, via an input / output I / F 905. The CPU 901 acquires data from the input device 910 via the input / output I / F 905, and outputs generated data to the output device 911. Note that a GPU (Graphics Processing Unit) or the like may be used as a processor together with the CPU 901.

[0162] The HDD 904 stores programs executed by the CPU 901 and data used by the programs, etc. The communication I / F 906 receives data from other devices via a communication network (e.g., NW (Network) 920) and outputs the data to the CPU 901, and also transmits data generated by the CPU 901 to other devices via the communication network.

[0163] The media I / F 907 reads a program or data stored in the recording medium 912 and outputs it to the CPU 901 via the RAM 903. The CPU 901 loads a program related to a target process from the recording medium 912 onto the RAM 903 via the media I / F 907, and executes the loaded program. The recording medium 912 is an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase change rewritable Disc), a magneto-optical recording medium such as an MO (Magneto Optical disk), a magnetic recording medium, a conductive memory tape medium, a semiconductor memory, or the like.

[0164] For example, when the computer 900 functions as the intra-server data transfer devices 200, 200A, 200B configured as one device according to this embodiment, the CPU 901 of the computer 900 executes a program loaded onto the RAM 903 to realize the function of the intra-server data transfer device 100. Furthermore, the HDD 904 stores data in the RAM 903. The CPU 901 reads and executes a program relating to a target process from the recording medium 912. Additionally, the CPU 901 may read a program relating to a target process from another device via a communication network (NW 920).

[0165] [effect] As described above, the intra-server data transfer device 200 performs data transfer control of the interface unit (accelerator 120, NIC 130) in user space, and the OS (OS 70) has a kernel (Kernel 171), a ring buffer (mbuf; a buffer with a ring structure to which the PMD 151 copies data by DMA) in the memory space of the server having the OS, and a driver (PMD 151) capable of selecting a polling mode or an interrupt mode for data arrival from the interface unit (accelerator 120, NIC 130), and a thread (polling The data transfer unit 220 is provided with a data transfer unit 220 that starts up a data arrival schedule (a sleep thread), and a sleep control management unit (sleep control management unit 210) that manages data arrival schedule information and distributes the data arrival schedule information to the data transfer unit 220 to perform sleep control of the data transfer unit 220. The data transfer unit 220 puts the thread to sleep based on the data arrival schedule information distributed from the sleep control management unit 210, and starts a timer just before data arrives to perform sleep release to wake up the thread.

[0166] In this way, the sleep control management unit 210 collectively controls the sleep / wake timing of each data transfer unit 220 in order to control the sleep of multiple data transfer units in accordance with the data arrival timing. When data arrives, the polling mode bypasses the kernel and transfers packets with low latency, thereby reducing latency. When no data arrives, the data arrival monitoring is stopped and the unit goes to sleep, thereby reducing power consumption. As a result, by controlling sleep through timer control that takes into account the data arrival timing, it is possible to achieve both low latency and power consumption.

[0167] The intra-server data transfer device 200 can achieve low latency by implementing data transfer delays within a server using a polling model rather than an interrupt model. That is, in the intra-server data transfer device 200, like DPDK, the data transfer unit 220 arranged in the user space 160 can bypass the kernel and refer to a ring-structured buffer. And, by having a polling thread constantly monitor this ring-structured buffer, it is possible to instantly grasp the arrival of a packet (this is a polling model rather than an interrupt model).

[0168] In addition, for data flows with fixed data arrival timings, such as time division multiplexed data flows, as in signal processing in vRAN, the data transfer unit 220 can be controlled to sleep in consideration of the data arrival schedule, thereby reducing CPU usage while maintaining low latency and achieving power saving. In other words, the problem of wasting CPU resources in the polling model can be solved by controlling sleep through timer control that takes into account the data arrival timing, thereby achieving both low latency and power saving.

[0169] Also, the Guest OS (Guest OS70) operating within the virtual machine has a kernel (Kernel171), a ring buffer (mbuf; a buffer with a ring structure to which PMD151 copies data by DMA) in the memory space of the server having the Guest OS, a driver (PMD151) capable of selecting a polling mode or an interrupt mode for data arrival from the interface unit (accelerator 120, NIC130), and a protocol processing unit 74 that performs protocol processing of the packet for which reaping has been executed, and is characterized in that it has a data transfer unit 220 that starts a thread (polling thread) that monitors packet arrival using a polling model, and a sleep control management unit (sleep control management unit 210) that manages data arrival schedule information and delivers the data arrival schedule information to the data transfer unit 220 to perform sleep control of the data transfer unit 220, and the data transfer unit 220 puts the thread to sleep based on the data arrival schedule information delivered from the sleep control management unit 210, and activates a timer just before data arrival to perform sleep release to wake up the thread.

[0170] In this way, in a system with a virtual server configuration of VMs, for a server having a Guest OS (Guest OS 70), it is possible to reduce CPU utilization while maintaining low latency, thereby achieving power savings.

[0171] A Host OS (Host OS20) on which a virtual machine and an external process formed outside the virtual machine can operate includes a kernel (Kernel91), a ring buffer (mbuf; a buffer with a ring structure to which PMD151 copies data by DMA) in a memory space in a server having the Host OS, a driver (PMD151) capable of selecting a polling mode or an interrupt mode for data arrival from an interface unit (accelerator 120, NIC130), and a tap device 222A which is a virtual interface created by the kernel (Kernel91), and a thread (polling The data transfer unit 220 is characterized by having a data transfer unit 220 that starts up a data arrival schedule (a sleep thread), and a sleep control management unit (sleep control management unit 210) that manages data arrival schedule information and distributes the data arrival schedule information to the data transfer unit 220 to control the sleep of the data transfer unit 220, and the data transfer unit 220 puts the thread to sleep based on the data arrival schedule information distributed from the sleep control management unit 210, and starts a timer just before data arrives to perform a sleep release that wakes up the thread.

[0172] By doing this, in a system with a VM virtual server configuration, for a server equipped with a kernel (Kernel 191) and a host OS (Host OS 20), it is possible to reduce CPU usage while maintaining low latency, thereby achieving power savings.

[0173] In addition, in the intra-server data transfer device 200B, the OS (OS 70) includes a kernel (Kernel 171) and a ring buffer (Ring the data transfer unit 220 includes a data arrival monitor unit 222 that monitors (polling) the poll list, a packet reaping unit (Rx data transfer unit 223) that, when a packet has arrived, refers to the packet held in the ring buffer and executes reaping to delete the corresponding queue entry from the ring buffer based on the next process to be performed, and a poll list (poll_list 86) that registers information on network devices indicating which device a hardware interrupt (hardIRQ) from the interface unit (NIC 11) belongs to, a data transfer unit 220 that starts up a thread (thread) for monitoring packet arrival using a polling model within the kernel, and a sleep control management unit (sleep control management unit 210) that manages a data arrival schedule, manages data arrival schedule information, and distributes the data arrival schedule information to the data transfer unit 220 to perform sleep control of the data transfer unit 220. The data transfer unit 220 includes a data arrival monitor unit 222 that monitors (polling) the poll list, a packet reaping unit (Rx data transfer unit 223) that, when a packet has arrived, refers to the packet held in the ring buffer and executes reaping to delete the corresponding queue entry from the ring buffer based on the next process to be performed, and a poll list (poll_list 86) that registers information on network devices indicating which device a hardware interrupt (hardIRQ) from the interface unit (NIC 11) belongs to, and a sleep control unit (sleep control unit 221) that puts the thread into sleep and performs sleep release by a hardware interrupt (hardIRQ) when the sleep is released.

[0174] In this way, the data transfer device 200B in the server can achieve low latency by implementing data transfer delay in the server using a polling model instead of an interrupt model. In particular, for data flows such as time division multiplexed data flows in which data arrival timing is fixed, such as signal processing in vRAN, the data transfer unit 220 can be controlled to sleep while maintaining low latency, thereby reducing CPU usage and achieving power saving. That is, the problem of wasteful use of CPU resources in the polling model can be solved by controlling sleep through timer control that takes into account data arrival timing, thereby achieving both low latency and power saving.

[0175] The data transfer unit 220 puts a thread (polling thread) to sleep based on the data arrival schedule information received from the sleep control management unit 210, and when waking up the sleep, performs sleep wakeup using a hardware interrupt (hardIRQ). This provides the following effects (1) and (2) in addition to the effects described above.

[0176] (1) Software interrupts (softIRQ) that cause delays when packets arrive are stopped, and a polling model is implemented in the kernel (Kernel 171). In other words, unlike the existing technology NAPI, the intra-server data transfer system 1000G implements a polling model rather than an interrupt model, which is the main cause of network delays. When a packet arrives, it is immediately harvested without waiting, making it possible to implement low-latency packet processing.

[0177] (2) The polling thread in the intra-server data transfer device 200 operates as a kernel thread and monitors packet arrivals in polling mode. The kernel thread (polling thread) that monitors packet arrivals sleeps while no packets have arrived. When no packets have arrived, the CPU is not used due to the sleep state, resulting in power saving effects.

[0178] Then, when a packet arrives, the polling thread that was sleeping is woken up (woke from sleep) by the hardIRQ handler at the time of packet arrival. By waking up from sleep by the hardIRQ handler, it is possible to immediately start the polling thread while avoiding softIRQ contention. A distinctive feature of this method is that sleep is not woken up by a timer, but by the hardIRQ handler. Note that if the traffic load is known in advance, for example, if a 30 ms sleep is known as the workload transfer rate shown in Figure 23, then the hardIRQ handler may be woken up at this timing.

[0179] In this way, the intra-server data transfer device 200B can achieve both low latency and power saving by performing sleep management of the polling thread that performs packet transfer processing.

[0180] The intra-server data transfer device 200A is characterized by comprising a CPU frequency setting section (CPU frequency / CPU idle control section 225) that sets the CPU operating frequency of the CPU core used by the thread to a low value during sleep.

[0181] In this way, the intra-server data transfer device 200A dynamically varies the CPU operating frequency in accordance with the traffic; in other words, if the CPU is not used due to sleep, the CPU operating frequency during sleep can be set low, thereby making it possible to further enhance the power saving effect.

[0182] The intra-server data transfer device 200A is characterized by comprising a CPU idle setting unit (CPU frequency / CPU idle control unit 225) that sets the CPU idle state of the CPU core used by the thread to a power saving mode during sleep.

[0183] By doing this, the intra-server data transfer device 200A can further enhance the power saving effect by dynamically varying the CPU idle state (power saving function according to the CPU model, such as changing the operating voltage) in accordance with the traffic.

[0184] In addition, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by a known method. In addition, the information including the processing procedures, control procedures, specific names, various data and parameters shown in the above documents and drawings can be changed arbitrarily unless otherwise specified. In addition, each component of each device shown in the figure is a functional concept, and does not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads, usage conditions, etc.

[0185] In addition, the above-mentioned configurations, functions, processing units, processing means, etc. may be realized in part or in whole by hardware, for example, by designing them as integrated circuits. In addition, the above-mentioned configurations, functions, etc. may be realized by software for a processor to interpret and execute a program that realizes each function. Information on the programs, tables, files, etc. that realize each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC (Integrated Circuit) card, SD (Secure Digital) card, or optical disk. [Explanation of symbols]

[0186] 1 Data processing APL (application) 2 Data flow time slot management scheduler 3 PHY(High) 4. MAC 5. RLC 6 FAPI (FAPI P7) 20,70 Host OS (OS) 50 Guest OS (OS) 86 poll_list 72 Ring Buffer 91,171,181 Kernel 110HW 120 Accelerator (Interface) 121 cores (Core processor) 122,131 Rx Queue 123,132 Tx Queues 130 NIC (physical NIC) (interface section) 140 OS 151 PMD (driver that allows selection of polling mode or interrupt mode for data arrival) 160 user space 200, 200A, 200B Data transfer device within server 210 sleep control management section 210A Container 211 Data Transfer Management Department 212 Data Arrival Schedule Management Department 213 Data Arrival Schedule Distribution Department 220 Data Transfer Unit 221 sleep control unit 222 Data Arrival Monitoring Unit 223 Rx data transfer unit (packet harvesting unit) 224 Tx Data transfer section 225 CPU frequency / CPU idle control unit (CPU frequency control unit, CPU idle control unit) 1000, 1000A, 1000B, 1000C, 1000D, 1000E, 1000F, 1000G Data transfer system within server Mbuf A ring-structured buffer to which PMD copies data using DMA.

Claims

1. An OS comprising: The kernel, A ring-structured buffer in a memory space in a server having the OS; a driver capable of selecting a polling mode or an interrupt mode for data arrival from an interface unit, and performing data transfer control of the interface unit in a user space, a data transfer unit that launches a thread that monitors for data arrival using a polling model; a sleep control management unit that manages data arrival schedule information and delivers the data arrival schedule information to the data transfer unit to perform sleep control of the data transfer unit; The data transfer unit is Based on the data arrival schedule information distributed from the sleep control management unit, the thread is put to sleep, and a timer is started just before the data arrives to perform a sleep release to wake up the thread.

2. A data transfer device in a server comprising:

2. A guest OS running in a virtual machine, The kernel, A ring-structured buffer in a memory space in a server having the Guest OS; a driver capable of selecting a polling mode or an interrupt mode for data arrival from an interface unit, and performing data transfer control of the interface unit in a user space, a data transfer unit that launches a thread that monitors for data arrival using a polling model; a sleep control management unit that manages data arrival schedule information and delivers the data arrival schedule information to the data transfer unit to perform sleep control of the data transfer unit; The data transfer unit is Based on the data arrival schedule information distributed from the sleep control management unit, the thread is put to sleep, and a timer is started just before the data arrives to perform a sleep release to wake up the thread.

2. A data transfer device in a server comprising:

3. A Host OS in which a virtual machine and an external process formed outside the virtual machine can operate, The kernel, A ring buffer in a memory space in a server having the Host OS; A driver that can select polling mode or interrupt mode for data arrival from the interface unit; a tap device which is a virtual interface created by the kernel, and which controls data transfer of the interface unit in a user space, a data transfer unit that launches a thread that monitors for data arrival using a polling model; a sleep control management unit that manages data arrival schedule information and delivers the data arrival schedule information to the data transfer unit to perform sleep control of the data transfer unit; The data transfer unit is Based on the data arrival schedule information distributed from the sleep control management unit, the thread is put to sleep, and a timer is started just before the data arrives to perform a sleep release to wake up the thread.

2. A data transfer device in a server comprising:

4. An OS comprising: The kernel, a poll list for registering information on a network device indicating which device a hardware interrupt caused by data arrival from an interface unit is from, and the data transfer control of the interface unit is performed within a kernel, a data transfer unit that launches a thread in the kernel to monitor data arrival using a polling model; a sleep control management unit that manages data arrival schedule information and delivers the data arrival schedule information to the data transfer unit to perform sleep control of the data transfer unit; The data transfer unit is A data arrival monitor for monitoring the poll list; a data reaping unit that, when data has arrived, refers to the data stored in the ring buffer and executes reaping to delete the corresponding queue entry from the ring buffer based on the next process to be performed; a sleep control unit that puts the thread to sleep based on the data arrival schedule information received from the sleep control management unit, and releases the thread from sleep by a hardware interrupt when the thread is released from sleep.

2. A data transfer device in a server comprising:

5. The data transfer unit is A CPU frequency control unit is provided for setting a lower CPU operating frequency of a CPU core used by the thread during the sleep state.

5. The server data transfer device according to claim 1, wherein the server data transfer device is a data transfer device for transferring data from the server to the server.

6. The data transfer unit is a CPU idle control unit that sets the CPU idle state of the CPU core used by the thread to a power saving mode during the sleep mode; 5. The server data transfer device according to claim 1, wherein the server data transfer device is a data transfer device for transferring data from the server to the server.

7. An OS comprising: The kernel, A ring-structured buffer in a memory space in a server having the OS; A data transfer method for an intra-server data transfer device having a driver capable of selecting a polling mode or an interrupt mode for data arrival, the data transfer control of an interface unit being performed in a user space, the method comprising: a data transfer unit that launches a thread that monitors for data arrival using a polling model; a sleep control management unit that manages data arrival schedule information and delivers the data arrival schedule information to the data transfer unit to perform sleep control of the data transfer unit, The data transfer unit is putting the thread to sleep based on the data arrival schedule information delivered from the sleep control management unit; and performing a wake-up step of waking up the thread by firing a timer just before the data arrives.

23. A method for transferring data within a server, comprising:

8. A program for causing a computer to function as a server-internal data transfer device described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Low-power-consumption adaptive polling

    JP2004199683A

  • Virtual communication path construction system, virtual communication path construction method, and virtual communication path construction program

    JP2015197874A

  • Connection control system of virtual machine and connection control method of virtual machine

    JP2018032156A

  • Techniques for Received Packet Processing and Associated Power Management in Network Devices

    JP2018507457A

  • Variable polling interval based on historical timing results

    US20090089784A1