Computation Processing Offloading System, Client, Server, and Computation Processing Offloading Method

By integrating ACC function/argument data packetization and parsing units within the OS and using packet processing in-line insertion units, the offloading system reduces latency and enhances versatility, addressing the challenges of overhead and dedicated NIC requirements in existing systems.

JP7687399B2Active Publication Date: 2025-06-03NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023532893
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-07-05
Publication Date
2025-06-03
Estimated Expiration
2041-07-05

AI Technical Summary

Technical Problem

Existing offloading systems face challenges in reducing latency due to overhead in the cooperation between the OS protocol stack and ACC function/argument data, and they require dedicated NICs, limiting versatility.

Method used

The proposed computing processing offloading system integrates ACC function/argument data packetization and parsing units within the OS, eliminating the need for multiple protocol processes and dedicated NICs by serializing and deserializing data according to a predetermined protocol, and using packet processing in-line insertion units to exchange data without going through the existing protocol stack.

Benefits of technology

This approach reduces latency by eliminating overhead in data cooperation between the OS protocol stack and ACC function/argument data, and enhances versatility by eliminating the need for dedicated NICs, thereby achieving low latency and high-speed processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007687399000001
    Figure 0007687399000001
  • Figure 0007687399000002
    Figure 0007687399000002
  • Figure 0007687399000003
    Figure 0007687399000003
Patent Text Reader

Abstract

A client (100) includes an L3 / L4-protocol ACC-function and parameter data packetization unit (121) that serializes a function name and a parameter, which are received by an OS (120) from an application, according to a format of a predetermined protocol and packetizes the serialized function name and parameter as a payload; and an L3 / L4-protocol ACC-function and return-value data parsing unit (122) that deserializes packet data received from a server (200) according to the format of the predetermined protocol to obtain a function name and an execution result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an arithmetic processing offloading system, a client, a server, and an arithmetic processing offloading method.

Background Art

[0002] With the development of cloud computing, from a client machine deployed at a user site, to a server at a remote site (such as a data center located near the user) via a network (hereinafter referred to as NW), by offloading some processes with a large amount of computation, it is becoming widespread to simplify the configuration of the client machine (see Non-Patent Document 1).

[0003] FIG. 12 is a diagram for explaining the device configuration of an offloading system via NW. As shown in FIG. 12, an offloading system via NW1 includes a client 10 deployed at a user site, and a server 50 connected to the client 10 via NW1. The client 10 is a terminal driven by a battery or the like and having limited computing power. The client 10 includes a client HW (hardware) 20, an OS (Operating System) 30, and an application (hereinafter, appropriately referred to as APL) 40. The APL 40 has a client application part 41, an ACC utilization IF 42, and middleware 43. The ACC utilization IF 42 is an ACC (Accelerator: computing accelerator device) utilization IF specification composed of OpenCL (Open Computing Language) or the like. The client application part 41 is a program executed in the user space. The offloading system via NW is constructed on the premise of using a specified API (Application Programming Interface) such as OpenCL, and has input / output with these APIs.

[0004] Client 10 does not have a computing accelerator device such as an FPGA (Graphics Processing Unit) / GPU (Field Programmable Gate Array) (hereinafter referred to as ACC). Client 10 has a NIC (Network Interface Card) 21 mounted on client HW20.

[0005] The client application unit 41 is an application that operates on client 10 and complies with a standard API (Application Programming Interface) for ACC access. Since the client application unit 41 that operates on client 10 assumes image processing and the like, it requires low operation latency.

[0006] Server 50 includes a server HW60, an OS70, an APL80, and an accelerator (ACC) 62 mounted on the server HW60. APL80 has offloading middleware 81. Server 50 mounts one or more accelerators 62. Server 50 has a NIC 61 mounted on the server HW60.

[0007] Client 10 and server 50 can communicate with each other via their respective NICs 21, 61 and NW1.

[0008] In the offloading system shown in FIG. 12, it is preferable to satisfy the following requirements 1 to 3. Requirement 1: Do not modify the client application unit 41 (transparency). Requirement 2: The client-side terminal (client 10) does not require special hardware such as a NIC (general-purpose). Requirement 3: Keep the overhead small in ACC operation offloading via NW1 (low latency).

[0009] As an existing technology for transparent accelerator processing offloading via NW, there is "Packetization of function names and arguments of accelerator standard IF functions and remote offloading by NW transfer" (see Non-Patent Document 1).

[0010] FIG. 13 is a diagram for explaining an accelerator standard IF offloading system using the OS protocol stack described in Non-Patent Document 1. In the description of FIG. 13, the same components as those in FIG. 12 are denoted by the same reference numerals. The solid arrows in FIG. 13 indicate the offloading forward path, and the dashed arrows in FIG. 13 indicate the offloading return path. As shown in FIG. 13, the accelerator standard IF offloading system includes a client 10 and a server 50 connected to the client 10 via NW1. The client 10 shown in FIG. 13 includes a client HW20, an OS 30, and an application (hereinafter, appropriately referred to as APL) 40. The OS 30 has an L4 / L3 protocol stack unit 31 and a NIC driver unit 32. The APL 40 has a client application unit 41, an ACC function proxy reception unit 44, an ACC function / return value packetization unit 45, an ACC function / argument data parsing unit 46, and an ACC function proxy response unit 47.

[0011] The server 50 shown in FIG. 13 includes a server HW60, an OS 70, an APL 80, a NIC 61 on the server HW60, and an accelerator 62. The OS 70 has an L4 / L3 protocol stack unit 71 and a NIC driver unit 72. The APL 80 has a function / argument data parsing unit 82, an ACC function proxy execution unit 83, and an ACC function / return value packetization unit 84.

[0012] ·Offloading forward path The client application unit 41 has input / output with a prescribed API such as OpenCL. The ACC function proxy reception unit 44 is implemented as middleware with an IF compatible with the specified API. It has an IF equivalent to the specified API such as OpenCL and receives API calls from the client application unit 41. The ACC function proxy reception unit 44 receives a function name and arguments from the client application unit 41 as inputs (see reference symbol a in FIG. 13). The ACC function proxy reception unit 44 passes the function name and arguments to the function / return value packetization unit 45 as outputs (see reference symbol b in FIG. 13).

[0013] Based on the received function name and arguments, the ACC function / return value packetization unit 45 passes a transmission packet to the L4 / L3 protocol stack unit 31 (see reference symbol c in FIG. 13). The L4 / L3 protocol stack unit 31 conforms the input packet to the L4 / L3 protocol, and the NIC driver unit 32 passes the transmission packet conforming to the L4 / L3 protocol to the NIC 21 (see reference symbol d in FIG. 13). The NIC 21 transmits a packet to the NIC 61 of the server 50 connected via the NW1.

[0014] The NIC driver unit 72 of the server 50 receives a packet from the NIC 61 (see reference symbol e in FIG. 13) and passes it to the L4 / L3 protocol stack unit 71. The L4 / L3 protocol stack unit 71 converts the received packet conforming to the L4 / L3 protocol into packet data that can be processed and passes it to the ACC function / argument data parsing unit 82 (see reference symbol f in FIG. 13). The ACC function / argument data parsing unit 82 deserializes the packet data and passes the function name and execution result to the ACC function proxy execution unit 83 (see reference symbol g in FIG. 13).

[0015] Based on the received function name and execution result, the ACC function proxy execution unit 83 offloads (see reference symbol h in FIG. 13) the accelerator function / argument data to the accelerator (ACC) 62 for execution.

[0016] · Offload return path The accelerator 62 executes the ACC function and passes the function name and function execution result to the ACC function proxy execution unit 83 (see reference symbol i in FIG. 13). The ACC function proxy execution unit 83 passes the function name and function execution result from the accelerator 62 to the function / return value packetization unit 84 (see reference symbol j in FIG. 13). The ACC function / return value packetization unit 84 packetizes the passed function name and function execution result and passes it to the L4 / L3 protocol stack unit 71 (see reference symbol k in FIG. 13). The L4 / L3 protocol stack unit 71 conforms the bucket data to the L4 / L3 protocol, and the NIC driver unit 72 passes the packet data conforming to the L4 / L3 protocol to the NIC 61 (see reference symbol l in FIG. 13). The NIC 61 transmits a packet from the NIC 21 of the client 10 connected via the NW1.

[0017] The NIC driver unit 32 of the client 10 receives a packet from the NIC 21 (see reference symbol m in FIG. 13) and passes it to the L4 / L3 protocol stack unit 31. The L4 / L3 protocol stack unit 31 converts the received packet conforming to the L4 / L3 protocol into packet data that can be processed and passes it to the ACC function / argument data parsing unit 46 (see reference symbol n in FIG. 13). The ACC function / argument data parsing unit 46 deserializes the function name and execution result into serial data and passes it to the ACC function execution proxy response unit 47 (see reference symbol o in FIG. 13). The ACC function proxy response unit 47 passes the received serial data to the client application unit 41 as accelerator processing data (see reference symbol p in FIG. 13).

[0018] In the above configuration, both the client 10 and the server 50 use a dedicated NIC (for example, RDMA HCA: Remote Direct Memory Access Host Channel Adapter) having a protocol stack processing function. Both the client 10 and the server 50 bypass the protocol stack of the OS kernel by providing the protocol stack function unit in the NICs 21 and 61.

Prior Art Documents

Non-Patent Documents

[0019]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0020] However, in the offloading system described in Non-Patent Document 1, as shown in FIG. 13, the L4 / L3 protocol stack 31 of the OS 30, the ACC function / argument data parsing unit 46, and the ACC function / return value packetizing unit 45 are independent. Similarly, the L4 / L3 protocol stack 31 of the OS 70, the ACC function / argument data parsing unit 82, and the ACC function / return value packetizing unit 84 are independent. For this reason, there is a problem that it is difficult to reduce the latency because overhead occurs in the cooperation (parsing, packet generation) between the L4 / L3 protocol stack function of the OS and the ACC function / argument data. In addition, as shown in FIG. 13, the client 10 requires a dedicated NIC (RDMA HCA), so there is a problem that it is difficult to achieve versatility.

[0021] In view of such a background, the present invention has been made, and an object of the present invention is to reduce the latency by eliminating the overhead in the cooperation between the “protocol stack” of the OS and the “ACC function / argument data”.

Means for Solving the Problems

[0022] To solve the above problems, the present invention provides a computing processing offloading system comprising a client and a server connected via a network, wherein the client offloads specific processing of an application to an accelerator arranged in the server for computing processing, and the client has an operating system (OS) that In the offload forward path, serializes the function name and arguments input from the application side according to the format of a predetermined protocol make it single data, and communication packet and its packets them as a payload in an accelerator function and argument data packetization unit, and an accelerator function and return value data parsing unit that deserializes the packet data input from the server side according to the format of a predetermined protocol and obtains the function name and execution result. The computing processing offloading system is characterized by having these components. In the offload return path,

Advantages of the Invention

[0023] According to the present invention, it is possible to reduce latency by eliminating the overhead in the cooperation between the "protocol stack" of the OS and the "ACC function and argument data".

Brief Description of the Drawings

[0024]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

[0025] Hereinafter, an arithmetic processing offloading system and the like in an embodiment for carrying out the present invention (hereinafter referred to as "the present embodiment") will be described with reference to the drawings. (Embodiment) [Outline] FIG. 1 is a schematic configuration diagram of an arithmetic processing offloading system according to an embodiment of the present invention. The present embodiment is an example applied to offloading processing using XDP (eXpress Data Path) / eBPF (Berkeley Packet Filter) of Linux (registered trademark). The same reference numerals are given to the same constituent parts as in FIGS. 12 and 13. As shown in FIG. 1, the arithmetic processing offloading system 1000 includes a client 100 and a server 200 connected to the client 100 via NW1. In the arithmetic processing offloading system 1000, the client 100 offloads specific processing of an application to an accelerator 212 arranged in the server 200 for arithmetic processing.

[0026] [Client 100] The client 100 includes a client HW110, an OS120, and an APL130.

[0027] 《Client HW110》 The client HW110 has a NIC111. The NIC111 is NIC hardware that realizes a NW interface. In the <transmission pattern>, the NIC111 receives a "transmission packet" as an input from the packet processing in-line insertion unit 123 via the NIC driver unit 124. In the <transmission pattern>, the NIC111 passes the "transmission packet" as an output to the NIC211 of the server 200 connected via NW1. In the <reception pattern>, the NIC111 receives a "reception packet" as an input from the NIC211 of the server 200 connected via NW1. In the <reception pattern>, the NIC111 passes the "reception packet" as an output to the packet processing in-line insertion unit 123 via the NIC driver unit 124.

[0028] 《OS120》 The OS120 has an L3 / L4 protocol · ACC function · argument data packetization unit (hereinafter referred to as the ACC function · argument data packetization unit) 121, an L3 / L4 protocol · ACC function · return value data parsing unit (hereinafter referred to as the ACC function · return value data parsing unit) 122, a packet processing in-line insertion unit 123, and a NIC driver unit 124.

[0029] Here, first, an overview of the ACC function / argument data packetization unit 121 and the ACC function / return value data parsing unit 122 will be described (detailed descriptions will be provided later).

[0030] (1) The ACC function / argument data packetization unit 121 is characterized by combining the function / return value packetization unit 45 on the APL40 side and the L4 / L3 protocol stack unit 31 of the OS 30 in the prior art accelerator standard IF offload system shown in FIG. 13 into a single dedicated function. The single dedicated function, as described in the existing technology shown in FIG. 5 to be described later, is the integration of each protocol processing function such as the Soft IRQ handler processing L2, L3 protocol processing 303, packet harvesting processing (NAPI) 304, Soft IRQ handler processing / L4 protocol processing 305, and ACC function parsing processing 311 of the OS 30 shown in FIG. 5 into a single function, which is one of the features of the present invention. In the prior art, there were multiple protocol processing functions, but in this embodiment, the ACC function / argument data packetization unit 121 is integrated as a single dedicated function.

[0031] Similarly to the case of the ACC function / argument data packetization unit 121, the ACC function / return value data parsing unit 122 is characterized by combining the function / argument data parsing unit 46 on the APL40 side and the L4 / L3 protocol stack unit 31 of the OS 30 in the prior art accelerator standard IF offload system shown in FIG. 13 into a single dedicated function.

[0032] Thus, in the prior art, there are multiple protocol processes (L2 and L3 protocol processes, packet harvesting process (NAPI), L4 protocol process, ACC function parsing process, etc.), and a process of selecting a protocol stack such as L4 / L3 is required. In contrast, the arithmetic processing offloading system 1000 is characterized in that it eliminates and specializes multiple protocol processes required in the prior art by including an ACC function / argument data packetizing unit 121 and an ACC function / return value data parsing unit 122 having a single dedicated function. In the following description, "eliminating the process of selecting a protocol stack of L4 / L3 with multiple processes" is referred to as "data cooperation (first cooperation)". As a result, on the client 100 side, the arithmetic processing offloading system 1000 can achieve high speed without overhead by reducing the number of selections and copies through data cooperation.

[0033] (2) The ACC function / argument data packetizing unit 121 and the ACC function / return value data parsing unit 122 are characterized by being provided within the OS 120. In the prior art accelerator standard IF offloading system shown in FIG. 13, the function / return value packetizing unit 45 and the function / argument data parsing unit 46 are provided on the APL 40 side.

[0034] As described above, the ACC function / argument data packetizing unit 121 and the ACC function / return value data parsing unit 122 are configured with a single dedicated function and are deployed within the OS 120.

[0035] (3) First, an overview of the packet processing in-line insertion unit 123 will be described (detailed description will be given later). The packet processing in-line insertion unit 123 is characterized by exchanging data with the NIC driver unit 124 that harvests data from the NIC 111 without passing through the existing protocol stack. The packet processing in-line insertion unit 123 refers to "exchanging data without passing through the existing protocol stack" as "data cooperation (second cooperation)".

[0036] <ACC function and argument data packetization unit 121> The ACC function and argument data packetization unit 121 serializes the function name and arguments input from the application side according to the format of a predetermined protocol and packetizes them as a payload.

[0037] The ACC function and argument data packetization unit 121 converts the input function name and arguments into data as a UDP / IP packet and its payload. The ACC function and argument data packetization unit 121 serializes the input function name and multiple arguments according to a default format and converts them into single data.

[0038] Figure 2 is a diagram showing a configuration example of the ACC function and argument data packet 450. The ACC function and argument data packet 450 is formatted with an L2 frame (0 - 14 bytes), an L3 header (~34 bytes), an L4 header (~42 bytes), control bits (~46 bytes), a function ID (~50 bytes), argument 1 (~54 bytes), and argument 2 (~58 bytes). The control bits add control information for the packet. The ACC function and argument data packetization unit 121 has, for example, a function to split into multiple packets when the argument size is large. At this time, control data for notifying the final packet is added to the "control bits" in the last packet after splitting. Note that the packet format shown in Figure 2 may include not only the function name and arguments but also an ID that can uniquely identify the accelerator to be used.

[0039] Returning to Figure 1, the ACC function and argument data packetization unit 121 receives "function name and arguments" as input from the ACC function proxy reception unit 132. The ACC function and argument data packetization unit 121 passes "transmission packet" as output to the packet processing in - line insertion unit 123.

[0040] Here, the L3 / L4 protocol may be other than TCP / IP, such as TCP / IP (Transmission Control Protocol / Internet Protocol), or it may eliminate a part of L3 / L4 and use only L3. The packet format may include not only function names and arguments but also an ID that can uniquely identify the accelerator to be used. Also, when the argument size is large, it may have a function to split into multiple packets. At this time, control data for notifying the final packet shown in FIG. 2 is added to the last packet after splitting.

[0041] <ACC function · return value data parsing unit 122> The ACC function · return value data parsing unit 122 deserializes the packet data input from the server 200 side according to the format of a predetermined protocol and obtains the function name and execution result.

[0042] The ACC function · return value data parsing unit 122 deserializes the input packet data to obtain the "function name · execution result" from the input data and passes it to the ACC function proxy response unit 133.

[0043] FIG. 3 is a diagram showing a configuration example of the ACC function · return value packet 500. The ACC function · return value packet 500 is the format of the data to be parsed by the ACC function · return value data parsing unit 122.

[0044] The ACC function · return value packet 500 is formatted with an L2 frame (0 to 14 bytes), an L3 header (~34 bytes), an L4 header (~42 bytes), control bits (~46 bytes), a function ID (~50 bytes), and a return value (~54 bytes). The control bits add control information for the packet. The ACC function · return value data parsing unit 122 has, for example, a function to split into multiple packets when the argument size is large. At this time, control data for notifying the final packet is added to the "control bits" of the last packet after splitting.

[0045] Returning to FIG. 1, the ACC function / return value data parsing unit 122 receives a "received packet" from the packet processing in-line insertion unit 123 as an input. The ACC function / return value data parsing unit 122 passes a "function name / execution result" to the ACC function proxy response unit 133 as an output.

[0046] The embodiment of the packet format conforms to the ACC function / return value packet 500 in FIG. 3. Also, when the ACC function / argument data packetization unit 121 has a function of splitting into a plurality of packets, the ACC function / return value data parsing unit 122 also has a combining process.

[0047] <Packet processing in-line insertion unit 123> The packet processing in-line insertion unit 123 is a transmission / reception function that exchanges data with a device driver (NIC driver unit 124) without going through an existing protocol stack for the input packet data ("transmission packet"). The packet processing in-line insertion unit 123 corresponds to, for example, a high-speed communication mechanism with a driver such as XDP / eBPF of Linux (registered trademark).

[0048] The packet processing in-line insertion unit 123 exchanges data with the ACC function / argument data packetization unit 121 and the ACC function / return value data parsing unit 122 and the NIC driver unit 124 that extracts data from the NIC 111 without going through a predetermined protocol stack.

[0049] In the <transmission pattern>, the packet processing in-line insertion unit 123 receives a "transmission packet" from the ACC function / argument data packetization unit 121 as an input. In the <transmission pattern>, the packet processing in-line insertion unit 123 passes a "transmission packet" to the NIC driver unit 124 as an output.

[0050] In the <reception pattern>, the packet processing in-line insertion unit 123 accepts a "received packet" as input from the NIC driver unit 124. In the <reception pattern>, the packet processing in-line insertion unit 123 passes the "received packet" as output to the L3 / L4 protocol · ACC function · argument data parsing unit 122.

[0051] <NIC driver unit 124> The NIC driver unit 124 is a device driver that abstracts the interface specific to each NIC type. The NIC driver unit 124 is composed of ordinary off-the-shelf device drivers. In the <transmission pattern>, the NIC driver unit 124 accepts a "transmitted packet" as input from the packet processing in-line insertion unit 123. In the <transmission pattern>, the NIC driver unit 124 passes the "transmitted packet" as output to the NIC 111.

[0052] In the <reception pattern>, the NIC driver unit 124 accepts a "received packet" as input from the NIC 111. In the <reception pattern>, the NIC driver unit 124 passes the "received packet" as output to the packet processing in-line insertion unit 123.

[0053] 《APL130》 The APL 130 has a user application unit 131, an ACC function proxy reception unit 132, and an ACC function proxy response unit 133.

[0054] <User application unit 131> The user application unit 131 is a program executed in the user space. The user application unit 131 is constructed on the premise of using specified APIs such as OpenCL and has input / output with these APIs. The user application unit 131 has a "function name · argument" for the ACC function proxy reception unit 132 as output. The user application unit 131 accepts the function execution result from the ACC function proxy response unit 133 as input. The user application unit 131 may also be configured to have a result output destination such as image drawing on a display as another output destination.

[0055] <ACC function proxy reception unit 132> The ACC function proxy reception unit 132 is implemented as middleware having an interface compatible with the specified API. The ACC function proxy reception unit 132 has an interface equivalent to the specified API such as OpenCL and receives API calls from the user. The ACC function proxy reception unit 132 is prepared as a binary file separate from the regular user application and is implemented in the form of a "dynamic library" in which dynamic link calls are made at runtime. The ACC function proxy reception unit 132 receives "function name · arguments" from the user application unit 131 as input. The ACC function proxy reception unit 132 passes "function name · arguments" to the ACC function · argument data packetization unit 121 as output. The ACC function proxy reception unit 132 may also be in the form of a "static library" that is linked to the user application at program generation time and executed integrally.

[0056] <ACC function proxy response unit 133> The ACC function proxy response unit 133 is implemented as middleware having an interface compatible with the specified API. The ACC function proxy response unit 133 is prepared as a binary file separate from the user application unit 131 and is implemented in the form of a "dynamic library" in which dynamic link calls are made at runtime.

[0057] The ACC function proxy response unit 133 exchanges data with the ACC function · argument data parsing unit 222, the ACC function · return value data packetization unit 221, and the NIC driver unit 124 that extracts data from the NIC 111 without going through a predetermined protocol stack.

[0058] The ACC function proxy response unit 133 receives the "function name · execution result" from the ACC function · return value data parsing unit 122 as input. The ACC function proxy response unit 133 passes the "return value" (response data) to the user application unit 131 as output.

[0059] The ACC function proxy response unit 133 may be in the form of a "static library" that is linked and executed together when the program is generated for the user application.

[0060] [Server 20] Server 200 includes server HW210, OS220, and APL230.

[0061] 《Server HW210》 Server HW210 has a NIC 211 and an accelerator 212.

[0062] <nic211> The NIC211 is NIC hardware that implements an NW interface. In the <transmission pattern>, the NIC211 receives a "transmission packet" as input from the packet processing in-line insertion unit 223. In the <transmission pattern>, the NIC211 passes the "transmission packet" as output to the NIC111 of the client 100 connected via NW1. In the <reception pattern>, the NIC211 receives a "reception packet" as input from the NIC111 of the client 100 connected via NW1. In the <reception pattern>, the NIC211 passes the "reception packet" as output to the packet processing in-line insertion unit 223 via the NIC driver 224.

[0063] <Accelerator 212> The accelerator 212 is computing unit hardware that performs specific operations at high speed based on input from the CPU. The GPU / FPGA connected to the server 200 corresponds to the accelerator 212. In the <transmission pattern>, the accelerator 212 receives "ACC instruction data" as input from the ACC function proxy execution unit 231. In the <transmission pattern>, the accelerator 212 passes the "execution result" as output to the ACC function proxy execution unit 231.

[0064] The accelerator 212 may be an integrated form where the CPU and the accelerator are integrated on one chip, such as a System on Chip (SoC). Note that when the accelerator 212 is not installed, the ACC function proxy execution unit 231 may not exist.

[0065] 《OS220》 The OS 220 includes an L3 / L4 protocol · ACC function · return value data packetization unit (hereinafter referred to as the ACC function · return value data packetization unit) 221, an L3 / L4 protocol · ACC function · argument data parsing unit (hereinafter referred to as the ACC function · argument data parsing unit) 222, a packet processing in-line insertion unit 223, and a NIC driver unit 224.

[0066] The OS 220 of the server 200 also has the following characteristics similar to the OS 120 of the client 100 described above. (1) The ACC function · return value data packetization unit 221 is characterized in that it combines the function · return value packetization unit 84 on the APL80 side and the L4 / L3 protocol stack unit 71 of the OS 70 in the prior art accelerator standard IF offloading system shown in FIG. 13 into a single dedicated function. Similarly, the ACC function · argument data parsing unit 222 is characterized in that it combines the function · argument data parsing unit 82 on the APL80 side and the L4 / L3 protocol stack unit 71 of the OS 70 in the prior art accelerator standard IF offloading system shown in FIG. 13 into a single dedicated function.

[0067] As a result, the operation processing offloading system 1000 can achieve high speed without overhead by reducing the number of selections and copies through data cooperation on the server 200 side.

[0068] (2) The ACC function · return value data packetization unit 221 and the ACC function · argument data parsing unit 222 are provided on the OS 220 side.

[0069] In this way, the ACC function · return value data packetization unit 221 and the ACC function · argument data parsing unit 222 on the server 200 side are configured as a single dedicated function, similar to the ACC function · argument data packetization unit 121 and the ACC function · return value data parsing unit 122 on the client 100 side, and the function deployment is changed on the OS 220 side.

[0070] (3) The packet processing in-line insertion unit 223, similar to the packet processing in-line insertion unit 123 on the client 100 side, is characterized by "data cooperation (second cooperation)" that exchanges data with the NIC driver unit 224 without going through the existing protocol stack.

[0071] <ACC function · return value data packetization unit 221> The ACC function · return value data packetization unit 221 serializes the function name and arguments input from the accelerator 212 according to the format of a predetermined protocol and packetizes them as the payload.

[0072] The ACC function · return value data packetization unit 221 is a function that converts the input function name and function execution result into data as a UDP / IP packet and its payload. The ACC function · return value data packetization unit 221 serializes the input function name and function execution result according to the default format and converts them into single data.

[0073] The ACC function · return value data packetization unit 221 receives "function name and arguments" from the accelerator 212 as input. The ACC function · return value data packetization unit 221 passes "transmission packet" to the packet processing in-line insertion unit 223 as output.

[0074] Similar to the ACC function · argument data packetization unit 121, the L3 / L4 protocol of the ACC function · return value data packetization unit 221 may be TCP / IP, SCTP (Stream Control Transmission Protocol) / IP, etc., other than UDP (User Datagram Protocol) / IP. Also, a configuration that uses only L3 instead of both L3 and L4 may be adopted. Specifically, a configuration that uses IP for L3 and a dedicated protocol defined by the user for L4 and above is conceivable.

[0075] Alternatively, only the ACC function / return value data packetization unit 221 may be integrated with the L4 protocol, and the L3 protocol may use the general protocol stack of the OS. Also, when the argument size is large, a function for dividing into multiple packets may be provided. In this case, control data for notifying the final packet is added to the last divided packet (see Fig. 3).

[0076] <ACC function / argument data parsing unit 222> The ACC function / argument data parsing unit 222 deserializes the packet data input from the client 100 side according to the format of a predetermined protocol, and acquires the function name and a plurality of arguments.

[0077] The ACC function / argument data parsing unit 222 deserializes the input packet data, acquires the function name and a plurality of arguments from the input data, and passes them to the ACC function proxy execution unit 231. The format of the data to be parsed by the ACC function / argument data parsing unit 222 is shown in Fig. 2. The ACC function / argument data parsing unit 222 receives the "received packet" from the packet processing in-line insertion unit 223 as input. The ACC function / argument data parsing unit 222 passes the "function name / argument data" to the ACC function proxy execution unit 231 as output.

[0078] Examples of the packet format of the ACC function / argument data parsing unit 222 follow those of the ACC function / argument data packetization unit 121. Also, when the ACC function / argument data packetization unit 121 has a function for dividing into multiple packets, the ACC function / argument data parsing unit 222 also has a combining process.

[0079] <Packet processing in-line insertion unit 223> The packet processing in-line insertion unit 223 is a transceiver function that communicates with a device driver without going through an existing protocol stack for the input packet data. For the packet processing in-line insertion unit 223, for example, a high-speed communication mechanism with a driver such as XDP / eBPF of Linux (registered trademark) is applicable.

[0080] In the <transmission pattern>, the packet processing in-line insertion unit 223 receives a "transmission packet" as input from the ACC function · return value data packetization unit 221. In the <transmission pattern>, the packet processing in-line insertion unit 223 passes the "transmission packet" as output to the NIC driver unit 224.

[0081] In the <reception pattern>, the packet processing in-line insertion unit 223 receives a "reception packet" as input from the NIC driver unit 224. In the <reception pattern>, the packet processing in-line insertion unit 223 passes the "reception packet" as output to the ACC function · argument data parsing unit 222.

[0082] <NIC driver unit 224> The NIC driver unit 224 is a device driver that abstracts the interface specific to each NIC type. The NIC driver unit 224 is composed of ordinary off-the-shelf device drivers. In the <transmission pattern>, the NIC driver unit 224 receives a "transmission packet" as input from the packet processing in-line insertion unit 223. In the <transmission pattern>, the NIC driver unit 224 passes the "transmission packet" as output to the NIC 211.

[0083] In the <reception pattern>, the NIC driver unit 224 receives a "reception packet" as input from the NIC 211. In the <reception pattern>, the NIC driver unit 224 passes the "reception packet" as output to the packet processing in-line insertion unit 223.

[0084] 《APL230》 APL230 has an ACC function proxy execution unit 231. The ACC function proxy execution unit 231 executes the ACC function based on the input function name and arguments, and coordinates the results with the accelerator 212. The ACC function proxy execution unit 231 assumes, for example, an existing accelerator utilization runtime such as the OpenCL runtime or the CUDA runtime.

[0085] In the <execution pattern>, the ACC function proxy execution unit 231 receives the "function name and arguments" from the ACC function / argument data parsing unit 222 as input. In the <execution pattern>, the ACC function proxy execution unit 231 passes the "ACC instruction data" to the accelerator 212 as output. In the <result response pattern>, the ACC function proxy execution unit 231 receives the "execution result" from the accelerator 212 as input. In the <result response pattern>, the ACC function proxy execution unit 231 passes the "function name and function execution result" to the ACC function / return value data packetization unit 221 as output.

[0086] Note that function execution without the accelerator 212 is also possible. Specifically, an RPC server or the like is applicable. In this case, there is no coordination with the accelerator 212, and the result of the calculation performed by the CPU is responded.

[0087] In this way, the operation processing offloading system 1000 of the present embodiment realizes the "parsing function" and the "packet generation function" as single dedicated functions (that is, the ACC function / argument data packetization unit 121, the ACC function / return value data parsing unit 122, the ACC function / return value data packetization unit 221, the ACC function / argument data parsing unit 222) for each of the functions inside the OS. In the existing technology shown in FIG. 13, only general-purpose and common functions are provided in the OS, but in the present embodiment, the above dedicated functions are arranged inside the OSs 120 and 220.

[0088] Specifically, in FIGS. 1 and 4, in the arithmetic processing offloading system 1000, the OS 120 of the client 100 includes an ACC function / argument data packetization unit 121, an ACC function / return value data parsing unit 122, and a packet processing in-line insertion unit 123, and the OS 220 of the server 200 includes an ACC function / return value data packetization unit 221, an ACC function / argument data parsing unit 222, and a packet processing in-line insertion unit 223.

[0089] By providing the dedicated functions (ACC function / argument data packetization unit 121, ACC function / return value data parsing unit 122, ACC function / return value data packetization unit 221, ACC function / argument data parsing unit 222) in the OSs 120 and 220, a configuration is adopted in which there is no overhead (explained by comparison in FIGS. 5 and 6) due to data cooperation between the APL and the OS.

[0090] Furthermore, by causing the dedicated functions (ACC function / argument data packetization unit 121, ACC function / return value data parsing unit 122, ACC function / return value data packetization unit 221, ACC function / argument data parsing unit 222) to cooperate with the NIC driver units 124 and 224 by the packet processing in-line insertion units 123 and 223, a configuration is adopted in which there is no overhead between the NIC driver units 124 and 224 and the dedicated functions (ACC function / argument data packetization unit 121, ACC function / return value data parsing unit 122, ACC function / return value data packetization unit 221, ACC function / argument data parsing unit 222). Since the above is an in-OS software implementation, it can be realized without special function provision to the NICs 111 and 211.

[0091] Hereinafter, the operation of the arithmetic processing offloading system 1000 configured as described above will be described. [Overview of the operation of the arithmetic processing offloading system 1000] FIG. 4 is a diagram for explaining the offloading processing flow of the arithmetic processing offloading system 1000 of FIG. 1. In the description of FIG. 4, the same offloading processing flow as in FIG. 13 is denoted by the same reference numerals. The solid arrows in FIG. 4 indicate the off-road forward path, and the dashed arrows in FIG. 4 indicate the off-road return path.

[0092] ·Off-road forward path As shown in FIG. 4, the ACC function proxy reception unit 132 of the APL 130 of the client 100 receives "function name · arguments" from the user application unit 131 as input (see reference symbol a in FIG. 4). The ACC function proxy reception unit 132 of the client 100 passes "function name · arguments" to the ACC function · argument data packetization unit 121 of the OS 120 as output (see reference symbol b in FIG. 4).

[0093] The ACC function · argument data packetization unit 121 of the OS 120 receives "function name · arguments" from the ACC function proxy reception unit 132 as input (see reference symbol b in FIG. 4). The ACC function · argument data packetization unit 121 converts the input function name · arguments into data as a UDP / IP packet and its payload. The ACC function · argument data packetization unit 121 serializes the input function name · multiple arguments according to a predetermined format and performs single data conversion. The ACC function · argument data packetization unit 121 passes "transmission packet" to the packet processing in-line insertion unit 123 as output (see reference symbol q in FIG. 4).

[0094] In the <transmission pattern>, the packet processing in-line insertion unit 123 of the OS 120 receives "transmission packet" from the ACC function · argument data packetization unit 121 as input. The packet processing in-line insertion unit 123 exchanges data with the device driver without going through the existing protocol stack for the input packet data. In the <transmission pattern>, the packet processing in-line insertion unit 123 passes "transmission packet" to the NIC driver unit 124 as output (see reference symbol q in FIG. 4).

[0095] The NIC driver section 124 of OS 120 receives a "transmission packet" as input from the packet processing in-line insertion section 123 in the <transmission pattern> (see reference q in FIG. 4). The NIC driver section 124 abstracts the interfaces specific to each NIC type. The NIC driver section 124 passes the "transmission packet" to NIC 111 as output in the <transmission pattern> (see reference d in FIG. 4). NIC 111 transmits a packet to NIC 211 of server 200 connected via NW1.

[0096] The NIC driver section 224 of server 200 receives a packet from NIC 211 (see reference e in FIG. 4) and passes it to the packet processing in-line insertion section 223.

[0097] The packet processing in-line insertion section 223 receives a "received packet" as input from the NIC driver section 224 in the <reception pattern>. The packet processing in-line insertion section 223 exchanges data with the device driver without going through the existing protocol stack for the input packet data. The packet processing in-line insertion section 223 passes the "received packet" to the ACC function / argument data parsing section 222 as output in the <reception pattern> (see reference r in FIG. 4).

[0098] The ACC function / argument data parsing section 222 of OS 220 of server 200 receives a "received packet" as input from the packet processing in-line insertion section 223. The ACC function / argument data parsing section 222 deserializes the input packet data to obtain the function name and multiple arguments from the input data. The ACC function / argument data parsing section 222 passes the "function name / argument data" to the ACC function proxy execution section 231 as output (see reference g in FIG. 4).

[0099] In the <execution pattern>, the ACC function proxy execution unit 231 of the APL 230 in the server 200 receives "function name · arguments" from the ACC function - argument data parsing unit 222 as input. The ACC function proxy execution unit 231 executes the ACC function based on the input function name · arguments and coordinates the result with the accelerator 212. In the <execution pattern>, the ACC function proxy execution unit 231 passes "ACC instruction data" to the accelerator 212 as output (see reference h in Fig. 4).

[0100] The accelerator 212 in the server HW 210 of the server 200 receives "ACC instruction data" from the ACC function proxy execution unit 231 as input. The accelerator 212 performs specific operations at high speed based on the input from the CPU.

[0101] · Offload return path In the <transmission pattern>, the accelerator 212 passes the "execution result" to the ACC function proxy execution unit 231 (see reference i in Fig. 4).

[0102] In the <result response pattern>, the ACC function proxy execution unit 231 receives the "execution result" from the accelerator 212 as input (see reference i in Fig. 4). The ACC function proxy execution unit 231 executes the ACC function based on the input function name · arguments and coordinates the result with the accelerator 212. In the <result response pattern>, the ACC function proxy execution unit 231 passes "function name · function execution result" to the ACC function - return value data packetization unit 221 as output (see reference j in Fig. 4).

[0103] The ACC function - return value data packetization unit 221 receives "function name · arguments" from the ACC function proxy execution unit 231 as input (see reference j in Fig. 4). The ACC function - return value data packetization unit 221 converts the input function name · function execution result into data as a UDP / IP packet and its payload. The ACC function - return value data packetization unit 221 passes the "transmission packet" to the packet processing in - line insertion unit 223 as output (see reference s in Fig. 4).

[0104] In the <transmission pattern>, the packet processing in-line insertion unit 223 receives a "transmission packet" as input from the ACC function / return value data packetization unit 221 (see reference symbol s in FIG. 4). The packet processing in-line insertion unit 223 communicates with the device driver without going through the existing protocol stack for the input packet data. In the <transmission pattern>, the packet processing in-line insertion unit 223 passes the "transmission packet" as output to the NIC driver unit 224 (see reference symbol s in FIG. 4).

[0105] In the <transmission pattern>, the NIC driver unit 224 receives a "transmission packet" as input from the packet processing in-line insertion unit 223 (see reference symbol l in FIG. 4). In the <transmission pattern>, the NIC driver unit 224 passes the "transmission packet" as output to the NIC 211 (see reference symbol l in FIG. 4). The NIC 211 transmits a packet to the NIC 111 of the client 100 connected via the NW1.

[0106] The NIC driver unit 124 of the client 100 receives a packet from the NIC 111 and passes it to the packet processing in-line insertion unit 123 (see reference symbol m in FIG. 13).

[0107] In the <reception pattern>, the packet processing in-line insertion unit 123 receives a "reception packet" as input from the NIC driver unit 124 (see reference symbol m in FIG. 4). The packet processing in-line insertion unit 123 communicates with the device driver without going through the existing protocol stack for the input packet data. In the <reception pattern>, the packet processing in-line insertion unit 123 passes the "reception packet" as output to the L3 / L4 protocol / ACC function / argument data parsing unit 122 (see reference symbol t in FIG. 4).

[0108] The ACC function / return value data parsing unit 122 deserializes the input packet data to obtain the function name / execution result from the input data and passes it to the ACC function proxy response unit 133 (see reference symbol o in FIG. 4).

[0109] The ACC function proxy response unit 133 receives the "function name · execution result" from the ACC function · return value data parsing unit 122 as input (see reference o in FIG. 4). The ACC function proxy response unit 133 executes the ACC function proxy response by middleware having an IF compatible with the specified API. The ACC function proxy response unit 133 passes the "return value" to the user application unit 131 as output (see reference p in FIG. 4). The user application unit 131 receives the function execution result from the ACC function proxy response unit 133.

[0110] In the arithmetic processing offloading system 1000 of the present embodiment, for the "parsing function" and "packet generation function" as functions inside the OSs 120 and 220, the "L3 / L4 protocol stack" and "ACC function · argument data" are dedicated functions (ACC function · argument data packetization unit 121, ACC function · return value data parsing unit 122, ACC function · return value data packetization unit 221, ACC function · argument data parsing unit 222). Thus, since the dedicated functions operate as functions inside the OS, there is no overhead due to data cooperation between the APL and the OS (which will be described by comparison in FIGS. 5 and 6 to be described later).

[0111] Furthermore, the dedicated functions are made to cooperate with the NIC driver units 124 and 224 by the packet processing in-line insertion units 123 and 223. Thus, since the dedicated functions cooperate with the NIC driver units 124 and 224 by the packet processing in-line insertion units 123 and 223 (see references q, t, s, and r in FIG. 4), there is no overhead between the NIC driver units 124 and 224 and the dedicated functions.

[0112] Next, the overhead due to data cooperation between the APL and the OS will be described. [Overhead due to Data Cooperation between APL and OS] FIGS. 5 and 6 are diagrams for explaining the overhead due to data cooperation between the APL and the OS. [Middleware Processing of Existing Technology] FIG. 5 is a diagram for explaining an outline of middleware processing in a receiving section of an existing technology client. This middleware processing takes as an example middleware processing using Socket-based remote ACC. As the Socket-based remote ACC use middleware, rCUDA (Remote Compute Unified Device Architecture) or the like can be used.

[0113] <os30> As shown in FIG. 5, in the OS 30, a NIC driver (HIRD handler) 301, which is a handler that is called by the occurrence of a processing request of a NIC 61 (physical NIC) that is a network interface card (see reference symbol u in FIG. 5) and executes the requested processing (hardware interrupt), is arranged.

[0114] In the OS 30, a Soft IRQ handler processing · packet reception processing 302, which is a handler that is called (queuing) by the occurrence of a processing request of the NIC driver (HIRD handler) 301 (see reference symbol v in FIG. 5) and executes the requested packet reception processing (software interrupt), a Soft IRQ handler processing · L2, L3 protocol processing 303, which is a handler that executes L2, L3 protocol processing (software interrupt) in response to the packet reception processing, and a packet harvesting processing (NAPI) 304 that repeatedly performs packet harvesting processing (NAPI) (see reference symbol x in FIG. 5) after the Soft IRQ handler processing · L2, L3 protocol processing 303 are arranged. Incidentally, queue harvesting means referring to the contents of the packets stored in the buffer and deleting the corresponding queue entry from the buffer in consideration of the next processing to be performed for the processing of the packet.

[0115] In the OS 30, a Soft IRQ handler processing · L4 protocol processing 305, which is a handler that executes Soft IRQ handler processing · L4 protocol processing (software interrupt) in response to the Soft IRQ handler processing · L2, L3 protocol processing 303 (see reference symbol w in FIG. 5), a SocketQueue 306 that stores the queue generated by the Soft IRQ handler processing · L4 protocol processing 305, and a socket reception processing 307 that performs socket reception processing based on the queue read from the SocketQueue 306 (see reference symbol y in FIG. 5) are arranged. Also, outside the kernel of the OS 30, a SocketAPI 308 is arranged.

[0116] <apl40> In APL40, there is a Socket library 309 that saves the output of the socket reception process 307 sent via the Socket API 308 (refer to symbol z in FIG. 5), a User space memory 310 for Socket reception that copies the data of the Socket library 309 (refer to symbol aa in FIG. 5) and temporarily stores it, an ACC function parsing process 311 that receives the data stored in the User space memory 310 (refer to symbol bb in FIG. 5) and performs ACC function parsing processing, a User space memory 312 for post-parsing combination that stores the result of the ACC function parsing process 311 (refer to symbol cc in FIG. 5), and an ACC function execution process 313 that executes ACC function processing based on the post-parsed combined data (refer to symbol dd in FIG. 5) from the User space memory 312.

[0117] In the existing technology Socket-based remote ACC utilization middleware process shown in FIG. 5, in the thick solid line blocks in FIG. 5, there are overheads (overheads caused by the process of selecting the L4 / L3 protocol stack with multiple processes and overheads caused by the NIC driver part exchanging data via the existing protocol stack) in the Soft IRQ handler process L2, L3 protocol process 303, Soft IRQ handler process · L4 protocol process 305 of the OS 30, and the User space memory 310 for Socket reception of APL40 and the ACC function parsing process 311 that performs ACC function parsing processing.

[0118] 《Middleware Processing of the Arithmetic Processing Offloading System 1000》 FIG. 6 is a diagram for explaining the outline of the Socket-based remote ACC utilization middleware process constructed based on XDP / eBPF in the arithmetic processing offloading system 1000 of the present embodiment. In the description of FIG. 6, the same symbols are assigned to the same processes as in FIG. 5, and the description of the overlapping parts is omitted.

[0119] <os120> As shown in FIG. 6, in the OS 120, instead of the Soft IRQ handler processing · L2, L3 protocol processing 303 and the Soft IRQ handler processing · L4 protocol processing 305 in FIG. 5, a Soft IRQ handler processing · XDP / eBPF L2 / L3 / L4 protocol processing · ACC function parsing processing 400 is arranged. The Soft IRQ handler processing · XDP / eBPF L2 / L3 / L4 protocol processing · ACC function parsing processing 400 is a single dedicated function that processes the "L3 / L4 protocol stack" and "ACC function · argument data" for each of the "parsing function" and "packet generation function" as an internal function of the OS.

[0120] Also, a packet harvesting process (NAPI) 304 is arranged that repeatedly performs a packet harvesting process (see reference numeral ff in FIG. 5) on the Soft IRQ handler processing · XDP / eBPF L2 / L3 / L4 protocol processing · ACC function parsing processing 400 and passes it to the Soft IRQ handler processing · packet reception processing 302 (see reference numeral ee in FIG. 5).

[0121] In the Soft IRQ handler processing · XDP / eBPF L2 / L3 / L4 protocol processing · ACC function parsing processing 400, the "L3 / L4 protocol stack" and "ACC function · argument data" are processed together in a single manner for each of the "parsing function" and "packet generation function" and passed to the SocketQueue 306 (see reference numeral gg in FIG. 5).

[0122] As a result, within the APL40, the Soft IRQ handler processing, XDP / eBPF L2 / L3 / L4 protocol processing, and ACC function parsing processing 400 are executed collectively and individually. Therefore, the arithmetic processing offloading system 1000 has no overhead due to data communication between the APL and the OS. More specifically, the arithmetic processing offloading system 1000 has no overhead caused by the process of selecting the L4 / L3 protocol stack with multiple processes and the overhead caused by the NIC driver section exchanging data through the existing protocol stack. In addition, since the Soft IRQ handler processing, XDP / eBPF L2 / L3 / L4 protocol processing, and ACC function parsing processing 400 are software implementations within the OS, they can be realized without special function provisioning for the NICs 111, 211 (see FIGS. 1 and 4).

[0123] <apl130> For APL130, the User space memory 310 for Socket reception of APL40 in FIG. 5 and the ACC function parsing process 311 for parsing the ACC function are deleted. The process of copying the data of the Socket library 309 in FIG. 5 to the User space memory 310 for Socket reception and the ACC function parsing process by the ACC function parsing process 311 are omitted, and the data of the Socket library 309 is directly passed to the User space memory 312 for post-parsing combination (see reference hh in FIG. 5). Based on this, the ACC function execution process 313 for executing the ACC function process is arranged.

[0124] As a result, the overhead (overhead caused by data copying) generated between the User space memory 310 for Socket reception of APL40 in FIG. 5 and the ACC function parsing process 311 for parsing the ACC function does not occur in the arithmetic processing offloading system 1000.

[0125] [Offloading Process of Arithmetic Processing Offloading System 1000] Next, the offloading process of the arithmetic processing offloading system 1000 will be described with reference to the control sequence in FIG. 7 and the flowcharts in FIGS. 8 to 10.

[0126] FIG. 7 is a control sequence showing the offloading process of the arithmetic processing offloading system 1000 in FIG. 1. As shown in FIG. 7, the client 100 (see FIGS. 1 and 4) executes the offloading process at the time of transmission (S100; see FIG. 8), and transmits data to the server 200 (see FIGS. 1 and 4) via NW1 (see FIGS. 1 and 4) (S1; see the data transmission sequence).

[0127] The server 200 receives the data from the client 100 transmitted via NW1 and executes the offloading process at the server (S200; see FIG. 9).

[0128] The server 200 transmits the ACC function processing result to the client 100 via the NW1 (refer to S2; data transmission sequence).

[0129] The client 100 executes offloading processing at the time of reception (refer to S300; FIG. 10).

[0130] FIG. 8 is a flowchart showing the offloading processing at the time of transmission of the client 100 of the arithmetic processing offloading system 1000 in FIG. 1 (processing of S100 in FIG. 7). In step S101, the user application unit 131 makes an API call and outputs "function name · arguments". In step S102, the ACC function proxy reception unit 132 receives "function name · arguments" from the user application unit 131 and passes "function name · arguments" to the ACC function · argument data packetization unit 121.

[0131] In step S103, the ACC function · argument data packetization unit 121 serializes the input "function name · multiple arguments" according to a predetermined format, makes it single data, and outputs it as a "transmission packet".

[0132] In step S104, the packet processing in-line insertion unit 123 communicates with the device driver (NIC driver unit 124) without going through the existing protocol stack for the input packet data ("transmission packet").

[0133] In step S105, the NIC driver unit 124 receives the "transmission packet" from the packet processing in-line insertion unit 123, abstracts it to the interface specific to each NIC type, and passes it to the NIC 111. In step S106, the NIC 111 transmits the packet to the NIC 211 of the server 200 connected via the NW1.

[0134] FIG. 9 is a flowchart showing the offloading processing of the server 200 of the arithmetic processing offloading system 1000 in FIG. 1 (processing of S200 in FIG. 7). In step S201, NIC211 receives a packet from NIC111 of client 100 connected via NW1. In step S202, the NIC driver section 224 abstracts the interface specific to each NIC type. In step S203, the packet processing in-line insertion section 223 communicates with the device driver without going through the existing protocol stack for the input packet data.

[0135] In step S204, the ACC function - argument data parsing section 222 deserializes the input packet data to obtain "function name · multiple arguments" from the input data and passes it to the ACC function proxy execution section 231.

[0136] In step S205, the ACC function proxy execution section 231 receives "function name · argument" from the ACC function - argument data parsing section 222, executes the ACC function based on the input "function name · argument", and coordinates with the accelerator 212 for the result. In step S206, the accelerator 212 performs specific operations at high speed based on the input from the CPU.

[0137] In step S207, the ACC function proxy execution section 231 receives the "execution result" from the accelerator 212 and passes "function name · function execution result" to the ACC function - return value data packetization section 221.

[0138] In step S208, the ACC function - return value data packetization section 221 serializes the input function name · function execution result according to the default format, converts it into single - data, and outputs it as a "transmission packet".

[0139] In step S209, the packet processing in-line insertion section 223 communicates with the device driver (NIC driver section 224) without going through the existing protocol stack for the input packet data ("transmission packet").

[0140] In step S210, the NIC driver unit 224 receives a "transmission packet" from the packet processing in-line insertion unit 223, abstracts it to the interface specific to each NIC type, and outputs it to the NIC 211.

[0141] In step S211, the NIC 211 transmits a packet to the NIC 111 of the client 100 connected via the NW1.

[0142] FIG. 10 is a flowchart showing the offloading process (the process of S300 in FIG. 7) at the time of reception of the client 100 of the arithmetic processing offloading system 1000 in FIG. 1. In step S101, the user application unit 131 makes an API call and outputs "function name · arguments". In step S301, the NIC 111 receives a packet from the NIC 211 of the server 200 connected via the NW1.

[0143] In step S302, the NIC driver unit 124 receives a "received packet" from the NIC 111, abstracts it to the interface specific to each NIC type, and passes it to the packet processing in-line insertion unit 123.

[0144] In step S303, the packet processing in-line insertion unit 123 exchanges the input packet data ("received packet") with the device driver (NIC driver unit 124) without going through the existing protocol stack, and passes the "received packet" to the L3 / L4 protocol · ACC function · argument data parsing unit 122.

[0145] In step S304, the ACC function · return value data parsing unit 122 deserializes the input packet data to obtain the function name · execution result from the input data, and passes it to the ACC function proxy response unit 133.

[0146] In step S305, the ACC function proxy response unit 133 receives "function name · execution result" from the ACC function · return value data parsing unit 122 and passes "return value" to the user application unit 131.

[0147] In step S306, the user application unit 131 receives the function execution result from the ACC function proxy response unit 133.

[0148] [Hardware Configuration] The client 100 of the arithmetic processing offloading system 1000 according to this embodiment is realized by a computer 900 having a configuration as shown in FIG. 11, for example. FIG. 11 is a hardware configuration diagram showing an example of the computer 900 that realizes the functions of the client 100. The computer 900 has a CPU 901, a ROM 902, a RAM 903, an HDD 904, a communication interface (I / F) 906, an input / output interface (I / F) 905, and a media interface (I / F) 907.

[0149] The CPU 901 operates based on a program stored in the ROM 902 or the HDD 904 and controls each part of the client 100 shown in FIG. 1. The ROM 902 stores a boot program executed by the CPU 901 when the computer 900 is started up, a program dependent on the hardware of the computer 900, and the like.

[0150] The CPU 901 controls an input device 910 such as a mouse and a keyboard, and an output device 911 such as a display via the input / output I / F 905. The CPU 901 acquires data from the input device 910 via the input / output I / F 905 and outputs the generated data to the output device 911. Note that, together with the CPU 901, a GPU (Graphics Processing Unit) or the like may be used as a processor.

[0151] The HDD 904 stores programs executed by the CPU 901, data used by the programs, and the like. The communication I / F 906 receives data from other devices via a communication network (e.g., NW (Network) 920) and outputs it to the CPU 901, and also transmits data generated by the CPU 901 to other devices via the communication network.

[0152] The media I / F 907 reads a program or data stored in the recording medium 912 and outputs it to the CPU 901 via the RAM 903. The CPU 901 loads a program related to the target process from the recording medium 912 onto the RAM 903 via the media I / F 907 and executes the loaded program. The recording medium 912 is an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto Optical disk), a magnetic recording medium, a conductor memory tape medium, or a semiconductor memory.

[0153] For example, when the computer 900 functions as the client 100 configured as one device according to the present embodiment, the CPU 901 of the computer 900 realizes the functions of the client 100 by executing the program loaded onto the RAM 903. Also, the data in the RAM 903 is stored in the HDD 904. The CPU 901 reads and executes a program related to the target process from the recording medium 912. In addition, the CPU 901 may read a program related to the target process from other devices via the communication network (NW 920).

[0154] As described above, the client 100 of the arithmetic processing offloading system 1000 according to the present embodiment has been described, but the server 200 can also be realized by a computer 900 having a similar configuration.

[0155] [Effect] As described above, the present invention provides an operation processing offloading system including a client 100 and a server 200 connected via a NW1, where the client 100 offloads specific processing of an application to an accelerator 212 disposed in the server 200 for arithmetic processing. The client 100 is characterized in that an OS 120 includes an ACC function / argument data packetization unit 121 that serializes a function name / argument input from the application side according to the format of a predetermined protocol and packetizes it as a payload, and an ACC function / return value data parsing unit 122 that deserializes packet data input from the server 200 side according to the format of a predetermined protocol and obtains a function name / execution result.

[0156] In this way, the client 100 is provided with an ACC function / argument data packetization unit 121 and an ACC function / return value data parsing unit 122 each having a single dedicated function in the OS 120, eliminating and specializing a plurality of protocol processes (L2 and L3 protocol processes, packet harvesting process (NAPI), L4 protocol process, ACC function parsing process, etc.) that were necessary in the prior art. As a result, on the client 100 side, the process of selecting an L4 / L3 protocol stack having a plurality of processes can be eliminated, and the overhead in data cooperation (first cooperation) between the "protocol stack" of the OS and the "ACC function / argument data" can be eliminated to achieve low latency. In the client 100, by reducing the number of selections and copies through data cooperation, overhead can be eliminated and high speed can be achieved. Since the above is a software implementation inside the OS, it can be realized without special function provision to the NIC 111.

[0157] Furthermore, it includes a client 100 and a server 200 connected via NW1. It is an arithmetic processing offloading system in which the client 100 offloads specific processing of an application to an accelerator 212 arranged in the server 200 for arithmetic processing. The server 200 is characterized in that an OS 220 has an ACC function - argument data parsing unit 222 that deserializes packet data input from the client 100 side according to the format of a predetermined protocol and obtains a function name and a plurality of arguments, and an ACC function - return value data packetizing unit 221 that serializes the function name and arguments input from the accelerator 212 according to the format of a predetermined protocol and packetizes them as a payload.

[0158] In this way, the server 200 specializes by eliminating multiple protocol processes that were necessary in the prior art through the OS 220 having the ACC function - argument data parsing unit 222 and the ACC function - return value data packetizing unit 221 each having a single dedicated function. As a result, on the server 200 side, the process of selecting an L4 / L3 protocol stack with multiple processes can be eliminated, and the overhead in data cooperation (the first cooperation) between the "protocol stack" of the OS and the "ACC function - argument data" can be eliminated to achieve low latency. On the server 200 side, by reducing the number of selections and copies through data cooperation, there is no overhead and high speed can be achieved. Since the above is a software implementation inside the OS, it can be realized without special function provision to the NIC 211.

[0159] In the arithmetic processing offloading system 1000, the OS 120 of the client 100 is characterized in that a packet processing in - line insertion unit 123, 223 in which an ACC function - argument data packetizing unit 121, an ACC function - return value data parsing unit 122, and a NIC driver unit 124 that extracts data from the NIC 111 exchange data without passing through a predetermined protocol stack.

[0160] By doing so, by implementing the packet processing in-line insertion unit 123 as software inside the OS 120, it is possible to enhance the versatility without special function provision to the NIC and eliminating the need for a dedicated NIC. Also, the packet processing in-line insertion unit 123 exchanges data with the NIC driver unit 124 that extracts data from the NIC 111 without going through the existing protocol stack, and through data cooperation (second cooperation), there is no overhead between the ACC function / argument data packetization unit 121 and the ACC function / return value data parsing unit 122 and the packet processing in-line insertion unit 123, enabling speed improvement. Since the above is an implementation of software inside the OS, it can be realized without special function provision to the NIC 111.

[0161] In the arithmetic processing offloading system 1000, the OS 220 of the server 200 is characterized in that it has a packet processing in-line insertion unit 223 in which the ACC function / argument data parsing unit 222 and the ACC function / return value data packetization unit 221 exchange data with the NIC driver unit 224 that extracts data from the NIC 211 without going through a predetermined protocol stack.

[0162] By doing so, by implementing the packet processing in-line insertion unit 223 as software inside the OS 220, it is possible to enhance the versatility without special function provision to the NIC and eliminating the need for a dedicated NIC. Also, the packet processing in-line insertion unit 223 exchanges data with the NIC driver unit 224 that extracts data from the NIC 211 without going through the existing protocol stack, and through data cooperation (second cooperation), there is no overhead between the ACC function / argument data parsing unit 222 and the ACC function / return value data packetization unit 221 and the packet processing in-line insertion unit 123, enabling speed improvement. Since the above is an implementation of software inside the OS, it can be realized without special function provision to the NIC 211.

[0163] [Modification Example] (1) This embodiment is an example in which the present invention is applied to both the client 100 and the server 200. As a result, there is no overhead in both the client 100 and the server 200, and the speed can be increased. However, the present invention may be applied to either the client 100 or the server 200. For example, the client 100 may adopt the configuration shown in FIG. 1, and the server 200 may adopt the configuration shown in FIG. 13. Alternatively, the server 200 may adopt the configuration shown in FIG. 1, and the client 100 may adopt the configuration shown in FIG. 13. By leaving the prior art on the system, a way to be generally applied to existing systems can be left.

[0164] (2) In this embodiment, the packet processing in-line insertion units 123 and 223 are arranged in both the client 100 and the server 200, but they may be arranged in either one of them. Alternatively, a mode in which the packet processing in-line insertion units 123 and 223 are not arranged may be adopted. Although there is no synergistic effect with the present invention, there is an advantage that the system configuration can be simplified.

[0165] Among the respective processes described in the above embodiment, all or part of the processes described as being automatically performed can also be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, regarding the processing procedures, control procedures, specific names, information including various data and parameters shown in the above documents and drawings, they can be arbitrarily changed unless otherwise specified. In addition, each component of each device shown in the drawings is a functional concept, and it is not necessarily physically configured as shown in the drawings. That is, the specific form of the distribution and integration of each device is not limited to that shown in the drawings, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads and usage situations.

[0166] In addition, each of the above-described configurations, functions, processing units, processing means, etc. may be implemented in hardware by designing a part or all of them, for example, by using an integrated circuit. Further, each of the above-described configurations, functions, etc. may be implemented by software for a processor to interpret and execute a program for realizing each function. Information such as a program, a table, and a file for realizing each function can be held in a memory, a recording device such as a hard disk or an SSD (Solid State Drive), or a recording medium such as an IC (Integrated Circuit) card, an SD (Secure Digital) card, or an optical disk.

Explanation of Signs

[0167] 1 Network (NW) 100 Client 110 Client HW 111,211 NIC 120 Client OS 121 L3 / L4 Protocol · ACC Function · Argument Data Packetization Unit (Accelerator Function · Argument Data Packetization Unit) 122 L3 / L4 Protocol · ACC Function · Return Value Data Parsing Unit (Accelerator Function · Return Value Data Parsing Unit) 123,223 Packet Processing Inline Insertion Unit 124,224 NIC Driver Unit 130,230 APL 131 User Application Unit 132 ACC Function Proxy Reception Unit 133 ACC Function Proxy Response Unit 200 Server 210 Server HW 212 Accelerator 220 Server OS 221 L3 / L4 Protocol · ACC Function · Return Value Data Packetization Unit (Accelerator Function · Return Value Data Packetization Unit) 222 L3 / L4 Protocol · ACC Function · Argument Data Parsing Unit (Accelerator Function · Argument Data Parsing Unit) 231 ACC Function Proxy Execution Unit 450 ACC Function - Argument Data Packet 500 ACC Function - Return Value Packet 1000 Computation Offloading System

Claims

1. An operation processing offloading system comprising a client and a server connected via a network, wherein the client offloads specific processing of an application to an accelerator arranged in the server for arithmetic processing, wherein the client has an OS that in the offloading forward path, an accelerator function / argument data packetization unit that serializes a function name / arguments input from the application side according to the format of a predetermined protocol into single data and packetizes it as a communication packet and its payload, in the offloading return path, an accelerator function / return value data parsing unit that deserializes packet data input from the server side according to the format of a predetermined protocol and obtains a function name / execution result, characterized in that it is an operation processing offloading system.

2. An operation processing offloading system comprising a client and a server connected via a network, wherein the client offloads specific processing of an application to an accelerator arranged in the server for arithmetic processing, wherein the server has an OS that in the offloading forward path, an accelerator function / argument data parsing unit that deserializes packet data input from the client side according to the format of a predetermined protocol and obtains a function name / plural arguments, in the offloading return path, an accelerator function / return value data packetization unit that serializes a function name / execution result input from the accelerator according to the format of a predetermined protocol and packetizes it as a communication packet and its payload, characterized in that it is an operation processing offloading system.

3. The OS has a packet processing in-line insertion unit in which the accelerator function / argument data packetization unit and the accelerator function / return value data parsing unit exchange data with a NIC driver unit that extracts data from a NIC (Network Interface Card) without passing through a predetermined protocol stack. The operation processing offloading system according to claim 1, characterized in that.

4. The OS The accelerator function / argument data parsing unit and the accelerator function / return value data packetizing unit, and the NIC driver unit that harvests data from the NIC, have a packet processing in-line insertion unit that exchanges data without going through a predetermined protocol stack. The arithmetic processing offloading system according to claim 2, characterized in that.

5. A client that includes a client and a server connected via a network, and the client offloads specific processing of an application to an accelerator arranged in the server for arithmetic processing. The client of the arithmetic processing offloading system is such that The OS is In the offloading forward path, an accelerator function / argument data packetizing unit that serializes the function name / arguments input from the application side according to the format of a predetermined protocol into single data and packetizes it as a communication packet and its payload. In the offloading return path, an accelerator function / return value data parsing unit that deserializes the packet data input from the server side according to the format of a predetermined protocol and obtains the function name / execution result. A client characterized by that.

6. A server that includes a client and a server connected via a network, and the client offloads specific processing of an application to an accelerator arranged in the server for arithmetic processing. The server of the arithmetic processing offloading system is such that The OS is In the offloading forward path, an accelerator function / argument data parsing unit that deserializes the packet data input from the client side according to the format of a predetermined protocol and obtains the function name / plural arguments. In the offloading return path, an accelerator function / return value data packetizing unit that serializes the function name / execution result input from the accelerator according to the format of a predetermined protocol and packetizes it as a communication packet and its payload. A server characterized by that.

7. An arithmetic processing offloading method of an arithmetic processing offloading system that includes a client and a server connected via a network, and the client offloads specific processing of an application to an accelerator arranged in the server for arithmetic processing. The OS of the client is In the offload forward path, the function name and arguments input from the application side are serialized according to the format of a predetermined protocol to form a single data, and are packetized as a communication packet and its payload. In the offload return path, the packet data input from the server side is deserialized according to the format of a predetermined protocol to obtain the function name and execution result, and the following steps are executed. A calculation processing offload method characterized by the above.

8. A calculation processing offload method for a calculation processing offload system including a client and a server connected via a network, where the client offloads specific processing of an application to an accelerator arranged in the server for calculation processing. The OS of the server is as follows: In the offload forward path, the packet data input from the client side is deserialized according to the format of a predetermined protocol to obtain the function name and a plurality of arguments. In the offload return path, the function name and execution result input from the accelerator are serialized according to the format of a predetermined protocol, and are packetized as a communication packet and its payload, and the following steps are executed. A calculation processing offload method characterized by the above.

Citation Information

Patent Citations

  • Topology aware grouping and provisioning of GPU resources in GPU-as-a-Service platform

    US10325343B1

  • Technologies for securely providing remote accelerators hosted on the edge to client compute devices

    US20190228166A1