Information processing device, information processing system, collective communication offloading method and program

By mapping computation process virtual memory to offload function virtual addresses for RDMA, the device enables efficient overlapping of communication and computation in collective communication, addressing inefficiencies in existing systems.

JP2025139396APending Publication Date: 2025-09-26NEC CORP

Patent Information

Application Number
JP2024038312
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies fail to enable overlapping of communication and computation during collective communication between computing processes, particularly when using MPI, due to the need for buffering communication data, which results in inefficient processing.

Method used

An information processing device and system that utilize a coprocessor-host memory conversion mechanism to map computation process virtual memory to an offload function virtual address, enabling remote direct memory access (RDMA) for collective communication offload without buffering on the host, allowing communication and computation to overlap.

Benefits of technology

This approach allows for seamless communication and computation overlap without interrupting the computation process, enhancing processing efficiency in systems with multiple hosts and coprocessors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025139396000001_ABST
    Figure 2025139396000001_ABST
Patent Text Reader

Abstract

To provide an information processing device that enables overlapping of communication and calculation without buffering the communication between calculation processes on a host.SOLUTION: In an information processing device, a coprocessor-host memory translation function maps a calculation process virtual memory for a calculation process in a coprocessor provided in the information processing device to an offloading function virtual address for collective communication offloading. The collective communication offloading function uses the mapped offloading function virtual address to perform remote direct memory access (RDMA) for another information processing device connected via a network, and implements communication between the calculation process and a calculation process of the other information processing device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing system, and a collective communication offload method and program. [Background technology]

[0002] In the field of High Performance Computing (HPC), computing systems consisting of multiple hosts and coprocessors are widely used. Furthermore, the computing processes executed on such computing systems generally communicate with each other using the Message Passing Interface (MPI), a standard for utilizing parallel computing.

[0003] In such computing systems, hosts are connected via a high-speed interconnect, and InfiniBand is often used, which allows the computing process to issue RDMA (Remote Direct Memory Access) directly. "RDMA" refers to the direct transfer of data from the memory of a local computer to the memory of a different remote computer without going through the operating systems of both computers. "InfiniBand" also refers to a high-speed I / O bus architecture and interconnect for servers / clusters.

[0004] In addition, GPUDirect RDMA is known as a technology that enables direct RDMA between coprocessors, which are circuits that assist or act on behalf of CPU processing, and the memory on the coprocessor via InfiniBand. In communication using InfiniBand, the instruction to start communication and the confirmation of communication completion can be performed separately. This allows the computation process to perform processing other than communication between the start and completion of communication, and it is common to overlap computation and communication to shorten the execution time of computation jobs.

[0005] By the way, in communication between computation processes using MPI, in addition to one-to-one communication between specific computation processes, collective communication between multiple computation processes is also used. Collective communication requires multiple communications to complete processing, and because this processing must be performed by the computation process, it is difficult to overlap communication and computation.

[0006] One technology that allows MPI collective communication and computation to overlap is the collective communication offload function using NVIDIA's BlueField DPU. The BlueField DPU implements a CPU that can be used for offload processing on the network interface, and the offload process running on this CPU performs RDMA READ, which reads communication data from the computation process, and RDMA WRITE, which writes data to other computation processes, thereby eliminating the need for communication processing in the computation process itself and achieving offloading. However, this technology requires that communication data be copied to memory on the BlueField DPU, which means that communication that would normally be completed in one go must be performed twice, making it an inefficient process.

[0007] For example, as a technology related to RDMA communication, Patent Document 1 discloses a technology for improving processing capacity by transmitting an RDMA reception response using a different route, rather than using the route used for transmitting the RDMA communication request and transfer data. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-187973 Summary of the Invention [Problem to be solved by the invention]

[0009] Although the technology described in Patent Document 1 contributes to improving the processing performance of RDMA, it does not enable overlapping of communication and computation when performing collective communication between computing processes running on a coprocessor, much less enables overlapping of communication and computation without buffering the communication between computing processes on the host.

[0010] An object of the present disclosure is to provide an information processing device, an information processing system, a collective communication offload method, and a program that solve the above-mentioned problems. [Means for solving the problem]

[0011] An information processing device according to one embodiment of the present disclosure comprises a coprocessor-host memory conversion means for mapping a computation process virtual memory for a computation process in a coprocessor provided in the information processing device to an offload function virtual address for collective communication offload, and a collective communication offload means for performing remote direct memory access (RDMA) to another information processing device connected via a network using the mapped offload function virtual address, thereby implementing communication between the computation process and the computation process of the other information processing device.

[0012] An information processing system according to an aspect of the present disclosure is configured by connecting a plurality of information processing devices according to the above-described aspect of the present disclosure.

[0013] A method for offloading collective communication by an information processing device according to one embodiment of the present disclosure maps a computation process virtual memory for a computation process in a coprocessor provided in the information processing device to an offload function virtual address for collective communication offload, and uses the mapped offload function virtual address to perform remote direct memory access (RDMA) to another information processing device connected via a network, thereby implementing communication between the computation process and the computation process of the other information processing device.

[0014] A program according to one embodiment of the present disclosure causes a computer to map a computing process virtual memory for a computing process in a coprocessor provided in an information processing device to an offload function virtual address for collective communication offload, and to use the mapped offload function virtual address to perform remote direct memory access (RDMA) to another information processing device connected via a network, thereby implementing communication between the computing process and the computing process of the other information processing device. [Effects of the Invention]

[0015] According to the above aspect, communication and computation can be overlapped without buffering communication between computing processes on the host. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 2 is a diagram illustrating functional blocks of a host that is an information processing apparatus according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a diagram illustrating a system configuration in which a plurality of hosts having the configuration shown in FIG. 1 are connected via a network. [Figure 3] 1 is a table illustrating an example of an arrangement of calculation processes executed in one embodiment of the present disclosure. [Figure 4] FIG. 10 is a diagram showing parameters of the MPI_Alltoallv collective communication executed by the computation process 111 executed by the coprocessor 11 of the host 1. [Figure 5] FIG. 10 is a diagram showing parameters of the MPI_Alltoallv collective communication executed by the computation process 211 executed by the coprocessor 21 of the host 2. [Figure 6] 1 is a diagram showing a correspondence table between virtual addresses of computing processes held by the collective communication offload function 17 of the host 1 and corresponding virtual addresses of the collective communication offload function 17. FIG. [Figure 7]10 is a diagram showing a correspondence table between the virtual addresses of the computing processes of the collective communication offload function 27 of the host 2 and the corresponding virtual addresses of the collective communication offload function 27. FIG. [Figure 8] 1 is a diagram showing a correspondence table between the virtual addresses of the collective communication offload function 17, the virtual addresses of the computing process, and the physical addresses of the memory window, which the coprocessor-host memory conversion function 16 of the host 1 has. [Figure 9] 10 is a diagram showing a correspondence table between the virtual addresses of the collective communication offload function 27, the virtual addresses of the computing process, and the physical addresses of the memory window, which are held by the coprocessor-host memory conversion function 26 of the host 2. FIG. [Figure 10] 10 is a diagram showing the contents of a memory translation table 131 for translating an offload function virtual address of the network interface 13 of the host 1 into a physical address. FIG. [Figure 11] 10 is a diagram showing the contents of a memory translation table 231 held by the network interface 23 of the host 2. FIG. [Figure 12] 10 is a flowchart illustrating an operation when each calculation process executes MPI_Alltoallv according to an embodiment of the present disclosure. [Figure 13] FIG. 10 is a diagram illustrating in detail communication processing by the collective communication offload function. [Figure 14] FIG. 1 is a diagram illustrating a configuration example of an information processing device serving as a host according to an embodiment of the present disclosure. [Figure 15] FIG. 2 is a block diagram showing an example of a hardware configuration of an information processing device serving as a host. DETAILED DESCRIPTION OF THE INVENTION

[0017] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In all drawings, the same or corresponding components are designated by the same reference numerals, and common descriptions will be omitted.

[0018] 1 is a diagram showing functional blocks of a host, which is an information processing device according to an embodiment of the present disclosure. As shown in FIG. 1, the host 1 of this embodiment is equipped with a coprocessor 11, a coprocessor 12, and a network interface 13. The coprocessors 11 and 12 and the network interface 13 are connected to an internal bus of the host 1, such as PCI-Express, a high-speed serial interface. The host 1 is a general computing server equipped with a CPU and memory, and an OS is running to enable program execution.

[0019] The coprocessor 11 is equipped with a CPU and memory and is capable of executing programs. The memory of the coprocessor 11 can be referenced as part of the physical address of the host 1 via a memory window 113. In this embodiment, the coprocessor 11 runs a calculation process 111 and has a communication buffer 112 in the memory. The coprocessor 12 also has a similar configuration to the coprocessor 11.

[0020] The network interface 13 connects the host 1 to a network and enables RDMA with another host connected to the network. The network interface 13 includes a memory translation table 131. The memory translation table 131 is composed of an identification number mkey, a virtual address, a physical address, and a size, as shown in FIG. 10 (described later). When a program communicates using the network interface 13, the program inputs the communication direction, communication partner, local virtual address, local mkey, remote virtual address, remote mkey, and communication size as parameters to the network interface 13. The network interface 13 converts the local virtual address to a physical address using the memory translation table 131 and issues DMA (Direct Memory Access) to the physical address. Similarly, the network interface of the host that is the communication partner of the host 1 converts the remote virtual address to a physical address and issues DMA to the physical address.

[0021] The host 1 has a network interface management function 14, a coprocessor management function 15, a coprocessor-to-host memory translation function 16, and a collective communication offload function 17. The network interface management function 14 is a function for making the network interface 13 available to the host 1, and has the function of receiving memory registration requests from programs and manipulating the memory translation table 131. The coprocessor management function 15 performs memory management and task management for the coprocessors 11 and 12. When the network interface management function 14 receives a memory registration request, the coprocessor-to-host memory translation function 16 has the function of requesting the coprocessor management function 15 to operate the memory windows 113 and 123 and returning the physical address if the virtual address is memory on the coprocessors 11 and 12. The collective communication offload function 17 has the function of performing RDMA using the mapped virtual address to perform communication between computation processes on behalf of the computation processes without buffering the communication in memory on the host 1.

[0022] 1 shows an example in which the host 1 includes two coprocessors 11 and 12, but the present invention is not limited to this. For example, the host 1 may include one or more coprocessors.

[0023] A plurality of hosts having a configuration similar to that of host 1 described above are connected via a network as shown in Figure 2 to form a system consisting of multiple hosts. Figure 2 shows an example in which, in addition to host 1 shown in Figure 1, hosts 2-4 having functions equivalent to host 1 are connected to the network. However, this is not limited to this, and it is sufficient if two or more hosts having functions equivalent to host 1 are connected via a network.

[0024] The following explanation will be given taking communication between host 1 and host 2 as an example. In this case, the corresponding functional blocks of host 2 to host 1 will be explained by replacing the second digit "1" of the reference numerals shown in Figure 1 with "2." For example, coprocessor 21 of host 2 is the function of host 2 that corresponds to coprocessor 11 of host 1.

[0025] Next, we will explain the tables used when each calculation process executes MPI_Alltoallv using Figures 3 to 11. "MPI_Alltoallv" is one of the functions used when performing parallel calculations using multiple computers (nodes), and is a function used to realize collective communication offload functionality; it is a type of function that collects data from members of a group and sends data to members of the group.

[0026] FIG. 3 is a table showing an example of the allocation of computational processes executed in one embodiment of the present disclosure. In FIG. 3, "rank number" is a unique value assigned to each process when MPI is used. "process ID" is a value assigned by the OS running on the host to identify the computational process, and is a unique value within the host. "host" is a value indicating which host the computational process is running on. This table is assumed to be shared by each computational process and the collective communication offload function. For example, in FIG. 3, computational process 111 in FIG. 1 indicates that the host running it with rank number "0" is "1."

[0027] 4 shows the parameters of the MPI_Alltoallv collective communication executed by the computation process 111 running on the coprocessor 11 of the host 1. As shown in Fig. 4, the parameters include a send buffer address, a receive buffer address, a rank number, a send offset, a send length, a receive offset, and a receive length. Although not shown, similar parameters are present as parameters of the MPI_Alltoallv collective communication executed by the computation process 121 running on the coprocessor 12 of the host 1.

[0028] Fig. 5 shows the parameters of MPI_Alltoallv collective communication executed by the computational process 211 executed by the coprocessor 21 of the host 2. This is the same as Fig. 4, and the explanation will be omitted. Note that the same parameters are used as the parameters of MPI_Alltoallv collective communication executed by the computational process 221 executed by the coprocessor 22 of the host 2. The parameters of MPI_Alltoallv collective communication are set in the same way for computational processes executed by other hosts.

[0029] FIG. 6 is a correspondence table between the virtual addresses of the computation processes of the collective communication offload function 17 of the host 1 and the corresponding virtual addresses of the collective communication offload function 17. FIG. 6 shows an example of a correspondence table prepared before the collective communication offload function 17 starts communication processing (A4) of FIG. 12, which will be described separately. As shown in FIG. 6, the correspondence table sets the computation process virtual address, offload function virtual address, size, and mkey for each process ID. In FIG. 6, the first line is a correspondence table related to the send buffer of the computation process 111 of the coprocessor 11, and the second line is a correspondence table related to the receive buffer of the computation process 111 of the coprocessor 11. In addition, in FIG. 6, the third line is a correspondence table related to the send buffer of the computation process 121 of the coprocessor 12, and the fourth line is a correspondence table related to the receive buffer of the computation process 121 of the coprocessor 12.

[0030] Fig. 7 is a correspondence table between the virtual addresses of the computation processes of the collective communication offload function 27 of the host 2 and the corresponding virtual addresses of the collective communication offload function 27. Fig. 7 is similar to Fig. 6, so its explanation will be omitted. Note that other hosts also have similar correspondence tables for the computation processes on their hosts.

[0031] FIG. 8 shows a correspondence table between the virtual addresses of the collective communication offload function 17, the virtual addresses of the compute processes, and the physical addresses of the memory windows for each compute process in the coprocessor-host memory translation function 16 of the host 1. FIG. 8 shows an example of a correspondence table prepared before the collective communication offload function 17 starts the communication process (A4) in FIG. 12 (described separately). As shown in FIG. 8, the correspondence table sets the compute process virtual address, the offload function virtual address, the memory window physical address, and the size for each process ID. As in FIG. 6, the first row in FIG. 8 is a correspondence table for the send buffer of the compute process 111 of the coprocessor 11, and the second row is a correspondence table for the receive buffer of the compute process 111 of the coprocessor 11. Also, in FIG. 8, the third row is a correspondence table for the send buffer of the compute process 121 of the coprocessor 12, and the fourth row is a correspondence table for the receive buffer of the compute process 121 of the coprocessor 12.

[0032] Figure 9 is a correspondence table between the virtual addresses of the collective communication offload function 27 for each computing process, the virtual addresses of the computing process, and the physical addresses of the memory window, which are held by the coprocessor-host memory conversion function 26 of the host 2. Since Figure 9 is similar to Figure 8, its explanation will be omitted. Note that other hosts also have similar correspondence tables for the computing processes on their hosts.

[0033] FIG. 10 shows the contents of the memory translation table 131 provided in the network interface 13 of the host 1 for translating an offload function virtual address into a physical address. FIG. 10 also shows an example of a correspondence table prepared before the collective communication offload function 17 starts the communication process (A4) of FIG. 12, which will be described separately. As shown in FIG. 10, the correspondence table sets a virtual address, a physical address, and a size for each mkey. As in FIG. 6, the first row in FIG. 10 is a translation table relating to the send buffer of the computation process 111 of the coprocessor 11, and the second row is a translation table relating to the receive buffer of the computation process 111 of the coprocessor 11. Also, in FIG. 10, the third row is a translation table relating to the send buffer of the computation process 121 of the coprocessor 12, and the fourth row is a translation table relating to the receive buffer of the computation process 121 of the coprocessor 12.

[0034] Fig. 11 shows the contents of the memory conversion table 231 held by the network interface 23 of the host 2. Fig. 11 is similar to Fig. 10, and therefore its explanation will be omitted. Note that other hosts also have similar correspondence tables relating to the computation processes on those hosts.

[0035] 12 is a flowchart showing the operation when each computing process executes MPI_Alltoallv according to an embodiment of the present disclosure. The operation when each computing process executes MPI_Alltoallv will be described with reference to FIG.

[0036] Each computing process sends the parameters of MPI_Alltoallv to the collective communication offload function on the same host (step A1). For example, the computing process 111 in the coprocessor 11 of host 1 sends the contents of Figure 4 to the collective communication offload function 17 of host 1, and the computing process 211 in the coprocessor 21 of host 2 sends the contents of Figure 5 to the collective communication offload function 27 of host 2. Similar processing is also performed in the computing processes on hosts 1 to 4.

[0037] Next, the collective communication offload function allocates a virtual address corresponding to the communication buffer of each computation process (step A2). The collective communication offload function 17 of the host 1 inputs the process ID = 111, computation process virtual address = 0x00100000, and size = 0x8000 for the send buffer of the computation process 111, and requests the coprocessor-host memory translation function 16 of the host 1 to allocate a virtual address.

[0038] The coprocessor-host memory translation function 16 of the host 1 assigns a virtual address of 0x40100000 to this request. It also sets the contents of the first line in Figure 8 as the process ID of 111, the calculation process virtual address of 0x00100000, and the size of 0x8000, which are input from the collective communication offload function 17 of the host 1, and sets the assigned virtual address to 0x40100000. At this point, it creates a correspondence table with the memory window physical address corresponding to the process ID of 111 as not set.

[0039] The collective communication offload function 17 of host 1 uses the virtual address information returned from the coprocessor-to-host memory translation function 16 to set the contents of the first line in Figure 6. Here, the process ID = 111, calculation process virtual address = 0x00100000, size = 0x8000 input to the coprocessor-to-host memory translation function, and the offload function virtual address 0x40100000 returned from the coprocessor-to-host memory translation function 16 are set as the contents of the first line in Figure 6. At this point, a correspondence table is created with mkey = not set corresponding to process ID = 111.

[0040] The collective communication offload function 17 of the host 1 performs the same process on the receive buffer of the computing process 111. The collective communication offload function 17 also performs the same process on the receive buffer and send buffer of the computing process 121, and creates the first to fourth rows as shown in Fig. 7. At this point, in the table of Fig. 7, the memory window physical address = not set.

[0041] The collective communication offload function 27 of the host 2 performs the same processing for the computation processes 211 and 221, and creates the first to fourth rows as shown in Fig. 7. At this point, in the table of Fig. 7, the memory window physical address = not set.

[0042] The same processing is performed in the collective communication offload functions of hosts 3 and 4.

[0043] Next, the collective communication offload function performs memory registration with the network interface management function (step A3). First, the collective communication offload function 17 of the host 1 performs memory registration of the send buffer. Because 0x40100000 was assigned as the virtual address of the collective communication offload function 17 for the send buffer in step A2, a memory registration request is made to the network interface management function 14 specifying virtual address = 0x40100000 and size = 0x8000 as memory registration parameters.

[0044] The network interface management function 14 queries the coprocessor / host memory translation function 16 for the physical address corresponding to the virtual address = 0x40100000 and size = 0x8000.

[0045] The coprocessor-host memory translation function 16, which has received an inquiry from the network interface management function 14, refers to the table shown in Fig. 8 and searches for a row that includes the range of virtual address = 0x40100000 and size = 0x8000. As a result, it acquires the first row in Fig. 8 (information about the send buffer of the calculation process 111).

[0046] The coprocessor-host memory conversion function 16 requests the coprocessor management function 15 to allocate a memory window and acquire its physical address using the process ID of the acquired line (=111), the calculation process virtual address (=0x00100000), and the size (=0x8000) as parameters.

[0047] In response to a request from the coprocessor-host memory translation function 16 , the coprocessor management function 15 assigns a physical address of 0x80100000 to the memory window 113 and returns it to the coprocessor-host memory translation function 16 .

[0048] The coprocessor-host memory conversion function 16 updates the first line of Figure 8 with the returned memory window physical address information. It also returns the physical address = 0x80100000 to the network interface management function 14.

[0049] The network interface management function 14 assigns a unique identification number 1000 as the mkey within the network interface 13, and updates the memory conversion table 131 as shown in the first line of Figure 10, with mkey = 1000, calculation process virtual address = 0x00100000, memory window physical address = 0x80100000, and size = 0x8000.

[0050] The network interface management function 14 returns the identification number=1000 as a result of memory registration to the collective communication offload function 17. The collective communication offload function 17 updates the content of the first line in Fig. 6 from mkey=unset to mkey=1000.

[0051] Next, the collective communication offload function 17 of the host 1 performs the same processing on the receive buffer of the computing process 111, and also performs the same processing on the send buffer and receive buffer of the computing process 121.

[0052] By the above processing of A1 to A3, the tables shown in FIGS. 6, 8 and 10 are prepared.

[0053] In addition, in host 2, similar processing is performed for calculation process 211 and calculation process 221, and rows 1 to 4 of FIG. 11 are created for memory conversion table 231 of host 2. Furthermore, at this point, the tables shown in FIGS. 7, 9, and 11 are prepared for host 2. Note that in hosts 3 and 4, tables similar to the tables shown in FIGS. 6, 8, and 10 are prepared by processing A1 to A3 of FIG. 12.

[0054] Next, the collective communication offload function performs communication processing (step A4), the details of which will be described later with reference to FIG.

[0055] Finally, the collective communication offload function notifies the computation process of the completion of the collective communication (step A5).

[0056] 13 is a diagram showing in detail the communication processing by the collective communication offload function (step A4 in FIG. 12). The communication processing by the collective communication offload function will be explained with reference to FIG.

[0057] First, the collective communication offload function sends a transmission request to all ranks (steps B1 and B2). As an example, we will explain a case where a transmission request is sent to a computing process 211 in host 2 of rank 2 while a collective communication is being processed by a computing process 111 in host 1.

[0058] The collective communication offload function 17 of host 1 references the parameters of the MPI_Alltoallv collective communication executed by the computation process 111 shown in Figure 4, and obtains a send offset of 0x2000 and a send length of 0x1000 for rank 2. As a result, it obtains 0x00102000, which is the send buffer address for rank 2, obtained by adding the send offset of 0x2000 to the send buffer address of 0x00100000. It also references the computation process arrangement shown in Figure 3, and obtains host 2 as the host for rank 2.

[0059] Next, the collective communication offload function 17 sends a send request to the collective communication offload function 27 of host 2. The collective communication offload function 17 refers to the correspondence table shown in FIG. 6 and searches for a row containing the send buffer address for rank 2 = 0x00102000 and the send length = 0x1000. As a search result, the collective communication offload function 17 obtains the virtual address on host 1 = 0x40102000 and mkey = 1000. Here, the virtual address on host 1 = 0x40102000 is obtained by subtracting the computing process virtual address = 0x00100000 in the first row of the table shown in FIG. 6 from the send buffer address for rank 2 = 0x00102000 and adding the offload function virtual address = 0x40100000 in the first row of the table shown in FIG. 6. The collective communication offload function 17 sends the following parameters of the send request to the collective communication offload function 27 of the host 2: send rank=0, send virtual address=0x40102000, send length=0x1000, mkey=1000, and receive rank=2.

[0060] The transmission of a transmission request by the collective communication offload function described above can also be performed in cases other than when a transmission request is sent to the computing process 211 in the host 2 of rank 2 while the computing process 111 in the host 1 is processing collective communication. That is, the collective communication offload function of each host performs the above-mentioned transmission request transmission processing (B1, B2) for communication between multiple computing processes.

[0061] Next, the collective communication offload function receives a transmission request from each rank (steps B3 and B4). In this embodiment, a case where a transmission request is received from a computing process 211 in a host 2 of rank 2 while a collective communication of a computing process 111 in a host 1 is being processed is described.

[0062] The send request from the computing process 211 of host 2 is sent from the collective communication offload function 27 of host 2, and its parameters are send rank = 2, send virtual address = 0x50100000, send length = 0x1000, mkey = 2000, and receive rank = 0. As a result, the receive buffer address for rank 2 is obtained as 0x00202000, which is the receive buffer address = 0x00200000 plus the receive offset = 0x2000.

[0063] The collective communication offload function 17 of the host 1 searches the row with rank number=2 in the table shown in FIG. 4, and obtains receive offset=0x2000 and receive length=0x1000.

[0064] Next, the collective communication offload function 17 searches the table shown in Fig. 6 for a row containing process ID = 111 corresponding to receive rank = 0, receive buffer address = 0x00202000, and receive length = 0x1000, and as a result, obtains offload function virtual address = 0x40202000 and mkey = 1001. Here, the offload function virtual address = 0x40202000 is obtained by subtracting the calculation process virtual address = 0x00200000 in the second row of Fig. 6 from the receive buffer address = 0x00202000 and adding the offload function virtual address = 0x40200000.

[0065] Next, the collective communication offload function issues an RDMA READ (step B5). "RDMA READ" refers to a process of directly transferring (reading) data from the memory of another host to the memory of the host itself, without going through the processing of the OS of each host. In this embodiment, an example will be described in which a send request is received from a rank 2 computing process 211 while a collective communication process of a computing process 111 in the host 1 is being processed.

[0066] Based on the contents of the send request received in step B4 and the contents obtained from Figure 6, the collective communication offload function 17 of host 1 instructs the network interface 13 of host 1 to start RDMA using parameters: communication direction = RDMA READ, communication partner = host 2, local side virtual address = 0x40202000, local side mkey = 1001, and remote side virtual address = 0x50100000, remote side mkey = 2000, communication size = 0x1000 in the send request sent from the collective communication offload function 27 of host 2.

[0067] The network interface 13 sends an RDMA READ request to the network interface 23 of the host 2. The network interface 23 of the host 2 refers to the memory translation table 231 in FIG. 11, obtains the physical address 0x90100000 from the remote virtual address 0x50100000 and the remote mkey 2000, and performs a DMA READ from the physical address 0x90100000. In this embodiment, the memory access to the physical addresses 0x90100000-0x90107fff as shown in FIG. 9 corresponds to the memory access to the virtual addresses 0x10100000-0x10107fff of the computation process 211 in the host 2. Therefore, the transmission data for rank 0 of MPI_Alltoallv performed by the computation process 211 as shown in FIG. 5 is read.

[0068] Next, the network interface 23 of host 2 transmits the read data to the network interface 13 of host 1. The network interface 13 of host 1 refers to the memory translation table 131 shown in FIG. 10, obtains the physical address 0x80202000 from the local virtual address 0x40202000 and the local mkey 1001, and performs a DMA WRITE to the physical address 0x80202000. "DMA WRITE" refers to the process of writing data directly to memory in the host itself, without going through processing by the OS.

[0069] In this embodiment, as shown in FIG. 8, the memory access to the physical addresses 0x80200000-0x80207fff is equivalent to the memory access to the virtual addresses 0x00200000-0x00207fff of the input buffer of the calculation process 111, and therefore the received data from rank 2 of MPI_Alltoallv performed by the calculation process 111 is written as shown in FIG. 4.

[0070] Finally, the collective communication offload function waits for all RDMA READs to be completed (step B6).

[0071] The above-mentioned reception processing of a transmission request by the collective communication offload function can also be performed when a transmission request is received from the computing process 211 in the host 2 of rank 2 while the collective communication processing of the computing process 111 in the host 1 is being performed. That is, when the collective communication offload function of each host receives a transmission request from a computing process in another host, it performs the same reception processing of the transmission request (B3 to B6) as described above.

[0072] As described above, this disclosure provides a collective communication offload function on the host. The collective communication offload function uses the coprocessor-to-host memory translation function to treat the communication buffer of the computation process running on the coprocessor as part of its own virtual address. This virtual address is registered in the memory address translation table of the network interface, and a physical address that can access the communication buffer of the computation process is set. This allows the collective communication offload function to communicate directly via RDMA without copying the communication buffer of the computation process onto the host. As a result, collective communication using the collective communication offload function can be performed without interrupting the computation processing of the computation process.

[0073] Furthermore, as shown in one embodiment of the present disclosure, a computing process can perform computational processing between transmitting collective communication parameters in step A1 and receiving a collective communication completion notification in step A5. This is because the computing process does not need to perform processing related to collective communication during this period. Therefore, in a computing system consisting of multiple hosts equipped with coprocessors, when collective communication is performed between computing processes running on the coprocessors, the communication processing is performed using the collective communication offload function on the host, eliminating the need for communication processing in the computing process and enabling overlap of communication and computation.

[0074] In this disclosure, the network interface management function, coprocessor-host memory conversion function, coprocessor management function, and collective communication offload function are software running on the host, but they may also be dedicated hardware or firmware.

[0075] Although this disclosure uses MPI_Alltoallv as an example, it can also be applied to other collective communications. In this case, the collective communication offload function does not necessarily need to perform RDMA between coprocessors; it can relay data through memory on the host accessible from the collective communication offload function, as in known offload technologies using BlueField DPUs. For example, when performing MPI_Bcast with rank 0 as the root, the collective communication offload function 17 does not send send requests to each rank, but instead issues a send request to each collective communication offload function on each host. Each collective communication offload function can issue an RDMA READ from the send buffer of rank 0 to memory on the host, and after completion, RDMA WRITE the received data to each computation process on the same host.

[0076] 14 is a diagram illustrating an example configuration of an information processing device serving as a host according to an embodiment of the present disclosure. The information processing device 1 includes a coprocessor-to-host memory conversion means 116 and a collective communication offload means 117. The coprocessor-to-host memory conversion means 116 maps a computing process virtual memory for a computing process in a coprocessor included in the information processing device to an offload function virtual address for collective communication offload. The collective communication offload means 117 uses the mapped offload function virtual address to perform remote direct memory access (RDMA) with another information processing device connected via a network, thereby implementing communication between the computing process and the computing process of the other information processing device.

[0077] FIG. 15 is a block diagram showing an example of the hardware configuration of a host information processing device. The hardware configuration of the information processing device 1 includes a CPU 51, a RAM (Random Access Memory) 52, a ROM (Read Only Memory) 53, a storage device 54, and the like. The ROM 53 and the storage device 54 store programs and information for implementing the network interface management function, coprocessor-host memory conversion function, coprocessor management function, and collective communication offload function. The RAM 52 is used as a work area for temporarily storing data used by the CPU 51 and the like during operation. The information processing device 1 also includes an input / output port 55 that serves as an interface with input / output devices such as a display device, mouse, and keyboard. The input / output port 55 also functions as a communication port for communication with other host information processing devices. The ROM 53 may be configured with an EEPROM (Electrically Erasable Programmable Read-Only Memory) or the like, and the recording device 54 may be configured with a hard disk, SSD, or the like, and the computer programs for realizing the network interface management function, coprocessor-host memory conversion function, coprocessor management function, and collective communication offload function may be updated in the ROM 53 or the recording device 54. The coprocessor is omitted from Figure 15.

[0078] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.

[0079] Some or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes.

[0080] (Appendix 1) a coprocessor-host memory conversion means for mapping a calculation process virtual memory for a calculation process in a coprocessor provided in the information processing device to an offload function virtual address for collective communication offload; a collective communication offload means for performing remote direct memory access (RDMA) to another information processing device connected via a network using the mapped offload function virtual address, and for performing communication between the computing process and a computing process of the other information processing device; An information processing device comprising:

[0081] (Appendix 2) the collective communication offload means has a correspondence table between the computing process virtual memory and the offload function virtual address prepared by using the coprocessor-host memory conversion means, converts the computing process virtual memory for the computing process into the offload function virtual address, and transmits a transmission request to the other information processing device, including as parameters the converted offload function virtual address, information identifying the computing process, and information identifying the computing process in the other information processing device; 2. The information processing device according to claim 1.

[0082] (Appendix 3) When the collective communication offload means receives a transmission request from the collective communication offload means of the other information processing device, the collective communication offload means identifies a computing process in its own information processing device that is the target of the transmission request from the received transmission request, obtains the offload function virtual address for the identified computing process, and issues a read request by RDMA to the other information processing device that has transmitted the transmission request, including the obtained offload function virtual address as a parameter. 10. The information processing device according to claim 1 or 2.

[0083] (Appendix 4) a network interface means having a memory translation table that associates physical addresses corresponding to the computing process virtual memory for the computing process with the offload function virtual addresses; When the network interface means receives the read request via RDMA, it uses the memory translation table to translate the offload function virtual address for the computing process that has made the transmission request into the physical address, reads data from the translated physical address, and transmits the read data to the information processing device that has issued the read request via RDMA. 4. The information processing device according to claim 3.

[0084] (Appendix 5) a network interface means having a memory translation table that associates physical addresses corresponding to the computing process virtual memory for the computing process with the offload function virtual addresses; The network interface means converts an offload function virtual address for the computing process that is the recipient of the transmission request into a physical address using the memory conversion table, and writes data received from the other information processing device in response to the read request by RDMA to the physical address converted by Direct Memory Access (DMA). 5. The information processing device according to claim 3 or 4.

[0085] (Appendix 6) the coprocessor-host memory conversion means maps a computation process virtual memory to an offload function virtual address for collective communication offload for each computation process in the coprocessor; the collective communication offload means performs the RDMA to another information processing device connected via a network, using the offload function virtual address mapped for the computing process that is the target of collective communication offload, and implements communication between the computing process of its own information processing device and the computing process of the other information processing device; An information processing device according to any one of Supplementary Note 1 to Supplementary Note 5.

[0086] (Appendix 7) the coprocessor-host memory conversion means maps a computation process virtual memory to an offload function virtual address for collective communication offload for each computation process in the coprocessor and for each send buffer and receive buffer; the collective communication offload means performs the RDMA to another information processing device connected via a network, using the offload function virtual address of the send buffer or receive buffer mapped for the computing process that is the target of collective communication offload, depending on the content of communication, and performs communication between the computing process of its own information processing device and the computing process of the other information processing device; 7. The information processing device according to claim 6.

[0087] (Appendix 8) An information processing system configured by connecting a plurality of information processing devices according to any one of Supplementary Note 1 to Supplementary Note 7.

[0088] (Appendix 11) mapping a calculation process virtual memory for a calculation process in a coprocessor provided in the information processing device to an offload function virtual address for collective communication offload; performing remote direct memory access (RDMA) to another information processing device connected via a network using the mapped offload function virtual address, and performing communication between the computing process and a computing process of the other information processing device; A method for offloading collective communication using an information processing device.

[0089] (Appendix 12) a correspondence table between the computing process virtual memory and the offload function virtual address prepared by using the coprocessor-host memory conversion means, converting the computing process virtual memory for the computing process into the offload function virtual address, and transmitting a transmission request to the other information processing device, including as parameters the converted offload function virtual address, information identifying the computing process, and information identifying the computing process in the other information processing device; 12. The method of offloading collective communications according to claim 11.

[0090] (Appendix 13) When receiving a transmission request from the collective communication offload means of the other information processing device, identify a computing process in the own information processing device that is the target of the transmission request from the received transmission request, obtain the offload function virtual address for the identified computing process, and issue a read request by RDMA to the other information processing device that has sent the transmission request, including the obtained offload function virtual address as a parameter. 13. The collective communication offloading method according to claim 11 or 12.

[0091] (Appendix 14) a memory translation table that associates physical addresses corresponding to the computing process virtual memory for the computing process with the offload function virtual addresses; When the read request by RDMA is received, the offload function virtual address for the computing process that made the transmission request is converted into the physical address using the memory conversion table, data is read from the converted physical address, and the read data is transmitted to the information processing device that issued the read request by RDMA. 14. The method of offloading collective communications according to claim 13.

[0092] (Appendix 15) a memory translation table that associates physical addresses corresponding to the computing process virtual memory for the computing process with the offload function virtual addresses; using the memory translation table, converting an offload function virtual address for the computing process that is the recipient of the transmission request into a physical address, and writing data received from the other information processing device in response to the read request by RDMA to the physical address converted by Direct Memory Access (DMA); 15. The method for offloading collective communication according to claim 13 or 14.

[0093] (Appendix 16) For each computation process in the coprocessor, mapping a computation process virtual memory to an offload function virtual address for collective communication offload; performing the RDMA to another information processing device connected via a network using the offload function virtual address mapped for the computing process that is the target of collective communication offload, and implementing communication between the computing process of the information processing device itself and the computing process of the other information processing device; 16. A collective communication offloading method according to any one of Supplementary Note 11 to Supplementary Note 15.

[0094] (Appendix 17) Mapping a computation process virtual memory to an offload function virtual address for collective communication offload for each computation process in the coprocessor and for each send buffer and receive buffer; According to the content of communication, the offload function virtual address of the mapped send buffer or receive buffer for the computing process that is the target of collective communication offload is used to perform the RDMA to another information processing device connected via a network, and communication is performed between the computing process of the own information processing device and the computing process of the other information processing device. 17. The method of offloading collective communications according to claim 16.

[0095] (Appendix 18) A collective communication offloading method in which the collective communication described in Supplementary Note 11 to Supplementary Note 17 is offloaded in an information processing system configured by connecting a plurality of information processing devices described in Supplementary Note 1 to Supplementary Note 7.

[0096] (Appendix 21) mapping a calculation process virtual memory for a calculation process in a coprocessor provided in the information processing device to an offload function virtual address for collective communication offload; performing remote direct memory access (RDMA) to another information processing device connected via a network using the mapped offload function virtual address, and performing communication between the computing process and a computing process of the other information processing device; A program that makes a computer do something.

[0097] (Appendix 22) a correspondence table between the computing process virtual memory and the offload function virtual address prepared by using the coprocessor-host memory conversion means, converting the computing process virtual memory for the computing process into the offload function virtual address, and transmitting a transmission request to the other information processing device, including as parameters the converted offload function virtual address, information identifying the computing process, and information identifying the computing process in the other information processing device; 22. The program according to claim 21, which causes a computer to execute the steps.

[0098] (Appendix 23) When receiving a transmission request from the collective communication offload means of the other information processing device, identify a computing process in the own information processing device that is the target of the transmission request from the received transmission request, obtain the offload function virtual address for the identified computing process, and issue a read request by RDMA to the other information processing device that has sent the transmission request, including the obtained offload function virtual address as a parameter. 23. The program according to claim 21 or 22, which causes a computer to execute the steps.

[0099] (Appendix 24) a memory translation table that associates physical addresses corresponding to the computing process virtual memory for the computing process with the offload function virtual addresses; When the read request by RDMA is received, the offload function virtual address for the computing process that made the transmission request is converted into the physical address using the memory conversion table, data is read from the converted physical address, and the read data is transmitted to the information processing device that issued the read request by RDMA. 24. The program according to claim 23, which causes a computer to execute the steps.

[0100] (Appendix 25) a memory translation table that associates physical addresses corresponding to the computing process virtual memory for the computing process with the offload function virtual addresses; using the memory translation table, converting an offload function virtual address for the computing process that is the recipient of the transmission request into a physical address, and writing data received from the other information processing device in response to the read request by RDMA to the physical address converted by Direct Memory Access (DMA); 25. The program according to claim 23 or 24, which causes a computer to execute the steps.

[0101] (Appendix 26) For each computation process in the coprocessor, mapping a computation process virtual memory to an offload function virtual address for collective communication offload; performing the RDMA to another information processing device connected via a network using the offload function virtual address mapped for the computing process that is the target of collective communication offload, and implementing communication between the computing process of the information processing device itself and the computing process of the other information processing device; 26. A program according to any one of appendices 21 to 25, which causes a computer to execute the following:

[0102] (Appendix 27) Mapping a computation process virtual memory to an offload function virtual address for collective communication offload for each computation process in the coprocessor and for each send buffer and receive buffer; According to the content of communication, the offload function virtual address of the mapped send buffer or receive buffer for the computing process that is the target of collective communication offload is used to perform the RDMA to another information processing device connected via a network, and communication is performed between the computing process of the own information processing device and the computing process of the other information processing device. 17. The program according to claim 16, which causes a computer to execute the steps.

[0103] (Appendix 28) A program that causes a computer to execute a method for offloading collective communication using the information processing device according to any one of Supplementary Notes 11 to 17 in an information processing system configured by connecting multiple information processing devices according to any one of Supplementary Notes 1 to 7. [Explanation of symbols]

[0104] 1 Host (information processing device) 11,12 Coprocessor 13 Network Interface 14 Network Interface Management Functions 15 Coprocessor Management Functions 16 Coprocessor-host memory conversion function 17 Collective communication offload function 111 Calculation Process 112 Communication Buffer 113 Memory Window 121 Calculation Process 122 communication buffer 123 Memory Window 131 Memory Conversion Table

Claims

1. a coprocessor-host memory conversion means for mapping a calculation process virtual memory for a calculation process in a coprocessor provided in the information processing device to an offload function virtual address for collective communication offload; a collective communication offload means for performing remote direct memory access (RDMA) to another information processing device connected via a network using the mapped offload function virtual address, and for implementing communication between the computing process and a computing process of the other information processing device; An information processing device comprising:

2. the collective communication offload means has a correspondence table between the computing process virtual memory and the offload function virtual address prepared by using the coprocessor-host memory conversion means, converts the computing process virtual memory for the computing process into the offload function virtual address, and transmits a transmission request to the other information processing device, including as parameters the converted offload function virtual address, information identifying the computing process, and information identifying the computing process in the other information processing device; The information processing device according to claim 1 .

3. When the collective communication offload means receives a transmission request from the collective communication offload means of the other information processing device, the collective communication offload means identifies a computing process in its own information processing device that is the target of the transmission request from the received transmission request, obtains the offload function virtual address for the identified computing process, and issues a read request by RDMA to the other information processing device that has transmitted the transmission request, including the obtained offload function virtual address as a parameter. The information processing device according to claim 1 .

4. a network interface means having a memory translation table that associates physical addresses corresponding to the computing process virtual memory for the computing process with the offload function virtual addresses; When the network interface means receives the read request via RDMA, it uses the memory translation table to translate the offload function virtual address for the computing process that has made the transmission request into the physical address, reads data from the translated physical address, and transmits the read data to the information processing device that has issued the read request via RDMA. The information processing device according to claim 3 .

5. a network interface means having a memory translation table that associates physical addresses corresponding to the computing process virtual memory for the computing process with the offload function virtual addresses; the network interface means uses the memory translation table to translate an offload function virtual address for the computing process that is the recipient of the transmission request into a physical address, and writes data received from the other information processing device in response to the read request by RDMA to the physical address translated by Direct Memory Access (DMA); The information processing device according to claim 3 .

6. the coprocessor-host memory conversion means maps a computing process virtual memory to an offload function virtual address for collective communication offload for each computing process in the coprocessor; the collective communication offload means performs the RDMA to another information processing device connected via a network, using the offload function virtual address mapped for the computing process that is the target of collective communication offload, and implements communication between the computing process of its own information processing device and the computing process of the other information processing device; The information processing device according to claim 1 .

7. the coprocessor-host memory conversion means maps a computing process virtual memory to an offload function virtual address for collective communication offload for each computing process in the coprocessor and for each transmission buffer and reception buffer; the collective communication offload means performs the RDMA to another information processing device connected via a network, using the offload function virtual address of the send buffer or receive buffer mapped for the computing process that is the target of collective communication offload, depending on the content of communication, and performs communication between the computing process of its own information processing device and the computing process of the other information processing device; The information processing device according to claim 6 .

8. 8. An information processing system configured by connecting a plurality of information processing devices according to claim 1.

9. mapping a calculation process virtual memory for a calculation process in a coprocessor provided in the information processing device to an offload function virtual address for collective communication offload; using the mapped offload function virtual address to perform remote direct memory access (RDMA) with another information processing device connected via a network, and performing communication between the computing process and a computing process of the other information processing device; A method for offloading collective communication using an information processing device.

10. mapping a calculation process virtual memory for a calculation process in a coprocessor provided in the information processing device to an offload function virtual address for collective communication offload; using the mapped offload function virtual address to perform remote direct memory access (RDMA) with another information processing device connected via a network, and performing communication between the computing process and a computing process of the other information processing device; A program that makes a computer do something.

Citation Information

Patent Citations

  • JP2017‐187973A

Cited By

  • Ensemble communication unloading method, system, equipment and medium

    CN121979690A