Efficient iterative collective operation using network attached memory
By allocating memory areas in the network-attached memory and updating only the sparse modification part of the data set using a new interface call, the inefficiency problem caused by sparse modification when performing collective operations between processing entities of multiple computing nodes is solved, and more efficient iterative collective operations are achieved.
Patent Information
- Application Number
- CN202410884170.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-06
- Filing Date
- 2024-07-03
- Publication Date
- 2025-06-06
AI Technical Summary
When performing collective operations between processing entities (ranks) of multi-computing nodes, although only a few data elements are modified (sparsely modified), the entire data set is still required to be transferred, resulting in inefficiency, especially in the case of iterative collective operations.
By allocating memory areas in network-attached memory and updating only sparsely modified portions of the dataset using a new interface call, avoiding transfer of the entire dataset in subsequent iterations of collective operations.
The number of communication times and data amounts with the network attached memory are reduced, and the system efficiency of performing iterative collective operations is improved.
Smart Images

Figure CN120104044A_ABST
Abstract
Description
Technical Field
[0001] Processing entities (or "ranks") on multiple compute nodes can perform collective operations, where each rank can distribute its portion of data across ranks, or each rank can send its portion of data to one rank, which performs the computation and returns the results to every other rank. In these use cases, the entire sequence of transfer and computation operations needs to be performed, even when only a few data elements are modified in a repeated sequence of operations (i.e., "sparse modifications"). Collective operations with sparse modifications still result in the original amount of data being transferred, even though only a portion of the data is modified in subsequent operations. Furthermore, this inefficiency is exacerbated in the case of iterative collective operations with sparse modifications, as each iteration also results in the original amount of data being transferred. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Figure 1 An environment for facilitating efficient iterative collective operations using network attached storage according to an aspect of the present application is illustrated.
[0003] Figure 2 A memory region with memory segments assigned to ranks according to an aspect of the present application is illustrated.
[0004] FIG. 3A to FIG. 3D An example of collective operation according to an aspect of the present application is illustrated.
[0005] Figure 4A A diagram illustrating communications for memory allocation in collective operations according to an aspect of the present application is presented.
[0006] Figure 4B A data structure storing a virtual mapping according to an aspect of the present application is depicted.
[0007] Figure 5A The reduce collective operation according to one aspect of the present application is illustrated.
[0008] Figure 5B An all-reduce collective operation according to one aspect of the present application is illustrated.
[0009] Fig. 6A A diagram illustrating communications for efficient iterative collective operations using network attached storage according to an aspect of the present application is presented, including a first iteration of an all-reduce collective operation.
[0010] Figure 6BA diagram illustrating communications for an efficient iterative collective operation using network attached storage, including an update operation and a second or subsequent iteration of an all-reduce collective operation, according to an aspect of the present application is presented.
[0011] Figure 7 A diagram illustrating updating a segment of a network attached storage based on an update operation according to an aspect of the present application is illustrated.
[0012] FIG. 8A to FIG. 8B A flow chart illustrating a method for facilitating efficient iterative collective operations using network attached storage according to an aspect of the present application is presented.
[0013] Fig. 9 A computer system for facilitating efficient iterative collective operations using network attached storage according to one aspect of the present application is illustrated.
[0014] Fig.10 A non-transitory computer-readable medium for facilitating efficient iterative collective operations using network attached storage according to an aspect of the present application is illustrated.
[0015] In the drawings, like reference numerals refer to like drawing elements. DETAILED DESCRIPTION
[0016] Aspects of the present application provide a method, computer system, and computer-readable medium that promotes efficient iterative collective operations using network attached memory. Collective operations (e.g., message passing interface (MPI) collective operations) can allocate memory areas for data in network attached memory (e.g., networked memory (FAM) nodes), and calculations can be unloaded to FAM nodes via remote procedure calls. The described aspects can transmit the entire data set to FAM in the first iteration of the collective operation. Using new interface calls, the described aspects can perform updates with sparse modifications by only updating specific sections of the data set in the memory area allocated in the network attached memory. Therefore, in subsequent iterations of the collective operation, the described aspects can avoid (bypass) transmitting the entire data set.
[0017] Thus, by using new interface calls to update only sparse modifications of the entire data set, the described aspects can reduce the number of communications with the network attached storage and the amount of data transferred to and from the network attached storage. This reduction in communications and data volume can make systems for performing iterative collective operations using network attached storage more efficient.
[0018] As described above, processing entities or ranks on multiple compute nodes can perform collective operations, where each rank can distribute its portion of data across ranks, or each rank can send its portion of data to a rank that performs the computation and returns the results to every other rank. Figure 1 The execution of a collective operation may involve a large number of messages passed between the involved processing entities, as described below with respect to FIG. 3A to FIG. 3D In addition, the execution of the collective operation may involve computations that consume processing resources of the computing nodes.
[0019] Furthermore, execution of a collective operation may require execution of the entire sequence of transfer and computation operations, even when only a few data elements are modified in a repeated sequence of operations (i.e., sparse modifications). A collective operation with sparse modifications still results in the original amount of data being transferred, despite only a portion of the data being modified in subsequent operations. This inefficiency is exacerbated in the case of iterative collective operations with sparse modifications, since each iteration also results in the original amount of data being transferred.
[0020] The described aspects provide a system that can allocate memory regions in a network attached memory, the system allowing processing entities or ranks in multiple computing nodes to access the allocated memory regions. This in turn can enable the computing nodes to use the allocated memory regions instead of passing messages between computing nodes to exchange data in order to perform collective operations. The computations involved in the collective operations can also be offloaded to processors in the network attached memory.
[0021] In addition, after the first iteration of the collective operation (including memory allocation and complete data set transfer) has been performed, the new interface call can be used to perform updates with sparse modifications. Therefore, all subsequent iterations of the collective operation do not need to transfer any data to the allocated memory area. Instead, subsequent iterations of the collective operation only need to read the resulting data from the memory area. Therefore, by using the new interface call, the described aspects can eliminate the need for complete data set transfers in subsequent iterations of the collective operation, which can significantly improve performance.
[0022] Parallel computing architecture (such as Figure 1The collective operations in (as depicted in ) can be based on an interface that passes messages between processing entities or ranks in a computing node. An example of such an interface is a message passing interface (MPI), which defines the syntax and semantics of library routines to support parallel computing (including MPI collective operations) across computing nodes. MPI can provide a communication interface by implementing an application programming interface (API), which can include functions, routines, or methods that can be called or enabled by processing entities or ranks executed on computing nodes. These functions, routines, or methods can include MPI collective operations. Other types of interfaces that facilitate communication between processing entities or ranks in computing nodes can also be used.
[0023] A "computing node" may refer to a computer, a portion of a computer, or a collection of computers.
[0024] A "memory server" may refer to a computing entity associated with persistent memory. The persistent memory may be located in the memory server or may be external to the memory server. The persistent memory may be implemented with one or more persistent memory devices, such as flash memory devices, disk-based storage devices, or other types of memory devices that are capable of retaining data when power is removed. A memory server may also be associated with volatile memory, which may be implemented with one or more volatile memory devices, such as dynamic random access memory (DRAM) devices, static random access memory (SRAM) devices, or other types of memory devices that retain data when power is on but lose data when power is removed. In some aspects, a memory server may include one or more processors and a network interface card (NIC), as described below with respect to Figure 1 as described.
[0025] Environment for facilitating efficient iterative collective operations using network attached storage
[0026] Figure 1 An environment 100 for facilitating efficient iterative collective operations using network attached storage according to an aspect of the present application is illustrated. The environment 100 may include a plurality of computing nodes 102-1 to 102-N (where N is greater than or equal to 2). The computing nodes 102-1 to 102-N may be part of a computing node cluster and may be interconnected to a network attached storage (e.g., FAM 106) via a network 104. The FAM 106 may include a plurality of memory servers 108-1 to 108-M (where M is greater than or equal to 2). M may be the same as N, or M may be different from N. The network 104 may include an interconnect (e.g., a throughput, low latency interconnect) between a processor and a memory, which may be based on an open source standard, a protocol defined by a standards body, or a proprietary implementation.
[0027] Each memory server may include a corresponding processor, memory, and network interface card (NIC). For example, memory server 108-1 may include processor 110-1, memory 112-1, and NIC 114-1. Similarly, memory server 108-M may include processor 110-M, memory 112-M, and NIC 114-M. The corresponding processor may include a central processing unit (CPU) that executes the operating system (OS) and other machine-readable instructions (including firmware such as basic input / output system (BIOS) and application programs) of the memory server. In some aspects, the corresponding processor may refer to another type of processor, such as a graphics processing unit (GPU) that handles specific calculations in the memory server. The corresponding memory may be a persistent memory, a volatile memory, or a combination of the two. The corresponding NIC may allow the memory server to communicate over the network 104 based on one or more protocol layers.
[0028] Each computing node may also include a corresponding processor, local memory, and NIC. For example, computing node 102-1 may include processor 115-1, local memory 116-1, and NIC 117-1. Similarly, computing node 102-N may include processor 115-N, local memory 116-N, and NIC 117-N. Typically, a computing node may include one or more processors. One or more memory devices may be used to implement the corresponding local memory, which may be any combination of persistent memory and volatile memory as described herein. The corresponding NIC may allow computing nodes to communicate over network 104 based on one or more protocol layers.
[0029] The processor of the computing node can execute one or more processing entities (PE). PE can refer to a "rank" associated with a program being executed in the computing node. The program can be an application or another type of program, such as an OS, firmware, and another type of machine-readable instructions. As shown in environment 100, processor 115-1 can execute PE 118-1, and processor 115-N can execute PE 118-P-1 to 118-P. Each PE can access data stored in FAM 106 by reading data from FAM 106 or writing data to the FAM. Each PE can also perform calculations on data read from FAM 106 or data to be written to FAM 106.
[0030] A total of P PEs may be executed across computing nodes 102-1 through 102-N. The P PEs may constitute P ranks that may collaborate to perform collective operations. The P ranks may be part of a communication group that performs collective operations together. Although a specific number of PEs are illustrated in computing nodes 102-1 through 102-N, in some aspects, a different or the same number of PEs may be executed in each computing node.
[0031] The compute node 102-1 may further include a collective operation programming interface 120-1, which may be an API, such as an MPI API or another type of API. The collective operation programming interface 120-1 may include: a memory allocation function 122-1; a collective operation function 124-1; and a sparse modification update function 126-1. The memory allocation function 122-1 may be called (initiated) by a PE (e.g., 118-1) in the compute node 102-1 to allocate a memory area in the FAM 106 to perform a collective operation, such as an MPI_Alloc_mem function. The collective operation function 124-1 may be called by a PE (e.g., 118-1) in the compute node 102-1 to initiate a corresponding collective operation that may use the allocated memory area in the FAM 106, such as described below with respect to FIG. 3A to FIG. 3D The MPI collective operation described. The sparse modification update function 126-1 may be called by a PE (e.g., 118-1) in the compute node 102-1 to update only certain specified portions of a segment of data previously written to the FAM 106. Figure 6B and Figure 7 A detailed description of the update function 126 - 1 is provided.
[0032] During operation, PE 118-1 may call a memory allocation function 122-1 that allocates memory regions in FAM 106 for use by ranks involved in the collective operation. The ranks involved in the collective operation may include PEs 118-1 to 118-P. In some aspects, the ranks involved in the collective operation may include a subset of PEs 118-1 to 118-P. After the memory regions are allocated by the memory allocation function 122-1, PEs 118-1 to 118-P may write data to the allocated memory regions and read data from the allocated memory regions. Information related to the allocated memory regions may be stored in the local memory 116-1 of the computing node 102-1 as, for example, allocated memory region information 128-1. The allocated memory region information 128-1 may include a data structure that stores a mapping of a descriptor (e.g., a FAM descriptor associated with one of the memories 112-1 to 112-M) of a virtual address to a physical location of the memory.
[0033] After a memory region has been allocated via a memory allocation function enabled by a given PE, any of the PEs 118 - 1 through 118 -P may enable a collective operation function to initiate a collective operation utilizing the allocated memory region.
[0034] The computing node 102-1 may further include a FAM programming interface 130-1 and a FAM client 132-1. The FAM programming interface 130-1 may include an API, a library, or any other program-accessible subsystem that is enabled in the computing node to perform operations on the FAM 106, including read operations, write operations, calculations offloaded to the FAM 106, or another operation involving data. In some aspects, the FAM programming interface 130-1 may include an API that includes functions that can be used by the computing node (or more specifically, by the PE in the computing node) to manage the FAM 106 and access the data of the FAM 106. An example of such an API is the OpenFAM API. The FAM client 132-1 may be a program that is executed in the computing node 102-1 to manage access to the FAM 106 (such as in response to a call to the FAM programming interface 130-1).
[0035] Compute node 102-N may include similar elements having functionality as described above with respect to compute node 102-1, including: a collective operation programming interface 120-N, the collective operation programming interface including a memory allocation function 122-N, a collective operation function 124-N, and a sparse modification update function 126-N; a FAM programming interface 130-N; and a FAM client 132-N.
[0036] Example of Network Attached Storage and Collective Operations
[0037] Figure 2 A memory region 200 with memory segments assigned to ranks according to an aspect of the present application is illustrated. As described above with respect to environment 100, memory allocation function 122-1, when enabled, may allocate memory region 200 in persistent memory (e.g., 112-1 to 112-M) of FAM 106. A corresponding segment or portion of memory region 200 may be allocated to each PE (rank 1 to P).
[0038] The segments assigned to a given rank may include a memory portion having a specified memory portion size at a corresponding fixed "offset". The offset may refer to a memory location of a persistent memory (e.g., represented by a memory address). For P ranks, the memory area 200 may be divided into P memory segments 210-1 to 210-P, which are assigned to corresponding P ranks (rank 1 to P, depicted as rank_1 202, rank_2 204, rank_3 206, and rank_P 208). Each of the memory segments 210-1 to 210-P may have the same memory segment size (e.g., memory segment size 222) or different memory segment sizes.
[0039] The memory segment 210-j (j=1 to P) assigned to rank j can be obtained from offset j ( Figure 2 Therefore, Figure 2 In the memory region 200 of FIG. 1 , the memory segment 210-1 assigned to rank_1 202 starts at offset_1 220-1; the memory segment 210-2 assigned to rank_2 204 starts at offset_2 220-2; the memory segment 210-3 assigned to rank_3 206 starts at offset_3 220-3; and the memory segment 210-P assigned to rank_P starts at offset_P 220-P. Although the examples discussed herein involve ranks starting at 1, in some aspects, ranks may start at other numbers (such as 0 (as described below with respect to FIG. 3A to FIG. 3D and Figure 7 In some aspects, the offset may be determined based on the product of the rank number (j) and the memory segment size. The rank number may be an identifier of the rank. For example, PEs 118-1 through 118-P may have rank numbers 1 through P.
[0040] One rank (e.g., rank 1) may call the memory allocation function 122-1 to allocate the memory region 200. Each of the other ranks (e.g., ranks 2 to P) may similarly call the memory lookup function of the collective operation programming interface to obtain information about the allocated memory region 200, which may allow the other ranks to access their allocated memory segments of the allocated memory region 200.
[0041] FIG. 3A to FIG. 3D An example of collective operation according to an aspect of the present application is illustrated. FIG. 3A to FIG. 3D The collective operations depicted in , including broadcast, scatter, gather, all-gather, and all-to-all, can be performed by Figure 1The ranking of the computer nodes 102-1 to 102-N can also be executed. FIG. 3A to FIG. 3D collective operations other than those shown in FIG. 5A to FIG. 5B and Fig. 6A To the reduction collective operation described in Figure 6C.
[0042] FIG. 3A to FIG. 3D Each of the 30000 ranks depicts two arrays, including a first array indicating a state before a collective operation and a second array indicating a state after a collective operation. Each array may include rows representing ranks 0-5 (302) as vertically increasing numbered rows, and may further include columns representing "data units" (304) as elements of each rank or numbered row from left to right. A data unit may refer to a portion of data to be communicated between other ranks. A data unit may include a message, an information element, or any other portion of data.
[0043] Figure 3A An example broadcast collective operation is illustrated. Figure 3A The broadcast collective operation broadcasts (306) data unit A0 from rank 0 (as depicted by array 300) to all other ranks 1-5, which causes all ranks 0-5 to include a copy of data unit A0 (as depicted by array 308).
[0044] Figure 3B An example scatter collective operation and an example gather collective operation are illustrated. Figure 3B The scatter collective operation scatters (312) the data units A0-A5 from rank 0 (as depicted by array 310) to ranks 0-5 (as depicted by array 314). The gather collective operation gathers (316) the data units A0-A5 from ranks 0-5 (as depicted by array 314) to rank 0 (as depicted by array 310). In some aspects, a scatter-gather collective operation may be initiated that performs a scatter and a gather of data units between ranks 0-5.
[0045] Figure 3CAn example all-gather collective operation is illustrated that aggregates (322) data units A0-A5 from ranks 0-5 (as depicted by array 320) to all other ranks 0-5 so that each rank has data units A0-A5 (as depicted by array 324). As depicted by array 320, data unit A0 is initially stored at rank 0, data unit A1 is initially stored at rank 1, data unit A2 is initially stored at rank 2, data unit A3 is initially stored at rank 3, data unit A4 is initially stored at rank 4, and data unit A5 is initially stored at rank 5. In the all-gather collective operation, rank 0 aggregates data units A0-A5 from ranks 0-5, rank 1 aggregates data units A0-A5 from ranks 0-5, rank 2 aggregates data units A0-A5 from ranks 0-5, rank 3 aggregates data units A0-A5 from ranks 0-5, rank 4 aggregates data units A0-A5 from ranks 0-5, and rank 5 aggregates data units A0-A5 from ranks 0-5.
[0046] Figure 3DAn example all-to-all collective operation is illustrated that distributes (332) data units A(0,0) to A(5,5) initially stored at ranks 0-5 (as depicted by array 330) among ranks 0-5 such that each rank has data units initially stored at other ranks (as depicted by array 334). As depicted by array 330, data units A(0,0) to A(0,5) are initially stored at rank 0, data units A(1,0) to A(1,5) are initially stored at rank 1, data units A(2,0) to A(2,5) are initially stored at rank 2, data units A(3,0) to A(3,5) are initially stored at rank 3, data units A(4,0) to A(4,5) are initially stored at rank 4, and data units A(5,0) to A(5,5) are initially stored at rank 5. In the all-to-all collective operation, rank 0 aggregates the data units A(0,0), A(1,0), A(2,0), A(3,0), A(4,0), and A(5,0) from ranks 0-5, rank 1 aggregates the data units A(0,1), A(1,1), A(2,1), A(3,1), A(4,1), and A(5,1) from ranks 0-5, rank 2 aggregates the data units A(0,2), A(1,2), A(2,2), A(3,2), A(4,2), and A(5,2) from ranks 0-5, rank 3 aggregates the data units A(0,3), A(1,3), A(2,3), A(3,3), A(4,3), and A(5,3) from ranks 0-5, and rank 4 aggregates the data units A(0,1), A(1,1), A(2,1), A(3,1), A(4,1), and A(5,1) from ranks 0-5. The data units of 0-5 are A(0,4), A(1,4), A(2,4), A(3,4), A(4,4), A(5,4), and rank 5 aggregates the data units A(0,5), A(1,5), A(2,5), A(3,5), A(4,5), A(5,5) from ranks 0-5.
[0047] Communication for memory allocation in collective operations
[0048] Figure 4A A diagram 400 is presented that illustrates communications for memory allocation in collective operation according to an aspect of the present application. Diagram 400 illustrates communications between the following entities: Figure 1 PE 118-1 to 118-P similar to rank 1-P402; Figure 1 collective operation programming interface 404 similar to collective operation programming interface 120-1 and 120-N of Figure 1The FAM programming interfaces 130 - 1 and 130 -N are similar to the FAM programming interface 406 .
[0049] To allocate a memory region (e.g., a memory segment in FAM 106) for a collective operation, each of ranks 1-P 402 may call a memory allocation function, such as MPI_Alloc_mem function or call 410 in collective operation programming interface 404. Calling the memory allocation function may cause the requested memory region (e.g., memory segments 210-1 to 210-P in memory region 200) to be allocated in FAM 106. Figure 2 As described, each memory segment may start from a corresponding offset (e.g., any one of offsets 220-1 to 220-P) and may be determined as a product of a rank number and a specified memory segment size (222). The MPI_Alloc_mem function may also allocate memory segments in a local memory (e.g., local memory 116-1 to 116-N) for each rank, wherein the memory segments allocated in the local memory may be used as a cache memory for corresponding memory segments allocated in the FAM 106.
[0050] In some aspects, in response to the MPI_Alloc_mem 410 function being called, the collective operation programming interface 404 may call an allocation function in the FAM programming interface 406, such as an OpenFam memory allocation function (e.g., fam_allocate()), which may cause the memory servers 108-1 to 108-M to allocate a memory region (indicated by the allocate memory region 414 operation) in the FAM 106. The allocated memory region in the FAM 106 may include P memory segments for the P ranks.
[0051] As a result of calling, for example, fam_allocate(), the FAM programming interface 406 can return a FAM descriptor 416 of the allocated memory area. The FAM descriptor 416 can include the starting address of the allocated memory area and other information related to the memory area. Since the MPI_Alloc_mem function has been called, the collective programming interface 404 can receive the FAM descriptor 416.
[0052] The collective programming interface 404 may send both the FAM descriptor 418 and the virtual address 420 to the corresponding rank. Thus, when rank j calls the MPI_Alloc_mem function, rank j may receive the virtual address and FAM descriptor corresponding to the allocated memory region for use by other ranks. The corresponding compute node (e.g., associated with a given rank or PE) may create an entry including a mapping of the virtual address 420 to the FAM descriptor 418 and store the entry in a data structure (indicated by the storage mapping 422 operation). Figure 4B to depict an example mapping table and entries.
[0053] Figure 4B A data structure 430 storing a virtual map according to an aspect of the present application is depicted. The data structure 430 may correspond to Figure 4A The data structure 430 may be stored in, for example, the allocated memory region information 128-1 to 128-N. The data structure 430 may include entries indicating at least a virtual address 432 and a FAM descriptor 434. The entry 436 may be stored in the memory allocation communication (e.g., Figure 4A The entry 436 may include at least a virtual address having a value of "VA_123" and a FAM descriptor having a value of "FAMD_123".
[0054] Example of reduction operation
[0055] Figure 5A A reduction collective operation according to an aspect of the present application is illustrated. Rank 502 may include ranks 1-6 storing data units 504, for example, ranks 1, 2, 3, 4, 5, and 6 may initially store data units A1, A2, A3, A4, A5, and A6, respectively. The reduction collective operation may collect data units A1, A2, A3, A4, A5, and A6 from ranks 1, 2, 3, 4, 5, and 6 (as shown by 510, 512, 514, 516, 518, and 520, respectively), and may apply a reduction operation or function F 522 to the collected data units. The reduction function F 522 may be, for example, a sum, a maximum, a minimum, a median, an average, or another operation that aggregates data on the collected data units. The reduction function F 522 may output a value based on the collected data units A1, A2, A3, A4, A5, and A6, which may be represented as “S” and stored at rank 1 (as shown by 524).
[0056] Figure 5BAn all-reduce collective operation according to an aspect of the present application is illustrated. Rank 532 may include ranks 1-6 storing data units 534, for example: rank 1 may initially store vector [8, 6]; rank 2 may initially store vector [7, 5]; rank 3 may initially store vector [3, 0]; rank 4 may initially store vector [9, 8]; rank 5 may initially store vector [4, 2]; and rank 6 may initially store vector [1, 5]. The all-reduce collective operation may collect data units 534 from ranks 1, 2, 3, 4, 5, and 6 (as shown by 540, 542, 544, 546, 548, and 550, respectively), and may apply a reduction operation or function F 552 to the collected data units. The reduction function F 552 may apply, for example, an indexed vector and operation in which the data values at the first index of the vector are added together (i.e., 8+7+3+9+4+1) to produce 32, and the data values at the second index are added together (i.e., 6+5+0+8+2+5) to produce 26. Thus, the reduction function F 552 may produce an output vector [32, 26] that is distributed to each of ranks 1, 2, 3, 4, 5, and 6.
[0057] Communication for efficient iterative all-reduce collective operations (including update operations)
[0058] Fig. 6A A diagram 600 illustrating communications for efficient iterative collective operations using network attached storage according to an aspect of the present application is presented, including a first iteration 610 of an all-reduce collective operation. The first iteration 610 may include communications 612-634, as described herein. Diagram 600 illustrates communications between the following entities: Figure 4A Rank 1-P 402 and Figure 1 PE 118-1 to 118-P similar to rank 1-P 602; Figure 4A The collective operation programming interface 404 and Figure 1 120-1 and 120-N similar collective operation programming interface 604; and Figure 4A The FAM programming interface 406 and Figure 1 FAM programming interface 606 similar to 130-1 and 130-N of FIG. 130; and Figure 1 The communication depicted in diagram 600 assumes that Figure 4A Memory allocation communication has occurred as a result of executing the all-reduce collective operation, and entry 446 has stored a mapping of virtual addresses to FAM descriptors of memory regions allocated to the all-reduce collective operation.
[0059] To perform the all-reduce collective operation, each of the ranks 1-P 602 may call a collective all-reduce function, such as the MPI_Allreduce function in the collective operation programming interface 604 or call 612. When the MPI programming interface is used as the collective operation programming interface 604, the syntax of the MPI_Allreduce function may be expressed as:
[0060] int MPI_Allreduce(const void*sendbuf,void*recvbuf,int count,
[0061] MPI_Datatype datatype,MPI_Op op,MPI_Comm comm),
[0062] in:
[0063] "*sendbuf" can indicate the starting address of the send buffer,
[0064] "*recvbuf" can indicate the starting address of the receive buffer.
[0065] "count" can represent the number of elements in the send buffer,
[0066] "datatype" can indicate the type of data in the send buffer.
[0067] "op" can represent the reduction operation or function to be performed, and
[0068] "comm" can mean either communicator or group.
[0069] In response to calling the MPI_Allreduce 612 function (or other collective operation), the collective operation programming interface 604 may use the virtual address *recvbuf to search (operation 614) the mapping data structure 430 to obtain the corresponding FAM descriptor. If the entry (e.g., entry 436) containing the virtual address *recvbuf exists in the mapping data structure 430, the corresponding FAM descriptor is returned as the search result. Returning the FAM descriptor as the search result may indicate that the memory area for the all-reduce collective operation has been allocated and can be used for the all-reduce collective operation.
[0070] The collective operation programming interface 604 can call a put function in the FAM programming interface 606, such as the fam_put 616 function, which can cause the FAM programming interface 606 to send a command to write (operation 618) the data in the send buffer of each rank (pointed to by *sendbuf) to the allocated memory area in the FAM 608. The FAM 608 can store the written data (operation 620). The data contributed by different ranks can be written to corresponding different memory segments (at different offsets) of the allocated memory area, as described above with respect to Figure 2 as described.
[0071] When multiple ranks write data from their corresponding send buffers to the corresponding memory segments of the allocated memory area of FAM 608, the system can implement a synchronization mechanism (such as a barrier operation) that forces the ranks to wait for the multiple ranks to complete the data writing. When all ranks have written their data to the allocated memory area of FAM 608, the collective operation programming interface 604 can call an unload function 622, which can cause the FAM programming interface 606 to offload the calculation (operation 624) to the memory server (e.g., 108-1 to 108-M) of FAM 608. The processor (e.g., 110-1 to 110-M) of the memory server can perform the reduction operation specified by the parameter "op". In some aspects, calling the unload function 622 can cause the FAM client (e.g., 130-1 to 130-N) to queue the calculation of the reduction operation in a queue, and after queuing the calculation, send the calculation to the memory server for execution (e.g., operation 624). The FAM 608 (via the memory server's processor) may perform the offloaded computations and store the results in a temporary buffer, for example, in an allocated memory area of the FAM 608 (operation 626).
[0072] When the results are available, the collective operation programming interface 604 can call a get function, such as the fam_get 628 function, which can cause the FAM programming interface 606 to send a command to read the results (operation 630) in the receive buffer (pointed to by *recvbuf) in the allocated memory area in the FAM 608. The FAM 608 can return the requested data as a result 632, which can be returned by the FAM programming interface 606 and the collective programming interface 604 as a result 634 to each of the ranks 1-P 602.
[0073] Other types of collective operations involving computations may be performed in a similar manner (eg, by offloading computations to memory servers of the FAM 608).
[0074] Figure 6B A diagram 650 is presented illustrating communications for an efficient iterative collective operation using a network attached storage according to an aspect of the present application, including an update operation 660 and a second or subsequent iteration 680 of an all-reduce collective operation. The update operation 660 may include communications 662-670, and the second or subsequent iteration 680 may include communications 682-698, as described herein. Diagram 650 illustrates communications between the same entities described above in diagram 600. The communications depicted in diagram 650 assume that Fig. 6A Communications in the first iteration 610 of diagram 600 have occurred as a result of executing an all-reduce collective operation, and entry 446 has stored a mapping of virtual addresses to FAM descriptors of memory regions allocated to the all-reduce collective operation.
[0075] If one or more ranks have sparse modifications to data previously transferred to and stored in the allocated memory region of FAM 608, the ranks may use a new MPI interface call (e.g., MPI_Update_data function or call 662) in the collective operation programming interface 604 to update only specific portions of their corresponding segments in the allocated memory region. The syntax of the new MPI_Update_data call may be represented as:
[0076] int MPI_Update_data(const void*sendbuf,uint64_t nelems,
[0077] uint64_t*index_array,void*new_value_array),
[0078] in:
[0079] "*sendbuf" can indicate the starting address of the send buffer, which is the virtual address used in MPI collective operations.
[0080] "nelems" can indicate the number of elements to be updated in "index_array",
[0081] "index_array" may indicate a list or array of indices corresponding to the data units to be modified, and
[0082] 'new_value_array' may indicate an array of new values of data cells to be written to or updated at the corresponding index.
[0083] Below about Figure 7Describes an example of an update operation in which two elements at indices 0 and 2 are updated by the first rank and three elements at indices 1, 3, and 5 are updated by the second rank.
[0084] In response to calling the MPI_Update_data 662 function, the collective operation programming interface 604 can use the virtual address *sendbuf to search (operation 664) the mapping data structure 430 to obtain the corresponding FAM descriptor. If the entry (e.g., entry 436) exists in the mapping data structure 430, the corresponding FAM descriptor is returned as the search result. Returning the FAM descriptor as the search result can indicate that the memory area for the update operation has been allocated.
[0085] The collective operation programming interface 604 can call an update function in the FAM programming interface 606, such as the fam_update666 function, which can cause the FAM programming interface 606 to send a command to write (operation 668) specific data (as indicated in "new_value_array") to the index (as indicated in "index_array") of the corresponding memory segment for the rank in the allocated memory area in the FAM 608. The FAM 608 can update and store the specific data (operation 670). The data contributed by different ranks can be written to corresponding different memory segments (at different offsets) of the allocated memory area, as described above with respect to Figure 2 as described.
[0086] After performing the update operation 660, subsequent MPI collective operation calls from the application program to perform the MPI collective operation do not need to transfer or write the entire data set. Figure 6B As shown, the all-reduction collective operation (i.e., Fig. 6A The second or subsequent iteration 680 of the first iteration 610 can simply offload the computation and read the results, thereby eliminating expensive data transfers for the entire dataset across all ranks.
[0087] That is, each of the ranks 1-P 602 may call a collective all-reduce function, such as the MPI_Allreduce 682 call in the collective operation programming interface 604. In response to calling the MPI_Allreduce 682 function, the collective operation programming interface 604 may search (operation 684) the mapping data structure 430 using the virtual address to obtain the corresponding FAM descriptor. If an entry (e.g., entry 436) containing the virtual address exists in the mapping data structure 430, the corresponding FAM descriptor is returned as a search result, which may indicate that a memory region for the all-reduce collective operation has been allocated and may be used for the all-reduce collective operation. Due to the new update interface call described above in the update operation 660, the data previously stored in the allocated memory region of the FAM 608 (as a result of the first iteration 610) has been updated with the sparse modification operation, and thus any subsequent calls to the all-reduce collective operation (i.e., the second or subsequent iteration 680) may eliminate operations or steps involving data writes (e.g., Fig. 6A 616, 618, and 620). Instead, the collective operation programming interface 604 may immediately call an offload function 686, which may cause the FAM programming interface 606 to offload (operation 688) the computation to a processor of a memory server of the FAM 608 (e.g., processors 110-1 through 110-N of memory servers 108-1 through 108-M of the FAM 106). The FAM 608 (via the processor of the memory server) may perform the offloaded computation and store the result in a temporary buffer, for example, in an allocated memory area of the FAM 608 (operation 690).
[0088] When the results are available, the collective operation programming interface 604 can call a get function, such as the fam_get 692 function, which can cause the FAM programming interface 606 to send a command to read the results (operation 694) in the receive buffer (pointed to by *recvbuf) in the allocated memory area in the FAM 608. The FAM 608 can return the requested data as a result 696, which can be returned by the FAM programming interface 606 and the collective programming interface 604 as a result 698 to each of the ranks 1-P 602.
[0089] Figure 7A diagram 700 is illustrated of updating a segment of a network attached storage based on an update operation according to an aspect of the present application. In diagram 700, the network attached storage is illustrated as FAM 710. FAM 710 may include an allocated memory area, which is illustrated as a send buffer ("sendbuf") 720. FAM 710 may also include a receive buffer ("recvbuf") 730 to which data or results of calculations performed by processors of the storage servers of FAM 710 may be stored and from which such data and results may be retrieved (as described above with respect to FIG. 6A to FIG. 6B as described in FAM 608).
[0090] The sendbuf 720 may include P memory segments, each of which may include six memory "portions" (i.e., each memory segment may store six data units). Each portion of the memory segment may be indicated by an index (as indicated by index 740), and each memory segment may correspond to one of the P ranks involved in a collective operation of the allocated memory region sendbuf 720 using the FAM 710. For example, rank_0 720 may have been assigned a sendbuf rank_0 722 memory segment having six memory portions (e.g., storing data units corresponding to indices [0]-[5]); sendbuf rank_1 (not shown) may have been assigned a sendbuf rank_1 724 memory segment having six memory portions (e.g., storing data units corresponding to indices [0]-[5]); and rank_n 704 may have been assigned a sendbuf rank_n 726 memory segment having six memory portions (e.g., storing data units corresponding to indices [0]-[5]).
[0091] After the first iteration of the all-reduce collective operation has been called, one or more of the ranks 0-n may perform sparse modifications on the data that has been transferred and stored in the allocated memory area of the FAM 710. For example, rank_0 702 may call MPI_Update_data(sendbuf, 2, [0, 2], [x0, y0]) (call 732), which may result in communications as described above with respect to the update operation 660. Specifically, the parameters of call 732 may indicate the following: two elements or data cells to be written to the corresponding memory segment (sendbuf rank_0 722) identified by looking up the "sendbuf" virtual address in the mapping data structure; an array or list of indices "0" and "2" corresponding to respective portions of the corresponding memory segment (722) to be written; and an array or list of new values to be written to the memory locations indicated by the index arrays, where the values are "x0" and "y0", respectively. Thus, as a result of call 732, data values "x0" and "y0" may be written to sendbuf rank — 0 memory segment 722 at indices [0] and [2].
[0092] In another example, rank_n 704 may call MPI_Update_data(sendbuf,3,[1,3,5],[xn,yn,zn]) (call 734), which may result in communications as described above with respect to update operation 660. Specifically, the parameters of call 734 may indicate the following: three elements or data cells to be written to the corresponding memory segment (sendbuf rank_n 726) identified by looking up the "sendbuf" virtual address in the mapping data structure; an array or list of indices "1", "3", and "5" corresponding to respective portions of the corresponding memory segment (726) to be written; and an array or list of new values to be written at the memory locations indicated by the index arrays, where the values are "xn", "yn", and "zn", respectively. Thus, as a result of call 734, data values "xn", "yn", and "zn" may be written to sendbuf rank_n memory segment 726 at indices [1], [3], and [5].
[0093] Method for facilitating efficient iterative collective operations using network attached storage
[0094] FIG. 8A to FIG. 8B Flowcharts 800 and 820 illustrating a method for facilitating efficient iterative collective operations using network attached storage according to an aspect of the present application are presented. Fig. 8A800 presents operations involved in performing a first iteration of a collective operation. The system receives, by a first computing node among a plurality of computing nodes, a first request to perform a collective operation, the first request indicating a first data unit to be written to a first segment of a memory region (operation 802). The first request may correspond to a first iteration of the collective operation, similar to the first iteration described above with respect to Fig. 6A The first iteration 610 of MPI_Allreduce 612 is described. The memory area to which the first data unit is to be written may correspond to Figure 1 FAM 106 or Fig. 6A FAM 608, and the first section can be similar to Figure 2 The system generates a first virtual address for the memory area (operation 804). The first virtual address may correspond to a virtual address of a receive buffer or a transmit buffer indicated in the first request. The system allocates a memory area based on the first virtual address, the memory area being associated with a descriptor of a physical location of the memory (operation 806). When called using the OpenFAM library interface, the descriptor may be a FAM descriptor (416, 418) and may be returned by the FAM programming interface (406) after allocating the memory area (414), as described above with respect to Figure 4A The system stores the mapping of the first virtual address to the descriptor in a data structure (operation 808). For example, a compute node associated with one of the ranks 1-P 402 may store the mapping in a mapping data structure, as described above with respect to Figure 4A 422 and Figure 4B As described by entry 436 of mapping table 430 .
[0095] The system performs a collective operation that includes at least writing a first data unit to a first segment and accessing data units from other segments of a memory area (operation 810). The collective operation may include writing all data to a network attached storage (including writing the first data unit to the first segment) ( Fig. 6A 616, 618 and 620), and may also include offloading any computation associated with the collective operation to a processor associated with the network attached memory ( Fig. 6A The collective operation may also include reading all data from the network attached storage (including accessing data units from other segments of the memory area) ( Fig. 6A 628, 630, 632 and 634). The operation is Figure 8B Continue from label A of .
[0096] Figure 8BFlowchart 820 of presents operations involved in performing an update operation (i.e., sparse modification) of a portion of data stored in a memory segment and a second iteration of the collective operation. The system receives a second request to perform an update operation, the second request indicating a second virtual address, one or more portions of the memory region segment to be updated, and corresponding data units to be written to the one or more portions (operation 822). The update operation can be a new MPI interface call, such as MPI_Update_data, as described above with respect to Figure 6B as described by the update operation 660, and may include parameters such as: a virtual address (e.g., sendbuf) of the allocated memory area segment to be updated; the number of portions of the segment to be updated; an array of indices corresponding to the one or more portions of the segment to be updated; and an array of new values corresponding to data units to be written at corresponding indices of the one or more portions of the segment to be updated.
[0097] The system performs a first search in the data structure for a mapping of the second virtual address (operation 824), the first search being similar to the above with respect to Figure 6B The system identifies a mapping of the first virtual address to the descriptor based on the first search (operation 826) in response to the second virtual address matching the first virtual address. A successful search in the data structure may indicate that the data to be updated has been transferred and therefore the corresponding memory area has been allocated and stores the previously transferred data. Figure 7 As described by the results of the MPI_Update_data calls 732 and 734, the system updates the one or more portions by writing corresponding data units to the one or more portions of the memory region segment based on the identified mapping (operation 828).
[0098] After performing the first iteration of the collective operation (operations 802-810) and the update operation (822-828), the system receives a third request to perform the collective operation, the third request indicating a third virtual address (operation 830). The third request may correspond to a second or subsequent iteration of the collective operation, similar to the above with respect to Figure 6B The system performs a second search in the data structure for a mapping of the third virtual address (operation 832), which is similar to the above with respect to MPI_Allreduce 682. Figure 6B If no mapping is found (decision 834), the operation returns to Fig. 8AOperation 804 of , ie, performing a memory allocation process, because the unsuccessful lookup indicates that the third request does not correspond to the second or subsequent iteration, but rather to the first iteration of the collective operation.
[0099] If a mapping is found (decision 834), the system identifies a mapping of the first virtual address to the descriptor based on the second search in response to the third virtual address matching the first virtual address (operation 836). A successful search in the data structure may indicate that the data associated with the collective operation has been transmitted and, therefore, has been stored in the corresponding allocated memory area. The system performs the collective operation, which further includes at least avoiding writing any data to the memory area and accessing data units only from the memory area based on the identified mapping (operation 838). Therefore, by using the new MPI interface call MPI_Update_data to perform sparse modifications on the data that has been transmitted after the first iteration of the collective operation, the second or subsequent iterations of the collective operation can eliminate the significant traffic cost of transmitting the entire data set. Instead, a given rank or processing entity of the computing node only needs to access data units from the allocated memory area, for example, by reading data from a receive buffer, as described above with respect to Figure 6B Communications 692-698 and Figure 7 The recvbuf 730 is described. The operation returns.
[0100] Computer system for facilitating efficient iterative collective operations using network attached storage
[0101] Fig. 9 A computer system 900 for facilitating efficient iterative collective operations using network attached storage according to one aspect of the present application is illustrated. The computer system 900 includes a processor 902, a memory 904, and a storage device 906. The memory 904 may include a volatile memory (e.g., random access memory (RAM)) that is used as a managed memory and can be used to store one or more memory pools. In addition, the computer system 900 can be coupled to peripheral I / O user devices 910 (e.g., a display device 911, a keyboard 912, and a pointing device 913). The storage device 906 includes a non-transitory computer-readable storage medium and stores an operating system 916, a content processing system 918, and data 932. The computer system 900 may include, for example, a processor 902, a memory 904, and a storage device 906. Fig. 9 Fewer or more entities or instructions may be shown.
[0102] The content processing system 918 may include instructions that, when executed by the computer system 900, may cause the computer system 900 to perform the methods and / or processes described in the present disclosure. Specifically, the content processing system 918 may include instructions 920 for receiving a request to perform a collective operation or an update operation, as described above with respect to FIG. 6A to FIG. 6B The content processing system 918 may include instructions 922 for performing a memory allocation process, for example, by allocating a memory region based on a first virtual address and determining a descriptor of a physical location of the memory, as described above with respect to FIG. 4A to FIG. 4B The content processing system 918 may include instructions 924 for storing a mapping of a descriptor of a virtual address to a physical location in memory, as described above with respect to Figure 1 , FIG. 4A to FIG. 4B and FIG. 6A to FIG. 6B as described.
[0103] The content processing system 918 may also include instructions 926 for performing a collective operation, which may include at least writing a first data unit to a first section and accessing data units from other sections of the memory area, as described above in FIG. 3A to FIG. 3D Example collective operations and Fig. 6A The first operation 610 is described in the communication.
[0104] The content processing system 918 may additionally include instructions 920 for receiving a second request to perform an update operation, the second request indicating a second virtual address, one or more portions of the memory area segment to be updated, and corresponding data units to be written to the one or more portions, as described above with respect to Figure 6B The content processing system 918 may include instructions 928 for searching a mapping in a data structure, including searching the data structure to obtain a mapping of a first virtual address to a descriptor based on a second virtual address, the second virtual address mapping the first virtual address, as described above with respect to Fig. 6A and Figure 6B As described by functions 614, 664 and 684.
[0105] The content processing system 918 may include instructions 930 for updating the one or more portions by writing corresponding data units to the one or more portions of the memory area segment based on the obtained mapping, as described above with respect to Figure 6B Operation 660 and Figure 7 The results of the MPI interface calls 732 and 734 are described.
[0106] Data 932 may include any data required as input or generated as output by the methods, operations, communications, and / or processes described in the present disclosure. Specifically, data 932 may store at least: requests; data units; indicators of segments of memory areas; virtual addresses; virtual addresses corresponding to send buffers or receive buffers; descriptors; FAM descriptors; functions; calls; interface calls; MPI collective calls; OpenFAM calls; indicators of one or more parts of memory area segments; data structures; mapping tables; mappings; mappings of virtual addresses to descriptors; results of searches or lookups in data structures or mapping tables; numbers of one or more parts of segments; indices; lists or arrays of indices; and lists or arrays of new values or data units.
[0107] Computer system 900 and content processing system 918 may include Fig. 9 For example, the content processing system 918 may also store instructions for performing the operations described above with respect to the following items: Figure 1 , FIG. 4A to FIG. 4B , FIG. 6A to FIG. 6B and Figure 7 communications in; FIG. 8A to FIG. 8B The operations depicted in the flowchart of ; and the instructions of the non-transitory computer-readable storage medium 1000 in Figure 1000.
[0108] Non-transitory computer-readable medium for facilitating efficient iterative collective operations using network attached storage
[0109] Fig.10 A non-transitory computer-readable medium 1000 for facilitating efficient iterative collective operations using network attached storage according to one aspect of the present application is illustrated. The storage medium 1000 may store instructions that, when executed by a computer, cause the computer to perform the methods, operations, and functions described herein. Specifically, the storage medium 1000 may store instructions 1002 for performing the following operations: receiving a first request to perform a collective operation by a first computing node, the first request indicating a first virtual address. The storage medium 1000 may store instructions 1004 for performing the following operations: storing a mapping of a descriptor of a first virtual address to a physical location of a memory area allocated for performing the collective operation in a data structure. The storage medium 1000 may store instructions 1006 for performing the following operations: performing a collective operation, the collective operation comprising writing a data unit to a memory area, and accessing a data unit from a memory area.
[0110] In addition, the storage medium 1000 may store instructions 1002 for performing the following operations: receiving a second request to perform an update operation, the second request indicating a second virtual address, one or more portions of the memory area segment to be updated, one or more indexes corresponding to the one or more portions, and one or more new values to be written to the one or more portions. The storage medium 1000 may store instructions 1008 for performing the following operations: searching the data structure based on the second virtual address to obtain a mapping of the first virtual address to the descriptor, the second virtual address matching the first virtual address. The storage medium 1000 may store instructions 1010 for performing the following operations: updating the one or more portions by writing the one or more new values to the one or more portions of the memory area segment at the indicated one or more indexes based on the obtained mapping.
[0111] The storage medium 1000 may store instructions 1002 for performing the following operations: receiving a third request to perform a collective operation, the third request indicating a third virtual address. The storage medium 1000 may store instructions 1008 for performing the following operations: searching the data structure based on the third virtual address to obtain a mapping of the first virtual address to the descriptor. The storage medium 1000 may store instructions 1012 for performing the following operations: based on the obtained mapping, determining that the third request is a subsequent iteration of the collective operation, and that the data unit associated with the collective operation has been stored and updated in the memory area. The storage medium 1000 may store instructions 1014 for performing the following operations: based on the obtained mapping, performing the collective operation by avoiding writing any data to the memory area and only reading data units from the memory area.
[0112] The storage medium 1000 may include Fig.10 For example, the storage medium 1000 may also store instructions for performing the operations described above with respect to the following items: Figure 1 , FIG. 4A to FIG. 4B , FIG. 6A to FIG. 6B and Figure 7 communications in; FIG. 8A to FIG. 8B The operations depicted in the flowchart of ; and Fig. 9 Instructions of the content processing system 918 in.
[0113] Aspects and variants
[0114] In general, the disclosed aspects provide a method, computer system, and non-transitory computer-readable storage medium for facilitating efficient iterative collective operations using network attached storage. In one aspect, the system receives a first request to perform a collective operation by a first computing node among a plurality of computing nodes, the first request indicating a first data unit to be written to a first segment of a memory area. The system generates a first virtual address of the memory area. The system allocates a memory area based on the first virtual address; the memory area is associated with a descriptor of a physical location of the memory. The system stores a mapping of the first virtual address to the descriptor in a data structure. The system performs a collective operation, the collective operation including at least writing the first data unit to the first segment and accessing data units from other segments of the memory area. The system receives a second request to perform an update operation, the second request indicating a second virtual address, one or more portions of the memory area segment to be updated, and corresponding data units to be written to the one or more portions. The system performs a first search in the data structure for a mapping of the second virtual address. The system identifies a mapping of the first virtual address to the descriptor based on the first search in response to the second virtual address matching the first virtual address. The system updates the one or more portions of the memory area segment by writing corresponding data units to the one or more portions based on the identified mapping.
[0115] In a variation of this aspect, the system receives a third request to perform the collective operation, the third request indicating a third virtual address. The system performs a second search in the data structure for a mapping for the third virtual address. The system identifies a mapping of the first virtual address to the descriptor based on the second search in response to the third virtual address matching the first virtual address. The system performs the collective operation, the collective operation further comprising at least avoiding writing any data to the memory region and accessing only data units from the memory region based on the identified mapping.
[0116] In a further variation of this aspect, the memory area includes a network attached storage. Figure 1 The storage servers 108-1 to 108-M describe network attached storage, and the storage areas may correspond to Figure 1 The memories 112-1 to 112-M. Figure 1 The FAM 106 and the FAM 608 of FIG. 6 may
[0117] In further variations, the collective operation is performed by a processor associated with the network attached storage. Figure 1 As described, the processor 110 - 1 of the memory server 108 - 1 and the processor 110 -M of the memory server 108 -M may perform collective operations.
[0118] In a further variation, the collective operation comprises a message passing interface (MPI) collective operation, as described above with respect to FIG. 3A to FIG. 3D as described.
[0119] In a further variation, the second request further indicates the following: a number of the one or more portions of the memory area segment to be updated; a list of indices corresponding to the one or more portions of the memory area segment to be updated, wherein a respective index corresponds to an ordered portion of the memory area segment; and a corresponding data unit to be written to the one or more portions, wherein a respective corresponding data unit includes a new value to be updated at a respective corresponding index. The second request may correspond to Figure 6B The update operation 660 (including the call 662 to MPI_Update_data). Figure 7 Calls 732 and 734 to MPI_Update_data describe an example of a second request.
[0120] In a further variation, each index in the index list corresponds to a respective data unit to be written to a respective portion of the memory area segment, as described above with respect to Figure 7 The calls to MPI_Update_data 732 and 734 are described.
[0121] In a further variation, the system receives, via a second computing node of the plurality of computing nodes, a fourth request to perform a collective operation, the fourth request indicating a second data unit to be written to a second segment of the memory area. The system performs the collective operation, the collective operation further comprising at least writing the second data unit to the second segment and accessing at least the first data unit from the first segment of the memory area. For example, as described above with respect to Fig. 6A As described, each of rank1-P 602 may call a collective operation (eg, MPI_Allreduce 612 function) as part of a first iteration 610, which may include both writing data to FAM 608 (eg, communications 616-620) and reading data from FAM 608 (eg, communications 628-634).
[0122] In another aspect, a computer system includes a processor and a storage device storing instructions, the instructions, when executed by the processor, cause the processor to perform a method, that is, the instructions are used to perform the following operations: receiving a first request to perform a collective operation by a first computing node among a plurality of computing nodes, the first request indicating a first virtual address and a first data unit to be written to a first segment of a memory area; performing a memory allocation process by allocating a memory area based on the first virtual address and determining a descriptor of a physical location of the memory; storing a mapping of the first virtual address to the descriptor in a data structure; performing the collective operation, the collective operation comprising at least writing the first data unit to the first segment and accessing data units from other segments of the memory area; receiving a second request to perform an update operation, the second request indicating a second virtual address, one or more portions of the memory area segment to be updated, and corresponding data units to be written to the one or more portions; searching the data structure based on the second virtual address to obtain a mapping of the first virtual address to the descriptor, the second virtual address matching the first virtual address; and updating the one or more portions by writing the corresponding data unit to the one or more portions of the memory area segment based on the obtained mapping. The instructions may further perform the operations described herein, including operations related to: Figure 1 , FIG. 4A to FIG. 4B , FIG. 6A to FIG. 6B and Figure 7 communications in; FIG. 8A to FIG. 8B The operations depicted in the flowchart of ; and the instructions of the non-transitory computer-readable storage medium 1000 in Figure 1000.
[0123] In yet another aspect, a non-transitory computer-readable storage medium stores instructions for performing the following operations: receiving a first request to perform a collective operation by a first computing node, the first request indicating a first virtual address; storing a mapping of the first virtual address to a descriptor of a physical location of a memory area allocated for performing the collective operation in a data structure; performing the collective operation, the collective operation comprising writing a first data unit to a first segment of the memory area and accessing data units from other segments of the memory area; receiving a second request to perform an update operation, the second request indicating a second virtual address, one or more portions of the memory area segment to be updated, one or more indexes corresponding to the one or more portions, and one or more new values to be written to the one or more portions; searching the data structure based on the second virtual address to obtain a mapping of the first virtual address to the descriptor, the second virtual address matching the first virtual address; and updating the one or more portions by writing the one or more new values to the one or more portions of the memory area segment at the indicated one or more indexes based on the obtained mapping. The instructions may further perform the operations described herein, including operations related to: Figure 1 , FIG. 4A to FIG. 4B , FIG. 6A to FIG. 6B and Figure 7 communications in; FIG. 8A to FIG. 8B The operations depicted in the flowchart of ; and Fig. 9 Instructions of the computer system 900 and the content processing system 918.
[0124] The foregoing description is presented to enable any person skilled in the art to make and use the various aspects and examples, and is provided in the context of a specific application and its requirements. Various modifications to the disclosed aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects and applications without departing from the spirit and scope of the present disclosure. Therefore, the various aspects described herein are not limited to the aspects shown, but are intended to conform to the maximum scope consistent with the principles and features disclosed herein.
[0125] In addition, the foregoing descriptions of various aspects have been presented for the purpose of illustration and description only. These descriptions are not intended to be exhaustive or to limit the various aspects described herein to the disclosed forms. Accordingly, many modifications and variations will be apparent to those of ordinary skill in the art. In addition, the above disclosure is not intended to limit the various aspects described herein. The scope of the various aspects described herein is defined by the appended claims.
Claims
1. A method comprising: receiving, by a first computing node of the plurality of computing nodes, a first request to perform a collective operation, the first request indicating a first data unit to be written to a first section of a memory region; generating a first virtual address of the memory area; allocating the memory region based on the first virtual address, the memory region being associated with a descriptor of a physical location of a memory; storing a mapping of the first virtual address to the descriptor in a data structure; performing the collective operation, the collective operation comprising at least writing the first data unit to the first segment and accessing data units from other segments of the memory area; receiving a second request to perform an update operation, the second request indicating a second virtual address, one or more portions of the memory region segment to be updated, and corresponding data units to be written to the one or more portions; performing a first search in the data structure for a mapping of the second virtual address; in response to the second virtual address matching the first virtual address, identifying a mapping of the first virtual address to the descriptor based on the first search; as well as Based on the identified mapping, the one or more portions are updated by writing the corresponding data units to the one or more portions of the memory area segment.
2. The method of claim 1, further comprising: receiving a third request to perform the collective operation, the third request indicating a third virtual address; performing a second search in the data structure for a mapping of the third virtual address; in response to the third virtual address matching the first virtual address, identifying a mapping of the first virtual address to the descriptor based on the second search; as well as The collective operation is performed, the collective operation further comprising at least avoiding writing any data to the memory region and accessing only data units from the memory region based on the identified mapping.
3. The method according to claim 1, in, The memory area includes a network attached storage.
4. The method according to claim 3, in, The collective operations are performed by a processor associated with the network attached storage.
5. The method according to claim 1, in, The collective operation comprises a message passing interface (MPI) collective operation.
6. The method of claim 1, wherein: The second request further indicates the following: a number of the one or more portions of the memory area segment to be updated; a list of indices corresponding to the one or more portions of the memory area segments to be updated, wherein respective indices correspond to ordered portions of the memory area segments; and Corresponding data units of the one or more portions are to be written, wherein the respective corresponding data units include new values to be updated at the respective corresponding indexes.
7. The method according to claim 6, in, The index in the list corresponds to a respective data unit to be written to a respective portion of the memory area segment.
8. The method of claim 1, further comprising: receiving, by a second computing node of the plurality of computing nodes, a fourth request to perform the collective operation, the fourth request indicating a second data unit to be written to a second section of the memory region; as well as The collective operation is performed, the collective operation further comprising at least writing the second data unit to the second segment and accessing at least the first data unit from the first segment of the memory area.
9. A computer system comprising: processor; as well as A storage device storing instructions for performing the following operations: receiving, by a first computing node of the plurality of computing nodes, a first request to perform a collective operation, the first request indicating a first virtual address and a first data unit to be written to a first segment of a memory region; performing a memory allocation process by allocating the memory region based on the first virtual address and determining a descriptor of a physical location of the memory; storing a mapping of the first virtual address to the descriptor in a data structure; performing the collective operation, the collective operation comprising at least writing the first data unit to the first segment and accessing data units from other segments of the memory area; receiving a second request to perform an update operation, the second request indicating a second virtual address, one or more portions of the memory region segment to be updated, and corresponding data units to be written to the one or more portions; searching a data structure to obtain a mapping of the first virtual address to the descriptor based on the second virtual address, the second virtual address matching the first virtual address; as well as Based on the obtained mapping, the one or more portions are updated by writing the corresponding data units to the one or more portions of the memory area segment.
10. The computer system of claim 9, wherein the instructions are further configured to perform the following operations: receiving a third request to perform the collective operation, the third request indicating a third virtual address and received after performing the collective operation indicated in the first request and the update operation indicated in the second request; searching the data structure to obtain a mapping of the first virtual address to the descriptor based on the third virtual address, the third virtual address matching the first virtual address; as well as The collective operation is performed, the collective operation further comprising at least avoiding writing any data to the memory region and accessing data units only from the memory region based on the obtained mapping.
11. The computer system of claim 9, in, The memory area includes a network attached storage.
12. The computer system of claim 11, in, The collective operations are performed by a processor associated with the network attached storage.
13. The computer system of claim 9, in, The collective operation comprises a message passing interface (MPI) collective operation.
14. The computer system of claim 9, wherein the second request further indicates the following: a number of the one or more portions of the memory area segment to be updated; a list of indices corresponding to the one or more portions of the memory area segments to be updated, wherein The respective indices correspond to ordered portions of the memory area segments; as well as Corresponding data units of the one or more portions are to be written, wherein the respective corresponding data units include new values to be updated at the respective corresponding indexes.
15. The computer system of claim 9, wherein the instructions are further configured to perform the following operations: receiving, by a second computing node of the plurality of computing nodes, a fourth request to perform the collective operation, the fourth request indicating a second data unit to be written to a second section of the memory region; and The collective operation is performed, the collective operation further comprising at least writing the second data unit to the second segment and accessing at least the first data unit from the first segment of the memory area.
16. A non-transitory computer-readable storage medium storing instructions for performing the following operations: Receiving, by a first computing node, a first request to perform a collective operation, the first request indicating a first virtual address; storing in a data structure a mapping of a descriptor of the first virtual address to a physical location of a memory region allocated for performing the collective operation; performing the collective operation, the collective operation comprising writing a first data unit to a first segment of the memory area and accessing data units from other segments of the memory area; receiving a second request to perform an update operation, the second request indicating a second virtual address, one or more portions of the memory area segment to be updated, one or more indexes corresponding to the one or more portions, and one or more new values to be written to the one or more portions; searching a data structure to obtain a mapping of the first virtual address to the descriptor based on the second virtual address, the second virtual address matching the first virtual address; as well as Based on the obtained mapping, the one or more portions are updated by writing the one or more new values to the one or more portions of the memory area segment at the indicated one or more indexes.
17. The non-transitory computer readable storage medium of claim 16, wherein the instructions are further configured to: After performing the collective operation indicated in the first request and the update operation indicated in the second request, receiving a third request to perform the collective operation, the third request indicating a third virtual address; searching the data structure based on the third virtual address to obtain a mapping of the first virtual address to the descriptor; determining, based on the obtained mapping, that the third request is a subsequent iteration of the collective operation and that data units associated with the collective operation have been stored and updated in the memory area; as well as Based on the obtained mapping, the collective operation is performed by avoiding writing any data to the memory area and only reading data units from the memory area.
18. The non-transitory computer readable storage medium of claim 16, The memory area includes a network attached storage, The network attached memory is associated with a processor that performs the collective operation, and The collective operation comprises a message passing interface (MPI) collective operation.
19. The non-transitory computer readable storage medium of claim 18, The network attached storage and associated processor comprise a memory server that performs computations offloaded by the first compute node.
20. The non-transitory computer readable storage medium of claim 16, The first request indicates that the first data unit is to be written to a first section of the memory area, and The instructions are further used to perform the following operations: receiving, by a second computing node, a fourth request to perform the collective operation, the fourth request indicating a second data unit to be written to a second section of the memory region; and The collective operation is performed, the collective operation further comprising at least writing the second data unit to the second segment and accessing at least the first data unit from the first segment of the memory area.
Citation Information
Patent Citations
Mapping storage of data in a dispersed storage network
US20140195846A1
Storage system using cloud based ranks as replica storage
US20190082009A1