A data aggregation method and apparatus
Patent Information
- Application Number
- CN202211128255.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-09-16
AI Technical Summary
由于每个计算设备在将各自的中间数据存储于内存时的排列方式的限制,该多个计算设备在进行数据聚合时,只能按照符合内存的排列方式的预设维度来聚合,而无法根据用户指定的维度来聚合,数据聚合方式不灵活
[0029]计算系统中第二计算设备实现上述第一方面或第一方面中的任意一种方法中第二计算设备执行的方法。
Smart Images

Figure CN117763376B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computing technology, and in particular to a data aggregation method and apparatus. Background Technology
[0002] Model training refers to providing a computing system with a large amount of training data so that the system can determine a suitable neural network architecture and the values of each parameter in the neural network architecture, thus obtaining a trained model. In this way, the neural network architecture can be used to more accurately identify or distinguish objects.
[0003] In practical applications, multiple computing devices can be combined into a computing system for training models. These devices include graphics processing units (GPUs), neural network processing units (NPUs), data processing units (DPUs), and tensor processing units (TPUs). Each of these computing devices can be input with different training data, or they can be used to train different sub-models of the model. After each iteration, these devices obtain their own intermediate data, which is then passed between them to obtain an aggregated result of all intermediate data for that iteration. Each device then uses this aggregated result as input for the next iteration. Through multiple rounds of iterative computation, these multiple computing devices can learn more crucial feature details, thus becoming more intelligent.
[0004] During model training, the intermediate data generated by computing devices is typically multi-dimensional. For example... Figure 1 The intermediate data generated by the exemplary computing device A has two dimensions: 3 data points in the first dimension and 2 data points in the second dimension, for a total of 6 data points. Due to limitations in how each computing device arranges its intermediate data in memory, these multiple computing devices can only aggregate data according to a preset dimension that conforms to the memory's arrangement, and cannot aggregate according to user-specified dimensions, resulting in an inflexible data aggregation method. Summary of the Invention
[0005] This application provides a data aggregation method and apparatus for enabling multiple computing devices to aggregate intermediate data according to user-specified dimensions during model training, thereby improving the flexibility of data aggregation methods.
[0006] In a first aspect, this application provides a data aggregation method applied to a computing system. The computing system includes R computing devices, each computing device including M data points, the M data points being represented by multiple dimensions, and each computing device having a number, where R and M are both integers greater than 1. The method includes: a first computing device among the R computing devices determining an aggregation dimension for the data, the aggregation dimension being one of the multiple dimensions of the M data points; the first computing device acquiring the M data points from each computing device respectively; and the first computing device aggregating the M data points from each computing device according to the aggregation dimension and in the order of the numbers of each computing device.
[0007] In the above technical solution, the first computing device determines the aggregation dimension of the data, and then aggregates the data according to the aggregation dimension in the order of the numbers of each computing device. In this way, multiple computing devices can aggregate intermediate data according to the user-specified dimension (i.e., the aggregation dimension), thereby improving the flexibility of the data aggregation method.
[0008] In one possible implementation, R computing devices are connected in a ring and aggregated for communication through an allgather operation using a message passing interface (MPI). In another possible implementation, a first computing device acquires M data points from each of the R computing devices, including: the first computing device acquiring M data points from the other R computing devices through a second computing device; wherein the second computing device is the preceding computing device in the ring connection.
[0009] In one possible implementation, the first computing device acquires M data points from one of the computing devices, including: the second computing device determining the number of data points sent each time during the process of sending the M data points to the first computing device based on the aggregation dimension; and the second computing device sending the M data points to the first computing device in multiple batches based on the number of data points sent each time.
[0010] In the above technical solution, the second computing device sequentially sends M data from a computing device other than the first computing device to the first computing device in multiple batches, and the number of data sent each time is determined based on the aggregation dimension. This enables the first computing device to directly aggregate the received data to obtain the aggregation result corresponding to the aggregation dimension without having to go through splitting and recombination operations, thereby speeding up the aggregation process.
[0011] In one possible implementation, the second computing device determines the number of data items sent each time during the process of sending M data items to the first computing device based on the aggregation dimension, including: the second computing device taking the number of consecutive data items of the M data items in the aggregation dimension as the number of data items sent each time.
[0012] In the above technical solution, the second computing device uses the number of consecutive data points of M data points in the aggregation dimension as the number of data points sent each time, so that the first computing device can directly aggregate the received data to obtain the aggregation result corresponding to the aggregation dimension without having to go through splitting and recombination operations, thereby speeding up the aggregation process.
[0013] In one possible implementation, the first computing device aggregates the M data points of each computing device according to the aggregation dimension and the serial number of each computing device, including: the first computing device determining the location in the first computing device where the M data points of each computing device are stored, based on the aggregation dimension and the serial number of each computing device; and aggregating the M data points of each computing device based on the location in the first computing device where the M data points of each computing device are stored.
[0014] In the above technical solution, after receiving M data from each computing device, the first computing device places the M data from each computing device in the position jointly indicated by the aggregation dimension and the number of each computing device to obtain the aggregation result corresponding to the aggregation dimension, without having to go through splitting and recombination operations, thus speeding up the aggregation process.
[0015] In one possible implementation, the first computing device determines the aggregation dimension of the data by: the first computing device receiving user parameters and obtaining the aggregation dimension from the user parameters.
[0016] In the above technical solution, users can specify the aggregation dimension used when aggregating data among multiple computing devices during model training, which helps to improve the flexibility of model training.
[0017] Secondly, this application provides a computing system comprising R computing devices, each computing device including M data points, the M data points being represented by multiple dimensions, and each computing device having a number, wherein R and M are both integers greater than 1; the R computing devices include a first computing device; the first computing device is used to: determine the aggregation dimension for the data, the aggregation dimension being one of the multiple dimensions of the M data points; acquire the M data points from each computing device respectively; and aggregate the M data points from each computing device according to the aggregation dimension and in the order of the numbers of each computing device.
[0018] In one possible implementation, R computing devices are connected in a ring and communicate in aggregate via an MPI allgather operation.
[0019] In one possible implementation, the R computing devices further include a second computing device; when the first computing device acquires M data points from each computing device, it specifically acquires M data points from other computing devices besides the first computing device through the second computing device. The second computing device is used to: determine the number of data points sent each time during the process of sending the M data points to the first computing device based on the aggregation dimension; and send the M data points to the first computing device in multiple batches based on the number of data points sent each time.
[0020] In one possible implementation, when the second computing device determines the number of data items sent each time during the process of sending M data items to the first computing device based on the aggregation dimension, it specifically uses the number of consecutive data items of the M data items in the aggregation dimension as the number of data items sent each time.
[0021] In one possible implementation, when the first computing device aggregates M data points from each computing device according to the aggregation dimension and the order of the device numbers, it specifically performs the following: determining the location in the first computing device where the M data points from each computing device are stored, based on the aggregation dimension and the device number; and aggregating the M data points from each computing device based on the location in the first computing device where the M data points from each computing device are stored.
[0022] In one possible implementation, the first computing device, when determining the aggregation dimension of the data, is specifically used to: receive user parameters and obtain the aggregation dimension from the user parameters.
[0023] In one possible implementation, one or more of the R computing devices are deployed on a server.
[0024] Thirdly, this application provides a computer-readable storage medium storing a computer program or instructions, which, when executed by a computing system,...
[0025] In a computing system, a first computing device implements the method executed by the first computing device in the first aspect or any one of the methods described above, and...
[0026] The second computing device in the computing system implements the method executed by the second computing device in the first aspect or any one of the methods in the first aspect.
[0027] Fourthly, this application provides a computer program product comprising a computer program or instructions, which, when executed by a computing system,
[0028] In a computing system, a first computing device implements the method executed by the first computing device in the first aspect or any one of the methods described above, and...
[0029] The second computing device in the computing system implements the method executed by the second computing device in the first aspect or any one of the methods in the first aspect.
[0030] The technical effects that can be achieved by any of the second to fourth aspects mentioned above can be referred to the description of the beneficial effects in the first aspect mentioned above, and will not be repeated here. Attached Figure Description
[0031] Figure 1 A schematic diagram of intermediate data with multiple dimensions generated by a computing device;
[0032] Figure 2 A schematic diagram of a neural network structure is provided in this application;
[0033] Figure 3 A schematic diagram of the aggregation communication algorithm provided in this application;
[0034] Figure 4 A schematic diagram of the architecture of a computing system provided in this application;
[0035] Figure 5 A schematic diagram of the architecture of another computing system provided in this application;
[0036] Figure 6 This application provides a schematic diagram of storing intermediate data in memory.
[0037] Figure 7 A schematic diagram illustrating the aggregation of multi-dimensional intermediate data provided in this application;
[0038] Figure 8 A flowchart illustrating a data aggregation method provided in this application;
[0039] Figure 9 A schematic diagram of a ring composed of multiple computing devices provided in this application;
[0040] Figure 10 A flowchart illustrating a data transmission method provided in this application;
[0041] Figure 11 A schematic diagram of a preset template provided for this application;
[0042] Figure 12 A schematic diagram illustrating how a computing device, as provided in this application, fills data to be aggregated into a preset template;
[0043] Figure 13 A schematic diagram illustrating an allgather aggregation method provided in this application;
[0044] Figure 14This is a schematic diagram of data transmission using the ring algorithm in an allgather aggregation method provided in this application. Detailed Implementation
[0045] To better explain the embodiments of this application, the relevant terms or technologies used in this application will be explained first:
[0046] I. Neural Networks
[0047] Neural networks (NNs) are mathematical models that mimic the behavioral characteristics of animal neural networks to perform distributed parallel information processing. They can process information by adjusting the connections between a large number of nodes within the neural network, and possess self-learning and adaptive capabilities.
[0048] Specifically, neural networks typically contain multiple interconnected layers, such as convolutional layers, fully connected layers (FC), activation layers, or pooling layers. Each layer can be expressed as a function y = f. w (x), where f is the function, f is derivative, w is the weight, x is the input, and y is the output.
[0049] Figure 2 This is a schematic diagram of a neural network structure, which may include m layers connected end-to-end, where m is an integer greater than or equal to 2. The 0th layer of the neural network can be expressed as a function f0, with x as the input, y0 as the output, and w0 as the weight; the 1st layer of the neural network can be expressed as a function f1, with y0 as the input, y1 as the output, and w1 as the weight, and so on.
[0050] II. Model Training
[0051] Suppose there exists a data set {(x0, l0), ..., (x... n-1 , l n-1 )}, where x0, ..., x n-1 There are n inputs, and corresponding l0, ..., l n-1 These are the expected outputs from the n inputs, often referred to as labels. Each (x... j , l j A sample dataset is defined as any input (which can be represented as x) in this dataset. j Enter as follows Figure 2 In the m layers of a neural network, the output of the neural network can be obtained. The output of the neural network can be represented as... The goal of model training is to solve for w0, ..., w m-1 So that under the loss function L, y jm-1 and l j The closest.
[0052] Furthermore, the solution process can utilize the stochastic gradient descent (SGD) method. The SGD method can include forward propagation and backward propagation methods.
[0053] Forward propagation method: Given any input from the dataset (which can be represented as x),... j The input is fed into function f0, and function f0 outputs the result. Then Input is given to function f1, and function f1 outputs... And so on, we obtain functions f0 to f m-1 The corresponding outputs are, respectively Then combine with x j The corresponding l j And the loss is calculated using the loss function L.
[0054] Backpropagation: Using the chain rule, calculate y for each layer sequentially. j gradient Δy j w j gradient Δw j Specifically, for example, through loss and y m-1 Determine the gradient Δy of the (m-1)th layer m-1 Then according to Δy m-1 and w m-1 Determine the gradient Δw of the (m-1)th layer m-1 ; and so on, to obtain Δy and Δw for each layer, that is, Δy0, Δw0, ..., Δy m-1 Δw m-1 .
[0055] III. Collective Communication
[0056] A computing system may include multiple computing devices, such as GPUs, NPUs, DPUs, and TPUs. These multiple computing devices are used to jointly train a model. Specifically, each of these multiple computing devices can perform the entire model training process using its own local data (which can be understood as data parallelism), or each of these multiple computing devices can perform a portion of the model training process using its own local data (which can be understood as model parallelism).
[0057] In each iteration of model training, each computing device can obtain its own intermediate data for that iteration. For example, the intermediate data for each computing device may include one or more of the following: features (or activations), gradients, and model parameters obtained from local model training. Features may be, for example, features learned from the training data by the model; model parameters may be, for example, the parameters of a function in a neural network; and gradients may be, for example, the gradient generated during backpropagation. j The difference Δw j Then, each computing device aggregates its local intermediate data with intermediate data from other computing devices via aggregation communication to obtain an aggregation result, which is then used for model training in the next iteration. The intermediate data obtained by each computing device in each iteration can be referred to as the data to be aggregated.
[0058] Common aggregation communication operations include MPI allreduce and MPI allgather. Hereinafter, MPI allreduce will be referred to as allreduce, and MPI allgather as allgather.
[0059] Specifically, `allreduce` processes the data to be aggregated across all computing devices according to a specified mapping function to obtain the final aggregation result, and then distributes the final aggregation result back to all computing devices. Figure 3 In example (a), the computing system includes computing devices 1 to 3, and the data to be aggregated corresponding to computing devices 1 to 3 are data 1 to data 3, respectively. The aggregation results obtained by computing devices 1 to 3 are all calculated based on data 1 to data 3. `allgather` performs the concatenation operation on all computing devices and distributes the aggregation results to all computing devices. Figure 3 In example (b), the computing system includes computing devices 1 to 3, and the data to be aggregated corresponding to computing devices 1 to 3 are data 1 to data 3 respectively. The aggregation results obtained by computing devices 1 to 3 are all obtained by concatenating data 1 to data 3.
[0060] like Figure 4 This application provides a schematic diagram of the architecture of a computing system.
[0061] The computing system includes client device 10, server 20, and computing device 205. Server 20 is a common computer device. Users can input data to server 20 through client device 10. Client device 10 is a terminal device, including but not limited to personal computers, mobile phones, tablets, or smart cars.
[0062] Server 20 includes an input / output (I / O) interface 204, a processor 201, and a memory 202. The I / O interface 204 is used to communicate with devices located outside the server 20. For example, client device 10 inputs data and sends model training tasks to server 20 through the I / O interface 204. After processing the input data, server 20 sends the processed output results back to client device 10 through the I / O interface 204.
[0063] Processor 201 is the computing and control core of server 20. It can be a central processing unit (CPU) or other specific integrated circuits. Processor 201 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In practical applications, server 20 can be configured with multiple processors 201. Processor 201 includes one or more processor cores. An operating system and other software programs are installed in processor 201, enabling processor 201 to access memory 202 and various high-speed serial computer expansion bus standard (Public Component Interconnect Express, PCIe) devices.
[0064] The processor 201 is connected to the memory 202 via a double data rate (DDR) bus or other types of bus. The memory 202 is the main memory of the server 20. The memory 202 is typically used to store various running software in the operating system, input data received from the client device 10, and output results to be sent to the client device 10 in the future. To improve the access speed of the processor 201, the memory 202 needs to have a high access speed. In traditional computer devices, dynamic random access memory (DRAM) is typically used as the memory 202. Besides DRAM, the memory 202 can also be other random access memories, such as static random access memory (SRAM). Alternatively, the memory 202 can also be read-only memory (ROM). For example, a read-only memory could be a programmable read-only memory (PROM) or an erasable programmable read-only memory (EPROM). This embodiment does not limit the number or type of memory 202.
[0065] Optionally, for persistent data storage, the computing system also includes a data storage system 203, which may be located outside the server 20 (e.g., Figure 4 As shown, the data storage system 203 exchanges data with the server 20 via a network. Alternatively, the data storage system 203 can be located inside the server 20 and exchange data with the processor 201 via the PCIe bus 206. In this case, the data storage system 203 functions as a hard disk.
[0066] The computing device 205 is used to perform model training tasks. The processor 201 sends the received model training task and input data to the computing device 205. The computing device 205 performs model training based on the input data and sends the processing result back to the processor 201 after completing the model training task. Figure 4 As shown, computing device 205 can be directly plugged into a slot on the motherboard of server 20 and exchanges data with processor 201 via PCIe bus 206. It can be understood that computing device 205 is located within server 20, or that server 20 includes a host and computing device 205. The host includes processor 201, memory 202, I / O interface 204, and PCIe bus 206, and computing device 205 exchanges data with processor 201 in the host via PCIe bus 206. It should also be noted that... Figure 4 The PCIe bus 206 in the computer device 205 can be replaced with a bus that uses the Compute Express Link (CXL), Universal Serial Bus (USB) protocol, or other protocols to enable data transfer.
[0067] In large-scale model training, the computing system may include multiple computing devices 205, which may be located in the same server 20 or in different servers 20.
[0068] like Figure 5 This is a schematic diagram of another computing system architecture. The computing system exemplarily shows three computing devices 205, of which two computing devices 205 are located in the same server 20, and the third computing device 205 is located in a different server 20. Furthermore, each computing device 205 may include a processing unit 2051 and memory 2052. The processing unit 2051 is used for model computation, and the memory 2052 is used to store the computation results obtained from the model computation. The computation results may be, for example, intermediate data obtained in each iteration during the model computation process, i.e., data to be aggregated.
[0069] Furthermore, two computing devices 205 located in the same server 20 can communicate via a high-speed internal link within the server 20. This high-speed internal link communication may be based on the System Direct Memory Access (SDMA) protocol. Alternatively, data can be directly copied between the two computing devices 205.
[0070] Two computing devices 205 located in different servers 20 can communicate via a switch 30. Specifically, each computing device 205 also includes a network interface card (NIC), and each computing device 205 connects to the switch 30 through its respective NIC, thereby enabling communication with computing devices 205 located in other servers 20 (i.e., communication between the two servers 20) through the NIC and the switch 30. Optionally, the communication between the two servers 20 can be based on the remote direct memory access (RDMA) protocol.
[0071] It should be noted that when the computing device 205 stores the data to be aggregated in memory 2052, it needs to arrange it according to a preset sorting method. Furthermore, when the data to be aggregated is multi-dimensional, it needs to be stored sequentially in memory 2052 according to a preset dimension. For example, if the data to be aggregated in computing device 205 is two-dimensional, this two-dimensional data refers to... Figure 1 As shown, computing device 205 aggregates the data to be aggregated according to... Figure 6 The information is stored in memory 2052 of computing device 205 in the manner shown.
[0072] When performing allgather data aggregation on multiple computing devices, due to the limitations of the arrangement of the data to be aggregated in memory, the multiple computing devices can only aggregate according to the preset dimensions that conform to the arrangement in memory, and cannot aggregate according to the dimensions specified by the user, making the data aggregation method inflexible.
[0073] In one possible approach, each computing device can obtain a user-specified dimension (i.e., the aggregation dimension). When performing data aggregation, each computing device first aggregates the data according to a preset dimension that conforms to the memory arrangement to obtain the aggregation result. Then, each computing device performs a split operation and a concat operation on the aggregation result according to the user-specified aggregation dimension to obtain the aggregation result corresponding to the aggregation dimension.
[0074] For example, the aggregated data A to C corresponding to computing devices A to C can be found in [reference needed]. Figure 7 As shown in (a), each piece of data to be aggregated has two dimensions. Furthermore, computing devices A through C need to aggregate the data according to the second dimension (i.e., the aggregation dimension is the second dimension). When computing devices A through C store their respective data to be aggregated in memory, their arrangement is as follows: Figure 7 As shown in (b). Taking computing device A as an example, when computing device A aggregates data with other computing devices, it first aggregates according to the first dimension to obtain the following... Figure 7 The aggregation result is shown in (c). Computing device A then performs... Figure 7 The aggregation result in (c) is split and recombined to obtain, as shown below. Figure 7 The aggregation result corresponding to the second dimension is shown in (d). The operations in computing devices B and C are similar to those in computing device A, and will not be described again. This results in significant aggregation latency and wasted computing resources.
[0075] This application provides a data aggregation method applicable to the aforementioned computing system, which includes R computing devices that need to perform aggregation communication. Furthermore, each computing device includes M data points to be aggregated, which can be represented by K dimensions. Here, R, M, and K are integers greater than 1.
[0076] For example, R computing devices participate in the training process of an image recognition model. In one iteration of model training, each computing device generates data to be aggregated, which corresponds to three dimensions. These three dimensions can be image color, image length, and image width. The image color dimension includes 3 data points, the image length dimension includes 256 data points, and the image width dimension includes 256 data points.
[0077] See Figure 8 The exemplary flowchart of the data aggregation method illustrates the execution actions of the first computing device among R computing devices in the data aggregation method, wherein the first computing device is any one of the R computing devices.
[0078] Step 801: The first computing device determines the aggregation dimension of the data, which is one of the K dimensions of the M data.
[0079] In this context, the aggregation dimension refers to the dimension that the first computing device follows when aggregating data with other computing devices. Specifically, this aggregation dimension is the k-th dimension out of K dimensions, where k is a positive integer less than or equal to K. For example, if the aggregation dimension is the first dimension (k=1), then the first computing device needs to aggregate the data to be aggregated with the data to be aggregated with the data from other computing devices according to the first dimension; if the aggregation dimension is the second dimension (k=2), then the first computing device needs to aggregate the data to be aggregated with the data from other computing devices according to the second dimension.
[0080] Optionally, the first computing device further determines the data shape of the M data points. The data shape can be used to indicate the total number of dimensions K corresponding to the M data points, and to indicate the distribution of the M data points in each dimension. For example, if the data shape of the data to be aggregated is 3×2, then the data to be aggregated includes a total of 6 data points, and the total number of dimensions K corresponding to these 6 data points is 2. Specifically, there are 3 data points in the first dimension and 2 data points in the second dimension. As another example, if the data shape of the data to be aggregated is 2×3×2, then the data to be aggregated includes a total of 12 data points, and the total number of dimensions K corresponding to these 12 data points is 3. Specifically, there are 2 data points in the first dimension, 3 data points in the second dimension, and 2 data points in the third dimension.
[0081] In one possible implementation, when issuing a model training task, the user can define the data shape and aggregation dimensions of the data to be aggregated during model training. Combined with... Figure 4In the example, the user inputs user parameters in the client device 10, which include data shape and aggregation dimension. Accordingly, the client device 10 inputs the user parameters to the processor 201 through the I / O interface 204 in the server 20. The processor 201 then inputs the user parameters to the computing device 205 through the PCIe bus 206. In this way, the computing device 205 can obtain the data shape and aggregation dimension.
[0082] Optionally, after determining the aggregation dimension and data shape, the first computing device determines the data group size based on the aggregation dimension and data shape. The data group size indicates the number of data points (m) in each group when the first computing device discretizes M data points into I groups, where I and m are both positive integers, and I × m = M.
[0083] It can be understood that the first computing device sends a packet of data to the next computing device each time, that is, it sends m data packets to the next computing device each time. Further, m is the number of consecutive data points in the k-th dimension when the first computing device stores the M data packets in its memory. Combined with... Figure 1 and Figure 6 For example, when the aggregation dimension k=1, the number of consecutive data points in the first dimension is equal to 3 when 6 data points are stored in memory, so m=3; when the aggregation dimension k=2, the number of consecutive data points in the second dimension is equal to 6 when 6 data points are stored in memory, so m=6.
[0084] In one possible implementation, the first computing device can specifically determine the data group size according to the expression m = shape[k] × shape[k+1] … × shape[K], where m is the data group size, k is the aggregation dimension, and shape[k] is the number of data in the k-th dimension. For example, if the data shape of the data to be aggregated is 2 × 3 × 2, then M = 12, K = 3. Furthermore, when the aggregation dimension k = 2, m = shape[2] × shape[3] = 3 × 2 = 6. For another example, if the data shape of the data to be aggregated is 2 × 3 × 2, then M = 12, K = 3. Furthermore, when the aggregation dimension k = 3, m = shape[3] = 2.
[0085] It should be added that the first computing device can also determine the number of data transmissions (repeat) based on the data shape and aggregation dimension. For example, the first computing device determines the number of data transmissions based on the expression repeat = shape[1] × shape[2] ... × shape[k], where repeat is the number of data transmissions, k is the aggregation dimension, and shape[k] is the number of data items in the k-th dimension. For example, if the data shape of the data to be aggregated is 2 × 3 × 2, when the aggregation dimension k = 2, repeat = shape[1] × shape[2] = 2 × 3 = 6; as another example, if the data shape of the data to be aggregated is 2 × 3 × 2, when the aggregation dimension k = 3, repeat = shape[1] × shape[2] × shape[3] = 12.
[0086] This data transmission count is used by the first computing device to determine whether the aggregation communication is complete. It can be understood that when the first computing device determines that the number of data transmissions has reached this count, it indicates that this round of aggregation communication has been completed. An explanation of this data transmission count can be found below. Figure 13 The description in the relevant embodiments.
[0087] Step 802: The first computing device acquires M data points from each computing device.
[0088] Optionally, the R computing devices are connected in a ring, and each computing device has its own identification information within the ring. Specifically, the identification information can be the serial number or identifier of the computing device within the ring.
[0089] For example, the aggregation communication system includes three computing devices, referred to as computing devices A through C, and the ring formed by computing devices A through C is as follows: Figure 9 As shown, computing device A can be considered to be numbered 1 in the ring, computing device B to be numbered 2 in the ring, and computing device C to be numbered 3 in the ring.
[0090] Furthermore, in one round of aggregation, the preceding computing device sends its data to be aggregated to the following computing device; that is, computing device A sends its data to be aggregated to computing device B, computing device B sends its data to be aggregated to computing device C, and computing device C sends its data to be aggregated to computing device A. Alternatively, it can be understood that the following computing device obtains the data to be aggregated from all other computing devices except itself through the preceding computing device; that is, computing device A obtains the data to be aggregated from computing devices B and C through computing device C, computing device B obtains the data to be aggregated from computing devices A and C through computing device A, and computing device C obtains the data to be aggregated from computing devices A and B through computing device B.
[0091] In this application, the computing device located before the first computing device among the R computing devices is designated as the second computing device. (Combined with...) Figure 9 In this context, when the first computing device is computing device A, the second computing device is computing device C. The first computing device can obtain M data points from other computing devices besides itself through the second computing device. Correspondingly, the second computing device sends the M data points from the other computing devices to the first computing device. The following example illustrates how the second computing device sends M data points from any one computing device to the first computing device; for a detailed process, please refer to [link to documentation]. Figure 10 An exemplary schematic diagram of the data transmission process is shown.
[0092] Step 1001: The second computing device determines the aggregation dimension.
[0093] The implementation method of the second computing device determining the aggregation dimension can be found in the description of the first computing device determining the aggregation dimension in step 801. The term "first computing device" can be replaced with "second computing device" for better understanding.
[0094] Step 1002: The second computing device determines the number of data items sent each time (i.e., the data group size m) during the process of sending M data items from the second computing device to the first computing device based on the aggregation dimension.
[0095] The second computing device also acquires the data shape and determines the data group size based on the aggregation dimension and the data shape. For details on how the second computing device determines the data group size, please refer to the description of the first computing device determining the data group size in step 801. You can replace "first computing device" with "second computing device" for better understanding.
[0096] Optionally, the second computing device also determines the number of data transmissions based on the data shape and aggregation dimension. This number of data transmissions is used by the second computing device to determine whether the aggregation communication is complete. For details, please refer to the description of the first computing device determining the number of data transmissions in step 801. You can replace "first computing device" with "second computing device" to understand this.
[0097] Step 1003: The second computing device sends M data points to the first computing device in multiple transmissions, based on the number of data points sent each time. This can be understood as the number of times the second computing device sends M data points to the first computing device, specifically equal to the ratio of M to m. When M = 12 and m = 2, the number of transmissions is 6.
[0098] When the second computing device sends M data to the first computing device, the second computing device can first determine the preset template of the second computing device based on the data group size and the identification information of the second computing device in R computing devices; then fill the M data in the memory of the second computing device into the M preset positions of the preset template, and send the M data located in the M preset positions to the first computing device in multiple times.
[0099] The following explains the two processes: determining a preset template based on the second computing device and sending M data points from the second computing device to the first computing device.
[0100] Step 1: The second computing device determines the preset template.
[0101] In one possible implementation, when determining the preset template, the second computing device can first determine R×M preset positions, and then select M padding positions from the R×M preset positions according to the data group size and the identification information of the second computing device, thereby obtaining the preset template. The M padding positions are used to sequentially place M data items. Further, the other positions in the R×M preset positions besides the M padding positions can be called non-padding positions, which include invalid data, such as initialized values. For example, when the second computing device selects M padding positions from the R×M preset positions, it specifically forms M padding positions from the selected m×R×(i-1)+(r-1)×m+1 to m×R×(i-1)+r×m preset positions, where r is the identification information (e.g., sequence number) of the second computing device in the R second computing devices, 1≤r≤R, and i is an integer in the range [1, 1].
[0102] For example, an aggregated communication system includes Figure 9The computing devices A through C shown have a data shape of 2×3×2, with the third aggregation dimension, i.e., R=3, M=12, k=3. Each of computing devices A through C first determines the data grouping size m=2 based on the data shape of 2×3×2 and the aggregation dimension k=3. The details are as follows:
[0103] When the second computing device is computing device A, computing device A first determines R×M = 36 preset positions, and then determines the 1st and 2nd, 7th and 8th, 13th and 14th, 19th and 20th, 25th and 26th, and 31st and 32nd preset positions from these 36 preset positions as 12 fill positions in the preset template. Furthermore, the other 24 positions among the 36 preset positions, excluding these 12 fill positions, can be used as non-fill positions. The preset template determined by computing device A can be found in [reference needed]. Figure 11 As shown in (a).
[0104] When the second computing device is computing device B, computing device B first determines R×M = 36 preset positions, and then determines the 3rd and 4th, 9th and 10th, 15th and 16th, 21st and 22nd, 27th and 28th, and 33rd and 34th preset positions from these 36 preset positions as 12 fill positions in the preset template. Furthermore, the other 24 positions among the 36 preset positions, excluding these 12 fill positions, can be used as non-fill positions. The preset template determined by computing device B can be found in [reference needed]. Figure 11 As shown in (b).
[0105] When the second computing device is computing device C, computing device C can first determine R×M = 36 preset positions, and then determine the 5th and 6th, 11th and 12th, 17th and 18th, 23rd and 24th, 29th and 30th, and 35th and 36th preset positions from these 36 preset positions as 12 fill positions in the preset template. Furthermore, the other 24 positions among the 36 preset positions, excluding these 12 fill positions, can be used as non-fill positions. The preset template determined by computing device C can be found in [reference needed]. Figure 11 As shown in (c).
[0106] Understandably, before performing model training, the second computing device can determine a preset template based on the data shape, aggregation dimension, and its identification information in the ring. Subsequently, the second computing device fills the preset template with the generated M data points. Furthermore, in the first iteration of model training, the second computing device can also determine the preset template based on the data shape, aggregation dimension, and its identification information in the ring, and then fill the preset template with the M data points generated in the first iteration. Further, in subsequent iterations, the second computing device can also fill the preset template with the generated M data points.
[0107] In process 2, the second computing device sends M data points to the first computing device.
[0108] Step a: The second computing device fills M data points from the second computing device into M preset positions in the preset template.
[0109] For example, the second computing device, according to a preset template, fills the (i-1)×m+1 to i×m data points from the M data points into the m×R×(i-1)+(r-1)×m+1 to m×R×(i-1)+r×m fill positions, where i is an integer in the range [1, 1]. Since the preset template also includes (R-1)×M non-fill positions (where the fill positions include invalid values), the second computing device obtains R×M data after filling the M data points into the M preset positions of the preset template.
[0110] In one possible implementation, the second computing device includes two buffers, which can be denoted as an input buffer and an output buffer, respectively. Further, the second computing device can store M data generated by itself in the input buffer, and then store the resulting R×M data (after filling the M data into M preset positions of a preset template) in the output buffer. It can be understood that the amount of data in the output buffer is R times the amount of data in the input buffer.
[0111] Combination Figure 12 The illustrative diagram illustrates how computing device A fills 12 data points into 12 preset positions of a preset template. Computing device A is a second computing device, and the preset template for computing device A can be found in [reference needed]. Figure 11As shown in (a), computing device A obtains 12 data points (denoted as A1 to A12) from the input buffer, and then sequentially fills these 12 data points into a preset template to obtain 36 data points. Computing device A then stores these 36 data points into the output buffer. When the second computing device is computing device B, the method by which computing device B sequentially fills the 12 data points into its corresponding preset template, and when the second computing device is computing device C, the method by which computing device C sequentially fills the 12 data points into its corresponding preset template, are similar to those of computing device A and will not be described further.
[0112] Step b: The second computing device sends M data points at M preset positions in the preset template.
[0113] Optionally, the second computing device sends M data points from the M preset locations to the first computing device. The second computing device and the first computing device may be located on the same server or on different servers.
[0114] For example, when both the second computing device and the first computing device are located in the first server, the second computing device can send M data items at M preset locations to the first computing device via the high-speed internal link of the first server. Alternatively, the second computing device can directly copy the M data items at the M preset locations to the first computing device.
[0115] As another example, when the second computing device is located on the second server and the first computing device is located on the first server, the second computing device and the first computing device can communicate through a switch. Specifically, the second computing device transmits M data points from M preset locations to the first computing device on the first server via the switch.
[0116] Furthermore, when the second computing device sends M data points at M preset positions in a preset template to the first computing device, it can be implemented in two ways as follows:
[0117] In implementation method 1, the second computing device sends all M×R data to the first computing device at once.
[0118] In implementation method 2, the second computing device first sends the first set of data at the M preset positions in the preset template to the first computing device, then sends the second set of data at the M preset positions in the preset template to the first computing device, then sends the third set of data at the M preset positions in the preset template to the first computing device, and so on, until the second computing device sends the I set of data at the M preset positions in the preset template to the first computing device. That is, the second computing device sends all the M data at the M preset positions in the preset template to the first computing device in I parts (the number of transmissions is equal to I). The offset between the previous set of data and the next set of data is (R-1)×m, and each set includes m data.
[0119] Step 803: The first computing device aggregates the M data points of each computing device according to the aggregation dimension and the order of the numbers of each computing device.
[0120] In one specific implementation, the first computing device determines the location within the first computing device where M data points from each computing device are stored, based on the aggregation dimension and the device number. Subsequently, the first computing device aggregates the M data points from each computing device based on their respective locations within the first computing device.
[0121] In one specific implementation, the first computing device also determines a preset template based on the data group size and its identification information (such as serial number) among the R computing devices. The implementation method for the first computing device to determine the preset template is similar to the implementation method for the second computing device to determine the preset template described above; "second computing device" can be replaced with "first computing device" for illustrative purposes. Furthermore, since the first and second computing devices have different numbers, the preset template determined by the second computing device is different from that determined by the first computing device. Subsequently, the first computing device can also fill its M data points into the M preset positions of its preset template.
[0122] Furthermore, since the first computing device is a computing device located after the second computing device in the ring, the padding bits in the preset template obtained by the second computing device exactly correspond to the non-padding bits in the preset template obtained by the first computing device. Thus, the first computing device can replace the invalid data in the non-padding bits of the preset template obtained by the first computing device with M data from the second computing device. The first computing device performs a concatenation operation on the M data from the second computing device and the M data from the first computing device.
[0123] Combination Figure 11For example, when the second computing device is computing device A and the first computing device is computing device B, the padding bits in the preset template obtained by computing device A correspond to the non-padding bits in the preset template obtained by computing device B. That is, after receiving M data from computing device A, computing device B can replace the M invalid data in the non-padding bits of its preset template with the received M data. Computing device B performs a concatenation operation on the M data from computing device A and the M data from computing device B.
[0124] In this application, one of the R computing devices located after the first computing device is designated as the third computing device. The third computing device also has a pre-defined template, and M data points from the third computing device are filled into its pre-defined template, thereby storing M×R data points in the output buffer. Combined with... Figure 9 In the context of the ring, when the first computing device is computing device B, the third computing device is computing device C. After the first computing device fills M data points into M preset positions of its preset template, it can also send those M data points located in the M preset positions of the preset template to the third computing device. Similarly, after receiving the M data points from the first computing device, the third computing device can perform a concatenation operation based on the M×R data points in its output buffer and the received M data points. Specifically, the third computing device can replace the M invalid data points in the non-filled positions of the preset template obtained by the third computing device with the M data points from the first computing device.
[0125] To better understand the embodiments of this application, the following is combined with Figure 13 The example illustrates how the first, second, and third computing devices achieve allgather data aggregation through ring transmission.
[0126] Among them, the second computing device, the first computing device, and the third computing device can be respectively Figure 13 The three computing devices are A, B, and C. The output buffers of computing devices are: A (A), B (B), and C (C). Figure 13 As shown.
[0127] In the first step of the ring algorithm:
[0128] It should be noted that each computing device can send a set of data to the next computing device in one round of data transmission, and each computing device's output buffer contains 6 sets of data. Therefore, the first step of the ring consists of a total of 6 rounds of data transmission.
[0129] Reference Figure 14 Explanation of the first round of data transmission in the ring algorithm:
[0130] Computing device A sends the first set of data (i.e., A1 and A2 located at the first and second positions) in output buffer A to computing device B. Correspondingly, computing device B receives A1 and A2 from computing device A and replaces the two invalid data at the first and second positions in output buffer B with A1 and A2.
[0131] Computing device B sends the first set of data (i.e., B1 and B2 located at the 3rd and 4th positions) in output buffer B to computing device C. Correspondingly, computing device C receives B1 and B2 from computing device B and replaces the two invalid data at the 3rd and 4th positions in output buffer C with B1 and B2.
[0132] Computing device C sends the first set of data (i.e., C1 and C2 located at the 5th and 6th positions) in output buffer C to computing device A. Correspondingly, computing device A receives C1 and C2 from computing device C and replaces C1 and C2 with two invalid data at the 5th and 6th positions in output buffer A.
[0133] In the second round of data transmission, computing device A sends the second set of data (A3 and A4 at positions 7 and 8) from output buffer A to computing device B; computing device B sends the second set of data (B3 and B4 at positions 9 and 10) from output buffer B to computing device C; and computing device C sends the second set of data (C3 and C4 at positions 11 and 12) from output buffer C to computing device A. For each computing device, the offset between the data sent in the later round and the data sent in the previous round is (R-1)×m.
[0134] The data transmission methods for rounds 3 to 6 are similar to those for rounds 1 and 2, and will not be repeated here.
[0135] The aggregation results obtained by computing devices A through C in the first step of the ring algorithm can be found in [reference]. Figure 13 As shown.
[0136] In step 2 of the ring algorithm:
[0137] It should be noted that the second step of the ring algorithm also includes 6 rounds of data transmission.
[0138] Let's take the first round of data transmission as an example:
[0139] Computing device A sends C1 and C2 located at positions 5 and 6 to computing device B. Computing device B receives C1 and C2 from computing device A and replaces C1 and C2 with two invalid data at positions 5 and 6 in output buffer B.
[0140] Computing device B sends A1 and A2 located at the first and second positions to computing device C. Computing device C receives A1 and A2 from computing device B and replaces A1 and A2 with two invalid data at the first and second positions in output buffer C.
[0141] Computing device C sends B1 and B2 located at the 3rd and 4th positions to computing device A. Computing device A receives B1 and B2 from computing device C and replaces B1 and B2 with two invalid data at the 3rd and 4th positions in output buffer A.
[0142] The data transmission methods for rounds 2 through 6 are similar to those for round 1, and will not be repeated here.
[0143] The aggregation results obtained by computing devices A through C in step 2 of the ring algorithm can be found in [reference]. Figure 13 As shown.
[0144] It should be added that the first step of the ring algorithm mentioned above includes 6 rounds of data transmission, and the second step of the ring algorithm also includes 6 rounds of data transmission. That is, when computing devices A through C perform data aggregation based on the third dimension, a total of 12 rounds of data transmission are involved. In a possible application scenario, computing device A, based on the data shape of 2×3×2 and the aggregation dimension k=3, determines the number of data transmission rounds (repeat=12). Once computing device A determines that it has completed 12 rounds of data transmission, it can determine that the current data aggregation is complete. Similarly, computing devices B and C can also determine the number of data transmission rounds (repeat=12), and thus, after determining that they have completed 12 rounds of data transmission, they can determine that the current data aggregation is complete.
[0145] In the above technical solution, the first computing device determines the aggregation dimension of the data, and then aggregates the data according to the aggregation dimension in the order of the device numbers. This allows multiple computing devices to aggregate intermediate data according to the user-specified dimension (i.e., the aggregation dimension), improving the flexibility of the data aggregation method. Furthermore, in the ring, the preceding computing device (e.g., the second computing device) sequentially sends M data points from one computing device (excluding the following computing device, e.g., the first computing device) to the following computing device in multiple batches. The number of data points sent each time is determined based on the aggregation dimension. This allows the following computing device to directly aggregate the received data to obtain the aggregation result corresponding to the aggregation dimension without further splitting and reorganization operations, thus accelerating the aggregation speed.
[0146] Based on the above content and the same concept, this application provides a computing system, which includes R computing devices, each computing device including M data, the M data being represented by multiple dimensions, and each computing device having a number, wherein R and M are both integers greater than 1; the R computing devices include a first computing device.
[0147] The first computing device is used to: determine the aggregation dimension of the data, which is one of the multiple dimensions of the M data; obtain the M data from each computing device; and aggregate the M data from each computing device according to the aggregation dimension and in the order of the device numbers.
[0148] In one possible implementation, R computing devices are connected in a ring and communicate in aggregate via an MPI allgather operation.
[0149] In one possible implementation, the R computing devices further include a second computing device; when the first computing device acquires M data points from each computing device, it specifically acquires M data points from other computing devices besides the first computing device through the second computing device. The second computing device is used to: determine the number of data points sent each time during the process of sending the M data points to the first computing device based on the aggregation dimension; and send the M data points to the first computing device in multiple batches based on the number of data points sent each time.
[0150] In one possible implementation, when the second computing device determines the number of data items sent each time during the process of sending M data items to the first computing device based on the aggregation dimension, it specifically uses the number of consecutive data items of the M data items in the aggregation dimension as the number of data items sent each time.
[0151] In one possible implementation, when the first computing device aggregates M data points from each computing device according to the aggregation dimension and the order of the device numbers, it specifically performs the following: determining the location in the first computing device where the M data points from each computing device are stored, based on the aggregation dimension and the device number; and aggregating the M data points from each computing device based on the location in the first computing device where the M data points from each computing device are stored.
[0152] In one possible implementation, the first computing device, when determining the aggregation dimension of the data, is specifically used to: receive user parameters and obtain the aggregation dimension from the user parameters.
[0153] In one possible implementation, one or more of the R computing devices are deployed on a server.
[0154] Based on the above content and the same concept, this application provides a computer-readable storage medium storing a computer program or instructions. When the computer program or instructions are executed by a computing system, a first computing device in the computing system implements the method executed by the first computing device in the above method embodiments, and a second computing device in the computing system implements the method executed by the second computing device in any one of the above method embodiments.
[0155] Based on the above content and the same concept, this application provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a computing system, a first computing device in the computing system implements the method executed by the first computing device in any of the above method embodiments, and a second computing device in the computing system implements the method executed by the second computing device in any of the above method embodiments.
[0156] It is understood that the various numerical designations used in the embodiments of this application are merely for descriptive convenience and are not intended to limit the scope of the embodiments of this application. The order of the process numbers described above does not imply the order of execution; the execution order of each process should be determined by its function and internal logic.
[0157] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of protection of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A data aggregation method, characterized in that, The method is applied to a computing system, which includes R computing devices, each computing device including M data, the M data being represented in multiple dimensions, and each computing device having a number, where R and M are both integers greater than 1; The method includes: The first computing device among the R computing devices determines the aggregation dimension of the data, and the aggregation dimension is one of the multiple dimensions of the M data. The first computing device acquires M data points from each of the computing devices; The first computing device aggregates M data points from each computing device according to the aggregation dimension and in the order of the device numbers; The R computing devices are connected in a ring and aggregated communication through the allgather operation; The first computing device aggregates M data points from each computing device according to the aggregation dimension and in the order of the device's serial number, including: The first computing device determines the location within the first computing device where M data points from each computing device are stored, based on the aggregation dimension and the number of each computing device. The first computing device stores M data points from each computing device in a location within the first computing device, and aggregates the M data points from each computing device.
2. The method as described in claim 1, characterized in that, The R computing devices are connected in a ring and aggregated for communication via an allgather operation, including: The R computing devices are connected in a ring and perform aggregated communication through the MPI allgather operation.
3. The method as described in claim 2, characterized in that, The first computing device acquires M data points from each computing device, including: The first computing device acquires M data points from the other computing devices besides the first computing device through the second computing device among the R computing devices; The first computing device acquires M data points from one of the computing devices, including: The second computing device determines the number of data items sent each time during the process of sending the M data items to the first computing device based on the aggregation dimension; The second computing device sends the M data points to the first computing device in multiple batches, based on the number of data points sent each time.
4. The method as described in claim 3, characterized in that, The second computing device determines the number of data items sent each time during the process of sending the M data items to the first computing device based on the aggregation dimension, including: The second computing device uses the number of consecutive data points of the M data points on the aggregation dimension as the number of data points sent each time.
5. The method according to any one of claims 1-4, characterized in that, The first computing device determines the aggregation dimensions of the data, including: The first computing device receives user parameters and obtains the aggregation dimension from the user parameters.
6. A computing system, characterized in that, The computing system includes R computing devices, each computing device includes M data points, the M data points are represented in multiple dimensions, and each computing device has a number, where R and M are both integers greater than 1; the R computing devices include a first computing device. The R computing devices are connected in a ring and aggregated communication through the allgather operation; The first computing device is used for: Determine the aggregation dimension for the data, where the aggregation dimension is one of the multiple dimensions of the M data; Obtain M data points from each computing device respectively; The M data points from each computing device are aggregated according to the aggregation dimension and in the order of the device numbers. When the first computing device aggregates M data points from each computing device according to the aggregation dimension and in the order of the device numbers, it specifically performs the following: Based on the aggregation dimension and the number of each computing device, determine the location of M data points from each computing device stored in the first computing device; Based on the location of M data points from each computing device stored in the first computing device, the M data points from each computing device are aggregated.
7. The computing system as described in claim 6, characterized in that, The R computing devices are connected in a ring and aggregated for communication via an allgather operation, including: The R computing devices are connected in a ring and perform aggregated communication through the MPI allgather operation.
8. The computing system as described in claim 7, characterized in that, The R computing devices also include a second computing device; When the first computing device acquires M data points from each computing device, it specifically performs the following tasks: The second computing device acquires M data points from other computing devices besides the first computing device. The second computing device is used for: The number of data items sent each time during the process of sending the M data items to the first computing device is determined based on the aggregation dimension. Based on the number of data items sent each time, the M data items are sent to the first computing device in multiple batches.
9. The computing system as described in claim 8, characterized in that, When the second computing device determines the number of data items sent each time during the process of sending the M data items to the first computing device based on the aggregation dimension, it is specifically used for: The number of consecutive data points of the M data points along the aggregation dimension is taken as the number of data points sent each time.
10. The computing system as described in any one of claims 6-9, characterized in that, When determining the aggregation dimension of the data, the first computing device is specifically used for: Receive user parameters and obtain the aggregation dimension from the user parameters.
11. The computing system as described in any one of claims 6-10, characterized in that, One or more of the R computing devices are deployed on a server.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer programs or instructions that, when executed by a computing system, will... The first computing device in the computing system implements the method executed by the first computing device in any one of claims 1 to 5, and, The second computing device in the computing system implements the method executed by the second computing device in the method of claim 3 or 4.
Citation Information
Patent Citations
Data processing system and data processing method
CN111886593A