A data processing apparatus, method, and related device
By introducing state machines and multipliers into the distributed training system, the processing flow of gradient data is optimized, the problem of frequent communication between computing nodes and external devices is solved, and the training efficiency is improved.
Patent Information
- Application Number
- CN202510028546.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-01-08
AI Technical Summary
In a distributed training system, each computing node needs to communicate frequently with external devices to obtain gradient data, resulting in large communication overhead and low training efficiency.
The data processing device is adopted, including a first state machine, a second state machine, a third state machine and a multiplier, through which the reading, dimensionality reduction and storage of gradient data are coordinated within the computing node, reducing communication with external devices.
提高了训练效率,减少了通信开销,使得计算节点可以直接从本地存储位置调用降维结果。
Smart Images

Figure CN119443194B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of distributed training, and particularly to a data processing device, method, and related equipment. Background Art
[0002] In the field of natural language processing, large language models based on deep learning are often used. In the process of processing natural language based on a deep learning model, the text to be processed is often input into the Embedding layer of the deep learning model, and each word vector of the text to be processed is obtained based on the word vector embedding matrix of the Embedding layer. Then, based on the word vectors, the output result is obtained based on the large language model.
[0003] The accuracy of obtaining each word vector of the text to be processed based on the word vector embedding matrix of the Embedding layer directly determines the accuracy of the output result of the large language model. Therefore, it is necessary to train the parameters of the word vector embedding matrix of the Embedding layer. However, the number of parameters in the Embedding layer is large, and it is difficult for a single computing node to store the complete parameters of the Embedding layer. Therefore, currently, a distributed training method is often used to distribute the complete parameters of the Embedding layer among multiple computing nodes, and multiple computing nodes perform parallel training.
[0004] In the prior art, when multiple computing nodes perform parallel training on the Embedding layer, it is necessary to use an external device to calculate the gradients of the weight parameters of the Embedding layer deployed on each node based on the output result of this round using the gradient descent algorithm, and send the gradients of the weight parameters of the Embedding layer deployed on each node to each computing node for each computing node to update its respective weight parameters until the gap between the output result and the target result reaches the minimum. In each round of updating the weight parameters, each computing node needs to communicate with the external device to request and obtain the gradient data for updating the weight parameters in this round. Therefore, it requires a huge communication overhead and low efficiency. Summary of the Invention
[0005] The main objective of this application is to propose a data processing device, method, and related equipment, aiming to solve the problem that when parallel training of a deep learning model is currently performed, each computing node needs to communicate with an external device to obtain the gradient data for updating the weight parameters in this round, resulting in a large communication overhead and low efficiency.
[0006] To achieve the above object, the present application proposes a data processing device, including: being applied to a distributed training system, the distributed training system includes a plurality of computing nodes and a memory module, each of the computing nodes is used to deploy partial operators of the model to be trained, and the memory module is used to store gradient data when the weight parameters of the operators deployed by each of the computing nodes are updated;
[0007] The data processing device includes: a first state machine, a second state machine, a third state machine, and a multiplier-accumulator. The first state machine, the second state machine, the third state machine, and the multiplier-accumulator are used to be deployed on any one of the computing nodes, and the first state machine, the second state machine, the third state machine, and the multiplier-accumulator are deployed on the same computing node;
[0008] The first state machine is configured to: in response to receiving a weight parameter update instruction, determine, based on the weight parameter update instruction, the target computing node information for updating the weight parameter in the weight parameter update instruction and the storage address of the gradient data corresponding to the target computing node in the memory module; and, based on the storage address, send a read request to the memory module, and send the target computing node information to the second state machine;
[0009] The second state machine is configured to: receive the target computing node information and the gradient data returned by the memory module; and, send the gradient data to the multiplier-accumulator, and send the target computing node information to the third state machine;
[0010] The multiplier-accumulator is configured to: perform dimensionality reduction on the gradient data to obtain a dimensionality reduction result, and send the dimensionality reduction result to the third state machine;
[0011] The third state machine is configured to: wrap the dimensionality reduction result based on the target computing node information, and write it to the storage location corresponding to the target computing node for the target computing node to call.
[0012] In an embodiment of the present application, when the weight parameter update instruction contains a plurality of target computing node information and corresponding storage addresses of a plurality of gradient data, the first state machine is further configured to: respectively send read requests to the memory module based on each storage address, and send a first identifier to the second state machine before the gradient data of each read request flows back, where the first identifier is used to identify the start position of the return data corresponding to each read request; and
[0013] In the order of the request sequence of each read request, sequentially send the target computing node information corresponding to each read request to the second state machine.
[0014] In an embodiment of the present application, the second state machine is further configured to: sequentially send each first identifier and the corresponding return data to the multiplier-accumulator; and send each of the target computing node information to the third state machine;
[0015] The multiplier-accumulator is further configured to: in response to receiving a first identifier, accumulate the data after the first identifier, and when receiving the next first identifier, output the current accumulation result as the dimensionality reduction result to the third state machine.
[0016] The third state machine is further configured to: respectively wrap the corresponding dimensionality reduction result based on each target computing node information, and write the wrapped dimensionality reduction results into their corresponding storage locations.
[0017] In an embodiment of the present application, the first state machine is further configured to: when all the gradient data corresponding to the storage addresses included in the weight parameter update instruction have been returned, send a second identifier to the second state machine, where the second identifier is used to identify that all the gradient data corresponding to the weight parameter update instruction have been returned;
[0018] The second state machine is further configured to: after sending all the first identifiers and the corresponding return data to the multiplier-accumulator, send the second identifier to the multiplier-accumulator;
[0019] The multiplier-accumulator is further configured to: in response to receiving the second identifier, end the dimensionality reduction operation, and send the current accumulation result to the third state machine.
[0020] In an embodiment of the present application, the first state machine is further configured to: when the gradient data in the storage address corresponding to any read request is empty, send a third identifier to the second state machine, where the third identifier is used to identify that the return data of the current read request is non-gradient data;
[0021] The second state machine is further configured to: sequentially send the third identifier and the corresponding return data to the multiplier-accumulator;
[0022] The multiplier-accumulator is further configured to: in response to receiving the third identifier, skip all the data after the third identifier until receiving other identifiers.
[0023] The present application also provides a data processing method, which is applied to the data processing device in any of the above embodiments. The data processing method includes:
[0024] In response to receiving a weight parameter update instruction, based on the first state machine, determine the target computing node information for updating the weight parameter in the weight parameter update instruction and the storage address of the gradient data corresponding to the target computing node in the memory module; and, based on the storage address and the first state machine, send a read request to the memory module, and send the target computing node information to the second state machine based on the first state machine;
[0025] Based on the second state machine, receive the target computing node information and the gradient data returned by the memory module; and, based on the second state machine, send the gradient data to the multiplier - adder, and send the target computing node information to the third state machine;
[0026] Based on the multiplier - adder, perform dimensionality reduction on the gradient data to obtain a dimensionality reduction result, and send the dimensionality reduction result to the third state machine;
[0027] Based on the third state machine and the target computing node information, wrap the dimensionality reduction result and write it to the storage location corresponding to the target computing node for the target computing node to call.
[0028] In an embodiment of the present application, when the weight parameter update instruction contains multiple target computing node information and corresponding multiple gradient data storage addresses, the data processing method further includes:
[0029] Based on the first state machine, send read requests to the memory module respectively based on each storage address, and send a first identifier to the second state machine before the gradient data of each read request flows back, where the first identifier is used to identify the starting position of the returned data corresponding to each read request;
[0030] Based on the first state machine, in the order of the request sequence of each read request, send the target computing node information corresponding to each read request to the second state machine in sequence.
[0031] In an embodiment of the present application, the data processing method further includes:
[0032] Based on the second state machine, send each first identifier and the corresponding returned data to the multiplier - adder in sequence; and based on the second state machine, send each target computing node information to the third state machine;
[0033] Based on the multiplier - adder, accumulate the data after the first identifier, and when receiving the next first identifier, output the current accumulated result as the dimensionality reduction result to the third state machine.
[0034] In an embodiment of the present application, the data processing method further includes:
[0035] Based on the third state machine, wrap the corresponding dimensionality reduction results based on each target computing node information respectively, and write the wrapped dimensionality reduction results into their corresponding storage locations respectively.
[0036] In the embodiment of the present application, the data processing method further includes:
[0037] After all the gradient data corresponding to the storage addresses included in the weight parameter update instruction have flowed back, send a second identifier to the second state machine based on the first state machine, where the second identifier is used to identify that all the gradient data corresponding to the weight parameter update instruction have flowed back;
[0038] After all the first identifiers and the corresponding returned data are sent to the multiplier-accumulator, send the second identifier to the multiplier-accumulator based on the second state machine;
[0039] When the multiplier-accumulator receives the second identifier, end the dimensionality reduction operation ended by the multiplier-accumulator, and send the current accumulated result to the third state machine.
[0040] In the embodiment of the present application, the data processing method further includes:
[0041] When the gradient data in the storage address corresponding to any read request is empty, send a third identifier to the second state machine based on the first state machine, where the third identifier is used to identify that the returned data of the current read request is non-gradient data;
[0042] Based on the second state machine, send the third identifier and the corresponding returned data to the multiplier-accumulator in sequence;
[0043] When the multiplier-accumulator receives the third identifier, skip all the data after the third identifier by the multiplier-accumulator until other identifiers are received.
[0044] The embodiment of the present application also proposes a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the data processing method described in any one of the above is implemented.
[0045] The embodiment of the present application also proposes a computing device, the computing device includes a processor, and when the processor executes a computer program stored in a memory, the data processing method described in any one of the above embodiments is implemented.
[0046] In the embodiments of the present application, by setting the first to third state machines and the multiplier-accumulator, the first state machine is used to send a read request to the memory module according to a weight parameter update instruction, the second state machine is used to receive the returned gradient data, the multiplier-accumulator reduces the dimension of the gradient data to obtain a dimensionality reduction result, and the third state machine wraps the dimensionality reduction result and stores it at the storage location corresponding to each computing node. Therefore, only the computing nodes deploying the first state machine, the second state machine, the third state machine and the multiplier-accumulator need to communicate with external devices, while the other computing nodes do not need to communicate with external devices and can directly call the dimensionality reduction result from their respective corresponding storage locations. During the training process, the first state machine can continuously send read requests based on consecutive storage addresses, the second state machine can continuously receive the returned data, the multiplier-accumulator can continuously reduce the dimension of the returned gradient data, and the third state machine can continuously store the wrapped dimensionality reduction result at the storage location corresponding to each computing node, thereby improving the training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0048] Figure 1 It is a schematic flowchart of a distributed training system in the prior art;
[0049] Figure 2 It is a module diagram of a data processing device in an embodiment of the present application;
[0050] Figure 3 It is a schematic flowchart of the data processing device in an embodiment of the present application when processing a weight parameter update instruction;
[0051] Figure 4 It is a step diagram of a data processing method in an embodiment of the present application;
[0052] Figure 5 It is a module diagram of a computer-readable storage medium in an embodiment of the present application;
[0053] Figure 6 It is a module diagram of a computing device in an embodiment of the present application.
[0054] The implementation, functional features and advantages of the objectives of the present application will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The principles and spirit of the present application will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present application, rather than limiting the scope of the present application in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0056] Those skilled in the art know that the embodiments of the present application can be implemented as a system, device, method, or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0057] According to the embodiments of the present application, a data processing device, method, and related equipment are proposed.
[0058] In this article, it should be understood that the number of any elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0059] The principles and spirit of the present application will be elaborated below with reference to several representative embodiments of the present application.
[0060] Exemplary Device
[0061] As Figure 1 shown, Figure 1 is a distributed training system in the prior art. Among them, the distributed training system includes computing nodes 1-6. Assume that the model to be trained is the Embedding layer of a deep learning model. The trainable parameters in the Embedding layer are mainly the word vector embedding matrix (operator), as Figure 1 shown, taking Taking the matrix of as an example, it is distributed on computing nodes 1-6, and each computing node is only responsible for updating the parameters of a part of the word vector embedding matrix. During training, the training sample X is input into computing nodes 1-6. Computing nodes 1-6 respectively obtain six word embedding vectors X1-X6 based on the part of the word vector embedding matrix deployed on each of them. After merging the six word embedding vectors, the total word embedding vector Y is obtained and output to the Transformer model, which can output the prediction result Z based on the word embedding vector Y. After obtaining the prediction result Z, using an external device (such as other computing nodes dedicated to calculating gradients), according to the error between the prediction result Z and the expected result, the loss function algorithm can be used to calculate the descending gradients of the weight parameters of each computing node respectively, and then the descending gradients of the weight parameters of each computing node are sent to each computing node respectively for each computing node to update its own weight parameters. After multiple rounds of iteration, until the error between the prediction result output by the Transformer model and the expected result reaches the minimum, the training ends.
[0062] During this process, in each iteration round, each computing node needs to request and obtain the gradient data of this round from the external device, and there are often many computing nodes in the distributed training system ( Figure 1 only for exemplary illustration). In each round, each computing node needs to communicate with the external device, so there will be a huge communication overhead, and it also limits the training efficiency.
[0063] Based on this, the embodiment of the present application proposes a data processing device to solve the problem that the communication overhead between each computing node of the distributed training system and the external device for calculating gradient data is relatively large and the training efficiency is relatively low when obtaining gradient data.
[0064] As Figure 2 shown, in the embodiment of the present application, the data processing device is applied to a distributed training system. The distributed training system includes multiple computing nodes and a memory module. Each computing node is used to deploy part of the operators of the model to be trained, and the memory module is used to store the gradient data when the weight parameters of the operators deployed by each computing node are updated;
[0065] The data processing device includes: a first state machine, a second state machine, a third state machine and a multiplier-accumulator. The first state machine, the second state machine, the third state machine and the multiplier-accumulator are used to be deployed on any one of the computing nodes, and the first state machine, the second state machine, the third state machine and the multiplier-accumulator are deployed on the same computing node;
[0066] The first state machine is configured to: in response to receiving a weight parameter update instruction, determine the target computing node information for updating the weight parameter in the weight parameter update instruction and the storage address of the gradient data corresponding to the target computing node in the memory module based on the weight parameter update instruction; and send a read request to the memory module based on the storage address and send the target computing node information to the second state machine;
[0067] The second state machine is configured to: receive the target computing node information and the gradient data returned by the memory module; and send the gradient data to the multiplier-accumulator and send the target computing node information to the third state machine;
[0068] The multiplier-accumulator is configured to: perform dimensionality reduction on the gradient data to obtain a dimensionality reduction result and send the dimensionality reduction result to the third state machine;
[0069] The third state machine is configured to: wrap the dimensionality reduction result based on the target computing node information and write it to the storage location corresponding to the target computing node for the target computing node to call.
[0070] Among them, the distributed training system can be Figure 1 the training system shown in the figure, or any other distributed training system with multiple computing nodes. The data processing device in the embodiments of the present application can be deployed on any computing node of the distributed training system.
[0071] Operators of the model to be trained are deployed on each computing node of the distributed training system. For example, partial word embedding vector matrices of the Embedding layer as shown in the figure are deployed on each computing node. The gradient data of the respective weight parameters of each computing node in each iteration round is calculated by an external device as shown in the figure. The specific calculation method of the gradient data of each computing node in each iteration round in the present application is not limited. Figure 1 Figure 1 Figure 1 In the embodiments of the present application, the first to third state machines are all composed of state registers and combinational logic circuits, and can perform state transitions according to control signals according to preset states, and are control centers for coordinating the actions of relevant signals and completing specific operations.
[0072] As shown in the figure
[0073] As Figure 3 shown Figure 3This is a flowchart of data processing in an embodiment of the present application. Among them, in each iteration round, the gradient data required for updating the weight parameters of the operators deployed on each computing node is calculated by an external device and stored in a memory module. The memory module can be the local memory of the computing nodes deployed by the data processing device in the embodiment of the present application. The weight parameter update instruction contains the information of the target computing node whose weight parameters need to be updated, and the storage address of the gradient data required for the target computing node to update the weight parameters in the memory module. The target computing node information can be the number of the target computing node, etc. The weight parameter update instruction is generally sent by the host in the distributed training system. The host pre-writes the information of the target computing node whose weight parameters need to be updated and the storage address of the corresponding gradient data into the weight parameter update instruction. When the first state machine receives the weight parameter update instruction, it can parse the weight parameter update instruction to determine the information of the target computing node whose weight parameters need to be updated and the storage address of the corresponding gradient data.
[0074] After determining the information of the target computing node and the storage address of the corresponding gradient data, the first state machine can send a read request to the memory module based on this storage address to request to read the gradient data at this storage address. When the memory module receives the read request, it returns the gradient data at this storage address to the second state machine according to the requested storage address. In addition, the first state machine also sends the target computing node information to the second state machine so that the second state machine can determine the ownership of the gradient data returned from the memory module.
[0075] When the second state machine receives the returned gradient data from the memory module and the target computing node information from the first state machine, it sends the returned gradient data to the multiplier-accumulator and sends the target computing node information to the third state machine. Among them, the gradient data is stored in the memory module in the form of data packets included. A data packet includes a data header, that is, multiple sub-data. When the gradient data is returned, it is also returned in the form of each sub-data. Therefore, when the second state machine sends it to the multiplier-accumulator, it also sends it to the multiplier-accumulator in the form of each sub-data. After receiving each sub-data in the gradient data, the multiplier-accumulator accumulates them to obtain a dimensionality reduction result and sends the dimensionality reduction result to the third state machine.
[0076] After receiving the dimensionality reduction result, the third state machine wraps the dimensionality reduction result according to the target computing node. The wrapping form can form a data packet in the form of a data header and data content. Among them, the data header can indicate the computing node and length to which the dimensionality reduction result belongs. After wrapping the dimensionality reduction result, the third state machine can store it in the storage location corresponding to the target computing node, so as to facilitate the target computing node to directly call it from this storage location.
[0077] In the embodiments of the present application, by setting the first to third state machines and the multiplier-accumulator, the first state machine is used to send a read request to the memory module according to the weight parameter update instruction, the second state machine is used to receive the returned gradient data, the multiplier-accumulator reduces the dimension of the gradient data to obtain a dimension-reduced result, and the third state machine wraps the dimension-reduced result and stores it at the storage locations corresponding to each computing node. Therefore, only the computing nodes deploying the first state machine, the second state machine, the third state machine, and the multiplier-accumulator need to communicate with external devices, while the other computing nodes do not need to communicate with external devices and can directly call the dimension-reduced result from their respective corresponding storage locations. During the training process, the first state machine can continuously send read requests based on consecutive storage addresses, the second state machine can continuously receive the returned data, the multiplier-accumulator can continuously reduce the dimension of the returned gradient data, and the third state machine can continuously store the wrapped dimension-reduced result at the storage locations corresponding to each computing node, thereby improving the training efficiency.
[0078] In the embodiments of the present application, a weight parameter update instruction may include gradient data requests of multiple computing nodes. For example, in the embodiments of the present application, taking Figure 1 the distributed training system shown as an example, it is possible that computing nodes 1-6 all need to update the weight parameters, that is, computing nodes 1-6 all need to obtain the gradient data for updating the weight parameters. At this time, the first state machine can parse the weight parameter update instruction to obtain six target computing nodes and six storage addresses. The six target computing nodes are respectively computing nodes 1-6. Assuming the six storage addresses are respectively addr1-addr6, at this time, the first state machine respectively sends six read requests to the memory module based on addr1-addr6 to respectively request to read the gradient data corresponding to computing nodes 1-6. After receiving the six read requests, the memory module respectively returns the six gradient data corresponding to the six read requests to the second state machine.
[0079] It should be noted that the read request and the gradient data are in one-to-one correspondence. Whenever the first state machine sends a read request to the memory module, a first identifier is sent to the second state. The first identifier is used to separate the gradient data corresponding to each read request. Among them, the first identifier can be, for example, the binary code 10.
[0080] In the embodiment of the present application, assume that the six read requests corresponding to the read addresses addr1 - addr6 are read1 - read6, and the corresponding reflux data of read1 - read6 are data1 - data6 respectively. Then, the second state machine will receive in sequence: 10, data1, 10, data2, 10, data3, 10, data4, 10, data5, 10, data6. At the same time, the second state machine will send them to the multiplier - adder in the order of 10, data1, 10, data2, 10, data3, 10, data4, 10, data5, 10, data6. Each time the multiplier - adder receives a binary code 10, it will output the previous accumulation result. If there is no accumulation result currently, it will not output.
[0081] For example, for "10, data1, 10, data2, 10, data3, 10, data4, 10, data5, 10, data6", when the multiplier - adder receives the first 10, there is no accumulation result currently, so it does not output. After receiving the first 10, it is the data in data1. At this time, the multiplier - adder will accumulate the data in data1. When receiving the second 10, it proves that all the data in data1 have been accumulated. At this time, the accumulation result is the dimensionality reduction result corresponding to data1, and it can be output to the third state machine. Then, the multiplier - adder continues to receive the data in data2 and accumulates them until receiving the third binary code 10. At this time, the output of the accumulation result is the dimensionality reduction result corresponding to data2, and it can be output to the third state machine. And so on, 6 dimensionality reduction results corresponding to data1 - data6 can be obtained respectively and output to the third state machine one after another.
[0082] In addition, it should be noted that after parsing the weight parameter update instruction to obtain the target calculation information, the first state machine can send the target calculation node information to the third state machine in the order of read1 - read6, or send the target calculation node information to the second state machine in the order of read1 - read6, and then the second state machine forwards it to the third state machine in sequence. Among them, the multiplier - adder outputs six dimensionality reduction results in the order of data1 - data6, which respectively correspond to the six target calculation node information received by the third state machine in sequence, such as the number information of calculation nodes 1 - 6. Then, based on the order of the received dimensionality reduction results and according to the order of the target calculation nodes, the third state machine can wrap and store each dimensionality reduction result. For example, when storing, wrap the dimensionality reduction result of data1 and store it in the storage location corresponding to calculation node 1, wrap the dimensionality reduction result of data2 and store it in the storage location corresponding to calculation node 2, and so on, wrap the dimensionality reduction result of data6 and store it in the storage location corresponding to calculation node 6.
[0083] In the embodiments of the present application, when a weight parameter update instruction includes multiple target computing nodes, by setting a delimiter, the gradient data corresponding to each target computing node can be separated, so as to avoid errors in the multiplier-accumulator.
[0084] Continue to refer to Figure 3 In the embodiments of the present application, the first state machine is further configured to: when the gradient data corresponding to each storage address included in the weight parameter update instruction has all flowed back, send a second identifier to the second state machine, where the second identifier is used to identify that the gradient data corresponding to the weight parameter update instruction has all flowed back;
[0085] The second state machine is further configured to: after sending all the first identifiers and the corresponding backflow data to the multiplier-accumulator, send the second identifier to the multiplier-accumulator;
[0086] The multiplier-accumulator is further configured to: in response to receiving the second identifier, end the dimensionality reduction operation and send the current accumulation result to the third state machine.
[0087] For example, still taking the computing nodes 1-6 shown in Figure 1 as an example, the target computing nodes in the weight parameter update instruction include computing nodes 1-6 and storage addresses addr1-addr6. When the first state machine sends read1-read6 to the memory module respectively based on addr1-addr6, when the data6 corresponding to read6 flows back, the first state machine also sends a second identifier, such as the binary code 11, to the second state machine. At this time, the data received by the second state machine is: 10, data1, 10, data2, 10, data3, 10, data4, 10, data5, 10, data6, 11, and it sends them to the multiplier-accumulator in sequence. When the multiplier-accumulator receives "10, data1, 10, data2, 10, data3, 10, data4, 10, data5, 10, data6, 11" in sequence, the processing process from the first binary code 10 to data6 is the same as that of the above embodiment and will not be elaborated here one by one; when the multiplier-accumulator receives the binary code 11, it means that all the gradient data of this batch has been dimensionally reduced, that is, all the gradient data corresponding to all the target computing nodes in this weight parameter update instruction has been dimensionally reduced. At this time, after the multiplier-accumulator outputs the current accumulation result, that is, the dimensionality reduction result of data6, to the third state machine, it can clear all the data to prepare for processing the next weight parameter update instruction.
[0088] In the embodiment of the present application, the first state machine is further configured to: when the gradient data in the storage address corresponding to any read request is empty, send a third identifier to the second state machine, where the third identifier is used to identify that the return data of the current read request is non-gradient data;
[0089] The second state machine is further configured to: sequentially send the third identifier and the corresponding return data to the multiplier-accumulator;
[0090] The multiplier-accumulator is further configured to: in response to receiving the third identifier, skip all the data after the third identifier until other identifiers are received.
[0091] In a distributed training system, there may be a situation where the stored data is empty at the storage location corresponding to the storage address in the weight parameter update instruction. For example, still taking Figure 1Taking the computing nodes 1-6 shown as an example, the target computing nodes in the weight parameter update instruction include computing nodes 1-6 and storage addresses addr1-addr6. Assume that the data in the storage location corresponding to addr4 is empty, that is, there is no corresponding gradient data for the read request of read4. At this time, the backflow data corresponding to read4 is not the gradient data corresponding to computing node 4. It may be a preset identifier representing empty data, such as the binary code 00. At this time, when the storage location corresponding to the storage address is empty, the first state machine sends the third identifier, such as the binary code 01, to the second state machine to indicate that the backflow data corresponding to the read request is not gradient data. At this time, the backflow data received by the second state machine is: 10, data1, 10, data2, 10, data3, 01, 00, 10, data5, 10, data6, 11, and it is sent to the multiplier-accumulator in sequence. When the multiplier-accumulator receives the first binary code 10, there is no accumulated data currently. Starting from the first binary code 10, the subsequent data of data1 are accumulated. When the second binary code 10 is received, the current accumulated result, that is, the dimensionality reduction result of data1, is output to the third state machine, and the subsequent data of data2 are accumulated. When the third binary code 10 is received, the current accumulated result, that is, the dimensionality reduction result of data2, is output to the third state machine, and the subsequent data of data3 are accumulated. When the binary code 11 is received, the current accumulated result, that is, the dimensionality reduction result corresponding to data3, is output to the third state machine, and the subsequent 00 data is skipped, that is, no dimensionality reduction operation is performed, until the fifth binary code 10 is received. When the fifth binary code 10 is received, since no accumulation operation has been performed currently, a preset identifier can be output to the third state machine at this time. For example, the binary code 00 is output to the third state machine to represent that there is no dimensionality reduction result currently, and the subsequent data of data5 are accumulated. When the sixth binary code 10 is received, the current accumulated result, that is, the dimensionality reduction result of data5, is output to the third state machine, and the subsequent data of data6 are accumulated. When 11 is received, it represents that all the gradient data of the current batch have been processed. The current accumulated result is the dimensionality reduction result of data6. After outputting it to the third state machine, all data can be cleared.
[0092] For the third state machine, the dimensionality reduction results of data1-data3 and data5-data6 can be wrapped and stored based on their respective corresponding target computing node information. For the identifier 00 corresponding to data4, which represents that there is no dimensionality reduction result of the gradient data corresponding to computing node 4, there is no need to wrap and store it at this time.
[0093] In addition, in the embodiments of the present application, multiple storage buckets can be preset in the local memory of the computing nodes where the data processing device is deployed. Each storage bucket corresponds to a different computing node. When the third state machine stores the dimension-reduced results corresponding to each computing node after wrapping, it can store them in the corresponding storage bucket, which is convenient for each computing node to call from the corresponding storage bucket.
[0094] In the embodiments of the present application, by setting the first to third state machines and the multiplier-accumulator, the first state machine is used to send a read request to the memory module according to the weight parameter update instruction, the second state machine is used to receive the backflow gradient data, the multiplier-accumulator reduces the dimension of the gradient data to obtain the dimension-reduced result, and the third state machine wraps the dimension-reduced result and stores it at the storage location corresponding to each computing node. Therefore, only the computing nodes deploying the first state machine, the second state machine, the third state machine and the multiplier-accumulator need to communicate with external devices, while the other computing nodes do not need to communicate with external devices and can directly call the dimension-reduced results from their respective corresponding storage locations. During the training process, the first state machine can continuously send read requests based on the storage address, the second state machine can continuously receive the backflow data, the multiplier-accumulator can continuously reduce the dimension of the backflow gradient data, and the third state machine can continuously store the wrapped dimension-reduced results at the storage location corresponding to each computing node, thereby improving the training efficiency.
[0095] Exemplary method
[0096] As Figure 4 shown, this exemplary embodiment proposes a data processing method applied to the data processing device described in any one of the above exemplary devices. In the embodiments of the present application, the data processing method includes the following steps S100-S400:
[0097] Step S100: In response to receiving the weight parameter update instruction, based on the first state machine, determine the target computing node information for updating the weight parameter in the weight parameter update instruction and the storage address of the gradient data corresponding to the target computing node in the memory module; and, based on the storage address, the first state machine sends a read request to the memory module, and sends the target computing node information to the second state machine based on the first state machine;
[0098] Step S200: Based on the second state machine, receive the target computing node information and the gradient data backflowed from the memory module; and, based on the second state machine, send the gradient data to the multiplier-accumulator and send the target computing node information to the third state machine;
[0099] Step S300: Based on the multiplier - adder, perform dimensionality reduction on the gradient data to obtain a dimensionality - reduced result, and send the dimensionality - reduced result to the third state machine;
[0100] Step S400: Based on the third state machine and the target computing node information, wrap the dimensionality - reduced result and write it to the storage location corresponding to the target computing node for the target computing node to call.
[0101] In the embodiment of the present application, when the weight parameter update instruction contains multiple target computing node information and corresponding multiple gradient data storage addresses, the data processing method further includes:
[0102] Based on the first state machine, send read requests to the memory module respectively based on each storage address, and send a first identifier to the second state machine before the gradient data of each read request flows back. The first identifier is used to identify the starting position of the returned data corresponding to each read request;
[0103] Based on the first state machine, in the order of the request sequence of each read request, send the target computing node information corresponding to each read request to the second state machine in sequence.
[0104] In the embodiment of the present application, the data processing method further includes:
[0105] Based on the second state machine, send each first identifier and the corresponding returned data to the multiplier - adder in sequence; and based on the second state machine, send each target computing node information to the third state machine;
[0106] Based on the multiplier - adder, accumulate the data after the first identifier, and when receiving the next first identifier, output the current accumulated result as the dimensionality - reduced result to the third state machine.
[0107] In the embodiment of the present application, the data processing method further includes:
[0108] Based on the third state machine, wrap the dimensionality - reduced result corresponding to each target computing node information respectively, and write the wrapped dimensionality - reduced results to their corresponding storage positions respectively.
[0109] In the embodiment of the present application, the data processing method further includes:
[0110] When all the gradient data corresponding to each storage address included in the weight parameter update instruction have flowed back, send a second identifier to the second state machine based on the first state machine. The second identifier is used to identify that all the gradient data corresponding to the weight parameter update instruction have flowed back;
[0111] After all the first identifiers and the corresponding feedback data are sent to the multiplier-adder, the second identifier is sent to the multiplier-adder based on the second state machine;
[0112] When the multiplier-adder receives the second identifier, the dimensionality reduction operation ended by the multiplier-adder is terminated, and the current accumulated result is sent to the third state machine.
[0113] In the embodiment of the present application, the data processing method further includes:
[0114] When the gradient data in the storage address corresponding to any read request is empty, a third identifier is sent to the second state machine based on the first state machine, and the third identifier is used to identify that the feedback data of the current read request is non-gradient data;
[0115] Based on the second state machine, the third identifier and the corresponding feedback data are sequentially sent to the multiplier-adder;
[0116] When the multiplier-adder receives the third identifier, all the data after the third identifier is skipped by the multiplier-adder until other identifiers are received.
[0117] For the specific implementation methods of the steps of the data processing method in the above embodiments, refer to the embodiments of each data processing device in the exemplary device, which will not be elaborated here one by one.
[0118] In addition, the data processing method proposed in the embodiment of the present application is applied to the data processing device in any of the above embodiments. Therefore, it has at least all the beneficial effects of the above data processing device, which will not be elaborated here one by one.
[0119] Exemplary medium
[0120] After introducing the methods, media, and systems of the exemplary embodiments of the present application, next, refer to Figure 5 A computer-readable storage medium of the exemplary embodiment of the present application will be described. Please refer to Figure 5, which shows that the computer-readable storage medium is an optical disc 70, on which a computer program (i.e., program product) is stored. When the computer program is run by a processor, it will implement the steps described in the above method embodiments. For example, in response to receiving a weight parameter update instruction, based on the first state machine, determine the target computing node information for updating the weight parameter in the weight parameter update instruction and the storage address of the gradient data corresponding to the target computing node in the memory module; and, based on the storage address and the first state machine, send a read request to the memory module, and send the target computing node information to the second state machine based on the first state machine; based on the second state machine, receive the target computing node information and the gradient data refluxed by the memory module; and, based on the second state machine, send the gradient data to the multiplier-accumulator and send the target computing node information to the third state machine; based on the multiplier-accumulator, perform dimensionality reduction on the gradient data to obtain a dimensionality reduction result, and send the dimensionality reduction result to the third state machine; based on the third state machine and the target computing node information, wrap the dimensionality reduction result and write it to the storage location corresponding to the target computing node for the target computing node to call. The specific implementation manners of each step will not be repeated here. It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated one by one here.
[0121] Exemplary computing device
[0122] After introducing the methods, systems, and media of the exemplary embodiments of the present application, next, reference is made to Figure 6 the computing device of the exemplary embodiments of the present application.
[0123] Figure 6 FIG. shows a block diagram of an exemplary computing device 80 suitable for implementing the embodiments of the present application. The computing device 80 may be a computer system or a server. Figure 6 The shown computing device 80 is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present application.
[0124] As Figure 6 shown, the components of the computing device 80 may include, but are not limited to: one or more processors or processing units 801, a system memory 802, and a bus 803 connecting different system components (including the system memory 802 and the processing unit 801).
[0125] Computing device 80 typically includes a variety of computer system readable media. These media can be any available media accessible by computing device 80, including volatile and non-volatile media, removable and non-removable media.
[0126] System memory 802 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 8021 and / or cache memory 8022. Computing device 80 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 8023 may be used to read and write to non-removable, non-volatile magnetic media ( Figure 6 not shown in the figure, commonly referred to as a "hard disk drive"). Although not shown in Figure 6 the figure, a disk drive for reading and writing to removable non-volatile disks (such as "floppy disks"), and an optical disk drive for reading and writing to removable non-volatile optical disks (such as CD-ROM, DVD-ROM or other optical media) may be provided. In these cases, each drive may be connected to bus 803 through one or more data media interfaces. System memory 802 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the various embodiments of the present application.
[0127] A program / utility 8025 having a set (at least one) of program modules 8024 may be stored, for example, in system memory 802, and such program modules 8024 include but are not limited to: an operating system, one or more application programs, other program modules, and program data, and the implementation of a network environment may be included in each or some combination of these examples. Program modules 8024 generally perform the functions and / or methods described in the embodiments of the present application.
[0128] Computing device 80 may also communicate with one or more external devices 804 (such as a keyboard, pointing device, display, etc.). Such communication may be through an input / output (I / O) interface. Also, computing device 80 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through network adapter 806. As Figure 6 shown, network adapter 806 communicates with other modules of computing device 80 (such as processing unit 801, etc.) through bus 803. It should be understood that although Figure 6 not shown in the figure, other hardware and / or software modules may be used in conjunction with computing device 80.
[0129] The processing unit 801 executes various functional applications and data processing by running the programs stored in the system memory 802. For example, in response to receiving a weight parameter update instruction, based on the first state machine, it determines the target calculation node information of the updated weight parameter in the weight parameter update instruction and the storage address of the gradient data corresponding to the target calculation node in the memory module; and, based on the storage address and the first state machine, it sends a read request to the memory module, and sends the target calculation node information to the second state machine based on the first state machine; based on the second state machine, it receives the target calculation node information and the gradient data returned by the memory module; and, based on the second state machine, it sends the gradient data to the multiplier-accumulator and sends the target calculation node information to the third state machine; based on the multiplier-accumulator, it reduces the dimension of the gradient data to obtain a dimension reduction result and sends the dimension reduction result to the third state machine; based on the third state machine and the target calculation node information, it wraps the dimension reduction result and writes it to the storage location corresponding to the target calculation node for the target calculation node to call. The specific implementation methods of each step will not be repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the arithmetic device are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.
[0130] In the description of the present application, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0131] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, systems, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0132] In several embodiments provided by the present application, it should be understood that the disclosed systems, systems, and methods can be implemented in other ways. The above-described system embodiments are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the systems or units can be in electrical, mechanical, or other forms.
[0133] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0134] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, can also exist physically alone for each unit, or two or more units can be integrated in one unit.
[0135] If the above function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0136] Finally, it should be noted that: the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solution of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0137] In addition, although the operations of the method of the present application are described in a specific order in the drawings, this does not require or imply that these operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
Claims
1. A data processing device is applied to a distributed training system. The distributed training system includes multiple computing nodes and a memory module. Each of the computing nodes is used to deploy partial operators of a model to be trained, and the memory module is used to store gradient data when weight parameters of the operators deployed by each of the computing nodes are updated; The gradient data of the respective weight parameters of each computing node in each iteration round is calculated by an external device; The data processing device includes: a first state machine, a second state machine, a third state machine, and a multiplier-accumulator. The first state machine, the second state machine, the third state machine, and the multiplier-accumulator are used to be deployed on any one of the computing nodes, and the first state machine, the second state machine, the third state machine, and the multiplier-accumulator are deployed on the same computing node; The computing node on which the first state machine, the second state machine, the third state machine, and the multiplier-accumulator are deployed is communicatively connected to an external device, and other computing nodes do not need to be communicatively connected to the external device; The first state machine is configured to: in response to receiving a weight parameter update instruction, determine, based on the weight parameter update instruction, the target computing node information for updating the weight parameter in the weight parameter update instruction and the storage address of the gradient data corresponding to the target computing node in the memory module; and, based on the storage address, send a read request to the memory module, and send the target computing node information to the second state machine; The second state machine is configured to: receive the target computing node information and the gradient data returned by the memory module; and, send the gradient data to the multiplier-accumulator, and send the target computing node information to the third state machine; The multiplier-accumulator is configured to: perform dimensionality reduction on the gradient data to obtain a dimensionality reduction result, and send the dimensionality reduction result to the third state machine; The third state machine is configured to: wrap the dimensionality reduction result based on the target computing node information, and write it to the storage location corresponding to the target computing node for the target computing node to call.
2. The data processing device according to claim 1, when the weight parameter update instruction includes multiple target computing node information and corresponding multiple gradient data storage addresses, the first state machine is further configured to: respectively send read requests to the memory module based on each storage address, and send a first identifier to the second state machine before the gradient data of each read request is returned. The first identifier is used to identify the start position of the returned data corresponding to each read request; and According to the request order of each read request, sequentially send the target computing node information corresponding to each read request to the second state machine.
3. The data processing device according to claim 2, the second state machine is further configured to: sequentially send each first identifier and the corresponding returned data to the multiplier-accumulator; and send each of the target computing node information to the third state machine; The multiplier-accumulator is further configured to: in response to receiving the first identifier, accumulate the data after the first identifier, and when receiving the next first identifier, output the current accumulated result as the dimensionality reduction result to the third state machine.
4. The data processing device according to claim 3, the third state machine is further configured to: respectively wrap the dimensionality reduction result corresponding to each target computing node information based on each target computing node information, and write the wrapped dimensionality reduction results to their corresponding storage locations respectively.
5. The data processing device according to claim 2, wherein the first state machine is further configured to: after all the gradient data corresponding to each storage address included in the weight parameter update instruction has been backflowed, send a second identifier to the second state machine, where the second identifier is used to identify that all the gradient data corresponding to the weight parameter update instruction has been backflowed; The second state machine is further configured to: after all the first identifiers and the corresponding backflow data have been sent to the multiplier-accumulator, send the second identifier to the multiplier-accumulator; The multiplier-accumulator is further configured to: in response to receiving the second identifier, end the dimensionality reduction operation and send the current accumulated result to the third state machine.
6. The data processing device according to claim 2, wherein the first state machine is further configured to: when the gradient data in the storage address corresponding to any read request is empty, send a third identifier to the second state machine, where the third identifier is used to identify that the backflow data of the current read request is non-gradient data; The second state machine is further configured to: sequentially send the third identifier and the corresponding backflow data to the multiplier-accumulator; The multiplier-accumulator is further configured to: in response to receiving the third identifier, skip all the data after the third identifier until other identifiers are received.
7. A data processing method, applied to the data processing device according to any one of claims 1-6, the data processing method comprising: In response to receiving a weight parameter update instruction, based on the first state machine, determine the target calculation node information for updating the weight parameter in the weight parameter update instruction and the storage addresses of the gradient data corresponding to the target calculation node in the memory module; And, based on the storage addresses, send a read request to the memory module by the first state machine, and send the target calculation node information to the second state machine based on the first state machine; Based on the second state machine, receive the target calculation node information and the gradient data backflowed from the memory module; And, send the gradient data to the multiplier-accumulator based on the second state machine, and send the target calculation node information to the third state machine; Based on the multiplier-accumulator, perform dimensionality reduction on the gradient data to obtain a dimensionality reduction result, and send the dimensionality reduction result to the third state machine; Based on the third state machine and the target calculation node information, wrap the dimensionality reduction result and write it to the storage location corresponding to the target calculation node for the target calculation node to call.
8. The data processing method according to claim 7, when the weight parameter update instruction includes multiple target calculation node information and corresponding multiple gradient data storage addresses, the data processing method further comprises: Based on the first state machine, send read requests to the memory module respectively based on each storage address, and send a first identifier to the second state machine before the gradient data of each read request is backflowed, where the first identifier is used to identify the starting position of the backflow data corresponding to each read request; Based on the first state machine, in accordance with the request order of each read request, the target computing node information corresponding to each read request is sequentially sent to the second state machine.
9. The data processing method according to claim 8, wherein the data processing method further comprises: Based on the second state machine, each first identifier and the corresponding feedback data are sequentially sent to the multiplier-accumulator; And based on the second state machine, the target computing node information is sent to the third state machine; Based on the multiplier-accumulator, the data after the first identifier is accumulated, and when the next first identifier is received, the current accumulated result is output to the third state machine as the dimensionality reduction result.
10. The data processing method according to claim 9, wherein the data processing method further comprises: Based on the third state machine, each dimensionality reduction result is wrapped based on each target computing node information, and the wrapped dimensionality reduction results are respectively written to their corresponding storage locations.
11. The data processing method according to claim 8, wherein the data processing method further comprises: After all the gradient data corresponding to the storage addresses included in the weight parameter update instruction are fed back, based on the first state machine, a second identifier is sent to the second state machine, and the second identifier is used to identify that all the gradient data corresponding to the weight parameter update instruction are fed back; After all the first identifiers and the corresponding feedback data are sent to the multiplier-accumulator, based on the second state machine, the second identifier is sent to the multiplier-accumulator; When the multiplier-accumulator receives the second identifier, the dimensionality reduction operation of the multiplier-accumulator is ended, and the current accumulated result is sent to the third state machine.
12. The data processing method according to claim 8, wherein the data processing method further comprises: When the gradient data in the storage address corresponding to any read request is empty, based on the first state machine, a third identifier is sent to the second state machine, and the third identifier is used to identify that the feedback data of the current read request is non-gradient data; Based on the second state machine, the third identifier and the corresponding feedback data are sequentially sent to the multiplier-accumulator; When the multiplier-accumulator receives the third identifier, all the data after the third identifier are skipped by the multiplier-accumulator until other identifiers are received.
13. A computer-readable storage medium, which includes instructions that, when running on a computer, cause the computer to execute the method according to any one of claims 7-12.
14. A computing device, the computing device includes a processor, and when the processor executes a computer program stored in a memory, the method according to any one of claims 7-12 is implemented.
Citation Information
Patent Citations
Enhanced pre-read capabilities for storage devices
CN114168495A
Model training method, computing device and system
CN118278540A
Data processing device and method, medium and computing equipment
CN118349213A