Data Communication Method, Apparatus and Electronic Device for Distributed Training
By adjusting cache size parameters in a distributed system to optimize data transmission, the problem of distributed training communication optimization on hardware devices of different models is solved, and efficient communication and performance improvement of deep learning models is achieved.
Patent Information
- Application Number
- CN202211678659.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-12-26
AI Technical Summary
How to optimize communications for unified distributed training on multiple hardware devices of different models to improve performance.
The backpropagation data obtained by the deep learning model for distributed training is obtained through the application layer, and the data is sent to the second node based on the first cache size parameter of the communication layer and the second cache size parameter of the network layer. After obtaining the data transmission performance information, adjust the cache size parameters to optimize data transmission.
The communication optimization of distributed training of deep learning models is realized, and data transmission performance is improved.
Smart Images

Figure CN115987914B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning technology, and particularly to a data communication method, apparatus, and electronic device for distributed training. Background Art
[0002] Distributed training refers to splitting the training task of an originally huge machine learning model into multiple subtasks, and each subtask is executed separately on an independent machine. Distributed training includes two training methods: data parallelism and model parallelism. Data parallelism means splitting the training data set and placing it on multiple hardware devices, performing calculations on each hardware device, and transmitting model parameters between multiple hardware devices. Model parallelism means distributing the calculation of a single operator of the model to multiple hardware devices for concurrent calculation to achieve the purpose of improving the calculation speed of a single operator.
[0003] Whether it is data parallelism or model parallelism, distributed training involves communication between multiple hardware devices and has different performance manifestations in different scenarios. Therefore, how to optimize the communication of unified distributed training on multiple hardware devices of different models and improve performance has become a problem to be solved. Summary of the Invention
[0004] Embodiments of this application provide a data communication method, apparatus, and electronic device for distributed training to achieve communication optimization for distributed training of deep learning models and thus improve performance.
[0005] In a first aspect, embodiments of this application provide a data communication method for distributed training. The method is applied to a first node in a distributed system, and the method includes:
[0006] Obtaining, through an application layer, backpropagation data obtained by performing distributed training on a deep learning model;
[0007] Based on a first cache size parameter of a communication layer and a second cache size parameter of a network layer, sending the backpropagation data to a second node in the distributed system;
[0008] Obtaining data sending performance information, adjusting the first cache size parameter and the second cache size parameter according to the data sending performance information, and sending the received backpropagation data to the second node according to the adjusted first cache size parameter and second cache size parameter.
[0009] In a second aspect, embodiments of this application provide a data communication apparatus for distributed training. The apparatus is applied to a first node in a distributed system, and the apparatus includes:
[0010] An obtaining module, configured to obtain, through an application layer, backpropagation data obtained by performing distributed training on a deep learning model;
[0011] A sending module, configured to send backpropagation data to a second node in a distributed system based on a first cache size parameter of a communication layer and a second cache size parameter of a network layer;
[0012] An adjustment module, configured to obtain data sending performance information, adjust the first cache size parameter and the second cache size parameter according to the data sending performance information, and send the received backpropagation data to the second node according to the adjusted first cache size parameter and second cache size parameter.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the method described in any one of the above is implemented.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method described in any one of the above is implemented.
[0015] Compared with the prior art, the present application has the following advantages:
[0016] The present application provides a data communication method, apparatus, and electronic device for distributed training. A first node obtains backpropagation data obtained by distributed training of a deep learning model through an application layer; based on a first cache size parameter of a communication layer and a second cache size parameter of a network layer, sends the backpropagation data to a second node; obtains data sending performance information, adjusts the first cache size parameter and the second cache size parameter according to the data sending performance information, and sends the received backpropagation data to the second node according to the adjusted first cache size parameter and second cache size parameter. In this embodiment, the first node realizes data communication with the second node based on the first cache size parameter of the communication layer and the second cache size parameter of the network layer, adjusts the first cache size parameter and the second cache size parameter according to the data sending performance information, and sends the received backpropagation data to the second node according to the adjusted first cache size parameter and second cache size parameter to realize communication optimization for distributed training of a deep learning model, thereby improving performance.
[0017] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. Description of the Drawings
[0018] In the accompanying drawings, unless otherwise specified, the same reference numerals in multiple drawings denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings merely depict some embodiments in accordance with the present application and should not be regarded as limiting the scope of the present application.
[0019] Figure 1 A schematic diagram of an application scenario of the data communication method for distributed training provided for the present application;
[0020] Figure 2 A flowchart of the data communication method for distributed training according to an embodiment of the present application;
[0021] Figure 3 A structural block diagram of the data communication device for distributed training according to an embodiment of the present application; and
[0022] Figure 4 A block diagram of an electronic device for implementing the embodiments of the present application. Detailed implementation manners
[0023] In the following, only some exemplary embodiments are briefly described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature and not restrictive.
[0024] To facilitate the understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.
[0025] Figure 1 A schematic diagram of an application scenario of the data communication method for distributed training provided for the present application. As Figure 1As shown, the method in this embodiment can be deployed in nodes of a distributed system. The nodes can be any computing devices, such as servers, etc. The method in this embodiment involves an application layer, a communication layer, and a network layer, and is applicable to both data parallelism and model parallelism in the distributed training of machine learning models. Among them, the application layer can be implemented through a Model-wrapper component to perform forward propagation operations and backward propagation operations on a deep learning model, and can also set and adjust fusion granularity parameters, etc. The communication layer can be implemented through a Comm-wrapper component, where cache size parameters, data unit size parameters (chunk_size), and data channel capacity (channel_size) can be set and adjusted, and operators in a communication function library (such as the NVIDIA Collective Communications Library (NCCL)) can be called to perform communication-related calculations. The collective communication algorithm of NCCL is executed through atomic operators, and each operator is combined according to the pipeline of tasks, and the cache size parameter is set to optimize the loss caused by context overhead. The network layer can be implemented through a Socket-wrapper component, which encapsulates socket interface protocols, etc., and can set and adjust read cache size parameters and write cache size parameters.
[0026] In practical applications, the first node in the distributed system performs distributed training of a deep learning model through the application layer, performs forward propagation calculations and backward propagation calculations, fuses the obtained multiple gradient data according to the fusion granularity parameters to obtain backward propagation data; based on the cache size parameter, data unit size parameter, data channel capacity of the communication layer, and the read cache size parameter and write cache size parameter of the network layer, transmits the backward propagation data to the second node in the distributed system; uses a hook function to obtain the interval time between two adjacent forward propagation calculations of the deep learning model as data transmission performance information, and adjusts the cache size parameter of the communication layer and the cache size parameter of the network layer according to the data transmission performance information. In this embodiment, based on the cache size parameter of the communication layer and the cache size parameter of the network layer, the backward propagation data is sent to the second node, the fusion granularity parameter of the application layer, the cache size parameter of the communication layer, and the cache size parameter of the network layer are adjusted according to the data transmission performance information, and the received backward propagation data is sent to the second node according to the adjusted first cache size parameter and second cache size parameter to achieve communication optimization for the distributed training of the deep learning model, thereby improving performance.
[0027] An embodiment of the present application provides a data communication method for distributed training. The method in this embodiment can be applied to a computing device, which may include a server, etc. As Figure 2 The following is a flowchart of a data communication method for distributed training according to an embodiment of the present application, including:
[0028] Step S201, obtain the backpropagation data obtained by distributed training of a deep learning model through the application layer.
[0029] Step S202, based on the first cache size parameter of the communication layer and the second cache size parameter of the network layer, send the backpropagation data to a second node in the distributed system.
[0030] Step S203, obtain data sending performance information, adjust the first cache size parameter and the second cache size parameter according to the data sending performance information, and send the received backpropagation data to the second node according to the adjusted first cache size parameter and the second cache size parameter.
[0031] The method in this embodiment can be applied to a first node in a distributed system. Among them, the application layer can be implemented through a Model-wrapper component to perform forward propagation operations and backpropagation operations on a deep learning model to obtain backpropagation data. The communication layer can be implemented through a Comm-wrapper component, which can set and adjust the initial value of the first cache size parameter, and can call operators in a communication function library (for example, NCCL) to perform communication-related calculations. The network layer can be implemented through a Socket-wrapper component, which encapsulates a socket interface protocol, etc., and can set and adjust the initial value of the second cache size parameter.
[0032] Among them, the first cache size parameter represents the size of the cache area corresponding to the communication layer; the second cache size parameter represents the size of the cache area corresponding to the network layer. The data sending performance information represents the performance of data sent from the first node to the second node, and the performance can be measured by various metrics, for example, data transmission time, etc. The shorter the data transmission time, the higher the performance.
[0033] Optionally, adjust the first cache size parameter and the second cache size parameter according to the data transmission performance information, including: preset the number of steps, the data volume threshold corresponding to the first cache size parameter, and the data volume threshold corresponding to the second cache size parameter. When the respective data volume thresholds are not reached, increase the number of steps. The number of steps can be consecutive integers. According to the data transmission performance information, increase or decrease the first cache size parameter, or increase or decrease the second cache size parameter. Send the received backpropagation data to the second node according to the adjusted first cache size parameter and second cache size parameter, which can improve the performance of data transmission.
[0034] The embodiment of the present application provides a data communication method for distributed training, which obtains the backpropagation data obtained by distributed training of a deep learning model through the application layer; based on the first cache size parameter of the communication layer and the second cache size parameter of the network layer, sends the backpropagation data to the second node; obtains the data transmission performance information, adjusts the first cache size parameter and the second cache size parameter according to the data transmission performance information, and sends the received backpropagation data to the second node according to the adjusted first cache size parameter and second cache size parameter. In this embodiment, the first node realizes data communication with the second node based on the first cache size parameter of the communication layer and the second cache size parameter of the network layer, adjusts the first cache size parameter and the second cache size parameter according to the data transmission performance information, and sends the received backpropagation data to the second node according to the adjusted first cache size parameter and second cache size parameter to achieve communication optimization for distributed training of the deep learning model, thereby improving the performance.
[0035] In one implementation, step S201, obtaining the backpropagation data obtained by distributed training of a deep learning model through the application layer, includes:
[0036] Step S2011, obtaining the data fusion granularity parameter of the application layer.
[0037] Step S2012, according to the data fusion granularity parameter, fuse multiple gradient data obtained by distributed training of the deep learning model to obtain backpropagation data.
[0038] Among them, the fusion granularity parameter represents the amount of data that each bucket can hold. After fusing and processing multiple gradient data, they are distributed into multiple buckets as backpropagation data. Based on the first cache size parameter of the communication layer and the second cache size parameter of the network layer, the backpropagation data is sent to the second node. Each gradient data corresponds to a tensor composed of the model parameters of the deep learning model. The initial value and tuning threshold of the fusion granularity parameter can be preconfigured. After receiving gradient data each time, the size of the fusion granularity parameter can be adjusted, and data fusion processing is performed according to the adjusted data fusion granularity parameter until the tuning threshold is reached.
[0039] In one implementation, in step S203, it further includes: adjusting the data fusion granularity parameter according to the data sending performance information.
[0040] When adjusting the data fusion granularity parameter, if the data fusion granularity parameter is large, it will lead to inefficiency in computing waiting for communication and slow down the overall performance; if the data fusion granularity parameter is small, it will lead to additional overhead in calling the communication layer and the overall performance will decline. By obtaining the parameters of the deep learning model through the application layer, the release and re-initialization of the deep learning model can be completed. By adjusting the size of the fusion granularity parameter, the data transmission performance can be improved.
[0041] Optionally, the initial values and tuning thresholds corresponding to the fusion granularity parameter, the first cache size parameter, and the second cache size parameter can be preconfigured. When the data sending performance information does not meet the preset requirements, the data fusion granularity parameter, the first cache size parameter, and the second cache size parameter are adjusted respectively according to the preset step number, the start step number of tuning, and the end step number of tuning until at least one of them reaches the corresponding tuning threshold, and the tuning ends. By adjusting the parameters corresponding to the application layer, the communication layer, and the network layer respectively, the data sending performance can be improved.
[0042] When adjusting the second cache size parameter, the size of the parameter can be adjusted according to the network state. If the current network latency is large, the second cache size parameter is adjusted to a larger value to reduce the loss caused by the request confirmation operation. If the current network latency is small, the second cache size parameter is adjusted to a smaller value to improve the response speed and reduce the latency.
[0043] In one implementation, step S202, sending the backpropagation data to the second node based on the first cache size parameter of the communication layer and the second cache size parameter of the network layer, includes:
[0044] In step S2021, cache the backpropagation data in the first buffer area corresponding to the communication layer. If the amount of data cached in the first buffer area reaches the data volume threshold corresponding to the first buffer size parameter, then cache the data cached in the first buffer area in the second buffer area corresponding to the network layer.
[0045] In step S2022, if the amount of data cached in the second buffer area reaches the data volume threshold corresponding to the second buffer size parameter, then send the data cached in the second buffer area to the second node.
[0046] Among them, the first buffer size parameter represents the size of the buffer area corresponding to the communication layer. The second buffer size parameter represents the size of the buffer area corresponding to the network layer. Cache the backpropagation data in the first buffer area corresponding to the communication layer. If the amount of data cached in the first buffer area reaches the data volume threshold corresponding to the first buffer size parameter, that is, the maximum amount of data that can be cached in the buffer area corresponding to the communication layer, then cache the data cached in the first buffer area in the second buffer area corresponding to the network layer. If the amount of data cached in the second buffer area reaches the data volume threshold corresponding to the second buffer size parameter, that is, the maximum amount of data that can be cached in the buffer area corresponding to the network layer, then send the data cached in the second buffer area to the second node.
[0047] In one implementation, the method further includes: obtaining the data unit size parameter and data channel capacity of the communication layer. Based on the first buffer size parameter of the communication layer and the second buffer size parameter of the network layer, sending the backpropagation data to the second node further includes: based on the first buffer size parameter, the second buffer size parameter, the data unit size parameter, and the data channel capacity, sending the backpropagation data to the second node.
[0048] Among them, the data unit size (chunk_size) parameter represents the amount of data for each data unit (chunk) when dividing the backpropagation data into multiple data units (chunks). The data channel (Channel) is bidirectional and can perform data reading and data writing.
[0049] In one implementation, sending the backpropagation data to the second node based on the first cache size parameter, the second cache size parameter, the data unit size parameter, and the data channel capacity includes: dividing the backpropagation data into multiple data units according to the data unit size parameter and the data channel capacity, where the multiple data units correspond to at least one data channel. Caching the multiple data units in the first cache area corresponding to the communication layer. If the amount of data cached in the first cache area reaches the data volume threshold corresponding to the first cache size parameter, then caching the data cached in the first cache area in the second cache area corresponding to the network layer. If the amount of data cached in the second cache area reaches the data volume threshold corresponding to the second cache size parameter, then sending the data cached in the second cache area to the second node.
[0050] In practical applications, the backpropagation data is divided into multiple data units according to the data unit size parameter and the data channel capacity, and the multiple data units (chunks) are used to read and write data through at least one data channel (Channel). Caching the multiple data units in the first cache area corresponding to the communication layer. If the amount of data cached in the first cache area reaches the data volume threshold corresponding to the first cache size parameter, then the first cache area is full, and the data cached in the first cache area is cached in the second cache area corresponding to the network layer. If the amount of data cached in the second cache area reaches the data volume threshold corresponding to the second cache size parameter, then the second cache area is full, and the data cached in the second cache area is sent to the second node.
[0051] Among them, the second cache size parameter may include a write cache size parameter and a read cache size parameter. The first node sends the backpropagation data to the second node or receives the data sent by the second node based on the write cache size parameter or the read cache size parameter.
[0052] In one implementation, in step S203, obtaining the data sending performance information includes: using a hook function to obtain the interval time between the forward propagation calculations of two adjacent trainings of the deep learning model as the data sending performance information.
[0053] Using a hook function to obtain the start time of the forward propagation calculation of the previous training of the deep learning model and the start time of the forward propagation calculation of the current training, and taking the interval time between these two start times as the data sending time, and taking the data sending time as the data sending performance information.
[0054] In an example, the communication layer calls an operator in the communication function library NCCL for communication-related calculations, and the network layer encapsulates the Socket protocol interface. The parameter adjustment processes of the application layer, the communication layer, and the network layer are as follows:
[0055] 1. Set the tuning thresholds, step numbers for each parameter, as well as the start step number and the maximum step number for the entire tuning.
[0056] 2. Set the initial values corresponding to the fusion granularity parameter buffer_fusing_size, the first buffer size parameter buffer_nccl, and the second buffer size parameter buffer_socket.
[0057] 3. Application layer: Start tuning buffer_fusing_size. Each time buffer_fusing_size is set, training needs to be restarted to avoid resource consumption such as additional video memory occupation. At the same time, maintain a daemon process to control the end of the loop to achieve parameter tuning during runtime.
[0058] 4. Insert a hook function to trigger the start of performance timing.
[0059] 5. When reaching the sampling point, end the timing and update the performance table:
[0060] dp[buffer_fusing_size][buffer_nccl][buffer_socket] = [xxx, xxx, xxx]. Here, x represents data transmission performance information.
[0061] 6. If the relevant parameter configuration already exists, directly return the performance result. Under different buffer_fusing_size, there will be data packets with different communication volumes, but at the underlying layer, there will be the same data unit size. Therefore, the same parameter configuration needs to be reused to calculate the performance and reduce the number of tuning steps.
[0062] 7. If the tuning threshold of buffer_fusing_size is reached, or the maximum number of iteration steps is reached, return the currently performance-optimal parameter configuration and exit the loop.
[0063] 8. Network layer: If the tuning threshold of buffer_fusing_size is not reached, perform parameter tuning for buffer_socket: Adjust the read buffer size parameter read_buffer and the write buffer size parameter write_buffer according to the progressive step number. Each adjustment needs to wait for the communication operator in NCCL to calculate and end, otherwise the setting cannot be successful.
[0064] 9. Communication layer: If the buffer_socket tuning threshold is reached, parameter tuning of buffer_nccl is performed. The parameters such as the data channel capacity channel_size, the data unit size parameter chunk_size, and buffer_nccl are adjusted according to the progressive steps. Each adjustment requires releasing the communication class in NCCL to reduce additional resource consumption.
[0065] 10. If the buffer_nccl tuning threshold is reached, the application layer is returned to adjust buffer_fusing_size. In this embodiment, the fusion granularity can be reset during the running phase to ensure accurate performance evaluation. Moreover, operations such as intercepting, releasing, and reconstructing the NCCL operator can be implemented to ensure accurate performance evaluation. Multi-level cache parameter unified optimization of the application layer, communication layer, and network layer is achieved, the optimal configuration in a specific environment is found within the limited solution space, and runtime tuning is supported to achieve dynamic optimization of data transmission.
[0066] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides a data communication device for distributed training. As Figure 3 shown in the structural block diagram of the data communication device for distributed training according to an embodiment of the present application, the device is applied to the first node in the distributed system, and the device includes:
[0067] An acquisition module 301, configured to acquire backpropagation data obtained by distributed training of a deep learning model through the application layer;
[0068] A sending module 302, configured to send the backpropagation data to a second node in the distributed system based on a first cache size parameter of the communication layer and a second cache size parameter of the network layer;
[0069] An adjustment module 303, configured to acquire data sending performance information, adjust the first cache size parameter and the second cache size parameter according to the data sending performance information, and send the received backpropagation data to the second node according to the adjusted first cache size parameter and the second cache size parameter.
[0070] An embodiment of the present application provides a data communication device for distributed training, which obtains backpropagation data obtained by distributed training of a deep learning model through an application layer; based on a first cache size parameter of a communication layer and a second cache size parameter of a network layer, sends the backpropagation data to a second node; obtains data transmission performance information, and adjusts the first cache size parameter and the second cache size parameter according to the data transmission performance information, and sends the received backpropagation data to the second node according to the adjusted first cache size parameter and the second cache size parameter. In this embodiment, the first node realizes data communication with the second node based on the first cache size parameter of the communication layer and the second cache size parameter of the network layer, adjusts the first cache size parameter and the second cache size parameter according to the data transmission performance information, and sends the received backpropagation data to the second node according to the adjusted first cache size parameter and the second cache size parameter, so as to optimize the communication for distributed training of the deep learning model, and further improve the performance.
[0071] In one implementation, an obtaining module 301 is configured to: obtain a data fusion granularity parameter of an application layer; and fuse multiple gradient data obtained by distributed training of a deep learning model according to the data fusion granularity parameter to obtain backpropagation data.
[0072] In one implementation, the adjusting module 303 is further configured to:
[0073] Adjust the data fusion granularity parameter according to the data transmission performance information.
[0074] In one implementation, a sending module 302 is configured to: cache the backpropagation data in a first cache area corresponding to the communication layer, and if the amount of data cached in the first cache area reaches the data amount threshold corresponding to the first cache size parameter, cache the data cached in the first cache area in a second cache area corresponding to the network layer; and if the amount of data cached in the second cache area reaches the data amount threshold corresponding to the second cache size parameter, send the data cached in the second cache area to the second node.
[0075] In one implementation, the device is further configured to: obtain a data unit size parameter and a data channel capacity of the communication layer;
[0076] The sending module 302 is further configured to: send the backpropagation data to the second node based on the first cache size parameter, the second cache size parameter, the data unit size parameter, and the data channel capacity.
[0077] In one implementation, the sending module 302 is specifically configured to: divide the backpropagation data into multiple data units according to the data unit size parameter and the data channel capacity, where the multiple data units correspond to at least one data channel; cache the multiple data units in a first cache area corresponding to the communication layer, and if the amount of data cached in the first cache area reaches the data volume threshold corresponding to the first cache size parameter, cache the data cached in the first cache area in a second cache area corresponding to the network layer; and if the amount of data cached in the second cache area reaches the data volume threshold corresponding to the second cache size parameter, send the data cached in the second cache area to the second node.
[0078] In one implementation, when obtaining the data sending performance information, the adjustment module 303 is configured to: use a hook function to obtain the interval time of the forward propagation calculation of two adjacent trainings of the deep learning model as the data sending performance information.
[0079] For the functions of the modules in each device of the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and they have the corresponding beneficial effects, which will not be elaborated here.
[0080] Figure 4 It is a block diagram of an electronic device for implementing the embodiments of the present application. As Figure 4 shown, the electronic device includes: a memory 410 and a processor 420, and a computer program that can run on the processor 420 is stored in the memory 410. When the processor 420 executes the computer program, the methods in the above embodiments are implemented. The number of the memory 410 and the processor 420 can be one or more.
[0081] The electronic device further includes:
[0082] a communication interface 430, configured to communicate with external devices and perform data interaction and transmission.
[0083] If the memory 410, the processor 420, and the communication interface 430 are independently implemented, the memory 410, the processor 420, and the communication interface 430 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0084] Optionally, in a specific implementation, if the memory 410, the processor 420, and the communication interface 430 are integrated on a single chip, the memory 410, the processor 420, and the communication interface 430 can communicate with each other through an internal interface.
[0085] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the method provided in the embodiment of the present application.
[0086] An embodiment of the present application further provides a chip, which includes a processor for calling and running instructions stored in a memory, so that a communication device equipped with the chip executes the method provided in the embodiment of the present application.
[0087] An embodiment of the present application further provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is used to execute the code in the memory, and when the code is executed, the processor is used to execute the method provided in the embodiment of the application.
[0088] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the advanced reduced instruction set machine (ARM) architecture.
[0089] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0090] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.
[0091] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0092] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.
[0093] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed.
[0094] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatus, or devices.
[0095] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.
[0096] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist independently as individual units physically, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. This storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.
[0097] As mentioned above, it is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope recorded in the present application can easily think of various changes or substitutions thereof, and these should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A data communication method for distributed training, characterized in that, the method is applied to a first node in a distributed system, and the method includes: obtaining, through an application layer, backpropagation data obtained by performing distributed training on a deep learning model; sending the backpropagation data to a second node in the distributed system based on a first cache size parameter of a communication layer and a second cache size parameter of a network layer; obtaining data sending performance information, and according to the data sending performance information, adjusting the first cache size parameter and the second cache size parameter, and sending the received backpropagation data to the second node according to the adjusted first cache size parameter and second cache size parameter.
2. The method according to claim 1, characterized in that, the obtaining, through the application layer, the backpropagation data obtained by performing distributed training on the deep learning model includes: obtaining a data fusion granularity parameter of the application layer; fusing a plurality of gradient data obtained by performing distributed training on the deep learning model according to the data fusion granularity parameter to obtain the backpropagation data.
3. The method according to claim 2, characterized in that, further includes: adjusting the data fusion granularity parameter according to the data sending performance information.
4. The method according to claim 1 or 2, characterized in that, the sending the backpropagation data to a second node based on a first cache size parameter of a communication layer and a second cache size parameter of a network layer includes: caching the backpropagation data in a first cache area corresponding to the communication layer, and if the amount of data cached in the first cache area reaches a data amount threshold corresponding to the first cache size parameter, caching the data cached in the first cache area in a second cache area corresponding to the network layer; if the amount of data cached in the second cache area reaches a data amount threshold corresponding to the second cache size parameter, sending the data cached in the second cache area to the second node.
5. The method according to claim 1 or 2, characterized in that, the method further includes: obtaining a data unit size parameter and a data channel capacity of the communication layer; the sending the backpropagation data to a second node based on a first cache size parameter of a communication layer and a second cache size parameter of a network layer further includes: sending the backpropagation data to the second node based on the first cache size parameter, the second cache size parameter, the data unit size parameter, and the data channel capacity.
6. The method according to claim 5, characterized in that, the sending the backpropagation data to the second node based on the first cache size parameter, the second cache size parameter, the data unit size parameter, and the data channel capacity includes: dividing the backpropagation data into a plurality of data units according to the data unit size parameter and the data channel capacity, and the plurality of data units correspond to at least one data channel; Cache the multiple data units in a first cache area corresponding to the communication layer. If the amount of data cached in the first cache area reaches the data volume threshold corresponding to the first cache size parameter, cache the data cached in the first cache area in a second cache area corresponding to the network layer; If the amount of data cached in the second cache area reaches the data volume threshold corresponding to the second cache size parameter, send the data cached in the second cache area to the second node.
7. The method according to claim 1, wherein, the obtaining of the data sending performance information includes: using a hook function to obtain the interval time of the forward propagation calculation between two adjacent trainings of the deep learning model as the data sending performance information.
8. A data communication device for distributed training, wherein, the device is applied to a first node in a distributed system, and the device includes: an obtaining module, configured to obtain backpropagation data obtained by distributed training of a deep learning model through an application layer; a sending module, configured to send the backpropagation data to a second node in the distributed system based on a first cache size parameter of a communication layer and a second cache size parameter of a network layer; an adjustment module, configured to obtain data sending performance information, adjust the first cache size parameter and the second cache size parameter according to the data sending performance information, and send the received backpropagation data to the second node according to the adjusted first cache size parameter and second cache size parameter.
9. An electronic device, wherein, it includes a memory, a processor, and a computer program stored on the memory. When the processor executes the computer program, the method according to any one of claims 1-7 is implemented.
10. A computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
A gradient transmission method and a distributed training system
CN109919313A
Distributed training method and device for machine learning model and computer equipment
CN111709533A