Method for Implementing Distributed Neural Network Training Based on the nGraph Framework
By integrating OpenCL and Intel IKL platform environments into the nGraph framework and adding necessary operator implementations to the FPGA device, the technical difficulties of distributed training of deep learning neural networks in FPGA devices are solved, and efficient distributed training and performance optimization are achieved.
Patent Information
- Application Number
- CN202111161608.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-09-30
AI Technical Summary
In the prior art, distributed training of deep learning neural networks cannot be implemented in FPGA backend devices, which limits the efficiency and scalability of model training.
By integrating OpenCL standard API library and Intel IKL platform environment into the nGraph framework, operators required for neural network training are determined, and corresponding kernel implementations are added to the FPGA backend device to realize multi-device management and input data distribution to support distributed neural network training.
It realizes distributed training of deep learning neural networks in FPGA backend devices through the nGraph framework, improves training performance and scalability, and simplifies performance optimization across frameworks and hardware platforms.
Smart Images

Figure CN113988287B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and more specifically, to a method and apparatus for implementing distributed neural network training based on the nGraph framework, an electronic device, and a computer-readable storage medium. Background Art
[0002] Currently, DNN (Deep Neural Network) has been widely applied, including image and video classification, speech recognition, and language translation. However, with the increasingly widespread development and use of deep neural networks, the model size has become larger and larger. For example, it can reach hundreds of layers, with a total of 10 million to 20 million parameters. This growth has made efficient model training more important. The emergence of deep learning frameworks such as TensorFlow and PyTorch, as well as various hardware accelerators such as GPUs, FPGAs, and ASIC chips, has made great contributions to the improvement of neural network training performance. However, the working principles, development, and optimization methods vary greatly between different deep learning frameworks and different hardware acceleration devices. When developers want to replace the deep learning framework or deploy the deep learning model to other more advanced devices during the development process, they need to spend a lot of energy and time on migration and optimization. To address the above problems, Intel has introduced the nGraph framework, which is a deep neural network model compiler for various devices and frameworks, and can greatly simplify the complexity of cross-framework and hardware platform implementation of deep learning performance optimization, expanding the applicability and portability of deep learning models. Currently, the nGraph framework already supports or is developing support for front-end deep learning frameworks such as TensorFlow, MXNet, and PaddlePaddle, and already supports or is developing support for back-end hardware acceleration devices such as CPUs, NNP, and various GPUs.
[0003] Adding support for FPGA devices in the nGraph framework can enable developers based on the nGraph framework to utilize FPGA devices to accelerate the training of deep learning neural networks and further improve performance. However, in the related art, only single-machine training of deep learning neural networks can be performed in the FPGA back-end device through the nGraph framework, and multi-machine and multi-card distributed training is not supported.
[0004] Therefore, how to implement distributed training of deep learning neural networks in the FPGA back-end device through the nGraph framework is a technical problem that needs to be solved by those skilled in the art. Summary of the Invention
[0005] The purpose of this application is to provide a method and device for implementing distributed neural network training based on the nGraph framework, an electronic device, and a computer-readable storage medium, which realizes distributed training of deep learning neural networks in FPGA backend devices through the nGraph framework.
[0006] To achieve the above purpose, this application provides a method for implementing distributed neural network training based on the nGraph framework, including:
[0007] Integrate the OpenCL standard API library and the Intel IKL platform environment into the nGraph framework;
[0008] Determine the operators required for neural network training, add the class definitions corresponding to the operators in the nGraph framework, and add the kernel implementations corresponding to the operators in the FPGA backend device;
[0009] Create corresponding processes for each server in the cluster to manage multiple FPGA backend devices in the corresponding server in a loop using the processes;
[0010] During the distributed neural network training process, distribute the input data of the neural network training to each server so that each server distributes the obtained input data to multiple FPGA backend devices included.
[0011] Among them, the determining the operators required for neural network training, adding the class definitions corresponding to the operators in the nGraph framework, and adding the kernel implementations corresponding to the operators in the FPGA backend device includes:
[0012] Determine the operators required for neural network training, determine the first target operators supported by the nGraph framework but not supported by the FPGA backend device among the operators, and determine the second target operators not supported by both the nGraph framework and the FPGA backend device among the operators;
[0013] Add the class definitions corresponding to the second target operators in the nGraph framework, and add the kernel implementations corresponding to the first target operators and the second target operators in the FPGA backend device.
[0014] Among them, determining the first target operators supported by the nGraph framework but not supported by the FPGA backend device among the operators, and determining the second target operators not supported by both the nGraph framework and the FPGA backend device among the operators includes:
[0015] Obtain the list of operators supported by the nGraph framework and the list of operators supported by the FPGA backend device;
[0016] Determine the first list of operators not supported by the nGraph framework by comparing the operators required for distributed neural network training with the list of operators supported by the nGraph framework;
[0017] Determine the second list of operators not supported by the FPGA backend device by comparing the operators required for distributed neural network training with the list of operators supported by the FPGA backend device;
[0018] Determine the first target operators supported by the nGraph framework but not supported by the FPGA backend device and the second target operators not supported by both the nGraph framework and the FPGA backend device by comparing the first list of operators and the second list of operators.
[0019] Among them, the operator at least includes a communication operator for synchronizing weight data between multiple FPGA backend devices.
[0020] Among them, it also includes:
[0021] According to the forward calculation and backward propagation process of the neural network, sequentially concatenate the operators required for neural network training to construct a distributed training graph.
[0022] Among them, the step of distributing the input data for neural network training to each of the servers includes:
[0023] Calculate the amount of input data obtained by each process corresponding to each server according to the number of FPGA backend devices included in each server and the number of samples selected for a single training of the FPGA backend device;
[0024] Determine the starting position of the input data obtained by each process corresponding to each server according to the current data file pointer position, the process number of each process corresponding to each server, and the amount of input data obtained by each process corresponding to each server;
[0025] Judge whether the data set has been read completely; if so, reset the current data file pointer to the starting position of the data set, and distribute the input data to each server according to the starting position of the input data obtained by each process corresponding to each server and the amount of input data obtained; if not, directly distribute the input data to each server according to the starting position of the input data obtained by each process corresponding to each server and the amount of input data obtained.
[0026] To achieve the above object, the present application provides an apparatus for implementing distributed neural network training based on the nGraph framework, including:
[0027] An integration module, configured to integrate the OpenCL standard API library and the Intel IKL platform environment into the nGraph framework;
[0028] An operator sorting and adding module, configured to determine the operators required for neural network training, add the class definitions corresponding to the operators in the nGraph framework, and add the kernel implementations corresponding to the operators in the FPGA backend device;
[0029] A multi-device management module, configured to create corresponding processes for each server in the cluster, so as to circularly manage multiple FPGA backend devices in the corresponding server by using the processes;
[0030] An input data distributed processing module, configured to distribute the input data for neural network training to multiple FPGA backend devices in the server during the distributed neural network training process.
[0031] Wherein, it further includes:
[0032] A graph construction module, configured to sequentially connect the operators required for neural network training according to the forward calculation and backward propagation processes of the neural network, so as to construct a distributed training graph.
[0033] To achieve the above object, the present application provides an electronic device, including:
[0034] A memory, configured to store a computer program;
[0035] A processor, configured to implement the steps of the method for implementing distributed neural network training based on the nGraph framework as described above when executing the computer program.
[0036] To achieve the above object, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for implementing distributed neural network training based on the nGraph framework as described above are implemented.
[0037] As can be seen from the above solution, a method for implementing distributed neural network training based on the nGraph framework provided by this application includes: integrating the OpenCL standard API library and the Intel IKL platform environment into the nGraph framework; determining the operators required for neural network training, adding class definitions corresponding to the operators in the nGraph framework, and adding kernel implementations corresponding to the operators in the FPGA backend devices; creating corresponding processes for each server in the cluster to manage multiple FPGA backend devices in the corresponding server in a loop; during the distributed neural network training process, distributing the input data of the neural network training to each server so that each server distributes the obtained input data to multiple FPGA backend devices included therein.
[0038] In this application, based on the Intel IKL platform environment, synchronous communication between FPGA devices is achieved. Class definitions corresponding to the operators required for neural network training are added to the nGraph framework, and kernel implementations corresponding to the operators are added to the FPGA backend devices. At the same time, multiple backend devices are managed through the OpenCL standard API library. Each process is responsible for managing multiple FPGA backend devices in the corresponding server in a loop, that is, obtaining the input data of the neural network training and then distributing it to multiple FPGA backend devices in this server. It can be seen that this application realizes that the nGraph framework supports multiple FPGA backend devices, and thus realizes distributed training of deep learning neural networks in the FPGA backend devices through the nGraph framework. This application also discloses a device for implementing distributed neural network training based on the nGraph framework, an electronic device, and a computer-readable storage medium, which can also achieve the above technical effects.
[0039] It should be understood that the above general description and the following detailed description are only exemplary and do not limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. The drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification, and are used together with the following specific implementation manners to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the drawings:
[0041] Figure 1Flowchart of a method for implementing distributed neural network training based on the nGraph framework shown according to an exemplary embodiment;
[0042] Figure 2 Flowchart of the sorting and addition of an OP operator shown according to an exemplary embodiment;
[0043] Figure 3 Schematic diagram of a multi-device management shown according to an exemplary embodiment;
[0044] Figure 4 Example diagram of the distributed processing of input data shown according to an exemplary embodiment;
[0045] Figure 5 Flowchart of the distributed processing of input data shown according to an exemplary embodiment;
[0046] Figure 6 nGraph framework diagram shown according to an exemplary embodiment;
[0047] Figure 7 Structural diagram of a device for implementing distributed neural network training based on the nGraph framework shown according to an exemplary embodiment;
[0048] Figure 8 Structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application. In addition, in the embodiments of the present application, "first", "second", etc. are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence.
[0050] The embodiments of the present application disclose a method for implementing distributed neural network training based on the nGraph framework, which realizes the distributed training of deep learning neural networks in the FPGA backend device through the nGraph framework.
[0051] See Figure 1 , the flowchart of a method for implementing distributed neural network training based on the nGraph framework shown according to an exemplary embodiment, as Figure 1 shown, includes:
[0052] S101: Integrate the OpenCL standard API library and the Intel IKL platform environment into the nGraph framework;
[0053] OpenCL (Open Computing Language) is a standard API (Application Programming Interface) and programming language for parallel computing on heterogeneous devices. Compared with traditional FPGA algorithm development and HLS development, developing FPGA backend devices based on OpenCL high-level synthesis programming software can greatly simplify the FPGA development process and shorten the development cycle. This step aims to integrate the OpenCL standard API library into the nGraph framework for subsequent development and use of FPGA backend devices.
[0054] Among them, integrating the OpenCL standard API library into the nGraph framework can include: adding the OpenCL standard API library to the source code of the nGraph framework; modifying the cmake compilation file of the nGraph framework to compile the OpenCL standard API library into a dynamic link library within the nGraph framework. Specifically, first add the OpencCL standard API library to the source code of the nGraph framework. Since the OpenCL standard API library is used for the development of FPGA backend devices, in the source code of the nGraph framework, add the OpenCL standard API library to the same directory as the FPGA backend device. After adding the OpenCL standard API library to the source code of the nGraph framework, further modify the cmake compilation file of the nGraph framework to compile the OpenCL standard API library into a dynamic link library within the nGraph framework. In this way, the OpenCL standard API library is integrated with the nGraph framework and can be used by other modules in the nGraph framework.
[0055] The IO Channel read-write technology provided in the Intel IKL platform environment can achieve point-to-point communication between kernels across devices. In this embodiment, the node communication operator FPGA_Allreduce and the average operator Average between FPGA backend devices are implemented through the Intel IKL platform environment, thereby realizing the weight data synchronization communication between multiple devices in a distributed cluster.
[0056] Among them, integrating the Intel IKL environment into the nGraph framework may include: First, complete the installation of the Intel IKL platform environment in the server, then modify the CMake compilation file under the nGraph framework, add the paths where the Intel IKL header files and library files are located therein, and at the same time add some necessary dependent libraries of IKL. Finally, recompile and install the nGraph framework.
[0057] S102: Determine the operators required for neural network training, add the class definitions corresponding to the operators in the nGraph framework, and add the kernel implementations corresponding to the operators in the FPGA backend device;
[0058] In this step, add the class definitions corresponding to the OP operators required for neural network training in the nGraph framework, and add the kernel implementations corresponding to these OP operators in the FPGA backend device.
[0059] As a preferred implementation manner, this step includes: determining the operators required for neural network training, and determining a first target operator supported by the nGraph framework and not supported by the FPGA backend device among the operators, and determining a second target operator not supported by both the nGraph framework and the FPGA backend device among the operators; adding the class definitions corresponding to the second target operator in the nGraph framework, and adding the kernel implementations corresponding to the first target operator and the second target operator in the FPGA backend device.
[0060] It can be understood that there are already class definitions corresponding to some operators in the nGraph framework, and there is no need to add them repeatedly when these operators are required for neural network training. Similarly, there are already kernel implementations corresponding to some operators in the FPGA backend device, and there is no need to add them repeatedly when these operators are required for neural network training. In specific implementation, the process of sorting out and adding OP operators is as Figure 2 shown. Determine a first target operator supported by the nGraph framework and not supported by the FPGA backend device among the operators required for neural network training, and determine a second target operator not supported by both the nGraph framework and the FPGA backend device. For the first target operator, only need to add the corresponding kernel implementation in the FPGA backend device. For the second target operator, not only need to add the corresponding kernel implementation in the FPGA backend device, but also need to add the corresponding class definition in the nGraph framework.
[0061] As a feasible implementation manner, determining a first target operator supported by the nGraph framework but not supported by the FPGA backend device, and determining a second target operator not supported by both the nGraph framework and the FPGA backend device in the operator includes: obtaining an operator list supported by the nGraph framework and an operator list supported by the FPGA backend device; determining a first operator list not supported by the nGraph framework by comparing the operators required for distributed neural network training with the operator list supported by the nGraph framework; determining a second operator list not supported by the FPGA backend device by comparing the operators required for distributed neural network training with the operator list supported by the FPGA backend device; determining the first target operator supported by the nGraph framework but not supported by the FPGA backend device and the second target operator not supported by both the nGraph framework and the FPGA backend device by comparing the first operator list and the second operator list.
[0062] In a specific implementation, a new operator sorting function is added. This function combines the operator list supported by the nGraph framework and the operator list supported by the FPGA backend device through the obtained neural network, and sequentially judges each operator required for neural network training, automatically sorting out the above first target operator and second target operator.
[0063] It should be noted that an OP operator that must be added in this embodiment is FPGA_Allreduce. The FPGA_Allreduce operator is a collective communication operator. Its goal is to integrate the data in different computing nodes and then distribute the results to each node, so that each computing node finally has the integrated data.
[0064] S103: Create corresponding processes for each server in the cluster to manage multiple FPGA backend devices in the corresponding server by using the processes in a loop.
[0065] In this embodiment, the management of multiple devices is implemented through the OpenCL standard API library. As Figure 3 shown, first obtain the number of servers in the cluster and the number of FPGA devices on each server, and then create multiple processes using the OpenCL standard API library according to the number of servers in the cluster. Each process is responsible for managing one server, and each server has one or more FPGA devices. The FPGA devices on the same server are uniformly managed by the process to which the current server belongs through a for loop.
[0066] S104: During the distributed neural network training process, distribute the input data for neural network training to each of the servers, so that each server distributes the obtained input data to multiple FPGA backend devices it contains.
[0067] During distributed training, each FPGA backend device needs to obtain different input data for calculation. In this step, divide the input data according to the number of server nodes. The distributed data processing solution is as Figure 4 shown. Each server process determines the amount of data to be obtained and the starting position of the data to be obtained in the data file according to its current process number and the number of FPGA backend devices on the server, and obtains the data in sequence. After obtaining the data on the server, distribute it to different FPGA backend devices. Figure 4 Taking two machines with four cards as an example, in the first iteration process, sever0 obtains data blocks AB, then distributes A to FPGA0 and B to FPGA1. Sever1 obtains data blocks CD, then distributes C to FPGA0 and D to FPGA1; in the second iteration process, sever0 obtains data blocks EF, sever0 obtains data blocks GH, and so on.
[0068] As a feasible implementation manner, this step includes: calculating the amount of input data obtained by each process corresponding to each server according to the number of FPGA backend devices contained in each server and the number of samples selected for a single training by a single FPGA backend device; determining the starting position of the input data obtained by each process corresponding to each server according to the current data file pointer position, the process number of each process corresponding to each server, and the amount of input data obtained by each process corresponding to each server; determining whether the data set has been read completely; if so, reset the current data file pointer to the starting position of the data set, and distribute the input data to each server according to the starting position of the input data obtained by each process corresponding to each server and the amount of input data obtained; if not, directly distribute the input data to each server according to the starting position of the input data obtained by each process corresponding to each server and the amount of input data obtained.
[0069] In a specific implementation, the distributed processing solution of the input data is as Figure 5As shown in the figure, first, obtain the total number of processes rank_size for the FPGA cluster startup and the current server process number rank. Calculate the amount of data to be fetched each time for each server process according to the single-node batch_size and the number of FPGA backend devices per server num_device_per_server: batch_size × num_device_per_server. Secondly, determine the starting position for each server process to fetch data according to the current process number rank: rank × batch_size × num_device_per_server. Then, determine whether the current epoch dataset has been read completely. If so, reset the data file pointer to the starting position, and reset the starting position for the current server process to fetch data, and then read the data; if not, directly read the data. Finally, distribute the data read by the current server process to each FPGA backend device on the current server, update the position for the current server process to fetch data next time, and enter the step of determining whether the current epoch dataset has been read completely.
[0070] On the basis of the above embodiments, as a preferred embodiment, it further includes: sequentially connecting the operators required for neural network training according to the forward calculation and backward propagation processes of the neural network to construct a distributed training graph. The process of constructing the graph for distributed training is a process of sequentially connecting the required op operators according to the forward calculation and backward propagation processes of the used neural network.
[0071] After completing the above development steps, the support of the nGraph framework for distributed training on the FPGA backend can be realized. The entire framework is as Figure 6As shown in the figure, it includes an external dependency basic module, a client front-end, and an FPGA back-end. The external dependency integration module mainly includes the integration of the OpenMPI library (i.e., the OpenCL standard API library) and the Intel IKL environment into the nGraph framework. The client front-end includes an input data distributed processing module, a Graph construction module, and a training start module. The input data distributed processing module is used to distribute the input data to each server. The Graph construction module is used to construct a distributed training graph. The training start module is used to start the training of the distributed neural network. The FPGA back-end includes a multi-device management module and a required OP sorting and adding module. The multi-device management module is used to create corresponding processes for each server. The required OP sorting and adding module is used to add the class definitions corresponding to the operators required for neural network training in the nGraph framework and add the kernel implementations corresponding to the operators in the FPGA back-end devices. nGraph client users can develop programs according to their original programming habits. They only need to specify the back-end device as the FPGA device when creating the backend, and at the same time set the total number of FPGA devices to be used in the distributed training cluster and the number of FPGA devices on each server, then they can use the FPGA back-end to accelerate the distributed training of the deep learning neural network constructed by the user.
[0072] In the embodiment of the present application, based on the Intel IKL platform environment, synchronous communication between FPGA devices is realized. The class definitions corresponding to the operators required for neural network training are added in the nGraph framework, and the kernel implementations corresponding to the operators are added in the FPGA back-end devices. At the same time, multiple back-end devices are managed through the OpenCL standard API library. Each process is responsible for circularly managing multiple FPGA back-end devices in the corresponding server, that is, obtaining the input data for neural network training and then distributing it to multiple FPGA back-end devices in this server. It can be seen that the embodiment of the present application realizes that the nGraph framework supports multiple FPGA back-end devices, and thus realizes the distributed training of deep learning neural networks in the FPGA back-end devices through the nGraph framework.
[0073] Next, a device for implementing distributed neural network training based on the nGraph framework provided by the embodiment of the present application will be introduced. The device for implementing distributed neural network training based on the nGraph framework described below can be referred to each other with the method for implementing distributed neural network training based on the nGraph framework described above.
[0074] See Figure 7 , the structural diagram of a device for implementing distributed neural network training based on the nGraph framework shown according to an exemplary embodiment, as Figure 7As shown in the figure, it includes:
[0075] An integration module 701, which is used to integrate the OpenCL standard API library and the Intel IKL platform environment into the nGraph framework;
[0076] An operator sorting and adding module 702, which is used to determine the operators required for neural network training, add class definitions corresponding to the operators in the nGraph framework, and add kernel implementations corresponding to the operators in the FPGA backend device;
[0077] A multi-device management module 703, which is used to create corresponding processes for each server in the cluster, so as to use the processes to circularly manage multiple FPGA backend devices in the corresponding servers;
[0078] An input data distributed processing module 704, which is used to distribute the input data of neural network training to multiple FPGA backend devices in the server during the distributed neural network training process.
[0079] In the embodiment of the present application, based on the Intel IKL platform environment, synchronous communication between FPGA devices is realized, class definitions corresponding to the operators required for neural network training are added to the nGraph framework, and kernel implementations corresponding to the operators are added to the FPGA backend device. At the same time, multiple backend devices are managed through the OpenCL standard API library, and each process is responsible for circularly managing multiple FPGA backend devices in the corresponding server, that is, obtaining the input data of neural network training and then distributing it to multiple FPGA backend devices in the server. It can be seen that the embodiment of the present application realizes that the nGraph framework supports multiple FPGA backend devices, and further realizes distributed training of deep learning neural networks in the FPGA backend device through the nGraph framework.
[0080] On the basis of the above embodiment, as a preferred implementation manner, the operator sorting and adding module 702 includes:
[0081] A first determination unit, which is used to determine the operators required for neural network training, determine the first target operators supported by the nGraph framework but not supported by the FPGA backend device among the operators, and determine the second target operators not supported by both the nGraph framework and the FPGA backend device among the operators;
[0082] An adding unit, which is used to add class definitions corresponding to the second target operators in the nGraph framework, and add kernel implementations corresponding to the first target operators and the second target operators in the FPGA backend device.
[0083] Based on the above embodiments, as a preferred implementation manner, the first determination unit includes:
[0084] A first determination subunit, configured to determine operators required for neural network training;
[0085] An acquisition subunit, configured to acquire an operator list supported by the nGraph framework and an operator list supported by the FPGA backend device;
[0086] A second determination subunit, configured to determine a first operator list not supported by the nGraph framework by comparing the operators required for distributed neural network training with the operator list supported by the nGraph framework;
[0087] A third determination subunit, configured to determine a second operator list not supported by the FPGA backend device by comparing the operators required for distributed neural network training with the operator list supported by the FPGA backend device;
[0088] A fourth determination subunit, configured to determine a first target operator supported by the nGraph framework but not supported by the FPGA backend device and a second target operator not supported by both the nGraph framework and the FPGA backend device by comparing the first operator list and the second operator list.
[0089] Based on the above embodiments, as a preferred implementation manner, the operator at least includes a communication operator for synchronizing weight data between multiple FPGA backend devices.
[0090] Based on the above embodiments, as a preferred implementation manner, the input data distributed processing module 704 includes:
[0091] A calculation unit, configured to calculate the amount of input data acquired by each process corresponding to each server according to the number of FPGA backend devices included in each server and the number of samples selected for a single FPGA backend device training at one time;
[0092] A second determination unit, configured to determine the starting position of the input data acquired by each process corresponding to each server according to the current data file pointer position, the process number of each process corresponding to each server, and the amount of input data acquired by each process corresponding to each server;
[0093] A distribution unit for determining whether the reading of the data set is completed; if so, resetting the current data file pointer to the starting position of the data set, and distributing the input data to each of the servers according to the starting position of the input data and the quantity of the acquired input data corresponding to the process of each server; if not, directly distributing the input data to each of the servers according to the starting position of the input data and the quantity of the acquired input data corresponding to the process of each server.
[0094] Based on the above embodiments, as a preferred implementation, it further includes:
[0095] A graph construction module for sequentially connecting the operators required for neural network training according to the forward calculation and backward propagation processes of the neural network to construct a distributed training graph.
[0096] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here in detail.
[0097] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiments of the present application, the embodiments of the present application further provide an electronic device, Figure 8 As shown in the structural diagram of an electronic device according to an exemplary embodiment, such as Figure 8 shown, the electronic device includes:
[0098] A communication interface 1 capable of interacting with other devices such as network devices.
[0099] A processor 2 connected to the communication interface 1 to implement information interaction with other devices, and when running a computer program, executing the method for implementing distributed neural network training based on the nGraph framework provided by one or more of the above technical solutions. And the computer program is stored on a memory 3.
[0100] Of course, in actual application, the various components in the electronic device are coupled together through a bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. The bus system 4 includes, in addition to the data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 8 all the various buses are labeled as the bus system 4.
[0101] The memory 3 in the embodiments of the present application is used to store various types of data to support the operation of the electronic device. Examples of these data include: any computer program for operating on the electronic device.
[0102] It can be understood that the memory 3 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM, Read Only Memory), a programmable read-only memory (PROM, Programmable Read-Only Memory), an erasable programmable read-only memory (EPROM, Erasable Programmable Read-Only Memory), an electrically erasable programmable read-only memory (EEPROM, Electrically Erasable Programmable Read-Only Memory), a ferromagnetic random access memory (FRAM, ferromagnetic random access memory), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM, Compact Disc Read-Only Memory); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM, Random Access Memory), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a static random access memory (SRAM, Static Random Access Memory), a synchronous static random access memory (SSRAM, Synchronous Static Random Access Memory), a dynamic random access memory (DRAM, Dynamic Random Access Memory), a synchronous dynamic random access memory (SDRAM, Synchronous Dynamic Random Access Memory), a double data rate synchronous dynamic random access memory (DDR SDRAM, Double Data Rate Synchronous Dynamic Random Access Memory), an enhanced synchronous dynamic random access memory (ESDRAM, Enhanced Synchronous Dynamic Random Access Memory), a sync link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), a direct rambus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 2 described in the embodiments of the present application is intended to include, but is not limited to, these and any other suitable types of memories.
[0103] The method disclosed in the embodiments of the present application above can be applied to the processor 2 or implemented by the processor 2. The processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 2 or by instructions in the form of software. The above-mentioned processor 2 may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 2 can implement or execute each method, step, and logic block diagram disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present application, it can be directly embodied as being executed and completed by the hardware decoding processor, or by a combination of the hardware and software modules in the decoding processor. The software module may be located in the storage medium, and this storage medium is located in the memory 3. The processor 2 reads the program in the memory 3 and combines its hardware to complete the steps of the foregoing method.
[0104] When the processor 2 executes the program, it implements the corresponding processes in each method of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.
[0105] In an exemplary embodiment, the embodiments of the present application also provide a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as the memory 3 including a stored computer program. The above computer program can be executed by the processor 2 to complete the steps of the foregoing method. The computer-readable storage medium may be a FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.
[0106] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as mobile storage devices, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0107] Alternatively, if the above integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application essentially or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
[0108] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for implementing distributed neural network training based on the nGraph framework, characterized in that, it includes: Integrate the OpenCL standard API library and the Intel IKL platform environment into the nGraph framework; Determine the operators required for neural network training, add class definitions corresponding to the operators in the nGraph framework, and add kernel implementations corresponding to the operators in the FPGA backend device; Create corresponding processes for each server in the cluster to manage multiple FPGA backend devices in the corresponding server using the processes in a loop; During the distributed neural network training process, distribute the input data of the neural network training to each of the servers, so that each server distributes the obtained input data to multiple FPGA backend devices included; Among them, the determining the operators required for neural network training, adding class definitions corresponding to the operators in the nGraph framework, and adding kernel implementations corresponding to the operators in the FPGA backend device includes: Determine the operators required for neural network training, and determine the first target operators supported by the nGraph framework but not supported by the FPGA backend device among the operators, and determine the second target operators not supported by both the nGraph framework and the FPGA backend device among the operators; Add class definitions corresponding to the second target operators in the nGraph framework, and add kernel implementations corresponding to the first target operators and the second target operators in the FPGA backend device; Among them, the distributing the input data of the neural network training to each of the servers includes: Calculate the amount of input data obtained by each process corresponding to each server according to the number of FPGA backend devices included in each server and the number of samples selected for one training by a single FPGA backend device; Determine the starting position of the input data obtained by each process corresponding to each server according to the current data file pointer position, the process number of each process corresponding to each server, and the amount of input data obtained by each process corresponding to each server; Judge whether the dataset has been read completely; if so, reset the current data file pointer to the starting position of the dataset, and distribute the input data to each server according to the starting position of the input data obtained by each process corresponding to each server and the amount of input data obtained; if not, directly distribute the input data to each server according to the starting position of the input data obtained by each process corresponding to each server and the amount of input data obtained.
2. The method according to claim 1, characterized in that, Determining the first target operators supported by the nGraph framework but not supported by the FPGA backend device among the operators, and determining the second target operators not supported by both the nGraph framework and the FPGA backend device among the operators includes: Obtain the operator list supported by the nGraph framework and the operator list supported by the FPGA backend device; Determine a first list of operators not supported by the nGraph framework by comparing the operators required for distributed neural network training with the list of operators supported by the nGraph framework; Determine a second list of operators not supported by the FPGA backend device by comparing the operators required for distributed neural network training with the list of operators supported by the FPGA backend device; Determine a first target operator supported by the nGraph framework but not supported by the FPGA backend device and a second target operator not supported by both the nGraph framework and the FPGA backend device by comparing the first list of operators and the second list of operators; 3. The method according to claim 1, wherein, The operator at least includes a communication operator for synchronizing weight data between multiple FPGA backend devices.
4. The method according to claim 1, wherein, It further includes: According to the forward calculation and backward propagation processes of the neural network, sequentially concatenate the operators required for neural network training to construct a distributed training graph.
5. A device for implementing distributed neural network training based on the nGraph framework, wherein, It includes: An integration module for integrating the OpenCL standard API library and the Intel IKL platform environment into the nGraph framework; An operator sorting and adding module for determining the operators required for neural network training, adding class definitions corresponding to the operators in the nGraph framework, and adding kernel implementations corresponding to the operators in the FPGA backend device; A multi-device management module for creating corresponding processes for each server in the cluster to manage multiple FPGA backend devices in the corresponding server using the processes in a loop; An input data distributed processing module for distributing the input data for neural network training to multiple FPGA backend devices in the server during the distributed neural network training process; wherein, the operator sorting and adding module includes: A first determination unit for determining the operators required for neural network training, determining a first target operator supported by the nGraph framework but not supported by the FPGA backend device among the operators, and determining a second target operator not supported by both the nGraph framework and the FPGA backend device among the operators; An adding unit for adding class definitions corresponding to the second target operator in the nGraph framework and adding kernel implementations corresponding to the first target operator and the second target operator in the FPGA backend device; wherein, the input data distributed processing module includes: A calculation unit for calculating the amount of input data obtained by each corresponding process of the server according to the number of FPGA backend devices included in each server and the number of samples selected for a single training of an FPGA backend device; A second determination unit, configured to determine the starting position of the input data obtained by the process corresponding to each server according to the current data file pointer position, the process number of the process corresponding to each server, and the amount of input data obtained by the process corresponding to each server each time. A distribution unit, configured to determine whether the dataset has been completely read; if so, reset the current data file pointer to the starting position of the dataset, and distribute the input data to each server according to the starting position of the input data obtained by the process corresponding to each server and the amount of input data obtained; if not, directly distribute the input data to each server according to the starting position of the input data obtained by the process corresponding to each server and the amount of input data obtained.
6. The apparatus according to claim 5, wherein, it further comprises: A graph construction module, configured to sequentially connect the operators required for neural network training according to the forward calculation and backward propagation processes of the neural network, so as to construct a distributed training graph.
7. An electronic device, wherein, it comprises: A memory, configured to store a computer program; A processor, configured to implement the steps of the method according to any one of claims 1 to 4 when executing the computer program.
8. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Digital circuits for evaluating neural engineering framework style neural networks
CA3051429A1
Method for realizing that nGraph framework supports FPGA rear-end equipment
CN112001494A