A method, system, device, equipment and medium for training a deep neural network

By transmitting the deep neural network training tasks to the target FPGA and performing them in parallel using the nGraph framework, the problem of high power consumption and low parallelism in deep neural network training is solved, and the training effect with low power consumption and high parallelism is achieved, and the applicability is improved.

CN114139702BActive Publication Date: 2025-06-20GUANGDONG INSPUR BIG DATA RES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111415573.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-25
Publication Date
2025-06-20
Estimated Expiration
2041-11-25

AI Technical Summary

Technical Problem

There are problems of high power consumption and low parallelism during deep neural network training, resulting in poor applicability.

Method used

The deep neural network training tasks are transmitted to the target FPGA through the nGraph framework, and the FPGA is used for parallel execution, thereby reducing power consumption and improving training efficiency.

Benefits of technology

Low-power consumption and high-parallel deep neural network training is realized, which improves the applicability of the training method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114139702B_ABST
    Figure CN114139702B_ABST
Patent Text Reader

Abstract

The present application discloses a method, system, device, equipment and computer medium for deep neural network training, which is applied to a connector to obtain a deep neural network training task sent by a deep learning framework; and transmits the deep neural network training task to a target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task. In the present application, the connector can transmit the deep neural network training task sent by the deep learning framework to the target FPGA for execution with the help of the nGraph framework, and can complete the deep neural network training with low power consumption and high parallelism, and has good applicability. A deep neural network training system, device, equipment and computer-readable storage medium provided by the present application also solve the corresponding technical problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and more specifically, to a method, system, device, equipment, and medium for training a deep neural network. Background Art

[0002] In the past few years, deep neural networks (DNNs) have been widely applied, including image and video classification, speech recognition, and language translation. However, as DNNs are developed and used more and more widely, the model size has become larger and larger, making efficient model training more important. The emergence of deep learning frameworks such as tensorflow and pytorch, as well as various hardware accelerators such as GPUs and ASIC chips, has made great contributions to the improvement of neural network training performance. However, the training process of deep neural networks has the disadvantages of high power consumption and low parallelism, and poor applicability.

[0003] In summary, how to improve the applicability of the deep neural network training method is an urgent problem to be solved by those skilled in the art currently. Summary of the Invention

[0004] The purpose of this application is to provide a method for training a deep neural network, which can, to a certain extent, solve the technical problem of how to improve the applicability of the deep neural network training method. This application also provides a system, device, equipment, and computer-readable storage medium for training a deep neural network.

[0005] To achieve the above purpose, this application provides the following technical solutions:

[0006] A method for training a deep neural network, applied to a connector, includes:

[0007] Obtain a deep neural network training task sent by a deep learning framework;

[0008] Transmit the deep neural network training task to a target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task.

[0009] Preferably, the obtaining of the deep neural network training task sent by the deep learning framework includes:

[0010] Obtain each deep neural network training subtask sent by the deep learning framework through each subprocess;

[0011] Concatenate all the deep neural network training subtasks into the deep neural network training task in the main process.

[0012] Preferably, transmitting the deep neural network training task to the target FPGA through the nGraph framework so that the target FPGA executes the deep neural network training task includes:

[0013] Based on the main process, transmitting the deep neural network training task to the target FPGA through the nGraph framework so that each FPGA in the target FPGA executes the corresponding deep neural network training subtask.

[0014] Preferably, after transmitting the deep neural network training task to the target FPGA based on the main process through the nGraph framework, it further includes:

[0015] Based on the main process, receiving the task execution result transmitted by the target FPGA through the nGraph framework, where the task execution result includes the sub-execution results of each FPGA for the corresponding deep neural network training subtask;

[0016] Splitting the task execution result into the sub-execution results corresponding to each sub-process;

[0017] Transmitting the sub-execution result to the deep learning framework based on the sub-process.

[0018] A deep neural network training system, applied to a connector, includes:

[0019] A first acquisition module, configured to acquire the deep neural network training task sent by the deep learning framework;

[0020] A first transmission module, configured to transmit the deep neural network training task to the target FPGA through the nGraph framework so that the target FPGA executes the deep neural network training task.

[0021] A deep neural network training device includes:

[0022] A deep learning framework, configured to acquire the deep neural network training task;

[0023] A connector connected to the deep learning framework, configured to transmit the deep neural network training task to the nGraph framework;

[0024] The nGraph framework connected to the connector, configured to transmit the deep neural network training task to the target FPGA;

[0025] The target FPGA connected to the nGraph framework, configured to execute the deep neural network training task.

[0026] Preferably, the connector includes a Bridge.

[0027] Preferably, the deep learning framework includes tensorflow.

[0028] A deep neural network training device includes:

[0029] A memory for storing a computer program;

[0030] A processor for implementing the steps of any of the above deep neural network training methods when executing the computer program.

[0031] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above deep neural network training methods are implemented.

[0032] A deep neural network training method provided by this application is applied to a connector, and a deep neural network training task sent by a deep learning framework is obtained; the deep neural network training task is transmitted to a target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task. In this application, the connector can transmit the deep neural network training task sent by the deep learning framework to the target FPGA for execution through the nGraph framework, and can complete the deep neural network training with low power consumption and high parallelism, and has good applicability. A deep neural network training system, device, equipment, and computer-readable storage medium provided by this application also solve the corresponding technical problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0034] Figure 1 It is a flowchart of a deep neural network training method provided by an embodiment of the present application;

[0035] Figure 2 It is a flowchart of a deep neural network training system provided by an embodiment of the present application;

[0036] Figure 3 It is a flowchart of a deep neural network training device provided by an embodiment of the present application;

[0037] Figure 4 It is a design diagram of the support module for distributed support of each layer of the nGraph framework;

[0038] Figure 5 It is a design diagram of the inter - process communication module for the Bridge middle layer;

[0039] Figure 6 It is a schematic diagram of the interaction between Bridge and tensorflow in this application;

[0040] Figure 7 It is a schematic diagram of the structure of a deep neural network training device provided by an embodiment of this application;

[0041] Figure 8 It is another schematic diagram of the structure of a deep neural network training device provided by an embodiment of this application. Detailed implementation manners

[0042] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0043] Please refer to Figure 1 , Figure 1 It is a flowchart of a deep neural network training method provided by an embodiment of this application.

[0044] A deep neural network training method provided by an embodiment of this application, which is applied to a connector, may include the following steps:

[0045] Step S101: Obtain a deep neural network training task sent by a deep learning framework.

[0046] In practical applications, the connector can first obtain a deep neural network training task sent by the deep learning framework. The specific information of the connector, the deep learning framework, and the deep neural network training task can be determined according to actual needs, and this application does not make specific limitations here.

[0047] Step S102: Transmit the deep neural network training task to the target FPGA through the nGraph framework so that the target FPGA executes the deep neural network training task.

[0048] In practical applications, after obtaining the deep neural network training task sent by the deep learning framework, the deep neural network training task can be transmitted to the target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task. That is, in this application, by connecting the deep learning framework, the connector, the nGraph framework, and the target FPGA, the function of parallelly executing the deep neural network training task with the target FPGA is realized, the power consumption is reduced, the deep neural network training efficiency is improved, and the applicability is good.

[0049] A deep neural network training method provided by this application is applied to a connector to obtain the deep neural network training task sent by the deep learning framework; and transmit the deep neural network training task to the target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task. In this application, the connector can transmit, through the nGraph framework, the deep neural network training task sent by the deep learning framework to the target FPGA for execution, and can complete the deep neural network training with low power consumption and high parallelism, and has good applicability.

[0050] A deep neural network training method provided by an embodiment of this application is applied to a connector. In the process of obtaining the deep neural network training task sent by the deep learning framework, since the deep learning framework sometimes needs to execute training tasks in parallel during deep neural network training, to meet this requirement, each deep neural network training subtask sent by the deep learning framework can be obtained through each subprocess; and all the deep neural network training subtasks are spliced into the deep neural network training task in the main process.

[0051] In a specific application scenario, in the process of transmitting the deep neural network training task to the target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task, based on the main process, the deep neural network training task can be transmitted to the target FPGA through the nGraph framework, so that each FPGA in the target FPGA executes the corresponding deep neural network training subtask.

[0052] In a specific application scenario, after transmitting the deep neural network training task to the target FPGA through the nGraph framework based on the main process, the task execution result transmitted by the target FPGA can also be received through the nGraph framework based on the main process. The task execution result includes the sub-execution results of each FPGA for the corresponding deep neural network training subtask; the task execution result is split into sub-execution results corresponding to each subprocess; and the sub-execution results are transmitted to the deep learning framework based on the subprocess.

[0053] Please refer to Figure 2 , Figure 2Flowchart of a deep neural network training system provided by an embodiment of the present application.

[0054] A deep neural network training system provided by an embodiment of the present application, which is applied to a connector, may include:

[0055] A first acquisition module 101, configured to acquire a deep neural network training task sent by a deep learning framework;

[0056] A first transmission module 102, configured to transmit the deep neural network training task to a target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task.

[0057] For the description of the corresponding modules in the deep neural network training system provided by an embodiment of the present application, reference may be made to the above description, which will not be elaborated herein.

[0058] Please refer to Figure 3 , Figure 3 Flowchart of a deep neural network training device provided by an embodiment of the present application.

[0059] A deep neural network training device provided by an embodiment of the present application may include:

[0060] A deep learning framework 11, configured to acquire a deep neural network training task;

[0061] A connector 12 connected to the deep learning framework 11, configured to transmit the deep neural network training task to the nGraph framework 13;

[0062] An nGraph framework 13 connected to the connector 12, configured to transmit the deep neural network training task to a target FPGA 14;

[0063] A target FPGA 14 connected to the nGraph framework 13, configured to execute the deep neural network training task.

[0064] For the description of the corresponding devices in the deep neural network training device provided by an embodiment of the present application, reference may be made to the above description, which will not be elaborated herein.

[0065] In the deep neural network training device provided by an embodiment of the present application, the connector may be a Bridge.

[0066] In the deep neural network training device provided by an embodiment of the present application, the deep learning framework may include tensorflow, etc.

[0067] In a specific application scenario, in the design of the support module solution for the nGraph FPGA backend device in the Bridge layer, it is necessary to modify the Bridge compilation file build_ngtf.py, add support for the FPGA backend in it, and then execute the build_ngtf.py script for compilation. Note to add the compilation condition "--build_fpga_backend" in the compilation command. After successful compilation, the support from the tensorflow front-end to the nGraph FPGA backend device is completed; at this time, the user enables the backend to be the FPGA device through the environment variable NGRAPH_TF_BACKEND, and adds the support statement import ngraph_bridge for nGraph in the tensorflow client-side training code, then the nGraph FPGA backend device can be used in the tensorflow front-end to accelerate the original tensorflow client training code, but only single-card training on a single FPGA is supported at this time.

[0068] In a specific application scenario, the design of the support module solution for distributed processing in each layer of the nGraph framework can be as Figure 4 shown. In the tensorflow front-end, the client distributed training code is implemented using the horovd multi-process method, and the mpirun command is used to start multiple processes. Each process runs the same TensorFlow client training code, that is, the tasks handled by each process are the same (only the input data to be processed is different), and it is required that the number of processes started here is the same as the number of backend FPGA devices. In the nGraph framework backend, since the FPGA OpenCL driver does not support the single-machine multi-process running mode, it is not possible to require multiple processes to handle the same task in this module. Therefore, multiple FPGA boards are placed in the same process (process 0) on a single server and managed and run in a for loop, while other processes do not involve processing related to the nGraph framework and FPGA devices. That is to say, in the nGraph framework backend, each process is responsible for handling different tasks. In the Bridge middle layer, the tasks handled by each process are also different. Among them, process 0, as the main process, includes all the processing procedures in the original Bridge code, and at the same time, the data interaction process with other processes is added in process 0, while other non-0 processes are simplified based on the original Bridge implementation, removing all interaction processes with the nGraph framework and backend devices, and adding the data interaction process with process 0.

[0069] In a specific application scenario, the design of the inter-process communication module solution in the Bridge middle layer can be as Figure 5As shown in the figure. Among them, each process communicates in the Bridge middle layer to exchange data, and the inter-process communication adopts the message queue method. Assume that there are N FPGA boards on a single server, then the tensorflow front-end starts N processes. Among them, the processing of FPGAs input data is as follows: input message queues in_queue_i (i is the corresponding non-zero process number) are respectively created between process 0 and other processes, which are used to sequentially transmit the input data to be processed from other processes to process 0. Specifically, process 0 receives the input data input_data_i of other processes through each input message queue, splices these data with the input data input_data_0 of process 0 itself to obtain the complete input data input_data_all, and then input_data_all is passed into the nGraph framework layer and distributed by the nGrpah framework layer to each FPGA device for processing. The processing of the FPGAs output result data is as follows: message queues out_queue_i are respectively created between process 0 and other processes, which are used to transmit the result data processed by the FPGA from process 0 to other processes respectively. Specifically, after the FPGA finishes processing, it returns the output result to the Brige layer. The Bridge layer obtains the complete result data res_data_all and splits it, and then sends the split result data to other processes correspondingly.

[0070] It should be noted that the information interaction process between Bridge and tensorflow can be determined according to actual needs. For example, its interaction process is as Figure 6 shown:

[0071] (1) Each process's Bridge obtains the Tensorflow graph graph and performs processing such as pass optimization, marking, and encapsulation on it;

[0072] (2) Each process obtains the pointers of each input TF tensors (i.e., the pointers to the positions of the input data to be processed) and the pointers of the TF output tensors (i.e., the pointers to the positions of the result data to be returned to the Tensorflow side);

[0073] (3) Obtain the current process number, and judge whether the current process is process 0. If it is, execute step (4). If not, jump to step (5) to execute;

[0074] (4) Processing related to process 0:

[0075] ① Map and translate the Tensorflow OP into the nGraph OP;

[0076] ② Call the nGraph framework for related processing, including: creating ops, constructing the function graph, creating the backend, setting the io_channel, compiling the function subgraph, and allocating buffers for each input and output op of the function subgraph on each FPGA board, etc.;

[0077] ③ Create or obtain the input message queues corresponding to each non-zero process;

[0078] ④ Read data from each input message queue;

[0079] ⑤ Concatenate the input data of each process. Note that the concatenated input data here includes the input data to be processed of process 0 and the input data to be processed of other non-zero processes;

[0080] ⑥ Call the nGraph framework's buffer writing function to write the concatenated input data into the buffers of each FPGA board, start the execution of each FPGA, and return the execution results to the Bridge layer;

[0081] ⑦ The Bridge layer obtains the result data and splits it;

[0082] ⑧ Create or obtain each output message queue;

[0083] ⑨ Write each result data into the corresponding output message queue

[0084] ⑩ Jump to step (6)

[0085] (5) Related processing for non-zero processes:

[0086] ① Create or obtain its input message queue in_queue_i according to the current process number (i is the current process number)

[0087] ② Write the input data into its input message queue input_queue_i

[0088] ③ Create or obtain its output message queue out_queue_i according to the current process number (i is the current process number)

[0089] ④ Read the result data from the output message queue out_queue_i

[0090] ⑤ Jump to step (6)

[0091] (6) Return to execute in the TF front end.

[0092] This application also provides a deep neural network training device and a computer-readable storage medium, both of which have the corresponding effects of a deep neural network training method provided by the embodiments of this application. Please refer to Figure 7 ,Figure 7 The structural schematic diagram of a deep neural network training device provided by an embodiment of the present application.

[0093] A deep neural network training device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented:

[0094] Obtain a deep neural network training task sent by a deep learning framework;

[0095] Transmit the deep neural network training task to a target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task.

[0096] A deep neural network training device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: Obtain each deep neural network training subtask sent by the deep learning framework through each subprocess; splice all the deep neural network training subtasks into the deep neural network training task in the main process.

[0097] A deep neural network training device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: Based on the main process, transmit the deep neural network training task to the target FPGA through the nGraph framework, so that each FPGA in the target FPGA executes the corresponding deep neural network training subtask.

[0098] A deep neural network training device provided by an embodiment of the present application includes a memory 201 and a processor 202. A computer program is stored in the memory 201. When the processor 202 executes the computer program, the following steps are implemented: After transmitting the deep neural network training task to the target FPGA through the nGraph framework based on the main process, receive the task execution result transmitted by the target FPGA through the nGraph framework based on the main process. The task execution result includes the sub-execution results of each FPGA for the corresponding deep neural network training subtask; split the task execution result into the sub-execution results corresponding to each subprocess; transmit the sub-execution results to the deep learning framework based on the subprocess.

[0099] Please refer to Figure 8, another deep neural network training device provided by an embodiment of the present application may further include: an input port 203 connected to the processor 202, configured to transmit an externally input command to the processor 202; a display unit 204 connected to the processor 202, configured to display the processing result of the processor 202 to the outside; a communication module 205 connected to the processor 202, configured to implement communication between the deep neural network training device and the outside. The display unit 204 may be a display panel, a laser scanning display, etc.; the communication methods adopted by the communication module 205 include but are not limited to Mobile High-Definition Link (HML), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), wireless connections: Wireless Fidelity (WiFi), Bluetooth communication technology, Bluetooth Low Energy communication technology, communication technology based on IEEE802.11s.

[0100] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented:

[0101] Obtain a deep neural network training task sent by a deep learning framework;

[0102] Transmit the deep neural network training task to a target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task.

[0103] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: obtain each deep neural network training subtask sent by the deep learning framework through each subprocess; splice all the deep neural network training subtasks into the deep neural network training task in the main process.

[0104] A computer-readable storage medium provided by an embodiment of the present application stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: based on the main process, transmit the deep neural network training task to the target FPGA through the nGraph framework, so that each FPGA in the target FPGA executes the corresponding deep neural network training subtask.

[0105] A computer-readable storage medium provided by an embodiment of the present application stores a computer program therein. When the computer program is executed by a processor, the following steps are implemented: After transmitting the deep neural network training task to the target FPGA through the nGraph framework based on the main process, receiving, based on the main process, the task execution result transmitted by the target FPGA through the nGraph framework, where the task execution result includes the sub-execution results of each FPGA for the corresponding deep neural network training subtask; splitting the task execution result into the sub-execution results corresponding to each sub-process; and transmitting the sub-execution results to the deep learning framework based on the sub-processes.

[0106] The computer-readable storage medium involved in the present application includes random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium well-known in the technical field.

[0107] For the description of the relevant parts in the deep neural network training system, device, equipment, and computer-readable storage medium provided by the embodiments of the present application, please refer to the corresponding detailed description in the deep neural network training method provided by the embodiments of the present application, which will not be elaborated here. In addition, the parts of the above technical solutions provided by the embodiments of the present application that are consistent with the corresponding technical solutions in the prior art in terms of implementation principles are not described in detail to avoid excessive elaboration.

[0108] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0109] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for training a deep neural network, characterized in that, Applied to a connector, including: Obtain a deep neural network training task sent by a deep learning framework; The connector transmits the deep neural network training task to a target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task; Among them, the connector includes Bridge, and the deep learning framework includes tensorflow; in the design of the support module solution for the nGraph FPGA backend device in the Bridge layer, modify the Bridge compilation file build_ngtf.py, add support for the FPGA backend in it, execute the build_ngtf.py script for compilation, add the compilation condition "--build_fpga_backend" in the compilation command. After successful compilation, complete the support from the tensorflow front-end to the nGraph FPGA backend device, and enable the backend to be used as the FPGA device through the environment variable NGRAPH_TF_BACKEND, and add the support statement import ngraph_bridge for nGraph in the tensorflow client-side training code to accelerate the original tensorflow client training code using the nGraph FPGA backend device in the tensorflow front-end.

2. The method according to claim 1, characterized in that, The obtaining of the deep neural network training task sent by the deep learning framework includes: Obtain each deep neural network training subtask sent by the deep learning framework through each subprocess; Concatenate all the deep neural network training subtasks into the deep neural network training task in the main process.

3. The method according to claim 2, characterized in that, The connector transmits the deep neural network training task to the target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task, including: Based on the main process, transmit the deep neural network training task to the target FPGA through the nGraph framework, so that each FPGA in the target FPGA executes the corresponding deep neural network training subtask.

4. The method according to claim 3, characterized in that, After transmitting the deep neural network training task to the target FPGA through the nGraph framework based on the main process, it further includes: Based on the main process, receive the task execution result transmitted by the target FPGA through the nGraph framework, and the task execution result includes the sub-execution results of each FPGA for the corresponding deep neural network training subtask; Split the task execution result into the sub-execution results corresponding to each subprocess; Transmit the sub-execution result to the deep learning framework based on the subprocess.

5. A system for training a deep neural network, characterized in that, Applied to a connector, including: A first obtaining module, configured to obtain a deep neural network training task sent by a deep learning framework; A first transmission module for transmitting the deep neural network training task to a target FPGA through the nGraph framework, so that the target FPGA executes the deep neural network training task; Wherein, the connector includes Bridge, and the deep learning framework includes tensorflow; In the design of the support module solution for the nGraph FPGA backend device by the Bridge layer, modify the Bridge compilation file build_ngtf.py, add support for the FPGA backend in it, execute the build_ngtf.py script for compilation, add the compilation condition "--build_fpga_backend" in the compilation command. After successful compilation, complete the support from the tensorflow front-end to the nGraph FPGA backend device, and enable the backend to be used as the FPGA device through the environment variable NGRAPH_TF_BACKEND, and add the support statement import ngraph_bridge for nGraph in the tensorflow client-side training code to accelerate the original tensorflow client training code using the nGraph FPGA backend device in the tensorflow front-end.

6. A device for training a deep neural network, characterized in that, Comprising: A deep learning framework for obtaining a deep neural network training task; A connector connected to the deep learning framework for transmitting the deep neural network training task to the nGraph framework; The nGraph framework connected to the connector for transmitting the deep neural network training task to the target FPGA; The target FPGA connected to the nGraph framework for executing the deep neural network training task; Wherein, the connector includes Bridge, and the deep learning framework includes tensorflow; In the design of the support module solution for the nGraph FPGA backend device by the Bridge layer, modify the Bridge compilation file build_ngtf.py, add support for the FPGA backend in it, execute the build_ngtf.py script for compilation, add the compilation condition "--build_fpga_backend" in the compilation command. After successful compilation, complete the support from the tensorflow front-end to the nGraph FPGA backend device, and enable the backend to be used as the FPGA device through the environment variable NGRAPH_TF_BACKEND, and add the support statement import ngraph_bridge for nGraph in the tensorflow client-side training code to accelerate the original tensorflow client training code using the nGraph FPGA backend device in the tensorflow front-end.

7. A device for training a deep neural network, characterized in that, Comprising: A memory for storing a computer program; A processor, configured to implement the steps of the deep neural network training method according to any one of claims 1 to 4 when executing the computer program.

8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the deep neural network training method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Neural network model processing method and device, reasoning method and device and electronic equipment

    CN113435565A

  • Data processing method and related products

    US20210334137A1