Multi-FPGA Cooperative Multi-Deep Neural Network Pipeline Acceleration Method

Through model operator recognition, task division and pipeline operation control, the model splitting and task allocation problems when multi-FPGAs jointly deploy deep neural networks are solved, and efficient collaborative acceleration of multi-model inference is achieved.

CN119808859BActive Publication Date: 2025-07-25NAT UNIV OF DEFENSE TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510002762.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-07-25
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

In the prior art, multiple FPGAs lack appropriate model splitting methods and task allocation strategies when co-deploying deep neural networks, resulting in the need to redeploy the accelerator after replacing the network model, and the high parallelism and pipeline design advantages of multiple FPGAs are not fully utilized.

Method used

Through model operator identification, task division and pipeline operation control, the operators are deployed to the corresponding FPGA computing unit according to the network structure and computing requirements of different deep neural network models, and pipeline operations are performed through the efficient scheduling strategy of the upper computer to achieve coordinated acceleration of multiple deep neural network models.

Benefits of technology

It effectively reduces the overall time of multi-model inference, makes full use of the high parallelism and flexibility of multi-FPGAs, and is suitable for a variety of deep neural network models to achieve the coordinated acceleration of multiple models on multi-FPGAs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808859B_ABST
    Figure CN119808859B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-FPGA collaborative multi-depth neural network pipeline acceleration method, and the method comprises the following steps: obtaining operators required for the inference of each depth neural network model according to the network structure and computing requirements of different depth neural network models to be deployed; dividing the inference process of each depth neural network model to be deployed into multiple task segments; sending corresponding scheduling instructions to multiple FPGA computing units according to the operators and the task segment division results, deploying the operators required for the inference of each depth neural network model to the corresponding FPGA computing units according to categories, and controlling the corresponding FPGA computing units to perform pipeline operations so as to perform the calculation of each task segment of the depth neural network model to be deployed and complete the inference of each depth neural network model to be deployed. The present invention can perform pipeline processing on different computing tasks in the inference processes of multiple models, effectively reducing the overall time of multi-model inference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of neural network accelerators, and particularly to a method for accelerating multiple deep neural network pipelines through multi-FPGA collaboration. Background Art

[0002] Currently, deep neural networks have been widely applied in application fields such as communication signal recognition, radio frequency fingerprint recognition, and radar interference recognition. In order to improve performance, during the application process of deep neural networks, the number of parameters and computational complexity have also increased rapidly. Correspondingly, the computing power requirements for the hardware platform during its deployment are also getting higher and higher.

[0003] In the prior art, GPUs (graphics processing units) are usually used to train and deploy deep neural networks. This technology has problems such as low energy usage efficiency and limited flexibility. Compared with GPUs, FPGAs (field programmable gate arrays) have higher energy efficiency ratios and more superior flexibility. FPGAs can utilize programmable logic units therein, such as LUTs (look-up tables) and DSPs (digital signal processors), for parallel computing, and through customized hardware designs, achieve higher energy efficiency ratios, computing efficiencies, and resource efficiencies. Therefore, the method of using FPGAs for deep neural network deployment has been favored by many researchers.

[0004] However, existing solutions for accelerating deep neural network inference based on FPGA chips mostly focus on single chips. When a single-chip FPGA processes a large-scale deep neural network, its overall computing power will be limited due to its relatively few computing resources and storage resources. Thus, related technologies for using multiple FPGA chips to collaboratively complete deep neural network deployment have emerged.

[0005] Deploying deep neural networks on multiple FPGAs can make full use of the parallel computing power of FPGAs to improve computing efficiency. In the existing literature, the Chinese invention patent "A Method for Collaborative Training of Neural Networks Based on Distributed Optimization Using Multiple FPGAs" (Application No.: CN202310598533) serially partitions a neural network and deploys it on multiple FPGAs connected in a star or ring shape, with the host computer serving as the core scheduling process to accelerate network training. However, multi-FPGA collaborative inference acceleration is more in line with the current situation of limited FPGA chip resources. The Chinese invention patent "A Method for Implementing Acceleration and Deployment of Large-Scale Convolutional Neural Networks Based on Multiple FPGAs" (Application No.: CN202410025778) uses multiple FPGAs in cascade and reduces idle time while balancing the workload of all FPGAs through task allocation to improve computing efficiency, but its operation is limited to supporting a single deep neural network model.

[0006] The deficiencies of existing solutions mainly include two aspects: on the one hand, when using multiple FPGAs for inference acceleration, there is a lack of appropriate model splitting methods and task allocation strategies, resulting in the need to redeploy the accelerator after changing the network model; on the other hand, existing methods are limited to deploying a single network model and do not fully utilize the high parallelism and pipeline design advantages of multiple FPGAs.

[0007] Therefore, the present invention proposes a method and system for multi-FPGA collaborative multi-depth neural network pipeline acceleration, which can execute multiple deep neural network models in parallel, determine task partitioning based on the hierarchical structure and computational volume of different network models, and through an efficient scheduling strategy of the host computer, accurately allocate computational tasks to the computational operators inside each FPGA, and uses a multi-FPGA pipeline design to reduce the data waiting time between operators, improving the computational parallelism and the overall computational efficiency of the system. Summary of the Invention

[0008] In view of this, the present invention provides a method and system for multi-FPGA collaborative multi-depth neural network pipeline acceleration, which is used to solve at least some of the technical problems in the background technology.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] On the one hand, the present invention discloses a method for multi-FPGA collaborative multi-depth neural network pipeline acceleration, including the following steps:

[0011] Model operator identification:

[0012] According to the network structure and computational requirements of different deep neural network models to be deployed, obtain the operators required for the inference of each deep neural network model.

[0013] Model task partitioning:

[0014] According to the network hierarchical structure and computational requirements of each deep neural network model to be deployed, perform task partitioning, and divide the inference process of each deep neural network model to be deployed into multiple task segments.

[0015] Pipeline operation control:

[0016] According to the operators required for the inference of each deep neural network model to be deployed and the task segment partitioning results, send corresponding scheduling instructions to multiple FPGA computing units, deploy the operators required for the inference of each deep neural network model to the corresponding FPGA computing units according to categories, and control the corresponding FPGA computing units to perform pipeline operations to calculate each task segment of the deep neural network model to be deployed and complete the inference of each deep neural network model to be deployed.

[0017] Further, in the above model operator recognition step, obtaining the operators required for the inference of each deep neural network model specifically includes:

[0018] Extract the hierarchical structure of each deep neural network model to be deployed, and identify and obtain various computational operators involved in the model inference process.

[0019] Further, the above computational operators include convolutional operators, pooling operators, and fully connected operators.

[0020] Further, the above model task division step specifically includes:

[0021] For each deep neural network model, perform a top-down task segment division according to the hierarchical structure category of the deep neural network model. During the division process, the hierarchies of adjacent same hierarchical structure categories are divided into the same task segment.

[0022] Further, in the above pipeline operation control step, sending corresponding scheduling instructions to multiple FPGA computing units is obtained through the following steps:

[0023] According to the operators required for the inference of each deep neural network model and the task segment division results, generate the allocation order of the weight parameter segments, the execution order of the task segments, and the data transmission path of each deep neural network model into scheduling instructions for each FPGA.

[0024] Further, in the above pipeline operation control step, controlling the corresponding FPGA computing unit to execute the pipeline operation specifically includes:

[0025] According to the execution order of the task segments of different deep neural network models to be deployed, control different task segments of the corresponding deep neural network models to be deployed to execute alternately in different FPGA computing units until the task segment calculations of all models are completed.

[0026] On the other hand, the present invention also discloses a multi-FPGA collaborative multi-deep neural network pipeline acceleration system, including a plurality of FPGA computing units, a main control unit, and a controller;

[0027] Wherein, each FPGA computing unit is used to calculate different network structures of the deep neural network model and execute the inference tasks of different network structures of the corresponding deep neural network model;

[0028] The main control unit is used to perform data exchange with each FPGA computing unit;

[0029] The controller can send corresponding scheduling instructions to multiple FPGA computing units according to the operators required for the inference of each deep neural network model to be deployed and the task segment division results, deploy the operators required for the inference of each deep neural network model to the corresponding FPGA computing units according to categories, and control the corresponding FPGA computing units to perform pipeline operations to support each FPGA computing unit to perform the calculation of each task segment of the deep neural network model to be deployed, and complete the inference task of each deep neural network model to be deployed.

[0030] Preferably, in the above system, the operators required for the inference of each deep neural network model to be deployed are obtained through the following steps:

[0031] The controller extracts the hierarchical structure of each deep neural network model to be deployed, and identifies and obtains various computing operators involved in the model inference process.

[0032] Preferably, in the above system, the controller sends corresponding scheduling instructions to multiple FPGA computing units through the following steps:

[0033] According to the operators required for the inference of each deep neural network model and the task segment division results, the execution order and data transmission path of each deep neural network model task segment are generated into scheduling instructions for each FPGA.

[0034] Preferably, in the above system, controlling the corresponding FPGA computing unit to perform pipeline operations specifically includes:

[0035] The controller controls the different task segments of the corresponding deep neural network model to be deployed to be alternately executed in different FPGA computing units according to the execution order of the task segments of different deep neural network models to be deployed until the calculation of the task segments of all models is completed.

[0036] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a multi-FPGA collaborative multi-deep neural network pipeline acceleration method and system, which has the following beneficial effects:

[0037] The present invention can achieve the simultaneous collaborative acceleration of multiple deep neural network models on multiple FPGAs. The technical solution of the present invention makes full use of the high parallelism and flexibility of multi-FPGA collaboration, can deploy specific operators on computing units according to the characteristics of the model network structure, is applicable to various deep neural network models, and can perform pipeline processing on different computing tasks in the inference processes of multiple models, effectively reducing the overall time of multi-model inference. Therefore, it can achieve the simultaneous collaborative acceleration of multiple deep neural network models on multiple FPGAs. Description of the Drawings

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0039] Figure 1 Figure 4 is the overall flowchart of the multi-FPGA collaborative multi-depth neural network pipeline acceleration method provided by the embodiments of the present invention.

[0040] Figure 2 Figure 8 is a schematic diagram of the overall system structure for implementing the multi-FPGA collaborative multi-depth neural network pipeline acceleration method provided by the embodiments of the present invention.

[0041] Figure 3 Figure 12 is the flowchart of the inter-board communication connection operation in the method provided by the embodiments of the present invention.

[0042] Figure 4 Figure 16 is the flowchart of the task allocation strategy operation in the method provided by the embodiments of the present invention.

[0043] Figure 5 Figure 20 is a schematic diagram of the task division of the AlexNet model in the method provided by the embodiments of the present invention.

[0044] Figure 6 Figure 24 is a schematic diagram of the task division of the VGG16 model in the method provided by the embodiments of the present invention.

[0045] Figure 7 Figure 28 is the flowchart of the calculation unit setting operation in the method provided by the embodiments of the present invention.

[0046] Figure 8 Figure 32 is the flowchart of the pipeline operation control in the method provided by the embodiments of the present invention.

[0047] Figure 9 Figure 36 is a schematic diagram of the timing of the pipeline operation execution in the method provided by the embodiments of the present invention. Detailed implementation manners

[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0049] An embodiment of the present invention first discloses a multi-FPGA collaborative multi-depth neural network pipeline acceleration method, which mainly includes three steps: model operator recognition, model task division, and pipeline operation control.

[0050] In the specific implementation process, as Figure 1 shown, the acceleration method may include processes such as inter-board communication connection, task allocation strategy, computing unit setting, and flow line operation control. The above method can be implemented based on the system structure as Figure 2 shown. The steps of the present invention in the specific implementation process will be described in detail below with reference to the accompanying drawings. The multi-FPGA collaborative multi-depth neural network pipeline acceleration method includes the following specific implementation steps:

[0051] S101. Inter-board communication connection: In this step, the host computer is used as the main control unit, responsible for the overall control and scheduling tasks. Each FPGA is used as a computing unit to execute specific computing tasks. The FPGAs achieve fast inter-board data transmission through high-speed interfaces; each FPGA is connected to the host computer equipped with a controller through a high-speed bus.

[0052] S102. Task allocation strategy: In this step, the host computer divides tasks according to the network layer structure and computing requirements of each model to be deployed, and divides the inference process of each model into multiple task segments.

[0053] S103. Computing unit setting: In this step, according to the network structure and computing requirements of each network model, the operators required for model inference are identified, and different operators are deployed on each FPGA computing unit; after each network model is divided into task segments, the parameters required for the computing task segments are stored in the off-chip memory of the main control unit.

[0054] S104. Pipeline operation control: In this step, the controller in the host computer sends corresponding scheduling instructions through the high-speed bus according to the settings of the computing units and the task division of each model to be deployed, and controls each computing unit FPGA to perform pipeline operations to calculate the task segments and infer each model.

[0055] In step S101 of the above embodiment, the main control unit refers to the host computer responsible for the overall system control and scheduling tasks, which can send data to other FPGAs and perform data exchange; the computing unit refers to the FPGA equipped with computing operators to calculate different network layer structures; the high-speed interface refers to a communication interface such as PCIE, MiniSAS, QSFP+ that can connect FPGAs to each other to achieve high-speed data transmission; the controller refers to a program in the host computer that can generate scheduling instructions for controlling each computing unit according to the division of computing tasks and generate code or configuration files that can be executed by the FPGA.

[0056] Further, as Figure 3 shown, the inter-board communication connection in step S101 includes the following detailed steps:

[0057] S1011: Hardware connection: Use the host computer as the main control unit, and each FPGA as a computing unit. The FPGAs achieve fast transmission of inter-board data through high-speed interfaces. Each FPGA is connected to the host computer through a high-speed bus.

[0058] For example, the host computer can be an ARM-based processor or CPU. Use the host computer as the main control unit and the FPGA as the computing unit. Connect the FPGA as the computing unit to the host computer as the main control unit through the PCIe high-speed communication bus, and also connect the computing units to each other to achieve inter-board data communication.

[0059] S1012: Data transfer module control: Design a driver and interface application program at the host computer end, and design data transfer modules that adapt to high-speed interfaces at the main control unit end and each computing unit end to achieve data transfer and transceiver between FPGAs.

[0060] For example, use the PCIe hardware module IP core provided by manufacturer Xilinx to achieve data communication between FPGAs.

[0061] S1013: Data scheduling: The controller of the host computer is responsible for scheduling the data transfer of the entire system, reading parameters from off-chip memory, sending the parameters required for each task segment calculation to the computing unit, and controlling the computing unit to transfer the calculation results of each task segment to the next computing unit to ensure the orderly transmission of the data stream.

[0062] In the above step S102, the network layer structure of the model refers to the types and connection methods of each network layer of the model; the computing requirements refer to the specific requirements of different network layer types for computing operations and algorithm implementation; the task segment refers to dividing the network layers of the same type into a segment of computing tasks in sequence according to the model inference process, and dividing the inference process of the model into multiple independent computing tasks.

[0063] Further, as Figure 4 shown, the task allocation strategy in step S102 includes the following detailed steps:

[0064] S1021: Model structure analysis: The host computer extracts the hierarchical structure of each neural network model and analyzes various computing operators involved in the model inference process.

[0065] For example, the hierarchical structure of a convolutional neural network model includes a convolutional layer, a pooling layer, and a fully connected layer, and the operators involved include convolutional operators, pooling operators, and fully connected operators.

[0066] S1022: Task Division Strategy: The host computer determines whether task division is required based on the network hierarchical structure of each model to be deployed. During the process of layer-by-layer inference of the model, if the hierarchical structure remains unchanged, it is determined that no division is required, and the inference continues to the next layer. If the hierarchical structure changes, it is determined that the division operation is to be performed, and each divided task segment is marked with the model to which it belongs, until the compiler divides each model to be deployed into multiple task segments.

[0067] For example, if the first network layer is a convolutional layer, the second network layer is a pooling layer, and the third network layer is again a convolutional layer, then the first layer is divided into one task segment, the second layer is divided into one task segment, and the third layer is divided into one task segment. To describe more vividly, the task division strategy disclosed in the present invention Figure 5 gives a schematic diagram of the task division result of the AlexNet network model, Figure 6 and gives a schematic diagram of the task division result of the VGG16 network model.

[0068] S1023: Weight Parameter Division: The weight parameters corresponding to the task segments obtained by dividing each model to be deployed are also correspondingly divided into parameter segments, and each divided parameter segment is marked with the model to which it belongs.

[0069] In the above step S103, the operator refers to the basic computing unit used to calculate different network layers in the network model; the off-chip memory refers to devices such as SDRAM and DDR3 that are used outside the chip to expand the storage capacity of the system.

[0070] Furthermore, as Figure 7 shown, the computing unit setting in step S103 includes the following detailed steps:

[0071] S1031: Operator Hardware Deployment: Determine the computing operators involved in all models to be deployed and classify them according to their functional characteristics. Deploy each type of functional operator on a specific computing unit, and each computing unit undertakes the computing tasks of different operators.

[0072] For example, a convolutional operator is deployed on computing unit 1 for convolutional layer calculation, a pooling operator is deployed on computing unit 2 for pooling layer calculation, a fully connected operator is deployed on computing unit 3 for fully connected layer calculation, etc.

[0073] S1032: Operator Resource Allocation: Evaluate the hardware resources of the FPGA. Based on the characteristics of the operator and the size of the computing load it undertakes, map low-load operators to the computing units composed of a single FPGA, and map high-load operators to the computing units composed of multiple FPGAs.

[0074] For example, the pooling operator is deployed on a computing unit composed of a single FPGA, and the convolution operator is deployed on a computing unit composed of two or more FPGAs to improve the computing speed of convolution operations.

[0075] S1033: Weight parameter configuration: Configure off-chip memory for the main control unit, mark and initialize the weight parameters of all models to be deployed, and store them on the off-chip memory according to the result of task division.

[0076] In the above step S104, the high-speed bus refers to a bus such as PCIe for communication between the host computer and each FPGA; the scheduling instruction refers to an instruction for controlling the collaborative computing of computing units to perform model inference; the pipelining operation refers to an operation in which the computing unit completes a specific computing task in each stage and passes the result to the computing unit in the next stage to achieve parallel computing in a pipeline form.

[0077] Further, as Figure 8 shown, the pipelining operation control in step S104 includes the following detailed steps:

[0078] S1041: Generate scheduling instructions: The controller in the host computer generates scheduling instructions for each FPGA based on the configuration of the computing unit and the result of task division, including the allocation order of each model weight parameter segment, the execution order of task segments, and the data transmission path.

[0079] S1042: Pipelining process: The host computer distributes weight parameter segments to the corresponding computing units through scheduling instructions, controls the execution process and data flow direction of each computing unit, and monitors the execution progress of task segments of each computing unit. The entire inference process is carried out in a pipeline form, and the task segment calculations of each model are executed alternately until the task segment calculations of all models are completed.

[0080] For example, in a certain stage of the computing inference process, the convolution computing unit is responsible for executing task segment x of model 1. When entering the next stage, the calculation result of task segment x of model 1 is passed to the pooling computing unit for the next calculation. At the same time, task segment y of model 2 is sent to the convolution computing unit for execution. When entering the next stage, the result of task segment y of model 2 is passed to the pooling computing unit, and task segment x + 1 of model 1 is sent to the convolution computing unit for execution, and so on alternately until the task segment calculations of all models are completed. The specific timing diagram of the system pipelining operation can be referred to Figure 9 shown.

[0081] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the apparatuses disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple. For related parts, reference can be made to the descriptions in the method part.

[0082] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A multi-FPGA collaborative multi-depth neural network pipeline acceleration method, characterized in that Including the following steps: Model operator identification: According to the network structure and computational requirements of different deep neural network models to be deployed, obtain the operators required for the inference of each deep neural network model; specifically including: Extract the hierarchical structure of each deep neural network model to be deployed, identify and obtain various computational operators involved in the model inference process, and the computational operators include convolution operators, pooling operators, and fully connected operators; Model task division: According to the network hierarchical structure and computational requirements of each deep neural network model to be deployed, perform task division, and divide the inference process of each deep neural network model to be deployed into multiple task segments; specifically including: For each deep neural network model, perform top-down task segment division according to the hierarchical structure category of the deep neural network model. During the division process, adjacent hierarchical structures of the same category are divided into the same task segment; Pipeline operation control: According to the operators required for the inference of each deep neural network model to be deployed and the task segment division results, send corresponding scheduling instructions to multiple FPGA computing units, deploy the operators required for the inference of each deep neural network model to the corresponding FPGA computing units according to the category, and control the corresponding FPGA computing units to perform pipeline operations to calculate each task segment of the deep neural network model to be deployed and complete the inference of each deep neural network model to be deployed; Controlling the corresponding FPGA computing unit to perform pipeline operations specifically includes: According to the execution order of the task segments of different deep neural network models to be deployed, control the different task segments of the corresponding deep neural network models to be deployed to execute alternately in different FPGA computing units until the task segment calculations of all models are completed.

2. The multi-FPGA collaborative multi-depth neural network pipeline acceleration method according to claim 1, wherein In the pipeline operation control step, the corresponding scheduling instructions are sent to multiple FPGA computing units and obtained through the following steps: According to the operators required for the inference of each deep neural network model and the task segment division results, generate the weight parameter segment allocation order, task segment execution order, and data transmission path of each deep neural network model into scheduling instructions for each FPGA.

3. An acceleration system for the multi-FPGA collaborative multi-depth neural network pipeline acceleration method according to claim 1, characterized in that, Including several FPGA computing units, a main control unit, and a controller; Among them, each FPGA computing unit is used to calculate different network structures of the deep neural network model and execute the inference tasks of different network structures of the corresponding deep neural network model; The main control unit is used for data exchange with each FPGA computing unit; The controller can, according to the operators required for the inference of each deep neural network model to be deployed and the task segment division results, send corresponding scheduling instructions to multiple FPGA computing units, deploy the operators required for the inference of each deep neural network model to the corresponding FPGA computing units according to the category, and control the corresponding FPGA computing units to perform pipeline operations to support each FPGA computing unit to calculate each task segment of the deep neural network model to be deployed and complete the inference task of each deep neural network model to be deployed.

4. An acceleration system according to claim 3, characterized in that, The operators required for the inference of each deep neural network model to be deployed are obtained through the following steps: The controller extracts the hierarchical structure of each deep neural network model to be deployed, and identifies and obtains various computing operators involved in the model inference process.

5. An acceleration system according to claim 3, characterized in that, In the controller, corresponding scheduling instructions are sent to multiple FPGA computing units, and obtained through the following steps: According to the operators required for the inference of each deep neural network model and the task segment division result, the execution order and data transmission path of each task segment of each deep neural network model are generated into scheduling instructions for each FPGA.

6. An acceleration system according to claim 3, wherein Control the corresponding FPGA computing unit to perform pipeline operations, specifically including: The controller controls different task segments of the corresponding deep neural network model to be deployed to execute alternately in different FPGA computing units according to the execution order of different task segments of the deep neural network model to be deployed, until the task segment calculations of all models are completed.

Citation Information

Patent Citations

  • Multi-FPGA cooperative training neural network method based on distributed optimization

    CN116842998A

  • Implementation method of large-scale convolutional neural network acceleration and deployment device based on multiple FPGAs (Field Programmable Gate Array)

    CN117709406A

  • Design method for deploying and optimizing operator library on FPGA and DSP

    CN113778459A

  • Model training method and apparatus based on hybrid parallelism mode, and device

    WO2024169906A1