A Heterogeneous Training Accelerator Based on Reverse Pipeline

By using a heterogeneous training accelerator based on reverse pipelines on the mobile terminal, the BP and WG processes in the neural network training process are processed in parallel, and the delay and power consumption problems of training acceleration under the limitation of mobile terminal resources are solved, achieving efficient and low-power local training.

CN114742216BActive Publication Date: 2025-06-10NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210412651.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-06-10
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

The prior art is difficult to achieve low-power local training acceleration in resource-constrained mobile terminals, especially under the demand for continuous learning in dynamic environments, the delay and privacy of cloud training and mobile terminal inference are insufficient.

Method used

A heterogeneous training accelerator based on reverse pipelines is adopted to process the BP and WG processes in neural network training process in parallel to reduce system delay, avoid additional data transmission between different levels of storage, and optimize different types of calculations through heterogeneous architecture design to improve energy consumption ratio and acceleration ratio.

Benefits of technology

It realizes low-power local training acceleration on the mobile terminal, reduces system delay and power consumption, and improves training efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114742216B_ABST
    Figure CN114742216B_ABST
Patent Text Reader

Abstract

The present application discloses a heterogeneous training accelerator based on a reverse pipeline. The accelerator includes a data module, a control module, and a computing module. The control module controls the forward and backward calculation kernels in the BP stage to perform convolution operations on the first channel, and controls the normalization pooling calculation kernel to process the convolution operation results, obtaining the first output error of the first channel of the current convolution layer, and transmitting the first output error to the WG calculation kernel. At the same time, the control module controls the WG calculation kernel to perform the convolution operation corresponding to the WG stage of the deep neural network according to the first output error and the input value of the first channel of the previous convolution layer. The pipeline parallel processing of the BP and WG processes of this accelerator reduces system latency, avoids additional transmission of data stored at different levels, reduces power consumption. At the same time, its heterogeneous architecture design can achieve separate optimization for different types of calculations, improving the energy consumption ratio and acceleration ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of convolutional computing of deep neural networks, and in particular to a heterogeneous training accelerator based on reverse pipeline. Background Art

[0002] In areas such as learnable autonomous driving systems and smart homes, mobile devices that deploy artificial intelligence systems are often required to have the ability to continuously learn in a dynamic environment. The previous cloud training and mobile inference methods are difficult to meet this requirement due to their high latency and low privacy. Therefore, it is necessary to achieve low-power local training acceleration on resource-constrained mobile terminals.

[0003] The article "A highly parallel FPGA implementation of sparse neural network training" published at the 2018 International Conference on Reconfigurable Computing and FPGA proposed a pipeline accelerator, which realizes parallel processing of different stages of sparse network training through fine-grained design of sparsity rate and network structure. However, the solution adopted by it is only applicable to simple and compact networks, such as three-layer MLP networks, because it requires fine-grained design, and is not applicable to commonly used large networks. The article "A Pipelined Direct Feedback Alignment-Based Deep Neural Network Learning Processor for Fast Online Learning" in IEEE Solid-State Circuits Magazine proposed a pipeline training accelerator, which relies on the DFA (direct feedback alignment) algorithm and removes the layer-by-layer restriction in the reverse direction, thereby realizing parallel processing of BP and WG. However, the solution adopted by it relies on a special DFA algorithm, which can complete the training of all layers at the cost of losing accuracy. Only when the algorithm only participates in the training of the last few layers can the accuracy be guaranteed, but it is predicted that the training efficiency will also be lost. Summary of the invention

[0004] This application discloses a heterogeneous training accelerator based on reverse pipeline, which adopts a new pipeline parallel processing method for BP and WG processes in the neural network training process, which reduces system latency, avoids additional data transmission between different levels of storage, and reduces system power consumption. The heterogeneous architecture design supporting reverse pipeline can achieve separate optimization of different types of calculations, thereby improving energy efficiency and acceleration ratio.

[0005] This application discloses a heterogeneous training accelerator based on a reverse pipeline. The accelerator includes a data module, a control module, and a computing module.

[0006] The computing module includes a forward-backward computing core, a WG computing core, and a normalization pooling computing core.

[0007] The control module is configured to perform the following steps:

[0008] Control the forward-backward computing core and the normalization pooling computing core to perform the operations corresponding to the FF stage of the deep neural network;

[0009] In the BP stage of the deep neural network, control the forward-backward computing core to obtain the input data corresponding to the first channel of the current convolutional layer from the data module, perform the corresponding convolutional operation, and control the normalization pooling computing core to process the result of the convolutional operation to obtain the first output error, and transmit the first output error to the WG computing core. The current convolutional layer is any network layer in the deep neural network, and the first channel is the computing channel corresponding to any convolutional kernel in the current convolutional layer;

[0010] Control the WG computing core to obtain the input value to be convolved with the first output error. The previous convolutional layer is the previous network layer of the current convolutional layer;

[0011] Control the WG computing core to perform the convolutional operation corresponding to the WG stage of the deep neural network according to the first output error and the input value;

[0012] When all the first channels of the current convolutional layer in the BP stage are calculated, the first output errors of all the first channels of the current convolutional layer form the output error of the current convolutional layer, and the output error is transmitted to the WG computing core of the previous convolutional layer.

[0013] In an implementable manner, the forward-backward computing core includes a first PE array and an adder tree, and the first PE array includes multiple first PEs.

[0014] The first PE is used to perform the convolutional operations in the FF stage and the BP stage. The convolutional operations in any one of the FF stage and the BP stage are evenly distributed among the first PEs.

[0015] The adder tree is used to perform an addition operation on the calculation results of the first PEs to obtain the convolutional calculation result.

[0016] The WG computing core includes a second PE array, and the second PE array includes multiple second PEs.

[0017] The second PE is used to perform the convolutional operation in the WG stage. The convolutional operation in the WG stage is evenly distributed among the second PEs.

[0018] In one implementable manner, the number of first PEs and the number of second PEs are determined as follows:

[0019] Obtain the utilization rate of the first PE in the BP stage and the utilization rate of the second PE in the WG stage;

[0020] According to the utilization rate of the first PE, the utilization rate of the second PE, the preset weight sparsity rate, the preset error sparsity rate, the number of channels of each convolutional layer in the BP stage, and the number of channels of each convolutional layer in the WG stage, determine the ratio of the parallelism of the first PE to the parallelism of the second PE;

[0021] According to the ratio of the parallelism of the first PE to the parallelism of the second PE, determine the number of first PEs and the number of second PEs.

[0022] In one implementable manner, obtaining the utilization rate of the first PE in the BP stage and the utilization rate of the second PE in the WG stage includes:

[0023] Perform hardware simulation on the heterogeneous training accelerator to obtain the utilization rate of the first PE in the BP stage and the utilization rate of the second PE in the WG stage.

[0024] In one implementable manner, according to the utilization rate of the first PE, the utilization rate of the second PE, the preset weight sparsity rate, the preset error sparsity rate, the amount of computation in each convolutional channel in the BP stage, and the amount of computation in each convolutional channel in the WG stage, determining the ratio of the parallelism of the first PE to the parallelism of the second PE includes:

[0025] Determine the ratio of the parallelism of the first PE to the parallelism of the second PE through the following formula:

[0026]

[0027] where p BP represents the parallelism of the first PE, p WG represents the parallelism of the second PE, u BP represents the utilization rate of the first PE, u WG represents the utilization rate of the second PE, sp e represents the preset error sparsity rate, M BP represents the amount of computation in each convolutional channel in the BP stage, M WG represents the amount of computation in each convolutional channel in the WG stage.

[0028] In one implementable manner, the data module includes a bus matrix module and a storage module;

[0029] The storage module is used to store the input data and output data of the current convolutional layer in each stage;

[0030] The bus matrix module is used to allocate and retrieve input data in the storage module.

[0031] In an implementable manner, the control module includes a configuration register and a general control unit;

[0032] The configuration register is used to initialize the configuration of data and parameters;

[0033] The general control unit is used to control the data module and the calculation module to perform corresponding operations.

[0034] This application discloses a heterogeneous training accelerator based on a reverse pipeline. The accelerator includes a data module, a control module, and a calculation module. The control module controls the forward and backward calculation kernels in the BP stage to perform convolution operations on the first channel, where the first channel is the calculation channel corresponding to any convolution kernel in the current convolution layer, and controls the normalization pooling calculation core to process the convolution operation result to obtain the first output error of the first channel of the current convolution layer, and transmits the first output error to the WG calculation kernel. At the same time, it controls the WG calculation kernel to perform the convolution operation corresponding to the WG stage of the deep neural network according to the first output error and the input value of the first channel of the previous convolution layer. The pipeline parallel processing of the BP and WG processes of this accelerator reduces system latency, avoids additional transmission of data stored at different levels, reduces power consumption. At the same time, its heterogeneous architecture design can achieve separate optimization of different types of calculations, improving the energy consumption ratio and acceleration ratio. Description of the Drawings

[0035] In order to more clearly illustrate the technical solutions of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0036] Figure 1 It is a schematic diagram of the three-stage calculation of the embodiment of this application;

[0037] Figure 2 It is a schematic diagram of the top-level architecture of the embodiment of this application;

[0038] Figure 3 It is a schematic diagram of the BPIP design of the embodiment of this application;

[0039] Figure 4 It is a schematic diagram of the parallelism of the BP and WG stages of the embodiment of this application. Detailed Embodiments

[0040] The embodiments will be described in detail below, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same letters in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application as detailed in the claims.

[0041] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the embodiments described next, and is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.

[0042] To facilitate the description of the technical solutions of the application, some concepts related to the present application will be described first below.

[0043] The training of a convolutional neural network generally involves three stages of calculations: FF (feed-forward, forward propagation stage), BP (backward propagation, backward propagation stage), and WG (weight gradient generation, weight update stage). In the FF process, forward propagation is performed on a batch of data to obtain their losses. In the BP process, the errors of each intermediate feature map are obtained by backpropagating the loss. In the WG process, the gradient value and update value of the weights are obtained using the errors of the intermediate features, and the weights are updated.

[0044] The specific calculations involved in the above three stages are as follows:

[0045] In the FF stage, the intermediate feature map a of the previous convolutional layer l-1 is convolved by the weights W of the current convolutional layer l and, after adding the bias b l , passes through the non-linear function σ to generate the intermediate feature map of the current convolutional layer. Let z l represent the convolution result, * represents the convolution operation, and the calculation of FF can be expressed as follows:

[0046] a l = σ(z l ) = σ(W l * a l-1 + b l ) Equation (1)

[0047] The operation of BP is similar to that of FF. The error δ of the next convolutional layer l+1 is convolved with the transposed weights W of this layer l+1 to obtain the error δ of the current convolutional layer l , σ'(z l) represents the derivative of a non - linear function, and this calculation is expressed by the formula as follows:

[0048] δ l = rot180(W l+1 ) * δ l+1 ·σ′(z l ) Formula (2)

[0049] WG convolves the current convolutional layer error δ l obtained by BP with the intermediate feature map a l-1 of the previous convolutional layer, to obtain the gradient of the current convolutional layer weight W l , and updates the weight with the gradient. Using α to represent the learning rate, WG is expressed by the formula as follows:

[0050]

[0051] As Figure 1 shows the schematic diagram of the three - stage calculation of the embodiment of the present application.

[0052] Based on the hardware implementation of an accelerator based on the above principle, the embodiment of the present application provides a heterogeneous training accelerator based on a reverse pipeline. As Figure 2 shown in the top - level architecture schematic diagram of the embodiment of the present application, this accelerator includes a data module, a control module, and a computing module.

[0053] The computing module includes a forward - backward computing core, a WG computing core, and a normalization pooling computing core.

[0054] Specifically, the forward - backward computing core includes a first PE array and an adder tree. The first PE array includes multiple first PEs. The first PE is used to perform the convolution operations in the FF stage and the BP stage. It should be noted that the convolution operations in either the FF stage or the BP stage are evenly distributed among the respective first PEs (processing unit, computing unit). The adder tree is used to perform an addition operation on the calculation results of each first PE to obtain the convolution calculation result.

[0055] In addition, the WG computing core includes a second PE array. The second PE array includes multiple second PEs. The second PE is used to perform the convolution operation in the WG stage. It should be noted that the convolution operation in the WG stage is evenly distributed among the respective second PEs.

[0056] The control module is configured to perform the following steps:

[0057] Step 1, control the forward - backward computing core and the normalization pooling computing core to perform the operations corresponding to the FF stage of the deep neural network.

[0058] In step one, the forward and backward calculation core is used to perform the calculation processes of the FF and BP stages. Among them, during the entire convolution calculation process, the forward and backward calculation core first performs the calculation process of the FF stage. At this time, the control module controls the WG calculation core to be in a low-power idle state. Optionally, clock gating technology is adopted.

[0059] Step two, in the BP stage of the deep neural network, control the forward and backward calculation core to obtain the input data corresponding to the first channel of the current convolution layer from the data module, perform the corresponding convolution operation, and control the normalization pooling calculation core to process the result of the convolution operation to obtain the first output error, and transmit the first output error to the WG calculation core. The current convolution layer is any network layer in the deep neural network, and the first channel is the calculation channel corresponding to any convolution kernel in the current convolution layer.

[0060] In step two, after the forward and backward calculation finishes the calculation process of the FF stage, it starts to perform the calculation process of the BP stage. The control module controls the forward and backward calculation core to obtain the input data corresponding to the first channel of the current convolution layer from the data module, perform the corresponding convolution operation, and control the normalization pooling calculation core to process the result of the convolution operation to obtain the first output error.

[0061] Specifically, since the BP stage performs the backpropagation process, when l represents the current convolution layer and l + 1 represents the next convolution layer of the current convolution layer, the BP stage performs the convolution of the transposed weights of the l + 1 layer and the error of the l + 1 layer, and the error of the l layer is obtained through the normalization pooling calculation core for processing. Refer to Figure 1 the three-stage calculation schematic diagram of the embodiment of the present application shown. It should be noted that the current convolution layer represents any convolution layer in the current neural network.

[0062] Furthermore, each convolution layer includes the convolution operations of multiple convolution kernels. The convolution operation of each convolution kernel is represented by one channel. In the network structure applied in the embodiment of the present application, all convolution layers of the deep neural network have the same number of convolution kernels. The number of convolution kernels in each convolution layer is represented by the letter C, that is, each convolution layer has C channels, and each convolution layer sequentially performs the convolution operations of C channels. It should be noted that the first channel in step two represents the calculation channel of any convolution kernel in the current convolution layer. Refer to Figure 3Schematic diagram of the BPIP (backward pipeline) design of the embodiments of the present application, showing the channel division of the convolutional layer in the embodiments of the present application and the representation of serial operations of each channel on the time line. At the same time, each channel performs corresponding operations to obtain a first output error, which is a part of the output error of the current convolutional layer. The first output errors of all channels of the current convolutional layer constitute the output error of the current convolutional layer.

[0063] In addition, the first output error obtained by the operation of one channel in the BP stage will be immediately transmitted to the corresponding channel in the WG stage and operate simultaneously with the next channel of the current convolutional layer in the BP stage.

[0064] Step 3: Control the WG computing core to obtain the input value to be convolved with the first output error, where the previous convolutional layer is the previous network layer of the current convolutional layer.

[0065] In Step 3, use the first output error as the convolution kernel in the WG stage and perform the corresponding convolution operation in the WG stage. At this time, the first channel in the WG stage is the computing channel corresponding to the first output error, and the control module controls the WG computing core to obtain the input value to be convolved with the first output error. Refer to Figure 1 The schematic diagram of the three-stage calculation of the embodiments of the present application. This input value is a part of the intermediate feature value of the previous convolutional layer during the forward propagation process in the FF stage.

[0066] Step 4: Control the WG computing core to perform the convolution operation corresponding to the WG stage of the deep neural network according to the first output error and the input value.

[0067] In Step 4, perform the convolution calculation process of the first output error and the input value. Specifically, refer to Figure 4 As shown in the parallel schematic diagram of the BP and WG stages of the embodiments of the present application, when the channel C of the l + 1 layer in the BP stage inputs the convolved error into the channel C of the l layer in the WG stage, the channel C + 1 of the l + 1 layer in the BP stage and the channel C in the WG stage run in parallel. In order to enable the computing core to immediately execute the calculation of the corresponding channel after each channel in the BP stage is calculated, so that the system has a smaller delay, it is necessary to configure the hardware in the forward and backward calculations and the WG computing core, that is, to configure the number of the first PE and the second PE. The number of the first PE and the number of the second PE are determined by the following method:

[0068] Obtain the utilization rate of the first PE in the BP stage and the utilization rate of the second PE in the WG stage.

[0069] Determine the ratio of the parallelism of the first PE to the parallelism of the second PE according to the first PE utilization rate, the second PE utilization rate, a preset weight sparsity rate, a preset error sparsity rate, the number of channels in each convolutional layer in the BP stage, and the number of channels in each convolutional layer in the WG stage. Specifically, the first PE utilization rate in the BP stage and the second PE utilization rate in the WG stage can be obtained by performing hardware simulation on the heterogeneous training accelerator.

[0070] Determine the number of the first PEs and the number of the second PEs according to the ratio of the parallelism of the first PE to the parallelism of the second PE. It should be noted that the first PE parallelism represents the number of the first PEs participating in the operation, and the second PE parallelism represents the number of the second PEs participating in the operation.

[0071] Further, the specific process of obtaining the number of the first PEs and the number of the second PEs will be explained below.

[0072] Specifically, referring to Figure 4 the parallel schematic diagrams of the BP and WG stages in the embodiments of the present application, to enable the convolutional layers in the BP stage and the WG stage to be executed in parallel, it is necessary to ensure that the channels in the BP stage and the channels in the WG stage have the same delay. Therefore, it is necessary to ensure that each channel in the BP stage and the WG stage has the same delay. The calculation method of this delay is as follows:

[0073] The delay is approximately equal to the amount of computation divided by the parallelism. In the case of considering the sparsity of the error data input to the BP stage, the computational delays of each channel in the BP stage and the WG stage are respectively expressed as follows:

[0074]

[0075]

[0076] Among them, L BP represents the computational delay of each channel in the BP stage, L WG represents the computational delay of each channel in the WG stage, p BP represents the parallelism of the first PE, p WG represents the second PE parallelism, u BP represents the first PE utilization rate, u WG represents the utilization rate of the second PE, sp e represents the error sparsity rate for the WG stage, and E, F, K, C respectively represent the width and height of the output error in the BP stage and the convolutional kernel size, M BP represents the amount of computation in each convolutional channel in the BP stage, M WG represents the amount of computation in each convolutional channel in the WG stage. It should be noted that in the case of considering other sparsities, the above formulas can be appropriately adjusted based on the original form.

[0077] Specifically, E×F×K×K represents the number of multiply-accumulate operations during the convolution of a one-dimensional convolution kernel with one-dimensional input data. In specific practice, the two such data in the BP stage and the WG stage are equal. Therefore, they are not distinguished in the embodiments of this application. The convolution calculation dimension of each channel is the dimension of the input data multiplied by the dimension of the convolution kernel.

[0078] Let L BP = L WG , and we can obtain:

[0079]

[0080] The above formula is the parallelism ratio of the first PE and the second PE. Through the above parallelism ratio and in combination with the hardware device and actual needs, the numbers of the first PE and the second PE can be set.

[0081] It should be noted that the pooling process in the BP stage will affect the size of the output error, that is, the sizes of E and F. However, the corresponding error sparsity rate will also be affected accordingly. In the specific experimental process, it is found that the growth of the two is positively correlated and can offset each other. Therefore, the impact of the pooling process on the operation can be ignored.

[0082] Step five, when all the first-channel calculations of the current convolutional layer in the BP stage are completed, the first output errors of all the first channels of the current convolutional layer form the output error of the current convolutional layer, and the output error is passed into the WG calculation kernel of the previous convolutional layer.

[0083] In step five, after each convolutional layer calculation in the BP stage is completed, the total output error is obtained. The calculated output error will continue the convolutional calculation process of the previous layer. It should be noted that the relevant calculations in the normalization pooling calculation kernel have a small delay and are ignored in the embodiments of this application.

[0084] In addition, the control module includes a configuration register and a total control unit.

[0085] Among them, the configuration register is used to initialize the configuration of data and parameters.

[0086] The total control unit is used to control the data module and the calculation module to perform corresponding operations.

[0087] In addition, the data module includes a bus matrix module and a storage module.

[0088] Among them, the storage module is used to store the input data and output data of the current convolutional layer in each stage.

[0089] The bus matrix module is used to allocate and retrieve the input data in the storage module.

[0090] The present application discloses a heterogeneous training accelerator based on a reverse pipeline. The accelerator includes a data module, a control module, and a computing module. The control module controls the forward and backward computing kernels in the BP stage to perform a convolution operation on the first channel, where the first channel is the computing channel corresponding to any convolution kernel in the current convolution layer, and controls the normalization pooling computing core to process the convolution operation result to obtain the first output error of the first channel of the current convolution layer, and transmits the first output error to the WG computing kernel. At the same time, the control module controls the WG computing kernel to perform the convolution operation corresponding to the WG stage of the deep neural network according to the first output error and the input value of the first channel of the previous convolution layer. The accelerator performs pipeline parallel processing in the BP and WG processes, reducing system latency, avoiding additional transmission of data stored at different levels, reducing power consumption. At the same time, its heterogeneous architecture design can achieve separate optimization for different types of computations, improving the energy consumption ratio and acceleration ratio.

[0091] The present application has been described in detail above in combination with specific embodiments and exemplary examples, but these descriptions should not be construed as limiting the present application. Those skilled in the art understand that without departing from the spirit and scope of the present application, various equivalent substitutions, modifications, or improvements can be made to the technical solutions and their implementation manners of the present application, and these all fall within the scope of the present application. The protection scope of the present application is subject to the appended claims.

Claims

1. A heterogeneous training accelerator based on a reverse pipeline, characterized in that, the accelerator includes a data module, a control module, and a computing module; the computing module includes a forward-backward computing core, a WG computing core, and a normalization pooling computing core; the control module is configured to perform the following steps: control the forward-backward computing core and the normalization pooling computing core to perform operations corresponding to the FF stage of the deep neural network; in the BP stage of the deep neural network, control the forward-backward computing core to obtain the input data corresponding to the first channel of the current convolutional layer from the data module, perform the corresponding convolutional operation, and control the normalization pooling computing core to process the result of the convolutional operation to obtain a first output error, and transmit the first output error to the WG computing core, the current convolutional layer is any network layer in the deep neural network, and the first channel is the computing channel corresponding to any convolutional kernel in the current convolutional layer; control the WG computing core to obtain the input value to be convolved with the first output error, and the previous convolutional layer is the previous network layer of the current convolutional layer; control the WG computing core to perform the convolutional operation corresponding to the WG stage of the deep neural network according to the first output error and the input value; when all the first channels of the current convolutional layer in the BP stage are calculated, the first output errors of all the first channels of the current convolutional layer form the output error of the current convolutional layer, and the output error is transmitted to the WG computing core of the previous convolutional layer.

2. The heterogeneous training accelerator according to claim 1, characterized in that, the forward-backward computing core includes a first PE array and an adder tree, and the first PE array includes a plurality of first PEs; the first PE is used to perform the convolutional operation in the FF stage and the BP stage, and the convolutional operation in any one of the FF stage and the BP stage is evenly distributed among the first PEs; the adder tree is used to perform an addition operation on the calculation results of the first PEs to obtain a convolutional calculation result; the WG computing core includes a second PE array, and the second PE array includes a plurality of second PEs; the second PE is used to perform the convolutional operation in the WG stage, and the convolutional operation in the WG stage is evenly distributed among the second PEs.

3. The heterogeneous training accelerator according to claim 2, characterized in that, the number of the first PEs and the number of the second PEs are determined by the following method: obtain the utilization rate of the first PEs in the BP stage and the utilization rate of the second PEs in the WG stage; determine the ratio of the parallelism of the first PEs to the parallelism of the second PEs according to the utilization rate of the first PEs, the utilization rate of the second PEs, a preset weight sparsity rate, a preset error sparsity rate, the number of channels of each convolutional layer in the BP stage, and the number of channels of each convolutional layer in the WG stage; determine the number of the first PEs and the number of the second PEs according to the ratio of the parallelism of the first PEs to the parallelism of the second PEs.

4. The heterogeneous training accelerator according to claim 3, characterized in that, the obtaining the utilization rate of the first PEs in the BP stage and the utilization rate of the second PEs in the WG stage includes: Perform hardware simulation on the heterogeneous training accelerator to obtain the first PE utilization rate in the BP stage and the second PE utilization rate in the WG stage.

5. The heterogeneous training accelerator according to claim 3, wherein, determining the ratio of the parallelism of the first PE to the parallelism of the second PE according to the first PE utilization rate, the second PE utilization rate, a preset weight sparsity rate, a preset error sparsity rate, the amount of computation in each convolutional channel in the BP stage, and the amount of computation in each convolutional channel in the WG stage, includes: determining the ratio of the parallelism of the first PE to the parallelism of the second PE through the following formula: Among them, p BP represents the parallelism of the first PE, p WG represents the parallelism of the second PE, u BP represents the utilization rate of the first PE, u WG represents the utilization rate of the second PE, sp e represents the preset error sparsity rate, M BP represents the amount of computation in each convolutional channel during the BP stage, M WG represents the amount of computation in each convolutional channel during the WG stage.

6. The heterogeneous training accelerator according to claim 1, wherein, the data module includes a bus matrix module and a storage module; the storage module is used to store the input data and output data of the current convolutional layer in each stage; the bus matrix module is used to allocate and retrieve the input data in the storage module.

7. The heterogeneous training accelerator according to claim 1, wherein, the control module includes a configuration register and a total control unit; the configuration register is used to perform initialization configuration on data and parameters; the total control unit is used to control the data module and the computing module to perform corresponding operations.

Citation Information

Patent Citations

  • Load-balanced sparse convolutional neural network accelerator and acceleration method thereof

    CN109993297A

  • FPGA-based convolutional neural network on-chip training accelerator

    CN113298237A