Deep neural network training acceleration method and accelerator
By converting non-standard convolutions in deep neural networks into standard convolutions, and weight updates are performed in Winograd domain, combined with interleaved gradient scheduling, the problem of device-side computing resources and memory bandwidth limitations is solved, and training efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202510352268.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-04
AI Technical Summary
The training of deep neural networks on the device side is limited by computing resources and memory bandwidth, and existing methods such as Winograd algorithms have not completely solved this problem.
The deep neural network training acceleration method based on Winograd algorithm is adopted. By converting non-standard convolution into standard convolution, convolution kernels greater than 3×3 are decomposed into multiple subconvolution kernels, and weight updates are performed in the Winograd domain, combining interleaved gradient scheduling strategy to optimize computing and memory access.
It improves the training efficiency and accuracy of deep neural network models, reduces the number of memory accesses, and improves the utilization rate of computing resources.
Smart Images

Figure CN120258064A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method and an accelerator for accelerating the training of deep neural networks. Background Art
[0002] Artificial Intelligence (AI) technology has achieved remarkable achievements in display applications such as face recognition, object detection, and autonomous driving, and has been widely used. The core of these applications lies in the training of Deep Neural Networks (DNN), and the training process of DNN is a computationally intensive and memory-intensive task. Traditional AI applications based on deep neural networks usually adopt the strategy of offline training and online inference, that is, first complete the training of the model on a high-performance server, and then deploy the trained model to the device side for inference. However, with the increasing demand for the DNN model to be able to continuously learn new data after deployment, as well as considerations in aspects such as privacy protection, data security, and communication overhead, directly performing DNN training on the device side has become increasingly important.
[0003] Implementing DNN training on the device side faces a series of challenges, among which the most prominent are the limitations of computing resources and memory bandwidth. The device side usually has limited computing power and memory resources, which limits the training efficiency and scale of the DNN model. To overcome these limitations, researchers have conducted a large number of explorations and studies, trying to find methods that can effectively improve the training efficiency of DNN on the device side.
[0004] As an efficient convolution calculation optimization method, the Winograd algorithm has achieved remarkable results in accelerating the convolution operation in DNN. The algorithm decomposes the standard convolution into a series of smaller sub-convolution blocks, and adopts a calculation process of pre-transformation, element-wise multiplication, and inverse transformation, significantly reducing the number of multiplications required for the convolution operation, thereby improving the computing efficiency. However, although the Winograd algorithm performs well in accelerating the convolution operation, it cannot completely solve the problems of computing resources and memory bandwidth limitations faced by DNN training on the device side. Summary of the Invention
[0005] This application provides a method and an accelerator for accelerating the training of deep neural networks to solve the problem that the improvement of training efficiency is restricted by computing resources and memory bandwidth in deep neural networks.
[0006] In the first aspect of this application, a method for accelerating the training of deep neural networks is provided, and the method includes:
[0007] Normalize the convolutions in a deep neural network based on the Winograd algorithm to convert non-standard convolutions in the convolutions into standard convolutions; the non-standard convolutions include non-unit stride convolutions, transposed convolutions, and dilated convolutions.
[0008] Decompose the convolution kernels larger than 3×3 in the convolutional layer into multiple sub-convolution kernels to accelerate the standard convolution without enumerating the transformation matrices required for each size of convolution kernel.
[0009] Replace the convolutional layer with a Winograd layer with predefined convolution kernel parameters to directly update the weights in the Winograd domain; the predefined convolution kernel parameters include the preset convolution kernel size, stride, and weights.
[0010] Adopt an interleaved gradient scheduling strategy to train the Winograd layer, and perform convolution acceleration calculations on the input data and convolution kernels in the Winograd layer through the Winograd algorithm to obtain convolution results.
[0011] Transmit the convolution result to the memory module as the input of the subsequent network layer to achieve the scheduling and operation of the deep neural network.
[0012] The above method, through convolution decomposition and Winograd domain training, can not only standardize the convolution to enable the use of the Winograd algorithm to accelerate various convolution operations, but also use the Winograd layer to replace the convolutional layer to directly update the weights in the Winograd domain to increase the trainable parameters in the Winograd domain, further improving the accuracy of the deep neural network model. In addition, it unifies the preprocessing of errors in the error backpropagation and weight gradient calculation processes, providing the possibility for finer-grained interleaved gradient scheduling and saving Winograd transformation-related overhead.
[0013] Optionally, the steps for converting the non-unit stride convolution into a standard convolution include:
[0014] Aggregate the convolution kernels and input data in an interleaved manner.
[0015] Divide the aggregated convolution kernels and input data into multiple sub-convolutions with unit stride.
[0016] Optionally, the steps for converting the transposed convolution into a standard convolution include:
[0017] Aggregate the convolution kernels into multiple sub-convolution kernels in an interleaved manner.
[0018] Perform convolution operations on each sub-convolution kernel and the corresponding input data to obtain multiple output data.
[0019] Disperse and splice multiple output data in an interleaved manner to obtain the final output data of the deconvolution.
[0020] Optionally, the steps of converting the dilated convolution into a standard convolution include:
[0021] Aggregate the input data into multiple sub-input data according to a preset dilation rate.
[0022] Successively use convolutional kernels to perform convolutional operations on each sub-input data to obtain multiple sub-output data.
[0023] Disperse and splice multiple sub-output data in an interleaved manner to ensure that the convolutional results are correctly distributed in the output space.
[0024] Optionally, while decomposing convolutional kernels larger than 3×3 in the convolutional layer into multiple sub-convolutional kernels, an aggregation layer and an accumulation layer are also generated.
[0025] The aggregation layer is used to downsample the input data during the forward propagation process to reduce the dimension of the input data and retain key features; during the backward propagation process, it is used to propagate the output gradient from the next layer back to the previous layer.
[0026] The accumulation layer is used to accumulate the input data and then output it during the forward propagation process; during the backward propagation process, it is used to directly propagate the output gradient.
[0027] Optionally, the steps of training the Winograd layer using an interleaved gradient scheduling strategy include:
[0028] Perform forward propagation on the Winograd layer to calculate the network output.
[0029] Calculate the error data based on the network output and the true label.
[0030] Use the interleaved gradient scheduling strategy to perform backward propagation with the error data as the input of the backward propagation to calculate the gradients of each layer.
[0031] Use the interleaved gradient scheduling strategy to calculate the weight gradient with the error data as the convolutional kernel to update the convolutional kernel parameters.
[0032] Optionally, the method further includes:
[0033] According to the capacity of the static random access memory, split the activation data and the error data into multiple SRAM-level activation blocks and error blocks respectively to ensure that each activation block and error block can be allocated to the static random access memory for storage.
[0034] The activation blocks and error blocks at each SRAM level are further divided into multiple activation granules and error granules at the register level according to the amount of computation that can be performed in each cycle and the capacity of the registers, so as to form a fused data stream oriented to Winograd.
[0035] The fused data stream oriented to Winograd not only realizes the reuse of errors at the register level and SRAM level, but also realizes the organic combination of the Winograd algorithm and the interleaved gradient scheduling, greatly reducing the on-chip and off-chip memory access while reducing the convolution calculation complexity.
[0036] Optionally, the steps of performing convolution acceleration calculation on the input data and the convolution kernel in the Winograd layer by the Winograd algorithm include:
[0037] Transform the input data to obtain a transformed input data matrix.
[0038] Transform the convolution kernel to obtain a transformed convolution kernel matrix.
[0039] Use the Winograd algorithm to calculate the transformed input data matrix and the transformed convolution kernel matrix to obtain an intermediate result matrix.
[0040] Convert the intermediate result matrix back to the standard convolution output format to obtain the convolution result.
[0041] The second aspect of this application also provides a deep neural network training accelerator, which is applicable to the deep neural network training acceleration method described in the first aspect. The training accelerator includes a preprocessing module, a Winograd module, a calculation execution module, and a memory module.
[0042] The preprocessing module is used to standardize the convolution in the deep neural network based on the Winograd algorithm, so that the non-standard convolution in the convolution is converted into standard convolution. The non-standard convolution includes non-unit stride convolution, deconvolution, and dilated convolution.
[0043] Decompose the convolution kernel larger than 3×3 in the convolution layer into multiple sub-convolution kernels to accelerate the standard convolution without enumerating the transformation matrices required for each size of convolution kernel.
[0044] The Winograd module is used to replace the convolution layer with a Winograd layer with predefined convolution kernel parameters, so as to directly update the weights in the Winograd domain.
[0045] The computing execution module is used to train the Winograd layer by adopting an interleaved gradient scheduling strategy, and perform convolution acceleration calculation on the input data and convolution kernel in the Winograd layer through the Winograd algorithm to obtain a convolution result.
[0046] The memory module is used to store the convolution result as the input of the subsequent network layer to realize the scheduling and operation of the deep neural network.
[0047] The above-mentioned deep neural network training accelerator, through convolution decomposition and Winograd domain training, can not only standardize convolution so as to be able to use the Winograd algorithm to accelerate various convolution operations, but also replace the convolution layer with the Winograd layer and directly update the weights in the Winograd domain to increase the trainable parameters in the Winograd domain, further improving the accuracy of the deep neural network model. In addition, it unifies the preprocessing of errors in the error backpropagation and weight gradient calculation processes, providing the possibility for finer-grained interleaved gradient scheduling and saving Winograd transformation-related overhead.
[0048] Optionally, it further includes a fused data module; the memory module includes a dynamic random access memory, a static random access memory, and registers.
[0049] The fused data module is used to divide the activation data and error data into multiple SRAM-level activation blocks and error blocks respectively according to the capacity of the static random access memory, so as to ensure that each activation block and error block can be allocated to the static random access memory for storage.
[0050] According to the amount of computation that can be executed in each cycle and the capacity of the registers, each SRAM-level activation block and error block are further divided into multiple register-level activation granules and error granules to form a Winograd-oriented fused data stream.
[0051] The dynamic random access memory includes activation data, weight data, and error data.
[0052] The static random access memory includes SRAM-level activation blocks and error blocks.
[0053] The registers include register-level activation granules and error granules.
[0054] The Winograd-oriented fused data stream not only realizes the reuse of errors at the register level and SRAM level, but also realizes the organic combination of the Winograd algorithm and interleaved gradient scheduling, greatly reducing on-chip and off-chip memory access while reducing the convolution calculation complexity.
[0055] As can be seen from the above technical solutions, the present application provides a method and an accelerator for accelerating the training of a deep neural network. The convolution in the deep neural network is standardized based on the Winograd algorithm, so that the non-standard convolution in the convolution is converted into a standard convolution; the non-standard convolution includes non-unit stride convolution, deconvolution, and dilated convolution; the convolution kernel larger than 3×3 in the convolution layer is decomposed into multiple sub-convolution kernels to accelerate the standard convolution without enumerating the transformation matrices required for each size of the convolution kernel; the convolution layer is replaced with a Winograd layer with predefined convolution kernel parameters, so as to directly update the weights in the Winograd domain; the predefined convolution kernel parameters include the preset convolution kernel size, stride, and weights; the Winograd layer is trained using an interleaved gradient scheduling strategy, and the convolution acceleration calculation is performed on the input data and the convolution kernel in the Winograd layer through the Winograd algorithm to obtain a convolution result; the convolution result is transmitted to the memory module as the input of the subsequent network layer to realize the scheduling and operation of the deep neural network. The problem of the limited training efficiency caused by the computing resources and memory bandwidth in the deep neural network is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0057] Figure 1 It is a schematic flowchart of the method for accelerating the training of a deep neural network according to an embodiment of the present application;
[0058] Figure 2 It is a schematic diagram of the execution order of the interleaved gradient scheduling and the traditional computing scheduling in an embodiment of the present application;
[0059] Figure 3 It is a schematic diagram of training a deep neural network by CWTS in the method for accelerating the training of a deep neural network according to an embodiment of the present application;
[0060] Figure 4 It is a schematic diagram of the fusion data stream oriented to the Winograd algorithm in the method for accelerating the training of a deep neural network according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] The embodiments will be described in detail below, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following embodiments do not represent all embodiments consistent with the present application. They are only examples of systems and methods consistent with some aspects of the present application.
[0062] AI has been widely adopted in display applications such as face recognition, object detection, and autonomous driving. Training deep neural networks for these applications involves a large amount of memory access and computation. Therefore, AI applications based on deep neural networks usually adopt an offline training and online inference strategy. However, with the growing demand for DNNs to be able to continuously learn from new data after deployment, and considering privacy, security, and communication overhead, it is best to be able to directly train on the device side. The most significant of these are the limitations of computing resources and memory bandwidth.
[0063] The Winograd algorithm is a mathematical algorithm for optimizing convolution operations, aiming to reduce the amount of computation and accelerate the convolution process. The algorithm decomposes the standard convolution into a series of smaller sub-convolution blocks and adopts a computational process of pre-transformation, element-wise multiplication, and inverse transformation, reducing the multiplication operations required for convolution. Taking the one-dimensional convolution F(2,3) as an example, ordinary matrix multiplication requires six multiplications, while the Winograd algorithm achieves one-dimensional convolution with only four multiplications, achieving a 1.5-fold acceleration. However, although the Winograd algorithm performs well in accelerating convolution operations, it cannot completely solve the problems of computing resources and memory bandwidth limitations faced by device-side DNN training.
[0064] See Figure 2 , staggered gradient scheduling is a scheduling method used in distributed training to optimize the utilization of computing resources and improve training efficiency. The training process of a deep neural network training model is divided into multiple stages. Staggered gradient scheduling interleaves these stages on multiple devices in a pipelined manner, making full use of the computing capabilities of different devices at different times, reducing the idle time of the devices, and thus improving the overall training efficiency. See specifically Figure 2 in (a) and (b) therein. Since error backpropagation (BP) and weight gradient calculation (WG) are two independent operations, these two calculations are usually carried out separately in traditional computational scheduling. Staggered gradient scheduling takes advantage of the fact that the error serves as an operand in both the BP and WG stages, and by interleaving the computational blocks of BP and the computational blocks of WG, it reuses the error and reduces memory access. Specifically, considering the limited on-chip storage, the computational tasks are decomposed into smaller computational blocks. Staggered gradient scheduling selects to first perform error backpropagation after loading a part of the error data, and then immediately use the same data for weight gradient calculation. This scheduling method achieves full reuse of the error, thus reducing the number of memory accesses.
[0065] This application aims to combine the Winograd algorithm with staggered gradient scheduling to further accelerate the training of deep neural networks. However, the error, as a common operand in the BP and WG stages, plays different roles in the two stages: as input data in BP and as a convolution kernel in WG. Therefore, when using the Winograd algorithm to accelerate training, the error needs to be preprocessed differently, including different data chunking and pre-transformations. At this time, although staggered gradient scheduling can still achieve SRAM-level data chunk reuse, thereby reducing off-chip memory access, it cannot achieve finer-grained data reuse, so it cannot reduce on-chip memory access. Therefore, directly combining the Winograd algorithm with staggered gradient scheduling will limit the acceleration effect of staggered gradient scheduling.
[0066] To solve the problem that the computing resources and memory bandwidth in deep neural networks limit the improvement of training efficiency, see Figure 1 , some embodiments of this application provide a method for accelerating the training of deep neural networks, and the method includes:
[0067] S100: Normalize the convolution in the deep neural network based on the Winograd algorithm so that non-standard convolutions in the convolution are converted into standard convolutions.
[0068] It should be understood that in order to use the Winograd algorithm to accelerate non-standard convolutions, non-standard convolutions need to be converted into standard convolutions. The non-standard convolutions include non-unit stride convolutions, transposed convolutions, and dilated convolutions.
[0069] In some embodiments, the steps of converting the non-unit stride convolution into a standard convolution include:
[0070] Aggregate the convolution kernel and input data in a staggered manner.
[0071] Divide the aggregated convolution kernel and input data into multiple sub-convolutions with unit stride.
[0072] It should be understood that the staggered manner can be channel dimension staggering or mixed staggering; channel dimension staggering can combine elements from different channels in a staggered manner. For example, elements from different channels can be staggered at specific spatial positions. Mixed staggering refers to a staggering method that combines the spatial dimension and the channel dimension. For example, the convolution kernel can be staggered first in the horizontal or vertical direction, and then further staggered in the channel dimension.
[0073] In some embodiments, the steps of converting the transposed convolution into a standard convolution include:
[0074] Aggregate the convolution kernel into multiple sub-convolution kernels in a staggered manner.
[0075] Perform a convolution operation on each sub-convolution kernel and the corresponding input data to obtain multiple output data.
[0076] Dispersedly splice the multiple output data in an interleaved manner to obtain the final output data of the transposed convolution.
[0077] It should be understood that the traditional method of converting the transposed convolution into a standard convolution is to pad zeros in the input data. However, in the present invention, the convolution kernels are aggregated into multiple sub-convolution kernels in an interleaved manner, and there is no longer a need to pad the input data.
[0078] In some embodiments, the steps of converting the dilated convolution into a standard convolution include:
[0079] Aggregate the input data into multiple sub-input data according to a preset dilation rate.
[0080] Successively perform convolution operations on each sub-input data using a convolution kernel to obtain multiple sub-output data.
[0081] Dispersedly splice the multiple sub-output data in an interleaved manner to ensure that the convolution result is correctly distributed in the output space.
[0082] It should be understood that the traditional method of converting the dilated convolution into a standard convolution requires zero-padding inside the convolution kernel according to the dilation rate, which will undoubtedly bring out a lot of invalid calculations (i.e., multiplications involving zeros). In the present invention, the input data is aggregated into multiple sub-input data according to a preset dilation rate, and then successively convolved with the convolution kernel. Finally, the multiple sub-output data are dispersedly spliced in an interleaved manner to ensure that the convolution result is correctly distributed in the output space.
[0083] S200: Decompose the convolution kernels larger than 3×3 in the convolutional layer into multiple sub-convolution kernels to accelerate the standard convolution without enumerating the transformation matrices required for each size of convolution kernel.
[0084] It should be understood that from the perspectives of efficiency and accuracy, decomposing the convolution kernels larger than 3×3 into multiple sub-convolution kernels can not only accelerate the standard convolution without enumerating the transformation matrices required for each size of convolution kernel, but also utilize the simple advantage of the transformation matrices of small-size Winograd blocks, and various transformations can be achieved through shift and addition operations, saving hardware overhead.
[0085] In some embodiments, while decomposing the convolution kernels larger than 3×3 in the convolutional layer into multiple sub-convolution kernels, an aggregation layer and an accumulation layer are also generated.
[0086] The aggregation layer is used to downsample the input data during forward propagation to reduce the dimension of the input data and retain key features; during backpropagation, it is used to propagate the output gradient from the next layer back to the previous layer.
[0087] The accumulation layer is used to accumulate the input data and then output it during forward propagation; during backpropagation, it is used to directly propagate the output gradient.
[0088] S300: Replace the convolutional layer with a Winograd layer with predefined convolutional kernel parameters so as to directly update the weights in the Winograd domain.
[0089] It should be understood that the predefined convolutional kernel parameters include the predefined convolutional kernel size, stride, and weights. Specifically, in order to further reduce the computational overhead associated with the Winograd transform, the convolutional layer can be replaced with a Winograd layer with a convolutional kernel size of 4×4 and a stride of ws so as to directly update the weights in the Winograd domain. This not only eliminates the weight transformation during the training process, but also improves the accuracy of the network to a certain extent because the size of the weights in the Winograd domain is larger than that of the weights in the spatial domain.
[0090] The above convolutional decomposition and Winograd-domain training scheme (Convolution-Decomposition Winograd-Domain Training Scheme, CWTS) unifies the preprocessing required for the error in the BP and WG stages. See Figure 3 that the Winograd module contains network layers that perform fixed operations and multiple Winograd layers. See Figure 3In (a), taking the use of CWTS for a convolutional layer with a convolutional kernel size of 5×5 and a stride of 2 as an example, the figure shows the connection relationship between the internal layers of the Winograd module. Among them, CONV is the abbreviation of the convolutional layer, Conv is the abbreviation of the standard convolutional layer with a stride of 1, and WinoLayer is the abbreviation of the Winograd layer. C is the number of input channels, M is the number of output channels, Hi and Wi are the height and width of the input, Ho and Wo are the height and width of the output, s is the stride of the convolutional layer, ws is the stride of the Winograd layer, Ah, Aw, Bh, and Bw are the transformation matrices involved in the Winograd algorithm, and T represents the matrix transpose. It can be seen that the normalization step and the decomposition step can be flexibly configured. After the normalization step, the convolutional layer is divided into four sub-convolutions of different sizes, accompanied by the generation of an aggregation layer and an accumulation layer. Since all the generated sub-convolutional kernels are smaller than 4×4, no further decomposition is required. Finally, the four sub-convolutional layers are replaced by four Winograd layers, and the stride of each Winograd layer is (5 - rh, 5 - rw), where rh and rw are the sizes of the sub-convolutional kernels before replacement. Figure 3 In (b) and (c), the forward propagation and backward propagation mechanisms of the aggregation layer and the block layer with a stride of 2 are respectively shown. The block layer in (c) corresponds to the block layer generated when a 4×4 convolution is divided into four 2×2 sub-convolutions. As for the accumulation layer, its forward propagation mechanism is to accumulate all the inputs and then output, and the backward propagation mechanism is to directly propagate the output gradient without any additional operations.
[0091] S400: Train the Winograd layer by adopting an interleaved gradient scheduling strategy, and perform convolution acceleration calculation on the input data and the convolutional kernel in the Winograd layer through the Winograd algorithm to obtain a convolution result.
[0092] It should be understood that training a deep neural network on an accelerator involves the frequent movement of activations (A), weights (W), and errors (E) between the hierarchical memories of the accelerator. Among them, the activation (A) represents the activation state of the neurons in each layer of the neural network. By reducing the frequent movement of activations between the hierarchical memories of the accelerator, the data transfer overhead can be reduced and the computing efficiency can be improved. The weight (W) is one of the core parameters of the neural network. In the fused data stream oriented to the Winograd algorithm, by reducing the frequent movement of weights between the hierarchical memories of the accelerator, the computing efficiency and the stability of model training can be improved. The error (E) refers to the difference between the predicted output of the neural network and the true label. In the fused data stream oriented to Winograd, the calculation and feedback of errors are crucial for model training. Reducing the frequent movement of errors between the hierarchical memories of the accelerator can improve the training efficiency and stability.
[0093] To reduce memory access, in some embodiments, referring to Figure 4 , the method further includes:
[0094] According to the capacity of the static random access memory, the activation data and error data are respectively divided into multiple activation blocks and error blocks at the SRAM level to ensure that each activation block and error block can be allocated to the static random access memory for storage.
[0095] According to the amount of computation that can be performed in each cycle and the capacity of the registers, each activation block and error block at the SRAM level are further divided into multiple activation granules and error granules at the register level to form a fused data stream oriented to Winograd.
[0096] It should be understood that Figure 4 in (a) and (b) are the block diagrams of the four-dimensional iteration space of the operation data for DNN training and the mapping of the corresponding data stream in time and space. Among them, the storage at the SRAM level is named Abuf and Ebuf, and the storage at the register level is named Areg, Ereg, and Wreg. Figure 4 in (c) is the schematic diagram of the hierarchical structure of the memory module. The dynamic random access memory (DRAM) is located off-chip, and the static random access memory (SRAM) and registers are located on-chip. DRAM, SRAM, and registers are interconnected. First, the activation data and error data are divided into multiple activation blocks and error blocks at the SRAM level and loaded into the on-chip SRAM at one time, and the relevant calculations are performed. Among them, Pi and Po respectively refer to the parallelism of the input channel and the output channel and can be freely configured. Since the on-chip computing resources are limited and the amount of computation that can be performed in one cycle is limited, it is necessary to further divide the activation blocks and error blocks at the SRAM level into multiple activation granules and error granules at the register level. The size (mh, mw) of the error granules is the minimum granularity of the data that can be interleaved and scheduled.
[0097] The fused data stream oriented to Winograd not only realizes the reuse of errors at the register level and SRAM level, but also realizes the organic combination of the Winograd algorithm and the interleaved gradient scheduling, greatly reducing the on-chip and off-chip memory access while reducing the convolution computing complexity.
[0098] In some embodiments, the steps of training the Winograd layer by using the interleaved gradient scheduling strategy include:
[0099] Performing forward propagation on the Winograd layer to calculate the network output.
[0100] Calculating the error data according to the network output and the true label.
[0101] The error data is used as the input for backpropagation using the staggered gradient scheduling strategy for backpropagation to calculate the gradients of each layer.
[0102] The error data is used as a convolution kernel using the staggered gradient scheduling strategy to calculate the weight gradients for updating the convolution kernel parameters.
[0103] It should be understood that Figure 3 In (d), the calculations of the Winograd layer with a stride of (2, 2) at different training stages are shown. It can be seen that the preprocessing of the error in the error backpropagation and weight gradient calculation processes of the Winograd layer is exactly the same. This consistency provides an opportunity for more fine-grained interleaving of BP and WG to reduce on-chip memory access.
[0104] In some embodiments, the steps of performing convolution acceleration calculation on the input data and the convolution kernel in the Winograd layer by the Winograd algorithm include:
[0105] Transform the input data to obtain the transformed input data matrix.
[0106] Transform the convolution kernel to obtain the transformed convolution kernel matrix.
[0107] Use the Winograd algorithm to calculate the transformed input data matrix and the transformed convolution kernel matrix to obtain the intermediate result matrix.
[0108] Convert the intermediate result matrix back to the standard convolution output format to obtain the convolution result.
[0109] S500: Transmit the convolution result to the memory module as the input for the subsequent network layer to achieve the scheduling and operation of the deep neural network.
[0110] The above method, through convolution decomposition and Winograd domain training, can not only standardize the convolution to enable the use of the Winograd algorithm to accelerate various convolution operations, but also replace the convolution layer with the Winograd layer to directly update the weights in the Winograd domain to increase the trainable parameters in the Winograd domain, further improving the accuracy of the deep neural network model. In addition, it unifies the preprocessing of the error in the error backpropagation and weight gradient calculation processes, providing the possibility for more fine-grained staggered gradient scheduling and saving the overhead related to the Winograd transformation.
[0111] Some embodiments of the present application also provide a deep neural network training accelerator applicable to the deep neural network training acceleration method described in the above embodiments. The training accelerator includes a preprocessing module, a Winograd module, a calculation execution module, and a memory module.
[0112] The preprocessing module is used to standardize the convolution in the deep neural network based on the Winograd algorithm, so that the non-standard convolution in the convolution is converted into a standard convolution; the non-standard convolution includes non-unit stride convolution, deconvolution, and dilated convolution.
[0113] Decompose the convolution kernels larger than 3×3 in the convolutional layer into multiple sub-convolution kernels to accelerate the standard convolution without enumerating the transformation matrices required for each size of convolution kernel.
[0114] The Winograd module is used to replace the convolutional layer with a Winograd layer with predefined convolution kernel parameters, so as to directly update the weights in the Winograd domain.
[0115] The computing execution module is used to train the Winograd layer by adopting an interleaved gradient scheduling strategy, and perform convolution acceleration calculation on the input data and convolution kernels in the Winograd layer through the Winograd algorithm to obtain a convolution result.
[0116] The memory module is used to store the convolution result as the input of the subsequent network layer to realize the scheduling and operation of the deep neural network.
[0117] The above-mentioned deep neural network training accelerator, through convolution decomposition and Winograd domain training, can not only standardize the convolution to enable the use of the Winograd algorithm to accelerate various convolution operations, but also use the Winograd layer to replace the convolutional layer to directly update the weights in the Winograd domain to increase the trainable parameters in the Winograd domain, further improving the accuracy of the deep neural network model. In addition, it unifies the preprocessing of errors in the error backpropagation and weight gradient calculation processes, providing the possibility for finer-grained interleaved gradient scheduling and saving Winograd transformation-related overhead.
[0118] In some embodiments, it further includes a fused data module; the memory module includes a dynamic random access memory, a static random access memory, and registers.
[0119] The fused data module is used to divide the activation data and error data into multiple SRAM-level activation blocks and error blocks respectively according to the capacity of the static random access memory, so as to ensure that each activation block and error block can be allocated to the static random access memory for storage.
[0120] Further divide each SRAM-level activation block and error block into multiple register-level activation granules and error granules according to the amount of computation that can be executed in each cycle and the capacity of the registers to form a Winograd-oriented fused data stream.
[0121] The dynamic random access memory includes activation data, weight data, and error data.
[0122] The static random access memory includes activation blocks and error blocks at the SRAM level.
[0123] The register includes activation grains and error grains at the register level.
[0124] It should be understood that since the dynamic random access memory, the static random access memory, and the register communicate with each other, the register also includes weight data.
[0125] The Winograd-oriented fusion data stream not only realizes the reuse of errors at the register level and the SRAM level, but also realizes the organic combination of the Winograd algorithm and the staggered gradient scheduling, greatly reducing the on-chip and off-chip memory access while reducing the convolution calculation complexity.
[0126] From the above technical solutions, the embodiments of the present application provide a method and an accelerator for accelerating the training of a deep neural network, which standardize the convolution in the deep neural network based on the Winograd algorithm, so that the non-standard convolution in the convolution is converted into a standard convolution; the non-standard convolution includes non-unit stride convolution, deconvolution, and dilated convolution; the convolution kernel larger than 3×3 in the convolution layer is decomposed into multiple sub-convolution kernels to accelerate the standard convolution without enumerating the transformation matrices required for each size of the convolution kernel; the convolution layer is replaced with a Winograd layer with predefined convolution kernel parameters to directly update the weights in the Winograd domain; the predefined convolution kernel parameters include the preset convolution kernel size, stride, and weights; the Winograd layer is trained using the staggered gradient scheduling strategy, and the input data and the convolution kernel in the Winograd layer are convolved and accelerated through the Winograd algorithm to obtain a convolution result; the convolution result is transmitted to the memory module as the input of the subsequent network layer to realize the scheduling and operation of the deep neural network. Solve the problem that the training efficiency improvement is limited by the computing resources and memory bandwidth in the deep neural network.
[0127] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other embodiments extended based on the solution of the present application without creative efforts belong to the protection scope of the present application.
Claims
1. A method for accelerating the training of a deep neural network, characterized in that, The method includes: Normalize the convolution in the deep neural network based on the Winograd algorithm to convert non-standard convolutions in the convolution into standard convolutions; the non-standard convolutions include non-unit stride convolutions, transposed convolutions, and dilated convolutions; Decompose the convolution kernels larger than 3×3 in the convolution layer into multiple sub-convolution kernels to accelerate the standard convolution without enumerating the transformation matrices required for each size of convolution kernel; Replace the convolution layer with a Winograd layer with predefined convolution kernel parameters to directly update the weights in the Winograd domain; the predefined convolution kernel parameters include the preset convolution kernel size, stride, and weights; Adopt an interleaved gradient scheduling strategy to train the Winograd layer, and perform convolution acceleration calculations on the input data and convolution kernels in the Winograd layer through the Winograd algorithm to obtain convolution results; Transmit the convolution results to the memory module as the input of the subsequent network layer to achieve the scheduling and operation of the deep neural network.
2. The method for accelerating the training of a deep neural network according to claim 1, wherein The steps for converting the non-unit stride convolution into a standard convolution include: Aggregate the convolution kernel and input data in an interleaved manner; Divide the aggregated convolution kernel and input data into multiple sub-convolutions with unit stride.
3. The deep neural network training acceleration method according to claim 1, characterized in that The steps for converting the transposed convolution into a standard convolution include: Aggregate the convolution kernel into multiple sub-convolution kernels in an interleaved manner; Perform convolution operations on each sub-convolution kernel and the corresponding input data to obtain multiple output data; Disperse and splice the multiple output data in an interleaved manner to obtain the final output data of the transposed convolution.
4. The deep neural network training acceleration method according to claim 1, characterized in that, The steps for converting the dilated convolution into a standard convolution include: Aggregate the input data into multiple sub-input data according to the preset dilation rate; Successively use the convolution kernels to perform convolution operations on each sub-input data to obtain multiple sub-output data; Disperse and splice the multiple sub-output data in an interleaved manner to ensure that the convolution results are correctly distributed in the output space.
5. The deep neural network training acceleration method according to claim 1, characterized in that When decomposing the convolution kernels larger than 3×3 in the convolution layer into multiple sub-convolution kernels, an aggregation layer and an accumulation layer are also generated; The aggregation layer is used for downsampling the input data during the forward propagation process to reduce the dimension of the input data and retain key features; During the backward propagation process, it is used to propagate the output gradient from the next layer back to the previous layer; The accumulation layer is used for accumulating the input data and then outputting it during the forward propagation process; during the backward propagation process, it is used to directly propagate the output gradient.
6. The deep neural network training acceleration method according to claim 1, wherein The steps for training the Winograd layer using the interleaved gradient scheduling strategy include: Perform forward propagation on the Winograd layer to calculate the network output; Calculate the error data based on the network output and the true labels; Adopt the interleaved gradient scheduling strategy to use the error data as the input for backward propagation for backward propagation to calculate the gradients of each layer; Adopt the interleaved gradient scheduling strategy to use the error data as the convolution kernel for weight gradient calculation to update the convolution kernel parameters.
7. The deep neural network training acceleration method according to claim 6, wherein The method also includes: According to the capacity of the static random access memory, the activation data and the error data are respectively divided into multiple activation blocks and error blocks at the SRAM level to ensure that each activation block and error block can be allocated to the static random access memory for storage; According to the amount of computation that can be executed in each cycle and the capacity of the registers, each activation block and error block at the SRAM level are further divided into multiple activation granules and error granules at the register level to form a fused data stream for Winograd.
8. The method for accelerating the training of a deep neural network according to claim 1, wherein The steps of performing convolution acceleration calculation on the input data and the convolution kernel in the Winograd layer by the Winograd algorithm include: Transforming the input data to obtain a transformed input data matrix; Transforming the convolution kernel to obtain a transformed convolution kernel matrix; Using the Winograd algorithm to calculate the transformed input data matrix and the transformed convolution kernel matrix to obtain an intermediate result matrix; Converting the intermediate result matrix back to the standard convolution output format to obtain the convolution result.
9. A deep neural network training accelerator, characterized in that, Applicable to the deep neural network training acceleration method described in any one of claims 1-8, the training accelerator includes a preprocessing module, a Winograd module, a calculation execution module, and a memory module; The preprocessing module is used to standardize the convolution in the deep neural network based on the Winograd algorithm, so that the non-standard convolution in the convolution is converted into a standard convolution; the non-standard convolution includes non-unit stride convolution, deconvolution, and dilated convolution; Decomposing the convolution kernel larger than 3×3 in the convolution layer into multiple sub-convolution kernels to accelerate the standard convolution without enumerating the transformation matrices required for each size of convolution kernel; The Winograd module is used to replace the convolution layer with a Winograd layer with predefined convolution kernel parameters for direct weight update in the Winograd domain; The calculation execution module is used to train the Winograd layer by adopting an interleaved gradient scheduling strategy and perform convolution acceleration calculation on the input data and the convolution kernel in the Winograd layer by the Winograd algorithm to obtain the convolution result; The memory module is used to store the convolution result as the input of the subsequent network layer to realize the scheduling and operation of the deep neural network.
10. The deep neural network training accelerator according to claim 9, characterized in that, It further includes a fused data module; the memory module includes a dynamic random access memory, a static random access memory, and registers; The fused data module is used to divide the activation data and the error data into multiple activation blocks and error blocks at the SRAM level according to the capacity of the static random access memory to ensure that each activation block and error block can be allocated to the static random access memory for storage; According to the amount of computation that can be executed in each cycle and the capacity of the registers, each activation block and error block at the SRAM level are further divided into multiple activation granules and error granules at the register level to form a fused data stream for Winograd; The dynamic random access memory includes activation data, weight data, and error data; The static random access memory includes an activation block and an error block at the SRAM level; The register includes activation particles and error particles at the register level.