Neural network training accelerator and acceleration method and device

By designing a neural network training accelerator that includes input registers, processing unit arrays, index matching units and controllers, the problem of low computing resource utilization in lightweight network training by existing accelerators is solved, and efficient training and resource utilization of standard and lightweight networks are achieved.

CN120146121APending Publication Date: 2025-06-13SUZHOU YUANYONG TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510217476.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When deploying lightweight networks, existing neural network training accelerators have low computing resource utilization and cannot efficiently support the training of both standard networks and lightweight networks.

Method used

A neural network training accelerator is designed, including input registers, processing unit arrays, index matching units and controllers, and the calculation of multiple output channels is processed in parallel, and the processing unit is fully utilized by setting the index matching units to improve the utilization of computing resources.

Benefits of technology

It realizes efficient training of standard neural networks and lightweight neural networks, improves the utilization rate of computing resources, and can flexibly meet the needs of different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146121A_ABST
    Figure CN120146121A_ABST
Patent Text Reader

Abstract

The invention provides a neural network training accelerator and an acceleration method and device.The training accelerator comprises an input register, a processing unit array, index matching units and a controller, the processing unit array comprises a plurality of processing units arranged in rows and columns, each index matching unit corresponds to one row of processing units, and each index matching unit corresponds to one row of processing units; each input register corresponds to a column of processing units, the controller is used for configuring a data path, the data path comprises a column direction path and a row direction path, the column direction path indicates that weight data is transmitted based on the direction of one column of processing units, and the row direction path indicates that the weight data is transmitted based on the direction of one row of processing units; the processing unit is further configured to perform a multiplication operation based on the valid data, the input data, and the weight data to output a feature map or gradient data. The training accelerator can realize less extra hardware through parallel processing of calculation of a plurality of output channels and an index matching unit, so as to improve the utilization rate of calculation resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and particularly to a neural network training accelerator, an acceleration method, and a device. Background Art

[0002] Standard deep neural networks use standard convolutions. As the model scale expands, the computational complexity and memory requirements increase, making it difficult to deploy and run these models on resource-constrained devices. This can be achieved through lightweight networks, which use depthwise separable convolutions to significantly reduce model parameters and improve computational speed.

[0003] Although lightweight networks can improve hardware efficiency, most training accelerators are designed based on standard networks. These accelerators improve throughput by designing a spatially parallel computing array in the channel dimension to achieve fast training. However, depth convolutions do not involve accumulation in the channel dimension, resulting in many computing units being idle when deploying lightweight network training on the accelerator, thus being unsuitable for the training of lightweight networks.

[0004] Some work specifically designs training accelerators for lightweight networks, but these accelerators have insufficient throughput problems when processing standard networks. To support the efficient training of both standard networks and lightweight networks simultaneously, some training accelerators design two independent computing engines for the training of standard networks and lightweight networks respectively, but these accelerators have low computational resource utilization problems. Summary of the Invention

[0005] The present application provides a neural network training accelerator, an acceleration method, and a device to solve the problem of low computational resource utilization of the accelerator.

[0006] In a first aspect, the present application provides a neural network training accelerator, including:

[0007] An input register, including a plurality of data channels for storing input data;

[0008] A processing unit array, including a plurality of processing units arranged in rows and columns;

[0009] An index matching unit, each index matching unit corresponding to a row of the processing units, and each input register corresponding to a column of the processing units;

[0010] The index matching unit is used to obtain valid data;

[0011] A controller for configuring a data path, the data path including a column - direction path and a row - direction path, the column - direction path indicating the direction of transfer of the weight data based on a column of processing units, and the row - direction path indicating the direction of transfer of the weight data based on a row of processing units;

[0012] The processing unit is further configured to perform a multiplication operation based on the valid data obtained by the index matching unit, the input data stored in the input register, and the weight data, and output a feature map or gradient data. The weight data is the weight data transferred by the previous processing unit in the data path where the processing unit is located. The feature map is a feature map generated based on the column - direction path or the row - direction path, and the gradient data is gradient data generated based on the row - direction path.

[0013] In some feasible embodiments, the index matching unit includes:

[0014] A sequence decoder for receiving bitmap data, retrieving the index of the data channel, and marking the valid index. The bitmap data is used to determine the validity of the data channel, and the valid index is the index of the valid channel;

[0015] A prefix - sum unit for generating an address sequence based on the bitmap data. The address sequence is a sequence calculated based on the total number of channels of the valid index;

[0016] A comparator for matching the valid index and the address sequence to determine the valid address. The valid address is the storage address of the valid data of the valid channel in the input register.

[0017] In some feasible embodiments, the number of columns of the processing unit is a first value, and the number of rows of the processing unit is a second value, and the second value is an integer multiple of the first value;

[0018] When the path of the weight data is configured in the column direction, the weight data is sequentially transferred within the processing unit and stored in the input register;

[0019] The processing units of the second value are configured to perform an element - wise multiplication operation on the input data and the weight data based on the valid data, and output a first feature map. The first feature map is a feature map obtained through standard convolutional forward propagation.

[0020] In some feasible embodiments, the index matching unit is configured to receive a pointer, and the pointer is used to record the number of valid channels;

[0021] When the path configuration of the weight data is in the row direction, configure the bitmap data as a one-hot code with a second number of bits, and configure the value of the pointer as 1. The processing unit is used to obtain the input data of multiple data channels from the input register, and the value of the pointer is used to represent the number of valid channels;

[0022] The comparator is used to compare the valid address and the pointer;

[0023] If the valid address is equal to the pointer, the comparator is further used to generate a switching signal, and the switching signal is used to control the index matching unit;

[0024] The processing unit is used to perform a convolution operation on the input data and the weight data of different data channels to generate a second feature map.

[0025] In some feasible embodiments, when the path configuration of the weight data is in the row direction, configure the bitmap data as binary bits of a second value, the value of the binary bit is 1, and configure the value of the pointer as the second value. The processing unit is used to obtain the input data of the same data channel from the input register;

[0026] The comparator is used to compare the valid address and the pointer;

[0027] If the valid address is equal to the pointer, the comparator is further used to generate a switching signal, and the switching signal is used to control the index matching unit;

[0028] The processing unit is used to obtain first gradient data, calculate second gradient data based on the first gradient data and the weight data, and update the weight data through the second gradient data. The first gradient data is data generated by the backpropagation of standard convolution.

[0029] In some feasible embodiments, a weight register is further included. When the path configuration of the weight data is in the column direction, the weight register is used to store the weight data.

[0030] In some feasible embodiments, a channel enhancement unit is further included;

[0031] The channel enhancement unit is used to obtain the number of output channels, set the number of input channel groups of the weight data based on the number of channels, and calculate the number of channels based on the number of channel groups;

[0032] The channel enhancement unit is further used to configure the value of the pointer as a third value, and configure the bitmap data as binary bits of a third value, the value of the binary bit is 1.

[0033] In some feasible embodiments, the channel enhancement unit is further configured to reorganize the weight data of the first dimension in the data channels of the first value to generate an input block of the second dimension;

[0034] The processing unit is configured to receive the input block of the second dimension to perform a convolution operation through the input block to generate a third feature map.

[0035] In a second aspect, the present application provides an acceleration method for a neural network training accelerator, including:

[0036] Obtain input data stored in an input register, where the input register includes a plurality of data channels for storing the input data;

[0037] Control an index matching unit to obtain valid data, where the index matching unit corresponds to a row of processing units, the input register corresponds to a column of the processing units, and the processing units are arranged in rows and columns;

[0038] Configure a data path, where the data path includes a column direction path and a row direction path, the column direction path indicates the direction of transmission of the weight data based on a column of processing units, and the row direction path indicates the direction of transmission of the weight data based on a row of processing units;

[0039] Control the processing unit to perform a multiplication operation based on the valid data obtained by the index matching unit, the input data stored in the input register, and the weight data to output a feature map or gradient data, where the weight data is the weight data transmitted by the previous processing unit in the data path where the processing unit is located, the feature map is a feature map generated based on the column direction path or the row direction path, and the gradient data is gradient data or weight gradient values generated based on the row direction path.

[0040] In a third aspect, the present application provides a neural network acceleration device, including an external memory and a neural network training accelerator.

[0041] As can be seen from the above technical solutions, the present application provides a neural network training accelerator, an acceleration method, and a device. The training accelerator includes an input register, a processing unit array, an index matching unit, and a controller. The input register includes multiple data channels for storing input data. The processing unit array includes multiple processing units arranged in rows and columns. Each index matching unit corresponds to a row of processing units, and each input register corresponds to a column of processing units. The index matching unit is used to obtain valid data. The controller is used to configure the data path. Among them, the data path includes a column direction path and a row direction path. The column direction path indicates the transfer of weight data in the direction of a column of processing units, and the row direction path indicates the transfer of weight data in the direction of a row of processing units. The processing unit is further used to perform a multiplication operation based on the valid data, input data, and weight data to output a feature map or gradient data. The training accelerator solves the problem of low utilization rate of computing resources in the accelerator by parallel processing the calculations of multiple output channels and making full use of the processing units with less additional hardware through the setting of the index matching unit. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0043] Figure 1 Structural schematic diagram of a training accelerator with two independent computing engines provided by an embodiment of the present application;

[0044] Figure 2 Structural schematic diagram of a neural network training accelerator provided by an embodiment of the present application;

[0045] Figure 3 Structural schematic diagram of an index matching unit provided by an embodiment of the present application;

[0046] Figure 4 Structural schematic diagram of a sequence decoder provided by an embodiment of the present application;

[0047] Figure 5 Structural schematic diagram of a prefix sum unit provided by an embodiment of the present application;

[0048] Figure 6 Schematic diagram of a channel enhancement process provided by an embodiment of the present application;

[0049] Figure 7 Acceleration method of a neural network training accelerator provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] Embodiments will be described in detail below, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following examples do not represent all embodiments consistent with the present application. They are merely examples of systems and methods consistent with some aspects of the present application detailed in the claims.

[0051] Convolutional neural networks can be applied in the field of computer vision. To address the problem of performance degradation due to domain shift and support user-specific services while considering privacy, security, and communication overhead, it can be achieved through an efficient device-side training accelerator. And deep neural network training is a highly computationally intensive task, but the resources on the device side are limited, resulting in difficulties in deploying deep neural network training on the device side.

[0052] Standard neural networks and lightweight neural networks are two types of models distinguished according to their complexity and resource consumption. Standard neural networks have a large number of parameters, a relatively complex structure, including multiple hidden layers, with thousands of neurons in each layer, and can learn very complex feature representations, requiring high computing resources, including memory, processor performance, etc.

[0053] Standard neural networks need to process a large amount of data and update a large parameter set, so the training process takes a long time and requires high-performance hardware support, such as GPU or TPU clusters. Standard neural networks are trained using large-scale datasets to ensure that they can learn the subtle patterns in the data and are prone to overfitting problems, so regularization techniques, such as Dropout, L2 regularization, etc., need to be used.

[0054] To address the problem of difficulties in deploying deep neural network training on the device side, it can be achieved by adopting a lightweight network. The design goal of a lightweight neural network is to reduce the model size, lower the computational cost, and improve the running efficiency. It has fewer parameters and a more simplified network structure and can be transformed from a large model through techniques such as pruning, quantization, knowledge distillation, etc.

[0055] The training time of a lightweight neural network is shorter, and the required computing resources are also less. Because the model is smaller, it can be more easily deployed to edge devices and can directly perform inference on these devices. However, its accuracy is lower than that of a standard neural network, and the efficiency is higher. In some cases, a lightweight model can quickly adapt to a new task through transfer learning, that is, by using the knowledge of a pre-trained standard model.

[0056] In summary, in terms of resource requirements, standard neural network training requires more computing resources and power, while lightweight neural networks are more energy-efficient and better adapted to low-power devices; in terms of training speed, due to the large number of parameters, the training of standard neural networks is slower, and lightweight networks train faster due to their simple structure; in terms of data dependence, standard neural networks require larger-scale datasets to avoid overfitting, while lightweight networks perform better on small datasets; in terms of deployment flexibility, lightweight networks are easier to deploy in different environments, and standard neural networks are more suitable for cloud or server-side applications with sufficient computing resources.

[0057] Standard neural networks tend to use standard convolutions, which can provide stronger feature extraction capabilities and are suitable for high-precision application scenarios. Lightweight neural networks employ depthwise convolutions or their variants, such as depthwise separable convolutions, to reduce parameters and computational burden, thus being suitable for application scenarios with limited resources.

[0058] However, some accelerators cannot support the efficient training of both standard networks and lightweight networks simultaneously, and cannot flexibly meet the requirements of different application scenarios.

[0059] The design of training accelerators is based on standard networks. These accelerators improve throughput by designing a spatially parallel computing array in the channel dimension to achieve fast training. However, depthwise convolutions do not involve accumulation in the channel dimension, resulting in most computing units being idle when deploying lightweight network training on these accelerators. Therefore, they are not suitable for lightweight network training. Some work focuses on designing training accelerators for lightweight networks, but these accelerators have insufficient throughput when processing standard networks.

[0060] To support the efficient training of both standard networks and lightweight networks simultaneously, as Figure 1 shown, in some embodiments, a system with two independent computing engines is set up, which are respectively used for the training of standard networks and lightweight networks.

[0061] Exemplarily, the two computing engines include a standard network computing engine and a lightweight network computing engine. The standard network computing engine is designed for standard neural networks, which have high computational complexity and large model sizes. The lightweight network computing engine is designed for lightweight neural networks, which usually have low computational complexity but require efficient resource utilization. For the two computing engines, containerization technologies, such as Docker or virtual machine technologies, can be used to create independent working environments for each computing engine, which can not only isolate different types of network training tasks but also facilitate environment configuration and management.

[0062] Among them, for the standard network computing engine, for the standard network model with high complexity and many parameters, a GPU with more memory and processing power can be used. For the lightweight network computing engine, for the lightweight network model with lower computing resource requirements, a relatively simple computing engine can be used.

[0063] In some embodiments, a task scheduling unit can also be set in the system. The task scheduling unit dynamically adjusts the resource allocation ratio between the standard network and the lightweight network according to the current computing load. For example, when the task volume of a certain type increases, the scheduling unit allocates more resources to the computing engine with the increased task volume.

[0064] However, if resources are statically allocated to each computing engine, then when the demand for a certain type of task, such as the standard network or the lightweight network, is low, the resources allocated to it may be idle, resulting in low overall resource utilization. The resource requirements of different types of tasks fluctuate over time, and a fixed resource allocation strategy cannot flexibly cope with this change.

[0065] Moreover, when using the data parallel method for training, if the batch size is too small, the parallel computing power of accelerator devices such as GPUs cannot be fully utilized. For some large models, although the model parallel technology can solve the memory limitation problem, its complexity and synchronization overhead may also affect the training efficiency, resulting in poor resource utilization.

[0066] In summary, setting independent computing engines for the standard network and the lightweight network still leads to the problem of low computing resource utilization.

[0067] To solve the problem of low computing resource utilization, some embodiments of the present application provide a neural network training accelerator that supports on-chip training of both standard neural networks and lightweight neural networks. By parallelly processing the calculations of multiple output channels and setting an index matching unit, the full utilization of the processing unit can be achieved with less additional hardware, so as to improve the computing resource utilization.

[0068] As Figure 2 shown, the neural network training accelerator provided by the embodiments of the present application includes an input register, a processing unit array, an index matching unit, and a controller.

[0069] The input register is a storage device including multiple data channels for storing input data. Among them, the input data is the output feature map from an external memory or the previous layer, and the input data is transmitted from the external memory to the input register through the data channels.

[0070] It can be understood that the input data can be data generated after preprocessing the original input data. For example, if the input data is image data, the preprocessing steps may include normalization, cropping, scaling, etc.; for another example, if the input data is audio data, the preprocessing steps may include sampling, filtering, etc.

[0071] The processing unit array is a computing structure formed by arranging multiple processing units in a row and column manner. The processing unit can be designed according to the neural network computing requirements. For example, multiple processing units are integrated on a chip through an integrated circuit manufacturing process.

[0072] Among them, the processing unit is composed of multiple processing elements, and each processing element has basic arithmetic operation capabilities, such as multiplication, addition, etc.

[0073] In some embodiments, the number of columns of the processing unit is a first value, and the number of rows of the processing unit is a second value. The second value is an integer multiple of the first value. In this embodiment, the computing array includes 32×16 processing units, and 16 processing units in each row are used to perform a set of 4×4 element-wise multiplications.

[0074] To match the processing unit, in this embodiment, the number of input registers is also 16. Each input register corresponds to a column of processing units. The processing unit can read the input data from the input register to perform calculations.

[0075] As Figure 3 shown, each index matching unit corresponds to a row of processing units. The index matching unit is used to obtain valid data. In a computing task, the data needs to be aligned in a specific manner. For example, in a convolutional neural network, data from different channels need to be aligned to perform convolutional operations. The index matching unit is used to select the correct subset from the dataset according to a predefined rule or pattern and pass them to the corresponding processing unit.

[0076] In some embodiments, the index matching unit includes a sequence decoder, a prefix sum unit, and a comparator. Among them, as Figure 4 shown, the sequence decoder is used to receive bitmap data. The bitmap data is used to represent a binary sequence of a set of data or states, where each bit represents the validity of a specific data item or state. In this embodiment, the bitmap data is used to determine the validity of data channels. For each data channel, 1 indicates that the channel is valid, and 0 indicates invalid.

[0077] The sequence decoder is also used to retrieve the indices of the data channels and mark the valid indices, where the valid indices are the indices of the valid channels. Exemplarily, bitmap data is first received, valid indices are generated based on the bitmap data, and a clear signal is generated to set the 1s at the retrieval positions to zero, which can efficiently identify the valid data channels, reduce the processing of invalid data, and thus improve the computing efficiency.

[0078] As Figure 5 shown, the prefix sum unit is used to generate an address sequence based on the bitmap data, where the address sequence is a sequence calculated based on the total number of channels of the valid indices. The prefix sum unit receives the valid index information from the sequence decoder and generates an address sequence according to the accumulated result by accumulating the number of valid channels before each valid index. By generating the address sequence, the storage position of the valid data in the register can be accurately located, ensuring the accuracy of data transmission.

[0079] The comparator is used to match the valid indices and the address sequence to determine the valid addresses, where the valid addresses are the storage addresses of the valid data of the valid channels in the input register.

[0080] In some embodiments, the index matching unit is used to receive a pointer input, where the pointer is used to record the number of valid channels. The comparator receives the address sequence and the pointer information from the prefix sum unit and performs a comparison operation. By comparing the pointer with the address generated by the prefix sum unit, it is determined whether all valid channels have been retrieved, and the processing unit is instructed to load new data.

[0081] The comparator compares the pointer with the address generated by the prefix sum unit. When the address is equal to the pointer, a switching signal is generated to instruct the index matching unit to prepare the next set of bitmap data and pointer input, and to instruct the processing unit to load new data. By generating the switching signal, the computing array can operate automatically and efficiently, improving flexibility and performance.

[0082] The training accelerator of the neural network is used to accelerate the forward propagation of standard convolution operations, the training process of depth convolution, the backward propagation of standard convolution operations, and the calculation of weight gradients. The forward propagation of standard convolution is a series of matrix multiplication and addition operations, and the accelerator can quickly process these large-scale matrix operations through its highly parallel architecture.

[0083] In the training of depth convolution, for very deep neural networks, the output of each layer is used as the input of the next layer, and the accelerator can effectively manage this inter-layer dependency and maintain high throughput.

[0084] In the backward propagation process, the accelerator can be parallelized through gradient calculation, which can reduce the calculation time. In addition to calculating gradients, the accelerator also optimizes the process of weight update.

[0085] Considering that depth convolution performs convolution operations separately on each input channel, and that the backpropagation stage involves the transpose requirements of the input and output channels of the weights, two data paths are configured in this embodiment. Among them, the controller is used to configure the data paths. The data paths include a column direction path, that is, a vertical direction path, and a row direction path, that is, a horizontal direction path. The column direction path indicates the direction of weight data transfer based on a column of processing units, and the row direction path indicates the direction of weight data transfer based on a row of processing units.

[0086] Among them, the controller is used to configure the data paths and coordinate the work of each component. The controller may include components such as an instruction register, a state machine, and a clock generator, which are used to store and execute control instructions, and control the flow of data and the working state of the processing units.

[0087] Since the training accelerator needs to execute various different types of computing tasks, the controller can match a suitable data path according to the type of computing task. For example, when performing standard convolution forward propagation, the column direction path is selected, and when performing depth convolution training, standard convolution backpropagation, and weight gradient calculation, the row direction path is selected.

[0088] The processing unit is also used to perform a multiplication operation based on the valid data obtained by the index matching unit, the input data stored in the input register, and the weight data, so as to output a feature map or gradient data. The weight data is the weight data passed by the previous processing unit in the data path where the processing unit is located. The feature map is a feature map generated based on the column direction path or the row direction path, and the gradient data is the gradient data or weight gradient value generated based on the row direction path.

[0089] The feature map is an intermediate result obtained after operations such as convolution in a neural network, which is used to characterize the feature information of the input data. The gradient data is the error gradient calculated during the backpropagation process, which is used to update the weights of the neural network.

[0090] The feature map is obtained by the processing unit performing multiplication operations and accumulation on the input data and the weight data. The gradient data is calculated according to the partial derivative of the loss function with respect to the weights. The gradient data starts from the output layer and backpropagates the error according to the chain rule to calculate the gradient of each weight.

[0091] For standard convolution forward propagation, in some embodiments, when the path of the weight data is configured as the column direction, the weight data is sequentially transferred within the processing unit and stored in the input register. The second numerical processing unit is used to perform an element-wise multiplication operation on the input data and the weight data based on the valid data, so as to output a first feature map. The first feature map is the feature map obtained through standard convolution forward propagation.

[0092] When the controller sets the data path to the column direction, the weight data is input from the bottommost processing unit in a column and passed upward until the topmost processing unit, ensuring the efficient transfer of weight data along the column direction. The input data and weight data are distributed to the corresponding processing units, and the processing units perform element-wise multiplication operations on the input data and weight data. The local output values are aggregated to form a feature map, i.e., the first feature map, which is the feature map obtained through standard convolution forward propagation.

[0093] During the transmission of weight data, the external memory first loads the weight data into the bottommost processing unit. Each processing unit receives the weight data passed from the previous processing unit and continues to pass it to the processing unit above until it reaches the topmost processing unit.

[0094] In some embodiments, the processing unit further includes a weight register. When the weight data reaches the topmost processing unit, it is stored in the weight register and keeps the weight data unchanged in the calculations involved in the processing unit.

[0095] For training deep convolution, since the weights of deep convolution are three-dimensional data, only the data for one clock cycle needs to be loaded. That is to say, within one clock cycle, the loading process of the convolution kernel weight data can be completed. Since there is a convolution kernel for each input channel in deep convolution, the convolution kernels can be loaded into the corresponding processing units in parallel within a single clock cycle, which can support efficient parallel processing and thus improve the computing efficiency.

[0096] The input data of different data channels is distributed and propagated to the processing units in different rows. Since deep convolution requires each input channel to perform convolution operations independently, the input data needs to be assigned to different processing units according to the channel it belongs to and stored in the input register.

[0097] Deep convolution is a special convolution operation in a convolutional neural network. A convolution kernel is applied separately to each input channel. Different from standard convolution, standard convolution uses the same convolution kernel to operate on all input channels and adds the results to produce the output, while deep convolution applies a specific convolution kernel to each input channel respectively. Therefore, each input channel generates an independent feature map.

[0098] In some embodiments, when the path of the weight data is configured to the row direction, the bitmap data is configured as a one-hot code with a second number of bits, and the value of the pointer is configured as 1. The processing unit is used to obtain the input data of multiple data channels from the input register. The value of the pointer is used to represent the number of valid channels. The processing unit is used to perform convolution operations on the input data of different data channels and the weight data to generate a second feature map.

[0099] During the training process of depth convolution, each convolution kernel only operates on one channel, and the computational complexity is relatively low. Therefore, only one channel needs to be marked as valid. One-hot encoding is a binary encoding method in which only one bit is 1 and the rest are 0. This means that only one input channel is considered valid at a time. The pointer value of 1 indicates that there is only one valid channel currently, which means that within the current cycle, only the processing units in this row will fetch the data of this channel from the input register for calculation.

[0100] Exemplarily, if there are 32 data channels, the one-hot bitmap [1, 0, 0,..., 0] indicates that the first channel is valid. Therefore, the bottommost processing unit will fetch the data of the first channel from the input register for calculation. Since the pointer value is 1, the processing units in each row will only process one different channel. For example, the processing unit in the first row processes the first channel, the processing unit in the second row processes the second channel, and so on.

[0101] For the backpropagation of standard convolution operations, in some embodiments, when the path configuration of the weight data is in the row direction, the bitmap data is configured as the binary bits of the second value, the value of the binary bit is 1, and the value of the pointer is configured as the second value. The processing unit is used to obtain the input data of the same data channel from the input register, the processing unit is used to obtain the first gradient data, and calculate the second gradient data based on the first gradient data and the weight data, and update the weight data through the second gradient data. The first gradient data is the data generated by the backpropagation of standard convolution.

[0102] Among them, the length of the bitmap data is 32 bits, such as [1, 1, 1,..., 1], indicating that all 32 channels are valid. The value of the pointer is the second value, that is, 32, indicating that there are 32 valid channels currently. The first gradient data is the error gradient data generated during the backpropagation of standard convolution and is used to update the weight data. The second gradient data is the weight gradient data calculated based on the first gradient data and the weight data.

[0103] For depth convolution training, the weight data is loaded and distributed to the processing units in different rows within one clock cycle. While during the backpropagation of standard convolution and weight gradient calculation, the weight data is gradually loaded in multiple clock cycles, and the data in each cycle will be distributed to the processing units in different rows. In this way, it can flexibly adapt to different types of task requirements and maximize the utilization rate of computing resources.

[0104] The controller configures the weight data path in the row direction, while configuring the bitmap data as 32 binary bits with a value of 1 and the pointer value as 32. The index matching unit determines the valid data channels from the input register according to the bitmap data and the pointer. The processing unit obtains the input data of the same data channel therefrom, and at the same time obtains the first gradient data. The processing unit calculates according to the backpropagation algorithm based on the obtained first gradient data and the weight data to obtain the second gradient data. The processing unit updates the weight data using the second gradient data. For example, using the gradient descent method, the weight data is subtracted by the result of the learning rate multiplied by the second gradient data.

[0105] It can effectively perform the backpropagation of standard convolution, accurately calculate the gradient of the error with respect to the weights, and update the weight data, which helps the neural network continuously adjust the parameters, improve the processing ability of the input data and the prediction accuracy, and enable the network to better fit the training data.

[0106] Since the parallelism of the output channels is configured as 32, and the calculation of the input channels is unfolded in the time dimension, the input data channels and depth convolution can be processed efficiently. However, when the total number of output channels of the network layer cannot be divisible by 32, it will cause the utilization rate of the processing unit to decrease. In addition, the input size of the deeper pointwise convolution layer is small and cannot be perfectly divided into 4×4 input blocks, which may cause the utilization rate of the processing unit to decrease to a certain extent.

[0107] To solve the above problems, as Figure 6 shown, in some embodiments, the training accelerator further includes a channel enhancement unit. The channel enhancement unit is used to obtain the number of output channels, set the number of input channel groups of the weight data based on the number of channels, and calculate the number of channels based on the number of channel groups. The channel enhancement unit is further used to configure the value of the pointer as a third value, and configure the bitmap data as the binary bits of the third value, and the value of the binary bit is 1.

[0108] When processing a network layer with an unconventional number of output channels, the channel enhancement unit can reorganize the reusable data in the input and weights from the input channel dimension to the output channel dimension. Specifically, as shown in Figure x, when the number of output channels M is less than 32, the utilization rate of the original computing unit is only M / 32. To make full use of the idle computing units, in this embodiment, the input channels of the weights are divided into 32 / d groups and mapped to the processing unit array. Each group contains p = d×P i / 32 channels, where d is the greatest common divisor of 32 and M, and P i is the parallelism of the input channels. Correspondingly, the pointer is configured as p, and the bitmap is configured as a sequence containing consecutive p 1s. By converting the time unfolding into spatial parallelism, the utilization rate of the computing unit is successfully increased to 100%, and no additional hardware resources are added.

[0109] When processing a deeper pointwise convolutional layer, the channel enhancement unit reorganizes the reusable data in the input and weights from the input channel dimension to the dimension of size (width and height). In some embodiments, the channel enhancement unit is also used to reorganize the weight data of the first size in the data channels of the first numerical value to generate an input block of the second size, and the processing unit is used to receive the input block of the second size to perform a convolutional operation through the input block to generate a third feature map.

[0110] A row of processing units processes 4×4 input blocks from the same input channel. The data enhancement unit reorganizes the 1×1 weights of 16 consecutive input channels into a 4×4 convolution. Since pointwise convolution involves the accumulation in the input channel dimension, when processing pointwise convolution, the problem of imperfect block division will no longer be faced, and the computing units can all be in an effective working state simultaneously.

[0111] The neural network training accelerator provided by the embodiments of the present application can process the calculations of multiple output channels in parallel. Through parallel processing and unfolding in the time dimension, the computing efficiency can be improved, and the hardware resources can be fully utilized. By introducing bitmaps and pointers and managing them through the corresponding index matching unit, the processing unit can only process the data of valid channels, reducing the processing overhead of invalid data, thereby improving the computing efficiency. It also utilizes the independence of element-wise multiplication in the Winograd algorithm to increase the parallelism of another dimension of the processing unit array, thereby promoting the efficient training of standard networks. It also solves the problem of mismatch between the number of output channels and the parallelism of the computing array and the imperfect block division problem in the Winograd algorithm through channel enhancement.

[0112] Moreover, it also supports various types of convolutional calculations, and also supports matrix multiplication and vector multiplication. In addition, it not only supports the training of dense standard neural networks and lightweight neural networks, but also supports the training of various deep neural networks with sparse channel dimensions without introducing load imbalance.

[0113] Based on the above neural network training accelerator, as Figure 7 shown, some embodiments of the present application also provide an acceleration method for a neural network training accelerator, including:

[0114] S100: Obtain the input data stored in the input register.

[0115] The input register includes multiple data channels, and the data channels are used to store input data;

[0116] S200: Control the index matching unit to obtain valid data.

[0117] The index matching unit corresponds to a row of processing units, the input register corresponds to a column of the processing units, and the processing units are arranged in rows and columns.

[0118] S300: Configure the data path.

[0119] The data path includes a column - direction path and a row - direction path. The column - direction path indicates the direction in which the weight data is transmitted based on a column of processing units, and the row - direction path indicates the direction in which the weight data is transmitted based on a row of processing units.

[0120] S400: Control the processing units to perform a multiplication operation based on the valid data obtained by the index matching unit, the input data stored in the input register, and the weight data, so as to output a feature map or gradient data.

[0121] The weight data is the weight data transmitted by the previous processing unit in the data path where the processing unit is located. The feature map is a feature map generated based on the column - direction path or the row - direction path, and the gradient data is gradient data generated based on the row - direction path.

[0122] For the effects achieved during the operation of the above - mentioned method embodiments, reference can be made to the effects of the above - mentioned accelerator embodiments, which will not be elaborated here.

[0123] Based on the above - mentioned neural network training accelerator, some embodiments of the present application further provide a neural network acceleration device, including an external memory and a neural network training accelerator.

[0124] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0125] For the similar parts between the embodiments provided in the present application, reference can be made to each other. The specific embodiments provided above are only several examples under the general concept of the present application and do not constitute a limitation on the protection scope of the present application. For those skilled in the art, any other implementation manners extended based on the solution of the present application without creative efforts belong to the protection scope of the present application.

Claims

1. A neural network training accelerator, characterized in that: include: An input register, comprising a plurality of data channels, wherein the data channels are used to store input data; A processing unit array, comprising a plurality of processing units arranged in rows and columns; Index matching units, each of the index matching units corresponds to a row of the processing units, and each of the input registers corresponds to a column of the processing units; The index matching unit is used to obtain valid data; A controller configured to configure a data path, the data path comprising a column direction path and a row direction path, the column direction path indicating that the weight data is transmitted based on a direction of a column of processing units, and the row direction path indicating that the weight data is transmitted based on a direction of a row of processing units; The processing unit is also used to perform a multiplication operation based on the valid data obtained by the index matching unit, the input data stored in the input register, and the weight data to output a feature map or gradient data, wherein the weight data is the weight data transmitted by the previous processing unit in the data path where the processing unit is located, and the feature map is a feature map generated based on the column-directional path or the row-directional path, and the gradient data is gradient data generated based on the row-directional path.

2. The neural network training accelerator according to claim 1, characterized in that: The index matching unit comprises: A sequence decoder, configured to receive bitmap data, retrieve an index of the data channel, and mark a valid index, wherein the bitmap data is used to determine the validity of the data channel, and the valid index is an index of a valid channel; A prefix sum unit, used to generate an address sequence based on the bitmap data, wherein the address sequence is a sequence calculated based on the sum of the number of channels of the valid indexes; A comparator is used to match the valid index and the address sequence to determine a valid address, where the valid address is a storage address of the valid data of the valid channel in the input register.

3. The neural network training accelerator according to claim 2, characterized in that: The number of columns of the processing unit is a first value, the number of rows of the processing unit is a second value, and the second value is an integer multiple of the first value; When the path of the weight data is configured as a column direction, the weight data is sequentially transmitted in the processing unit and stored in the input register; The processing unit of the second numerical value is used to perform an element-by-element multiplication operation on the input data and the weight data based on the valid data to output a first feature map, where the first feature map is a feature map obtained by standard convolution forward propagation.

4. The neural network training accelerator according to claim 2, characterized in that: The index matching unit is used to receive a pointer, and the pointer is used to record the number of valid channels; When the path of the weight data is configured as a row direction, the bitmap data is configured as a one-hot code of the second number of bits, and the value of the pointer is configured as 1, the processing unit is used to obtain input data of the plurality of data channels from the input register, and the value of the pointer is used to represent the number of the valid channels; The comparator is used to compare the effective address and the pointer; If the effective address is equal to the pointer, the comparator is further used to generate a switching signal, and the switching signal is used to control the index matching unit; The processing unit is used to perform a convolution operation on the input data and weight data of different data channels to generate a second feature map.

5. The neural network training accelerator according to claim 4, characterized in that: When the path of the weight data is configured as a row direction, the bitmap data is configured as a binary bit of a second value, the value of the binary bit is 1, and the value of the pointer is configured as the second value, and the processing unit is used to obtain input data of the same data channel from the input register; The comparator is used to compare the effective address and the pointer; If the effective address is equal to the pointer, the comparator is further used to generate a switching signal, and the switching signal is used to control the index matching unit; The processing unit is used to obtain first gradient data, calculate second gradient data based on the first gradient data and weight data, and update the weight data through the second gradient data, wherein the first gradient data is data generated by back propagation of standard convolution.

6. The neural network training accelerator according to claim 1, characterized in that: It also includes a weight register, and when the path of the weight data is configured in the column direction, the weight register is used to store the weight data.

7. The neural network training accelerator according to claim 1, characterized in that: Also included is a channel enhancement unit; The channel enhancement unit is used to obtain the number of output channels, set the number of input channel groups of the weight data based on the number of channels, and calculate the number of channels based on the number of channel groups; The channel enhancement unit is further used to configure the value of the pointer to be a third value, and to configure the bitmap data to be a binary bit of the third value, where the value of the binary bit is 1.

8. The neural network training accelerator according to claim 7, characterized in that: The channel enhancement unit is further used to reorganize the weight data of the first size in the data channel of the first value to generate an input block of the second size; The processing unit is used to receive an input block of the second size to perform a convolution operation on the input block to generate a third feature map.

9. A method for accelerating a neural network training accelerator, characterized in that: include: Acquire input data stored in an input register, wherein the input register includes a plurality of data channels, and the data channels are used to store the input data; Controlling the index matching unit to obtain valid data, the index matching unit corresponds to a row of processing units, the input register corresponds to a column of the processing units, and the processing units are arranged along the rows and columns; Configure a data path, the data path includes a column direction path and a row direction path, the column direction path indicates that the weight data is transmitted based on a direction of a column of processing units, and the row direction path indicates that the weight data is transmitted based on a direction of a row of processing units; Control the processing unit to perform a multiplication operation based on the valid data obtained by the index matching unit, the input data stored in the input register, and the weight data to output a feature map or gradient data, wherein the weight data is the weight data transmitted by the previous processing unit in the data path where the processing unit is located, the feature map is a feature map generated based on the column-direction path or the row-direction path, and the gradient data is the gradient data or weight gradient value generated based on the row-direction path.

10. A neural network acceleration device, characterized in that: The invention comprises an external memory and a neural network training accelerator as claimed in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Neural network processor based on systolic array

    CN107578098A

  • Neural network computing module, method and communication device

    CN113792868A

  • Neural network accelerator for dynamically matching non-zero values and oriented to unstructured sparseness

    CN116258188A