Converting neural network operation based on hardware configuration

By reshaping tensors to align with hardware configurations, the method optimizes resource utilization in DNN accelerators, reducing waste and enhancing performance and efficiency.

WO2026090973A1PCT designated stage Publication Date: 2026-05-07INTEL CORP +5
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INTEL CORP
Filing Date
2024-10-31
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current techniques for aligning neural network operations with hardware configurations often result in wasted computational and memory resources due to zero-padding, leading to impaired performance in DNN accelerators.

Method used

A reshaping method is applied to input tensors and kernels in neural network operations, adjusting the number of channels and spatial dimensions to align with hardware capabilities without changing the memory layout, thereby optimizing resource utilization.

Benefits of technology

This approach enhances hardware performance by minimizing unnecessary computations and conserving memory resources, improving efficiency and reducing data transfer requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128834_07052026_PF_FP_ABST
    Figure CN2024128834_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A convolution in a neural network may be converted to a new convolution for improving hardware performance. A kernel length of the convolution may be compared with a stride length of the convolution. The kernel length may be a length of the kernel of the convolution. The stride length may indicate a number of rows or columns by which the kernel moves across an input tensor of the convolution. The kernel length and stride length may be along the same spatial axis. A reshaping factor may be determined based on the comparison. The convolution may be converted by reshaping the input tensor or kernel based on the reshaping factor. The conversion of the convolution may also include padding the kernel or change of padding parameters or stride size of the original convolution. The neural network may be executed by carrying out the converted convolution in lieu of the original convolution.
Need to check novelty before this filing date? Find Prior Art

Description

CONVERTING NEURAL NETWORK OPERATION BASED ON HARDWARE CONFIGURATIONTechnical Field

[0001] This disclosure relates generally to neural networks (also referred to as “deep neural networks” or “DNN” ) , and more specifically, converting neural network operations (such as convolutions) in DNNs based on hardware configuration.Background

[0002] DNNs are used extensively for a variety of artificial intelligence (AI) applications ranging from computer vision to speech recognition and natural language processing due to their ability to achieve high accuracy. However, the high accuracy comes at the expense of significant computation cost. DNNs have extremely high computing demands as there can be a large number of operations as well as a large amount of data to read and write. Therefore, techniques to improve efficiency of DNNs are needed.Brief Description of the Drawings

[0003] Embodiments will be readily understood by the following detailed description in conjunction with the accompanying drawings. To facilitate this description, like reference numerals designate like structural elements. Embodiments are illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings.

[0004] FIG. 1 illustrates an example DNN, in accordance with various embodiments.

[0005] FIG. 2 illustrates an example convolution, in accordance with various embodiments.

[0006] FIG. 3 is a block diagram of a DNN system, in accordance with various embodiments.

[0007] FIG. 4 is a block diagram of a DNN module, in accordance with various embodiments.

[0008] FIGS. 5A-5C illustrate reshaping a tensor without changing the memory layout for storing the tensor, in accordance with various embodiments.

[0009] FIG. 6A illustrates an input tensor and kernel of a convolution, in accordance with various embodiments.

[0010] FIG. 6B illustrates converting the input tensor and kernel in FIG. 6A when a kernel length of the convolution equals a stride length of the convolution, in accordance with various embodiments.

[0011] FIG. 6C illustrates converting the input tensor and kernel in FIG. 6A when a kernel length of the convolution is smaller than a stride length of the convolution, in accordance with various embodiments.

[0012] FIG. 7A illustrates an input tensor, kernel, and output tensor of a convolution, in accordance with various embodiments.

[0013] FIG. 7B illustrates a converted input tensor generated by reshaping the input tensor in FIG. 7A for converting the convolution.

[0014] FIG. 7C illustrates a converted kernel generated by converting the kernel in FIG. 7A for converting the convolution, in accordance with various embodiments.

[0015] FIG. 7D illustrates an output tensor of the converted convolution, in accordance with various embodiments.

[0016] FIG. 8A illustrates an example kernel for a single output channel, in accordance with various embodiments.

[0017] FIGS. 8B-8E illustrate a kernel generated by converting the kernel in FIG. 8A, in accordance with various embodiments

[0018] FIG. 9A illustrates an example sparse cell, in accordance with various embodiments.

[0019] FIG. 9B illustrates an example sparse cell array, in accordance with various embodiments.

[0020] FIG. 10 illustrates an example processing element (PE) , in accordance with various embodiments.

[0021] FIG. 11 is a flowchart of a method of making an executable DNN, in accordance with various embodiments.

[0022] FIG. 12 is a block diagram of an example computing device, in accordance with various embodiments.Detailed Description

[0023] Overview

[0024] The last decade has witnessed a rapid rise in AI based data processing, particularly based on DNNs. DNNs are widely used in the domains of computer vision, speech recognition, image, and video processing mainly due to their ability to achieve beyond human-level accuracy. A DNN typically includes a sequence of layers. A DNN layer may  include one or more deep learning operations (also referred to as “neural network operations” ) , such as convolution, layer normalization, batch normalization, SoftMax operation, pooling, elementwise operation, linear operation, nonlinear operation, and so on.

[0025] Input or output data of neural network operations may be arranged in data structures called tensors. Taking a convolutional layer for example, the input tensors include an activation tensor (also referred to as “input feature map (IFM) ” or “input activation tensor” ) including one or more activations (also referred to as “input elements” ) and a weight tensor. The weight tensor may be a kernel (a2D weight tensor) , a filter (a3D weight tensor) , or a group of filters (a4D weight tensor) . A convolution may be performed on the input activation tensor and weight tensor to compute an output activation tensor in the convolutional layer.

[0026] A tensor is a data structure having multiple elements across one or more dimensions. Examples of tensors include vector (which is one-dimensional (1D) tensor) , matrix (which is two-dimensional (2D) tensor) , three-dimensional (3D) tensors, four-dimensional (4D) tensors, and even higher dimensional tensors. A dimension of a tensor may correspond to an axis, e.g., an axis in a coordinate system. A dimension may be measured by the number of data points along the axis. The dimensions of a tensor may define the shape of the tensor. A DNN layer may receive one or more input tensors and compute an output tensor from the one or more input tensors. In some embodiments, a 3D tensor may have an X-dimension, a Y-dimension, and Z-dimension. The X-dimension of a tensor may be the horizontal dimension, the length of which may be the width of the tensor; the Y-dimension may be the vertical dimension, the length of which may be the height of the tensor; and the Z-dimension may be the channel dimension, the length of which may be the number of channels. The coordinates of the elements along a dimension may be integers in an inclusive range from 0 to (L-1) , where L is the length of the tensor in the dimension. For instance, the x coordinate of the first element in a row may be 0, the x coordinate of the second element in a row may be 1, and so on. Similarly, the y coordinate of the first element in a column may be 0, the y coordinate of the second element in a column may be 1, and so on. A 4D tensor may have a fourth dimension, which may indicate the number of batches in the operation.

[0027] Tensors in DNNs can be saved in X-major (e.g., XYZ or XZY format) , Y-major formats (e.g., YXZ or YZX format) , or Z-major formats (e.g., ZXY or ZYX format) . The format of a tensor  may define the order in which the data points in the tensor are stored, written, or read. The first character may represent the dimension in which data points are contiguous in memory. The second character may represent the dimension in which data points can be accessed after the contiguous data points are accessed in memory. The third character may represent the dimension in which data points are accessed after the data points in the dimension represented by the second character are exhausted. Taking the ZXY format for example, the access order first starts in the Z-dimension, then moves to the X-dimension, and finally moves to the Y-dimension. Data points in the tensor are contiguous in memory in the Z-dimension, meaning data points having the same (x, y) coordinates are contiguous in memory. Using tensor permutation, the tensor may be read from memory in a different format.

[0028] The significant improvements in DNN model size and accuracy coupled with the rapid increase in computing power of execution platforms have led to the adoption of DNN applications even within resource constrained mobile and edge devices that have limited energy availability. DNN models may be executed, e.g., for training or inference, by DNN accelerators. A DNN accelerator may be or include one or more data processing units (DPUs) . A DPU may also be referred to as a compute block or compute tile. A DPU may include PEs that can carry out neural network operations.

[0029] In deep learning, such as within the structure of Convolutional Neural Networks (CNN) , the choice of the number of channels can significantly impact the performance and computational efficiency of the hardware devices (e.g., DNN accelerators) running the CNNs. Some deep learning frameworks and hardware-accelerating libraries might perform better when the number of channels matches a power of two. This may enhance the parallel computing abilities of the hardware device.

[0030] However, the channel dimension of many neural network operations does not align with the stipulations of the hardware. Many currently available techniques use expand solutions, in which zero-padding is used to expand the channel dimension of the input tensor and kernel for the purpose of aligning the channel dimension with the hardware configuration. Unfortunately, such expand solutions can lead to wasted resources and poor performance. In an example of a convolution whose input tensor size is 1×4×1080×2048 and kernel size is 4×4×4×4. The expand solution would expand the input tensor  to 1×16×1080×2048 and expand the kernel to 16×16×4×4.1 / 16 of the computations in the expanded convolution would produce the wanted output activations, and 15 / 16 of the computations would produce nothing meaningful, which can significantly waste the computational resources, memory resources, power, and so on. Therefore, the performance of the hardware device would be impaired.

[0031] Embodiments of the present disclosure may improve on at least some of the challenges and issues described above by converting neural network operations for improving hardware performance by reshaping tensors in a manner that can minimize the kernel size. In an example, the input tensor of a convolution may be reshaped without any change to the memory layout of the input activation. The kernel of the convolution may be redesigned to ensure that the correct output activations would be produced. The memory layout of the output activations, which are generated by performing the converted convolution, may also the same as if the original convolution is performed.

[0032] In various embodiments of the present disclosure, a reshaping factor may be determined for a convolution. The reshaping factor may be used to decrease a spatial length (e.g., width or height) of the input tensor and increase the number of input channels. The total number of input activations would remain the same despite the reshaping. The reshaping factor may be determined based on a kernel length and a stride length. The kernel length may be a length of the kernel of the convolution. The stride length may indicate a number of rows or columns by which the kernel moves across an input tensor of the convolution. The kernel length and stride length may be along the same spatial axis. The kernel length may be compared with the stride length. When the kernel length is greater than the stride length, a preliminary factor may be computed from the kernel length and the stride length. The reshaping factor may be determined by multiplying the preliminary factor with the stride length. When the kernel length is not greater than the stride length, the reshaping factor may equal the stride length. The kernel may be redesigned. For instance, the shape of the kernel may be changed. Additionally or alternatively, one or more zeros may be added to the kernel. The reshaped input tensor and redesigned kernel would be used for performing the converted convolution. The output tensor of the converted convolution may have a different shape from the output tensor of the original convolution, but the two output tensors include the same output activations. The memory layout of the  output activations may remain the same.

[0033] The present disclosure provides a reshape solution that is more advantageous than currently available expand solutions. The reshape solution in the present disclosure avoids the meaningless competition in the currently available expand solutions. Thus, the performance and efficiency of the hardware device can be improved. Also, as the memory layout of input activation and output activations can remain the same, data transfer is avoided and memory bandwidth can be saved.

[0034] For purposes of explanation, specific numbers, materials and configurations are set forth in order to provide a thorough understanding of the illustrative implementations. However, it will be apparent to one skilled in the art that the present disclosure may be practiced without the specific details or / and that the present disclosure may be practiced with only some of the described aspects. In other instances, well known features are omitted or simplified in order not to obscure the illustrative implementations.

[0035] Further, references are made to the accompanying drawings that form a part hereof, and in which is shown, by way of illustration, embodiments that may be practiced. It is to be understood that other embodiments may be utilized, and structural or logical changes may be made without departing from the scope of the present disclosure. Therefore, the following detailed description is not to be taken in a limiting sense.

[0036] Various operations may be described as multiple discrete actions or operations in turn, in a manner that is most helpful in understanding the claimed subject matter. However, the order of description should not be construed as to imply that these operations are necessarily order dependent. In particular, these operations may not be performed in the order of presentation. Operations described may be performed in a different order from the described embodiment. Various additional operations may be performed or described operations may be omitted in additional embodiments.

[0037] For the purposes of the present disclosure, the phrase “A or B” or the phrase "A and / or B" means (A) , (B) , or (A and B) . For the purposes of the present disclosure, the phrase “A, B, or C” or the phrase "A, B, and / or C" means (A) , (B) , (C) , (A and B) , (A and C) , (B and C) , or (A, B, and C) . The term "between, " when used with reference to measurement ranges, is inclusive of the ends of the measurement ranges.

[0038] The description uses the phrases "in an embodiment" or "in embodiments, " which  may each refer to one or more of the same or different embodiments. The terms "comprising, " "including, " "having, " and the like, as used with respect to embodiments of the present disclosure, are synonymous. The disclosure may use perspective-based descriptions such as "above, " "below, " "top, " "bottom, " and "side" to explain various features of the drawings, but these terms are simply for ease of discussion, and do not imply a desired or required orientation. The accompanying drawings are not necessarily drawn to scale. Unless otherwise specified, the use of the ordinal adjectives “first, ” “second, ” and “third, ” etc., to describe a common object, merely indicates that different instances of like objects are being referred to and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking or in any other manner.

[0039] In the following detailed description, various aspects of the illustrative implementations will be described using terms commonly employed by those skilled in the art to convey the substance of their work to others skilled in the art.

[0040] The terms “substantially, ” “close, ” “approximately, ” “near, ” and “about, ” generally refer to being within + / -20%of a target value as described herein or as known in the art. Similarly, terms indicating orientation of various elements, e.g., “coplanar, ” “perpendicular, ” “orthogonal, ” “parallel, ” or any other angle between the elements, generally refer to being within + / -5-20%of a target value as described herein or as known in the art.

[0041] In addition, the terms “comprise, ” “comprising, ” “include, ” “including, ” “have, ” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a method, process, device, or DNN accelerator that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such method, process, device, or DNN accelerators. Also, the term “or” refers to an inclusive “or” and not to an exclusive “or. ”

[0042] The systems, methods and devices of this disclosure each have several innovative aspects, no single one of which is solely responsible for all desirable attributes disclosed herein. Details of one or more implementations of the subject matter described in this specification are set forth in the description below and the accompanying drawings.

[0043] FIG. 1 illustrates an example DNN 100, in accordance with various embodiments. The DNN 100 may be executed by a DNN accelerator, e.g., the DNN accelerator 302 in FIG. 3. In an example, the DNN 100 may be a convolution-based DNN. In other examples, the DNN 100  may be other types of DNNs. For the purpose of illustration, the DNN 100 includes a sequence of layers comprising a plurality of convolutional layers 110 (individually referred to as “convolutional layer 110” ) , a plurality of pooling layers 120 (individually referred to as “pooling layer 120” ) , and a plurality of fully-connected layers 130 (individually referred to as “fully-connected layer 130” ) . In other embodiments, the DNN 100 may include fewer, more, or different layers. In an execution of the DNN 100, the layers of the DNN 100 execute tensor computation that includes many tensor operations, such as convolutions, pooling operations, elementwise operations (e.g., elementwise addition, elementwise multiplication, etc. ) , other types of tensor operations, or some combination thereof.

[0044] The convolutional layers 110 summarize the presence of features in inputs to the DNN 100. The convolutional layers 110 function as feature extractors. The first layer of the DNN 100 is a convolutional layer 110. In an example, a convolutional layer 110 performs a convolution on an input tensor 140 (also referred to as IFM 140) and a filter 150. As shown in FIG. 1, the IFM 140 is represented by a 7×7×3 three-dimensional (3D) matrix. The IFM 140 includes 3 input channels, each of which is represented by a 7×7 two-dimensional (2D) matrix. The 7×7 2D matrix includes 7 input elements (also referred to as input points) in each row and 7 input elements in each column. The filter 150 is represented by a 3×3×3 3D matrix. The filter 150 includes 3 kernels, each of which may correspond to a different input channel of the IFM 140. A kernel is a 2D matrix of weights, where the weights are arranged in columns and rows. A kernel can be smaller than the IFM. In the embodiments of FIG. 1, each kernel is represented by a 3×3 2D matrix. The 3×3 kernel includes 3 weights in each row and 3 weights in each column. Weights can be initialized and updated by backpropagation using gradient descent. The magnitudes of the weights can indicate importance of the filter 150 in extracting features from the IFM 140.

[0045] The convolution includes multiply-accumulate (MAC) operations with the input elements in the IFM 140 and the weights in the filter 150. The convolution may be a standard convolution 163 or a depthwise convolution 183. In the standard convolution 163, the whole filter 150 slides across the IFM 140. All the input channels are combined to produce an output tensor 160 (also referred to as output feature map (OFM) 160) . The OFM 160 is represented by a 5×5 2D matrix. The 5×5 2D matrix includes 5 output elements (also referred to as output points) in each row and 5 output elements in each column. For the  purpose of illustration, the standard convolution includes one filter in the embodiments of FIG. 1. In embodiments where there are multiple filters, the standard convolution may produce multiple output channels in the OFM 160.

[0046] The multiplication applied between a kernel-sized patch of the IFM 140 and a kernel may be a dot product. A dot product is the elementwise multiplication between the kernel-sized patch of the IFM 140 and the corresponding kernel, which is then summed, always resulting in a single value. Because it results in a single value, the operation is often referred to as the “scalar product. ” Using a kernel smaller than the IFM 140 is intentional as it allows the same kernel (set of weights) to be multiplied by the IFM 140 multiple times at different points on the IFM 140. Specifically, the kernel is applied systematically to each overlapping part or kernel-sized patch of the IFM 140, left to right, top to bottom. The result from multiplying the kernel with the IFM 140 one time is a single value. As the kernel is applied multiple times to the IFM 140, the multiplication result is a 2D matrix of output elements. As such, the 2D output matrix (i.e., the OFM 160) from the standard convolution 163 is referred to as an OFM.

[0047] In the depthwise convolution 183, the input channels are not combined. Rather, MAC operations are performed on an individual input channel and an individual kernel and produce an output channel. As shown in FIG. 1, the depthwise convolution 183 produces a depthwise output tensor 180. The depthwise output tensor 180 is represented by a 5×5×3 3D matrix. The depthwise output tensor 180 includes 3 output channels, each of which is represented by a 5×5 2D matrix. The 5×5 2D matrix includes 5 output elements in each row and 5 output elements in each column. Each output channel is a result of MAC operations of an input channel of the IFM 140 and a kernel of the filter 150. For instance, the first output channel (patterned with dots) is a result of MAC operations of the first input channel (patterned with dots) and the first kernel (patterned with dots) , the second output channel (patterned with horizontal strips) is a result of MAC operations of the second input channel (patterned with horizontal strips) and the second kernel (patterned with horizontal strips) , and the third output channel (patterned with diagonal stripes) is a result of MAC operations of the third input channel (patterned with diagonal stripes) and the third kernel (patterned with diagonal stripes) . In such a depthwise convolution, the number of input channels equals the number of output channels, and each output channel corresponds to a different input  channel. The input channels and output channels are referred to collectively as depthwise channels. After the depthwise convolution, a pointwise convolution 193 is then performed on the depthwise output tensor 180 and a 1×1×3 tensor 190 to produce the OFM 160.

[0048] The OFM 160 is then passed to the next layer in the sequence. In some embodiments, the OFM 160 is passed through an activation function. An example activation function is rectified linear unit (ReLU) . ReLU is a calculation that returns the value provided as input directly, or the value zero if the input is zero or less. The convolutional layer 110 may receive several images as input and calculate the convolution of each of them with each of the kernels. This process can be repeated several times. For instance, the OFM 160 is passed to the subsequent convolutional layer 110 (i.e., the convolutional layer 110 following the convolutional layer 110 generating the OFM 160 in the sequence) . The subsequent convolutional layers 110 perform a convolution on the OFM 160 with new kernels and generate a new feature map. The new feature map may also be normalized and resized. The new feature map can be kernelled again by a further subsequent convolutional layer 110, and so on.

[0049] In some embodiments, a convolutional layer 110 has four hyperparameters: the number of kernels, the size F kernels (e.g., a kernel is of dimensions F×F×D pixels) , the S step with which the window corresponding to the kernel is dragged on the image (e.g., a step of one means moving the window one pixel at a time) , and the zero-padding P (e.g., adding a black contour of P pixels thickness to the input image of the convolutional layer 110) . The convolutional layers 110 may perform various types of convolutions, such as 2-dimensional convolution, dilated or atrous convolution, spatial separable convolution, depthwise separable convolution, transposed convolution, and so on. The DNN 100 includes 16 convolutional layers 110. In other embodiments, the DNN 100 may include a different number of convolutional layers.

[0050] The pooling layers 120 down-sample feature maps generated by the convolutional layers, e.g., by summarizing the presence of features in the patches of the feature maps. A pooling layer 120 is placed between two convolution layers 110: a preceding convolutional layer 110 (the convolution layer 110 preceding the pooling layer 120 in the sequence of layers) and a subsequent convolutional layer 110 (the convolution layer 110 subsequent to the pooling layer 120 in the sequence of layers) . In some embodiments, a pooling layer 120  is added after a convolutional layer 110, e.g., after an activation function (e.g., ReLU, etc. ) has been applied to the OFM 160.

[0051] A pooling layer 120 receives feature maps generated by the preceding convolution layer 110 and applies a pooling operation to the feature maps. The pooling operation reduces the size of the feature maps while preserving their important characteristics. Accordingly, the pooling operation improves the efficiency of the DNN and avoids over-learning. The pooling layers 120 may perform the pooling operation through average pooling (calculating the average value for each patch on the feature map) , max pooling (calculating the maximum value for each patch of the feature map) , or a combination of both. The size of the pooling operation is smaller than the size of the feature maps. In various embodiments, the pooling operation is 2×2 pixels applied with a stride of two pixels, so that the pooling operation reduces the size of a feature map by a factor of 2, e.g., the number of pixels or values in the feature map is reduced to one quarter the size. In an example, a pooling layer 120 applied to a feature map of 6×6 results in an output pooled feature map of 3×3. The output of the pooling layer 120 is inputted into the subsequent convolution layer 110 for further feature extraction. In some embodiments, the pooling layer 120 operates upon each feature map separately to create a new set of the same number of pooled feature maps.

[0052] The fully-connected layers 130 are the last layers of the DNN. The fully-connected layers 130 may be convolutional or not. The fully-connected layers 130 receive an input operand. The input operand defines the output of the convolutional layers 110 and pooling layers 120 and includes the values of the last feature map generated by the last pooling layer 120 in the sequence. The fully-connected layers 130 apply a linear combination and an activation function to the input operand and generate a vector. The vector may contain as many elements as there are classes: element i represents the probability that the image belongs to class i. Each element is therefore between 0 and 1, and the sum of all is worth one. These probabilities are calculated by the last fully-connected layer 130 by using a logistic function (binary classification) or a SoftMax function (multi-class classification) as an activation function. In some embodiments, the fully-connected layers 130 multiply each input element by weight, make the sum, and then apply an activation function (e.g., logistic if N=2, SoftMax if N>2) . This is equivalent to multiplying the input operand by the matrix containing the weights.

[0053] FIG. 2 illustrates an example convolution, in accordance with various embodiments. The convolution may be a deep learning operation in a convolutional layer of a DNN, e.g., a convolutional layer 110 in FIG. 1. The convolution can be executed on an activation tensor 210 and filters 220 (individually referred to as “filter 220” ) . The filters may constitute a weight tensor of the convolution. The result of the convolution is an output tensor 230. In some embodiments, the convolution is performed by a DNN accelerator. An example of the DNN accelerator may be the DNN accelerator 302 in FIG. 3. For instance, the convolution may be performed by one or more DPUs 330 in the DNN accelerator 302.

[0054] The activation tensor 210 may be computed in a previous layer of the DNN. In some embodiments (e.g., embodiments where the convolutional layer is the first layer of the DNN) , the activation tensor 210 may be an image. In the embodiments of FIG. 2, the activation tensor 210 includes activations (also referred to as “input activations, ” “elements, ” or “input elements” ) arranged in a 3D matrix. The activation tensor 210 may also be referred to as an input tensor of the convolution. An input element is a data point in the activation tensor 210. The activation tensor 210 has a spatial size Hin×Win×Cin, where Hin is the height of the 3D matrix (i.e., the length along the Y axis, which indicates the number of activations in a column in the 3D matrix of each input channel) , Win is the width of the 3D matrix (i.e., the length along the X axis, which indicates the number of activations in a row in the 2D matrix of each input channel) , and Cin is the depth of the 3D matrix (i.e., the length along the Z axis, which indicates the number of input channels) . For the purpose of simplicity and illustration, the activation tensor 210 has a spatial size of 7×7×3, i.e., the activation tensor 210 includes three input channels and each input channel has a 7×7 2D matrix. Each input element in the activation tensor 210 may be represented by a (X, Y, Z) coordinate. In other embodiments, the height, width, or depth of the activation tensor 210 may be different.

[0055] Each filter 220 includes weights arranged in a 3D matrix. The values of the weights may be determined through training the DNN. A filter 220 has a spatial size Hf×Wf×Cf, where Hf is the height of the filter (i.e., the length along the Y axis, which indicates the number of weights in a column in each kernel) , Wf is the width of the filter (i.e., the length along the X axis, which indicates the number of weights in a row in each kernel) , and Cf is the depth of the filter (i.e., the length along the Z axis, which indicates the number of  channels) . In some embodiments, Cf equals Cin. For purpose of simplicity and illustration, each filter 220 in FIG. 2 has a spatial size of 2×3×3, i.e., the filter 220 includes 2 convolutional kernels with a spatial size of 2×3. In other embodiments, the height, width, or depth of the filter 220 may be different. The spatial size of the convolutional kernels is smaller than the spatial size of the 2D matrix of each input channel in the activation tensor 210.

[0056] An activation or weight may take one or more bytes in a memory. The number of bytes for an activation or weight may depend on the data format. For example, when the activation or weight has an INT8 format, the activation takes one byte. When the activation or weight has a FP16 format, the activation or weight takes two bytes. Other data formats may be used for activations or weights.

[0057] In the convolution, each filter 220 slides across the activation tensor 210 and generates a 2D matrix for an output channel in the output tensor 230. In the embodiments of FIG. 2, the 2D matrix has a spatial size of 5×5. The output tensor 230 includes activations (also referred to as “output activations, ” “elements, ” or “output element” ) arranged in a 3D matrix. An output activation is a data point in the output tensor 230. The output tensor 230 has a spatial size Hout×Wout×Cout, where Hout is the height of the 3D matrix (i.e., the length along the Y axis, which indicates the number of output activations in a column in the 2D matrix of each output channel) , Wout is the width of the 3D matrix (i.e., the length along the X axis, which indicates the number of output activations in a row in the 2D matrix of each output channel) , and Cout is the depth of the 3D matrix (i.e., the length along the Z axis, which indicates the number of output channels) . Cout may equal the number of filters 220 in the convolution. Hout and Wout may depend on the heights and weights of the activation tensor 210 and each filter 220. In an example where the kernel size is 1×1, Hout and Woutmay equal to Hin and Win, respectively.

[0058] As a part of the convolution, MAC operations can be performed on a 2×3×3 subtensor 215 (which is highlighted with a dotted pattern in FIG. 2) in the activation tensor 210 and each filter 220. The result of the MAC operations on the subtensor 215 and one filter 220 is an output activation. In some embodiments (e.g., embodiments where the convolution is an integral convolution) , an output activation may include 8 bits, e.g., one byte. In other embodiments (e.g., embodiments where the convolution is a floating-point convolution) , an output activation may include more than one byte. For instance, an output  element may include two bytes.

[0059] After the MAC operations on the subtensor 215 and all the filters 220 are finished, a vector 235 is produced. The vector 235 is highlighted with a dotted pattern in FIG. 2. The vector 235 includes a sequence of output activations, which are arranged along the Z axis. The output activations in the vector 235 have the same (x, y) coordinate, but the output activations correspond to different output channels and have different Z coordinates. The dimension of the vector 235 along the Z axis may equal the total number of output channels in the output tensor 230. After the vector 235 is produced, further MAC operations are performed to produce additional vectors till the output tensor 230 is produced. In the embodiments of FIG. 2, the output tensor 230 is computed in a Z-major format. When the output tensor 230 is computed in the ZXY format, the vector that is adjacent to the vector 235 along the X axis may be computed right after the vector 235. When the output tensor 230 is computed in the ZYX format, the vector that is adjacent to the vector 235 along the Y axis may be computed right after the vector 235. The output tensor 230 may be permuted, e.g., by the drain module 390, and stored in a memory (e.g., the local memory 340) in an X-major format or Y-major format.

[0060] In some embodiments, the MAC operations on a 3×3×3 subtensor (e.g., the subtensor 215) and a filter 220 may be performed by a plurality of MAC units. One or more MAC units may receive an input operand (e.g., an activation operand 217 shown in FIG. 2) and a weight operand (e.g., the weight operand 227 shown in FIG. 2) . The activation operand 217 includes a sequence of activations having the same (x, y) coordinate but different z coordinates. The activation operand 217 includes an activation from each of the input channels in the activation tensor 210. The weight operand 227 includes a sequence of weights having the same (x, y) coordinate but different z coordinates. The weight operand 227 includes a weight from each of the channels in the filter 220. Activations in the activation operand 217 and weights in the weight operand 227 may be sequentially fed into a MAC unit. The MAC unit may receive an activation and a weight ( “an activation-weight pair” ) at a time and multiple the activation and the weight. The position of the activation in the activation operand 217 may match the position of the weight in the weight operand 227. The activation and weight may correspond to the same channel.

[0061] Activations or weights may be floating-point numbers. Floating-point numbers may  have various data formats, such as FP32, FP16, BF16, and so on. A floating-point number may be a positive or negative number with a decimal point. A floating-point number may be represented by a sequence of bits that includes one or more bits representing the sign of the floating-point number (e.g., positive or negative) , bits representing an exponent of the floating-point number, and bits representing a mantissa of the floating-point number. The mantissa is the part of a floating-point number that represents the significant digits of that number. The mantissa is multiplied by the base raised to the exponent to give the actual value of the floating-point number.

[0062] In some embodiments, the output activations in the output tensor 230 may be further processed based on one or more activation functions before they are written into the memory or inputted into the next layer of the DNN. The processing based on the one or more activation functions may be at least part of the post processing of the convolution. In some embodiments, the post processing may include one or more other computations, such as offset computation, bias computation, and so on. The results of the post processing may be stored in a local memory of the compute block and be used as input to the next DNN layer. In some embodiments, the input activations in the activation tensor 210 may be results of post processing of the previous DNN layer.

[0063] FIG. 3 is a block diagram of a DNN system 300, in accordance with various embodiments. The whole DNN system 300 or a part of the DNN system 300 may be implemented in one or more computing devices, such as the computing device 2000 in FIG. 12. The DNN system 300 can generate and execute DNNs. As shown in FIG. 3, the DNN system 300 includes a DNN module 301 and a DNN accelerator 302. In other embodiments, alternative configurations, different or additional components may be included in the DNN system 300. For instance, the DNN system 300 may include multiple DNN modules or multiple DNN accelerators. Further, functionality attributed to a component of the DNN system 300 may be accomplished by a different component included in the DNN system 300 or a different system. In some embodiments, the DNN module 301 and DNN accelerator 302 may include different types of processing units. In an example, the DNN module 301 may be implemented by one or more central processing units (CPUs) . The DNN accelerator 302 may also be referred to as an AI accelerator or an AI processor. The DNN module 301 and DNN accelerator 302 may be implemented in the same chip or separate chips.

[0064] The DNN module 301 facilitates generation and deployment of DNNs. In some embodiments, the DNN module 301 may generate and train DNNs. For instance, the DNN module 301 can define the layered architecture of a DNN. The DNN module 301 can also determine the internal parameters of the DNN through a DNN training process. The DNN module 301 may also determine one or more hyperparameters that define how the DNN is trained. An example hyperparameter is a sparsity ratio that defines the sparsity level of one or more deep learning tensors for the DNN. The DNN module 301 may also compress DNNs, e.g., during or after training. In some embodiments, the DNN module 301 may prune weights in one or more layers of a DNN by changing nonzero valued weight to zeros. The DNN module 301 may prune weights based on a target weight sparsity ratio. A weight sparsity ratio may be the ratio of the number of zero-valued weights to the total number of weights. In an example where the DNN module 301 prunes weight during DNN training, the DNN module 301 may prune weight of a layer to achieve a target sparsity ratio after one or more epochs. The DNN module 301 may prevent the pruned weights from changing values during the rest of the training process. Alternatively, the DNN module 301 may allow the pruned weights to change values so that a pruned, zero-valued weight may have a nonzero value after further training. The DNN module 301 may prune weights of the layer again after one or more additional epochs.

[0065] The DNN module 301 may deploy trained, compressed, or validated DNNs for use in neural network applications. In some embodiments, the DNN module 301 may distribute trained, compressed, or validated DNNs to devices or systems which may use the DNNs to perform tasks (e.g., image classification, motion planning, etc. ) for which the DNNs were trained. In other embodiments, the DNN module 301 may facilitate deployment of the DNNs using the DNN accelerator 302. For instance, the DNN module 301 may receive data from a device or system coupled with the DNN system 300 and input the received data (or data generated by the DNN module 301, e.g., based on the received data) into a DNN. The DNN module 301 may generate instructions (e.g., computer program instructions) that can be executed by the DNN accelerator 302 for DNN execution. The DNN module 301 may receive an output of the DNN from the DNN accelerator 302. The DNN module 301 may transmit the output of the DNN (or a result of processing the output of the DNN by the DNN module 301) to the device or system. In some embodiments, the DNN module 301 may control execution  processes of trained, compressed, or validated DNNs. The DNN module 301 may function as a complier for DNNs executed by the DNN accelerator 302. The DNN module 301 may perform compilation of DNNs and generate compilation descriptors, based on which the DNNs may be executed.

[0066] The DNN module 301 facilitates converting neural network operations to optimize or improve the performance of the DNN accelerator 302. For instance, the DNN module 301 may reshape the input tensor or kernel of a convolution to change the number of input channels to a power of two to optimize parallel computation and efficiency of the hardware device. The DNN module 301 may also pad the input tensor or kernel to convert the convolution. The DNN module 301 may may then generate instructions that can be executed by one or more components of the DNN accelerator 302 (e.g., the direct memory access (DMA) engine 320) for carrying out the converted convolution in lieu of the original convolution. Certain aspects of the DNN module 301 are provided below in conjunction with FIG. 4.

[0067] The DNN accelerator 302 executes DNNs provided by the DNN module 301. For instance, the DNN accelerator 302 can execute a DNN by carrying out neural network operations in the DNN. The process of carrying out a neural network operation is also referred to as a process of executing the neural network operation or performing the neural network operation. The execution of the DNN may be for training the DNN or for using the DNN to perform AI tasks. As shown in FIG. 3, the DNN accelerator 302 includes a memory 310, a DMA engine 320, and DPUs 330 (individually referred to as “DPU 330” ) . In other embodiments, alternative configurations, different or additional components may be included in the DNN accelerator 302. For example, the DNN accelerator 302 may include more than one memory 310 or DMA engine 320. As another example, the DNN accelerator 302 may include a single DPU 330. Further, functionality attributed to a component of the DNN accelerator 302 may be accomplished by a different component included in the DNN accelerator 302 or by a different system. A component of the DNN accelerator 302 may be implemented in hardware, software, firmware, or some combination thereof.

[0068] The memory 310 stores data associated with neural network operations performed by the DNN accelerator 302. In some embodiments, the memory 310 may store data to be used by the DPUs 330 for executing neural network operations. The memory 310 may store  IFMs. The memory 310 may also store weights, such as weights in kernels of convolutions, which are determined by training DNNs. The memory 310 may further store outputs of neural network operations, such as OFMs. In some embodiments, the memory 310 includes one or more dynamic random-access memories (DRAMs) .

[0069] The DMA engine 320 facilitates data transfer between the memory 310 and local memories of the DPUs 330. For example, the DMA engine 320 can read data from the memory 310 and write data into a local memory of a DPU 330. As another example, the DMA engine 320 can read data from a local memory of a DPU 330 and write data into the memory 310. For instance, the DMA engine 320 may read data from the memory 310 and load the data to one or more DPUs 330. The DMA engine 320 may also write data computed by one or more DPUs 330 to the memory 310. The DMA engine 320 provides a DMA feature that allows the DPU 330 to initiate data transfer between the memory 310 and the local memories of the DPUs 330 and to perform other operations while the data transfer is being conducted. In some embodiments, the DMA engine 320 may read tensors from the memory 310, modify the tensors in a way that is optimized for the DPU 330 before it writes the tensors into the local memories of the DPUs 330.

[0070] The DPUs 330 perform neural network operations in DNNs. For instance, a DPU 330 may execute a DNN layer by running one or more deep learning operations in the DNN layer. A DPU 330 may execute a layer, or a portion of a layer, at a time. In some embodiments, the operations of the DNN layers may be run by multiple DPUs 330 in parallel. For instance, multiple DPUs 330 may each perform a portion of a workload for a neural network operation. Data may be shared between the DPUs 330. A DPU 330 may also be referred to as a neural processing unit, a compute block, or a compute tile.

[0071] The DPUs 330 may be capable of running various types of neural network operations, such as convolution (including depthwise convolutions) , layer normalization, SoftMax operation, pooling, elementwise operation, linear operation, nonlinear operation, and so on. N=Neural network operations performed by the DPUs 330 include tensor operations, i.e., operations whose inputs are tensors or operations whose outputs are tensors. In an example, the DPU 330 receives an input tensor and one or more convolutional kernels and performs a convolution with the input tensor and convolutional kernels. The result of the convolution may be an output tensor, which can be further computed, e.g., by the DPU 330  or another DPU 330.

[0072] In the embodiments of FIG. 3, each DPU 330 includes a local memory 340, a sparsity mode module 350, a load module 360, a processing engine 370, a post-processing engine 380, and a drain module 390. Some or all the components of the DPU 330 can be implemented on the same chip. In other embodiments, alternative configurations, different or additional components may be included in the DPU 330. Further, functionality attributed to a component of the DPU 330 may be accomplished by a different component included in the DPU 330, a different DPU 330, another component of the DNN accelerator 302, or a different system. A component of the DPU 330 may be implemented in hardware, software, firmware, or some combination thereof.

[0073] The local memory 340 is local to the corresponding DPU 330. In the embodiments of FIG. 3, the local memory 340 is inside the DPU 330. In other embodiments, the local memory 340 may be outside the DPU 330. Data in the local memory 340 may be transferred to or from the memory 310, e.g., through the DMA engine 320. In some embodiments, data in the local memory 340 may be transferred to or from the local memory of another DPU 330. The local memory 340 may store data received, used, or generated by the sparsity mode module 350, the load module 360, the processing engine 370, the post-processing engine 380, or the drain module 390. Examples of the data may include input activations, weights, output activations, sparsity bitmaps, and so on.

[0074] In some embodiments, the local memory 340 may store tensors to be processed by the processing engine 370 or the post-processing engine 380. The tensors may be input tensors of deep learning operations. The local memory 340 may also store tensors generated by the processing engine 370 or the post-processing engine 380. The tensors may be output tensors of deep learning operations. The layout of data points of a tensor in the local memory 340 may depend on the format in which the tensor is stored. In some embodiments, the local memory 340 may store tensors in various formats, including Z-major (e.g., ZXY or ZYX) format, X-major (e.g., XYZ or XZY) format, and Y-major (e.g., YXZ or YZX) format. For a tensor with Z-major format, the local memory 340 may store data points having the same (x, y) coordinate contiguously. For instance, the data points having the same (x, y) coordinate may be stored at a sequence of memory addresses in the local memory 340. For a tensor with the ZXY format or ZYX format, the local memory 340 may  store data points having the same (x, y) coordinate contiguously. For instance, the data points having the same (x, y) coordinate may be stored at a sequence of memory addresses in the local memory 340. For a tensor with X-major format, the local memory 340 may store data points having the same (y, z) coordinate contiguously. For a tensor with Y-major format, the local memory 340 may store data points having the same (x, z) coordinate contiguously.

[0075] In some embodiments, the local memory 340 may store dense tensors (e.g., dense activation tensors, dense weight tensors, etc. ) , sparse tensors (e.g., sparse activation tensors, sparse weight tensors, etc. ) , and so on. A dense tensor may be a tensor from which zero-valued elements (if any) are not removed. A dense tensor may be converted to a sparse tensor by removing one or more zero-valued elements in the dense tensor. A sparse tensor may also be referred to as a compressed tensor or packed tensor. The process of converting a dense tensor to a sparse tensor may be referred to as sparsity encoding. Sparsity encoding may also generate a sparsity tensor. Each element in the sparsity tensor may correspond to a different element in the dense tensor and indicate whether the element in the dense tensor is zero or not. The sparsity tensor may indicate positions of elements of the sparse tensor in the dense tensor. The sparsity tensor may be a sparsity bitmap, each element of which is a bit. A sparse tensor may be converted to a dense tensor through a densifying process, in which one or more zeros may be added to the sparse tensor based on the sparsity tensor.

[0076] In some embodiments, the local memory 340 includes one or more static random-access memories (SRAMs) . The local memory 340 may be byte-addressable, and each memory address identifies a single byte (eight bits) of storage. In some embodiments, the local memory 340 may include memory banks. The number of data banks in the local memory 340 may be 16, 64, 128, 356, 512, 1024, 2048, or other numbers. A memory bank may include a plurality of storage units. In an example, a data bank may include 8, 16, 64, or a different number of storage units. A memory bank or a storage unit in a memory bank may have a memory address. In an example, a storage unit may store a single byte, and data larger than a single byte may be stored in storage units with consecutive memory addresses, i.e., adjacent storage units. For instance, a storage unit can store an integer number in the INT8 format, versus two storage units may be needed to store a number in the FP16 or BF16 format, which has 16 bits. In some embodiments, 16 bits can be transferred from the local  memory 340 in a single read cycle. In other embodiments, 16 bits can be transferred from the local memory 340 in multiple read cycles, such as two cycles.

[0077] The sparsity mode module 350 determines sparsity modes in which the DPU 330 operates to execute DNN layers. For instance, the sparsity mode module 350 may determine whether to accelerate a layer based on weight sparsity, activation sparsity, or both. The sparsity mode module 350 select the sparsity mode for a layer from a group of sparsity modes that includes, for example, combined sparsity mode in which the layer is accelerated based on both weight sparsity and activation sparsity, activation sparsity mode in which the layer is accelerated based on activation sparsity but not based on weight sparsity, weight sparsity mode in which the layer is accelerated based on weight sparsity but not based on activation sparsity, and a dense mode in which the layer is not accelerated based on sparsity. In some embodiments (e.g., embodiments where a layer is executed by multiple DPUs 330) , the sparsity mode module 350 may determine the sparsity mode for all the DPUs 330 that executes the layer. In some embodiments, the sparsity mode module 350 may receive configuration parameters from the DNN module 301. A configuration parameter may correspond to a layer and indicate whether to accelerate the layer based on weight sparsity. The sparsity mode module 350 may determine the sparsity mode of the layer based on the configuration parameter.

[0078] The load module 360 loads data from the local memory 340 to the processing engine 370 or to the post-processing engine 380. The load module 360 may read tensors from the local memory 340. The tensors may include sparse activation tensors, sparse weight tensors, activation sparsity tensors, weight sparsity tensors, and so on. In some embodiments, the load module 360 may load data based on the sparsity mode determined by the sparsity mode module 350. The load module 360 may select different data to transmit to the processing engine 370 in different sparsity modes. For instance, the load module 360 may transmit an activation sparsity tensor and a weight sparsity tensor of a layer to the processing engine 370 in the combined sparsity mode, while transmit the activation sparsity tensor but not the weight sparsity tensor to the processing engine 370 in the activation sparsity mode and transmit the weight sparsity tensor but not the activation sparsity tensor to the processing engine 370 in the weight sparsity mode. In the dense mode, the load module 360 does not transmit either the activation sparsity tensor or the weight sparsity  tensor to the processing engine 370.

[0079] In some embodiments, the load module 360 may process (e.g., densify) data stored in the local memory 340 before providing the data to the processing engine 370. In an example, the load module 360, while operating in the weight sparsity mode, may densify sparse activation tensors to generate dense activation tensors based on corresponding activation sparsity tensors. For instance, the load module 360 may add one or more zeros into a sparse activation tensor based on an activation sparsity tensor associated with the sparse activation tensor to generate the dense activation tensor. The dense activation tensor includes one or more elements than the sparse activation tensor. The additional element (s) are zero-valued. The load module 360 may identify one or more elements in the activation sparsity tensor that correspond to the zero-valued element (s) , determine the position of each of the zero-valued element (s) in the dense activation tensor, and insert the zero-valued element (s) into the sparse activation tensor based on the determined positions. After the densification, the load module 360 may transmit the dense activation tensors to the processing engine 370. The load module 360 may also transmit corresponding sparse weight tensors and weight sparsity tensors to the processing engine 370. Activation sparsity tensor of the dense activation tensors may not be loaded to the processing engine 370.

[0080] In another example, the load module 360, while operating in the activation sparsity mode, may densify sparse weight tensors to generate dense weight tensors based on corresponding weight sparsity tensors by inserting zeros into sparse weight tensors. The densification of sparse weight tensors may be similar to the densification of sparse activation tensors described above. After the densification, the load module 360 may transmit the dense weight tensors to the processing engine 370. The load module 360 may also transmit corresponding sparse activation tensors and activation sparsity tensors to the processing engine 370. Weight sparsity tensor of the dense weight tensors may not be loaded to the processing engine 370.

[0081] In yet another example, the load module 360, while operating in the dense mode, may densify both sparse weight tensors and sparse activation tensors. The load module 360 may generate the input tensor and weight tensor of the layer and transmit the tensors to the processing engine 370 for executing the layer without sparsity acceleration.

[0082] The processing engine 370 performs operations in DNNs. The processing engine 370  may accelerate neural network operations based on sparsity in data. In some embodiments, the processing engine 370 may operate in a dense mode in which sparsity acceleration is not performed. The processing engine 370 may include one or more processing cells. In some embodiments, the processing cells may be arranged in one or more rows and one or more columns in the processing engine 370. Each processing cell may include PEs that may be arranged in an array that includes rows and columns. All the PEs in the processing engine 370 may constitute a bigger array that includes more rows and columns.

[0083] An example PE may be or may include one or more MAC units that can perform MAC operations. In some embodiments (e.g., embodiments where the DPU 330 executes a convolutional layer) , a computation in an MAC unit may be an MAC operation on an activation operand and a weight operand. The activation operand may be an activation tensor that may include one or more activations in the input tensor of the convolution. Different activations may be in different input channels. The weight operand may be a weight tensor that may include one or more weights in the filter of the convolution. The values of the weights are determined through training the DNN. The weights in the weight operand may be in different input channels.

[0084] In some embodiments, an MAC unit includes one or more multipliers for performing multiplications. An MAC unit may also include one or more accumulators ( “adders” ) for performing accumulations. A column of MAC units is referred to as an MAC column. An MAC column may be associated with one or more MAC lanes. A MAC lane is a path for loading data e.g., by the load module 360, into an MAC column. A MAC lane may be also referred to as a data transmission lane or data loading lane. An MAC column may have multiple MAC lanes. The loading bandwidth of the MAC column is an aggregation of the loading bandwidths of all the MAC lanes associated with the MAC column. With a certain number of MAC lanes, data can be fed into the same number of independent MAC units simultaneously. In some embodiments where an MAC column has four MAC lanes for feeding activations or weights into the MAC column and each MAC lane may have a bandwidth of 16 bytes, the four MAC lanes can have a total loading bandwidth of 64 bytes.

[0085] In some embodiments, the processing engine 370 may be capable of depthwise convolution, standard convolution, or both. In a depthwise convolution, an MAC unit may perform an MAC operation that includes a sequence of multiplications for an input operand  and a weight operand. Each multiplication in the sequence (also referred to as a cycle) is a multiplication of a different activation in the input operand with a different weight in the weight operand. The activation and weight in the same cycle may correspond to the same channel. The sequence of multiplication produces a product operand that includes a sequence of products. The MAC operation may also include accumulations in which multiple product operands are accumulated to produce an output operand of the MAC unit. The processing engine 370 may output multiple output operands at a time, each of which is generated by a different MAC unit. In a standard convolution, MAC operations may include accumulations across the channels. For instance, as opposed to generating an output operand, a MAC unit may accumulate products across different channels to generate a single output point.

[0086] In some embodiments, the processing engine 370 may perform MAC operations in quantized deep learning operations, such as MAC operations in a quantized convolution. In some embodiments, an MAC unit in the processing engine 370 may receive quantized activation and quantized weights and compute a quantized MAC result. The quantized MAC result may be a quantized value in an integer format and may be the output of the MAC unit. In some embodiments, the MAC unit may also include a quantization multiplier that can multiply a quantization scale with the quantized MAC result, and the output of the MAC unit may be a real value in a floating-point format. The MAC unit may include no quantization subtractors as zero-point offsetting is not needed for the MAC operations in quantized deep learning operations.

[0087] In some embodiments, the processing engine 370 may include sparsity acceleration logic for facilitating sparsity acceleration. For instance, each processing cell in the processing engine 370 may include one or more sparsity modules. In an example, each MAC column or each MAC row may have a corresponding sparsity module that accelerates MAC operations in the MAC column or MAC row. In some embodiments, a sparsity module accelerates computations in the processing engine 370 based on sparsity in activations, sparsity in weights, or both. The sparsity module may include a storage unit that stores a sparsity tensor, which may be loaded to the storage unit by the load module 360. The sparsity tensor may be an activation sparsity tensor, a weight sparsity tensor, or a combined sparsity tensor.

[0088] An activation sparsity tensor may be the sparsity tensor of an activation tensor and  has the same number of elements as the activation tensor. An element in the activation sparsity tensor may indicate whether the corresponding element in the activation tensor is zero or not. For instance, a zero-valued in the activation sparsity tensor may indicate that the corresponding element in the activation tensor is zero. A one-valued in the activation sparsity tensor may indicate that the corresponding element in the activation tensor is nonzero. A weight sparsity tensor may be the sparsity tensor of a weight tensor and has the same number of elements as the weight tensor. An element in the weight sparsity tensor may indicate whether the corresponding element in the weight tensor is zero or not. For instance, a zero-valued in the weight sparsity tensor may indicate that the corresponding element in the weight tensor is zero. A one-valued in the weight sparsity tensor may indicate that the corresponding element in the weight tensor is nonzero. The sparsity module may generate a combined sparsity tensor using an activation sparsity tensor and a weight sparsity tensor. For instance, the sparsity module may multiply an element of the activation sparsity tensor with a corresponding element of the weight sparsity tensor to compute an element of the combined sparsity tensor. The positions of the three elements in their corresponding sparsity tensors may match. In some embodiments, each element in a sparsity tensor may be a bit, and the sparsity tensor may be referred to as a sparsity bitmap.

[0089] The sparsity module may use the sparsity tensor to identify activations and weights to be used in MAC operations by the MAC units. In an embodiment where the processing engine 370 operates in the combined sparsity mode, the sparsity module may identify activations and weights that correspond to nonzero valued elements of a combined sparsity tensor. In an embodiment where the processing engine 370 operates in the activation sparsity mode, the sparsity module may identify activations and weights that correspond to nonzero valued elements of an activation sparsity tensor. In an embodiment where the processing engine 370 operates in the weight sparsity mode, the sparsity module may identify activations and weights that correspond to nonzero valued elements of a weight sparsity tensor. The sparsity module may be bypassed in the dense mode as no sparsity acceleration would be conducted.

[0090] The post-processing engine 380 processes outputs of the processing engine 370. The post-processing engine 380 may include one or more post-processing elements. In some embodiments, the post-processing elements in the post-processing engine 380 may be  arranged in an arrange that has rows and columns. In some embodiments, the post-processing engine 380 computes activation functions. The post-processing engine 380 may receive outputs of the processing engine 370 as inputs to the activation functions. In addition or alternative to activation functions, the post-processing engine 380 may perform other types of post processing on outputs of the processing engine 370. For instance, the post-processing engine 380 may apply a bias on an output of the processing engine 370. In some embodiments, the post-processing engine 380 may be bypassed for certain neural network operations.

[0091] The drain module 390 drains data from the processing engine 370 or from the post-processing engine 380. The drain module may write the data to the local memory 340. The drained data may be tensors, such as output tensors of neural network operations. In some embodiments, the drain module 390 may perform tensor permutation to change storage formats of tensors. For instance, the drain module 390 may permute tensors before writing the tensors to the local memory 340 so that the drain module 390 may write the tensors to the local memory 340 in the new formats. In some embodiments, the drain module 390 may perform tensor permutation to change Z-major formats to X-major formats or Y-major formats. For instance, a tensor drained by the drain module 390 from the processing engine 370 or from the post-processing engine 380 may be in a Z-major format. The drain module 390 may change the Z-major format to a X-major or Y-major format and write the tensor to the local memory 340 in the X-major or Y-major format.

[0092] In some embodiments, the drain module 390 may drain data on a cell level. For each processing cell, the drain module 390 may drain outputs of PEs in the processing cell based on a row index or column index of each PE. For instance, the drain module 390 may use a sequence of cycles to drain data from a processing cell. The drain module 390 may drain the output of some of the PE s in each cycle. The sequence of the cycles may be configured based on a configuration parameter indicating the operation mode of the load module 360.

[0093] In some embodiments, the drain module 390 includes sparsity encoding logic that can convert outputs of the processing engine 370 from a dense format to a sparse format. For instance, the drain module 390 may be implemented with one or more sparsity encoders. A sparsity encoder converts dense data to compressed data based on sparsity in the dense data. For instance, the sparsity encoder may remove zeros in an activation tensor  computed by the processing engine 370 to convert the activation tensor to a compressed activation tensor. The sparsity encoder may also generate sparsity tensors, including activation sparsity tensors.

[0094] In some embodiments, the data drained from the processing engine 370 may be at least part of an output tensor (e.g., the output tensor 230 in FIG. 2) of a deep learning operation. The sparsity encoder may generate a compressed version of the output tensor. The sparsity encoder may identify every zero-valued activation in the output tensor and remove these activations from the output tensor to generate a compressed activation tensor (aka “sparse activation tensor” ) . The sparsity encoder may also generate one or more sparsity tensors for the output tensor. A sparsity tensor may correspond to a portion of the output tensor (e.g., the vector 235 in FIG. 2) . The sparsity tensor may include sparsity elements (e.g., bits) , each of which corresponds to a different activation in the vector and indicates whether the corresponding activation is zeroed or not.

[0095] The drain module 390 may write the compressed activation tensor and the one or more sparsity tensors into the local memory 340. The sparse activation tensor and the one or more sparsity tensors may be further loaded to the memory 310, e.g., through the DMA engine 320. Additionally or alternatively, the sparse activation tensor and the one or more sparsity tensors may be loaded by the load module 360 to the processing engine 370 for further computation, e.g., for performing a deep learning operation in the next layer.

[0096] FIG. 4 is a block diagram of a DNN module 400, in accordance with various embodiments. The DNN module 400 may be an embodiment of the DNN module 301 in FIG. 3. As shown in FIG. 4, the DNN module 400 includes an interface module 410, a training module 420, a compressing module 430, a converting module 440, a compiler 450, and a datastore 460. In other embodiments, alternative configurations, different or additional components may be included in the DNN module 400. Further, functionality attributed to a component of the DNN module 400 may be accomplished by a different component included in the DNN module 400 or a different module or system. For instance, certain functionality attributed to the converting module 440 may be accomplished by the compiler 450, or vice versa.

[0097] The interface module 410 facilitates communications of the DNN module 400 with other modules or systems. For example, the interface module 410 establishes  communications between the DNN module 400 with an external database to receive data that can be used to train DNNs or input into DNNs to perform tasks. As another example, the interface module 410 may distribute trained DNNs to other systems, e.g., computing devices configured to apply DNNs to perform tasks.

[0098] The training module 420 trains DNNs by using a training dataset. The training module 420 forms the training dataset. In an example where the training module 420 trains an DNN to recognize objects in images, the training dataset includes training images and training labels. The training labels describe ground-truth classifications of objects in the training images. In some embodiments, each label in the training dataset corresponds to an object in a training image. In some embodiments, a part of the training dataset may be used to initially train the DNN, and the rest of the training dataset may be held back as a validation subset used by the training module 420 to validate performance of a trained DNN. The data portion of the training dataset not including the tuning subset and the validation subset may be used to train the DNN.

[0099] The training module 420 also determines hyperparameters for training the DNN. Hyperparameters are variables specifying the DNN training process. Hyperparameters are different from parameters inside the DNN (e.g., weights of filters) . In some embodiments, hyperparameters include variables determining the architecture of the DNN, such as number of hidden layers, etc. Hyperparameters also include variables which determine how the DNN is trained, such as batch size, number of epochs, etc. A batch size defines the number of training samples to work through before updating the parameters of the DNN. The batch size is the same as or smaller than the number of samples in the training dataset. The training dataset can be divided into one or more batches. The number of epochs defines how many times the entire training dataset is passed forward and backwards through the entire network. The number of epochs defines the number of times that the deep learning algorithm works through the entire training dataset. One epoch means that each training sample in the training dataset has had an opportunity to update the parameters inside the DNN. An epoch may include one or more batches. The number of epochs may be 1, 5, 10, 50, 100, 500, 1000, or even larger.

[0100] The training module 420 defines the architecture of the DNN, e.g., based on some of the hyperparameters. The architecture of the DNN includes an input layer, an output layer,  and a plurality of hidden layers. The input layer of an DNN may include tensors (e.g., a multidimensional array) specifying attributes of the input image, such as the height of the input image, the width of the input image, and the depth of the input image (e.g., the number of bits specifying the color of a pixel in the input image) . The output layer includes labels of objects in the input layer. The hidden layers are layers between the input layer and output layer. The hidden layers include one or more convolutional layers and one or more other types of layers, such as pooling layers, fully-connected layers, normalization layers, SoftMax or logistic layers, and so on. The convolutional layers of the DNN abstract the input image to a feature map that is represented by a tensor specifying the feature map height, the feature map width, and the feature map channels (e.g., red, green, blue images include 3 channels) . A pooling layer is used to reduce the spatial volume of input image after convolution. It is used between two convolution layers. A fully-connected layer involves weights, biases, and neurons. It connects neurons in one layer to neurons in another layer. It is used to classify images between different categories by training.

[0101] In the process of defining the architecture of the DNN, the training module 420 also adds an activation function to a hidden layer or the output layer. An activation function of a layer transforms the weighted sum of the input of the layer to an output of the layer. The activation function may be, for example, a ReLU activation function, a tangent activation function, or other types of activation functions.

[0102] After the training module 420 defines the architecture of the DNN, the training module 420 inputs a training dataset into the DNN. The training dataset includes a plurality of training samples. An example of a training sample includes an object in an image and a ground-truth label of the object. The training module 420 modifies the parameters inside the DNN ( “internal parameters of the DNN” ) to minimize the error between labels of the training objects that are generated by the DNN and the ground-truth labels of the objects. The internal parameters include weights of filters in the convolutional layers of the DNN. In some embodiments, the training module 420 uses a cost function to minimize the error.

[0103] The training module 420 may train the DNN for a predetermined number of epochs. The number of epochs is a hyperparameter that defines the number of times that the deep learning algorithm will work through the entire training dataset. One epoch means that each sample in the training dataset has had an opportunity to update internal parameters of the  DNN. After the training module 420 finishes the predetermined number of epochs, the training module 420 may stop updating the parameters in the DNN. The DNN having the updated parameters is referred to as a trained DNN.

[0104] The training module 420 may also verify accuracy of DNNs after training. In some embodiments, the training module 420 inputs samples in a validation dataset into a trained DNN and uses the outputs of the DNN to determine the model accuracy. In some embodiments, a validation dataset may be formed of some or all the samples in the training dataset. Additionally or alternatively, the validation dataset includes additional samples, other than those in the training sets. In some embodiments, the training module 420 may determine an accuracy score measuring the precision, recall, or a combination of precision and recall of the DNN. The training module 420 may use the following metrics to determine the accuracy score: Precision = TP  /  (TP + FP) and Recall = TP  /  (TP + FN) , where precision may be how many the DNN correctly predicted (TP or true positives) out of the total it predicted (TP + FP or false positives) , and recall may be how many the DNN correctly predicted (TP) out of the total number of objects that did have the property in question (TP +FN or false negatives) . The F-score (F-score = 2 *PR  /  (P + R) ) unifies precision and recall into a single measure.

[0105] The training module 420 may compare the accuracy score with a threshold score. In an example where the training module 420 determines that the accuracy score of the DNN is less than the threshold score, the training module 420 instructs the training module 420 to re-train the DNN. In one embodiment, the training module 420 may iteratively re-train the DNN until the occurrence of a stopping condition, such as the accuracy measurement indication that the DNN may be sufficiently accurate, or a number of training rounds having taken place.

[0106] The compressing module 430 compresses DNNs. For instance, the compressing module 430 may add pruning operations to DNN layers to reduce computational complexity or memory usage. A pruning operation may prune weight tensors of a DNN layer by changing one or more nonzero valued weights of the layer to zeros. The modification may be done before, during, or after training. Weights may be pruned during training, during inference, or a combination of both. The compressing module 430 may determine a sparsity ratio for a DNN layer. The sparsity ratio may be a ratio of the number of zero-valued weight  to the total number of weights in the layer. The compressing module 430 may perform the pruning operation till the sparsity ratio of the DNN layer meets a target sparsity ration, such as 10%, 20%, 30%, 40%, 50%, and so on.

[0107] In some embodiments, the compressing module 430 may select one or more layers in a DNN and modify each selected layer with a pruning operation. For instance, the compressing module 430 may select computationally complex layers, such as layers with large filters. For a pruning operation of a layer or of a type of layer, the compressing module 430 may determine a weight threshold that would not cause a loss of the accuracy of the DNN to exceed an accuracy loss constraint. A pruning operation may modify weights having absolute values above the weight threshold to zeros and leave the other weights unchanged. The weight pruning can reduce memory storage as zero-valued weights may not be stored. Also, the number of operations in the layer can be reduced as computations on zero-valued weights can be skipped without impacting the output of the layer. In some embodiments, the compressing module 430 may also measure energy saving, final DNN accuracy, or layer-wise sparsity caused by pruning operations.

[0108] After compressing a DNN, the compressing module 430 may fine tune the DNN, e.g., through a retraining process. The compressing module 430 may fine tunes DNNs after weights are pruned. In some embodiments, the fine-tuning process is a retraining or further training process. For instance, after weights in a DNN are pruned, the compressing module 430 may further train the DNN by inputting a training dataset into the DNN. The values of the unpruned weights in the DNN may be modified based on outputs of the DNN and ground-truth labels of the training samples in the training dataset. In some embodiments, the values of the pruned weights (i.e., zero) are not changed during the fine-tuning process. For instance, the compressing module 430 may place a mask over a pruned weight block and the mask can prevent values in the pruned weight blocks from being changed during the fine-tuning process. In other embodiments, the values of all weights, including the pruned weights, may be changed during the fine-tuning process. After one or more cycles of retraining and weight changing by the compressing module 430, the compressing module 430 may perform a new pruning process, e.g., by selecting weight blocks and pruning the selected weight blocks. In some embodiments, the weight pruning process may be repeated multiple times before the fine-tuning process is done. In some embodiments, the number of  epochs in the fine-tuning process may be different from the number of epochs in the training process in which the pre-pruning values of the weights are determined. For instance, the fine-tuning process may have less epochs than the training process. In an example, the number of epochs in the fine-tuning process may be relatively small, such as 2, 3, 4, 5, and so on.

[0109] The converting module 440 converts neural network operations in DNNs for optimizing the performance of hardware devices executing the DNNs. The neural network operations include convolutions. In some embodiments, the converting module 440 may identify or analyze one or more configurations of the hardware device that will be used to perform a neural network operation in a DNN and convert the neural network operation based on the one or more configurations of the hardware device. The hardware device may be the DNN accelerator 302 or part of the DNN accelerator 302. The configurations may be characteristics of the hardware device itself or configurations of the hardware device that are determined by the compiler 450. Examples of the configurations of the hardware device may include arrangements of available PEs or MAC units in the DNN accelerator 302, such as the number of PEs or MAC units in a sparce cell or sparse cell array, the number of PEs or MAC units in a row of a sparce cell or sparse cell array, the number of PEs or MAC units in a column of a sparce cell or sparse cell array, other hardware configurations, or some combination thereof.

[0110] In some embodiments, the converting module 440 may convert neural network operations to increase or optimize the utilization of PEs. For instance, the converting module 440 may change (e.g., increase) the number of input channels of a convolution to be a power of two, such as 4, 8, 16, 32, and so on. Activations and weights in different input channels may be distributed to different PEs for performing MAC operations. In an example in which a sparse cell has 16 PEs, the PE utilization may be optimized when there are 16 (or a multiple of 16) input channels.

[0111] To convert a neural network operation, the converting module 440 may determine a reshaping factor m and use the reshaping factor to convert one or more tensors of the neural network operation. In some embodiments, the converting module 440 may determine the reshaping factor for a convolution based on characteristics of the convolution, such as kernel length, stride length, or other characteristics of the convolution.

[0112] An example convolution may be denoted as:

[0113] where input is the input tensor of the convolution; kernel is the weight tensor of the convolution; padsbegin and padsend specifies padding of the input tensor, in which PBx is the number of columns to be added at the beginning of the X axis, PBy is the number of rows to be added at the beginning of the Y axis, PEx is the number of columns to be added at the end of the X axis, and PEy the number of rows to be added at the end of the Y axis; and strides denotes the stride size of the convolution, in which strideX is the number of columns that the kernel moves across the input tensor and strideY is the number of rows that the kernel moves across the input tensor.

[0114] The input tensor of the convolution may be denoted as and may be denoted as input= tensor<N×Y×X×IC>, where N represents the number of batches in the input tensor, Y represent the spatial length of the input tensor along the Y axis (which may be the height of the input tensor) , and X represent the spatial length of the input tensor along the X axis (which may be the width of the input tensor) , and IC represents the number of input channels (which may be the depth of the input tensor and kernel) . The kernel may be denoted as kernel=tensor<OC×IC×Ky×Kx>, where OC is the number of output channels (which may be the number of batches in the kernel) , Ky represent the kernel’s spatial length along the Y axis (which may be the height of the kernel) , and Kx represent the kernel’s spatial length along the X axis (which may be the width of the kernel) . The result of the convolution (e.g., the output tensor) may be denoted as output=tensor<N×OC×Oy×Ox>, where Oy represent the output tensor’s spatial length along the Y axis (which may be the height of the output tensor) , and Ox represent the output tensor’s spatial length along the X axis (which may be the width of the output tensor) .

[0115] In some embodiments, the converting module 440 may determine a preliminary factor n based on Kx and strideX. For instance, the converting module 440 may compare Kx and strideX. In embodiments where strideX≥Kx, the converting module 440 may determine that n=1. In embodiments where strideX<Kx, the converting module 440 may compute n=DivUp (Kx, strideX) , in which DivUp (Kx, strideX) =  (Kx+strideX-1) ÷strideX. The converting module 440 may determine the reshaping factor from the preliminary factor by multiplying the preliminary factor with strideX, i.e., m=strideX when strideX≥Kx and m=n×strideX when strideX<Kx.

[0116] After the reshaping factor is determined, the converting module 440 may adjust the number input channels using the reshaping factor. For instance, the converting module 440 may determine the new number of input channels by multiplying the original number of input channels with the reshaping factor, i.e., IC′=m×IC, where IC′denotes the number of input channels of the converted convolution. Additionally or alternatively, the converting module 440 may adjust the number of output channels by multiplying the original number of output channels with the preliminary factor, i.e., OC′=n×OC, where OC′denotes the number of output channels of the converted convolution.

[0117] As the number of activations in the input tensor or output tensor remains the same despite the conversion, the converting module 440 may decrease the width of the input tensor or output tensor. For instance, the converting module 440 may convert the input tensor by adjusting the width of the input tensor using the reshaping factor. In an example, the conversion of the input tensor may be denoted as X′=X÷m, where X′denotes the length of the converted input tensor. Additionally or alternatively, the converting module 440 may convert the output tensor by adjusting the width of the input tensor using the preliminary tensor. In an example, the conversion of the output tensor may be denoted as OX′=OX÷n, where OX′denotes the length of the converted output tensor.

[0118] In some embodiments, the converting module 440 may apply a power of two on the number of input channels of the converted convolution or the number of output channels of the converted convolution. The converting module 440 may ensure that the width of the converted input tensor satisfies the following criterion:

[0119] The converted convolution has the converted input tensor, converted kernel, and the converted output tensor. In some embodiments, the converted input tensor may be denoted as input′=tensor<N×IC′×Y×X′>, the converted kernel may be denoted as kernel′=tensor<OC′×IC′×Ky×3>, and the converted output tensor may be denoted as output′=tensor<N×OC′×Oy×Ox′>. The physical memory layout of the output activations of the converted convolution may be the same as the physical memory  layout of the output activations of the original convolution. The converting module 440 may also adjust one or more padding parameters and striding lengths. In an example, the converted convolution may be denoted as:

[0120] The physical memory layout of the output activations of the converted convolution may be the same as the physical memory layout of the output activations of the original convolution. The total number of output activations may be denoted as OC×IC×Kx×Ky×Ox×Oy×N, in which Ox= (X+padX-Kx+1) ÷strideX and Oy= (Y+padY-Ky+1) ÷strideY.

[0121] In an example, PadsBegin of the original convolution is [1, 1] , PadsEnd of the original convolution is [1: 1] , the stride size of the original convolution is [2, 2] , the shape of the input tensor of the original convolution is 1×4×1080×2048, the shape of the kernel of the original convolution is 4×4×4×4, and the shape of the output tensor of the original convolution is to 1×4×540×1024, in which 4 is the new number of output channels. PadsBegin and PadsEnd may remain the same after the conversion. The stride size may be changed to [2, 1] . The shape of the input tensor may be changed to 1×32×1080×256. The shape of the kernel is changed to 16×32×4×3. The shape of the output tensor is changed to 1×16×540×256, in which 16 is the new number of output channels.

[0122] Compared with currently available expand solutions, the performance of the DNN accelerator can be better with the reshape / convert solution. In some embodiments, the converting module 440 may estimate the improvement in the performance of the DNN accelerator executing the converted convolution compared with currently available expand solutions. For currently available expand solutions, the converting module 440 may compute the following:

[0123] ExpandCalcCnt= OCalign×ICalign×Kx×Ky×Ox×Oy×N, OCalign=DivUp (OC, alignedChannel) ×alignedChannel, and

[0124] ICalign=DivUp (IC, alignedChannel) ×alignedChannel,

[0125] in which ExpandCalcCnt denotes the number of output activations computed in the expanded convolution, OCalign denotes the number of output channels of the expanded convolution, ICalign denotes the number of input channels of the expanded convolution, Kx, Ky Ox, Oy, and N are the  same as the origin convolution. For the reshape solution, the converting module 440 may compute:

[0126] ReshapeCalcCnt= OC′×C′×3×Ky×Ox′×Oy×N,

[0127] in which ReshapeCalcCnt denotes the number of output activations computed in the reshaped convolution.

[0128] The difference in the number of output activations for the two solutions may be denoted as:

[0129] When strideX=1, n=Kx, Ox=X. So:

[0130] In an example in which the number of input channels and the number of output channels are both 3 and the hardware required channel alignment is 16, the performance difference can be denoted as:

[0131] When strideX=2 and Kx=4, n=2 and m=4. The performance difference would be:

[0132] In an example in which the number of input channels and the number of output channels are both 4 and the hardware required channel alignment is 16, the performance difference can be denoted as:

[0133] Therefore, the reshape solution has less calculation than expand solutions. In some embodiments, the converting module 440 may reshape the output of the previous convolution, which can decrease the expand / slice stride copies before / after the convolution.

[0134] The compiler 450 compiles information of DNNs to executable instructions that can be executed, e.g., by the DNN accelerator 302, to carry out neural network operations in DNNs. In some embodiments, the compiler 405 may generate a graph representing a DNN. The graph may include nodes and edges. A node may represent a specific neural network operation in the DNN. An edge may connect two nodes and represent a connection between the two corresponding neural network operations. In an example, an edge may encode a tensor that flows from one of the neural network operations to the other neural network operation. The tensor may be an output tensor of the first neural network operation and an input tensor of the second neural network operation. The edge may encode one or more attributes of the tensor, such as size, shape, storage format, and so on. The compiler 450 may use the graph to generate executable DNNs. For instance, the compiler may generate computer program instructions for executing DNNs.

[0135] In some embodiments, the compiler 450 may generate configuration parameters that may be used to configure components of the DNN accelerator 302 for DNN executions. The configuration parameters may be stored in one or more configuration registers associated with the components of the DNN accelerator 302. In some embodiments, the compiler 450 may compile a DNN after the converting module 440 converts neural network operations in the DNN. For instance, the compiler 450 may generate configuration parameters that cause a load module to load input activations and weights into PEs for performing convolutions converted by the converting module 440. The compiler 450 may also generate configuration parameters that cause a drain module to write output activations computed by the PEs into memory.

[0136] The datastore 460 stores data received, generated, used, or otherwise associated with the DNN module 400. For example, the datastore 460 stores the datasets used by the training module 420. The datastore 460 may also store data generated by the training module 420, such as the hyperparameters for training DNNs, internal parameters of trained DNNs (e.g., weights, etc. ) , data for sparsity acceleration (e.g., sparsity bitmap, etc. ) , and so  on. The datastore 460 may store instructions, configuration parameters, or other data generated by the compiler 450. The datastore 460 may further store data received, processed, or generated by the converting module 440. The datastore 460 may include one or more memories. In the embodiment of FIG. 4, the datastore 460 is a component of the DNN module 400. In other embodiments, the datastore 460 may be external to the DNN module 400 and communicate with the DNN module 400 through a network.

[0137] FIGS. 5A-5C illustrate reshaping a tensor 510 without changing the memory layout for storing the tensor, in accordance with various embodiments. The tensor 510 is shown in FIG. 5A. For the purpose of illustration and simplicity, the tensor 510 is a 3D tensor that includes 108 data elements, each of which is represented by a box in FIG. 5A. As shown in FIG. 5A, the 108 data elements are arranged as a data structure having a shape of 3×12×3, meaning the height of the tensor 510 (i.e., the length along the Y axis) is 3, the width of the tensor 510 (i.e., the length along the X axis) is 12, and the depth of the tensor 510 (i.e., the length along the Z axis) is 3. In some embodiments (e.g., embodiments where the tensor 510 is an input activation tensor of a convolution) , the depth may indicate the number of input channels of the convolution.

[0138] FIG. 5B shows a converted tensor 520 that is generated by reshaping the tensor 510, e.g., by the converting module 440 in FIG. 4. As shown in FIG. 5B, the converted tensor 520 is a 3D tensor having a shape of 3×6×6, meaning the height of the converted tensor 520 (i.e., the length along the Y axis) is 3, the width of the converted tensor 520 (i.e., the length along the X axis) is 6, and the depth of the converted tensor 520 (i.e., the length along the Z axis) is 6. In some embodiments, the converted tensor 520 may be generated by splitting every channel in the tensor 510 into two channels so that the converted tensor 520 has a smaller width by greater depth than the tensor 510. The converted tensor 520 has the same data elements at the tensor 510.

[0139] FIG. 5C shows a converted tensor 530 that is generated by reshaping the tensor 510 or converted tensor 520, e.g., by the converting module 440 in FIG. 4. As shown in FIG. 5B, the converted tensor 530 is a 3D tensor having a shape of 3×3×12, meaning the height of the converted tensor 530 (i.e., the length along the Y axis) is 3, the width of the converted tensor 530 (i.e., the length along the X axis) is 3, and the depth of the converted tensor 530 (i.e., the length along the Z axis) is 12. In some embodiments, the converted tensor 530 may  be generated by splitting every channel in the tensor 510 into four channels so that the converted tensor 530 has a smaller width by greater depth than the tensor 510 and converted tensor 520. The converted tensor 530 has the same data elements as the tensor 510 and converted tensor 520.

[0140] Even though the tensor 510, converted tensor 520, and converted tensor 530 have different shapes, they have the same 108 data elements. Also, the memory layout of the 108 data elements may be the same for all the three tensors. The memory layout may define the order in which the 108 data elements are stored in a memory, such as the memory 310 or local memory 340. In some embodiments, the tensor 510, converted tensor 520, or converted tensor 530 may be 4D tensors, in which three of the four dimensions are spatial dimension and the other dimension is a batch dimension that indicates the number of batches. For instance, the shape of the tensor 510 may be denoted as 1×3×12×3, the shape of the converted tensor 520 may be denoted as 1×3×6×6, and the shape of the converted tensor 530 may be denoted as 1×3×3×12, in which 1 may indicate that each of these tensors has one batch.

[0141] FIG. 6A illustrates an input tensor 610 and kernel 615 of a convolution, in accordance with various embodiments. For the purpose of illustration and simplicity, the input tensor 610 is a 3D tensor that includes 108 data elements arranged as a data structure. Each data element may be an input activation of the convolution. The input tensor 610 has a shape of 1×3×12×3, meaning the number of batches is 1, the height of the input tensor 610 (i.e., the length along the Y axis) is 3, the width of the input tensor 610 (i.e., the length along the X axis) is 12, and the depth of the input tensor 610 (i.e., the length along the Z axis) is 3. The kernel 615 is a 3D tensor that includes 18 data elements arranged as a data structure. Each data element may be a weight of the convolution. The kernel 615 has a shape of 1×3×2×3, meaning the number of batches is 1, the height of the kernel 615 (i.e., the length along the Y axis) is 3, the width of the kernel 615 (i.e., the length along the X axis) is 2, and the depth of the kernel 615 (i.e., the length along the Z axis) is 3. Both the depth of the input tensor 610 and the depth of the kernel 615 may be the number of input channels of the convolution. As the number of batches is 1, the output tensor of the convolution may have a single output channel. The output tensor of the convolution may have a shape of 1×1×6×1.

[0142] FIG. 6B illustrates converting the input tensor 610 and kernel 615 in FIG. 6A when a kernel length of the convolution equals a stride length of the convolution, in accordance with various embodiments. The convolution may have two stride lengths: a stride length along the Y axis (StrideY) and another stride length along the X axis (StrideX) . StrideY may indicate the number of rows that the kernel 615 moves across the input tensor 610. StrideX may indicate the number of columns that the kernel 615 moves across the input tensor 610. In an example in which StrideX is 1, in the first cycle of the convolution, the kernel 615 may be applied on the first two columns of the input tensor 610; in the second cycle of the convolution, the kernel 615 may be applied on the second column and third column of the input tensor 610. In an example in which StrideX is 2, in the first cycle of the convolution, the kernel 615 may be applied on the first two columns of the input tensor 610; in the second cycle of the convolution, the kernel 615 may be applied on the third column and fourth column of the input tensor 610. In some embodiments, the stride lengths that equals the kernel length may be StrideX.

[0143] In the embodiments of FIG. 6B, the stride size of the convolution is 1×2, meaning StrideY is 1 and StrideX is 2. Therefore, the kernel length of the convolution along the X axis (i.e., the width of the kernel 615) equals StrideX. StrideX can be used as a reshaping factor to reshape the input tensor 610 and kernel 615 for converting the convolution. The conversion of the convolution includes reshaping the input tensor 610 to generate a converted tensor 620 and reshaping the kernel 615 to generate a converted kernel 625. In some embodiments, the reshaping of the converted tensor 620 or converted kernel 625 includes increasing the number of input channels in the converted tensor 620 or converted kernel 625 and reducing the width of the converted tensor 620 or converted kernel 625 by 2, i.e., the value of StrideX.

[0144] As shown in FIG. 6B, the converted tensor 620 has a shape of 1×3×6×6, meaning the number of batches is 1, the height of the converted tensor 620 is 3, the width of the converted tensor 620 is 6, and the depth of the converted tensor 620 is 6. The converted kernel 625 includes 18 data elements arranged as a data structure. Each data element may be a weight of the convolution. The converted kernel 625 has a shape of 1×3×1×6, meaning the number of batches is 1, the height of the converted kernel 625 is 3, the width of the converted kernel 625 is 1, and the depth of the converted kernel 625 is 6. Both the  depth of the converted tensor 620 and the depth of the converted kernel 625 may be the number of input channels of the converted convolution, which is 6. The stride size of the converted convolution may be 1×1, meaning both StrideY and StrideX are 1. The output tensor of the converted convolution may have the same output activations and the same shape as the output tensor of the convolution. Therefore, the conversion of the convolution (e.g., the reshaping of the input tensor 610 and kernel 615) does not change the output.

[0145] FIG. 6C illustrates converting the input tensor and kernel in FIG. 6A when a kernel length of the convolution is smaller than a stride length of the convolution, in accordance with various embodiments. In the embodiments of FIG. 6C, the stride size of the original convolution is 1×3, meaning StrideY is 1 and StrideX is 3. Therefore, the width of the kernel 615 is smaller than StrideX. StrideX can be used as a reshaping factor to reshape the input tensor 610 and kernel 615 for converting the convolution. The conversion of the convolution in FIG. 6C includes reshaping the input tensor 610 to generate a converted tensor 620 and reshaping the kernel 615 to generate a converted kernel 625. In some embodiments, the reshaping of the converted tensor 620 or converted kernel 625 includes increasing the number of input channels in the converted tensor 620 or converted kernel 625 and reducing the width of the converted tensor 620 or converted kernel 625 by 3, i.e., the value of StrideX.

[0146] As shown in FIG. 6C, the converted tensor 630 has a shape of 1×3×4×9, meaning the number of batches is 1, the height of the converted tensor 630 is 3, the width of the converted tensor 630 is 4, and the depth of the converted tensor 630 is 9. In addition to reshaping the kernel 615, the conversion of the convolution in the embodiments of FIG. 6C also includes kernel padding. The kernel padding may include adding one or more zeros as additional weights. The converted kernel 635 includes 27 data elements including the 18 weights in the kernel 615 and 9 zeros. The 27 data elements are arranged as a data structure with a shape of 1×3×1×9, meaning the number of batches is 1, the height of the converted kernel 635 is 3, the width of the converted kernel 635 is 1, and the depth of the converted kernel 635 is 9. The kernel padding may be done to make sure the converted kernel 635 has the same number of channels as the converted tensor 630., which is 9. The stride size of the converted convolution may be 1×1, meaning both StrideY and StrideX are 1. The output tensor of the converted convolution may have the same output activations  and the same shape as the output tensor of the convolution. Therefore, the conversion of the convolution may not affect the output.

[0147] FIG. 7A illustrates an input tensor 710, kernel 720, and output tensor 730 of a convolution, in accordance with various embodiments. For the purpose of illustration and simplicity, the input tensor 710 includes 96 input activations and has a shape of 1×4×8×3, meaning the number of batches is 1, the height of the input tensor 710 is 4, the width of the input tensor 710 is 8, and the depth of the input tensor 710 is 3. The kernel 720 includes 18 weights and has a shape of 1×3×3×3, meaning the number of batches is 1, the height of the kernel 720 is 3, the width of the kernel 720 is 3, and the depth of the kernel 720 is 3. Both the depth of the input tensor 710 and the depth of the kernel 720 may be the number of input channels of the convolution.

[0148] The convolution may have padding. The padding process may include adding new data elements to one or more boundaries of the input tensor 710. In some embodiments, the new data elements may be zeros. The padding may be defined by padding parameters, such as padding parameters indicating the beginning and ending along each spatial axis. There may be a padding parameter indicating the number of activations added at the beginning of the X axis, padding parameter indicating the number of activations added at the beginning of the Y axis, padding parameter indicating the number of activations added at the end of the X axis, padding parameter indicating the number of activations added at the end of the Y axis, or some combination thereof. In an example, the convolution in FIG. 7A has a padding with PadsEnd [1: 1] , meaning a new column is added to the right edge of the input tensor 710 for each input channel and a new row is added to the bottom edge of the input tensor 710 for each input channel. The padding of the input tensor 710 produces a padded tensor 711. The padded tensor 711 has a shape of 1×5×9×3. The new activation in the padded tensor 711 are represented by boxes highlighted by a dotted pattern. The kernel 720 is applied on the padded tensor 711 with a stride size of 1×2. The output tensor 730 has a shape of 1×2×4×1, meaning the number of batches is 1, the height of the output tensor 730 is 2, the width of the output tensor 730 is 4, and the depth of the output tensor 730 is 1. The depth of the output tensor 730 may be the number of output channels of the convolution.

[0149] The convolution may be converted to optimize the performance and efficiency of the  DNN accelerator performing the convolution. The conversion of the convolution may include converting the input tensor 710 or the kernel 720. The output of the converted convolution may have a different shape from the output tensor 730, but the output activations would be the same.

[0150] FIG. 7B illustrates a converted input tensor 715 generated by reshaping the input tensor 710 in FIG. 7A for converting the convolution, in accordance with various embodiments. As shown in FIG. 7B, the converted input tensor 715 has the same input activations as the input tensor 710, but the input activations are rearranged into a data structure with a shape of 1×4×2×12. The converted convolution has 12 input channels.

[0151] FIG. 7C illustrates a converted kernel 725 generated by converting the kernel 720 in FIG. 7A for converting the convolution, in accordance with various embodiments. The conversion of the kernel 720 includes reshaping and padding. The padding may include adding zeros, which are shown in FIG. 7C. The zeros may be skipped from storing in one or more storage units (e.g., the local memory 340, register files in sparse cell, etc. ) or being used in MAC operations, so that the padding may introduce none or minimal overhead. The converted kernel 725 has a shape of 2×3×3×12. The converted kernel 725 has two batches 725A and 725B, each of which is a 3D tensor having a shape of 3×3×12.

[0152] The conversion of the convolution also includes change of the stride size and padding. The stride size may be changed to 1×1. The new padding may be represented by PadsBegin [1, 0] and PadsEnd [1: 1] . PadsBegin [1, 0] indicates that a new column is added to the left edge of the converted input tensor 715 for each input channel but no new row is added.

[0153] FIG. 7D illustrates an output tensor 735 of the converted convolution, in accordance with various embodiments. The output tensor 735 may be generated by perform the converted convolution using the converted input tensor 715 and the converted kernel 725 with the new stride size and new padding. The output tensor 735 has the same output activations as the output tensor 730 even though it has a different shape. The output activations may be stored in a memory, e.g., the memory 310 or local memory 340. The memory layout of the output activation may not need to be changed for performing the next neural network operation in the DNN.

[0154] FIG. 8A illustrates an example kernel 810 for a single output channel, in accordance  with various embodiments. The kernel 810 may be a kernel of a constitution. For the purpose of illustration, the kernel 810 includes 64 weights and has a shape of 1×4×4×4, meaning the number of batches is 1, the height of the kernel 810 is 4, the width of the kernel 810 is 4, and the depth of the kernel 810 is 4. FIGS. 8B-8E illustrate a kernel generated by converting the kernel 810 in FIG. 8A, in accordance with various embodiments. The conversion of the kernel 810 may be for increasing the performance and efficiency of the DNN accelerator executing the convolution. The new kernel has four batches 820A-820D, which are shown in FIGS. 8B-8E, respectively. Each of the batches 820A-820D has a shape of 4×3×32. The shape of the new kernel is 4×4×3×32. The number of input channels is increased from 4 to 32, which can improve the utilization rate of PEs in the DNN accelerator.

[0155] FIG. 9A illustrates an example sparse cell 900, in accordance with various embodiments. The sparse cell 900 may be a processing cell in a processing engine, e.g., the processing engine 370 in FIG. 3. The sparse cell 900 includes 16 MAC units 910 (individually referred to as “MAC unit 910” ) , which constitutes a MAC array having four rows and four columns. The MAC array has a spatial shape of 9x4, meaning the height of the MAC array is four and the width of the MAC array is also 9. The sparse cell 900 also includes 16 weight register files 920 (individually referred to as “weight register file 920” ) , 16 activation register files 930 (individually referred to as “activation register file 930” ) , four row buffers 940 (individually referred to as “row buffer 940” ) , and sparsity modules 960 (individually referred to as “sparsity module 960” ) . In other embodiments, the sparse cell 900 may include fewer, more, or different components. For example, the sparse cell 900 may include a different number of MAC units 910, weight register files 920, activation register files 930, row buffers 940, or sparsity modules 960. As another example, the sparse cell 900 may include column buffers in lieu of or in addition to the row buffers 940. Also, the shape (e.g., the height or width) of the MAC array may be different.

[0156] The MAC units 910 are configured to perform MAC operations. Each MAC unit 910 may include one or more multipliers and one or more adders. A multiplier may multiply an activation with a weight at a time to compute a product. In some embodiments (e.g., embodiments where the MAC unit 910 includes multiple multipliers) , the multipliers may operate simultaneously to process multiple activation-weight pairs and compute multiple products in one cycle. An adder may accumulate products computed by the multipliers. Even  though not shown in FIG. 9A, the sparse cell may include an adder tree including a plurality of adder tiers. The first tier may receive outputs of a plurality of MAC units 910. The number of adders in the first tier may be half of the number of the MAC units 910, and each adder may accumulate the outputs of two MAC units 910. The second tier may receive outputs of adders in the first tier. The number of adders in the second tier may be half of the number of adders in the first tier, and each adder in the second tier may accumulate the outputs of two adders in the first tier. The adder tree may include one or more other tiers. The last tier may include a single adder that accumulates outputs of adders in the second last tier to compute a partial sum of the sparse cell 900.

[0157] The weight register files 920 store weights to be processed in MAC operations. In the embodiments of FIG. 9A, four weight register files 920 are grouped into a storage set that stores data to be used by a column of MAC units 910. There are four storage sets corresponding to the four columns of MAC units 910. In some embodiments, a weight register file 920 may correspond to a MAC unit 910 and store data to be processed by the MAC unit. In some embodiments, all the 16 weight register files 920 constitute a weight storage unit.

[0158] The activation register files 930 stores activations to be processed in MAC operations. In the embodiments of FIG. 9A, four activation register files 930 are grouped into a storage set that stores data to be used by a row of MAC units 910. There are four storage sets corresponding to the four rows of MAC units 910. In some embodiments, an activation register file 930 may correspond to a MAC unit 910 and store data to be processed by the MAC unit. In some embodiments, all the 16 activation register files 930 constitute an activation storage unit. The row buffers 940 store outputs of the MAC units 910. Each row buffer 940 may drain outputs of a single row of MAC units 910.

[0159] The sparsity module 960 facilitates dynamic sparsity-based acceleration in the sparse cell 900. In the embodiments of FIG. 9A, each sparsity module 960 includes a sparsity tensor storage unit 965 and a control logic 967. The sparsity tensor storage unit 965 stores combined sparsity tensors. A combined sparsity tensor stored in the sparsity tensor storage unit 965 may correspond to an activation tensor and a weight tensor. A nonzero element in the combined sparsity tensor may correspond to a nonzero activation-weight pair that includes a nonzero activation and a nonzero weight. The position of the nonzero activation  in the activation tensor may match the position of the nonzero weight in the weight tensor. The product of the nonzero activation and nonzero weight would be nonzero.

[0160] The control logic 967 may control transmission of activations and weights stored from the weight register files 920 and the activation register files 930 to the MAC units 910 based on sparsity tensors. For instance, the control logic 967 may select a subset of the weights stored in the weight register files 920 and select a subset of activations stored in the activation register files 930 based on a combined sparsity tensor. The selected weights and activations constitute nonzero activation-weight pairs. The control logic 967 may transmit the selected weights and activations to the MAC units 910 for performing MAC operations. The other weights stored in the weight register files 920 and the other activations stored in the activation register files 930 are skipped from computation. In the embodiments of FIG. 9A, each sparsity module 960 controls sparsity acceleration in a respective MAC unit 910. As the sparsity acceleration is either based on both weight sparsity and activation sparsity, 16 sparsity modules 960 are used for acceleration computations in the 16 MAC units 910.

[0161] As shown in FIG. 9A, the sparse cell 900 is associated with multiplexers (MUXs) 903, 904, 905, and 906. In other embodiments, the sparse cell 900 may be associated with a different number of MUXs or other devices. The MUX 903 facilitates loading weights, e.g., from the local memory 340, into the weight register files 920. The MUX 904 facilitates loading activations, e.g., from the local memory 340, into the activation register files 930. The MUX 905 facilitates loading sparsity tensors into the sparsity tensor storage unit 965. The MUX 906 may be a drain MUX that can facilitate draining outputs of the MAC units 910, e.g., to the local memory 340.

[0162] FIG. 9B illustrates a sparse cell array 970, in accordance with various embodiments. The sparse cell array 970 may be an example of the processing engine 370 in FIG. 3. In FIG. 9B, the sparse cell array 970 includes sparse cells 980 (individually referred to as “sparse cell 980” ) arranged in four columns and four rows, an activation memory 990, and a weight memory 995. In other embodiments, the sparse cell array 970 may include fewer, more, or different components. For instance, the sparse cell array 970 may include a different number of columns, rows, or sparse cells 980.

[0163] Each sparse cell 980 may perform sparsity accelerated MAC operations. The sparse cells 980 may facilitate dynamic sparsity mode. For instance, the sparsity modes of a sparse  cell 980 may be dynamically changed between a combined sparsity mode, an activation sparsity mode, a weight sparsity mode, and a dense mode. An embodiment of a sparse cell 980 may be the sparse cell 900 in FIG. 9A. The activation memory 990 stores activations, such as activations in input tensors of neural network operations. Activations may be loaded from the activation memory 990 to sparse cells 980. The weight memory 995 stores weights, such as weights in filters of neural network operations. Weights may be loaded from the weight memory 995 to sparse cells 980. The activation memory 990 or weight memory 995 may be a buffer. In other embodiments, the sparse cell array 970 may include a dense data memory and a sparse data memory in lieu of the activation memory 990 and weight memory 995. The dense data memory may store dense tensors. The sparse data memory may store sparse tensors.

[0164] FIG. 10 illustrates an example PE 1000, in accordance with various embodiments. The PE 1000 may be a unit component of a processing cell, e.g., a processing cell in the processing engine 370. In the embodiments of FIG. 10, the PE 1000 includes an MAC unit 1005, an activation register file 1010, a weight register file 1020, an output register file 1050, and a sparsity accelerator 1060. The MAC unit 1005 includes a multiplier 1030 and an adder 1040. In other embodiments, the PE 1000 may include fewer, more, or different components.

[0165] The activation register file 1010 stores an activation operand, which may be a context. The activation register file 1010 may be an example of the activation register files 930 in FIG. 9. The weight register file 1020 stores a weight operand. The weight register file 1020 may be an example of the weight register files 920 in FIG. 9. The activation operand and weight operand may be loaded from a memory (e.g., the memory 340) into the activation register file 1010 and the weight register file 1020, respectively. The sparsity accelerator 1060 receives a sparsity bitmap 1015 that corresponds to the sparse tensor in the weight register file 1020. The sparsity bitmap 1015 may be a combined sparsity bitmap when the MAC unit 1005 operates in a combined sparsity mode. The sparsity bitmap 1015 may be an activation sparsity bitmap when the MAC unit 1005 operates in an activation sparsity mode. The sparsity bitmap 1015 may be a weight sparsity bitmap when the MAC unit 1005 operates in a weight sparsity mode. The sparsity bitmap 1015 may have the same size (e.g., the same number of elements) as or a larger size than the activation operand or  the weight operand.

[0166] Using the sparsity bitmap 1015, the sparsity accelerator 1060 selects four activations from the activation register file 1010 and selects four weights from the weight register file 1020. The sparsity accelerator 1060 transmits the selected activations and weights to the multiplier 1030. These selected data elements correspond to the nonzero valued elements of the sparsity bitmap 1015. The four selected activations and the four selected weights may constitute four activation-weight pairs. The multiplier 1030 may compute a product based on each activation-weight pair and therefore, compute four products in total. The four products may be provided to the adder 1040. Even though FIG. 10 shows a single multiplier 1030, the MAC unit 1005 may include multiple multipliers that can perform multiple multiplication operations at the same time.

[0167] The adder 1040 accumulates the four products and computes a unit-level internal partial sum. The four unselected elements of the dense tensor are not processed to save power and time, which would not impact the value of the unit-level internal partial sum. For instance, when the dense tensor is a dense activation tensor, the weights corresponding to the unselected activations are zeros so the products of the unselected activations and the weights would all be zero and have no contribution to the unit-level internal partial sum or other partial sums computed by the sparse cell. Similarly, when the dense tensor is a dense weight tensor, the activations corresponding to the unselected weights are zeros so the products of the unselected weights and the activations would all be zero and have no contribution to the unit-level internal partial sum or other partial sums computed by the sparse cell. In other embodiments, the MAC unit 1005 may operate in a dense mode in which the sparsity bitmap 1015 is not used and the sparsity accelerator 1060 is inactive. The MAC unit 1005 may process all the activations in the activation operand and all the weights in the weight operand.

[0168] The unit-level internal partial sum may be stored in the output register file 1050. In some embodiments, the unit-level internal partial sum may be used multiple times. For instance, the activation operand may represent N data blocks in the input tensor of the convolution, where N is an integer greater than 1. Instead of processing all the N data blocks to compute N unit-level internal partial sums, the unit-level internal partial sum is computed once and used N times in the convolutional layers as N unit-level internal partial sums.

[0169] In some embodiments, the PE 1000 receives one or more PE-level internal partial sums from one or more other PEs. The adder 1040 or an accumulator (not shown in FIG. 10) can accumulate the one or more PE-level internal partial sums with the PE-level internal partial sum of the PE 1000 and store the result of the accumulation (i.e., a multi-PE internal partial sum) in the output register file 1050. The one or more other PEs may be in the same column as the PE 1000 in a sparse cell. The multi-unit internal partial sum may be a column-level internal partial sum. In some embodiments, the PE-level internal partial sum of the PE 1000 or the multi-unit internal partial sum may be sent to one or more other PEs for further accumulation.

[0170] FIG. 11 is a flowchart of a method 1100 of making an executable DNN, in accordance with various embodiments. The method 1100 may be performed by the DNN module 301 in FIG. 3. Although the method 1100 is described with reference to the flowchart illustrated in FIG. 11, many other methods for making executable DNNs may alternatively be used. For example, the order of execution of the steps in FIG. 11 may be changed. As another example, some of the steps may be changed, eliminated, or combined.

[0171] The DNN module 301 identifies 1110 a kernel length of a convolution in a neural network. The kernel length is a length of a weight tensor of the convolution.

[0172] The DNN module 301 identifies 1120 a stride length of the convolution. The stride length indicates a number of rows or columns by which the weight tensor moves across an input tensor of the convolution

[0173] The DNN module 301 performs 1130 a comparison of the kernel length and the stride length. In some embodiments, the DNN module 301 determines whether the kernel length is greater than the stride length.

[0174] The DNN module 301 determines 1140 a reshaping factor for the convolution based on the comparison. In some embodiments, in response to determining that the kernel length is not greater than the stride length, the DNN module 301 uses the stride length as the reshaping factor. In other embodiments, the DNN module 301 computes a preliminary factor from the kernel length and the stride length. In response to determining that the kernel length is greater than the stride length, the DNN module 301 determines the reshaping factor from the preliminary factor and the stride length.

[0175] The DNN module 301 converts 1150 the convolution by reshaping the input tensor  based on the reshaping factor. In some embodiments, the DNN module 301 increases the number of input channels of the convolution. In some embodiments, the DNN module 301 reshapes the weight tensor by adding one or more zeros into the weight tensor. In some embodiments, the DNN module 301 increases the number of output channels of the convolution. In some embodiments, the DNN module 301 converts the weight tensor for a single output channel of the convolution to weight tensors for multiple output channels of the converted convolution.

[0176] The DNN module 301 causes 1160 execution of the neural network. The execution of the neural network includes performing the converted convolution. In some embodiments, the DNN module 301 converts the convolution (e.g., increases the number of input channels of the convolution) based on a configuration of a hardware device that executes the neural network. In some embodiments, the hardware devices may be the DNN accelerator 302 or part of the DNN accelerator 302. In some embodiments the configuration of the hardware device may be the arrangement of PEs in the hardware device, such as the number of PEs in a row, the number of PEs in a column, and so on. In some embodiments, the increased number of input channels is a power of two.

[0177] Example Computing Device

[0178] FIG. 12 is a block diagram of an example computing device 2000, in accordance with various embodiments. In some embodiments, the computing device 2000 can be used as at least part of the DNN system 300. A number of components are illustrated in FIG. 12 as included in the computing device 2000, but any one or more of these components may be omitted or duplicated, as suitable for the application. In some embodiments, some or all of the components included in the computing device 2000 may be attached to one or more motherboards. In some embodiments, some or all of these components are fabricated onto a single system on a chip (SoC) die. Additionally, in various embodiments, the computing device 2000 may not include one or more of the components illustrated in FIG. 12, but the computing device 2000 may include interface circuitry for coupling to the one or more components. For example, the computing device 2000 may not include a display device 2006, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which a display device 2006 may be coupled. In another set of examples, the computing device 2000 may not include an audio input device 2018 or an audio output  device 2008 but may include audio input or output device interface circuitry to which an audio input device 2018 or audio output device 2008 may be coupled.

[0179] The computing device 2000 may include a processing device 2002 (e.g., one or more processing devices) . The processing device 2002 processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. The computing device 2000 may include a memory 2004, which may itself include one or more memory devices such as volatile memory (e.g., DRAM) , nonvolatile memory (e.g., read-only memory (ROM) ) , high bandwidth memory (HBM) , flash memory, solid state memory, and / or a hard drive. In some embodiments, the memory 2004 may include memory that shares a die with the processing device 2002. In some embodiments, the memory 2004 includes one or more non-transitory computer-readable media storing instructions executable to perform operations for making executable DNNs (e.g., the method 1100 described in conjunction with FIG. 11) or some operations performed by one or more components of the DNN system 300 (such as the DNN module 301) . The instructions stored in the one or more non-transitory computer-readable media may be executed by the processing device 2002.

[0180] In some embodiments, the computing device 2000 may include a communication chip 2012 (e.g., one or more communication chips) . For example, the communication chip 2012 may be configured for managing wireless communications for the transfer of data to and from the computing device 2000. The term "wireless" and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data through the use of modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not.

[0181] The communication chip 2012 may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.10 family) , IEEE 802.16 standards (e.g., IEEE 802.16-2005 Amendment) , Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as "3GPP2" ) , etc. ) . IEEE 802.16 compatible Broadband Wireless Access (BWA) networks are generally referred to as WiMAX networks, an acronym that stands for  worldwide interoperability for microwave access, which is a certification mark for products that pass conformity and interoperability tests for the IEEE 802.16 standards. The communication chip 2012 may operate in accordance with a Global System for Mobile Communication (GSM) , General Packet Radio Service (GPRS) , Universal Mobile Telecommunications System (UMTS) , High Speed Packet Access (HSPA) , Evolved HSPA (E-HSPA) , or LTE network. The communication chip 2012 may operate in accordance with Enhanced Data for GSM Evolution (EDGE) , GSM EDGE Radio Access Network (GERAN) , Universal Terrestrial Radio Access Network (UTRAN) , or Evolved UTRAN (E-UTRAN) . The communication chip 2012 may operate in accordance with Code-division Multiple Access (CDMA) , Time Division Multiple Access (TDMA) , Digital Enhanced Cordless Telecommunications (DECT) , Evolution-Data Optimized (EV-DO) , and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. The communication chip 2012 may operate in accordance with other wireless protocols in other embodiments. The computing device 2000 may include an antenna 2022 to facilitate wireless communications and / or to receive other wireless communications (such as AM or FM radio transmissions) .

[0182] In some embodiments, the communication chip 2012 may manage wired communications, such as electrical, optical, or any other suitable communication protocols (e.g., the Ethernet) . As noted above, the communication chip 2012 may include multiple communication chips. For instance, a first communication chip 2012 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second communication chip 2012 may be dedicated to longer-range wireless communications such as global positioning system (GPS) , EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some embodiments, a first communication chip 2012 may be dedicated to wireless communications, and a second communication chip 2012 may be dedicated to wired communications.

[0183] The computing device 2000 may include battery / power circuitry 2014. The battery / power circuitry 2014 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 2000 to an energy source separate from the computing device 2000 (e.g., AC line power) .

[0184] The computing device 2000 may include a display device 2006 (or corresponding  interface circuitry, as discussed above) . The display device 2006 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD) , a light-emitting diode display, or a flat panel display, for example.

[0185] The computing device 2000 may include an audio output device 2008 (or corresponding interface circuitry, as discussed above) . The audio output device 2008 may include any device that generates an audible indicator, such as speakers, headsets, or earbuds, for example.

[0186] The computing device 2000 may include an audio input device 2018 (or corresponding interface circuitry, as discussed above) . The audio input device 2018 may include any device that generates a signal representative of a sound, such as microphones, microphone arrays, or digital instruments (e.g., instruments having a musical instrument digital interface (MIDI) output) .

[0187] The computing device 2000 may include a GPS device 2016 (or corresponding interface circuitry, as discussed above) . The GPS device 2016 may be in communication with a satellite-based system and may receive a location of the computing device 2000, as known in the art.

[0188] The computing device 2000 may include another output device 2010 (or corresponding interface circuitry, as discussed above) . Examples of the other output device 2010 may include an audio codec, a video codec, a printer, a wired or wireless transmitter for providing information to other devices, or an additional storage device.

[0189] The computing device 2000 may include another input device 2020 (or corresponding interface circuitry, as discussed above) . Examples of the other input device 2020 may include an accelerometer, a gyroscope, a compass, an image capture device, a keyboard, a cursor control device such as a mouse, a stylus, a touchpad, a bar code reader, a Quick Response (QR) code reader, any sensor, or a radio frequency identification (RFID) reader.

[0190] The computing device 2000 may have any desired form factor, such as a handheld or mobile computer system (e.g., a cell phone, a smart phone, a mobile internet device, a music player, a tablet computer, a laptop computer, a netbook computer, an ultrabook computer, a personal digital assistant (PDA) , an ultramobile personal computer, etc. ) , a desktop computer system, a server or other networked computing component, a printer, a scanner,  a monitor, a set-top box, an entertainment control unit, a vehicle control unit, a digital camera, a digital video recorder, or a wearable computer system. In some embodiments, the computing device 2000 may be any other electronic device that processes data.

[0191] Select Examples

[0192] The following paragraphs provide various examples of the embodiments disclosed herein.

[0193] Example 1 provides a method, the method including identifying a kernel length of a convolution in a neural network, the kernel length being a length of a weight tensor of the convolution; identifying a stride length of the convolution, the stride length indicating a number of rows or columns by which the weight tensor moves across an input tensor of the convolution; performing a comparison of the kernel length and the stride length;

[0194] determining a reshaping factor for the convolution based on the comparison; converting the convolution by reshaping the input tensor based on the reshaping factor; and causing execution of the neural network, the execution of the neural network comprising performing the converted convolution.

[0195] Example 2 provides the method of example 1, in which performing the comparison of the kernel length and the stride length includes determining whether the kernel length is greater than the stride length.

[0196] Example 3 provides the method of example 2, in which determining the reshaping factor includes in response to determining that the kernel length is not greater than the stride length, using the stride length as the reshaping factor.

[0197] Example 4 provides the method of example 2, in which determining the reshaping factor includes computing a preliminary factor from the kernel length and the stride length; and in response to determining that the kernel length is greater than the stride length, determining the reshaping factor from the preliminary factor and the stride length.

[0198] Example 5 provides the method of any one of examples 1-4, in which converting the convolution includes increasing a number of input channels of the convolution.

[0199] Example 6 provides the method of example 5, in which increasing the number of input channels of the convolution includes increasing the number of input channels of the convolution based on a configuration of a hardware device that executes the neural network.

[0200] Example 7 provides the method of example 5 or 6, in which the increased number of input channels is a power of two.

[0201] Example 8 provides the method of any one of examples 1-7, in which converting the convolution includes reshaping the input tensor by reducing a length of the input tensor.

[0202] Example 9 provides the method of any one of examples 1-8, in which converting the convolution includes converting the weight tensor by adding one or more zeros into the weight tensor.

[0203] Example 10 provides the method of any one of examples 1-9, in which converting the convolution includes increasing a number of output channels of the convolution.

[0204] Example 11 provides one or more non-transitory computer-readable media storing instructions executable to perform operations, the operations including identifying a kernel length of a convolution in a neural network, the kernel length being a length of a weight tensor of the convolution; identifying a stride length of the convolution, the stride length indicating a number of rows or columns by which the weight tensor moves across an input tensor of the convolution; performing a comparison of the kernel length and the stride length; determining a reshaping factor for the convolution based on the comparison; converting the convolution by reshaping the input tensor based on the reshaping factor; and causing execution of the neural network, the execution of the neural network comprising performing the converted convolution.

[0205] Example 12 provides the one or more non-transitory computer-readable media of example 11, in which performing the comparison of the kernel length and the stride length includes determining whether the kernel length is greater than the stride length.

[0206] Example 13 provides the one or more non-transitory computer-readable media of example 12, in which determining the reshaping factor includes in response to determining that the kernel length is not greater than the stride length, using the stride length as the reshaping factor.

[0207] Example 14 provides the one or more non-transitory computer-readable media of example 12, in which determining the reshaping factor includes computing a preliminary factor from the kernel length and the stride length; and in response to determining that the kernel length is greater than the stride length, determining the reshaping factor from the preliminary factor and the stride length.

[0208] Example 15 provides the one or more non-transitory computer-readable media of any one of examples 11-14, in which converting the convolution includes increasing a number of input channels of the convolution.

[0209] Example 16 provides the one or more non-transitory computer-readable media of example 15, in which increasing the number of input channels of the convolution includes increasing the number of input channels of the convolution based on a configuration of a hardware device that executes the neural network.

[0210] Example 17 provides the one or more non-transitory computer-readable media of example 15 or 16, in which the increased number of input channels is a power of two.

[0211] Example 18 provides the one or more non-transitory computer-readable media of any one of examples 11-17, in which converting the convolution includes reshaping the input tensor by reducing a length of the input tensor.

[0212] Example 19 provides the one or more non-transitory computer-readable media of any one of examples 11-18, in which converting the convolution includes converting the weight tensor by adding one or more zeros into the weight tensor.

[0213] Example 20 provides the one or more non-transitory computer-readable media of any one of examples 11-19, in which converting the convolution includes increasing a number of output channels of the convolution.

[0214] Example 21 provides an apparatus, including a computer processor for executing computer program instructions; and a non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations including identifying a kernel length of a convolution in a neural network, the kernel length being a length of a weight tensor of the convolution, identifying a stride length of the convolution, the stride length indicating a number of rows or columns by which the weight tensor moves across an input tensor of the convolution, performing a comparison of the kernel length and the stride length, determining a reshaping factor for the convolution based on the comparison, converting the convolution by reshaping the input tensor based on the reshaping factor, and causing execution of the neural network, the execution of the neural network comprising performing the converted convolution.

[0215] Example 23 provides the apparatus of example 22, in which determining the reshaping factor includes in response to determining that the kernel length is not greater  than the stride length, using the stride length as the reshaping factor.

[0216] Example 24 provides the apparatus of example 22, in which determining the reshaping factor includes computing a preliminary factor from the kernel length and the stride length; and in response to determining that the kernel length is greater than the stride length, determining the reshaping factor from the preliminary factor and the stride length.

[0217] Example 25 provides the apparatus of any one of examples 21-24, in which converting the convolution includes increasing a number of input channels of the convolution based on a configuration of a hardware device that executes the neural network.

[0218] The above description of illustrated implementations of the disclosure, including what is described in the Abstract, is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. While specific implementations of, and examples for, the disclosure are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the disclosure, as those skilled in the relevant art will recognize. These modifications may be made to the disclosure in light of the above detailed description.

Claims

1.A method, comprising:identifying a kernel length of a convolution in a neural network, the kernel length being a length of a weight tensor of the convolution;identifying a stride length of the convolution, the stride length indicating a number of rows or columns by which the weight tensor moves across an input tensor of the convolution;performing a comparison of the kernel length and the stride length;determining a reshaping factor for the convolution based on the comparison;converting the convolution by reshaping the input tensor based on the reshaping factor; andcausing execution of the neural network, the execution of the neural network comprising performing the converted convolution.2.The method of claim 1, wherein performing the comparison of the kernel length and the stride length comprises:determining whether the kernel length is greater than the stride length.3.The method of claim 2, wherein determining the reshaping factor comprises:in response to determining that the kernel length is not greater than the stride length, using the stride length as the reshaping factor.4.The method of claim 2, wherein determining the reshaping factor comprises:computing a preliminary factor from the kernel length and the stride length; andin response to determining that the kernel length is greater than the stride length, determining the reshaping factor from the preliminary factor and the stride length.5.The method of any one of claims 1-4, wherein converting the convolution comprises:increasing a number of input channels of the convolution.6.The method of claim 5, wherein increasing the number of input channels of the convolution comprises:increasing the number of input channels of the convolution based on a configuration of a hardware device that executes the neural network.7.The method of claim 5 or 6, wherein the increased number of input channels is a power of two.8.The method of any one of claims 1-7, wherein converting the convolution comprises:reshaping the input tensor by reducing a length of the input tensor.9.The method of any one of claims 1-8, wherein converting the convolution comprises:converting the weight tensor by adding one or more zeros into the weight tensor.10.The method of any one of claims 1-9, wherein converting the convolution comprises:increasing a number of output channels of the convolution.11.One or more non-transitory computer-readable media storing instructions executable to perform operations, the operations comprising:identifying a kernel length of a convolution in a neural network, the kernel length being a length of a weight tensor of the convolution;identifying a stride length of the convolution, the stride length indicating a number of rows or columns by which the weight tensor moves across an input tensor of the convolution;performing a comparison of the kernel length and the stride length;determining a reshaping factor for the convolution based on the comparison;converting the convolution by reshaping the input tensor based on the reshaping factor; andcausing execution of the neural network, the execution of the neural network comprising performing the converted convolution.12.The one or more non-transitory computer-readable media of claim 11, wherein performing the comparison of the kernel length and the stride length comprises:determining whether the kernel length is greater than the stride length.13.The one or more non-transitory computer-readable media of claim 12, wherein determining the reshaping factor comprises:in response to determining that the kernel length is not greater than the stride length, using the stride length as the reshaping factor.14.The one or more non-transitory computer-readable media of claim 12, wherein determining the reshaping factor comprises:computing a preliminary factor from the kernel length and the stride length; andin response to determining that the kernel length is greater than the stride length, determining the reshaping factor from the preliminary factor and the stride length.15.The one or more non-transitory computer-readable media of any one of claims 11-14, wherein converting the convolution comprises:increasing a number of input channels of the convolution.16.The one or more non-transitory computer-readable media of claim 15, wherein increasing the number of input channels of the convolution comprises:increasing the number of input channels of the convolution based on a configuration of a hardware device that executes the neural network.17.The one or more non-transitory computer-readable media of claim 15 or 16, wherein the increased number of input channels is a power of two.18.The one or more non-transitory computer-readable media of any one of claims 11-17, wherein converting the convolution comprises:reshaping the input tensor by reducing a length of the input tensor.19.The one or more non-transitory computer-readable media of any one of claims 11-18, wherein converting the convolution comprises:converting the weight tensor by adding one or more zeros into the weight tensor.20.The one or more non-transitory computer-readable media of any one of claims 11-19, wherein converting the convolution comprises:increasing a number of output channels of the convolution.21.An apparatus, comprising:a computer processor for executing computer program instructions; anda non-transitory computer-readable memory storing computer program instructions executable by the computer processor to perform operations comprising:identifying a kernel length of a convolution in a neural network, the kernel length being a length of a weight tensor of the convolution,identifying a stride length of the convolution, the stride length indicating a number of rows or columns by which the weight tensor moves across an input tensor of the convolution,performing a comparison of the kernel length and the stride length,determining a reshaping factor for the convolution based on the comparison,converting the convolution by reshaping the input tensor based on the reshaping factor, andcausing execution of the neural network, the execution of the neural network comprising performing the converted convolution.22.The apparatus of claim 21, wherein performing the comparison of the kernel length and the stride length comprises:determining whether the kernel length is greater than the stride length.23.The apparatus of claim 22, wherein determining the reshaping factor comprises:in response to determining that the kernel length is not greater than the stride length, using the stride length as the reshaping factor.24.The apparatus of claim 22, wherein determining the reshaping factor comprises:computing a preliminary factor from the kernel length and the stride length; andin response to determining that the kernel length is greater than the stride length, determining the reshaping factor from the preliminary factor and the stride length.25.The apparatus of any one of claims 21-24, wherein converting the convolution comprises:increasing a number of input channels of the convolution based on a configuration of a hardware device that executes the neural network.

Citation Information

Patent Citations

  • Learner high-order cognitive activity state identification method, device and system

    CN115630272A

  • Optimized neural network input stride method and apparatus

    US20220237461A1

  • Sparse machine learning acceleration

    US20220318604A1

  • Deep neural network (DNN) accelerators with weight layout rearrangement

    US20230017662A1