Dilated Convolution Using Systolic Arrays
By inserting zeros into the weight data matrix to reduce the number of strides of the convolution operation and optimizing zero-correlation operations in the hardware accelerator, the problem of high computational cost of convolution operations is solved and more efficient calculations are achieved.
Patent Information
- Application Number
- CN202080045813.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-28
- Filing Date
- 2020-06-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-06-26
AI Technical Summary
The prior art is costly to perform convolution operations, especially large image processing, and although the expansion convolution operations can reduce the total number of arithmetic operations, it may lead to reduced efficiency by increasing memory operations.
The coverage area of the weighted data matrix is amplified by inserting multiple zeros between elements of the weighted data matrix, thereby reducing the number of strides required to complete the input data matrix traversal and bypassing multiplication and addition operations involving zeros in the hardware accelerator.
The computational cost of convolution operations is effectively reduced, and the efficiency of expansion convolution operations is improved by reducing memory operations.
Smart Images

Figure CN114026569B_ABST
Abstract
Description
Background Art
[0001] Artificial neural networks are computing systems with an architecture based on biological neural networks. Artificial neural networks can be trained using training data to learn how to perform a computing task for an application.
[0002] Hardware accelerators such as neural network processors can be programmed to implement artificial neural networks to perform computing tasks. A common computing task is the convolution operation between a weight data matrix and an input data matrix. In a convolution operation, the weight data matrix can traverse the input data matrix in multiple strides and overlap with the input data matrix until the entire input data matrix has been traversed. For each stride, the sum of the multiplications between the superimposed portions of the weight data matrix and the input data matrix can be generated as the output of the convolution operation, and multiple outputs of the convolution operation can be generated at multiple strides. Convolution operations have many applications, such as extracting features from images, performing image recognition, etc.
[0003] In order to reduce the computational cost of the convolution operation between the weight data matrix and the large input data matrix (e.g., an image with a large number of pixels), an expansion convolution operation can be performed. In the expansion convolution operation, the coverage area of the weight data matrix can be enlarged by inserting multiple zeros between the elements of the weight data matrix. In the case of the enlarged weight data matrix, the number of strides required to complete the traversal of the input data matrix can be reduced. If the hardware accelerator also includes hardware to bypass multiplication and addition operations involving zeros, the total number of arithmetic operations can also be reduced. However, the reduction in arithmetic operations may be offset by additional delays introduced by other operations associated with the expansion convolution operation (e.g., memory operations), which may reduce the efficiency improvement caused by the expansion convolution operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Various embodiments according to the present disclosure will be described with reference to the drawings, in which:
[0005] Figure 1 An example of a classifier device that processes data using the techniques disclosed herein is shown;
[0006] Figure 2A-2E is a simplified block diagram illustrating predictive models and calculations using the techniques disclosed herein, according to certain aspects of the present disclosure;
[0007] Figure 3 shows an example implementation of a dilated convolution;
[0008] Figures 4A-4C illustrates an example neural network processor and its operation according to certain aspects of the present disclosure;
[0009] Figures 5A-5DFIG. 1 shows a method for performing a normal convolution operation according to certain aspects of the present disclosure. Figures 4A-4C Operations at an example neural network processor;
[0010] Figures 6A-6C FIG. 1 shows a method for performing a dilated convolution operation according to certain aspects of the present disclosure. Figures 4A-4C Operations at an example neural network processor;
[0011] Figure 7 includes a block diagram illustrating an example of a host system according to certain aspects of the present disclosure;
[0012] Figure 8 An example method of performing a dilated convolution operation at a neural network processor according to certain aspects of the present disclosure is shown;
[0013] Fig. 9 An example method of preparing instructions for a neural network processor to perform a dilated convolution operation according to certain aspects of the present disclosure is shown; and
[0014] Fig.10 Includes a diagram of an example network. DETAILED DESCRIPTION
[0015] Examples of the present disclosure relate to neural network processing, and more particularly, to performing dilated convolution operations at a neural network processor, and to hardware and / or software systems that support dilated convolution operations at a neural network processor.
[0016] Hardware accelerators such as neural network processors can be programmed to implement artificial neural networks to perform computing tasks. A common computing task is a convolution operation between a weight data matrix configured as a filter and an input data matrix. The input data matrix may correspond to pixels of an image, and the filter may include a filter coefficient array configured to extract target features from the image. As part of the convolution operation, the filter may traverse different positions of the image in multiple strides. At each stride position, the sum of the products between each filter coefficient and the overlapping pixels of the image may be generated as the convolution output of the stride position. The convolution output may indicate, for example, whether the image contains the target feature, the image position of the target feature, etc. Since the convolution operation requires a large number of arithmetic operations (e.g., multiplication and addition operations) to be performed at each stride position, the convolution operation can be quite computer intensive, especially when it comes to large images with a large number of pixels.
[0017] In order to reduce the computational cost of the convolution operation, an expansion convolution operation can be performed. In the expansion convolution operation, the coverage area of the weight data matrix can be enlarged by inserting multiple zeros between the elements of the weight data matrix. The number of zeros can be based on the rate of the expansion convolution. For example, when the rate is two, one zero is inserted between the elements of the weight data matrix, and when the rate is four, three zeros are inserted between the elements of the weight data matrix. In the case of an enlarged weight data matrix, the number of strides required to complete the traversal of the input data matrix can be reduced. If the hardware accelerator also includes hardware to bypass multiplication and addition involving zeros, the number of arithmetic operations at each stride can be the same or similar to a normal (non-expanded) convolution. As the number of strides decreases, the total number of arithmetic operations involving the expansion convolution operation can be reduced, and the efficiency of the expansion convolution operation can be improved.
[0018] Although the dilated convolution operation can reduce the total number of arithmetic operations, the reduction in computational cost may be offset by other operations such as memory operations associated with the dilated convolution operation. As an example, in order to perform the dilated convolution, groups of elements of an image superimposed with non-zero elements of an extended weight data matrix of different step positions in the dilated convolution operation can be read from a memory and stored at different locations in the memory. Each group of elements can be read from the memory and multiplied with the non-zero elements of the extended weight data matrix to produce a product, and the product can be summed to produce an output. The output can be rearranged in the memory to construct an output data array. Additional memory read / write operations can add latency to the dilated convolution operation and offset the efficiency improvement caused by the dilated convolution operation.
[0019] In some examples of the present disclosure, a neural network processor includes a memory, a systolic array, and a controller. The memory may store input data elements of an input data array and weight data elements of a weight data array. Both the input data array and the weight data array may be multidimensional. For example, the input data array may include one or more two-dimensional input data matrices, each of which corresponds to an input channel. In addition, the weight data array may include one or more two-dimensional weight data matrices, each of which corresponds to an input channel and an output channel. The input data elements may be stored in addresses in the memory based on their coordinates in the input data array, and the weight data elements may be stored in addresses in the memory based on their coordinates in the weight data array.
[0020] To perform a dilated convolution operation, the controller may obtain a first weight data element from a memory and load it into a systolic array, and then load a first subset of input data elements into the systolic array to be multiplied with the first weight data element. The first subset of input data elements loaded into the systolic array may represent input data elements that overlap with the first weight data element when the weight data array is at a different step position on the input data array. The systolic array may repeat these operations for other weight data elements to generate other partial sums, and control the summing buffer to accumulate the partial sums to generate an output data element of the dilated convolution. The summing buffer may then store the output data elements back to the memory to construct an output data array of the dilated convolution.
[0021] As described above, each subset of input data elements may be selected based on determining the input data elements that overlap with the weight data elements in the dilated convolution operation ("overlapping input data elements"). The determination of the overlapping input data elements of the weight data elements may be performed by a compiler, which may encode the overlapping input data element information in the instruction. The controller may then execute the instruction to perform the selection. The compiler may determine the overlapping input data elements based on the projection operation. The size of the summing buffer (e.g., the number of columns and rows) may define an output tile (tile) of output data elements including the first zone in the output data array. The first zone may be defined by the range of actual coordinates in the output data array. Based on the projection operation of the first zone of the output data array coordinates and the stride of the dilated convolution, the compiler may determine a second zone including the input data elements to be convolved with the first weight data element. The second zone may be defined by the range of target coordinates of the input data elements. The second zone (and the range of target coordinates) may be shifted by an offset based on the coordinates of the first weight data element in the weight data array, and by a proportional factor based on the rate of the dilated convolution operation. The compiler may then align the stride pattern with the shifted second region to identify the location of the second region that overlaps with the stride pattern. The stride pattern defines the location of the input data elements that overlap with the weight data elements and reflects the stride of the dilated convolution operation. Based on the overlap, the compiler may determine a set of target coordinates for the overlapping input data elements. Projection operations may also be performed for different input channels to identify overlapping input data elements in other input channels.
[0022] After determining a set of target coordinates of overlapping input data elements of the first weight data element, the compiler may perform a coordinate-to-address mapping operation to determine the addresses of the first subset of input data elements in the memory, and provide the addresses in the instructions to allow the controller to obtain the first subset of input data elements from the memory. The compiler may perform the aforementioned projection operation for each weight data element to generate instructions for the corresponding weight data element, wherein each instruction includes the address of the subset of input data elements to be convolved with the corresponding weight data element.
[0023] Examples of the present disclosure can improve the efficiency of a neural network processor when performing an expansion convolution operation by reducing memory operations. For example, compared with a case where overlapping input data elements are read from a memory and then stored at different locations in the memory, and then read again from different locations in the memory to perform an expansion convolution, a neural network processor according to an example of the present disclosure has addresses of overlapping input data elements and uses the addresses to selectively read overlapping input data elements from the memory, and extract the input data elements to a systolic array. The neural network processor does not need to perform additional memory write operations to write overlapping input data elements back to the memory. In addition, since the summing buffer can store the output data elements at predetermined locations in the memory to reconstruct the output data array, it is not necessary to rearrange the output data elements in the memory to construct the output data array. All of this can reduce the number of memory operations, thereby reducing memory access latency and increasing the speed of the expansion convolution operation.
[0024] In the following description, various examples will be described. For the purpose of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the examples. However, it is apparent to those skilled in the art that the examples can be practiced without the specific details. In addition, well-known features may be omitted or simplified to avoid obscuring the described embodiments.
[0025] Figure 1 An example classifier device 100 is shown for processing data using the techniques disclosed herein. The classifier device 100 can be, for example, a computing device that operates a software application 102 and a prediction model 103 to predict information included in a data sequence and performs a predetermined function based on the prediction. For example, the classifier device 100 can be part of an image recognition service that is configured to identify certain objects (e.g., text, people, etc.) from an image. It should be understood that the image recognition service is provided only as an illustrative example, and the techniques disclosed herein can be used for other data processing applications, including, for example, text-based data processing (e.g., processing of search queries), audio data processing, etc. In addition, the classifier device 100 can operate multiple different prediction models to process different input data simultaneously or at different times.
[0026] In some instances, the image recognition service may be provided in a multi-tenant computing service system. A multi-tenant computing service system may typically include multiple servers that can host data and can be used by multiple clients or organizations to run instances, such as virtual machine instances or bare metal instances (e.g., operating systems running directly on server hardware). In most cases, multi-tenant computing service systems such as bare metal or virtual machine instances can be allocated to clients when they need them and exit when they are no longer needed so that resources can be reallocated to other clients. In this disclosure, the terms "tenant," "client," and "customer" are used interchangeably, but such terms do not necessarily imply the existence of any particular business arrangement. The term "instance" may refer to, for example, an instance running directly on server hardware or as a virtual machine. Different types of instances typically correspond to different hardware capabilities and / or hardware arrangements (e.g., different amounts of available memory and / or processing hardware). In Figure 1 In an example, the multi-tenant computing service system can provide image recognition services when the client needs them, and the services can be deactivated when they are no longer needed, so that the resources supporting the image recognition services (e.g., access to the software application 102 and the underlying hardware resources for processing the software application 102) can be reallocated to other clients. Different clients (or one client) can request the application 102 to perform processing of different input data using the same or different prediction models including the prediction model 103.
[0027] exist Figure 1 In an example of, software application 102 may receive pixel data of image 104 from a user. Image 104 may include an array of pixels. Software application 102 may perform analysis on the pixel data and predict one or more objects 106 depicted in image 104. The analysis may include, for example, comparing the pixel data with a set of predetermined feature data. The predetermined feature data may include data associated with a set of predetermined visual image features, such as a nose object, a mouth object, etc. The predetermined feature data may also include data associated with non-visual image features or a combination of visual and non-visual image features. As will be discussed in more detail below, software application 102 may employ prediction model 103 to calculate a set of scores based on pixel data of image 104. The set of scores may represent, for example, the probability that image 104 includes image features represented by feature data. Software application 102 may then determine other information about the content of image 104 based on the scores. For example, based on the scores, software application 102 may determine that image 104 is an image of, for example, a panda, a cat, or other object. The present disclosure provides an example of a technology that allows a trade-off between speed and accuracy of operating prediction model 103, as will be discussed below.
[0028] The prediction model 103 may be in the form of an artificial neural network. The artificial neural network may include a plurality of processing nodes, wherein each processing node is configured to process a portion of the input pixel data, or to further process intermediate outputs from other processing nodes. Figure 1 An example of a prediction model 103 using the techniques disclosed herein is shown. Figure 1 In the example, the prediction model 103 may be a multi-layer neural network, such as a deep neural network (DNN), a convolutional neural network (CNN), etc. The prediction model 103 may include an input layer 207, a set of intermediate layers including intermediate layers 209 and 211, and an output layer ( Figure 2A ). It should be understood that the prediction model 103 may also include other different types of neural networks, including, for example, long-term memory (LSTM), multi-layer perception (MTP), multi-scale dense network (MSDNET), etc.
[0029] Layer 207 may process pixel data representing different portions of image 104. For example, Figure 2A In the example of , layer 207 can process pixel data of image 204. Each processing node of layer 207 is designated to receive a pixel value (e.g., x) corresponding to a predetermined pixel within image 104. 0 、x 1 、x 2 ……x n ), and transmits one or more weights with the received pixel values to layer 209. In the case where the prediction model 203 is a DNN, a set of weights defined based on the matrix W1 may be specified for each processing node of layer 207. Each processing node of layer 207 may send the received pixel value and the specified weight to each processing node of layer 209. In the case where the prediction model 103 is a CNN, a group of processing nodes of layer 207 may share a set of weights, and each group may send a set of weights and pixel values received by the group of processing nodes to a single processing node of layer 209. Different neural network models may include different topologies (e.g., including different numbers of layers, different connections between layers, etc.), and / or include different sets of weights for each layer.
[0030] Layer 209 may process the scaled outputs from layer 207 to generate a set of intermediate outputs. For example, assuming that processing node 210a of layer 209 is connected to n processing nodes in layer 207, processing node 210a may generate the sum of the scaled outputs received from layer 207 based on the following equation:
[0031]
[0032] Here, the sum 210a W1 represents the intermediate output generated by processing node 210a. i ×i The processing nodes that pass through layer 207 have associated weights (e.g., W1 0 ) of a particular pixel value (e.g., x 0 ). In the case where the prediction model 103 is a DNN, each processing node of layer 209 may generate a sum based on the scaling of the pixel values from each processing node of layer 207, and then generate a sum by summing the scaled pixel values (e.g., the sum 210a ). The sum may also represent a dot product between an input vector including multiple elements (eg, pixel values) and a weight vector (eg, W1). In some examples, a bias may also be added to the scaled output to produce an intermediate output.
[0033] In the case where the prediction model 103 is a CNN, each processing node of layer 209 may generate an intermediate output based on the scaling of the pixel values from the processing node group of layer 207. The intermediate output may represent the convolution result between the pixel value group and the filter including the weight value. Figure 2B An example of a convolution operation that layer 209 may perform is shown. Figure 2B In the embodiment of the present invention, filter 230 may include a two-dimensional array of weights. The weights in filter 230 may represent the spatial distribution of pixels of certain features to be detected based on the image. The two-dimensional array may have a height of R rows and a width of S columns, and is typically smaller than the input image having a height of H pixels and a width of W pixels. Each weight may be mapped to a rectangular block of pixels having the same R rows and S columns of pixel values. A processing node of layer 209 (e.g., processing node 210a) may receive a group 240 of pixel values corresponding to a first rectangular block of pixels from the input image from a group of processing nodes of input layer 207, the group 240 of pixel values corresponding to the first step position of filter 230, and generate a convolution output 242 based on the sum of the multiplication results between each weight of filter 230 and each corresponding pixel in group 240 according to equation 1 to generate a dot product between the matrix represented by filter 230 and the matrix represented by group 240. Another processing node of layer 209 may also receive group 244 of pixel values from another group of processing nodes of input layer 207, the group 244 of pixel values corresponding to a second rectangular block of pixels corresponding to a second stride position of filter 230 from the input image, and generate convolution output 246 based on the sum of multiplication results between each weight of filter 230 and each corresponding pixel in group 244 to generate a dot product between the matrix of filter 230 and the matrix represented by group 240 according to Equation 1. In some examples, Figure 2BEach convolution output in (e.g., convolution output 242, convolution output 346, etc.) may correspond to the output of a processing node of layer 209. In some examples, the pixel data in the input image may be referred to as an input feature map to indicate that the pixels are processed by the same filter (or the same set of filters) corresponding to a feature. The convolution output may be referred to as an output feature map to indicate that the output is the result of processing the input feature map through the filter.
[0034] like Figure 2B As shown, the convolution operation can be arranged in a sliding window so that the second rectangular block overlaps with the first rectangular block in the input image, or is otherwise adjacent to the first rectangular block in the input image. Figure 2B In the example of , D can be the distance of the stride (in pixels) of the sliding window used for each convolution operation, so that the pixel block corresponding to group 244 can be located at a distance D (in pixels) from the pixel block corresponding to group 240, and the next block of pixels can also be located at the same distance D from group 244. Other processing nodes of layer 209 can also receive pixel groups corresponding to other rectangular blocks and generate other intermediate outputs. The convolution output can be part of the convolution output array. The convolution output array can have a smaller height and a smaller width than the input image. The rectangular blocks of convolution outputs can be further grouped, and a convolution operation can be performed at layer 211 between the convolution output groups and another set of filter weights to generate another set of convolution outputs.
[0035] In some examples, a convolution operation may be performed between multiple images and multiple filters. Figure 2C , a set of C filters 260 may correspond to the number (C) of images 270, and a convolution operation may be performed between each filter of the set of filters 260 and a pixel block on a corresponding image of the image 270. Each of the images 270 may correspond to an input channel. The convolution results of each filter image pair may be summed to produce the following convolution output:
[0036]
[0037] Here, the convolution operation involves an image (or pixel array). c eD+r,fD+s may refer to the pixel value at index c in the image, whose row coordinate is eD+r and column coordinate is fD+s, within the number (C) of the image 270. For the remainder of this disclosure, the element X may be represented in the form of (eD+r, fD+s). c eD+r,fD+sThe index c may represent a specific input channel. D is the sliding window stride distance, while e and f correspond to the position of the data element in the convolution output array, which may also correspond to a specific sliding window. In addition, r and s correspond to specific positions within the sliding window. The image of a pixel at position (r, s) and index c may also correspond to the weight W in the corresponding filter of the same index c at the same (r, s) position. c r,s Equation 2 indicates that in order to calculate the convolution output O e,f , each pixel in the sliding window (indexed by (e, f)) can be multiplied by the corresponding weight W c r,s The partial sum of the multiplication products within each sliding window in each image within the image set may be calculated. And then the sum of the partial sums of all images of the image set may be calculated.
[0038] In addition, in some examples, multiple sets of filters can be used to perform a convolution operation with a set of images to generate a set of convolution output arrays, where each convolution output array corresponds to a set of filters. Each set of filters can correspond to an output channel. For example, multiple sets of filters can correspond to multiple features to be detected from the set of images, and each convolution output array can correspond to a detection result for each feature from the set of images. For example, where M sets of filters are applied to C images to generate M convolution output arrays, Equation 2 can be updated as follows:
[0039]
[0040] Here, the convolution output O e,f m and weight W c,m r,s With an index m corresponding to one of the M sets of filters. The index m may represent a particular output channel.
[0041] Figure 2D An example of C sets of input data sets (where C=3) to be convolved with M sets of filters (where M=2) is shown. Each set of input data corresponds to an entry of a pixel array. Each of the M sets of filters includes a set of C filters corresponding to the C sets of input pixel arrays. The convolution operation produces M sets of output data elements, where each set of output data elements corresponds to a convolution output array. Each convolution output array corresponds to convolving a set of filters (in the M sets) with the input pixel array. For example, 0,0 0 It may be generated by the sum of the dot products between the group of pixels 282 and filter array 284 , the dot products between the group of pixels 286 and filter array 288 , and the dot products between the group of pixels 289 and filter array 292 .
[0042] Return to reference Figure 2A , a processing node of layer 209 may be configured to generate a convolution output element of a convolution output array, and the set M of processing nodes of layer 209 may correspond to the set M of convolution output arrays. The processing nodes of layer 209 may also process each convolution output with an activation function to generate an activation output. The activation function may convert the convolution output into a decision (similar to the firing of a biological neuron) as to whether to forward the convolution output to the intermediate layer 211 to affect the classifier decision. An example of an activation function may be a rectified linear unit (ReLU) defined according to the following equation:
[0043]
[0044] In addition to ReLU, other forms of activation functions can also be used, including, for example, a soft add function (which can be a smooth approximation of the ReLU function), a hyperbolic tangent function (tanh), an arc tangent function (arctan), a S-type function, a Gaussian function, etc.
[0045] A processing node of layer 209 (eg, processing node 210a) may process the sum with the ReLU function to generate a first output of layer 209 based on the following equation:
[0046] first_output 210a =ReLU(sum 210a ) (Equation 5)
[0047] Layer 211 may further process the scaled intermediate output from layer 209 by, for example, performing additional convolution operations based on different sets of filters. The output from each processing node of layer 211 may be forwarded to other higher intermediate layers, or to the output layer ( Figure 2A 104 and / or the image 204 includes an image of a panda. For example, the output vector may be compared with a reference vector associated with a nose object of a panda or a reference vector associated with a panda. A decision as to whether image 104 is an image of a panda may be determined based on the comparison result.
[0048] Figure 2B-Figure 2D The convolution operation requires multiple arithmetic operations (e.g., multiplications and additions) to be performed at each stride position to find the dot product. The convolution operation can be quite computer intensive, especially when large images with a large number of pixels are involved, because a large number of strides may be required to fully traverse the filter over the image.
[0049] In order to reduce the computational cost of the convolution operation, a dilated convolution operation may be performed. Figure 2E An example of a dilated convolution operation is shown. Figure 2E As shown in Figure 2C Inserting zeros between the elements of filter 260 may produce filter 290 from filter 290. For example, zeros may be inserted between weight elements W 0,0 With W 0,1 to expand the width of the first line to S E In addition, zeros are inserted in the weight elements W 0,0 With W 1,0 to extend the height of the first column to R E The number of zeros inserted between elements can be based on the rate of the dilated convolution. Figure 2E In the example of the dilated convolution, the rate is two. In the case of a dilated convolution with a rate of four, three zeros may be inserted between each element. A row of zeros is also inserted between the dilated rows to accommodate the insertion of zeros in the weight element W. 0,0 With W 1,0 A column of zeros is inserted between the extended columns to accommodate the insertion of zeros between pixels X 0,0 With X 0,1 between.
[0050] As part of the dilated convolution operation, filter 290 may traverse image 270 at different step positions, such as Figure 2B As shown. At each stride position, each element of filter 290 may be multiplied with the overlapping pixels of image 270 to produce a product, and the products may be summed to produce an output at convolution array 280. The stride position may determine the coordinates of the output in convolution array 280. As a result of the enlarged coverage area of filter 290, the number of pixels overlapping filter 290 increases, but every other pixel overlaps with zero and is not represented in the sum. For example, at Figure 2E In the stride position shown, the weight element W 0,0 The image 270 pixels x 0,0 Multiply, and the weight element W 0,1 The image 270 pixels x 0,2 Multiply, skipping X 0,1 Because the coverage area of filter 290 is larger, the number of strides required for filter 290 to complete a pass over image 270 can be reduced. If the hardware accelerator also includes hardware to bypass multiplications and additions involving zero, the number of arithmetic operations at each stride can be reduced compared to the number of arithmetic operations involving filter 260. Figure 2C As the number of strides decreases, the total number of arithmetic operations involved in the dilated convolution operation can be reduced.
[0051] Figure 3 An example sequence 300 for implementing a dilated convolution operation in a computing environment is shown. Figure 3In the example of , a dilated convolution operation with a rate of two and a stride of two may be performed between image 270 and filter 290, where filter 290 may be generated from filter 260 by inserting zeros between adjacent weight elements as described above. To perform the dilated convolution operation, the pixel data of image 270 may be split into subsets of pixel data that overlap with non-zero weight elements of filter 290 at each stride position. Figure 3 As shown, the pixel data of the image 270 may be split into an input data subset 302, an input data subset 304, and the like. Each of the input data subsets 302 and 304 includes pixels that overlap with the filter 290 at a plurality of stride positions. A convolution operation 310 may be performed between the original filter 290 (which is expanded to the filter 290 for the dilated convolution operation) and each subset of pixels to produce a subset of output elements of the convolution array 280, including output data subsets 306 and 308. The output elements of the output data subsets 306 and 308 may be interleaved and / or rearranged to assemble the convolution array 280 according to the stride positions represented by the input data subsets.
[0052] Sequence 300 may involve a large number of memory read and write operations. For example, splitting the pixels of image 270 into subsets of pixels may be performed by reading the pixels of image 270 from memory locations and storing the subsets of pixels at other memory locations. In addition, after the output data subsets 306 and 308 are generated from convolution operation 310, the output data subsets 306 and 308 may be stored in memory, and as part of the interleave / shuffle operation, the output data subsets may be read from memory and stored back in memory to assemble convolution array 280. All of these additional memory read and write operations may add latency and increase the time required to complete the dilated convolution operation in the computing environment.
[0053] Figure 4A 4 illustrates an example of an integrated circuit device that can be configured to perform various types of convolution operations including normal convolution operations and dilated convolution operations. The example of FIG. 4 illustrates an accelerator 402. In various examples, for a set of input data (e.g., input data 450), the accelerator 402 can perform calculations using a processing engine array 410, an activation engine 416, and / or a pooling engine 418. In some examples, the example accelerator 402 can be an integrated circuit component of a processor such as a neural network processor. The processor can have other integrated circuit components, including additional accelerator engines. The accelerator 402 may include a controller 422 to control the operation of the processing engine array 410, the activation engine 416, and / or the pooling engine 418.
[0054] In various embodiments, the memory subsystem 404 may include multiple memory groups 414. In these embodiments, each memory group 414 is independently accessible, which means that the reading of one memory group is not dependent on the reading of another memory group. Similarly, writing to one memory group does not affect or limit writing to a different memory group. In some cases, each memory group can be read and written at the same time. Various technologies can be used to have independently accessible memory groups 414. For example, each memory group can be a physically independent memory component, which has an address space that is separate and independent of the address space of each other memory group. In this example, each memory group can have at least one read channel and can have at least one separate write channel that can be used simultaneously. In these examples, the memory subsystem 404 can allow simultaneous access to the read or write channels of multiple memory groups. As another example, the memory subsystem 404 may include arbitration logic so that, for example, arbitration between the outputs of multiple memory groups 414 can cause the outputs of more than one memory group to be used. In these and other examples, although the overall is managed by the memory subsystem 404, each memory group can operate independently of any other memory group.
[0055] Making memory banks 414 independently accessible can increase the efficiency of accelerator 402. For example, values can be read and provided to each row of processing engine array 410 simultaneously so that the entire processing engine array 410 can be used in one clock cycle. As another example, memory banks 414 can be read while results calculated by processing engine array 410 are being written to memory subsystem 404. In contrast, a single memory may only be able to service one read or write at a time. In the case of a single memory, multiple clock cycles may be required, for example, to read input data for each row of processing engine array 410 before processing engine array 410 can be started.
[0056] In various embodiments, the memory subsystem 404 may be configured to simultaneously serve multiple clients, including the processing engine array 410, the activation engine 416, the pooling engine 418, and any external client accessing the memory subsystem 404 through the communication structure 420. In some embodiments, being able to serve multiple clients may mean that the memory subsystem 404 has at least as many memory banks as the clients. In some cases, each row of the processing engine array 410 may be counted as a separate client. In some cases, each column of the processing engine array 410 may output a result, so that each column may be counted as a separate write client. In some cases, the output from the processing engine array 410 may be written to the memory bank 414, which may then provide input data to the processing engine array 410. As another example, the activation engine 416 and the pooling engine 418 may include multiple execution channels, each of which may be a separate memory client. For example, the memory bank 414 may be implemented using a static random access memory (SRAM).
[0057] In various embodiments, the memory subsystem 404 may include control logic. For example, the control logic may record the address space of each of the memory banks 414, identify the memory banks 414 to be read or written, and / or move data between the memory banks 414. In some embodiments, the memory banks 414 may be hardwired to specific clients. For example, a group of memory banks 414 may be hardwired to provide values to the rows of the processing engine array 410, where one memory bank serves one row. As another example, a group of memory banks may be hardwired to receive values from the columns of the processing engine array 410, where one memory bank receives data for each column.
[0058] Processing engine array 410 is a computational matrix of example accelerator 402. For example, processing engine array 410 may perform parallel integration, convolution, correlation, and / or matrix multiplication, etc. Processing engine array 410 includes a plurality of processing engines 411 arranged in rows and columns so that a result output by one processing engine 411 may be directly input into another processing engine 411. Processing engines 411 that are not on the outer edge of processing engine array 410 may therefore receive data for operation from other processing engines 411 rather than from memory subsystem 404.
[0059] In various examples, the processing engine array 410 uses systolic execution, where data arrives at each processing engine 411 at regular intervals from different directions. In some examples, input data may flow into the processing engine array 410 from the left, and weight values may be loaded at the top. In some examples, weights and input data may flow from the left, and partial sums may flow from top to bottom. In these and other examples, the multiply-accumulate operation moves through the processing engine array 410 as a diagonal wave front, where data moves right and down on the array. Control signals may be input on the left at the same time as the weights, and may flow through and down along with the calculations.
[0060] In various embodiments, the number of columns in processing engine array 410 determines the computational capacity of processing engine array 410, and the number of rows determines the required memory bandwidth for achieving maximum utilization of processing engine array 410. Processing engine array 410 may have, for example, 64 columns and 428 rows, or some other number of columns and rows.
[0061] An example of a processing engine 411 is illustrated in FIG4 . As shown by this example, the processing engine 411 may include a multiplier-accumulator circuit. The input from the left may include, for example, input data i and weight values w, wherein the input data is a value taken from a set of input data or a set of intermediate results, and the weight values come from a set of weight values that connect a layer of a neural network to the next layer. A set of input data may, for example, be an image submitted for identification or object recognition, an audio clip provided for speech recognition, a string of text for natural language processing or machine translation, or a current state of a game that needs to be analyzed to determine the next move, etc. In some instances, the input data and weight values are output to the right for input to the next processing engine 411.
[0062] In the example shown, the input from above may include a partial sum p_in provided from another processing engine 411 or from a previous round of calculations performed by the processing engine array 410. When calculations are started for a new set of input data, the top row of the processing engine array 410 may receive a fixed value for p_in, such as zero. As shown in this example, i and w are multiplied together, and the result is summed with p_in to produce a new partial sum p_out that can be input into another processing engine 411. Various other implementations of the processing engine 411 are possible.
[0063] The output from the last row in the processing engine array 410 may be temporarily stored in the result buffer 412. The result may be an intermediate result that may be written to the memory bank 414 to be provided to the processing engine array 410 for additional calculations. Alternatively, the result may be a final result that once written to the memory bank 414 may be read from the memory subsystem 404 through the communication fabric 420 to be output by the system.
[0064] In some embodiments, accelerator 402 includes activation engine 416. In these embodiments, activation engine 416 can combine results from processing engine array 410 into one or more output activations. For example, for a convolutional neural network, convolutions from multiple channels can be summed to produce an output activation for a single channel. In other instances, it may be necessary to accumulate results from one or more columns in processing engine array 410 to produce an output activation for a single node in the neural network. In some instances, activation engine 416 can be bypassed.
[0065] In various examples, the activation engine 416 may include multiple separate execution channels. In these examples, the execution channels may correspond to columns of the processing engine array 410, and operations may be performed on the outputs of the columns, the results of which may be stored in the memory subsystem 404. In these examples, the activation engine 416 is capable of performing between 1 and n parallel calculations, where n is equal to the number of columns in the processing engine array 410. In some cases, one or more of the calculations may be performed simultaneously. Examples of calculations that each execution channel may perform include exponential, square, square root, identity, binary step, bipolar step, sigmoidal, and ramp, among other examples.
[0066] In some embodiments, accelerator 402 may include pooling engine 418. Pooling is a combination of the outputs of the columns of processing engine array 410. Combination may include, for example, calculating maximum, minimum, average, median, sum, multiplication or another logical or mathematical combination. In various examples, pooling engine 418 may include multiple execution channels, which may operate on the values of the corresponding columns from processing engine array 410. In these examples, pooling engine 418 is capable of performing parallel calculations between 1 and n, where n equals the number of columns in processing engine array 410. In various examples, the execution channels of pooling engine 418 may operate in parallel and / or simultaneously. In some examples, pooling engine 418 may be bypassed.
[0067] Here, activation engine 416 and pooling engine 418 may be collectively referred to as an execution engine. Processing engine array 410 is another example of an execution engine. Another example of an execution engine is a direct memory access (DMA) engine that may be located outside of accelerator 402.
[0068] Input data 450 may arrive at the communication structure 420. The communication structure 420 may connect the accelerator 402 to other components of the processor, such as a DMA engine, a storage driver, or a network interface that may obtain the input data 450 from an input / output (I / O) device. The input data 450 may be, for example, one-dimensional data, such as a string or a sequence of values, or may be two-dimensional data, such as an array of pixel values for an image or frequency and amplitude over time for an audio signal. In some instances, the input data 450 may be three-dimensional, such as the case for contextual information or virtual reality data used by an autonomous vehicle. In some embodiments, the memory subsystem 404 may include a separate buffer for the input data 450. In some embodiments, when the accelerator 402 receives the input data 450, the input data 450 may be stored in the memory group 414.
[0069] In some examples, the accelerator 402 may implement a neural network processing engine. In these examples, the accelerator 402 may run a neural network to perform a task for which the neural network has been trained for a set of input data 450. Executing a neural network on a set of input data may be referred to as inference or performing inference.
[0070] The weights of the neural network may be stored in the memory subsystem 404 along with the input data 450 on which the neural network will operate. The addresses of the weights and input data 450 in the memory subsystem 404 may be based on or mapped to the coordinates of the weights and input data 450 in the weight data array and the input data array, respectively, which allows the weights and input data to be retrieved based on addresses derived from their coordinates. The neural network may also include instructions that can be executed by the controller 422 to control the processing engine array 410 to perform various calculations on the weights and input data. The instructions may be generated by a compiler and may also be stored in the memory subsystem 404, in the memory bank 414, or in a separate instruction buffer. The processing engine array 410 may output intermediate results that represent the outputs of various layers of the neural network. In some cases, the activation engine 416 and / or the pooling engine 418 may be enabled for calculations called by certain layers of the neural network. The accelerator 402 may store the intermediate results in the memory subsystem 404 to be input into the processing engine array 410 to calculate the results of the next layer of the neural network. The processing engine array 410 may also output the final result from the last layer of the neural network. The final result may be stored in the memory subsystem 404 and then copied to the host processor memory or to another location.
[0071] Figure 4B and Figure 4C An example of the operation of the accelerator 402 is shown. Figure 4BAs shown, the memory subsystem 404 can be organized into multiple rows, such as memory rows 425, 426, etc. Each memory row can store input data elements of a specific input channel. The memory access circuit (e.g., memory access circuit 427) can be controlled to sequentially extract the input data elements to the processing engine array 410 based on a set of memory extraction parameters 430 including a starting address, a step size, and a number of elements. The starting address parameter can define the location of the first input data element to be read from the memory row, the step size parameter can define the number of input data elements skipped between the extracted input data elements, and the number of extracted elements parameter can define the total number of input data elements to be extracted. When the input data elements are stored in a continuous space, the access circuit 427 can determine the address of the input data element to be obtained, and update the counter based on the step size. For example, the access circuit 427 can start to obtain the first input data element from the starting address, add the address offset based on the step size to the starting address to obtain the next input data element when skipping multiple input data elements, and repeat until the number of elements to be obtained is reached. As will be described in more detail below, the memory fetch parameters 430 may be included in instructions for calculating a set of partial sums. The instructions may be generated by a compiler and parsed by the controller 422 to extract the memory fetch parameters 430. The controller 422 may then control the fetching of input data elements from the memory subsystem 400 based on the fetched memory fetch parameters 430. As will be described in more detail below, the starting address, stride, and number of element parameters may be configured to support different types of convolution operations, such as normal convolution operations, dilated convolution operations, etc.
[0072] The processing engines 411 of the processing engine array 410 may be organized into rows, such as row 431, and columns, such as column 432. Each row of the processing engines 411 is mapped to an input channel and may sequentially receive input data elements from a memory row of the memory system 404 mapped to the input channel, while each column of the processing engines 411 may be mapped to an output channel. The input data elements are stored in a continuous address space and follow an order based on their coordinates in the input data array. Each processing engine 411 may store weight data elements for the input channel and the output channel to which the processing engine is mapped. Each column of the processing engine 411. Reference Figure 4A and Figure 4B , the processing engine 411 within the engine may receive input data elements of an input channel (e.g., Figure 4A The input data i) is multiplied by the stored weights (e.g., Figure 4AThe weight data w) is added to the input partial sum p_in to generate a product, the product is added to the input partial sum p_in to generate a new partial sum p_out, and the new partial sum p_out is passed to the processing engine 411 below the same column. The bottom processing engine 411 of the column can generate a partial sum representing the sum of the products between the weight data elements stored in the column of the processing engine 411 and the input data elements of the different input channels received from the memory substation 404.
[0073] Where the memory fetch parameters 430 indicate that the starting address is at the rightmost input data element of each row, a stride will be fetched (which may indicate a skip in this example), and a certain number of input data elements will be fetched, in the first iteration column 432 of the processing engine 411, a first partial sum may be generated based on the stored weight data elements and the input data elements provided by the memory subsystem 404, as shown below:
[0074] The first part and = X 0 0,0 ×W 0,0 0,0 +X 0 0,0 ×W 1,0 0,0 +...+X C 0,0 ×W C,0 0,0 (Equation 6)
[0075] In a second iteration, column 432 of processing engine 411 may generate a second partial sum based on the stored weight data elements and the input data elements provided by memory subsystem 404 as follows:
[0076] The second part and = X 0 0,1 ×W 0,0 0,0 +X 0 0,1 ×W 1,0 0,0 +...+X C 0,1 ×W C,0 0,0 (Equation 7)
[0077] Each column of processing engine 411 may provide partial sums generated in an iteration to a column sum buffer, such as column sum buffers 442, 443, etc., both of which are part of sum buffer 412. The partial sums are generated based on weight data elements at the same coordinates of different filter arrays associated with different input channels and output channels, and the partial sums correspond to different output data elements. Figure 4CEach of the column summing buffers 442 and 443 includes a plurality of entries, for example, E 0,0 、E 0,1 、E 0,2 Etc. Each entry may have coordinates that map to coordinates of an output tile that may represent a region of the output array. Each entry has an adder ( Figure 4C ), which allows the entry to add the received partial sum to the stored partial sum to produce an accumulated partial sum. The entry may then store the accumulated partial sum.
[0078] The operation at the column summing buffers 442 and 443 may be controlled by a set of buffer write parameters 452 including a destination offset, a step size, and the number of write elements. The destination offset parameter may indicate the first partial sum (of the first iteration) to be added to the entry. The step size parameter may indicate the number of entries to be skipped between adjacent entries receiving the partial sum. The step size parameter may correspond to the gap between non-zero input data elements that overlap with the weight data elements when the weight data array is in a different step position. In addition, the number of write elements indicates the number of partial sums to be added to the entries of the summing buffer starting from the starting address, wherein adjacent entries are separated based on the step size parameter as described above.
[0079] As an illustrative example, where the destination offset is 2 and the stride is 1, the first partial sum from column 432 may be stored in entry E 0,2 The second part and can be stored in E 0,3 The third part of the sum can be stored in E 0,4 , and so on until the number of partial sums specified by the number of written elements is stored. As will be described in more detail below, the buffer write parameters 452 may be included in the instructions for calculating a set of partial sums. The instructions may be parsed by the controller 422 to extract the buffer write parameters 452. The controller 422 may then control the operation of the summing buffer based on the extracted buffer write parameters 452. As will be described below, the buffer write parameters 452 may be configured to support convolution operations.
[0080] After calculating the partial sums from the first set of weight data elements (whose coordinates in the corresponding filter array are the same but the input and output channels are different), the processing engine array 410 may load a new set of weight data elements from different coordinates and repeat the partial sum calculations. The new partial sums may be added to the partial sums calculated from the first set of weight data elements stored in the summing buffer 412. The calculation and accumulation of the partial sums for the remaining weight data elements may continue to generate the data elements of the output tile. After generating the data elements of the output tile, the summing buffer 412 may provide the data elements of the output tile to the activation engine 416 and / or the pooling engine 418 for post-processing, and the post-processed output data elements may be stored in the memory subsystem 404. The post-processed output data may be sent from the memory subsystem 404 to the chip interconnect 420 and / or extracted to the processing engine array 410 as input data for subsequent neural network layer processing.
[0081] Figure 5A-5D An example configuration of accelerator 402 is shown for performing normal convolution operations. Figure 5A The overlap between the different weight data elements of the 3x3 filter array 504 and the input data elements of the input data array 502 is shown. Figure 5A , the input data array 502 may be padded with a row 506 of zeros on the top and a column 508 of zeros on the left. The number of rows padded with zeros may be specified by a pad_north parameter, where pad_north is equal to one, indicating that a row of zeros is padded on the top of the input data array 502. Additionally, the number of columns padded with zeros may be specified by a pad_west parameter, where pad_west is equal to one, indicating that a column of zeros is padded on the left of the input data array 502. A normal convolution operation may be performed between the zero-padded input data array 502 and the filter array 504, with a stride of 2. Some input data elements / zero padding that overlap with specific weight data elements at different stride positions are shaded. As Figure 5A As shown, some zero padding and input data elements at coordinates (1, 1) may overlap with weight data elements (0, 0) at different step positions. In addition, input data elements (0, 0), (0, 2) and (2, 2) may overlap with weight data elements (1, 1) at different step positions. In addition, input data elements (1, 1), (1, 3) and (3, 3) may overlap with weight data elements (2, 2) at different step positions. In each case, there is a gap between each input data element that overlaps with the weight data element. The gap can be defined based on the stride distance of the convolution operation. When the stride distance is two, the gap includes one input data element.
[0082] Return to reference Figure 4B and Figure 4C , in order to execute Figure 5A For normal convolution operation of the convolution operation, the controller 422 may be provided with memory extraction parameters 430, which define a set of overlapping non-zero input data elements for loading into the weight data elements in the processing engine 411. The set of overlapping non-zero input data elements may be defined based on the starting address, the step size, and the number of extraction element parameters. The starting address may be the address of the first overlapping non-zero input data element in the memory subsystem 404. The address may be determined based on a mapping between the memory subsystem 404 address and the coordinates of the input data elements stored in the memory subsystem 404. In addition, the step size may correspond to the gap described above and may be based on the stride distance. In addition, the number of extraction elements may be based on the size of the output tile block for the convolution operation, which in turn may be based on the size of the summing buffer, as will be described below. Based on the memory extraction parameters 430 for each weight data element, the controller 422 may extract the correct subset of input data elements to the processing engine array 410 to be multiplied with the weight data element to produce a partial sum.
[0083] In addition, the controller 422 may be provided with a buffer write parameter 452 to store non-zero partial sums for different stride positions at entries of a column sum buffer (eg, column sum buffer 442) corresponding to the stride position. Figure 4C , the partial sum at stride position (0, 0) can be accumulated in entry E 0,0 At the stride position (0, 1), the partial sum can be accumulated in entry E 0,1 The destination offset of the buffer write parameter 452 may also be based on a set of overlapping non-zero input data elements for the weight data element. Specifically, a starting offset may be provided to the column sum buffer to skip a number of entries corresponding to zero partial sums corresponding to a number of padding zeros that overlap with the weight data element until the entry corresponds to a stride position where the weight data element overlaps with the input data element. Figure 5A In the example of , at the first few stride positions, the weight data element (0,0) overlaps with the zero padding until it overlaps with the input data element at (1,1). The destination offset of the buffer write parameter 452 can be configured to ensure that the partial sum generated from the input data element at (1,1) is stored at the entry that reflects the stride position of the filter array 504 when the weight data element (0,0) overlaps with the input data element (1,1).
[0084] Figure 5B Example operations for determining overlapping input data elements for a particular weight data element that may be performed by a compiler are shown. Figure 5BBased on the size of the column sum buffer (e.g., the number of rows and columns of entries), the compiler can determine the size of the output tile and the total number of entries of the column sum buffer. As described above, the output tile includes the output data elements of the region 520 in the output data array, and the output tile can be defined by the coordinate range of the first region in the output data array.
[0085] In operation 522, the compiler may perform a projection operation from the region 520 represented by the output tile to determine a region 530 in the input data array 502 that may provide input data elements to be convolved with the weight data elements to produce the output tile. The projection operation may take into account the size of the first region and the stride distance of the convolution operation. As described above, the stride distance may define a gap between each overlapping input data element, where a stride distance of two produces a gap between the input data elements. The size of the region 530 may be determined based on the size of the region 520 and the gap. For example, referring to Figure 5B , based on an output tile having 10 output data elements per row and having three rows (30 entries total), the size of region 530 may be determined by scaling the size of region 520 by two, such that region 520 has 20 input data elements per row and has six rows. With such an arrangement, when one input data element is skipped between two input data elements, the total number of input data elements (and the resulting partial sum) may be equal to the total number of output data elements in the output tile, and the number of entries of the column sum buffer.
[0086] After determining the size of region 520, the compiler may align region 530 with padded input data array 502. The alignment may be based on the coordinates of the weight data elements and the pad_north parameter and the pad_west parameter. The coordinates of the upper left inflection point of region 530 relative to the original input data array 502 may be based on the following equation:
[0087] Start_coordinates=(weight_r-pad_west, weight_s-pad_north) (Equation 8)
[0088] In Equation 8, start_coordinates refers to the coordinates of the upper left inflection point of region 530, weight_r refers to the row coordinates of the weight data element, weight_s refers to the column coordinates of the weight data element, pad_west refers to the number of zero columns added to the left of the input data array 502, and pad_north refers to the number of zero rows added to the top of the input data array 502.
[0089] like Figure 5AAs shown, for the weight data element (0, 0), the upper left corner of the region 530 can be aligned with the upper left corner of the zero-padded input data array 502. With this alignment, the weight data element (0, 0) having the filter array 504 at the stride position (0, 0) overlaps with the uppermost element of the region 530, which represents the first input data element to be multiplied by the weight data element and is at coordinates (-1, -1) relative to the original input data array 502. Based on the alignment operation, the compiler can determine the range of the target coordinates of the region 520 relative to the original input data array 502. Figure 5B , the target coordinates of region 530 may range from (-1, -1) to (4, 18) relative to the upper left inflection point of the original input data array 502 having coordinates (0, 0). The position of the upper leftmost element of region 530 may be a reference position and may be based on the position of the output tile in the output data array. For example, for the second output tile immediately below region 520, the reference position may be offset from (-1, -1) by the height of region 530 and may be at (5, -1).
[0090] In operation 540, after determining the target coordinates of the region 530, the compiler may overlay a stride pattern 550 on the region 530. The stride pattern 550 may define gaps between overlapping input data elements based on a stride distance. Each dark box in the stride pattern 550 may represent an overlapping input data element. As described above, in the case where the stride distance is two, the gap includes an input data element. When the stride pattern 550 is overlaid on the region 530, the upper left corner of the stride pattern 550 is aligned with the upper left corner of the region 530. Based on the alignment and gap in the stride pattern 550, the compiler may calculate the coordinates of the stride pattern 550 relative to the original input data array 502 based on the coordinates of the upper leftmost element of the region 530 (-1, -1) and the gap information. For example, the first element of the stride pattern overlaps the upper leftmost element of the region 530 and has the coordinates (-1, -1), the second element of the stride pattern on the same row as the first element has a gap of 1 with the first element and may have the coordinates (-1, 1), and so on. Based on the coordinates of the stride pattern and the size of the original input data array 502 that may define a range of coordinates of input data elements included in the original input data array 502, the compiler may identify a first subset of coordinates within the zero-padding region and a second subset of coordinates within the original input data array 502. The first subset of coordinates is within the zero-padding region and may represent zero input data elements that produce a zero partial sum (due to multiplication of zeros), while the second subset of coordinates may represent non-zero input data elements that produce a non-zero partial sum. The non-zero input data elements represented by the second subset of coordinates may be overlapping input data elements with the weight data element (0, 0) in the original input data array 502.
[0091] The compiler may determine the starting address, stride, and number of elements to extract parameters of the memory fetch parameters 430 based on the first subset of coordinates in the padded region and the second subset of coordinates in the original input data array 502. Specifically, the starting address parameter of the memory fetch parameters 430 may correspond to the first coordinate in the second subset of coordinates. Figure 5B In the example of , the starting address may correspond to coordinate (1, 1) of the input data array 502, which may be translated into an address in the memory subsystem 404. Furthermore, the stride is based on the stride distance as described above. The number of extract element parameters may be based on the size of the second subset of coordinates. Figure 5B In the example of , the size of the second subset of coordinates may be 18, since there are 18 non-zero input data elements in the original input data array 502 that may overlap with the weight data element (0, 0). Therefore, the number of extracted elements parameter may also be set to 18.
[0092] The compiler may also determine the destination offset, stride, and number of write element parameters of the buffer write parameter 452 based on the first subset of coordinates in the zero padding area and the second subset of coordinates in the original input data array 502. Specifically, based on the first subset of coordinates, the compiler may determine that the first 11 input data elements are zero, which means that the first 11 entries of the column sum buffer need to be skipped, and the destination offset parameter may be set to 11. In addition, when 18 input data elements are obtained, 18 partial sums will be generated, and the number of write element parameters may be set to 18. In addition, the non-zero input data elements overlapping the weight data elements are separated by stride, which means that when the weight data array is in various stride positions in the input data array, there is no gap between the overlapping non-zero input data elements. Therefore, the stride parameter of the buffer write parameter 452 may be set to one. In order to maintain a rectangular shape, the 18 partial sums may be stored in a 9x2 area in the column sum buffer (which may be based on the number of input data elements in a row), and the column sum buffer is shifted to the right, wherein the first column entry is skipped.
[0093] Referring back to operation 522 and Equation 8, the compiler may adjust the alignment of region 530 relative to pad input data array 502 based on the coordinates of the weight data elements by adding offsets along the row and column dimensions. Figure 5C, for the weight data element (1, 1), the compiler can use equation 8 to calculate the coordinates of the upper left inflection point of zone 530 and obtain (0, 0). That is, compared with the weight data element (0, 0), zone 530 is shifted one unit to the right and bottom from the upper left inflection point of the filled input data array 502 and relative to the reference position (-1, -1). The coordinates of the upper left inflection point of zone 530 can be changed to (0, 0), and the coordinate range of zone 530 can be changed to (0, 0) to (5, 19). Through this alignment, the upper leftmost element of zone 530 representing the first input data element to be multiplied by the weight data element (1, 1) overlaps with the weight data element when the filter array 504 is in the stride position (0, 0). For the weight data element (1, 1), zone 530 and the stride pattern 550 overlap with the original input data array 502 instead of zero padding. The first input data element starts at coordinates (0, 0), and a total of 30 input data elements can be extracted. Furthermore, since there are no zero input data elements, there are no entries to skip the column summation buffer.
[0094] In addition, for the weight data element (2, 2), the compiler can use equation 8 to calculate the coordinates of the upper left inflection point of zone 530 and obtain (1, 1). That is, the compiler can shift zone 530 two units to the right and bottom from the upper left inflection point of the padding input data array 502 (reference position (-1, -1)). The coordinate range of zone 530 becomes (1, 1) to (6, 20). Through this alignment, the upper leftmost element of zone 530 representing the first input data element to be multiplied with the weight data element (2, 2) overlaps with the weight data element when the filter array 504 is in the stride position (0, 0). For the weight data element (2, 2), zone 530 and stride pattern 550 overlap with the original input data array 502 instead of zero padding. The first input data element starts at coordinates (1, 1), and a total of 27 input data elements can be extracted. In addition, since there is no zero input data element, there is no entry for skipping the column summation buffer. To maintain the rectangular shape, the 27 partial sums may be stored in 9x3 regions in the column sum buffer (which may be based on the number of input data elements in a row), and the column sum buffer is shifted to the left, skipping the last column entry.
[0095] Figure 5D Shown for Figure 5A-5C An example of memory fetch parameters 430 and buffer write parameters 452 for weight data elements (0,0), (1,1), and (2,2) of a convolution operation is shown, as described above.
[0096] Figure 6A-6C An example configuration of an accelerator 402 for performing a normal convolution operation is shown. Figure 6A-6CIn the example of FIG. 5 , a dilated convolution operation is shown between the filter array 504 and the input data array 502 at a rate of 2 and a stride of 2. Fig. 6A As shown, the effect of the dilated convolution operation can be to pad zeros between adjacent weight data elements of the filter array 504 to expand from 3x3 to 5x5, and perform a normal convolution between the zero-padded filter array 504 and the zero-padded input data array 502.
[0097] Figure 5B and Figure 5C The convolution operation described in can be modified to support a dilated convolution operation. Specifically, to perform a dilated convolution operation, the processing engine array 410 can load the same set of weight data elements of the filter array 504, but the selection of a subset of input data elements for each weight data element of the filter array 504 based on the target coordinate range of the region 530 will be modified to account for the rate of the dilated convolution operation. Specifically, referring to Figure 6B , the coordinates of the upper left corner element of region 530 relative to the original input data array 502 may be based on the following equation:
[0098] Start_coordinates dilated =(weight_r×rate-pad_west, weight_s×rate-pad_north) (Equation 9)
[0099] In Equation 9, start_coordinates dilated 502, weight_r refers to the row coordinates of the weight data elements in the original filter array without zero padding, weight_s refers to the column coordinates of the weight data elements in the original filter array without zero padding, pad_west refers to the number of zero columns added to the left of the input data array 502, and pad_north refers to the number of zero rows added to the top of the input data array 502. The effect of Equation 9 is that for different weight data elements, the shift magnitude of the coordinates of the upper left corner element of the region 530 that defines the coordinates of the input data elements that overlap with the weight data elements of the extended filter array 504 is proportionally adjusted at the rate.
[0100] For example, refer to Figure 6B , the weight data element (0,0) of the extended filter array 504 is at the same coordinate (0,0) in the original filter array 504, so the upper left corner of the region 530 is aligned with the upper left corner of the zero-padded input data array 502, and has a range of target coordinates from (-1,-1) to (4,18) relative to the upper leftmost element of the original input data array 502 having coordinates (0,0). Figure 5B Same as in, the number of elements to extract parameters and the number of elements to write parameter may be set to 18, and the destination offset parameter may be set to 11. The stride is set to 2 based on the stride.
[0101] In addition, for the weight data element (2, 2) of the extended filter array 504 corresponding to the weight data element (1, 1) of the original filter array 504, the compiler can use equation 9 to calculate the coordinates of the upper left inflection point of the zone 530 and obtain (1, 1). That is, compared with the weight data element (0, 0) (of both the extended filter array and the original filter array 504), the zone 530 is shifted two units to the right and bottom from the upper left corner of the padded input data array 502. The coordinate range of the zone 530 becomes (1, 1) to (6, 20). Through this alignment, the upper leftmost element of the zone 530 representing the first input data element to be multiplied with the weight data element (2, 2) of the extended filter array 504 overlaps with the weight data element when the filter array 504 is in the stride position (0, 0). For the weight data element (2, 2), the zone 530 and the stride pattern 550 overlap with the original input data array 502 instead of zero padding. The first input data element starts at coordinates (1, 1), and a total of 27 input data elements can be extracted. Furthermore, since there are no zero input data elements, there are no entries for the skip column summation buffer.
[0102] In addition, for the weight data element (4, 4) of the extended filter array 504 corresponding to the weight data element (2, 2) of the original filter array 504, the compiler can use equation 9 to calculate the coordinates of the upper left inflection point of the zone 530 and obtain (3, 3). That is, compared with the weight data element (0, 0) (of both the extended filter array and the original filter array 504), the zone 530 is shifted four units to the right and bottom from the upper left corner of the padded input data array 502. The coordinate range of the zone 530 becomes (3, 3) to (8, 22). Through this alignment, the upper leftmost element of the zone 530 representing the first input data element to be multiplied with the weight data element (4, 4) of the extended filter array 504 overlaps with the weight data element when the filter array 504 is in the stride position (0, 0). For the weight data element (4, 4), the zone 530 and the stride pattern 550 overlap with the original input data array 502 instead of zero padding. The first input data element starts at coordinates (3, 3), and a total of 24 input data elements can be extracted. Furthermore, since there are no zero input data elements, there are no entries for the skip column summation buffer. Figure 6C Shown for Figure 6A-6BAn example of memory fetch parameters 430 and buffer write parameters 452 for the weight data elements (0,0), (2,2), and (4,4) of the zero-padding filter array 504 for a dilated convolution operation is shown, as described above.
[0103] Figure 7 A block diagram is included that illustrates an example of a host system 700 on which a compiler 730, such as described herein, may run. The illustrated host system 700 is an example of a computing device and includes a processor 702, a processor memory 704, at least one storage device 706, various input / output (I / O) devices 708, and at least one network interface 710. Figure 7 In the example of the host system 700, the host system 700 further includes an acceleration engine 712, which may include Figure 4A-4C 402 of the accelerator. In various examples, the host system 700 may be implemented as a server in a data center, a desktop computer, a laptop computer, a tablet computer, or a smart phone, among other examples. In some examples, operations or components as discussed below or included in the host system 700 may be performed or included in other computer devices. For example, the compiler 730 may be executed on the host system 700, while the acceleration engine 712 is located at a different host system.
[0104] The processor 702 is an integrated circuit device that can execute program code in the form of instructions. The program code can be used for various software applications or tools, such as an operating system 720 or a compiler 730 shown. When the processor 702 executes a program, the instructions for the program can be stored in the processor memory 704. The instructions can also be stored elsewhere, such as on a storage device 706, and can be loaded into the processor memory 704 when the processor 702 needs it. The processor 702 can also use the processor memory 704 for temporary storage of other data that the processor 702 operates. In various examples, the processor memory 704 is a volatile memory type, such as a random access memory type, but alternatively or additionally, a non-volatile memory type can be used for the processor memory 704.
[0105] Storage device 706 is an example of a device that may include non-volatile memory. For example, storage device 706 may be a disk drive, a solid-state drive, or an optical drive, among other examples. Storage device 706 may further be non-transitory, such that when storage device 706 is not powered on, program code and other data stored on storage device 706 still exist.
[0106] Storage device 706 is an example of a peripheral device, which is a component that can be coupled to host system 700 to add functionality to host system 700. Other examples of peripheral devices include input / output device 108 and network interface 710. Input / output device 708 may include user input and output devices, such as keyboard, mouse, touch screen, microphone, display screen, speaker, printer and scanner, among other examples. Network interface 710, which may be implemented using a network interface card, may provide access to one or more networks. Network interface 710 may include, for example, a physical port for connecting a network cable and / or a wireless antenna to communicate with WiFi and / or cellular networks. Network interface 710 may also be described as an I / O device.
[0107] Acceleration engine 712 is also another type of peripheral device or I / O device. Acceleration engine 712 is a device that is specifically built to perform certain operations that can be performed by processor 702 but can be performed faster by acceleration engine 712. For example, acceleration engine 712 can be a neural network accelerator and can therefore perform massively parallel computations of neural networks more efficiently than if the computations were performed by processor 702. As another example, acceleration engine 712 can be a graphics processing unit (GPU) and can be optimized to perform the computations required for graphics rendering. Other examples of devices that can be implemented by acceleration engine 712 include cryptographic accelerators, compression and decompression accelerators, 3D accelerators, regular expression accelerators, security accelerators, etc.
[0108] In various examples, the acceleration engine 712 can execute program code to perform certain operations. For example, when the acceleration engine 712 is a neural network accelerator, the acceleration engine 712 can be programmed to execute a specific neural network, such as a neural network that performs image recognition or a neural network that performs machine translation. As another example, to support the execution of the neural network, the acceleration engine 712 can be programmed to perform operations such as copying data for the neural network from (for example) the processor memory 704 to the acceleration engine 712, copying input data for the neural network from the processor memory 704 to the acceleration engine 712, and / or copying results from the acceleration engine 712 to the processor memory 704, as well as other examples.
[0109] To generate program code for the acceleration engine 712, in various examples, the host system 700 may execute a compiler 730. Generally, a compiler is a software program that translates program code written in a human-readable language into a format (e.g., machine instructions) that can be read and processed by an integrated circuit device. Figure 7In the example of , acceleration engine 712 is a neural network accelerator, and compiler 730 is used to compile the neural network description into instructions to be executed by acceleration engine 712. When acceleration engine 712 implements a different type of accelerator, another compiler may be used.
[0110] The compiler 730 may be activated, for example, when the operating system 720 receives keyboard, mouse, touch screen, voice command, or other input from the input / output device 708. The input may also include parameters for the compiler 730, such as input code 742 to compile and configure options for the compilation process. After activating the compiler 730, the processor 702 may load instructions for the compiler 730 into the processor memory 704 and may execute the instructions.
[0111] exist Figure 7 In the example of , compiler 730 includes a first stage 732, a second stage 736, and a third stage 740 that each perform different operations to generate compiled code 144. In other examples, compiler 730 may combine the operations of first stage 732, second stage 736, and / or third stage 740 into fewer stages, or may divide the operations of one of the stages into multiple stages.
[0112] The first stage 732 can receive and process input code 742. The input code 742 can describe a program in a high-level programming language such as Java, C++ or Tensorflow and many other examples. The input code 742 can describe steps such as performing image recognition, speech recognition, machine translation or other operations. The input code 742 can be obtained, for example, from a storage device 706. Alternatively, although not described here, the input code 742 can be located in the processor memory 704 or can be obtained from a network location using a network interface 710. The processing of the input code 742 can include classifying the operations described in the input code 742 into layers, where the output of one layer provides input to the next layer. The processing can also include identifying steps to be performed by the processor 702 rather than by the acceleration engine 712. For example, through the execution of the driver 722, the processor 702 may need to perform steps such as configuring a direct memory access (DMA) descriptor for moving data into or out of the acceleration engine 712 and other examples.
[0113] The output 734 of the first stage 732 may be organized, for example, in layers, nodes, and connections between neural network nodes. The second stage 736 may perform intermediate processing on this output 734. For example, the operations performed in any one layer or in any one node in a layer may be too many for the acceleration engine 712 to perform simultaneously. The acceleration engine 712 may, for example, have a limited amount of local storage space for the data required for the calculation, or the calculation may exceed the amount that the acceleration engine 712 can perform at one time. In this example, the first stage 732 may decompose the operations of the layer or node into smaller operations that can fit into the local memory of the acceleration engine and / or can fit into the computational capacity of the acceleration engine 712. The processing of the output 734 of the first stage 732 may include other steps, such as arranging or determining the order in which the acceleration engine 712 and / or the processor 702 will perform the operations, as well as other examples.
[0114] In various examples, the output 738 of the second stage 736 includes various steps to be performed in the order in which the steps are performed by the components of the acceleration engine 712. The output 738 can be represented, for example, as a data flow graph, where the nodes in the graph represent memory operations, calculations, and other operations, and the edges or connections between the nodes represent dependencies between the nodes, such as data dependencies, memory dependencies, or operational dependencies, among other examples.
[0115] The third stage 740 may operate on the output 738 of the second stage 736 and perform various steps before generating instructions to be executed by the acceleration engine 712. These steps may include, for example, removing redundant dependencies, resolving or handling dependencies between nodes by inserting synchronization instructions into the code, identifying possible optimizations in memory usage or memory bandwidth usage, and other operations.
[0116] In some instances, the third stage 740 may include a data scheduler 750 to schedule the movement of data, such as input data and weight data, in the acceleration engine 712 to support various operations, such as the convolution operations and dilated convolutions described above. For example, the data scheduler 750 may obtain instructions (e.g., from a dataflow graph) to perform a convolution operation (e.g., normal convolution, dilated convolution, etc.) between an input data array and a filter array to produce a convolution output array. Based on the size of the summing buffer at the acceleration engine 712, the data scheduler 750 may determine the output tiles that fit into the summing buffer, and may determine a sequence of instructions for preparing the convolution operation to produce one output tile at a time. For each instruction, the data scheduler 750 may determine a sequence for loading the weight data elements of the filter array into the processing engine array 410, and based on the instructions described above in Figures 5A-6C430 , the data scheduler 750 may also determine a set of buffer write parameters 452 including a destination offset and a number of elements to be written. The data scheduler 750 may then generate instructions to control the acceleration engine 712 to load the weight data elements and the corresponding subset of the input data elements and the rate of the dilated convolution. The data scheduler 750 may also determine a set of buffer write parameters 452 including a destination offset and a number of elements to be written, based on the projection operation. The data scheduler 750 may then generate instructions to control the acceleration engine 712 to load the weight data elements and the corresponding subset of the input data elements to perform the convolution operation.
[0117] The output of the third stage 740 is compiled code 744, which may include machine instructions in binary format. In some examples, the compiled code 744 may be stored in the processor memory 704. Alternatively or additionally, the compiled code 744 may be copied to the storage device 706 or to a network location. As noted above, the acceleration engine 712 may be located at a different host system, in which case the compiled code 744 may be sent to another host system via the network interface 710.
[0118] exist Figure 7In the example of , the host system 700 can execute a driver 722, which can also be referred to as a device driver or a runtime driver, which manages the acceleration engine 712. The driver 722 can provide an interface between an application executing on the host system 700 (or on another host system) and the acceleration engine 712. For example, the driver 722 can provide an application program interface (API) that defines functions for feeding input data to the acceleration engine 712 and defining operations to be performed on the input data. In this example and other examples, the driver 722 can configure the acceleration engine 712 to perform the operations. For example, the driver 722 can identify the neural network that the acceleration engine 712 will execute, and the location of the compiled code 744 for the neural network in the processor memory 704 or on the storage device 706. The driver 722 can also load the compiled code 144 into the acceleration engine 712 or cause the acceleration engine 712 to load the compiled code, can load or cause the acceleration engine 712 to load the input data that the neural network will operate on, and / or can cause the acceleration engine 712 to begin executing the input data. After the acceleration engine 712 finishes, the acceleration engine 712 may notify the driver 722, and the driver 722 may pass the result back to the application that requested the result.
[0119] Figure 8 A flow chart of an example method 800 for performing a dilated convolution operation is shown. The method 800 may be performed by various components of the accelerator 402, such as the memory subsystem 404, the processing engine array 410, the summing buffer 412, and the controller 422.
[0120] Method 800 begins with step 802, where a controller (e.g., controller 422) may load a first weight data element of an array of weight data elements from a memory (e.g., memory subsystem 404) into a systolic array (e.g., processing engine array 410), the first weight data element being at a first coordinate within the array of weight data elements. The first weight data element may be obtained from memory subsystem 404 based on the first coordinate. The weight data element may be stored in addresses in memory subsystem 404 that reflect the coordinates of the weight data element in the array of weight data elements. Controller 422 may have the address of the first weight data element in a first calculation instruction, and may obtain the first weight data element based on the address from memory subsystem 404 after executing the first calculation instruction. In addition, as described above, each processing engine 411 may store weight data elements, and the controller may send the first weight data element to processing engine 411 for storage.
[0121] In step 804, the controller may select a first subset of input data elements of the input data element array based on the first coordinate of the first weight data element in the weight data element array, the stride of the dilated convolution operation, and the rate of the dilated convolution operation. The first subset of input data elements will be multiplied with the first weight data element at the processing engine 411 to produce a first partial sum, which may be forwarded to the column summing buffer (e.g., column summing buffer 442) of the summing buffer 412. The first subset may be selected based on a first calculation instruction including a first group of memory acquisition parameters 430, and the first subset may include a starting address, a stride, and the number of elements. When the weight data element array is in various stride positions within the input data element array in the dilated convolution operation, the starting address and the number of elements may reflect the input data elements that overlap with the first weight data element. The determination of the first subset of input data elements may be based on a projection operation. Specifically, the size of the summing buffer (e.g., the number of columns and rows) may define an output tile block of output data elements including the first zone in the output data array. The first zone may be defined by the range of actual coordinates in the output data array. Based on the projection operation of the first area of the output data array coordinates and the stride of the dilated convolution, the compiler can determine a second area including the input data elements to be convolved with the first weight data element. The second area can be limited by the range of the target coordinates of the input data elements. The second area (and the range of the target coordinates) can be shifted by an offset based on the coordinates of the first weight data element in the weight data array, and a proportional factor based on the rate of the dilated convolution operation. The compiler can then align the stride pattern with the shifted second area to identify the position of the second area overlapping with the stride pattern. The stride pattern defines the position of the input data elements overlapping with the weight data elements, and reflects the stride of the dilated convolution operation. Based on the overlap, the compiler can determine a set of target coordinates of the overlapping input data elements. The starting address of the first subset can be determined based on the coordinates of the first overlapping input data element, and the count of the overlapping input data elements can set the number of extraction element parameters. The stride parameter is set to two based on the stride pattern, and the stride pattern can reflect the stride of the dilated convolution operation.
[0122] In step 806, the controller may stream each input data element of the first subset into the systolic array starting from the first address from the memory to be multiplied by the first weight data element to calculate a first partial sum. In step 802, the input data elements may be sent in sequence to the processing engine 411 storing the first weight data elements. The processing engine 411 may multiply each input data element by the first weight data element to produce a first partial sum. The first address may be the starting address of the first subset described above. The first partial sum may be sent to a first destination address in the column sum buffer based on a first calculation instruction including a first set of buffer write parameters 452, which may include a destination offset, a step size, and a number of elements to be written. The first partial sum may be added to the data stored at the first destination address. The first set of buffer write parameters 452 may also be based on a shift stride pattern, such as Figure 6B described.
[0123] In step 808, the controller may load a second weight data element of the weight data element array from the memory subsystem 404 into the processing engine array 410, the second weight data element being at a second coordinate within the weight data element array. The second weight data element may be obtained from the memory subsystem 404 based on the second coordinate. The controller 422 may have an address of the second weight data element in the second calculation instruction, and may obtain the second weight data element based on the address from the memory subsystem 404 after executing the second calculation instruction. In addition, the second weight data element may replace the first weight data element stored in the processing engine 411.
[0124] In step 810, the controller may select a second subset of input data elements in the input data element array based on the second coordinate of the second weight data element in the weight data element array, the stride of the dilated convolution operation, and the rate of the dilated convolution operation. The second subset of input data elements will be multiplied with the second weight data element at the processing engine 411 to produce a second partial sum, which can be forwarded to the column summing buffer. The second subset can be selected based on a second calculation instruction including a second group of memory fetch parameters 430. When the weight data element array is in various stride positions within the input data element array in the dilated convolution operation, the second subset can be those input data elements that overlap with the second weight data element. The determination of the second subset of input data elements can be similar to that described in step 804.
[0125] In step 812, the controller may stream each input data element of the second subset into the systolic array starting from the second address from the memory to be multiplied by the second weight data element to calculate the second partial sum. In step 812, the input data elements may be sent in sequence to the processing engine 411 storing the second weight data elements. The processing engine 411 may multiply each input data element by the second weight data element to produce a second partial sum. The second address may be the starting address of the second subset described above. The second partial sum may be sent to a second destination address in the column sum buffer based on a second calculation instruction including a second set of buffer write parameters 452. The second partial sum may be added to the data stored at the second destination address, some or all of which may overlap with the first destination address. The second set of buffer write parameters 452 may be based on a shift stride pattern, such as Figure 6B described.
[0126] In step 814, the controller may generate an output data array of the transposed convolution operation based on the first partial sum and the second partial sum. As described above, some of the first destination address and the second destination address may overlap, and some of the output data elements of the output data array may include the sum of the first partial sum and the second partial sum. Other output data elements of the output data array may be formed by superposition of the first partial sum and the second partial sum.
[0127] Fig. 9 Flowchart showing an example method 900 for generating instructions for a neural network processor to perform a dilated convolution operation. The method 900 may be performed by, for example Figure 8 The compiler 830 or the like is executed by a compiler.
[0128] The method 900 begins at step 902, where the compiler may receive first information indicating a stride and rate of a dilated convolution operation to be performed by a systolic array (e.g., the processing engine array 410) based on a weight data array and an input data array to produce an output data array. The first information may be received from, for example, input code 842 of an application (an image processing operation, etc.) that may represent a result using a dilated convolution operation.
[0129] In step 904, the compiler may receive second information indicating a size of a summing buffer, such as column summing buffer 442. The summing buffer accumulates and stores partial sums from the systolic array for use in the dilated convolution operation. The second information may also be received from, for example, input code 842. As described above, the size information may be used to determine an output tile and, via a projection operation, may be used to determine a subset of input data elements of the input data array for each weight data element of the weight data array.
[0130] In step 906, the compiler may determine, for each weight data element of the weight data array, a subset of input data elements of the input data array to be multiplied with each weight data element to calculate the partial sum. The determination of the subset of input data elements may be based on a projection operation. Return to Reference Figure 5B and Figure 5C , the size of the summing buffer (e.g., the number of columns and rows) may define an output tile of output data elements including a first region in the output data array. The first region may be defined by a range of actual coordinates in the output data array. Based on the projection operation of the first region into the output data array coordinates and the stride of the dilated convolution, the compiler may determine a second region including input data elements to be convolved with the first weight data element. The second region may be defined by a range of target coordinates of the input data elements. The second region (and the range of target coordinates) may be shifted by an offset based on the coordinates of the first weight data element in the weight data array, and by a scale factor based on the rate of the dilated convolution operation. The compiler may then align the stride pattern with the shifted second region to identify the position of the second region overlapping the stride pattern. The stride pattern defines the position of the input data elements overlapping the weight data elements, and reflects the stride of the dilated convolution operation. Based on the overlap, the compiler may determine a set of target coordinates for the overlapping input data elements. After determining a set of target coordinates of overlapping input data elements of the first weight data element, the compiler may perform a coordinate-to-address mapping operation to determine addresses of a first subset of input data elements in a memory including a starting address and a count of addresses, wherein the count may indicate a count of the first subset of input data elements.
[0131] In step 908, the compiler may determine a destination address of a summing buffer for each weight data element of the weight data array to receive the partial sum, the destination address being determined based on a subset of the input data elements of each weight data element determined in step 906. For example, referring to Figure 6A-6C Based on the coordinates of the subset of input data elements, the compiler can determine the number of zero partial sums, the number of partial sums to be sent to the summing buffer, etc. Based on this information, the compiler can determine the starting destination address and the count of destination addresses for receiving partial sums.
[0132] In step 910, the compiler may generate a calculation instruction for each weight data element of the weight data array to include third information indicating a destination address and a subset of input data elements. The third information may include, for example, a starting source address and a count of input data elements based on the projection operation in step 906. The information may also include, for example, a starting destination address, a step size of one indicating no gaps between adjacent destination addresses, and a count of destination addresses based on the operation in step 908. The calculation instruction may also include an address of each weight data element in a memory based on the coordinates of each weight data element in the weight data array.
[0133] Embodiments of the present disclosure may also be described in terms of the following:
[0134] 1. A method comprising:
[0135] loading a first weight data element of an array of weight data elements from a memory into a systolic array, the first weight data element being at a first coordinate within the array of weight data elements;
[0136] receiving a selection of a first subset of input data elements of an array of input data elements, the first subset selected based on the first coordinate of the first weight data element, a stride of a dilated convolution operation, and a rate of the dilated convolution operation;
[0137] streaming each input data element of a selected first subset into the systolic array starting from a first address from the memory to be multiplied with the first weight data element to compute a first partial sum;
[0138] loading a second weight data element from the memory into the systolic array, the second weight data being at a second coordinate within the array of weight data elements;
[0139] receiving a selection of a second subset of input data elements of the array of input data elements, the second subset selected based on the second coordinate of the second weight data element, the stride of the dilated convolution operation, and the rate of the dilated convolution operation;
[0140] streaming each input data element of a selected second subset into the systolic array starting from a second address from the memory to be multiplied with the first weight data element to compute a second partial sum; and
[0141] An output data array of the dilated convolution operation is generated based on the first partial sum and the second partial sum.
[0142] 2. The method according to clause 1, further comprising:
[0143] accumulating a first partial sum with first data stored at first consecutive addresses of the summing buffer;
[0144] Accumulating a second partial sum with second data stored at second consecutive addresses of the summing buffer; and
[0145] after accumulating the first partial sum and the second partial sum, generating the output data array based on the data stored at the first consecutive addresses and the second consecutive addresses,
[0146] Wherein the count of the first subset of input data elements and the count of the second subset of input data elements are based on a size of the summation buffer and the rate of the dilated convolution operation.
[0147] 3. A method according to clause 2, wherein the dilated convolution is performed between zero-padded two-dimensional image data comprising the input data element array and the weight data element array;
[0148] Wherein the method further comprises: storing the first partial sum at the summing buffer starting at a third address based on the number of zero-padded elements of the zero-padded image data that overlap with the first weight data element in the dilated convolution operation.
[0149] 4. A method according to clause 3, wherein the dilated convolution between the zero-padded two-dimensional image data and the array of weight data elements does not involve multiplication between the zero-padded elements of the zero-padded image data and the array of weight data elements; and
[0150] wherein the second partial sum is added to data elements stored in the summation buffer starting at a fourth address, the difference between the fourth address and the third address corresponding to the number of zero-padded elements that overlap with the first weight data element in the dilated convolution operation.
[0151] 5. The method according to clause 4, further comprising:
[0152] executing a set of instructions to perform the dilated convolution operation,
[0153] The set of instructions includes:
[0154] first information indicating the first address of the memory, the count of the first subset of input data elements, and gaps between input data elements of the first subset of input data elements; and
[0155] second information indicating the second address of the memory, the count of the second subset of input data elements, and gaps between input data elements of the second subset of input data elements;
[0156] wherein receiving the selection of the first subset of input data elements comprises extracting the first information from the set of instructions; and
[0157] Wherein receiving the selection of the second subset of input data elements comprises extracting the second information from the set of instructions.
[0158] 6. A non-transitory computer-readable medium storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to:
[0159] loading a first weight data element of the array of weight data elements from a memory into the systolic array;
[0160] selecting a subset of input data elements, the subset selected based on information indicating a rate of a dilated convolution operation and a coordinate of the first weight data element within the array of weight data elements;
[0161] loading the subset of input data elements from the memory into the systolic array to perform a first computation of a dilated convolution operation; and
[0162] The systolic array is controlled to perform the first calculation based on the first weight data element and the subset to produce a first output data element of an output data array.
[0163] 7. The non-transitory computer-readable medium of clause 6, wherein the subset is selected based on a stride of the dilated convolution operation.
[0164] 8. The non-transitory computer-readable medium of clause 7, wherein the input data elements are stored in a contiguous address space within the memory.
[0165] 9. The non-transitory computer-readable medium of clause 8, wherein the subset of the input data elements is selected based on skipping a number of input data elements between each selected input data element, the number of skipped input data elements being based on the stride.
[0166] 10. The non-transitory computer-readable medium of clause 8 or 9, wherein the instruction comprises a plurality of parameters, the parameters comprising:
[0167] a starting source address of said subset of input data elements in said memory;
[0168] a skip parameter indicating said number of input data elements that are skipped; and
[0169] The number of the subset of the input data elements.
[0170] 11. The non-transitory computer-readable medium of clause 10, wherein the starting source address is based on the rate of the dilated convolution operation.
[0171] 12. The non-transitory computer-readable medium of any one of clauses 6 to 11, wherein the subset is a first subset;
[0172] wherein the instructions, when executed by one or more hardware processors, cause the one or more hardware processors to:
[0173] controlling the systolic array to perform the first calculation based on the first weight data element and the subset to produce a first partial sum;
[0174] controlling the summing buffer to accumulate the first partial sum;
[0175] loading a second weight data element of the array of weight data elements from a memory into the systolic array;
[0176] selecting a second subset of the input data elements from the memory into the systolic array;
[0177] controlling the systolic array to perform a second calculation based on the second weight data element and the second subset to produce a second partial sum; and
[0178] The summing buffer is controlled to accumulate the second partial sum to produce the first output data element.
[0179] 13. The non-transitory computer-readable medium of clause 12, wherein a first size of the first subset of the first input data elements and a second size of the second subset of the first input data elements are selected based on a size of the summation buffer.
[0180] 14. The non-transitory computer-readable medium of clause 13, wherein the instructions, when executed by one or more hardware processors, cause the one or more hardware processors to store the first partial sum at an address of a buffer based on a number of zero output data elements included in the first output data element.
[0181] 15. The non-transitory computer-readable medium of clause 14, wherein the instructions include a set of parameters, the set of parameters including:
[0182] a first starting source address of the first subset of the input data elements in the memory; and
[0183] a second starting source address of said second subset of said input data elements in said memory; and
[0184] Wherein the first starting source address and the second starting source address are offset from each other based on the rate of the dilated convolution.
[0185] 16. The non-transitory computer-readable medium of clause 15, wherein the set of parameters comprises a first starting destination address of the summing buffer for receiving the first partial sum and a second starting destination address of the summing buffer for receiving the second partial sum.
[0186] 17. A device comprising:
[0187] a memory storing a set of instructions; and
[0188] One or more hardware processors configured to execute the set of instructions to:
[0189] receiving first information indicating a rate and a stride of a dilated convolution operation to be performed by the systolic array based on the weight data array and the input data array to produce an output data array;
[0190] receiving second information indicating a size of the summing buffer;
[0191] determining, for each weight data element of the weight data array, a subset of input data elements of the input data array to be multiplied with the each weight data element to calculate a partial sum, the subset of input data elements being determined based on a projection operation from the summing buffer and based on the rate of the dilated convolution, the stride of the dilated convolution, the size of the summing buffer, and the coordinates of the each weight data element in the weight data array;
[0192] determining, for each of the weight data elements of the weight data array, a destination address of the summing buffer to receive the partial sum based on the subset of input data elements; and
[0193] A calculation instruction for each weight data element of the weight data array is generated to include third information indicative of the destination address and the subset of input data elements.
[0194] 18. The apparatus of clause 17, wherein the one or more hardware processors are configured to execute the set of instructions on each weight data element to:
[0195] determining a projection region aligned with a first input data element of the input data array based on the size information of the summing buffer and the stride of the dilated convolution operation;
[0196] shifting the projection region relative to the first input data element of the input data array by an offset based on the coordinates of the each weight data element in the weight data array and the rate of the dilated convolution;
[0197] aligning a stride pattern with the shifted projection area, the stride pattern being based on a stride of the dilated convolution operation;
[0198] determining, based on the size information, coordinates of input data elements of the input data array that overlap the stride pattern; and
[0199] The subset of input data elements for each weight data element is determined based on the coordinates.
[0200] 19. An apparatus according to clause 18, wherein the one or more hardware processors are configured to execute the set of instructions to determine the offset based on scaling the coordinates of the first weight data element in the weight data elements at the rate of the dilated convolution.
[0201] 20. The apparatus of clause 19, wherein the one or more hardware processors are configured to execute the set of instructions on each weight data element to:
[0202] determining a source address of a first one of the subset of the input data elements in the memory; and
[0203] Information indicating the address, the number of input data elements included in the subset of the input data array, and a stride is included in the computation instruction to enable the systolic array to load the subset of input data elements from the memory.
[0204] Fig.10 A diagram of an example network 1000 is included, which may include one or more host systems, such as Figure 7 For example, Fig.10 The example network 1000 includes a plurality of nodes 1002a-1002h, one or more of which may be, for example, Figure 7 The host system shown in FIG. Other nodes in the nodes 1002a-1002h may be other computing devices, each of which includes at least a memory for storing program instructions, a processor for executing instructions, and a network interface for connecting to the network 1000.
[0205] In various examples, the network 1000 can be used to process data. For example, input data can be received at one of the nodes 1002a-1002h or from other networks 1008 with which the network 1000 can communicate. In this example, the input data can be directed to a node in the network 1000 that includes an acceleration engine for the acceleration engine to operate and produce a result. The result can then be transmitted to the node or other network from which the input data is received. In various examples, the input data can be accumulated from various sources, including one or more of the nodes 1002a-1002h and / or a computing device located in other networks 1008, and the accumulated input data can be directed to one or more host systems in the network 1000. The results from the host system can then be distributed back to the source from which the input data was collected.
[0206] In various instances, one or more of nodes 1002a-1002h may be responsible for operations such as accumulating input data for host system operations, recording which host systems are busy and which host systems can accept more work, determining whether host systems are operating correctly and / or most efficiently, monitoring network security, and / or other management operations.
[0207] exist Fig.10 In the example of , nodes 1002a-1002h are connected to each other using a switching architecture with point-to-point links. The switching architecture includes multiple switches 1004a-1004d, which can be arranged in a multi-layer network such as a Clos network. A network device that screens and forwards packets between local area network (LAN) segments can be referred to as a switch. Switches typically operate at the data link layer (layer 2) and sometimes operate at the network layer (layer 3) of the open system interconnection (OSI) reference model, and can support several packet protocols. Fig.10 The switches 1004a-1004d may be connected to the nodes 1002a-1002h and provide multiple paths between any two nodes.
[0208] The network 1000 may also include one or more network devices, such as routers 1006, for connecting to other networks 1008. Routers use headers and forwarding tables to determine the best path for forwarding packets, and use protocols such as the Internet Control Message Protocol (ICMP) to communicate with each other and configure the best route between any two devices. Fig.10 The router 1006 may be used to connect to other networks 1008, such as a subnet, a LAN, a wide area network (WAN), and / or the Internet.
[0209] In some examples, the network 1000 may include any one or combination of many different types of networks, such as cable networks, the Internet, wireless networks, cellular networks, and other private and / or public networks. The interconnected switches 1004a-1004d and routers 1006 (if present) may be referred to as a switch fabric 1010, a fabric, a network fabric, or simply a network. In the context of computer networks, the terms "fabric" and "network" are used interchangeably herein.
[0210] Nodes 1002a-1002h may be any combination of host systems, processor nodes, storage subsystems, and I / O chassis representing user devices, service provider computers, or third-party computers.
[0211] The user device may include a computing device that accesses an application 1032 (e.g., a web browser or a mobile device application). In some aspects, the application 1032 may be hosted, managed, and / or provided by a computing resource service or service provider. The application 1032 may allow a user to interact with a service provider computer to, for example, access network content (e.g., web pages, music, videos, etc.). The user device may be a computing device, such as a mobile phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a netbook computer, a desktop computer, a thin client device, a tablet computer, an electronic book (e-book) reader, a game console, etc. In some examples, the user device may communicate with the service provider computer via other networks 1008. In addition, the user device may be part of a distributed system that is managed, controlled, or otherwise part of a service provider computer (e.g., a console device integrated with a service provider computer).
[0212] Fig.10The node may also represent one or more service provider computers. One or more service provider computers may provide native applications configured to run on user devices, and users may interact with the native applications. In some instances, service provider computers may provide computing resources, such as but not limited to client entities, low-latency data storage, durable data storage, data access, management, virtualization, cloud-based software solutions, electronic content performance management, etc. Service provider computers may also be used to provide web hosting, database, computer application development and / or implementation platform, a combination of the foregoing, etc. to users. In some instances, service provider computers may be provided as one or more virtual machines implemented in a hosted computing environment. The hosted computing environment may include one or more computing resources that are quickly provided and released. These computing resources may include computing, network connections, and / or storage devices. The hosted computing environment may also be referred to as a cloud computing environment. Service provider computers may include one or more servers, which may be arranged in a cluster, arranged as a server cluster, or as individual servers that are not associated with each other, and may host application 1032 and / or cloud-based software services. These servers may be configured as part of an integrated distributed computing environment. In some aspects, the service provider computer may additionally or alternatively include a computing device such as a cell phone, a smart phone, a personal digital assistant (PDA), a laptop computer, a desktop computer, a netbook computer, a server computer, a thin client device, a tablet computer, a gaming console, etc. In some cases, the service provider computer may communicate with one or more third-party computers.
[0213] In one example configuration, nodes 1002a-1002h may include at least one memory 1018 and one or more processing units (or processors 1020). Processor 1020 may be implemented in hardware, computer executable instructions, firmware, or a combination thereof. Computer executable instructions or firmware implementations of processor 1020 may include computer executable or machine executable instructions written in any suitable programming language for performing the various functions described.
[0214] In some cases, hardware processor 1020 may be a single-core processor or a multi-core processor. A multi-core processor may include multiple processing units within the same processor. In some instances, multi-core processors may share certain resources, such as a bus and a second or third level cache. In some cases, each core in a single-core or multi-core processor may also include multiple execution logical processors (or execution threads). In such cores (e.g., a core with multiple logical processors), several stages of the execution pipeline and lower level caches may also be shared.
[0215] The memory 1018 may store program instructions that may be loaded and executed on the processor 1020, as well as data generated during the execution of these programs. Depending on the configuration and type of the nodes 1002a-1002h, the memory 1018 may be volatile (e.g., RAM) and / or non-volatile (e.g., ROM, flash memory, etc.). The memory 1018 may include an operating system 1028, one or more data storage areas 1030, one or more application programs 1032, one or more drivers 1034, and / or services for implementing the features disclosed herein.
[0216] The operating system 1028 can support the basic functions of the nodes 1002a-1002h, such as scheduling tasks, executing applications, and / or controlling peripheral devices. In some embodiments, the service provider computer can host one or more virtual machines. In these embodiments, each virtual machine can be configured to execute its own operating system. Examples of operating systems include Unix, Linux, Windows, Mac OS, iOS, Android, etc. The operating system 1028 can also be a proprietary operating system.
[0217] The data storage area 1030 may include permanent or temporary data used and / or operated on by the operating system 1028, the application 1032, or the driver 1034. Examples of such data include web pages, video data, audio data, images, user data, etc. In some embodiments, the information in the data storage area 1030 may be provided to the user device via the network 1008. In some cases, the data storage area 1030 may additionally or alternatively include stored applications and / or drivers. Alternatively or in addition, the data storage area 1030 may store standard and / or proprietary software libraries, and / or standard and / or proprietary application user interface (API) libraries. The information stored in the data storage area 1030 may be machine-readable object code, source code, interpreted code, or intermediate code.
[0218] Driver 1034 includes programs that can provide communication between components in the node. For example, some drivers 1034 can provide communication between operating system 1028 and extra storage area 1022, network device 1024 and / or I / O device 1026. Alternatively or in addition, some drivers 1034 can provide communication between application program 1032 and operating system 1028 and / or application program 1032 and peripheral devices that can be accessed by service provider computer. In many cases, driver 1034 can include driver program (for example, printer driver program, display driver program, hard disk driver program, solid state device driver program) that provides easy-to-understand function. In other cases, driver 1034 can provide exclusive or special function.
[0219] The service provider computer or server may also include an additional storage area 1022, which may include a removable storage area and / or a non-removable storage area. The additional storage area 1022 may include a magnetic storage area, an optical disk, a solid state disk, a flash memory, and / or a tape storage area. The additional storage area 1022 may be housed in the same chassis as the nodes 1002a-1002h, or may be in an external housing. The memory 1018 and / or the additional storage area 1022 and its associated computer-readable medium may provide a non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computing device. In some embodiments, the memory 1018 may include a variety of different types of memory, such as SRAM, DRAM, or ROM.
[0220] Removable and non-removable memory 1018 and additional storage area 1022 are examples of computer-readable storage media. For example, a computer-readable storage medium may include volatile or non-volatile, removable or non-removable media implemented with a method or technology for storing information, the information including, for example, computer-readable instructions, data structures, program modules, or other data. Memory 1018 and additional storage area 1022 are examples of computer storage media. Additional types of computer storage media that may be present in nodes 1002a-1002h may include, but are not limited to, PRAM, SRAM, DRAM, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, DVD or other optical storage device, cassette, magnetic tape, disk storage device or other magnetic storage device, solid-state drive, or some other medium that can be used to store the desired information and can be accessed by nodes 1002a-1002h. Computer-readable media also include any combination of the above media types, including multiple units of one media type.
[0221] Alternatively or additionally, computer-readable communication media may include computer-readable instructions, program modules, or other data transmitted within a data signal such as a carrier wave or other transmission. However, as used herein, computer-readable storage media does not include computer-readable communication media.
[0222] The nodes 1002a-1002h may also include I / O devices 1026, such as keyboards, mice, pens, voice input devices, touch input devices, displays, speakers, printers, etc. The nodes 1002a-1002h may also include one or more communication channels 1036. The communication channels 1036 may provide a medium through which the various components of the nodes 1002a-1002h may communicate. The one or more communication channels 1036 may be in the form of a bus, a ring, a switching structure, or a network.
[0223] Nodes 1002a - 1002h may also contain a network device 1024 that allows nodes 1002a - 1002h to communicate with a stored database, another computing device or server, a user terminal, and / or other devices on the network 900 .
[0224] In some embodiments, the network device 1024 is a peripheral device, such as a PCI-based device. In these embodiments, the network device 1024 includes a PCI interface for communicating with a host device. The term "PCI" or "PCI-based" can be used to describe any protocol in the PCI family of bus protocols, including the original PCI standard, PCI-X, accelerated graphics port (AGP) and fast PCI (PCIe) or any other improvements or derivative protocols based on the PCI protocol discussed herein. The PCI-based protocol is a standard bus protocol for connecting devices such as local peripheral devices to a host device. The standard bus protocol is a data transmission protocol whose specifications have been defined and adopted by various manufacturers. The manufacturer ensures that the compatible device is compatible with the computing system implementing the bus protocol, and vice versa. As used herein, the PCI-based device also includes a device using non-volatile memory high speed (NVMe) communication. NVMe is a device interface specification for accessing non-volatile storage media attached to a computing system using PCIe. For example, a bus interface module can implement NVMe, and the network device 1024 can be connected to a computing system using a PCIe interface.
[0225] The modules described herein may be software modules, hardware modules, or a suitable combination thereof. If the module is a software module, the module may be embodied on a non-transitory computer-readable medium and processed by a processor in any computer system described herein. It should be noted that the described processes and architectures may be executed in real time or in an asynchronous mode prior to any user interaction. The modules may be configured in the manner recommended in the preceding figures, and / or the functions described herein may be provided by one or more modules that exist as separate modules, and / or the module functions described herein may be distributed over multiple modules.
[0226] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will, however, be evident that various modifications and changes may be made thereto without departing from the broader spirit and scope of the disclosure as set forth in the claims.
[0227] Other variations are within the spirit of the present disclosure. Therefore, while the disclosed technology is susceptible to various modifications and alternative constructions, certain illustrative examples thereof have been shown in the drawings and described in detail above. However, it should be understood that there is no intention to limit the present disclosure to the particular forms disclosed, but on the contrary, it is intended to cover all modifications, alternative constructions, and equivalents that fall within the spirit and scope of the present disclosure as defined in the appended claims.
[0228] Unless otherwise indicated herein or clearly contradicted by context, the use of the words "a / an" and "the" and similar indicators in the context of describing the disclosed examples (especially in the context of the appended claims) should be understood to cover both the singular and the plural. Unless otherwise indicated, the terms "comprising", "having", "including" and "containing" should be interpreted as open terms (i.e., meaning "including but not limited to") unless otherwise indicated. The term "connected" should be understood as partially or completely included, attached to or joined together, even in the presence of intermediates. Unless otherwise indicated herein, the recitation of the range of values herein is merely intended to serve as a shorthand method of individually referring to each individual value falling within the range, and each individual value is incorporated into the specification as if it were individually recited herein. Unless otherwise indicated herein or otherwise clearly contradicted by context, all methods described herein can be performed in any suitable order. Unless otherwise claimed, the use of any and all examples or exemplary language (e.g., "such as") provided herein is intended only to better illustrate the examples of the present disclosure and does not constitute a limitation on the scope of the present disclosure. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the present disclosure.
[0229] Unless specifically stated otherwise, disjunctive language such as the phrase "at least one of X, Y, or Z" is intended to be understood within the context as used to generally present that an item, term, etc. may be X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is generally not intended to, and should not, imply that certain instances require that at least one of X, at least one of Y, or at least one of Z each be present.
[0230] Various examples of the present disclosure are described herein, including the best mode known to the inventor for performing the present disclosure. After reading the above description, it will be clear to those skilled in the art that the variations of those examples will be present. The inventor expects that those skilled in the art will adopt such variations as needed, and the inventor intends to practice the present disclosure in other ways different from the specific description herein. Therefore, the present disclosure includes all modifications and equivalents of the subject matter described in the claims attached hereto as allowed by applicable law. In addition, unless otherwise indicated herein or otherwise clearly contradicted with the context, the present disclosure covers any combination of the elements described above with all possible variations thereof.
Claims
1. A non-transitory computer-readable medium storing instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to: loading a first weight data element of the array of weight data elements from a memory into the systolic array; selecting a subset of input data elements, the subset selected based on information indicating a rate of a dilated convolution operation and a coordinate of the first weight data element within the array of weight data elements, in, The subset is further selected based on a set of parameters including a starting source address of the subset in the memory, a gap between input data elements of the subset in the memory, and a number of elements of the subset; loading the subset of input data elements from the memory into the systolic array to perform a first computation of a dilated convolution operation; and The systolic array is controlled to perform the first calculation based on the first weight data element and the subset to produce a first output data element of an output data array.
2. The non-transitory computer-readable medium of claim 1, wherein the subset is selected based on a stride of the dilated convolution operation.
3. The non-transitory computer-readable medium of claim 2, wherein the input data elements are stored in a contiguous address space within the memory. 4 . The non-transitory computer-readable medium of claim 3 , wherein the subset of the input data elements is selected based on skipping a number of input data elements between each selected input data element, the number of skipped input data elements being based on the stride.
5. The non-transitory computer readable medium of claim 4, wherein the instruction includes a plurality of parameters, the plurality of parameters include: a starting source address of said subset of input data elements in said memory; a skip parameter indicating the number of input data elements to skip; as well as The number of input data elements selected.
6. The non-transitory computer-readable medium of claim 5, wherein the starting source address is based on the rate of the dilated convolution operation.
7. The non-transitory computer-readable medium of any one of claims 1 to 6, wherein the subset is a first subset; wherein the instructions, when executed by one or more hardware processors, cause the one or more hardware processors to: controlling the systolic array to perform the first calculation based on the first weight data element and the first subset to produce a first partial sum; controlling the summing buffer to accumulate the first partial sum; loading a second weight data element of the array of weight data elements from a memory into the systolic array; selecting a second subset of the input data elements from the memory into the systolic array; controlling the systolic array to perform a second calculation based on the second weight data element and the second subset to produce a second partial sum; and The summing buffer is controlled to accumulate the second partial sum to produce the first output data element.
8. The non-transitory computer-readable medium of claim 7, wherein a first size of the first subset of the input data elements and a second size of the second subset of the input data elements are selected based on a size of the summation buffer.
9. The non-transitory computer-readable medium of claim 8, wherein the instructions, when executed by one or more hardware processors, cause the one or more hardware processors to store the first partial sum at an address of a buffer based on a number of zero output data elements included in the first output data element.
10. The non-transitory computer readable medium of claim 9, wherein the instruction includes a set of parameters, the set of parameters include: a first starting source address of said first subset of said input data elements in said memory; as well as a second starting source address of said second subset of said input data elements in said memory; and Wherein the first starting source address and the second starting source address are offset from each other based on the rate of the dilated convolution.
11. The non-transitory computer-readable medium of claim 10, wherein the set of parameters includes a first starting destination address of the summing buffer for receiving the first partial sum and a second starting destination address of the summing buffer for receiving the second partial sum.
12. A device, include: a memory storing a set of instructions; One or more hardware processors configured to execute the set of instructions to: receiving first information indicating a rate and a stride of a dilated convolution operation to be performed by the systolic array based on the weight data array and the input data array to produce an output data array; receiving second information indicating a size of the summing buffer; determining, for each weight data element of the weight data array, a subset of input data elements of the input data array to be multiplied with the each weight data element to calculate a partial sum, the subset of input data elements being determined based on a projection operation from the summing buffer and based on the rate of the dilated convolution, the stride of the dilated convolution, the size of the summing buffer, and the coordinates of the each weight data element in the weight data array; determining, for each of the weight data elements of the weight data array, a destination address of the summing buffer to receive the partial sum based on the subset of input data elements; and generating a computation instruction for each weight data element of the weight data array to include third information indicating the destination address and the subset of input data elements, wherein the third information includes information indicating a starting source address of the subset of input data elements in the memory, a number of input data elements included in the subset of the input data array, and a stride in the computation instruction to enable the systolic array to load the subset of input data elements from the memory.
13. The apparatus of claim 12, wherein the one or more hardware processors are configured to execute the set of instructions on each weight data element to: determining a projection region aligned with a first input data element of the input data array based on the size information of the summing buffer and the stride of the dilated convolution operation; shifting the projection region relative to the first input data element of the input data array by an offset based on the coordinates of the each weight data element in the weight data array and the rate of the dilated convolution; aligning a stride pattern with the shifted projection area, the stride pattern being based on a stride of the dilated convolution operation; determining, based on the size information, coordinates of input data elements of the input data array that overlap the stride pattern; and The subset of input data elements for each weight data element is determined based on the coordinates.
14. The apparatus of claim 13, wherein the one or more hardware processors are configured to execute the set of instructions to determine the offset based on scaling the coordinates of each of the weight data elements at the rate of the dilated convolution.
Citation Information
Patent Citations
Image preprocessing for generalized image processing
US20190114499A1