System and Method for Memory Compression for Deep Learning Networks
By adapting bit widths and utilizing on-chip compression schemes that maintain data in a compressed state, the method efficiently addresses the inefficiencies in existing compression approaches for deep learning workloads, achieving substantial memory and energy reductions.
Patent Information
- Application Number
- JP2022569452
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-05-14
- Filing Date
- 2021-05-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-05-14
AI Technical Summary
Existing compression approaches for deep learning workloads face challenges due to the need to support random and fine-grained access patterns, which are not typical in neural networks, and fail to leverage the specific data distribution and access patterns of deep learning workloads, leading to inefficient memory usage.
A method and system for memory compression in deep learning networks that adapt the number of bits used for each element based on the value's magnitude, grouping values to use only the necessary number of bits, and employing on-chip compression schemes that maintain data in a compressed state until processing, utilizing the regular access patterns of neural networks to organize memory efficiently.
This approach significantly reduces the memory footprint and energy consumption by up to 49%, improves performance by 1.4 times, and decreases energy usage by 28%, while maintaining high processing unit utilization without requiring changes to the neural network model.
Smart Images

Figure 0007706173000006 
Figure 0007706173000007 
Figure 0007706173000008
Abstract
Description
Technical Field
[0001] The following generally relates to deep learning networks, and more specifically, to systems and methods for memory compression of deep learning networks.
Background Art
[0002] Compression of the memory hierarchy has received considerable attention, particularly in the context of general-purpose systems. However, there are various technical challenges in compression approaches for deep learning workloads. For example, general-purpose compression approaches typically need to support random and fine-grained access. Furthermore, programs in general-purpose systems typically exhibit patterns of values and various data types that do not exist in neural networks.
Summary of the Invention
[0003] In one aspect, a method for compressing the memory of a deep learning network is provided. The method includes defining, for a first memory of the deep learning network, a plurality of rows, each row having a specified number of columns, each column having a column width; receiving an input data stream processed by one or more layers of the deep learning network, the input data stream having a plurality of values with a fixed bit width; splitting the input data stream into subsets, the number of values in each subset being equal to the number of columns; compressing the data stream by sequentially compressing each subset, including identifying, for the values within a subset, the compressed bit width required to accommodate the value of the largest magnitude, storing the bit width in a bit width register associated with a row, and storing, in each column of the memory starting from the first available bit, the least significant bit of each value within the subset, the number of bits being equal to the bit width, and writing the remaining bits to each column of subsequent rows if more bits are required in each column of each row than those currently unused for storing the number of bits; and decompressing the compressed data stream to reproduce the input data stream by identifying the position of the first unread bit of each column of the compressed data stream, obtaining the bit width of each subset from each bit width register, taking out, from each column of the first memory, starting from the first unread bit of the column, the number of bits corresponding to the bit width, outputting the taken-out bits to the least significant bits of the output, updating the position of the first unread bit of each column so as to correspond to the bit position following the taken-out bits, and obtaining the reproduced input data value by zero-extending or sign-extending the remaining most significant bits of the output and sequentially outputting them.
[0004] In a particular case of this method, the position of a block of compressed values can be specified by one or more pointers.
[0005] In certain cases of this method, the block is a filter map data block or an input or output activation data block.
[0006] In certain cases of this method, the position is for the first compressed value of the block.
[0007] In certain cases of this method, one or more pointers include a first set of pointers to the data of the input or output activation map and a second set of pointers to the data of the filter map.
[0008] In certain cases of this method, receiving the input data stream includes sequentially receiving portions of the block starting at the position of one or more pointers, compressing the portions of the block, and updating an offset pointer to call the next portion to be received.
[0009] In certain cases of this method, receiving the input data stream includes sequentially receiving portions of the block, where the position of each portion is identified by one of the pointers.
[0010] In certain cases of this method, the portions of the compressed data values are forced to be stored starting from the least significant bit of the column by padding the most significant bit of the empty space of the previous data value.
[0011] In certain cases of this method, some rows of bit-width registers store a binary representation of the length of the bit width.
[0012] In certain cases of this method, other rows of bit-width registers store a single bit that specifies whether the bit width of the corresponding row is the same as or different from the previous row.
[0013] In a specific case of this method, this method is used to store floating-point values, where the floating-point value includes a sign part, an exponent part, and a mantissa part, the input data stream consists of the exponent part of the floating-point value, and compressing further includes storing, for each floating-point value, the sign part and the mantissa part adjacent to the compressed exponent part.
[0014] In a specific case of this method, during decompression, a pointer is established for the position of a specific one of the blocks known to be needed in the future.
[0015] In a specific case of this method, this method further includes tracking the next available position within each column of the first memory while compressing and storing the values.
[0016] In a specific case of this method, this method further includes initializing the first storage position of the first memory as available before compressing the data stream.
[0017] In a specific case of this method, a plurality of values have a fixed bit width that is less than or equal to the column width.
[0018] In a specific case of this method, the reproduced data stream is directly output to the arithmetic / logic unit.
[0019] In a specific case of this method, the reproduced data stream is output to a second memory having a plurality of rows each having a plurality of columns corresponding to the first memory.
[0020] In a specific case of this method, compressing further includes evaluating a function regarding the value of the input data stream to reduce the compressed bit width before identifying the compressed bit width, and reversing the function for decompression.
[0021] In another aspect, a method for memory decompression for a deep learning network is provided. The method includes obtaining a compressed data stream representing an input data stream, where the compressed data stream defines a plurality of rows with respect to a first memory of the deep learning network, each row having a specified number of columns and each column having a column width; receiving an input data stream to be processed by one or more layers of the deep learning network, where the input data stream has a plurality of values with a fixed bit width; splitting the input data stream into subsets, where the number of values in each subset is equal to the number of columns; compressing the data stream by sequentially compressing each subset, including identifying, for the values within a subset, the compressed bit width required to accommodate the value of the largest magnitude, storing the bit width in a bit width register associated with the row, and storing, in each column of the memory starting from the first available bit, the least significant bit of each value within the subset, where the number of bits is equal to the bit width and if more bits are required to store the number of bits than what is currently unused in each column of each row, the remaining bits are written to each column of subsequent rows; decompressing the compressed data stream to reproduce the input data stream, including identifying the first unread bit of each column of the compressed data stream, obtaining, from each bit width register, the bit width of each subset, extracting, from each column of the first memory starting from the first unread bit of the column, the number of bits corresponding to the bit width, outputting the extracted bits to the least significant bits of the output, updating the first unread bit of each column to correspond to the bit position following the extracted bits, and obtaining the reproduced input data value by zero-extending or sign-extending the remaining most significant bits of the output and sequentially outputting them.
[0022] In yet another aspect, a system for memory compression of a deep learning network is provided, the system comprising: a first memory having a plurality of rows, each row of the plurality of rows having a specified number of columns, each column having a column width; an input module configured to receive an input data stream processed by one or more layers of the deep learning network, the input data stream having a plurality of values of a fixed bit width, and to divide the input data stream into subsets, wherein the number of values in each subset is equal to the number of columns; a width detector module having a plurality of bit-width registers, each of the plurality of bit-width registers being associated with a row, and configured to identify, for the values within a subset, a compressed bit width required to accommodate the largest value and store the bit width in the bit-width register associated with the row; a compression module configured to store, in each column of the memory starting from the first available bit, the least significant bit of each value within the subset, wherein the number of bits is equal to the bit width, and if more bits are required to store the number of bits in each column of each row than are currently unused, writing the remaining bits to each column of subsequent rows; and a decompression module configured to decompress the compressed data stream and sequentially output the reproduced input data stream by: identifying the first unread bit of each column of the compressed data stream; obtaining the bit width of each subset from each bit-width register for the reproduced input data; extracting, from each column of the first memory starting from the first unread bit of the column, a number of bits corresponding to the bit width and outputting the extracted bits to the least significant bits of the output; updating the first unread bit of each column to correspond to the bit position following the extracted bits; and obtaining the reproduced input data value by zero-extending or sign-extending the remaining most significant bits of the output.
[0023] In certain cases of this system, the system further includes a pointer module having one or more pointers for tracking the positions of blocks of compressed values.
[0024] In certain cases of this system, the block is a filter map data block or an input or output activation data block.
[0025] In certain cases of this system, the position is for the first compressed value of the block.
[0026] In certain cases of this system, the one or more pointers include a first set of pointers to the data of the input or output activation map and a second set of pointers to the data of the filter map.
[0027] In certain cases of this system, the system further includes an offset pointer, and receiving an input data stream includes sequentially receiving portions of a block starting at the positions of the one or more pointers, compressing the portions of the block, and updating the offset pointer to call the next portion to be received.
[0028] In certain cases of this system, receiving an input data stream includes sequentially receiving portions of a block, and the position of each portion is identified by one of the pointers.
[0029] In certain cases of this system, portions of the compressed data values are forced to be stored starting from the least significant bit of a column by padding the most significant bit of the empty space of the previous data value.
[0030] In certain cases of this system, bit-width registers for several rows store binary representations of the bit-width length.
[0031] In a specific case of this system, the bit-width registers of other rows store a single bit that specifies whether the bit-width of the corresponding row is the same as or different from that of the previous row.
[0032] In a specific case of this system, this system is for storing floating-point values, where the floating-point value includes a sign part, an exponent part, and a mantissa part. The input data stream consists of the exponent part of the floating-point value, and compression further includes storing, for each floating-point value, the sign part and the mantissa part adjacent to the compressed exponent part.
[0033] In a specific case of this system, during decompression, a pointer is established for a specific one of the positions of the blocks known to be needed in the future.
[0034] In a specific case of this system, the compression module is configured to track the next available position within each column of the first memory while compressing and storing values.
[0035] In a specific case of this system, the compression module is configured to initialize the first storage position of the first memory as available before compressing the data stream.
[0036] In a specific case of this system, a plurality of values have a fixed bit-width that is less than or equal to the column width.
[0037] In a specific case of this system, the reproduced data stream is directly output to the arithmetic / logic unit.
[0038] In a specific case of this system, the reproduced data stream is output to a second memory having a plurality of rows each having a plurality of columns corresponding to the first memory.
[0039] In certain cases of this system, compressing further includes evaluating a function on the values of the input data stream to reduce the compressed bitwidth before identifying the compressed bitwidth and reversing the function for decompression.
[0040] These and other aspects are contemplated and described herein. It will be understood that the foregoing summary presents representative aspects of the embodiments to assist one of ordinary skill in the art in understanding the following detailed description.
[0041] A deeper understanding of the embodiments will be obtained by reference to the drawings.
Brief Description of the Drawings
[0042]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 8C
Figure 9A
Figure 9B
Figure 9C
Figure 10A
Figure 10B
Figure 11A
Figure 11B
Figure 12A
Figure 12B
Figure 13A
Figure 13B
Figure 14A
Figure 14B
Figure 15A
Figure 15B
Figure 16A
Figure 16B
Figure 17
[0043] Here, embodiments will be described with reference to the drawings. For the sake of simplicity and clarity of the description, when considered appropriate, reference numerals may be repeatedly used between the drawings to indicate corresponding or similar elements. In the following description, numerous specific details will be set forth in order to provide a thorough understanding of the various embodiments described. However, it will be well understood by those skilled in the art that the embodiments described herein can be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the embodiments described herein. Also, this description should not be regarded as limiting the scope of the embodiments described herein.
[0044] Any module, unit, component, server, computer, terminal, or device that executes a command can include or access a computer-readable medium, such as a storage medium, computer storage medium, or data storage device (removable and / or non-removable), for example, a magnetic disk, optical disk, or tape. Computer storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by an application, module, or both. Such computer storage media may be part of the device or may be accessible or connectable to the device. Any application or module described herein can be implemented using computer-readable / executable instructions that can be stored or held by such a computer-readable medium.
[0045] Compression in the memory hierarchy is particularly attractive for deep learning workloads and accelerators where memory access accounts for the majority of the overall energy consumption. Compression can bring technical advantages to the operation of the computer, particularly in the case of the deep learning network in the present invention. First, for example, compression can increase the effective capacity and bandwidth of the hierarchy, improve energy efficiency, and reduce the overall access latency. Specifically, compressing data at any level of the hierarchy can increase the effective capacity because fewer physical bits are required for each value during encoding. Second, accesses to higher levels of the hierarchy, which require more energy and time per access, are reduced, improving the effective latency and energy efficiency. Third, compression reduces the number of bits read and written per value, improving the effective bandwidth and energy efficiency. Furthermore, it complements data flow and blocks for reuse, which are the forefront techniques for enhancing the energy efficiency of the memory hierarchy. These advantages have motivated research on off-chip memory compression for neural networks. Embodiments of the present disclosure advantageously provide compression in the on-chip memory hierarchy.
[0046] Compression in the memory hierarchy has received considerable attention in the context of general-purpose computing systems. In compression for general-purpose computing systems, it is necessary to support any access pattern and generally depends on common value patterns (e.g., pointers or iterative values) in computer programs. However, the inventors have determined that deep learning workloads exhibit specific behaviors that present additional opportunities and technical challenges. For example, the access patterns of deep learning workloads are typically regular and composed of long sequential accesses. This reduces the merit of supporting random access patterns. In addition, the values of neural networks generally consist of feature maps and filter maps that do not exhibit the properties of common program variables. Furthermore, neural network hardware tends to be data-parallel and requires wide accesses.
[0047] To support random and fine-grained access, the ability to quickly and fine-grainedly find compressed values in memory is required. This requires using small blocks in general-purpose compression methods, which generally significantly suppresses the effective capacity. As a result, many compression approaches reduce the amount of data transferred, but do not reduce the size of the containers used in storage. For example, since the data within a cache line needs to be encoded, the number of bits for reading or writing needs to be reduced. However, the entire cache line still remains reserved. Alternatively, the method needs to carefully balance between flexible placement and metadata overhead using a certain level of indirection to identify where the data is currently in memory.
[0048] Typical programs tend to exhibit full or partial value redundancy. For example, due to the use of memory pointers, some values tend to share prefixes (e.g., pointers to structures allocated on the stack or heap). Programs often use aggregation data structures that tend to exhibit patterns of partially repeated values (such as flag fields). Usually, compression approaches need to handle various data types, such as integers and floating-point numbers, or characters in various character sets (e.g., UTF-16). Additionally, programs manage data types with various powers-of-two data widths, such as 8-bit, 16-bit, 32-bit, or more. Finally, programmers often use "default" integer or floating-point data types (such as 32-bit or 64-bit). Compression techniques can utilize these characteristics to reduce the footprint of the data.
[0049] In contrast, the inventors have determined that deep learning workloads tend to exhibit long sequential accesses even when blocking for reuse is being used. This reduces the need to support random access to fine-grained blocks. Further, the values of deep learning workloads generally do not exhibit repetitive patterns of typical computer programs. Most of the memory footprint is for storing large arrays of short data types such as 8-bit or 16-bit. Generally, given large amounts of data and computation, deep learning models carefully select data types to make them as small as possible. Quantization techniques to even smaller data types such as 4-bit can also be used. In some cases, there are still models that require 16-bit, for example, in the case of certain segmentation models, even a slight reduction in accuracy is converted to very noticeable artifacts. Further, programs tend to execute with narrow memory requirements, while neural networks generally exhibit data parallelism and prefer wide references.
[0050] Embodiments of the present disclosure advantageously provide an on-chip compression scheme in which data remains encoded as much as possible. In some cases, the data can be decompressed before processing elements of a deep learning approach, which supports a scheme with simple implementation, especially for decoding. Many of the compression techniques for general-purpose systems generally operate between the final level of cache and other caches of the on-chip hierarchy, where latency is not as critical and additional complexity can be tolerated. Advantageously, embodiments of the present disclosure can, for example, (1) support relatively long sequential accesses generally required by neural networks, (2) support multiple wide accesses to maintain a high utilization rate of processing units, (3) perform decoding immediately before the processing unit, so that the data can be kept compressed as long as possible, and (4) utilize the behavior of values typical of neural networks, to provide a lossless on-chip compression scheme.
[0051] Embodiments of the present disclosure (which may be referred to as "Boveda" for brevity) provide an on-chip memory hierarchy compression scheme that advantageously utilizes the typical distribution of values in a neural network that operates on fixed-point values. In particular, in each layer, most values tend to approach zero, so there are few high-magnitude values. Thus, embodiments of the present disclosure do not store all values using the same number of bits, but instead adjust the data width according to the content of the values and use only the necessary number of bits. If each value could select its data width individually, an unacceptable metadata overhead (a width field per value) would occur. Instead, embodiments of the present disclosure group the values and select a common data width that is wide enough to accommodate the largest-magnitude value within the group. For example, for a group of eight 8-bit (8-bit) values where the largest-magnitude value is 0x12, an 8×5-bit container can be used, while for another group where the largest-magnitude value is 0x0a, 8×4 bits can be used. In either case, a 3-bit metadata field specifies the number of bits used per value (5 and 4 respectively). Since variable data-width containers can be used, a wide crossbar is typically required to decode the values and properly align and feed them to the processing unit. For example, a processing element operating on eight 8-bit values each requires not only a 64-bit to 64-bit crossbar, but also additional logic to handle values that span two memory rows. Embodiments of the present disclosure utilize the regular access patterns of neural networks to organize the compressed data in memory and instead require multiple and much smaller "crossbars".
[0052] Advantageously, embodiments of the present disclosure can increase effective on-chip capacitance without requiring changes to the neural network model. This results in energy and / or performance advantages depending on whether the model is off-chip or compute-bound. An architect can deploy this embodiment at design time to reduce the amount of on-chip memory and thus the cost required to meet the desired performance goals. For neural network developers, this embodiment provides an approach that rewards quantization without the need to go off-chip and without requiring quantization for all models. In the present disclosure, experiments are applied to an accelerator for high-density models, an accelerator for sparse convolutional neural networks (SCNNs), and an accelerator targeting a pruned model to demonstrate that this approach is not specific to a particular accelerator architecture. For SCNNs, the experiments show that this embodiment can operate on top of zero compression of SCNNs. For illustrative purposes, the experiments use computer vision tasks, particularly image classification, to show the effectiveness of this embodiment. While this is only a small part of the vast number of domains where deep learning can be applied, the importance and value are very high due to the diversity and volume of applications where image processing systems are used. The experiments measured that this embodiment is as follows. · Reduced the footprint of the entire model by 49%. For models quantized using a specialized method, an almost ideal compression rate was achieved. In one method, by leveraging the value content, the compression rate was almost doubled compared to what is provided by specialized hardware. · Reduced the amount of bits accessed on-chip by 50%. · For a high-density accelerator with a 96 KB global buffer, the performance improved by 1.4 times and the energy improved by 28%. · Reduced the overall model footprint for zero compression of SCNNs by 66%. · Compared with the average 20% of the investigated configuration, the energy was reduced by 26% when combined with SCNN.
[0053] Referring now to FIGS. 1 and 2, there is shown a system 100 for memory compression for a deep learning network, according to one embodiment. In this embodiment, system 100 is executed on computing device 26 and accesses content on server 32 via a network 24 such as the Internet. In further embodiments, system 100 can be executed only on device 26, or only on server 32, or can be executed and / or distributed on other computing devices, such as, for example, a desktop computer, a laptop computer, a smartphone, a tablet computer, a server, a smartwatch, a distributed or cloud computing device(s), etc. In some embodiments, the components of system 100 are stored by and executed on a single computer system. In other embodiments, the components of system 100 are distributed among two or more computer systems that can be locally or remotely distributed.
[0054] FIG. 1 shows various physical and logical components of an embodiment of system 100. As shown, system 100 includes a central processing unit (“CPU”) 102 (including one or more processors), a random access memory (“RAM”) 104, an input interface 106, an output interface 108, a network interface 110, non-volatile storage 112, and a local bus 114 that enables the CPU 102 to communicate with the other components. The CPU 102 executes an operating system and various modules, as described in more detail below. The RAM 104 provides relatively responsive volatile storage to the CPU 102. The input interface 106 enables an administrator or user to provide input via input devices such as a keyboard and a mouse. The output interface 108 outputs information to output devices such as a display and / or a speaker. The network interface 110 enables communication with other systems such as other computing devices and servers located remotely from system 100, such as in a typical cloud-based access model. The non-volatile storage 112 stores an operating system and programs including computer-executable instructions for implementing the operating system and modules, as well as any data used by these services. As described below, additional stored data can be stored in the database 116. During operation of system 100, for ease of execution, the operating system, modules, and related data can be retrieved from the non-volatile storage 112 and placed in the RAM 104.
[0055] In one embodiment, system 100 includes several functional modules such as input module 120, decompression module 122, width detector module 126, compression module 124, deep learning (DL) module 128, and pointer module 130. In further embodiments, the functions of the modules can be combined or executed on other modules. In some cases, the functions of the modules can be executed at least partially on dedicated hardware, and in other cases, at least some of the functions of the modules can be executed on CPU 102.
[0056] The value distributions of the input feature map (imap) and the filter map (fmap) are generally highly skewed towards the smaller values. This behavior is what the system 100 can utilize to build a low-cost and energy-efficient compression technique. To utilize these distributions, the system 100 can adapt the number of bits (data width) used for each element to a length that is just sufficient to fit its current value. Since the fmap is usually static, the data width used varies by fmap element but is independent of the input. On the other hand, since the imap values are input-dependent, the data width used by the system 100 can adapt to the values taken by each element. In contrast, other memory hierarchies store all imap or fmap elements using a data width that is long enough to accommodate the possible values. However, as the inventors have empirically determined, this has proven to be excessive for most elements. For illustrative purposes, two models, ResNet18 (image classification) and SSD MobileNet (object detection), are highlighted, both of which are quantized to 8 bits. Figures 4A, 4B, 5A, and 5B show the regular and cumulative distributions of the imap values and fmap values for some representative convolutional and fully-connected layers. Figures 4A and 5A respectively show the imap value distribution and cumulative distribution for 64 randomly selected input batches, and Figures 4B and 5B show the distribution and cumulative distribution of the input-independent fmap values.
[0057] Figures 4A and 5A show that in res2a_branch1 of ResNet18, most imap values can be represented with 5 bits, which translates to a 37.5% reduction in footprint compared to the 8 bits used under ideal conditions. For substantially all imap values in the fully connected layer fc, only 4 bits, a 50% reduction compared to 8 bits, are sufficient. SSD Mobilenet shows similar behavior. In its 2D convolutional layer in the depth direction 12, 90% of the values require 6 bits or less, which is also sufficient to represent virtually all imap values in its object detection SSD module layer at point-by-point 13_2_2. Figures 4B and 5B show a similar trend for fmap. In res2 branch1 of ResNet18, only 5 bits are sufficient for most fmap values, but 6 bits are sufficient for virtually all values in the fc layer. However, in fc, up to 5 bits are required for 95% of the fmap values. The same is true for the fmap of SSD-MobileNet. Substantially all values fit within 6 bits, 90% fit within 5 bits, and over 80% fit within 4 bits.
[0058] In some cases, system 100 can be applied to an SCNN accelerator, which is an accelerator for the convolutional layers of a pruned CNN model. For illustrative purposes, system 100 is described as being applied to the convolutional layers of SCNN. However, it should be understood that it can be applied to other data-parallel deep learning accelerators and other types of layers such as fully connected layers.
[0059] Figure 6 is a diagram showing an example of a convolutional layer for illustrative purposes. The input is K fmaps of dimension S×R×C (height, width, channels) and an imap of H×W×C, where typically H≫S and W≫R, and the stride is s. The fmaps are statically known values (weights), while the imap is a value calculated at runtime (activation). The output is [Number] It is an omap (activation). In this example, s = 1 is assumed. Each omap value is determined as a 3D convolution of the fmap using a window of the same size as the imap for the fmap. Each fmap generates an omap value for one channel by sliding a window over the imap using a stride s along the H and W dimensions. The 3D convolution includes the multiplication of pairs of fmap elements and their corresponding imap elements, and then all these products are accumulated into the omap value. Each 3D convolution is equivalent to the sum of C 2D convolutions for each input channel.
[0060] SCNN stores values in the order of N samples-channel-height-width (NCHW), and the omap is determined by spatial input static convolution. This enables SCNN to process the imap and fmap one channel at a time, thereby allowing the utilization of sparsity. FIG. 7 is a diagram showing an example of the organization of SCNN tiles. SCNN scales up performance using a grid of such tiles. For the purpose of easy explanation and understanding, it can be assumed that there is only one tile. However, it is understood that the present embodiment of system 100 can be used for multiple tiles.
[0061] The tile has three buffers that hold imap (and omap), fmap, and an accumulator respectively. The accumulator accumulates the omap values. The SCNN uses a spatial data flow that executes all 2D convolutions for all windows of a single channel of imap at once. The SCNN is based on the observation that in the convolutional layer, the product of any imap value from the same channel as any fmap value contributes to some omap value. Thus, at maximum throughput, the tile processes four imap values and four fmap values from the same channel and computes the product of all 16 possible (imap, fmap) pairs. Next, all these products are sent via a crossbar to the corresponding accumulators. The accumulator buffer is organized into 32 banks to reduce the contention that occurs when multiple products are mapped to accumulators within the same bank. To exploit sparsity, imap and fmap store non-zero values as ((value), (skip)) pairs, omitting zero values. Here, (skip) is the number of zero values omitted after each. By using these (skip) fields, the SCNN infers the original position of each value and maps the products to the respective accumulators. For ease of explanation and understanding, the skip fields are omitted and 8-bit values are assumed. As described herein, 16-bit (original) and 8-bit SCNN configurations are considered.
[0062] Typically, the SCNN processes two consecutive blocks as follows. Consider four imap values (I0, ..., I3) and (I4, ..., I7), called BBlock0 and BBlock1 respectively. Note that the values within each block are conceptually ordered. I0 is the first value within BBlock0, and I7 is the first value within BBlock1. Initially, these can be assumed to be unsigned numerical values. FIG. 8A shows an example of a fixed data width buffer. In this example, the imap buffer of the SCNN uses 8-bit containers per value and supports a four-value width read (32 bits). In this configuration, the values read from the imap buffer align directly with the inputs to the multiplier. However, all values in BBlock0 have at least two zero-bit prefixes, and the values in BBlock1 have three bits. In contrast, one of the goals of system 100 is to avoid storing these prefix bits.
[0063] FIG. 8B shows an example of a simple approach that supports variable data widths. This approach is a simple way to store compressed values but is generally not desirable. For each BBlock of four values, a width field specifies the number of bits per value. In this example, it is 5 for BBlock0 and 6 for BBlock1 (encoded as 4(100) and 5(101)). One width field per BBlock amortizes its overhead for multiple values. In this example, the values are stored sequentially in BBlock0, and when BBlock0 is fully occupied, the values are stored sequentially in BBlock1.
[0064] Unfortunately, the values become misaligned with the multiplier input and may even span two rows, resulting in a high cost for decompression. For each column of the multiplier, it is necessary to extract the bits of the width (which varies for each BBlock), extend them to 8 bits, and then route them to the multiplier input. This routing requires an interconnection such as a 32-bit to 8-bit crossbar. Since there are four multiplier columns, four such crossbars are needed, significantly increasing the area and energy costs. If there are 8×8 multipliers in the multiplier grid, a 64-bit to 8-bit crossbar is required.
[0065] System 100 can advantageously implement an approach with much less complexity and cost. In one embodiment, the values can be treated as belonging to one of four groups called hileras corresponding to the multipliers, where the first value of each BBlock belongs to hilera0, the second value belongs to hilera1, and so on. The approach in FIG. 8B breaks this mapping and allows the compressed values to flow freely between hileras.
[0066] Instead, System 100 restricts the values to remain within their original hileras, as illustrated in the diagram of FIG. 8C. In this example, it shows that I0 and I4 are packed together in the hilera mapped to the first 8 bits of the buffer, while I3 and I7 are packed in the hilera mapped to the last 8 bits. For the sake of illustration, the values are used as bovedillas (bricks) to fill their hileras. The "crossbar" required in this example is 8-bit to 8-bit, and its size depends only on the maximum data width and not on the number of values read per cycle. For an 8×8 multiplier grid, eight 8-bit to 8-bit crossbars are required instead of eight 64-bit to 8-bit crossbars.
[0067] FIG. 9A shows an example diagram of the decompression module 122. The decompression module 122 decompresses a single value per cycle and properly processes those values that span two rows. The illustrated decompression module 122 shows one decompression block, but continuing with the example of FIG. 8C, four decompression blocks of the decompression module 122 operate in parallel and decompress four imap values per cycle. Reads from the imap buffer remain 32-bit wide. From each read, each block receives 8 bits corresponding to its hilera. Within each block, two 8-bit registers, register L and register R, hold the compressed data. Each time a new set of 8 bits is read, it is written to register L and simultaneously the current contents of register L are "copied" to register R. In some cases, instead of physically copying register L to register R, two bit pointers can be "exchanged" for use. In the steady state, since registers L and R contain two consecutive rows from one hilera of the imap buffer, they contain all the bits necessary to decompress 8-bit values regardless of the width. A 16-bit to 8-bit shifter extracts the current value from the value formed by concatenating the outputs of registers L and R. In this 16-bit example, the shifter only needs to support a shift specified by a 3-bit "offset" register ("OFS") that shifts up to 7 bits to the left. A 3-bit register ("W") holds the data width of the current BBlock. OFS and W, and the associated control logic, can be shared among all four decompression blocks. Initially, in this example, OFS = 0 and W = 7, both corresponding to the maximum data width. The "Bit-Extend" block passes the W least significant bits from the output of the shifter and sign-extends them to 8 bits. The compression block can operate as a two-stage pipeline, loading values into registers L and R in the first stage and extracting the next decompressed 8-bit value from the contents of registers L and R in the second stage. In this example, the decompression block requires a total of 3 cycles to decompress the first multiplier inferior I0 and I4 (additional cycles are required for the start interval). In the steady state, the decompression block can output a value per cycle.
[0068] Figures 9B and 9C show examples of cycles 2 and 3 for the above example of the decompression module 122. In cycle 1, the imap buffer provides the first set of 8 bits of the input data stream 0110 1100, which is written into register L. At the same time, W is loaded from the width memory with the data width 101 of BBlock0. OFS is updated to OFS=(OFS+W+1)mod8 = 0. Since OFS+W+1 exceeds 8 (carry from the adder), and register R contains no useful bits, the positions of register L and register R are swapped at the end of cycle 1, triggering a read from the imap buffer in the next cycle. In cycle 2, as shown in Figure 9B, the decompression block reads the next 8 bits and copies them into register L at the end of the cycle. Here, register L and register R contain two consecutive rows of compressed values from the same hilera, and thus are now in a steady state. During cycle 2, since OFS is 0, the 16-bit output of (register L, register R) is shifted by 0, thus aligning the least significant bit of the compressed I0 with the least significant bit of the output. The bit extension block passes the lower 6 bits and fills in the upper 2 bits accordingly under the guidance of W. In this example, since it is known that this imap only has positive values, the value is zero-extended to 8 bits. If the layer had signed imap values, the extender block would sign-extend instead. As a result, the value 0010 1100, the original I0, is sent to the multiplier. OFS is updated as before: OFS=(0+5+1)mod8 = 6. Since this does not exceed 8, the system does not read a value from the imap buffer in the next cycle. By the end of the cycle, a new width field is read into W. This is the width of BBlock1. In cycle 3, as shown in Figure 9C, OFS is used to instruct the shifter to slide (register L, register R) by 6 digits. Since W is zero-extended to 8 bits with 100, the extender block passes the 5 least significant bits. Next, OFS can be updated to (6+4+1)mod8. Since this exceeds 8, register L and register R are swapped, and the next imap row is loaded into L in the next cycle.
[0069] When the imap values and fmap values of all channels of the layer are processed, the output map is included in the accumulator. In most cases, the SCNN reads these values, passes them to the activation function, deletes the zero values, and copies the rest to the omap buffer (in some cases, the pointers are exchanged so that the omap buffer becomes the imap buffer for the next layer). The system 100 uses the output of zero compression. The number of values per BBlock can be selected by the user and / or designer. FIG. 10B shows an example of a compression block of the compression module 124 that processes four input values per cycle and has a B block size of 4 according to this embodiment.
[0070] FIG. 10A shows that an example of the compression module 124 can include three main components: (1) a width detector, (2) four compressor units (CUs), and (3) a 32-bit output register. The compression module reads four 8-bit values per cycle, encodes them into a BBlock, and stores them in the output register. When all 32 bits of the register are filled, it is sent to the omap buffer. A complete row can be generated per cycle. In some cases, the buffer rows can be output at a slower pace, because more values can be packed per row by compression. For this reason, the number of bits that need to be copied to the imap buffer is reduced, and energy can be saved.
[0071] The width detector module 126 identifies the bit width required to accommodate the value of the largest magnitude. For example, if the values are assumed to be positive (when using ReLU), the width detector module 126 first generates eight signals, one for each bit plane, which are the OR of all corresponding bits spanning four values. Next, the eight signals pass through a leading 1 detector module that identifies the most significant bit that is 1 among all the values. This is the width required by the BBlock. If there can be signed values in the layer, they can be inverted before the leading 1 detector (in the case of negative numbers, the detector determines whether the most significant bit is zero). In this case, one more bit is required for the sign. Whether the map can be negative is known statically. The detected width can be written to a width buffer. Thus, for data values that can potentially include negative values, the values can be sign-extended after unpacking based on the value of this sign bit. Positive values can be extended to full width by adding a zero bit in the most significant position, while negative values determined by a sign bit of 1 can be extended using the bit of value 1.
[0072] Figure 10B shows the structure of the compression block of the compression module 124 that substantially reflects the decompression module 122. Optionally, there may be one compression block per hilera. Register L and register R hold the current and next rows of the hilera. The compression block processes values every cycle. It extracts the least significant bits of its width (detector) and stores them in register R at the appropriate position via the "shift and mask" block. If the value requires more bits than the bits that are currently unused in register R, the remaining bits are written to register L. When register R is full, it is copied to the output row register (component (3)), and the two registers are swapped using a single-bit pointer (not shown). The 3-bit continuation register specifies at which bit position to continue filling register R. The shift and mask block includes an 8-bit to 16-bit shifter that needs to support a maximum shift of 7 digits to the right. In most cases, the system does not need to shift more than 7 bits. This is because there are no free bits left in register R and they would be written out.
[0073] In some cases, SCNN can store values in the order of N.SamplesChannel-Height-Width (NCHW). In this way, SCNN adjusts the size of the on-chip buffer so that the imap and omap for each layer fit into the on-chip buffer, and reads the fmap from off-chip in channel order. If there are multiple tiles, each imap channel is mapped to the tiles in parts of the same size, and the fmap is broadcast. The part of the imap assigned to each tile depends only on the dimensions of the layer. However, since SCNN uses zero compression, the number of imap values contained in each part is different. System 100 can use these characteristics for compression that can be used to further compress the data. The processing can still start from the beginning of the imap buffer. When a value is written to the output of a layer, the value is placed from the first position of the local omap buffer (SCNN for each layer exchanges the imap with the omap so that the omap of the previous layer becomes the imap of the next layer).
[0074] The DL module 128 operating with SCNN first stores the fmap channels, packs all the fmap values together, starting with fmap0, the values of channel 0, then fmap1, the values of channel 0, and so on. During processing, the tile cycles through all the fmap values of channel 0, then all of channel 1, and so on. Since the dimensions and count of the fmap are statically known to the DL module 128, it can determine when it reaches the end of each channel and count the number of values processed and the number of zeros skipped.
[0075] SCNN uses skip fields per value to eliminate zeros. Since skip fields are only used in the control logic of the tile (e.g., to determine the original position of the value), it may be better to store them in a separate structure next to the control logic rather than near the data path. DL module 128 extends this buffer to also store the width field per BBlock. In one example, when skip fields for 3-bit and 8-bit values are assumed, the width field requires 3 bits of overhead per BBlock, or less than 7% bit-wise overhead when using a BBlock of 4 values. The overhead for an 8-value BBlock is halved.
[0076] FIG. 11A shows an example of a system 100 used in a data parallel accelerator targeting a high-density model (i.e., not leveraging sparsity to improve performance). The accelerator has a global buffer to avoid off-chip access and a grid of processing elements (PEs). The PE shown in the example of FIG. 11B can process 16 (imap, fmap) value pairs per cycle, all accumulated into the same omap. Each PE has its own local imap, fmap, and omap buffers. Optionally, the conversion block first converts the values read from off-chip into a usable form and then writes them to the global on-chip buffer (and vice versa). The local buffers of the PEs read values from the global buffer and are thawed at that point. Omap values are compressed before being written to the global buffer. The width field is stored in a separate bank and address space of the global buffer.
[0077] There are advantageous differences when compared to a mere SCNN implementation. This is due in part to the need to support a set of diverse data flows and in part to the need to support primarily high-density models. In the model, (a) the on-chip implementation does not implement zero compression, and (b) to support a diverse set of data flows, support is needed to block access to imap and fmap at various levels, and thus, depending on the data flow requirements, the starting point of each reuse block can be found.
[0078] To support data flows other than zero compression, additional support is needed because system 100 changes the mapping of values to memory. If all values are of the same length, system 100 can directly index any value within imap, fmap, and omap. Since system 100 compresses these values, their positions in memory become content-dependent. Pointer module 130 can use pointers to support the blocking scheme of the selected data flow. Generally, the number of pointers required is small, and when the data is compressed on-chip or off-chip, only a few pointers need to be explicitly stored. Most pointers can be generated in a timely manner during processing and discarded once used. This is possible because (a) the data flow uses blocking to maximize reuse, and (b) as processing proceeds according to the data flow, system 100 naturally encounters the starting position of the next reuse block to be processed. This approach will first be described in the context of the fully connected layer and then for the convolutional layer. It is understood that it can be applied to any suitable layer type.
[0079] In most cases, a fully-connected layer receives one imap and K fmaps as inputs and generates an omap with the same number of elements as the fmap. Both the imap and the fmaps have the same number of elements C. Each of the K omap elements is the inner product of one of the imap and one of the fmaps. The system can utilize on-chip imap reuse access. For the purpose of explanation, consider that the PE has only one accelerator. If the imap fits on-chip, after reading the imap from off-chip once, the fmaps can be cycled through. In this case, the access to the imap and each fmap becomes sequential. If the imap is too large to fit on-chip, the system 100 can use blocking, where only a portion (reuse block) of the imap is loaded on-chip at any time while the system cycles through the corresponding portion of the fmaps. The resulting on-chip access pattern remains continuous for each reuse block. When the system 100 finishes processing the current imap reuse block, it can move on to the next imap reuse block. Thus, for a fully-connected layer, the system 100 generally only needs to support sequential access to relatively long blocks of the imap or fmaps. If the values are not compressed, the start position of each reuse block is a linear function of the block size and its relative position. In most cases, these positions depend on the content of the values. Since the access pattern is sequential, the DL module 128 reaches the start of each reuse block in order as required by the data flow. Thus, in most cases, the pointer module 130 only needs to maintain a single access pointer for each fmap and for the imap. If there are multiple PEs, the map can be divided into smaller reuse blocks that the DL module 128 can process simultaneously. Next, the system 100 requires the same number of pointers as the number of reuse blocks that need to be processed simultaneously, which can be stored as additional metadata for the layer.
[0080] Using N.Samples-Height-Width-Channel (NHWC) memory mapping can enhance the locality of data in convolutional layers. Compared to fully connected layers, an additional challenge in convolutional layers is the need to be able to initiate access to multiple, often overlapping windows. Without loss of generality, consider a channel-first output stationary data flow where each window is processed in the order of channel, width, and height. The term "column" can be used to refer to all imap values with the same (width, height) coordinates. To determine a single omap, the data flow can sequentially access the values within a column and then access other columns in the order of width and height. Boveda can sequentially group the values into BBlocks along each column according to the NHWC mapping.
[0081] A technical challenge for system 100 is that the starting position of each column generally ceases to be a linear function of its (width, height) coordinates. A simple solution is to hold pointers to each column (the 2D coordinates of the first channel). This is excessive because (a) each column is needed during the processing of some windows (for example, in the case of a 3×3 fmap, each column is accessed 9 times), and (b) the windows usually overlap, so the starting position of each column is detected during the processing of the previous window. Therefore, the pointer module 130 reduces the number of pointers explicitly stored as metadata while "restoring" the rest during processing and holding them only for the necessary period. The number of pointers that need to be stored along the imap depends on the dimensions of the imap and fmap, and the number of windows. In one example,
Number
[0082] To maintain the function of performing reads in a wide enough range as required for high PE utilization, the start positions of some BBlocks can be restricted to align with the lines in the on-chip memory. In some cases, the first values of all S columns (S is the stride) of all fmaps and imaps can be restricted to align with the beginning of the memory line. Therefore, padding may be required. However, this padding does not increase the footprint compared to the case where values are not compressed in order to minimize the effective compression rate.
[0083] The system 100 can be applied to other layers, such as depthwise separable convolution or pooling. Since each BBlock can be decoded in parallel, the system 100 may need to store pointers for parallel processing × block size to start parallel processing in parallel.
[0084] In addition to reducing the pointer overhead, system 100 can also reduce the group overhead. In the original design, log2(bit width) bits of value are used to store the BBlock size, but this can be further reduced based on the observation that the value of the BBlock size tends to repeat. System 100 can use extra bits for each BBlock to detect whether the size of the BBlock is the same as the previous one. In that case, there is no need to read the new size from memory. Thus, the new BBlock size has an overhead of 1 bit + log2(bit width) bits, and the repeated size has an overhead of 1 bit.
[0085] Advantageously, in various embodiments, system 100 can be inference-targeted and lossless and transparent. It can depend on the expected distribution of all values and benefit from sparsity, but does not require it.
[0086] Some neural networks exhibit spatial correlation of values, resulting in values within the same BBlock having similar magnitudes. In such cases, it is advantageous to perform a function on this value to reduce the amount of data that needs to be stored. For example, it may be advantageous to first represent all values as the difference from a common bias value. A suitable choice for the bias can be, for example, the maximum value or a constant within the BBlock. If the difference is much smaller than the original value, this approach results in fewer bits being used per packed value. The bias can be stored in an additional optional field. Other functions than the difference may be used.
[0087] In some neural networks, a floating-point representation of numbers is used. This representation uses a triplet (sign, exponent, mantissa). For example, in a common representation, 32 bits are used where the sign is 1 bit, the exponent is 8 bits, and the mantissa is 23 bits. Using this method, after removing the bias, the length of the exponent can be adjusted dynamically. For example, in the case of a block of four floating-point values (a, b, c, d) where the exponents are Ea, Eb, Ec, and Ed respectively, the encoded block can instead store (Ea - bias, Eb - Ea, Ec - Ea, Ec - Ed). The width field in this case encodes the number of bits required to represent the maximum value of the values within the encoded block. The bias is a constant defined by the floating-point standard. A set of adders after decoding can restore the original block (Ea, Eb, Ec, Ed) after decoding (Ea - bias, Eb - Ea, Ec - Ea, Ec - Ed). During compression, a subtractor before the compression unit can be given the original (Ea, Eb, Ec, Ed) and the bias to calculate (Ea - bias, Eb - Ea, Ec - Ea, Ec - Ed). Optionally, the mantissa can be stored using a global common width without requiring an additional width field.
[0088] In other approaches such as the Efficient Inference Engine (EIE), deep compression is used to significantly reduce the fmap size of fully connected layers. Deep compression is very specialized as it modifies the fmap to use a limited set of values (e.g., 16) and uses Huffman coding and look-up tables to decode the values at runtime. In contrast, this system can operate on "ready-to-use" neural networks.
[0089] In other approaches such as compression of DMA, zero values off-chip can be removed using per-block bit vectors. In contrast, in various embodiments, the system can target on-chip compression and all values. In other approaches such as Extended BitPlane Compression (EBPC), off-chip compression that combines zero-length encoding and bit-plane compression can be used, especially for pruned models. The EBPC decompression module requires 8 cycles per block of 8 eight-bit values. In contrast, in various embodiments, the system can benefit from both high-density and sparse networks and decompresses blocks every cycle. In other approaches such as ShapeShifter, off-chip compression can be used that adapts the data container to the content of the values and uses zero bit vectors. The ShapeShifter containers are stored sequentially in memory space regardless of alignment. Decompression per block is performed sequentially on the values one block at a time. Thus, ShapeShifter is not suitable for on-chip compression. Other approaches such as Diffy extend ShapeShifter by storing values as deltas. Diffy targets computational imaging neural networks that exhibit high spatial correlation of imap values. Diffy has a significantly higher computational cost than embodiments of this system because delta calculations are required for encoding and decoding. In other approaches such as Proteus, per-layer data widths derived from the profile can be used to store values both on-chip and off-chip. Thus, the skewed distribution of values within a layer cannot be utilized, and the maximum size per layer determines the width of all values. Embodiments of this system can be used to adapt the data width at a substantially finer granularity.
[0090] Figure 3 shows a flowchart of a method 300 for memory compression for a deep learning network, according to one embodiment.
[0091] At block 302, input module 120 receives an input data stream that is processed by one or more layers of the deep learning model.
[0092] At block 304, width detector module 126 determines the bit width required to accommodate the values from the input data stream having the largest size.
[0093] At block 306, compression module 124 stores the least significant bits of the input data stream in a first memory store (such as register "R"). The number of bits is equal to the bit width. If the value requires more bits than the bits that are currently unused in the first memory store, the remaining bits are written to a second memory store (e.g., register "L").
[0094] At block 308, when the first memory store is full, compression module 124 outputs the values of the first memory store as a consecutive portion of the compressed data stream, along with the relevant width of the data within the first memory store. Compression module 124 copies the values of the second memory store to the first memory store.
[0095] At block 310, decompression module 122 receives data from the compressed data stream having respective widths, moves the data from the first memory store to the second memory store, and the first memory store contains the previously stored data from the compressed data stream.
[0096] At block 312, decompression module 122 stores each bit of the compressed data stream in the first memory store having a length equal to the width of the first memory store.
[0097] At block 314, decompression module 122 concatenates the data in the first memory store and the second memory store.
[0098] At block 316, the decompression module 122 outputs the concatenated data, and the concatenated data has a width equal to the relevant width of the concatenated values received from the compressed data stream.
[0099] The inventors conducted experimental examples to evaluate the technical advantages of this embodiment. In the experimental examples, a custom cycle-accurate simulator was used to model the execution time and energy. The simulator modeled off-chip memory access using DRAMSim2. All accelerators and hardware modules were implemented in Verilog, synthesized with Synopsys Design Compiler, and placed with Cadence Innovus for TSMC's 65nm cell library due to the licensee's constraints. The power was estimated via Innovus using the circuit activity reported by Mentor Graphics ModelSim. CACTI was used to model the area and power consumption of the on-chip memory. All accelerators operated at 1GHz, which matched the CACTI speed estimate of the on-chip memory. Table 1 shows the investigated network models and the footprints of fmap and imap. Most models were quantized to 8 bits. Some models used more aggressive quantization. Originally, these models were developed in combination with specialized architectures.
Table 1
[0100] The experimental examples demonstrated that this embodiment provides the highest possible memory benefit without requiring method-specific hardware. These models include the following: · Intel's INQ. Its fmap values are limited to signed 2 to the 16th power or zero. 16 bits are required to represent the weights as magnitudes, but 5 bits were sufficient with specialized hardware. · PACT. A modified ReLU with configurable saturation thresholds was required, and 4-bit imap and fmap were used for all but the first and last layers which used 8 bits. Outlier-aware quantization aggressively reduced the number of bits for most values (e.g., 4 bits), except for some large values (8-bit outliers) that were processed individually. · Intel's Skim Caffe repository and the MIT's Eyeriss group (since SCNN is generally excellent for pruned models).
[0101] The experimental examples included validating the system with respect to a high-density model accelerator with 256 processing engines organized in 16×16 rows. Each processing engine executed 8 MACs in parallel and generated a single value. Each PE had 64-entry imap, fmap, and omap buffers. The system used 8 BBlock sizes. A 32-bank global buffer fed the processing engines.
[0102] Figure 12A shows a chart reporting the memory footprint of the entire neural network. The footprint is measured in bits, and the figure reports the footprint of the system relative to the baseline. Boveda uses memory to store a) the encoded values, b) the metadata per BBlock width, c) padding due to memory alignment, and d) pointers. On average, the system reduces the footprint by 49%. The merits of SSD-MobileNet and MobileNet are the least at 16%, but considering that off-chip access is prohibitively expensive, that is still a significant amount. The models with specialized quantization are highlighted in Figure 12A and demonstrate an ideal memory footprint, where the memory hierarchy is specially designed for them. This system reduces the footprint to within 4% of what is ideally possible. In the case of ResNet18-PACT, the system reduces the footprint much more than what was possible with 4-bit hardware. This is because the system exploits the actual value content.
[0103] The system increases the information content per bit of on-chip storage. Thus, less data needs to be fetched by the processing engine from the on-chip hierarchy. Figure 12B shows a chart depicting this reduction in traffic. Without this system, access reads only data, but with this system, access can also read metadata. Thus, two measurements are shown: a) access, and b) bits transferred, both normalized to the baseline. The system executed 62% less transfers on average and transferred 50% fewer bits in total. As expected, most of the accesses were to fmap and imap. The reduction in bit traffic was less than the reduction in accesses due to metadata. The observed trend is similar to the overall footprint trend. This reduction can be directly translated into energy savings.
[0104] The main design choice when designing an accelerator is the amount of on-chip storage to use. Increasing the on-chip memory reduces the frequency of data fetches from off-chip. For example, the on-chip buffer of the SCNN is sized such that there is little need to offload the feature maps. In the experimental example, four policies regarding the sizing of the on-chip capacity were investigated. a) imap, omap, and fmap of the largest layer, b) complete rows of fmap and windows from imap, and c) complete rows of windows from imap and fmap per processing engine. With policy (a), only the input and the final output were off-chip. With policy (b), it was guaranteed that each value was accessed once from off-chip per layer. With policy (c), only one access per layer was guaranteed for imap and omap. Also, (d) layer fusion that processes a subset of multiple layers without going off-chip for intermediate i / omap values was considered.
[0105] Figure 13A is a chart showing the on-chip memory capacity required under each of the above sizing policies. The capacity is normalized to the baseline (different for each policy) under the same policy. Overall, the reduction in required storage was closely linked to the compression ratio. In one case, with SSD-MobileNet using the first policy (all layers on-chip), no reduction was possible. There was a single layer where the system did not reduce the overall on-chip data volume. Even so, there were energy and performance benefits because the system reduced the overall model traffic and footprint. Regardless of the access policy used, the system reduced the frequency with which the accelerator had to go off-chip. Figure 13B shows a chart of off-chip traffic for each model with (solid line) and without (dotted line) the system. For clarity, only a subset of the network is shown. The traffic was normalized so that, where possible, all values were accessed once per layer. As the on-chip memory size increased, the traffic approached this minimum value. With this system, a smaller on-chip memory can be used. Furthermore, for a given memory capacity, the system reduces off-chip traffic. For example, in the case of SegNet, 512 KB of on-chip storage was insufficient to achieve minimum traffic without using this system. Using 32 KB of on-chip storage, the system reduces off-chip traffic by 3.8 times for ResNet18 (1.48 times for traffic with the system and 5.66 times without the system for one read of the value), and by 2.6 times for ResNet50S OA.
[0106] In the experimental example, the performance of three configurations using on-chip global buffers of 96 KB, 192 KB, and 256 KB was measured. All used DDR4-3200 dual-channel off-chip memory. Figure 14A shows a chart indicating the speedup normalized with respect to the baseline with a 96 KB global buffer. This system improves the performance by 1.4 times, 1.2 times, and 1.1 times on average respectively. The improvement is highest in SegNet where the convolutional layer is quite large and the system compresses the data significantly. The advantages of this system are also prominent for MobileNetV2-OA, MobileNet, and ResNet18-INQ, and the system can avoid the outflow from the chip in some layers. Since the on-chip layer of the system can maintain the peak execution bandwidth of the baseline, the performance advantage of the system is obtained from the reduction of off-chip traffic. Figure 14A also shows the relative energy of the same memory configuration. The system saves an average of 28%, 16%, and 10% of energy in the 96 KB, 192 KB, and 256 KB configurations respectively. These advantages are due to the low off-chip and on-chip traffic. As the on-chip capacity increases, off-chip access decreases, and their overall energy cost also decreases.
[0107] Table 2 shows the area and power for compression and decompression. The width detector module 126 is shared for each BBlock. The total area overheads of the on-chip configurations of 96 KB, 192 KB, and 256 KB are 6.7%, 3.8%, and 3.2% respectively. However, even if this area is spent on additional memory for the baseline, the system is 1.29 times, 1.15 times, and 1.1 times faster on average, and since the cost of on-chip access for the baseline is small, the energy efficiency is slightly higher.
Table 2
[0108] SCNN used zero compression both on-chip and off-chip. For the 16-bit network, SCNN used 4-bit zero skip indexes. In the experimental example, the system used 3-bit indexes instead of the 8-bit network to reduce the metadata overhead. It was found that this did not affect the number of zeros removed. In this case, the system does not compress the zero skip index. Figure 14B shows a chart demonstrating the reduction of the overall model footprint using a system that goes beyond SCNN's zero compression. This system reduces the memory footprint by an average of 34% compared to zero compression. SCNN usually sets the size of the on-chip memory to fit all the imap on-chip of AlexNet and GoogLeNet. With this configuration, larger networks such as ResNet50 let data flow out off-chip. Furthermore, the size of the accumulator limits the number of resulting omap values and the number of concurrent filters. By amplifying the on-chip storage capacity, the system reduces the outflow. These effects were investigated for three different configurations per PE imap / accumulator, 10KB / 6KB, 4KB / 4KB, and 2KB / 2KB like SCNN. The off-chip memory used two channels of DDR4-3200. The area overheads of these configurations were 3.1%, 2.3%, and 1.8% respectively for SCNN 16-bit. The overhead is smaller for SCNN 8-bit.
[0109] Figure 15A shows a chart indicating the speedup for the 2KB / 2KB configuration with and without using the embodiment of the present system. In the 2KB / 2KB configuration, the system performance was improved by 29%. In the recent ResNet50 model, since its imap is large, the improvement by the system was more remarkable. In the 10KB / 6KB configuration, the system performance was improved by 15%. Figure 15A shows that the energy decreased by 26%, 24%, and 20% on average for the three configurations respectively. The experimental examples show that the system always reduced energy. In compute-bound models such as GoogLeNet and ResNet50, more advantages were seen because on-chip traffic accounted for a higher proportion of the total energy.
[0110] In the experimental examples, it was demonstrated that the system can also bring benefits to the first-generation tensor processing unit (TPU). The TPU incorporated 28MB of on-chip imap memory and streamed fmap from off-chip DRAM using weight-fixed dataflow. A 256×256×8-bit systolic array computed the omap. The fmaps were left compressed in DRAM and the on-chip buffer decompressed them just before the systolic array. Similarly, the imaps remained compressed in on-chip DRAM and were decompressed by the systolic dataset setup unit. Figure 16A is a chart showing the breakdown of the TPU's memory energy with and without the system for 16 BBlocks. The system on the TPU had a negligible area overhead of less than 0.1%.
[0111] Initially, the model used 16-bit fixed-point numbers, but now 8 bits have become the standard for many models. To further investigate the potential effectiveness of the system for narrower data types across a broader set of models, in the experimental example, synthetic 6-bit, 4-bit, and 3-bit networks were generated by scaling existing 8-bit layers to fewer bits while maintaining the original relative distribution of values (linear quantization). Figure 16B is a chart showing the ideal compression rate for a representative subset of these layers compressed with 8 BBlocks. The results show that the system remains effective for 4-bit layers. For 3-bit layers, sometimes the system may not be able to reduce the footprint or may even increase it, but generally there are still computational advantages.
[0112] Generally, the compression rate of the system depends on the distribution of values and is given by the following formula.
Equation
[0113] Figure 17 is a chart showing the footprint reduction for optimized BBlock sizes against overhead. On average, repeating the group optimization reduces the BBlock size overhead by an average of 28%. ResNet18-PACT has an optimal reduction of 58% because the BBlock size for 4-bit values is likely to be repeated.
[0114] Figure 15B is a chart showing the footprint reduction of frequently occurring pattern compression (FPC) and Base-Delta-Immediate (BΔI), which are cache compression schemes for general-purpose systems. Both target the width of values, among other properties. FPC was motivated by the observation that programmers tend to use 32-bit variables regardless of the actual range of values needed. FPC detects whether a value can be stored in a container of size a power of two (4 bits minimum). On average, it reduces the footprint by 18%. This is due to almost zero deletion. BΔI takes advantage of the low dynamic range of values within a program (adjacent values tend to be close in value). It operates on 64-byte chunks and shrinks the width at byte granularity. It represents values as differences of 4, 2, or 1 byte from an initial value of zero or 8, 4, or 2 bytes. All-zero chunks are represented as 1 byte and metadata. This byte granularity is too large for neural networks. At best, it reduces the footprint of ResNet50S-OA, which uses zero values, by 7%.
[0115] In the experimental example, System B△I, which is a variation of the system incorporating elements of B△I, was evaluated. This applied a compression method for each value of B△I, but at a smaller granularity. The compression options had all bits zero, and delta sizes of 8 bits, 4 bits, and 2 bits. This enabled packing the values into a hilera so that decompression could be processed in parallel without requiring a large crossbar in the output. The base was always set to 1 byte, and the values in the working set were reduced to BBlocks of 8. The system achieved an average compression of 44% using B△I, ignoring the overhead of width and pointer metadata. This is close to what a system without using B△I achieves. However, decompression of values in the system using B△I was considerably more complex and required more energy. For example, to decompress a block, 8 additions had to be performed in parallel, and the base had to be broadcast to all of them. Compression was also more involved, running all compression possibilities in parallel before selecting the optimal one. Systems not using B△I have a high compression rate and are also simple to implement.
[0116] Furthermore, the experimental example was compared with run-length encoding and dictionary-based compression that utilize the content of the values. Run-length encoding was limited to 8 values, the dictionary table was limited to 8 entries, and the overhead of 8-bit values was avoided. Both of these approaches achieved a lower compression rate compared to this system, but required an expensive crossbar for decompression.
[0117] The experimental example shows that this embodiment is easy to implement and provides an effective on-chip compression technique for neural networks. This reduces on-chip traffic while increasing the effective on-chip capacity. As a result, the amount of on-chip storage required to avoid excessive off-chip access is reduced. Furthermore, for a given on-chip storage configuration, the frequency of off-chip access is decreased.
[0118] Although the present invention has been described with reference to specific embodiments, various changes and modifications will be apparent to those skilled in the art without departing from the spirit and scope of the invention as set forth in the claims appended hereto.
Claims
1. A method for memory compression for a deep learning network, the method comprising: defining a plurality of rows for a first memory of the deep learning network, each of the plurality of rows having a number of columns, each column having a column width; receiving an input data stream processed by one or more layers of the deep learning network, the input data stream having a plurality of values of a fixed bit width; dividing the input data stream into subsets, the number of values in each subset being equal to the number of columns; compressing the input data stream by sequentially compressing each subset; identifying, for the values within the subset, a compressed bit width required to accommodate the value of the largest magnitude; storing the compressed bit width in a bit width register associated with the row; storing the least significant bit of each value within the subset in each column of the first memory starting from the first available bit, the number of bits in which the value can be stored being equal to the compressed bit width, and if more bits are required to store the bits in which the value can be stored in each column of each row than those currently unused, the remaining ones of the bits in which the value can be stored are written to each column of subsequent rows; comprising; comprising The compressed data stream is decompressed to identify the position of the first unread bit of each column of the compressed data stream; a reproduced input data stream; obtaining the compressed bit width of each subset from each bit width register; retrieving, from each column of the first memory, starting from the first unread bit of the column, a number of bits corresponding to the compressed bit width and outputting the retrieved bits to the least significant bit of the output; updating the position of the first unread bit of each column to correspond to the bit position following the retrieved bits; zero-extending or sign-extending the remaining most significant bits of the output to obtain the reproduced input data value; sequentially outputting by; A method by which the input data stream can be reproduced.
2. The method according to claim 1, wherein the position of the block of compressed values can be specified by one or more pointers.
3. The method according to claim 2, wherein the block is a filter map data block or an input or output activation data block.
4. The method according to claim 2, wherein the position is for the first compressed value of the block.
5. The method according to claim 2, wherein the one or more pointers include a first set of pointers to data of an input or output activation map and a second set of pointers to data of a filter map.
6. The method according to claim 2, wherein receiving an input data stream includes sequentially receiving portions of the block starting at the positions of the one or more pointers, compressing the portions of the block, and updating an offset pointer to call the next portion to be received.
7. The method according to claim 2, wherein receiving an input data stream includes sequentially receiving portions of the block, and the position of each portion is identified by one of the pointers.
8. The method according to claim 2, wherein the portion of the compressed data value is forced to be stored starting from the least significant bit of the column by padding the most significant bit of the empty space of the previous data value.
9. The method according to claim 1, wherein several rows of bit-width registers store a binary representation of the length of the compressed bit width.
10. The method according to claim 9, wherein other rows of bit-width registers store a single bit specifying whether the compressed bit width of the corresponding row is the same as or different from the previous row.
11. The method is used to store floating-point values, the floating-point values include a sign part, an exponent part and a mantissa part, the input data stream consists of the exponent part of the floating-point values, and compressing further includes storing, for each floating-point value, the sign part and the mantissa part adjacent to the compressed exponent part. The method according to claim 1.
12. The method according to claim 1, wherein during decompression, a pointer is established for a specific one of the positions of a plurality of blocks of compressed values known to be needed in the future.
13. The method according to claim 1, further including tracking the next empty position within each column of the first memory while compressing and storing the values.
14. The method according to claim 1, further comprising initializing a first storage location of the first memory as free space before compressing the data stream.
15. The method according to claim 1, wherein the plurality of values have a fixed bit width that is less than or equal to the column width.
16. The method according to claim 1, wherein the reproduced data stream is directly output to an arithmetic / logical unit.
17. The method according to claim 1, wherein the reproduced data stream is output to a second memory having a plurality of rows each having a plurality of columns corresponding to the first memory.
18. The method according to claim 1, wherein compressing further comprises evaluating a function regarding values of the input data stream to reduce the compressed bit width before identifying the compressed bit width, and reversing the function for decompression.
19. A method for memory decompression for a deep learning network, the method comprising: obtaining a compressed data stream representing an input data stream, the compressed data stream defining a plurality of rows for a first memory of a deep learning network, the plurality of rows each having a number of columns, each column having a column width; receiving an input data stream processed by one or more layers of the deep learning network, the input data stream having a plurality of values of a fixed bit width; dividing the input data stream into subsets, the number of values in each subset being equal to the number of columns; compressing the input data stream by sequentially compressing each subset, identifying, for values within the subset, a compressed bit width required to accommodate the value of the largest magnitude; storing the compressed bit width in a bit width register associated with the row; storing the least significant bit of each value within the subset in each column of the first memory starting from the first available bit, the number of bits in which the value can be stored being equal to the compressed bit width, and if more bits are required in each column of each row than are currently unused to store the bits in which the value can be stored, the remaining bits of the bits in which the value can be stored are written in each column of subsequent rows. comprising; prepared by; uncompressing the compressed data stream to obtain the input data stream as; identifying the first unread bit of each column of the compressed data stream; the reproduced input data stream; obtaining the compressed bit widths of each subset from the respective bit width registers; extracting, from each column of the first memory, starting from the first unread bit of the column, the number of bits corresponding to the compressed bit width, and outputting the bits extracted to the least significant bit of the output; updating the position of the first unread bit of each column so as to correspond to the bit position following the extracted bit; zero-extending or sign-extending the remaining most significant bits of the output to obtain the reproduced input data value; sequentially outputting by; reproducing by; A method comprising. [
20. ] A system for memory compression of a deep learning network, the system comprising: a first memory having a plurality of rows each having a number of columns, each column having a column width; an input module, receiving an input data stream processed by one or more layers of the deep learning network, the input data stream having a plurality of values of a fixed bit width; dividing the input data stream into subsets, the number of values in each subset being equal to the number of columns; for; an input module; a compression module for storing the least significant bit of each value in the subset in each column of the first memory starting from the first available bit, the number of bits capable of storing the value being equal to the compressed bit width, and if more bits are required to store the bits capable of storing the value than those currently unused in each column of each row, the remaining bits of the bits capable of storing the value being written to each column of subsequent rows; Decompress the compressed data stream to obtain the input data stream, identify the first unread bit of each column of the compressed data stream, obtain the reproduced input data stream, obtain the compressed bit widths of each subset from the respective bit width registers, extract from each column of the first memory, starting from the first unread bit of the column, the number of bits corresponding to the compressed bit width, and output the bits extracted to the least significant bit of the output, update the first unread bit of each column so as to correspond to the bit position following the extracted bit, zero or sign-extend the remaining most significant bits of the output to obtain the reproduced input data value, sequentially output by, a decompression module for reproducing by, A system comprising.
21. The system according to claim 20, further comprising a pointer module having one or more pointers for tracking the position of a block of compressed values.
22. The system according to claim 21, wherein the block is a filter map data block or an input or output activation data block.
23. The system according to claim 21, wherein the position is for the first compressed value of the block.
24. The system according to claim 21, wherein the one or more pointers include a first set of pointers to the data of the input or output activation map and a second set of pointers to the data of the filter map.
25. The system further includes an offset pointer, and receiving the input data stream includes sequentially receiving the portion of the block starting at the position of the one or more pointers, compressing the portion of the block, and updating the offset pointer to call the next portion to be received. The system according to claim 21.
26. The system according to claim 21, wherein receiving the input data stream includes sequentially receiving the portions of the block, and the position of each portion is identified by one of the pointers.
27. The system according to claim 20, wherein the portion of the compressed data value is forced to be stored starting from the least significant bit of the column by padding the empty most significant bits of the previous data value.
28. The system according to claim 20, wherein bit-width registers of several rows store a binary representation of the length of the compressed bit-width.
29. The system according to claim 20, wherein bit-width registers of other rows store a single bit specifying whether the compressed bit-width of the corresponding row is the same as or different from that of the previous row.
30. The system is for storing floating-point values, the floating-point values include a sign part, an exponent part and a mantissa part, the input data stream consists of the exponent part of the floating-point values, and compressing further includes storing, for each floating-point value, the sign part and the mantissa part adjacent to the compressed exponent part. The system according to claim 20.
31. The system according to claim 20, wherein during decompression, a pointer is established for a specific one of the positions of a plurality of blocks of compressed values known to be needed in the future.
32. The system according to claim 20, wherein the compression module is configured to track the next empty position within each column of the first memory while compressing and storing the values.
33. The system according to claim 20, wherein the compression module is configured to initialize a first storage position of the first memory as empty before compressing the data stream.
34. The system according to claim 20, wherein the plurality of values have a fixed bit-width less than or equal to the column width.
35. The system according to claim 20, wherein the reproduced data stream is directly output to an arithmetic / logic unit.
36. The system according to claim 20, wherein the reproduced data stream is output to a second memory having a plurality of rows each having a plurality of columns corresponding to the first memory.
37. Compressing further includes evaluating a function regarding the values of the input data stream to reduce the compressed bit-width before identifying the compressed bit-width, and reversing the function for decompression. The system according to claim 20.
Citation Information
Patent Citations
Information processing device performing convolution arithmetic processing in layer of convolution neural network
JP2020013455A
Data inspection for compression / decompression configuration and data type determination
US20190081637A1
Lossless compression of neural network weights
US20200143249A1