Apparatus and method for representing a sparse matrix in a neural network
A buffer and sparse engine architecture with hierarchical bitmaps and arrays addresses the inefficiencies of sparse matrices in neural networks, reducing storage and improving computational efficiency.
Patent Information
- Application Number
- CN202180012162.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-05
- Filing Date
- 2021-01-21
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-01-21
AI Technical Summary
Sparse matrices in existing neural networks lead to reduced storage and computing efficiency, occupying a lot of storage space and computing involves a lot of unnecessary operations.
The sparse matrix is represented as a format that includes a first-level bitmap, a second-level bitmap, and an array of elements, and decompresses through a sparse engine to determine non-zero elements, using the processing array to perform neural network operations.
Reduces storage space requirements, improves computing efficiency, and enables simple decompression processes and higher throughput.
Smart Images

Figure CN115066692B_ABST
Abstract
Description
[0001] This disclosure claims the priority of a U.S. application with application number 16 / 783,069, filed on February 5, 2020, and is incorporated herein by reference. Background Art
[0002] Current neural networks often include multiple nodes and multiple layers. However, this reduces execution efficiency and increases latency. Thus, input sparsity, output sparsity, weight sparsity, or a combination thereof has been proposed for neural networks to improve execution efficiency and reduce latency. Indeed, the sparse performance of artificial neural networks can more accurately reflect how neurons in the human brain process information. However, sparse matrices in neural networks can lead to a significant reduction in storage and computational efficiency. For example, they require a large amount of unnecessary storage space, most of which is occupied by zero elements. In addition, the calculation of sparse matrices involves a large number of unnecessary operations (such as addition and multiplication) on zero elements. Summary of the Invention
[0003] In some embodiments, an exemplary operation unit includes a buffer for storing a representation of a sparse matrix in a neural network, a sparse engine communicatively coupled to the buffer, and a processing array communicatively coupled to the sparse engine. The sparse engine includes circuitry for performing the following operations: reading the representation of the sparse matrix from the buffer, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array; decompressing the first-level bitmap to determine whether each block of the sparse matrix includes non-zero elements; and in response to the block including non-zero elements, decompressing the second-level bitmap using the element array to obtain the block of the sparse matrix including non-zero elements. The processing array includes circuitry for performing a neural network using the sparse matrix.
[0004] In some embodiments, an exemplary processing core includes local memory and an operation unit communicatively coupled to the local memory. The operation unit includes a buffer for storing a representation of a sparse matrix in a neural network, a sparse engine communicatively coupled to the buffer, and a processing array communicatively coupled to the sparse engine. The sparse engine includes circuitry for performing the following operations: reading the representation of the sparse matrix from the buffer, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array; decompressing the first-level bitmap to determine whether each block of the sparse matrix includes non-zero elements; and in response to the block including non-zero elements, decompressing the second-level bitmap using the element array to obtain the block of the sparse matrix including non-zero elements. The processing array includes circuitry for performing a neural network using the sparse matrix.
[0005] In some embodiments, an exemplary method of executing a neural network includes: reading a representation of a sparse matrix, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array; decompressing the first-level bitmap to determine whether each block of the sparse matrix includes non-zero elements; and in response to a block including non-zero elements, using the element array to decompress the second-level bitmap to obtain the blocks of the sparse matrix that include non-zero elements; and executing a neural network using the sparse matrix.
[0006] In some embodiments, an exemplary non-transitory computer-readable storage medium stores a set of instructions that, when executed by one or more processing devices, cause an operating unit to perform a method that includes: reading a representation of a sparse matrix in a neural network, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array; decompressing the first-level bitmap to determine whether each block of the sparse matrix includes non-zero elements; in response to a block including non-zero elements, using the element array to decompress the second-level bitmap to obtain the blocks of the sparse matrix that include non-zero elements; and executing a neural network using the sparse matrix.
[0007] Other features and advantages of the present disclosure will be described in detail in part in the following description, and in part will be obvious based on the description or learned through practice of the embodiments of the present disclosure. The features and advantages of the present disclosure will be realized and obtained by the elements and combinations particularly pointed out in the appended claims.
[0008] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The drawings forming a part of the present disclosure are used to illustrate multiple embodiments and, together with the description, explain the principles and features of the embodiments of the present disclosure. In the figures:
[0010] Figure 1A is a schematic diagram of sparsifying a matrix in a neural network according to some embodiments of the present disclosure;
[0011] Figure 1B is another schematic diagram of sparsifying a matrix in a neural network according to some embodiments of the present disclosure;
[0012] Figure 2A is a schematic diagram of the architecture of an exemplary neural network heterogeneous acceleration processing unit (HAPU) according to some embodiments of the present disclosure;
[0013] Figure 2BIt is a schematic diagram of the architecture of a core of an exemplary neural network heterogeneous acceleration processing unit provided according to some embodiments of the present disclosure;
[0014] Figure 2C It is a schematic diagram of an exemplary cloud system including a neural network heterogeneous acceleration processing unit architecture provided according to some embodiments of the present disclosure;
[0015] Figure 3 It is a schematic diagram of an exemplary method for representing a sparse matrix in a neural network provided according to some embodiments of the present disclosure;
[0016] Figure 4 It is a schematic diagram of another exemplary method for representing a sparse matrix in a neural network provided according to some embodiments of the present disclosure;
[0017] Figure 5A It is a schematic diagram of an exemplary operation unit provided according to some embodiments of the present disclosure;
[0018] Figure 5B It is a schematic diagram of an exemplary sparse engine provided according to some embodiments of the present disclosure;
[0019] Figure 6 It is a flowchart of an exemplary method for representing a sparse matrix in a neural network provided according to some embodiments of the present disclosure;
[0020] Figure 7 It is a flowchart of an exemplary method for decompressing the representation of a sparse matrix in a neural network. Detailed implementation manners
[0021] Exemplary embodiments will be described in detail below, and examples thereof are shown in the drawings. The following description of the exemplary embodiments is set forth with reference to the drawings, in which like numbers represent the same or similar elements in different drawings unless otherwise specified. The implementations set forth in the following description of the exemplary embodiments do not represent all implementations consistent with the present disclosure; rather, they are merely examples of apparatuses, systems, and methods consistent with aspects related to the present disclosure as set forth in the appended claims.
[0022] In a neural network, there are many traditional sparse representations of sparse matrices. For example, Compressed Sparse Row (CSR) or Compressed Sparse Column (CSC) uses a format of non-zero elements plus indices to represent a sparse matrix. If the sparsity of the matrix is not high, these representations will occupy a huge storage space to store the indices. In addition, it is difficult for hardware to decompress the representations in these formats. Embodiments of the present disclosure can improve the traditional sparse representations. For example, in some embodiments of the present disclosure, the storage space for storing the representation of the sparse matrix can be reduced. In some embodiments of the present disclosure, the format of the representation is hardware-friendly, and it has a simple decompression process and a higher decompression throughput.
[0023] It can be understood that the matrix or sparse matrix of the present disclosure can be any matrix or sparse matrix. For example, the matrix or sparse matrix can be a weight matrix, an activation matrix, etc. associated with a neural network. In some embodiments of the present disclosure, a sparse matrix in a neural network can be obtained by sparsifying (e.g., pruning) the matrix. Quantization can also be used to reduce the precision of the element values (e.g., weights or activations), thereby reducing the storage cost and the amount of computation.
[0024] Figure 1A FIG. 110 is a schematic diagram of an exemplary sparsification of a matrix in a neural network according to some embodiments of the present disclosure. For example, sparsification 110 reduces matrix 111 to sparse matrix 115 to reduce the amount of computation required to execute the neural network. Although matrix 111 is described as a 4×4 two-dimensional (2D) matrix, matrix 111 can be of any size or any dimension, such as one-dimensional (1D) or three-dimensional (3D).
[0025] Therefore, as Figure 1A shown, sparsification 110 includes selecting one or more elements from matrix 111, such as elements 113a, 113b, 113c, and 113d. Although sparsification 110 is described as selecting four elements, sparsification 110 can use any predetermined number of elements. Elements 113a, 113b, 113c, and 113d are selected because they have the four largest absolute values. Sparsification 110 also includes zeroing out the unselected elements, as shown in sparse matrix 115. Therefore, as Figure 1A shown, sparsification 110 implements 75% sparsification in matrix 111. In addition, the degree of sparsification depends on the predetermined number of elements and the size of matrix 111.
[0026] Figure 1BFIG. 0 is another exemplary diagram of sparsifying matrix 120 in a neural network according to some embodiments of the present disclosure. For example, sparsification 120 reduces matrix 121 to a sparse matrix 125 to reduce the amount of computation required to execute the neural network. Although matrix 121 is depicted as a 4×4 matrix, matrix 121 can be of any size or any dimension, such as one-dimensional (1D) or three-dimensional (3D).
[0027] Thus, as Figure 1B shown, sparsification 120 includes selecting one or more elements from matrix 121, such as elements 123a, 123b, 123c, and 123d. Although sparsification 120 is described as selecting four elements, sparsification 120 can use any predetermined number of elements. Elements 123a, 123b, 123c, and 123d are selected because they are located in the selected columns. Although sparsification 120 is described as selecting one column, sparsification 120 can select any predetermined number of blocks, such as vectors, columns, rows, filters, channels, etc. Sparsification 120 also includes zeroing out the unselected elements, as shown in sparse matrix 125. Thus, as Figure 1B shown, sparsification 120 implements 75% sparsification on matrix 121. Additionally, the degree of sparsification depends on the predetermined number of columns and the size of matrix 121.
[0028] Since the elements with the largest absolute values may be distributed anywhere in matrix 111, Figure 1A sparsification 110 cannot provide spatial predictability when selecting the elements not to be set to zero. Thus, Figure 1A sparse matrix 115 has a random non-zero element distribution and unstructured sparsity. However, Figure 1B sparsification 120 provides spatial predictability because its non-zero elements are selected according to a specific rule (e.g., the first element in each row). Thus, Figure 1B sparse matrix 125 has a regular non-zero element distribution and structured sparsity. It can be understood that Figure 1A sparsification 110 of Figure 1B and
[0029] Figure 2A Figure 1B sparsification 120 are examples of generating sparse matrices, rather than limitations. Sparse matrix 115 and sparse matrix 125 are exemplary sparse matrices. For example, since there is a trade-off between using less aggressive sparse techniques for more accurate results and using more aggressive sparse techniques to save computation, the degree of sparsity of the matrix depends on the purpose of the result. Embodiments of the present disclosure can also use other sparse matrices with different sparsity levels and non-zero element distributions, as well as other sparsification methods.FIG. 200 shows the architecture of an exemplary neural network heterogeneous acceleration processing unit (HAPU) provided according to some embodiments of the present disclosure. In the context of the present disclosure, the neural network heterogeneous acceleration processing unit may also be referred to as a machine learning accelerator or a deep learning accelerator. In some embodiments, the architecture 200 of the HAPU is referred to as the architecture 200 of a neural network processing unit (NPU). As Figure 2A shown, the architecture 200 of the HAPU includes a plurality of cores 202, a command processor 204, a direct memory access (DMA) unit 208, a Joint Test Action Group / Test Access Port (JTAG / TAP) controller 210, a peripheral interface 212, a bus 214, and so on.
[0030] It can be understood that the cores 202 perform algorithm operations based on communication data. The cores 202 may include one or more processing elements, and the one or more processing elements may include a single instruction multiple data (SIMD) architecture. The single instruction multiple data architecture may include one or more processing units for performing one or more operations (such as multiplication, addition, multiply-accumulation, etc.) based on commands received from the command processor 204. In order to perform operations on the transmitted data packets, the cores 202 include one or more processing elements for processing the information in the data packets. Each processing element may include any number of processing units. According to some embodiments of the present disclosure, the architecture 200 of the HAPU includes a plurality of cores 202, such as four cores. In some embodiments, the plurality of cores 202 are communicatively coupled to each other. For example, the plurality of cores 202 are connected to a unidirectional ring bus that supports an efficient pipeline for large neural network models. The architecture of the cores 202 will be described in detail below with reference to Figure 2B FIGS.
[0031] The command processor 204 interacts with the host unit 220 and transmits relevant commands and data to the corresponding cores 202. In some embodiments, the command processor 204 interacts with the host unit under the supervision of a kernel mode driver (KMD). In some embodiments, the command processor 204 modifies the relevant commands transmitted to each core 202 so that the plurality of cores 202 can work in parallel as much as possible. The modified commands may be stored in an instruction buffer (not shown). In some embodiments, the command processor 204 is used to coordinate the parallel execution of one or more cores 202.
[0032] The DMA unit 208 can assist in transferring data between the host memory 221 and the architecture 200 of the HAPU. For example, the DMA unit 208 can assist in loading data or instructions from the host memory 221 into the local memory of the core 202. The DMA unit 208 can also assist in transferring data between multiple HAPUs. The DMA unit 208 can allow off-chip devices to access on-chip and off-chip memories without causing a host CPU interruption. In addition, the DMA unit 208 can assist in transferring data between components within the architecture 200 of the HAPU. For example, the DMA unit 208 can assist in transferring data between multiple cores 202 or within each core. Therefore, the DMA unit 208 can also generate memory addresses and initiate memory read or write cycles. The DMA unit 208 can also include several hardware registers that can be written to and read by one or more processors. These several hardware registers include a memory address register, a byte count register, one or more control registers, and other types of registers. These registers can specify the source, destination, direction (reading data from or writing data to an input / output (I / O) device), size of the transfer unit, and number of bytes transferred in one burst. It can be understood that the architecture 200 of the HAPU can include a second DMA unit for transferring data between other HAPU architectures to allow direct communication between multiple HAPU architectures without involving the host CPU.
[0033] The JTAG / TAP controller 210 can specify a dedicated debug port to implement a serial communication interface (such as a JTAG interface) to access the HAPU with low overhead and without directly accessing the external system address and data buses. The JTAG / TAP controller 210 can also have an on-chip test access interface (such as a TAP interface), which implements a protocol for accessing a set of test registers used to present the chip logic levels and device capabilities of each component.
[0034] If there is a peripheral interface 212 (such as a PCIe interface), it is typically used as an inter-chip bus to provide communication between the HAPU and other devices.
[0035] The bus 214 (e.g., I2C bus) may include an on-chip bus and an inter-chip bus. The on-chip bus interconnects all internal components according to the requirements of the system architecture. Although not all components are connected to each other, all components have a certain connection to other components with which they need to communicate. The inter-chip bus connects the HAPU to other devices (e.g., off-chip memory or peripherals). For example, the bus 214 can provide high-speed communication across cores and can also connect the cores 202 to other units (e.g., off-chip memory or peripherals). Generally, if there is a peripheral interface 212 (e.g., inter-chip bus), the bus 214 is only related to the on-chip bus, but in some implementations, the bus 214 is still related to dedicated inter-bus communication.
[0036] The accelerator architecture 200 may also communicate with the host unit 220. The host unit 220 may be one or more processing units (e.g., X86 central processing units). As Figure 2A shown, the host unit 220 may be associated with the host memory 221. In some embodiments, the host memory 221 is an integrated memory or an external memory associated with the host unit 220. In some embodiments, the host memory 221 includes a host disk, which is an external memory for providing additional memory to the host unit 220. The host memory 221 may be a double data rate synchronous dynamic random access memory (e.g., DDR-SDRAM), etc. Compared with the on-chip memory integrated within the HAPU chip, the host memory 221 serves as a higher-level cache and stores a large amount of data at a slower access speed. The data stored in the host memory 221 may be transferred to the HAPU architecture 200 for executing neural network models.
[0037] In some embodiments, the host system having the host unit 220 and the host memory 221 includes a compiler (not shown). The compiler is a program or computer software that converts computer code written in one programming language into instructions for the HAPU architecture 200 to create an executable program. In applications of machine learning, the compiler can perform various operations, such as preprocessing, lexical analysis, parsing, semantic analysis, converting the input program into an intermediate representation, neural network initialization, code optimization, code generation, or a combination of the above. For example, the compiler compiles a neural network to generate static parameters, such as the connections between neurons and the weights of neurons.
[0038] In some embodiments, the host system including the compiler pushes one or more commands to the architecture 200 of the HAPU. As described above, these commands can be further processed by the command processor 204 of the architecture 200 of the HAPU, and they can be temporarily stored in the instruction buffer of the HAPU architecture 200 and distributed to the corresponding one or more cores (e.g.,Figure 2A the core 202) or processing element therein. Some commands may instruct the DMA unit (e.g., Figure 2A the DMA unit 208) to load instructions and data from the host memory (e.g., Figure 2A the host memory 221) into the architecture 200 of the HAPU. Then, these loaded instructions are distributed to each core (e.g., Figure 2A the core 202) assigned with the corresponding task, and these instructions are processed by these one or more cores.
[0039] It can be understood that the first few instructions received by the core 202 may instruct the core 202 to load / store data from / to one or more local memories of the core (e.g., Figure 2B the local memory 2032). Then, each core 202 can start an instruction pipeline, which involves (e.g., through an sequencer) extracting instructions from an instruction buffer, (e.g., through Figure 2A the DMA unit 208) decoding the instructions, generating local memory addresses (e.g., corresponding to operands), reading source data, performing load / store operations, and then writing back the results.
[0040] According to some embodiments, the architecture 200 of the HAPU further includes a global memory (not shown), which has storage blocks (e.g., 4 blocks of 8GB second-generation high-bandwidth memory (HBM2)) serving as main memories. In some embodiments, the global memory stores instructions and data obtained from the host memory 221 via the DMA unit 208. Then, these instructions are distributed to the instruction buffers of each core assigned with the corresponding task, and these instructions are processed by the cores accordingly.
[0041] In some embodiments, the architecture 200 of the HAPU further includes a memory controller (not shown), which is used to manage reading and writing data from and to specific memory blocks (e.g., HBM2) within the global memory. For example, the memory controller manages reading / writing data from the core of another HAPU (e.g., from the DMA unit 208 or the DMA unit corresponding to another HAPU), or manages reading / writing data from the core 202 (e.g., from the local memory within the core 202). It can be understood that multiple memory controllers can be provided in the architecture 200 of the HAPU. For example, each storage block (e.g., HBM2) in the global memory has a memory controller.
[0042] A memory controller can generate memory addresses and initiate memory read or write cycles. The memory controller can include a number of hardware registers that are written to and read by one or more processors. These registers can include a memory address register, a byte count register, one or more control registers, and other types of registers. These registers can specify the source, destination, direction (read from or written to an input / output (I / O) device), size of the transfer unit, number of bytes transferred in a single burst, or other typical characteristics of the memory controller.
[0043] According to some embodiments of the present disclosure, Figure 2A the architecture 200 of the HAPU is for a Convolutional Neural Network (CNN), but it should be understood that Figure 2A the architecture 200 of the HAPU can be used for various neural networks, such as, for example, a Deep Neural Network (DNN), a Recurrent Neural Network (RNN), etc. In addition, some embodiments can be configured for various processing architectures, such as, for example, a neural processing unit (NPU), a graphic processing unit (GPU), a tensor processing unit (TPU), any other type of HAPU, etc.
[0044] Figure 2B An exemplary architecture of a core according to some embodiments of the present disclosure is shown. As Figure 2B shown, the core 202 includes one or more operation units, such as a first operation unit 2020 and a second operation unit 2022, a memory engine 2024, an sequencer 2026, an instruction buffer 2028, a constant buffer 2030, a local memory 2032, etc.
[0045] The one or more operation units include a first operation unit 2020 and a second operation unit 2022. The first operation unit 2020 can be used to perform operations on received data (e.g., a matrix). In some embodiments, the first operation unit 2020 includes one or more processing units for performing one or more operations (e.g., multiplication, addition, multiply-accumulate, element-wise operations, etc.). In some embodiments, the first operation unit 2020 is used to accelerate the execution of convolutional operations or matrix multiplication operations.
[0046] The second operation unit 2022 can be used to perform pooling operations, interpolation operations, region-of-interest (ROI) operations, etc. In some embodiments, the second operation unit 2022 includes an interpolation unit, a pooling data path, etc.
[0047] The memory engine 2024 can be used to perform data copying inside the corresponding core 202 or between two cores. The DMA unit 208 can assist in copying data inside the corresponding core or between two cores. The DMA unit 208 can assist in copying data inside the corresponding core or between two cores. For example, the DMA unit 208 supports the memory engine 2024 in copying data from the local memory (e.g., Figure 2B the local memory 2032) to the corresponding operation unit. The memory engine 2024 can also be used to perform matrix transposition to make the matrix suitable for use in the operation unit.
[0048] The sequencer 2026 is coupled to the instruction buffer 2028 and can be used to extract commands and distribute the commands to the various components of the core 202. For example, the sequencer 2026 distributes convolution commands or multiplication commands to the first operation unit 2020, pooling commands to the second operation unit 2022, or data copy commands to the memory engine 2024. The sequencer 2026 can also be used to monitor the execution of neural network tasks and parallelize multiple subtasks of the neural network tasks to improve execution efficiency. In some embodiments, the first operation unit 2020, the second operation unit 2022, and the memory engine 2024 run in parallel under the control of the sequencer 2026 according to the instructions stored in the instruction buffer 2028.
[0049] The instruction buffer 2028 can be used to store the instructions belonging to the corresponding core 202. In some embodiments, the instruction buffer 2028 is coupled to the sequencer 2026 and provides instructions to the sequencer 2026. In some embodiments, the instructions stored in the instruction buffer 2028 are transmitted or modified by the command processor 204.
[0050] The constant buffer 2030 can be used to store constant values. In some embodiments, the constant values stored in the constant buffer 2030 can be used by operation units such as the first operation unit 2020 or the second operation unit 2022 for batch normalization, quantization, dequantization, etc.
[0051] The local memory 2032 can provide a storage space with fast read / write speeds. To reduce the possible interaction with the global memory, the storage space of the local memory 2032 can be built into a large capacity. With a huge storage space, most data accesses are executed within the core 202, thereby reducing the latency caused by data access. In some embodiments, to minimize data loading latency and energy consumption, the static random access memory (SRAM) integrated on the chip is used as the local memory 2032. In some embodiments, the local memory 2032 has a capacity of 192 MB or more. According to some embodiments of the present disclosure, the local memory 2032 is evenly distributed on the chip to mitigate the problems of dense wiring and heat generation.
[0052] Figure 2C FIG. shows a schematic diagram of an exemplary cloud system including the HAPU architecture 200 provided according to some embodiments of the present disclosure. As Figure 2C shown, the cloud system 230 can provide cloud services with artificial intelligence (AI) capabilities and can include multiple computing servers (such as computing servers 232 and 234). In some embodiments, the computing server 232, for example, integrates Figure 2A the architecture 200 of the neural network HAPU in. For simplicity and clarity, Figure 2C the architecture 200 of the neural network HAPU is shown in a simplified manner in.
[0053] With the architecture 200 of the neural network HAPU, the cloud system 230 can provide extended AI functions such as image recognition, face recognition, translation, 3D modeling, etc. It can be understood that the architecture 200 of the neural network HAPU can be deployed in computing devices in other forms. For example, the architecture 200 of the neural network HAPU can also be integrated in computing devices such as smartphones, tablets, and wearable devices.
[0054] In addition, it can be understood that although Figures 2A - 2B the architecture of the neural network HAPU is shown, the present disclosure can use any HAPU that provides the ability to perform parallel computing.
[0055] Figure 3 FIG. is a schematic diagram of a method 300 for representing a sparse matrix in an exemplary neural network provided according to some embodiments of the present disclosure. As Figure 3As shown, the representation method 300 includes transforming or compressing the sparse matrix 301 into a representation 303. It can be understood that the sparse matrix 301 can be any sparse matrix, such as a weight matrix, activation matrix, etc. associated with a neural network. Although the matrix 301 is described as a 6×9 two-dimensional (2D) matrix, it can be of any size or any dimension.
[0056] The sparse matrix 301 includes many non-zero (NZ) elements, such as NZ0, NZ1, ……, NZ6, which are randomly or regularly distributed in the matrix. The sparse matrix 301 is divided into multiple blocks. The blocks can be of any size or shape, including but not limited to squares, rectangles, cubes, cuboids, vectors, columns, rows, filters, channels, etc. In some embodiments, the sparse matrix 301 is divided into blocks according to hyperparameters (such as compression ratio, distribution of non-zero elements, sparsification method, storage requirements, etc.). As Figure 3 shown, the sparse matrix 301 is divided into six 3×3 blocks, namely block 0, block 1, ……, block 5.
[0057] The sparse matrix 301 can be represented by the representation 303. The representation 303 can include multiple levels of bitmaps (BM). For example, the representation 303 includes a first-level bitmap 3031 to indicate the sparsity at the block level, e.g., to indicate whether the blocks of the sparse matrix 301 include non-zero elements. As Figure 3 shown, the first-level bitmap 3031 includes six elements B0, B1, ……, B5, and each element indicates whether the corresponding block (block 0, block 1, ……, or block 5) includes non-zero elements. B0, B1, ……, and B5 correspond to block 0, block 1, ……, and block 5 respectively. B0 is set to 1 indicating that block 0 has at least one non-zero element, while B1 is set to 0 indicating that block 1 has no non-zero elements, and so on. Although the first-level bitmap 3031 is shown as a vector, it can be of any other size or form, such as a 2D matrix.
[0058] In addition, the representation 303 further includes a second-level bitmap 3032 to indicate the sparsity of each block, e.g., for each block having non-zero elements, to indicate whether each element of the block is a non-zero element. For example, the second-level bitmap 3032 includes one or more parts (such as part 3032a), and each part corresponds to a block having non-zero elements. For blocks having no non-zero elements (such as block 1 or block 5) or for zero elements in the first-level bitmap 3031 (such as B1 or B5), the second-level bitmap 3032 has no corresponding parts. The one or more parts are arranged in the structure of the second-level bitmap 3032 corresponding to the structure of the first-level bitmap 3031. Each part of the second-level bitmap 3032 includes multiple elements, and each element indicates whether an element in the corresponding block is a non-zero element. As Figure 3As shown, for example, a portion 3032a of the second-level bitmap 3032 corresponds to B3 of the first-level bitmap 3031 or block 3 of the sparse matrix 301. Block 3 of the sparse matrix 301 includes two non-zero elements: NZ3 at position (0, 0) and NZ4 at position (2, 0). The portion 3032a of the second-level bitmap 3032 includes two elements, namely E(0, 0) at position (0, 0) and E(2, 0) at position (2, 0), which are set to 1 to indicate the presence of non-zero elements at these positions. Although the second-level bitmap 3032 is shown as a 3×N matrix, the second-level bitmap 3032 can be of any other size or form, such as a 1D matrix.
[0059] In some embodiments, the representation 303 includes an array of elements 3033 that has all the non-zero elements of the sparse matrix 301, such as NZ0, NZ1, ……, NZ6. The non-zero elements are arranged in a structure corresponding to the structure of the second-level bitmap 3032. As Figure 3 shown, for example, NZ3, NZ4 are arranged to correspond to the portion 3032a of the second-level bitmap 3032, and the portion 3032a corresponds to B3 of the first-level bitmap 3031, and thus corresponds to block 3 of the sparse matrix 301. Specifically, NZ3 and NZ4 respectively correspond to E(0, 0) and E(2, 0) of the portion 3032a of the second-level bitmap 3032.
[0060] In some embodiments, a portion of the second-level bitmap 3032 (e.g., portion 3032a) can be further divided into multiple sub-blocks and represented by a first sub-level bitmap and a second sub-level bitmap. Similar to the first-level bitmap 3031, the first sub-level bitmap 3032a-1 includes multiple elements to indicate the sparsity at the sub-block level, e.g., to indicate whether each sub-block of the portion 3032a of the second-level bitmap 3032 includes a non-zero element. For example, the portion 3032a of the second-level bitmap 3032 is divided into three columns. The first sub-level bitmap 3032a-1 can include three elements C0, C1, and C2 to indicate whether each column of the portion 3032a includes a non-zero element. C0, C1, and C2 respectively correspond to the left column, middle column, and right column of the portion 3032a. Since the left column of the portion 3032a includes two non-zero elements E(0, 0) and E(2, 0), while the middle column and right column of the portion 3032a include only zero elements, C0 is set to 1, while C1 and C2 are both set to 0. Although the first sub-level bitmap 3032a-1 is shown as a vector with three elements, the first sub-level bitmap 3032a-1 can be of any other size or form.
[0061] Similar to the second-level bitmap 3032, the second sub-level bitmap 3032a-2 includes one or more parts to indicate whether each element in a sub-block having non-zero elements is a non-zero element. In some embodiments, for a sub-block that does not have non-zero elements (e.g., the middle column or the right column of part 3032a) or for a zero element in the first sub-level bitmap (e.g., C1 or C2 of the first sub-level bitmap 3032a-1), the second sub-level bitmap (e.g., the second sub-level bitmap 3032a-2) does not have a corresponding part. Each part of the second sub-level bitmap 3032a-2 includes a plurality of elements, and each element indicates whether each element in the corresponding sub-block of part 3032a of the second-level bitmap 3032 is a non-zero element. The structure of the first sub-level bitmap 3032a-1 corresponds to the structure of the second-level bitmap 3032, and the structure of the second sub-level bitmap 3032a-2 corresponds to the structure of the first sub-level bitmap 3032a-1. For example, as Figure 3 shown, the second sub-level bitmap 3032a-2 has a part corresponding to C0 in the first sub-level bitmap 3032a-1 and the left column of the second-level bitmap 3032. This part of the second sub-level bitmap 3032a-2 includes three elements (1, 0, 1) to correspond to the three elements in the left column of part 3032a of the second-level bitmap 3032 and indicate whether they are non-zero elements. Although the second sub-level bitmap 3032a-2 is shown as a vector having three elements, it can be of any other size or form.
[0062] It can be understood that in some embodiments, there can be deeper sub-level bitmaps 3032a-n. For example, similar to the sparse matrix 301 or part 3032a of the second-level bitmap 3032, the second sub-level bitmap 3032a-2 can be divided into multiple sub-blocks and represented by other deeper sub-level bitmaps.
[0063] Although described as three separate data, in some embodiments, the first-level bitmap 3031, the second-level bitmap 3032, and the element array 3033 may be combined with each other. For example, the first-level bitmap 3031 and the second-level bitmap 3032 may be combined in such a way that an element of the first-level bitmap 3031 (e.g., B3) is followed by the corresponding part of the second-level bitmap 3032 (e.g., 3032a). As another example, the second-level bitmap 3032 may be combined with the element array 3033 in such a way that a part of the second-level bitmap 3032 (e.g., part 3032a) is followed by the corresponding non-zero elements of the element array 3033 (e.g., NZ3 and NZ4). As another example, the first-level bitmap 3031, the second-level bitmap 3032, and the element array 3033 may be combined in such a way that an element of the first-level bitmap 3031 (e.g., B3) is followed by the corresponding part of the second-level bitmap 3032 (e.g., part 3032a), and then followed by the corresponding non-zero elements of the element array 3033 (e.g., NZ3 and NZ4).
[0064] Figure 4 is a schematic diagram of another exemplary method 400 for representing a sparse matrix in a neural network provided according to some embodiments of the present disclosure. As Figure 4 shown, the representation method 400 includes transforming or compressing the sparse matrix 401 into a representation 403. It can be understood that the sparse matrix 401 can be any sparse matrix, such as a weight matrix, an activation matrix, etc. associated with a neural network. Although the matrix 401 is depicted as a 3D matrix, it can have any size or dimension.
[0065] The sparse matrix 401 is a 3D matrix of W×H×C, including many non-zero (NZ) elements randomly or regularly distributed in the matrix 401. The sparse matrix 401 can be divided into multiple blocks. The blocks can be of any size or shape, including but not limited to square, rectangle, cube, cuboid, vector, column, row, filter, channel, etc. As Figure 4 shown, the sparse matrix 401 is divided into multiple columns, for example, N columns, where N = W×H. For example, column (0, 0) includes two non-zero elements NZ0 and NZ1, and column (W-1, H-1) includes three non-zero elements NZ (n-2) 、NZ (n-1) and NZ n .
[0066] The sparse matrix 401 can be represented by the representation 403. The representation 403 includes multiple levels of bitmaps. For example, the representation 403 includes a first-level bitmap 4031, a second-level bitmap 4032, etc. The first-level bitmap 4031 indicates whether each column of the sparse matrix 401 includes non-zero elements. As Figure 4As shown, the first-level bitmap 4031 includes a plurality of elements (e.g., N elements), each element indicating whether the corresponding column includes a non-zero element. If there is one or more non-zero elements in the column corresponding to an element of the first-level bitmap 4031, it is set to 1, and if there are no non-zero elements in the column corresponding to an element of the first-level bitmap 4031, it is set to 0.
[0067] The second-level bitmap 3032 indicates whether each element in each column containing one or more non-zero elements is a non-zero element. For example, the second-level bitmap 3032 includes one or more parts, each part corresponding to a column of the sparse matrix 401. For columns without non-zero elements or for zero elements in the first-level bitmap 4031, the second-level bitmap 4032 has no corresponding part. In some embodiments, one or more parts of the second-level bitmap 4032 are arranged in a structure corresponding to the structure of the first-level bitmap 4031. Each part of the second-level bitmap 4032 includes a plurality of elements, and each element indicates whether each element in the corresponding column is a non-zero element or a zero element.
[0068] It should be understood that although shown as a combined vector, the first-level bitmap 4031 and the second-level bitmap 4032 can be of any other size or form. For example, the first-level bitmap 4031 or the second-level bitmap 4032 can be a separate vector or a 2D matrix, or can be combined in another structure, such as each element of the first-level bitmap 4031 followed immediately by the corresponding part of the second-level bitmap 4032.
[0069] As Figure 4 shown, the representation 403 further includes an element array 4033 having all non-zero elements of the sparse matrix 401, such as NZ0, NZ1, ……, NZ (n-2) 、NZ (n-1) 、NZ n . The non-zero elements are arranged in a structure corresponding to the structure of the second-level bitmap 4032.
[0070] In some embodiments, a part (e.g., a vector) of the second-level bitmap 4032 can be further divided into a plurality of sub-blocks (e.g., sub-vectors), and represented by a first sub-level bitmap 4032-1 and a second sub-level bitmap 4032-2. Similar to the first-level bitmap 4031, the first sub-level bitmap 4032-1 includes a plurality of elements, each element indicating whether a sub-block of the corresponding part of the second-level bitmap 4032 includes a non-zero element. For example, the vector of the second-level bitmap 4032 is divided into a plurality of sub-vectors. The first sub-level bitmap 3032-1 includes the same number of elements as the number of sub-vectors, each element indicating whether a sub-vector of the vector includes a non-zero element.
[0071] Similar to the second - level bitmap 4032, the second - sub - level bitmap 4032 - 2 includes one or more parts for indicating whether each element of a sub - block having non - zero elements is a non - zero element. In some embodiments, for a sub - block having no non - zero elements or for zero elements in the first - sub - level bitmap 4032 - 1, the second - sub - level bitmap 3032 - 2 does not include corresponding parts. Each part of the second - sub - level bitmap 4032 - 2 includes a plurality of elements, and each element can indicate whether an element in the corresponding sub - block of the corresponding part of the second - level bitmap 4032 is a non - zero element.
[0072] In some embodiments, the structure of the first - sub - level bitmap 4032 - 1 corresponds to the structure of the second - level bitmap 4032, and the structure of the second - sub - level bitmap 4032 - 2 corresponds to the structure of the first - sub - level bitmap 4032 - 1. Although shown as vectors, the first - sub - level bitmap 4032 - 1 and the second - sub - level bitmap 4032 - 2 can be of any other size or form.
[0073] It can be understood that in some embodiments, there can be deeper sub - level bitmaps. For example, similar to a part of the sparse matrix 401 or the second - level bitmap 4032, the second - sub - level bitmap 4032 - 2 can be divided into multiple sub - blocks and represented by other deeper sub - level bitmaps.
[0074] In some embodiments, the element array 4033 is combined with the first - level bitmap 4031 or the second - level bitmap 4032. For example, the element array 4033 and the second - level bitmap 4032 are combined in such a way that a part of the second - level bitmap 4032 corresponding to a column of the sparse matrix 401 is followed by the corresponding non - zero elements of the element array 4033.
[0075] Figure 5A is a schematic diagram of an exemplary operation unit 500 provided according to some embodiments of the present disclosure. In some embodiments, the operation unit 500 is the first operation unit (e.g., Figures 2A - 2B the first operation unit 2020 of Figure 2B included in the core 502 (e.g.,
[0076] The first buffer 510 can be used to store input data. In some embodiments, the data stored in the first buffer 510 is the input data (e.g., input features) for the processing array 530 to execute. In some embodiments, from the local memory 5032 (e.g., Figure 1BExtract the input data from the local memory 2032). The first buffer 510 can be used to support the reuse or sharing of data used in the processing array 530. In some embodiments, the input data stored in the first buffer 510 is activation data for convolution operations.
[0077] The second buffer 520 can be used to store matrix data, such as the representation of a sparse matrix (e.g., a weight matrix). For example, the operation unit 500 reads, extracts, or receives the representation of the sparse matrix from the local memory 5032 through a memory engine (not shown, e.g., Figure 2B the memory engine 2024) of and stores the representation in the second buffer 520. In some embodiments, the second buffer 520 is part of the first buffer 510 or is separate from the first buffer 510. The second buffer 520 is any suitable memory that provides storage space for data such as matrices or representations, such as registers, Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), etc.
[0078] The operation unit 500 also includes a sparse engine 590, which is communicatively coupled to the second buffer 520 and is used to read data from or write data to the second buffer 520. As Figure 5B shown, the sparse engine 590 may include one or more decompressors, such as a first-stage decompressor 591 and a second-stage decompressor 592, to decompress the representation of the sparse matrix. In some embodiments, the sparse engine 590, the first-stage decompressor 591, or the second-stage decompressor 592 is implemented as a circuit with a high processing speed. The sparse engine 590 reads the representation of the sparse matrix in the neural network from the second buffer 520 (e.g., Figure 3 the representation 303 of, Figure 4 the representation 403 of, etc.), and the representation of the sparse matrix includes a first-stage bitmap 521 (e.g., Figure 3 the first-stage bitmap 3031 of, Figure 4 the first-stage bitmap 4031 of, etc.), a second-stage bitmap 522 (e.g., Figure 3 the second-stage bitmap 3032 of, Figure 4 the second-stage bitmap 4032 of, etc.) and an element array 523 (e.g., Figure 3 the element array 3033 of, Figure 4 the element array 4033 of, etc.). In some embodiments, the representation of the sparse matrix also includes one or more sub-level bitmaps (e.g., Figure 3 the first sub-level bitmap 3032a-1, the second sub-level bitmap 3032a-2, more sub-level bitmaps 3032a-n in, andFigure 4 The first - level child bitmaps 4032 - 1, the second - level child bitmaps 4032 - 2, etc.). The multiple decompressors of the sparse engine 590 include circuitry for communication - coupling between them and allowing them to cooperate with each other to decompress the representation of the sparse matrix. For example, the first - level decompressor 591 and the second - level decompressor 592 are coupled and communicate with each other. The first - level decompressor 591 decompresses the first - level bitmap 521 to determine whether each block of the sparse matrix includes non - zero elements. For example, the elements of the first - level bitmap are set to 1 to indicate that their corresponding blocks include non - zero elements. If a block includes non - zero elements, the second - level decompressor 592 uses the element array 523 to decompress the second - level bitmap 522 to obtain the blocks of the sparse matrix that include non - zero elements. For example, the second - level decompressor 592 uses the second - level bitmap as a bit - mask and uses the element array to reconstruct the sparse matrix. One element of the second - level bitmap is set to 1 to indicate that its corresponding element is a non - zero element included in the element array. In some embodiments, the obtained blocks of the sparse matrix are the same as the original blocks. In some embodiments, the obtained sparse - matrix blocks are variants of the original blocks, but the variants are more hardware - friendly for computation. The first - level decompressor 591 can combine the obtained blocks to obtain the sparse matrix 593 or its variant. Although shown as two separate units, in some embodiments, the first - level decompressor 591 and the second - level decompressor 592 are combined into one decompressor.
[0079] For example, in combination with Figure 3, the sparse engine 590 reads a representation 303 from the second buffer 520 that includes a first-level bitmap 3031, a second-level bitmap 3032, and an element array 3033. The first-level decompressor 591 can decompress the first-level bitmap 3031. For example, it checks each element B0, B1, …, or B5 of the first-level bitmap 3031 to determine whether each block of the sparse matrix 301 includes non-zero elements. Elements B0, B2 - B4 are set to 1, indicating that blocks 0 and 2 - 4 of the sparse matrix 301 have non-zero elements, while B1 and B5 are set to 0, indicating that blocks 1 and 5 of the sparse matrix 301 do not have any non-zero elements. For elements B0, B2 to B4 of the first-level bitmap 3031 or blocks 0 and 2 to 4 of the sparse matrix 301, the second-level decompressor 592 can use the element array 3033 to decompress the second-level bitmap 3032 to obtain blocks 0 and 2 to 4 of the sparse matrix 301. For example, the second-level decompressor 592 reads the second-level bitmap 3022 and checks each element of the second-level bitmap 3022. For a portion 3032a corresponding to B3 or block 3, elements E(0, 0) and E(2, 0) are set to 1 and respectively correspond to non-zero elements NZ3 and NZ4 of the element array 3033. The second-level decompressor 592 can reconstruct block 3 by placing NZ3 and NZ4 at positions (0, 0) and (2, 0) respectively and filling all other positions with zeros. Similarly, blocks 0, 2, and 4 can also be reconstructed. Then, the sparse engine 590 can combine the reconstructed blocks to convert the representation 303 into the sparse matrix 301 or a variant of the sparse matrix 301.
[0080] In some embodiments, the representation read by the sparse engine 590 further includes one or more sub-level bitmaps (not shown), such as a first sub-level bitmap and a second sub-level bitmap. The sparse engine 590 may include more decompressors to process one or more sub-level bitmaps. In some embodiments, the first-level decompressor 591 and the second-level decompressor 592 can process one or more sub-level bitmaps. For example, the first-level decompressor 591 (or a first sub-level decompressor (not shown)) can decompress the first sub-level bitmap to determine whether a sub-block of a portion of the second-level bitmap 522 includes non-zero elements. If the sub-block includes non-zero elements, the second-level decompressor 592 (or a second sub-level decompressor (not shown)) can decompress the second sub-level bitmap to obtain the corresponding sub-block of the corresponding portion of the second-level bitmap 522.
[0081] For example, in combination with Figure 3, the sparse engine 590 may read a representation 303 including a first sub - level bitmap 3032a - 1 and a second sub - level bitmap 3032a - 2 from the second buffer 520. The first - level decompressor 591 may decompress the first sub - level bitmap 3032a - 1. For example, it checks each element C0, C1, or C2 of the first sub - level bitmap 3032a - 1 to determine whether each sub - block (e.g., column) of the portion 3032a includes non - zero elements. The element C0 is set to 1, indicating that the left column of the portion 3032a has one or more non - zero elements, while C1 and C2 are set to 0, indicating that the middle and right columns of the portion 3032a do not have any non - zero elements. For the element C0 of the first sub - level bitmap 3032a - 1 or for the left column of the portion 3032a, the second - level decompressor 592 may decompress the second sub - level bitmap 3032a - 2 to obtain the left column of the portion 3032a. For example, the second - level decompressor 592 reads the second sub - level bitmap 3032a - 2 and reconstructs the left column of the portion 3032a using the second sub - level bitmap 3032a - 2. The second - level decompressor 592 may reconstruct the middle and right columns of the portion 3032a using multiple zeros. Similarly, the sparse engine 590 may reconstruct other parts of the second - level bitmap 3032 and decompress the second - level bitmap 3032 to obtain the sparse matrix 301.
[0082] The operation unit 500 also includes a processing array 530, which may have multiple layers (e.g., K layers). According to some embodiments of the present disclosure, each layer of the processing array 530 may include multiple processing strings that can perform calculations in parallel. For example, the first processing string included in the first layer of the processing array 530 includes a first multiplier (e.g., dot - product) 540_1 and a first accumulator (ACC) 550_1, and the second processing string includes a second multiplier 540_2 and a second accumulator 550_2. Similarly, the i - th processing string in the first layer includes the i - th multiplier 540_i and the i - th accumulator 550_i.
[0083] In some embodiments, the processing array 530 performs calculations under the control of SIMD. For example, when performing a convolution operation, each layer of the processing array 530 uses different data to execute the same instruction.
[0084] According to some embodiments of the present disclosure, Figure 5A the illustrated processing array 530 may be included in a core (e.g., Figure 2A or Figure 2BThe core 202). In some embodiments, when the number of processing strings included in a layer of the processing array 530 (e.g., i processing strings) is less than the number of work items (e.g., B work items), the i work items are first executed by the processing array 530, and then the remaining work items (B - i work items) are executed by the processing array 530. In some other embodiments, the i work items may be executed by the processing array 530, and the remaining work items may be executed by another processing array 530 in another core.
[0085] According to some embodiments of the present disclosure, the processing array 530 further includes an element operation unit 560. In some embodiments, the element operation unit 560 is disposed at the end of the processing string. In some embodiments, the processing strings in each layer of the processing array 530 share the element operation unit 560. For example, the i processing strings in the first layer of the processing array 530 share the element operation unit 560. In some embodiments, the element operation unit 560 in the first layer of the processing array 530 sequentially performs element operations on each output value of the accumulators 550_1 to 550_i. Similarly, the element operation unit 560 in the k-th layer of the processing array 530 sequentially performs element operations on each output value of the accumulators 550_1 to 550_i. In some embodiments, the element operation unit 560 can be used to perform multiple element operations. In some embodiments, the element operations performed by the element operation unit 560 include activation functions, such as ReLU function, ReLU6 function, Leaky ReLU function, Sigmoid function, Tanh function, etc.
[0086] In some embodiments, the multiplier 540 or the accumulator 550 is used to perform its operations on a data type different from the data type executed by the element operation unit 560. For example, the multiplier 540 or the accumulator 550 performs its operations on integer type data (e.g., Int 8, Int 16, etc.), while the element operation unit 560 performs its operations on floating point type data (e.g., FP24, etc.). Therefore, according to some embodiments of the present disclosure, the processing array 530 further includes an inverse quantizer 570 and a quantizer 580, and the element operation unit 560 is disposed between the two. In some embodiments, the batch normalization operation is incorporated into the inverse quantizer 570 because both the inverse quantizer 570 and the batch normalization operation can be performed by multiplication and addition operations of constants, and the constants can be provided by the constant buffer 5030 (e.g., Figure 2B the constant buffer 2030). In some embodiments, the compiler combines the batch normalization operation and the inverse quantization operation into one operation. As Figure 5A shown, the constant buffer 5030 can provide constants to the inverse quantizer 570 for inverse quantization or batch normalization.
[0087] The sparse engine 590 may provide the decompressed sparse matrix 593 to the processing array 530, and the processing array 530 may perform calculations (such as addition, multiplication, multiply-accumulate, convolution, etc.) on the decompressed sparse matrix. In some embodiments, the processing array 530 reads input features from the first buffer 510 and uses them for calculations.
[0088] Figure 6 is a flowchart of an exemplary method 600 for representing a sparse matrix in a neural network provided according to some embodiments of the present disclosure. The method 600 may be implemented by Figure 2A or Figure 2C the heterogeneous acceleration processing unit architecture 200 of the neural network. In addition, the method 600 may also be implemented as a computer program product, which is embodied in a computer-readable medium and includes computer-executable instructions executed by a computer, such as program code. In some embodiments, the host unit (such as Figure 2A or the host unit 220 with 2C) compiles the software code to generate instructions provided to one or more HAPUs to execute the method 600.
[0089] As Figure 6 shown, in step 601, the sparse matrix is divided into multiple blocks. The sparse matrix can be of any size or dimension. In some embodiments, the sparse matrix is a sparse weight matrix or a sparse activation matrix. The blocks can be of any size or shape, including but not limited to square, rectangle, cube, cuboid, vector, column, row, filter, channel, etc. In some embodiments, the sparse matrix is divided into blocks according to hyperparameters (such as compression ratio, distribution of non-zero elements, sparsification method, storage requirements, etc.). The sparse matrix may include many non-zero elements randomly or regularly distributed in the matrix. The blocks of the sparse matrix may include one or more non-zero elements or may not include any non-zero elements. Referring to Figure 3 , for example, the sparse matrix 301 is a 6×9 2D matrix, including many non-zero elements, such as NZ0, NZ1,..., NZ6.. The sparse matrix 301 is divided into six 3×3 blocks, namely block 0, block 1,..., block 5.
[0090] In step 603, a first-level bitmap is determined. The first-level bitmap indicates whether each of the multiple blocks includes non-zero elements. The first-level bitmap may include multiple elements, such as the same number of elements as the number of blocks. Each element corresponds to a block and indicates whether the block contains non-zero elements. Referring to Figure 3, for example, the first - level bitmap 3031 includes six elements B0, B1, …, B5, and each element indicates whether the corresponding block (block 0, block 1, …, or block 5) includes non - zero elements. B0, B1, …, and B5 correspond to block 0, block 1, …, and block 5 respectively. B0 is set to 1, indicating that block 0 has at least one non - zero element, while B1 is set to 0, indicating that block 1 has no non - zero elements, and so on. Although the first - level bitmap 3031 is shown as a vector, it can be of any other size or form.
[0091] In step 605, the second - level bitmap is determined. The second - level bitmap indicates, for a block that contains non - zero elements, whether each element in that block is a non - zero element. The second - level bitmap can include multiple parts. Each part corresponds to a block that has non - zero elements. In some embodiments, each part of the second - level bitmap and its corresponding block have the same size or form. Each element of each part indicates whether the corresponding element of the corresponding block is a non - zero element. For a block that has no non - zero elements or a zero element in the first - level bitmap, the second - level bitmap has no corresponding part. One or more parts of the second - level bitmap can be arranged in a structure corresponding to the structure of the first - level bitmap. The elements of each part are arranged in a structure corresponding to the structure of the corresponding block.
[0092] Reference Figure 3 , for example, part 3032a of the second - level bitmap 3032 corresponds to B3 of the first - level bitmap 3031 or block 3 of the sparse matrix 301. Block 3 of the sparse matrix 301 includes two non - zero elements: NZ3 at position (0, 0) and NZ4 at position (2, 0). Part 3032a of the second - level bitmap 3032 includes two elements: E(0, 0) at position (0, 0) and E(2, 0) at position (2, 0), which are set to 1 to indicate the presence of non - zero elements at these positions. Although shown as a 3×N matrix, the second - level bitmap 3032 can be of any other size or form.
[0093] In step 607, an element array including the non - zero elements in the sparse matrix is formed. The non - zero elements can be arranged in a structure corresponding to the structure of the second - level bitmap. Reference Figure 3 , for example, NZ3 and NZ4 are arranged corresponding to part 3032a of the second - level bitmap 3032, and part 3032a of the second - level bitmap 3032 corresponds to B3 of the first - level bitmap 3031, so that NZ3 and NZ4 correspond to block 3 of the sparse matrix 301. Specifically, NZ3 and NZ4 correspond to E(0, 0) and E(2, 0) of part 3032a of the second - level bitmap 3032 respectively.
[0094] In some embodiments, method 600 further includes determining the size or form of each of the plurality of blocks according to hyperparameters such as compression ratio, distribution of non-zero elements, sparsification method, storage requirements, and the like.
[0095] In some embodiments, method 600 further includes pruning a matrix associated with a neural network to obtain a sparse matrix. For example, by using Figure 1A sparsification 110 of Figure 1B sparsification 120 of, etc. to prune the matrix.
[0096] In some embodiments, method 600 further includes dividing a portion of a second-level bitmap into a plurality of sub-blocks, determining a first sub-level bitmap including a plurality of elements to indicate whether each sub-block of the portion of the second-level bitmap includes non-zero elements, and determining a second sub-level bitmap including one or more portions to indicate whether each element in the sub-blocks having non-zero elements is non-zero. In some embodiments, for sub-blocks having no non-zero elements or for zero elements in the first sub-level bitmap, there is no corresponding portion in the second sub-level bitmap. The structure of the first sub-level bitmap corresponds to the structure of the second-level bitmap, and the structure of the second sub-level bitmap corresponds to the structure of the first sub-level bitmap. Referring to Figure 3 , for example, a portion 3032a of a second-level bitmap 3032 is divided into three columns. A first sub-level bitmap 3032a-1 includes three elements C0, C1, and C2 to indicate whether each column of the portion 3032a includes non-zero elements. C0, C1, and C2 respectively correspond to the left, middle, and right columns of the portion 3032a. Since the left column of the portion 3032a includes two non-zero elements E(0, 0) and E(2, 0), while the middle and right columns of the portion 3032a include only zero elements, C0 is set to 1, while C1 and C2 are both set to 0. The second sub-level bitmap 3032a-2 has a portion corresponding to C0 of the first sub-level bitmap 3032a-1 and the left column of the second-level bitmap 3032. This portion of the second sub-level bitmap 3032a-2 includes three elements (1, 0, 1), corresponding to three elements in the left column of the portion 3032a of the second-level bitmap 3032, and indicates whether they are non-zero elements.
[0097] It should be understood that in some embodiments, method 600 includes determining bitmaps at deeper sub-levels. For example, the second sub-level bitmap can be divided into a plurality of sub-blocks and represented by other deeper sub-level bitmaps.
[0098] In some embodiments, method 600 includes combining at least two of the first-level bitmap, the second-level bitmap, and the element array with each other.
[0099] Figure 7is a flowchart of an exemplary method 700 for decompressing a representation of a sparse matrix in a neural network provided according to some embodiments of the present disclosure. Method 700 can be implemented by Figure 2A or Figure 2C the architecture 200 of a neural network heterogeneous acceleration processing unit or Figures 5A - 5B the sparse engine 590 therein. In addition, method 700 can also be implemented as a computer program product embodied in a computer-readable medium, including computer-executable instructions executed by a computer, such as program code.
[0100] As Figure 7 shown, at step 701, a representation of a sparse matrix in a neural network is read. The representation of the sparse matrix includes a first-level bitmap, a second-level bitmap, and an element array. In some embodiments, the representation of the sparse matrix further includes one or more sub-level bitmaps. For example, Figure 5A the sparse engine 590 Figure 3 reads a representation of a sparse matrix in a neural network (e.g., Figure 4 representation 303, Figure 3 the first-level bitmap 3031 of Figure 4 representation 403, etc.), the representation of the sparse matrix includes a first-level bitmap 521 (e.g., Figure 3 the first-level bitmap 3032 of Figure 4 representation 403, etc.), a second-level bitmap 522 (e.g., Figure 3 the element array 3033 of Figure 4 representation 403, etc.) and an element array 523 (e.g., Figure 3 the first sub-level bitmap 3032a-1, the second sub-level bitmap 3032a-2, more sub-level bitmaps 3032a-n of Figure 4 representation 403, the first sub-level bitmap 4032-1, the second sub-level bitmap 4032-2, etc.). The representation of the sparse matrix may further include one or more sub-level bitmaps (e.g.,
[0101] At step 703, the first-level bitmap is decompressed to determine whether each block of the sparse matrix includes non-zero elements. The sparse matrix includes a plurality of blocks. Decompressing the first-level bitmap includes checking each element of the first-level bitmap. In some embodiments, an element of the first-level bitmap is set to 1 to indicate that its corresponding block includes non-zero elements. For example, Figure 5A the sparse engine 590 (e.g., the first-level decompressor 591) of Figure 3The first - level bitmap 3031, for example, checks each element B0, B1, …, or B5 of the first - level bitmap 3031 to determine whether each block of the sparse matrix 301 includes non - zero elements. Elements B0, B2 - B4 are set to 1, indicating that blocks 0 and 2 to 4 of the sparse matrix 301 have non - zero elements, while B1 and B5 are set to 0, indicating that blocks 1 and 5 of the sparse matrix 301 do not have any non - zero elements.
[0102] In step 705, in response to a block including non - zero elements, the second - level bitmap is decompressed using the element array to obtain the blocks of the sparse matrix that include non - zero elements. In some embodiments, the second - level bitmap can be used as a bit - mask and the sparse matrix is reconstructed using the element array. One element of the second - level bitmap is set to 1 to indicate that its corresponding element is a non - zero element included in the element array. In some embodiments, the obtained blocks are combined to obtain the sparse matrix or a variant thereof. For example, Figure 5B the sparse engine 590 (e.g., the second - level decompressor 592) of Figure 3 decompresses the second - level bitmap 3032 of Figure 3 using the element array 3033 to obtain blocks 0 and 2 to 4 of the sparse matrix 301. Referring to Figure 5B for the portion 3032a corresponding to B3 or block 3, elements E(0, 0) and E(2, 0) are set to 1 and respectively correspond to the non - zero elements NZ3 and NZ4 of the element array 3033. Figure 5B the sparse engine 590 (e.g., the second - level decompressor 592) of
[0103] can reconstruct block 3 by placing NZ3 and NZ4 at positions (0, 0) and (2, 0) respectively and filling all other positions with zeros. Similarly, blocks 0, 2, and 4 can also be reconstructed.
[0104] the sparse engine 590 (e.g., the first - level decompressor 591) of
[0103] combines the obtained blocks (blocks 0 and 2 - 4) to obtain the sparse matrix 301.
[0104] In step 707, a sparse matrix is provided. In some embodiments, the provided sparse matrix is the same as the original sparse matrix or is a variant of the original sparse matrix.
[0104] In some embodiments, method 700 further includes decompressing a first sub - level bitmap to determine whether each sub - block of each portion of the second - level bitmap includes non - zero elements, and in response to a sub - block including non - zero elements, decompressing a second sub - level bitmap to obtain the sub - blocks of the corresponding portion of the second - level bitmap that include non - zero elements. For example, in combination with Figure 3 and Figures 5A - 5B, the representation 303 includes a first sub - level bitmap 3032a - 1 and a second sub - level bitmap 3032a - 2. The sparse engine 590 (e.g., the first - level decompressor 591) decompresses the first sub - level bitmap 3032a - 1. For example, it checks each element C0, C1, and C2 of the first sub - level bitmap 3032a - 1 to determine whether the sub - blocks (e.g., columns) of the portion 3032a include non - zero elements. The element C0 is set to 1, indicating that the left column of the portion 3032a has one or more non - zero elements, while C1 and C2 are set to 0, indicating that the middle and right columns of the portion 3032a have no non - zero elements. For the element C0 of the first sub - level bitmap 3032a - 1 or the left column of the portion 3032a, the sparse engine 590 (e.g., the second - level decompressor 592) can reconstruct the left column of the portion 3032a using the second sub - level bitmap 3032a - 2 and reconstruct the middle and right columns of the portion 3032a using multiple zeros. Similarly, the sparse engine 590 can reconstruct other parts of the second - level bitmap 3032 and decompress the second - level bitmap 3032 to obtain the sparse matrix 301.
[0105] Embodiments of the present disclosure can bring many technical advantages. For example, the hierarchical representation of the sparse matrix in some embodiments of the present disclosure can reduce storage space. In addition, the hierarchical representation is hardware - friendly, thus enabling a simple decompression process and higher decompression throughput.
[0106] In some embodiments of the present disclosure, hyperparameters are used to control the granularity of sparsity in a neural network. In some embodiments, the hierarchical representation of the sparse matrix characterizes structured sparsity at one level (e.g., the first - level BM) and unstructured sparsity at another level (e.g., the second - level BM). This can improve the storage and computational efficiency of the sparse matrix.
[0107] It should be understood that the embodiments disclosed herein can be used in various application environments, such as artificial intelligence (AI) training and inference, database and big data analysis acceleration, video compression and decompression, etc. Applications related to artificial intelligence may involve machine learning (ML) or deep learning (DL) based on neural networks. Therefore, the embodiments of the present disclosure can be used in various neural network architectures, such as deep neural network (DNN), convolutional neural network (CNN), recurrent neural network (RNN), etc. For example, some embodiments of the present disclosure can be used for artificial intelligence inference of DNN.
[0108] Embodiments of the present disclosure can be applied to many products. For example, some embodiments of the present disclosure can be applied to Ali NPU (such as HanGuang NPU), Ali-Cloud, Ali PIM-AI, Ali-DPU, Ali-AI platform, Ali-Data Center AI Inference Chip, IoT Edge AI Chip, GPU, TPU, etc.
[0109] The embodiments are further described below using the following claims.
[0110] 1. A method for representing a sparse matrix in a neural network, comprising:
[0111] Dividing the sparse matrix into a plurality of blocks;
[0112] Determining a first-level bitmap to indicate whether each of the plurality of blocks includes non-zero elements;
[0113] Determining a second-level bitmap to indicate whether each element in the blocks containing non-zero elements is a non-zero element; and
[0114] Forming an element array including the non-zero elements in the sparse matrix.
[0115] 2. The method according to claim 1, wherein the sparse matrix is a sparse weight matrix or a sparse activation matrix, or the block is a square, rectangle, cube, cuboid, vector, column, row, filter, or channel.
[0116] 3. The method according to claim 1 or 2, further comprising:
[0117] Trimming a matrix related to the neural network to obtain the sparse matrix.
[0118] 4. The method according to any one of claims 1-3, further comprising:
[0119] Determining the size or form of each of the plurality of blocks according to at least one of a compression ratio, non-zero element distribution, sparsification method, and storage requirements.
[0120] 5. The method according to any one of claims 1 to 4, wherein determining the first-level bitmap includes:
[0121] Determining a plurality of elements, each element corresponding to a block and indicating whether the block includes non-zero elements.
[0122] 6. The method according to any one of claims 1 to 5, wherein determining the second-level bitmap comprises:
[0123] For each block containing non-zero elements, determining each part having a plurality of elements, wherein each element of each part indicates whether the corresponding element of the corresponding block is a non-zero element; and
[0124] Combining one or more of the determined parts to form the second-level bitmap.
[0125] 7. The method according to any one of claims 1 to 6, further comprising:
[0126] Dividing each part of the second-level bitmap into a plurality of sub-blocks;
[0127] Determining a first sub-level bitmap including a plurality of elements to indicate whether each sub-block of the corresponding part of the second-level bitmap includes non-zero elements; and
[0128] Determining a second sub-level bitmap including one or more parts to indicate, for a sub-block having non-zero elements, whether each element of the sub-block is a non-zero element.
[0129] 8. The method according to any one of claims 1 to 7, further comprising:
[0130] Combining at least two of the first-level bitmap, the second-level bitmap, and the element array with each other.
[0131] 9. A method for executing a neural network, comprising:
[0132] Reading a representation of a sparse matrix, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array;
[0133] Decompressing the first-level bitmap to determine whether each block of the sparse matrix includes non-zero elements;
[0134] In response to a block including non-zero elements, using the element array to decompress the second-level bitmap to obtain the blocks of the sparse matrix including non-zero elements; and
[0135] Using the sparse matrix to execute a neural network.
[0136] 10. The method according to claim 9, wherein decompressing the first-level bitmap comprises:
[0137] Checking each element of the first-level bitmap to determine whether the corresponding block of the sparse matrix contains non-zero elements.
[0138] 11. The method according to claim 9 or 10, wherein decompressing the second-level bitmap comprises:
[0139] Using the second-level bitmap as a bitmask, reconstruct the sparse matrix using the element array.
[0140] 12. The method according to any one of claims 9 to 11, wherein the representation of the sparse matrix includes a first-level bitmap and a second-level bitmap, and the method further includes:
[0141] Decompress the first-level bitmap to determine whether each sub-block of each part of the second-level bitmap includes non-zero elements; and
[0142] In response to a sub-block containing non-zero elements, decompress the second-level bitmap to obtain the sub-blocks including non-zero elements of the corresponding part of the second-level bitmap.
[0143] 13. An operation unit, comprising:
[0144] A buffer for storing the representation of a sparse matrix in a neural network;
[0145] A sparse engine communicatively coupled to the buffer;
[0146] A processing array communicatively coupled to the sparse engine, the processing array including circuitry for performing the neural network using the sparse matrix;
[0147] Wherein the sparse engine includes circuitry for performing the following operations:
[0148] Read the representation of the sparse matrix from the buffer, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array;
[0149] Decompress the first-level bitmap to determine whether each block of the sparse matrix includes non-zero elements;
[0150] In response to a block including non-zero elements, use the element array to decompress the second-level bitmap to obtain the blocks of the sparse matrix including non-zero elements.
[0151] 14. The operation unit according to claim 13, wherein the sparse engine includes circuitry for performing the following operations:
[0152] Check each element of the first-level bitmap to determine whether the corresponding block of the sparse matrix contains non-zero elements.
[0153] 15. The operation unit according to claim 13 or 14, wherein the sparse engine includes circuitry for performing the following operations:
[0154] Using the second-level bitmap as a bitmask, reconstruct the sparse matrix using the element array.
[0155] 16. The operation unit according to any one of claims 13-15, wherein the representation of the sparse matrix includes a first-level bitmap and a second-level bitmap, and the sparse engine includes circuitry for performing the following operations:
[0156] Decompressing the first-level bitmap to determine whether each sub-block of each part of the second-level bitmap includes a non-zero element; and
[0157] In response to a sub-block including a non-zero element, decompressing the second-level bitmap to obtain the sub-blocks including non-zero elements of the corresponding part of the second-level bitmap.
[0158] 17. The operation unit according to any one of claims 13-16, wherein the sparse engine includes:
[0159] A first-level decompressor including circuitry for decompressing the first-level bitmap; and
[0160] A second-level decompressor communicatively coupled to the first-level decompressor, including circuitry for decompressing the second-level bitmap.
[0161] 18. The operation unit according to any one of claims 13-17, wherein the processing array includes a plurality of layers, and at least one layer of the plurality of layers includes circuitry for performing the neural network using the sparse matrix.
[0162] 19. A processing core, comprising:
[0163] A local memory; and
[0164] An operation unit communicatively coupled to the local memory, the operation unit including:
[0165] A buffer for storing the representation of the sparse matrix in the neural network;
[0166] A sparse engine communicatively coupled to the buffer;
[0167] A processing array communicatively coupled to the sparse engine, the processing array including circuitry for performing the neural network using the sparse matrix;
[0168] Wherein the sparse engine includes circuitry for performing the following operations:
[0169] Reading the representation of the sparse matrix from the buffer, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array;
[0170] Decompressing the first-level bitmap to determine whether each block of the sparse matrix includes a non-zero element; and
[0171] In response to a block including non-zero elements, decompress the second-level bitmap using the element array to obtain the block of the sparse matrix including non-zero elements.
[0172] 20. The processing core according to claim 19, wherein the sparse engine includes circuitry for performing the following operations:
[0173] Examine each element of the first-level bitmap to determine whether the corresponding block of the sparse matrix contains non-zero elements.
[0174] 21. The processing core according to claim 19 or 20, wherein the sparse engine includes circuitry for performing the following operations:
[0175] Using the second-level bitmap as a bitmask, reconstruct the sparse matrix using the element array.
[0176] 22. The processing core according to any one of claims 19 to 21, wherein the representation of the sparse matrix includes a first sub-level bitmap and a second sub-level bitmap, and the sparse engine includes circuitry for performing the following operations:
[0177] Decompress the first sub-level bitmap to determine whether each sub-block of each part of the second-level bitmap includes non-zero elements; and
[0178] In response to a sub-block containing non-zero elements, decompress the second sub-level bitmap to obtain the sub-block of the corresponding part of the second-level bitmap that contains non-zero elements.
[0179] 23. The processing core according to any one of claims 19 to 22, wherein the sparse engine includes:
[0180] A first-level decompressor including circuitry for decompressing the first-level bitmap; and
[0181] A second-level decompressor communicatively coupled to the first-level decompressor, including circuitry for decompressing the second-level bitmap.
[0182] 24. The processing core according to any one of claims 19 to 23, wherein the processing array includes multiple layers, and at least one layer of the multiple layers includes circuitry for performing the neural network using the sparse matrix.
[0183] 25. A non-transitory computer-readable storage medium storing a set of instructions that, when executed by one or more processing devices, cause an operating unit to perform a method, the method including:
[0184] Partition a sparse matrix into multiple blocks;
[0185] Determine a first-level bitmap to indicate whether each of the plurality of blocks includes non-zero elements;
[0186] Determine a second-level bitmap to indicate whether each element in the blocks containing non-zero elements is a non-zero element; and
[0187] Form an element array including the non-zero elements in the sparse matrix.
[0188] 26. The computer-readable storage medium according to claim 25, wherein the sparse matrix is a sparse weight matrix or a sparse activation matrix, or the block is a square, a rectangle, a cube, a cuboid, a vector, a column, a row, a filter, or a channel.
[0189] 27. The computer-readable storage medium according to claim 25 or 26, wherein the set of instructions can be executed by one or more processing devices to cause the operation unit to further perform:
[0190] Prune the matrix related to the neural network to obtain the sparse matrix.
[0191] 28. The computer-readable storage medium according to any one of claims 25 to 27, wherein the set of instructions can be executed by one or more processing devices to cause the operation unit to further perform:
[0192] Determine the size or form of each of the plurality of blocks according to at least one of the compression ratio, non-zero element distribution, sparsification method, and storage requirements.
[0193] 29. The computer-readable storage medium according to any one of claims 25 to 28, wherein determining the first-level bitmap includes:
[0194] Determine a plurality of elements, each element corresponding to a block and indicating whether the block includes non-zero elements.
[0195] 30. The computer-readable storage medium according to any one of claims 25 to 29, wherein determining the second-level bitmap includes:
[0196] For the blocks containing non-zero elements, determine a portion having a plurality of elements, the plurality of elements indicating whether the corresponding elements of the block are non-zero elements; and
[0197] Combine one or more of the determined portions to form the second-level bitmap.
[0198] 31. The computer-readable storage medium according to any one of claims 25 to 30, wherein the set of instructions can be executed by one or more processing devices to cause the operation unit to further perform:
[0199] Divide each part of the second-level bitmap into a plurality of sub-blocks;
[0200] Determine a first sub-level bitmap including a plurality of elements to indicate whether each sub-block of each part of the second-level bitmap includes a non-zero element; and
[0201] Determine a second sub-level bitmap including one or more parts to indicate whether each element of a sub-block is a non-zero element for a sub-block having non-zero elements.
[0202] 32. The computer-readable storage medium according to any one of claims 25 to 31, wherein the set of instructions is executable by one or more processing devices to cause the operation unit to further execute:
[0203] Combine at least two of the first-level bitmap, the second-level bitmap, and the element array with each other.
[0204] 33. A non-transitory computer-readable storage medium storing a set of instructions to cause an operation unit to execute a method, the method comprising:
[0205] Read a representation of a sparse matrix, the representation including a first-level bitmap, a second-level bitmap, and an element array;
[0206] Decompress the first-level bitmap to determine whether each block of the sparse matrix includes a non-zero element;
[0207] In response to a block including a non-zero element, use the element array to decompress the second-level bitmap to obtain the blocks of the sparse matrix including non-zero elements; and
[0208] Execute the neural network using the sparse matrix.
[0209] 34. The computer-readable storage medium according to claim 33, wherein decompressing the first-level bitmap includes:
[0210] Check each element of the first-level bitmap to determine whether the corresponding block of the sparse matrix contains a non-zero element.
[0211] 35. The computer-readable storage medium according to claim 33 or 34, wherein decompressing the second-level bitmap includes:
[0212] Using the second-level bitmap as a bit mask, reconstruct the sparse matrix using the element array.
[0213] 36. The computer-readable storage medium according to any one of claims 33 to 35, wherein the representation includes a first sub-level bitmap and a second sub-level bitmap, and the set of instructions is executable by one or more processing devices to cause the operating unit to further perform:
[0214] Decompress the first sub-level bitmap to determine whether each sub-block of each part of the second-level bitmap includes a non-zero element; and
[0215] In response to the sub-block including a non-zero element, decompress the second sub-level bitmap to obtain the sub-block including non-zero elements of the corresponding part of the second-level bitmap.
[0216] The various exemplary embodiments have been described above in the context of steps or processes of a method. On the one hand, these steps or processes of the method can be implemented by a computer program product included in a computer-readable medium, which computer program product includes computer-executable instructions such as program code executed by a computer in a network environment. The computer-readable medium can include removable storage devices and non-removable storage devices, including but not limited to read-only memory (ROM), random access memory (RAM), optical discs, digital versatile discs (DVDs), etc. Generally, program modules can include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The computer-executable instructions, related data structures, and program modules represent examples of program code for performing the steps of the methods disclosed herein. Such a specific sequence of executable instructions or related data structures represents an example of the corresponding actions for implementing the functions described in such steps or processes.
[0217] The present disclosure provides the above description for illustrative purposes. It is not exhaustive and is not limited to the precise forms or embodiments disclosed. Modifications and adaptations of the embodiments will be apparent by considering the specification and practice of the embodiments of the present disclosure. For example, the described implementations include hardware, but systems and methods consistent with the present disclosure can be implemented with hardware and software. Additionally, although some components have been described as being coupled to each other, these components can be integrated with each other or deployed in any suitable manner.
[0218] Moreover, although illustrative embodiments have been described herein, the scope of the present disclosure includes any and all embodiments having equivalent elements, modifications, omissions, combinations (e.g., combining aspects of different embodiments), adaptations, or alterations based on the present disclosure. The elements in the claims will be broadly interpreted based on the language used in the claims and are not limited to the examples described in this specification or during the application process, which examples will be interpreted as non-exclusive. Additionally, the steps of the disclosed methods can be modified in any way, including reordering steps and / or inserting or deleting steps.
[0219] The features and advantages of the present disclosure will be apparent from the detailed description. Accordingly, the appended claims are intended to cover all systems and methods that fall within the spirit and scope of the present disclosure. As used herein, the indefinite articles "a" and "an" mean "one or more." Additionally, since many modifications and variations will readily occur to those of skill in the art upon a study of this disclosure, it is not desired to limit the disclosure to the exact construction and operation shown and described. Accordingly, all suitable modifications and equivalents fall within the scope of the present disclosure.
[0220] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations, unless it is not feasible. For example, if it is stated that a component can include A or B, then the component can include A, B, or A and B, unless specifically stated otherwise or it is not feasible. As a second example, if it is stated that a component can include A, B, or C, then the component can include A, B, C, A and B, A and C, B and C, or A and B and C, unless specifically stated otherwise or it is not feasible.
[0221] Other embodiments will be apparent to those of skill in the art from consideration of the specification and practice of the embodiments disclosed herein. The specification and examples are to be considered as exemplary only, with the true scope and spirit of the disclosed embodiments being indicated by the claims.
Claims
1. An operation unit, comprising: A buffer for storing a representation of a sparse matrix in a neural network; A sparse engine communicatively coupled to the buffer; A processing array communicatively coupled to the sparse engine, the processing array including circuitry for performing the neural network using the sparse matrix; Wherein the sparse engine includes circuitry for performing the following operations: Reading the representation of the sparse matrix from the buffer, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array; Decompressing the first-level bitmap to determine whether each block of the sparse matrix includes non-zero elements, each block of the sparse matrix including a plurality of elements; and In response to the block including non-zero elements, decompressing the second-level bitmap using the element array to obtain the block of the sparse matrix including non-zero elements.
2. The operating unit according to claim 1, wherein, The sparse engine includes circuitry for performing the following operations: Checking each element of the first-level bitmap to determine whether the corresponding block of the sparse matrix contains non-zero elements.
3. The operating unit according to claim 1, wherein The sparse engine includes circuitry for performing the following operations: Using the second-level bitmap as a bitmask and reconstructing the sparse matrix using the element array.
4. The operating unit according to claim 1, wherein, The representation of the sparse matrix includes a first sub-level bitmap and a second sub-level bitmap, and the sparse engine includes circuitry for performing the following operations: Decompressing the first sub-level bitmap to determine whether each sub-block of each part of the second-level bitmap includes non-zero elements; And In response to the sub-block including non-zero elements, decompressing the second sub-level bitmap to obtain the sub-block of the corresponding part of the second-level bitmap including non-zero elements.
5. The operating unit according to claim 1, wherein The sparse engine includes: A first-level decompressor including circuitry for decompressing the first-level bitmap; and A second-level decompressor communicatively coupled to the first-level decompressor, including circuitry for decompressing the second-level bitmap.
6. The operating unit according to claim 1, wherein, The processing array includes a plurality of layers, and at least one of the plurality of layers includes circuitry for performing the neural network using the sparse matrix.
7. A processing core, comprising: Local memory; And An operation unit communicatively coupled to the local memory, the operation unit including: A buffer for storing a representation of a sparse matrix in a neural network; A sparse engine communicatively coupled to the buffer; A processing array communicatively coupled to the sparse engine, the processing array including circuitry for performing the neural network using the sparse matrix; Wherein the sparse engine includes circuitry for performing the following operations: Reading the representation of the sparse matrix from the buffer, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array; Decompressing the first-level bitmap to determine whether each block of the sparse matrix includes non-zero elements, each block of the sparse matrix including a plurality of elements; and In response to the block including non-zero elements, decompressing the second-level bitmap using the element array to obtain the block of the sparse matrix including non-zero elements.
8. The processing core according to claim 7, wherein, The sparse engine includes circuitry for performing the following operations: Checking each element of the first-level bitmap to determine whether the corresponding block of the sparse matrix contains non-zero elements.
9. The processing core according to claim 7, wherein, The sparse engine includes circuitry for performing the following operations: Using the second-level bitmap as a bitmask, reconstruct the sparse matrix using the element array.
10. The processing core according to claim 7, wherein, The representation of the sparse matrix includes a first-level bitmap and a second-level bitmap, and the sparse engine includes circuitry that performs the following operations: Decompress the first-level bitmap to determine whether each sub-block of each part of the second-level bitmap includes a non-zero element; and In response to the sub-block including a non-zero element, decompress the second-level bitmap to obtain the sub-blocks of the corresponding part of the second-level bitmap that include non-zero elements.
11. The processing core according to claim 7, wherein, The sparse engine includes: A first-level decompressor, including circuitry for decompressing the first-level bitmap; and A second-level decompressor, communicatively coupled to the first-level decompressor, including circuitry for decompressing the second-level bitmap.
12. The processing core according to claim 7, wherein, The processing array includes multiple layers, and at least one layer of the multiple layers includes circuitry for performing the neural network using the sparse matrix.
13. A method for performing a neural network, including: Read a representation of a sparse matrix, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array; Decompress the first-level bitmap to determine whether each block of the sparse matrix includes a non-zero element, each block of the sparse matrix including multiple elements; In response to the block including a non-zero element, decompress the second-level bitmap using the element array to obtain the blocks of the sparse matrix that include non-zero elements; and Perform the neural network using the sparse matrix.
14. The method according to claim 13, wherein, The decompressing the first-level bitmap includes: Checking each element of the first-level bitmap to determine whether the corresponding block of the sparse matrix contains a non-zero element.
15. The method according to claim 13, wherein, The decompressing the second-level bitmap includes: Using the second-level bitmap as a bitmask, reconstruct the sparse matrix using the element array.
16. The method according to claim 13, wherein, The representation of the sparse matrix includes a first-level bitmap and a second-level bitmap, and the method further includes: Decompress the first-level bitmap to determine whether each sub-block of each part of the second-level bitmap includes a non-zero element; and In response to the sub-block including a non-zero element, decompress the second-level bitmap to obtain the sub-blocks of the corresponding part of the second-level bitmap that include non-zero elements.
17. A non-transitory computer-readable storage medium storing a set of instructions that, when executed by one or more processing devices, cause an operating unit to perform a method that includes: Read a representation of a sparse matrix in a neural network, the representation of the sparse matrix including a first-level bitmap, a second-level bitmap, and an element array; Decompress the first-level bitmap to determine whether each block of the sparse matrix includes a non-zero element, each block of the sparse matrix including multiple elements; In response to the block including a non-zero element, decompress the second-level bitmap using the element array to obtain the blocks of the sparse matrix that include non-zero elements; and Perform the neural network using the sparse matrix.
18. The computer-readable storage medium according to claim 17, wherein, The set of instructions, when executed by one or more processing devices, cause the operating unit to perform: Checking each element of the first-level bitmap to determine whether the corresponding block of the sparse matrix contains a non-zero element.
19. The computer-readable storage medium according to claim 17, wherein, When the set of instructions is executed by one or more processing devices, cause the operation unit to perform: Using the second-level bitmap as a bitmask, reconstruct the sparse matrix using the element array.
20. The computer-readable storage medium according to claim 17, wherein, The representation of the sparse matrix includes a first sub-level bitmap and a second sub-level bitmap. When the set of instructions is executed by one or more processing devices, cause the operation unit to perform: Decompress the first sub-level bitmap to determine whether each sub-block of each part of the second-level bitmap includes a non-zero element; and In response to a sub-block including a non-zero element, decompress the second sub-level bitmap to obtain the sub-blocks including non-zero elements of the corresponding part of the second-level bitmap.
Citation Information
Patent Citations
Compression system and method for accelerating sparse matrix computations
US20070198621A1
Exploiting input data sparsity in neural network compute units
US20200012608A1